A Defense Method Against Synonym Substitution Attacks

Generating high-quality synonym adversarial samples through Counter-fitting and TextRank algorithms solves the problem of low text quality in the existing methods, and achieves efficient training and robustness improvement of the target model.

CN115719061BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211422598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-05-27
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

Existing methods are difficult to generate effective synonym adversarial samples that are closely related to the context, resulting in low quality of adversarial text and low training value.

Method used

The word embedding vector in the news text is represented as a sentence implicit vector through Counter-fitting, keywords are extracted using the TextRank algorithm, candidate synonyms are generated, and high-quality adversarial text is generated through grammatical checking and cosine similarity calculation.

Benefits of technology

The generated adversarial text has a high semantic similarity and high quality to the original text, which can effectively train the target model and make it highly robust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719061B_ABST
    Figure CN115719061B_ABST
Patent Text Reader

Abstract

The present invention relates to a defense method against synonym replacement attacks, including embedding news texts into vector representations to obtain a sequence of sentence vectors of the news texts; extracting the keywords of each sentence in the news texts to generate a candidate keyword set; calculating the cosine similarity between the word vectors of each candidate keyword and the glove-trained word vectors to obtain the synonyms of the candidate keywords; performing a grammar check on the synonyms of the candidate keywords; using the synonyms with the same part of speech as the candidate keywords as the final candidate words, and replacing the corresponding keywords in the news texts with the final candidate words to obtain a second news text; embedding the second news text into a vector representation to obtain a first news text vector; inputting the sequence of sentence vectors of the news texts into an LSTM to calculate a second news text vector; mapping the first news text vectors with a cosine similarity exceeding a set threshold with the second news text vector into texts and using them as adversarial samples of the news texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to a defense method against synonym replacement attacks. Background Art

[0002] With the advent of the big data era and the rapid development of artificial intelligence technology, deep learning has been widely applied in the fields of computer image processing, signal analysis, natural language processing, etc. In the field of natural language processing, text processing technology has made breakthrough progress in the fields of reading comprehension, sentiment analysis, text classification, speech recognition, etc.

[0003] At the same time, deep learning models are also facing various security problems. As a kind of publicity text with great public influence, the security problems of news texts are particularly important. Among them, adversarial attack is a hot issue. In the field of natural language processing, adversarial attack refers to adding tiny perturbations to the original text to obtain adversarial texts, and by inputting the adversarial texts, the deep learning model outputs a wrong semantic label, while users cannot perceive the semantic changes of the adversarial texts compared with the original texts. To solve this security problem, it is necessary to use adversarial texts to train the target model to defend against adversarial attacks, so that the target model can effectively resist external adversarial attacks.

[0004] The fundamental purpose of adversarial training is to use the generated adversarial samples to enhance the robustness of the deep learning model, so as to defend against the malicious interference of text attacks on the model.

[0005] During the process of adversarial training, synonym replacement attacks generally add some adversarial texts with tiny perturbations to the training samples, and then make the deep neural network slowly adapt to this kind of interference, and then be robust to the adversarial attack texts. The most important feature of adversarial training is the generation of adversarial samples. Often, whether adversarial training can really improve the robustness of the model against a certain type of attack depends on whether there is relatively perfect text data under the corresponding attack method. Therefore, the generation of adversarial samples is crucial for adversarial training.

[0006] However, in the existing methods, a large number of word vectors are difficult to be mapped back to the text format. Especially for synonyms, it is difficult to generate effective synonym adversarial samples that are closely related to the context, or the generated adversarial texts have a large semantic difference from the original texts, and the quality of the adversarial texts is low, resulting in a low training value of the adversarial texts. Summary of the Invention

[0007] To solve the problems existing in the background art, the present invention provides a defense method against synonym replacement attacks, including:

[0008] S1: Obtain news text data with label information; the label information is the category of the news;

[0009] S2: Use Counter-fitting to obtain the hidden vectors of each sentence in the news text by embedding the word vectors in the news text, and generate a sentence vector sequence of the news text;

[0010] S3: Use the TextRank algorithm to extract the keywords of each sentence in the news text to generate a candidate keyword set;

[0011] S4: Calculate the cosine similarity between the word vector of each keyword in the candidate keyword set and the glove training word vector respectively, and take the words corresponding to the M glove training word vectors with the cosine similarity greater than the set threshold as the candidate synonyms of the keyword;

[0012] S5: Conduct a grammar check on the candidate synonyms of the keyword; take the candidate synonyms with the same part of speech as the keyword as the final candidate words of the keyword, and replace the corresponding keywords in the news text with the final candidate words to obtain the second news text; represent the second news text in vector form to obtain the first news text vector;

[0013] S6: Input the sentence vector sequence of the news text into the LSTM for calculation to obtain a second news text vector containing context information;

[0014] S7: According to the cosine similarity between the first news text vector and the second news text vector, map the first news text vectors with the cosine similarity exceeding the set threshold into text and use them as adversarial samples of the news text;

[0015] S8: Use the adversarial samples of the news text and the news text as training samples to train the target model, create a target loss function according to the label information of the news text, and update the parameters of the target model through the mechanism of backpropagation.

[0016] The present invention has at least the following beneficial effects

[0017] The present invention extracts features from the sentence-level state and multiple word-level sub-states, and represents sentence vectors containing context semantic relationships; by replacing the corresponding keywords in the original text with synonyms, and combining the news text vectors containing context semantic relationships, adversarial texts with high semantic similarity to the original text can be generated, that is, high-quality adversarial texts can be obtained. Furthermore, the target model can be effectively trained using the high-quality adversarial texts, so that the trained target model has high robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the method of the present invention;

[0019] Figure 2 is a schematic diagram of the framework process of the present invention. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described with reference to the accompanying drawings of the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention and cannot be used to limit the scope of the present invention.

[0021] Please refer to Figure 1 and Figure 2 , the present invention provides a defense method against synonym replacement attacks, including:

[0022] S1: Obtain news text data with tag information; the tag information is the category of the news;

[0023] In the present invention, AG’s News (AG), DBPediaontology, and Yahoo! Answers are obtained. Among them, AG’s News consists of world news, sports news, business news, and technology news. Each category of news includes 30,000 training samples and 1,900 test samples. DBPedia is constructed by selecting 14 non-overlapping classes from DBPedia2014, and structured information is extracted from Wikipedia. For each of the 14 ontology classes, there are 40,000 training samples and 5,000 test samples. Yahoo! Answers is a topic classification dataset containing 10 classes, and each class contains 140,000 training samples and 5,000 test samples. The news text data is obtained by providing an API interface through a social platform.

[0024] S2: Use Counter-fitting to represent the word embedding vectors in the news text to obtain the hidden vectors of each sentence in the news text and generate a sentence vector sequence of the news text;

[0025] s = {w 1 , w 2 ,... w i ,... w n}

[0026]

[0027] where s represents the news text, w i represents the i-th sentence in the news text, represents the hidden vector of the i-th sentence in the news text, and s h represents the sentence vector sequence of the news text.

[0028] S3: Use the TextRank algorithm to extract the keywords of each sentence in the news text to generate a candidate keyword set;

[0029] Preferably, the extraction of candidate word sets by generating keywords for each sentence in the news text using the TextRank algorithm includes:

[0030] Segment the current news text into complete sentences. For each sentence, perform word segmentation and part-of-speech tagging, and filter out stop words, only retaining words of specified parts of speech, such as nouns, verbs, and adjectives. Extract these retained words as candidate keywords for our news text. Using these candidate keywords as nodes, construct an edge between two nodes based on the co-occurrence relationship between keywords to build a keyword graph G=(V, E). The co-occurrence relationship between keywords includes: two keywords co-occur in a window of length k. Iteratively propagate the weights of each node until convergence, and then sort in reverse order to select the top T keywords with the highest node weights and store them in the candidate word set KW:

[0031]

[0032] Among them, WS(V i ) represents the weight of the word w i in the news text. d is the damping coefficient, generally 0.85, used to overcome the inherent defect of the following summation formula. w ji represents the similarity between word j and word i. In(V i ) represents the set of predecessor nodes of node V i . Out(V j ) represents the set of successor nodes of node V j . WS(V j ) represents the weight of word j obtained in the previous iteration. The summation represents the contribution of each word-sentence pair to this word. Further measure the importance of keywords in their current context. For each keyword in KW, delete the keyword in the sentence s 1 to which it belongs in the news text, and use the corresponding sentence s 2 after deletion as the input and input it into the RNN classification model to perform a classification task on it Among them, is the classification objective function, s is the input sentence, y is the classification result. If , then retain the candidate word deleted in s 2 as the candidate word to be replaced.

[0033] S4: Calculate the cosine similarity between the word vectors of each keyword in the candidate keyword set and the glove training word vectors respectively, and use the words corresponding to the M glove training word vectors with a cosine similarity greater than the set threshold as the candidate synonyms of this keyword;

[0034] Select the glove pre-trained word vectors of 300d, calculate the cosine values of the candidate words and all the words in the glove pre-trained word vectors, select the synonyms with cosine values greater than δ, and increase or decrease the candidate synonym set of the word w by setting the value of δ, and filter the synonyms from the semantic similarity.

[0035]

[0036] Among them, w represents the word vector of the keyword in the candidate word set, and V i represents the pre-trained word vectors in glove, δ represents the set cosine similarity threshold, and cosθ represents the cosine similarity between the word vector of the keyword in the candidate word set and the glove word vector.

[0037] S5: Conduct a grammar check on the candidate synonyms of the keyword; use the candidate synonyms with the same part of speech as the keyword as the final candidate words of the keyword, and replace the corresponding keyword in the news text with the final candidate words to obtain the second news text; embed the second news text to obtain the first news text vector;

[0038] Using the candidate synonyms of the keyword as the final candidate words of the keyword includes:

[0039] Use the spaCy grammar checker to conduct a grammar check on the candidate synonyms of the keyword. spaCy first tokenizes the candidate synonyms through the Tokenizer to generate a doc object, and the doc object contains the sequence of Tokens and the annotations of Tokens. Then the doc object will be processed through the pipeline, which includes a tagger, a lemmatizer, a parser, and an entity recognizer. After these steps of processing, the final doc object is formed, and then the candidate synonyms with the same part of speech as the keyword are selected as the final candidate words of the keyword.

[0040] S6: Input the sentence vector sequence s = {x 1 , x 2 ,..., x n} of the news text into the LSTM (Long Short-Term Memory Network) to calculate and obtain the second news text vector containing context information;

[0041] The LSTM includes: a forget gate, a cell unit, an input gate, and an output gate;

[0042] The forget gate is as follows:

[0043]

[0044]

[0045] Among them, f i t is a gate that controls the information from . represents the state of the LSTM cell at time t, x i represents the word vector, g t-1 represents the sentence vector at time t-1. is a gate that controls the information from . represents the state of the sentence cell at time t. is obtained by taking the average of the hidden vectors of all words at the previous time . represents the concatenation of the hidden vectors of the words within the current word context window. W x , U x , V x , b x are all parameters of the model, and σ is the activation function.

[0046] The input gate is as follows:

[0047]

[0048]

[0049]

[0050] Where controls the information from the input word vector x of the current news text i , g t-1 represents the sentence vector at time t-1. represents the state vector of the LSTM cell at time t-1. represents the state vector of the sentence cell at time t-1. controls the information from the left text state cell , r i t controls the information from the right text state cell , f i t controls the information from . controls the information from the sentence context cell state . We normalize so that their sum is 1. W x , U x , V x , b xThey are all parameters of the model, and σ is the activation function.

[0051] The output gate is as follows:

[0052]

[0053]

[0054]

[0055]

[0056] Among them is an output gate that is responsible for controlling the information output from the current state to the hidden state . is obtained by taking the average of the hidden vectors of all words at the previous moment . o t is the output gate that controls the output of information from the recurrent unit to g t . represents the state vector of the LSTM cell at time step t, represents the state vector of the sentence unit at time step t, g t represents the second news text vector output at the t-th time step.

[0057] S7: According to the cosine similarity between the first news text vector and the second news text vector, map the first news text vector with a cosine similarity exceeding the set threshold to text and use it as an adversarial sample of the news text;

[0058] cos(V adv , g ori ) > λ

[0059] cos([v 1 , v 2 , v 3 ,...., v n , [g 1 , g 2 , g 3 ,... g n ) > λ

[0060]

[0061]

[0062] Among them, V adv represents the first news text vector, g ori represents the second news text vector, v iDenote the sentence vector in the first news text as g i Denote the sentence vector in the second news text vector, and λ represents the set threshold.

[0063] S8: Use the adversarial sample of the news text and the news text as training samples to train the target model, create a target loss function according to the label information of the news text, and update the parameters of the target model through the mechanism of backpropagation:

[0064] J(θ,x,y)=αL(θ,x,y)+(1-α)L(θ,x adv ,y)+λ||F y (s)-F y (s adv )||

[0065] Among them, L(θ,x,y) represents the loss function of the original word x being classified as y in the classification task, and L(θ,x adv ,y) represents the loss function of the replaced synonym x adv being classified as y in the classification task. F y (s) and F y (s adv ) respectively represent the classification output confidence scores of the original sentence and the sentence after replacing the synonym.

[0066] The target model includes: a news classification model based on a neural network, for example, a CNN or RNN neural network.

[0067] The present invention extracts features from the sentence-level state and multiple word-level sub-states, representing a sentence vector containing context semantic relationships; by replacing the corresponding keywords in the original text with synonyms and combining the news text vector containing context semantic relationships, an adversarial text with a high semantic similarity to the original text can be generated, that is, a high-quality adversarial text can be obtained. Furthermore, the target model can be effectively trained using this high-quality adversarial text, so that the trained target model has high robustness.

[0068] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A defense method against synonym replacement attacks, characterized in that, it includes: S1: Obtain news text data with label information; the label information is the category of the news; S2: Use Counter-fitting to represent the word embedding vectors in the news text to obtain the hidden vectors of each sentence in the news text and generate the sentence vector sequence of the news text; S3: Use the TextRank algorithm to extract the keywords of each sentence in the news text to generate a candidate keyword set; S4: Calculate the cosine similarity between the word vector of each keyword in the candidate keyword set and the glove training word vector respectively, and take the words corresponding to the M glove training word vectors with cosine similarity greater than the set threshold as the candidate synonyms of the keyword; S5: Conduct a grammar check on the candidate synonyms of the keyword; take the candidate synonyms with the same part of speech as the keyword as the final candidate words of the keyword, and replace the corresponding keywords in the news text with the final candidate words to obtain the second news text; Embed the second news text into a vector representation to obtain the first news text vector; S6: Input the sentence vector sequence of the news text into the LSTM to calculate and obtain the second news text vector containing context information; S7: According to the cosine similarity between the first news text vector and the second news text vector, map the first news text vector with cosine similarity exceeding the set threshold to text and use it as the adversarial sample of the news text; S8: Use the adversarial sample of the news text and the news text as training samples to train the target model, create a target loss function according to the label information of the news text, and update the parameters of the target model through the mechanism of backpropagation.

2. The defense method against synonym replacement attacks according to claim 1, characterized in that, the cosine similarity between the word vector of the keyword and the glove training word vector includes: Among them, w represents the word vector of the keyword in the candidate word set, V i represents the pre-trained word vector in glove, δ represents the set cosine similarity threshold, and cosθ represents the cosine similarity between the word vector of the keyword in the candidate word set and the glove word vector.

3. The defense method against synonym replacement attacks according to claim 1, characterized in that, the mapping of the first news text vector with cosine similarity exceeding the set threshold to text and using it as the adversarial sample of the news text includes: cos(V adv ,g ori ) > λ cos([v 1 ,v 2 ,v 3 ,....,v n ,[g 1 ,g 2 ,g 3 ,...g n ) > λ Among them, V adv represents the first news text vector, g ori represents the second news text vector, v i represents the sentence vector in the first news text, g i represents the sentence vector in the second news text vector, and λ represents the set threshold.

Citation Information

Patent Citations

  • Text information recognition method and device, server and storage medium

    CN112989810A

  • Keyword-based abstract generation system and method

    CN114637840A