Chinese text confrontation attack method based on beam search
By combining bundle search and adaptive replacement methods of Chinese text features, the problem of poor Chinese text adversarial sample generation in the prior art is solved, and high-quality adversarial sample generation and model misleading effects are achieved.
Patent Information
- Application Number
- CN202510033047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
It is difficult for the prior art to effectively generate Chinese text adversarial samples, and the existing text adversarial sample search methods cannot effectively consider the unique characteristics of Chinese text, resulting in poor anti-sample generation results.
A beam search-based method is adopted, combining the glyph, pinyin and synonyms of Chinese text, and filtering through adaptive substitution and semantic similarity, high-quality adversarial samples are generated, and a larger potential adversarial sample space is explored through the beam search algorithm.
Effectively generate Chinese text adversarial samples that can guide natural language processing models to generate mispredictions, and retain the semantics of the original text, improving the quality and readability of the adversarial samples.
Smart Images

Figure CN119940362A_ABST
Abstract
Description
(I) Technical field
[0001] The invention relates to natural language processing technology, which is a Chinese text counterattack method based on beam search. (II) Background technology
[0002] The essence of text adversarial samples is to perturb the input text of the model without changing the semantics of the original text, but causing the natural language processing model to make incorrect predictions.
[0003] It is well known that deep neural networks for processing images are vulnerable to adversarial examples. Recent studies have shown that natural language processing using deep neural networks is also vulnerable to adversarial examples. Given that natural language processing plays an important role in information processing, such as sentiment analysis of comments used in recommendation systems and identification of toxic content in online governance. Studying adversarial examples for natural language processing models is crucial and is also a key step in achieving robustness in natural language processing. However, generating text adversarial examples is more challenging than generating image adversarial examples because text data is discrete and even small perturbations to word vectors may produce non-existent word vectors, i.e., not associated with any valid word in the language. In addition, these non-existent word vectors will result in meaningless text and also affect the readability of the text. Therefore, the adversarial example generation method in deep neural networks for image processing cannot be directly transplanted to deep neural network models for natural language processing.
[0004] The existing text adversarial sample search method is a greedy search method. For a diverse potential search space, greedy search can only obtain local optimal solutions. Unlike the existing methods that use greedy selection strategies, we use beam search technology. The beam search algorithm records k states instead of just one. It starts with k randomly generated states. In each step, all successor states of all states are generated. If one of them is the target state, the algorithm stops. Otherwise, it will select the k best successors from the list of all successor states and repeat the above operation.
[0005] In addition, existing text adversarial samples are mainly concentrated in English, and the attack methods include insertion, deletion, exchange and replacement. But they have not considered the possibility of transfer to Chinese. And Chinese itself also has some unique characteristics. Therefore, the present invention proposes an adaptive replacement method of glyphs, pinyin, and synonyms based on Chinese text features and a beam search search strategy. Through adaptive replacement, fluent and semantically preserved adversarial samples are found. Then, a larger potential adversarial sample space is explored using beam search, thereby effectively guiding the model to make incorrect predictions. (III) Summary of the invention
[0006] The goal of this invention is to provide a Chinese text anti-attack method based on beam search to solve the problem that the existing technology lacks consideration of Chinese text features. The technical content is as follows:
[0007] Step 1: Obtain a Chinese dataset for counterattack; preprocess the collected Chinese dataset. Use a Chinese word segmentation tool to segment Chinese text; then set the size b of the sample retained in each iteration of the beam search algorithm.
[0008] Step 2: Find candidate replacement words for each word in each position after word segmentation; we use Chinese glyphs, pinyin, and synonyms to find replacement words. Chinese glyphs include the splitting and replacement of radicals, the replacement of Chinese and traditional characters, and the replacement of similar characters. For pinyin replacement, we use the pinyin database to obtain homophones or characters with similar pinyin, including the replacement of front nasal sounds and back nasal sounds, and homophone replacement. Synonym replacement includes Chinese word forests and semantic elements to replace.
[0009] Step 3: We use candidate words to construct adversarial samples: For each candidate word at each position, replace the candidate word with the original word to form a new adversarial sample. Note that each adversarial sample here only changes the word once.
[0010] Step 4: Calculate the semantic similarity s between the new adversarial sample and the original text. Filter the adversarial samples whose semantic similarity exceeds the threshold. We use a universal sentence encoder for semantic similarity. The calculation of semantic similarity is as follows: x: represents the original input sentence x′: represents the sentence of the adversarial sample we constructed A: represents the semantic vector of the original input sentence after being encoded by the universal sentence encoder B: represents the semantic vector of the adversarial sample we constructed after being encoded by the universal sentence encoder n: represents the dimension of the semantic vector
[0011] Step 5: Obtain the probability P(y|x′) that the generated adversarial sample makes the model output the correct label. We select b adversarial samples that minimize the probability of outputting y. Perform the next round of iteration. Where y represents the actual label of the original sentence. The probability calculation formula for the output model predicted label is: m: represents the number of model output labels z: represents the logit vector of the output prediction of the last layer of the model i: represents the position of the true label y in the logit vector z i: is the i-th element in vector z
[0012] Step 6: Expand the obtained b adversarial samples using the method of glyph, pinyin and synonyms, then calculate the probability of the model outputting the correct label and sort them from small to large in terms of probability. If the adversarial sample with the smallest probability of outputting the target label can lead the model to make an incorrect prediction, then the process ends. Otherwise, repeat this step until the number of queries reaches the preset value or an adversarial sample that can cause the model to make an incorrect prediction is found.
[0013] The present invention utilizes the features of Chinese text such as glyphs, pinyin, synonyms, etc., and can retain more semantics of the original sentence. At the same time, based on the beam search method, it can effectively retain potential candidate adversarial samples and expand the potential search space, which promotes the effect of adversarial sample attacks. (IV) Description of the drawings
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 The figure is a flow chart of a Chinese text anti-attack method based on beam search of the present invention.
[0016] Figure 2 It is the overall framework diagram of the present invention. (V) Specific implementation methods
[0017] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.
[0018] See also Figure 1 The present invention proposes a Chinese text anti-attack method based on beam search, which includes the following steps: obtaining a Chinese data set for anti-attack; using a word segmentation tool to segment the Chinese text; performing word replacement on the segmented text; and using a beam search algorithm to find adversarial samples.
[0019] The specific steps to obtain the Chinese dataset for adversarial attacks are as follows: Step 1: Download the dataset from major open source platforms, including GitHub, huggingface, etc. Step 2: Organize the dataset to meet the format required by our attack framework, that is, it consists of Chinese sentences and label pairs.
[0020] Combination Figure 2 , we segment the original Chinese text. For example, for the Chinese sentence "This movie is very good", we use the jieba tool to obtain the word sequence after segmentation: "this", "movie", "very", "good looking".
[0021] For the word sequence after word segmentation, we use candidate words to replace the original text to obtain Chinese text adversarial samples; for each word, a pool of candidate words with similar shapes, similar pronunciations, and synonyms is generated. The specific steps of word replacement are as follows:
[0022] Step 1: Initialize the word replacement strategy: For glyph replacement, we use similar glyph mapping tables, radical split mapping tables, and traditional Chinese mutual mapping tables. These tables are eventually processed into a dictionary format. For the input word, we can get the corresponding characters with similar glyphs; for homophones, we use the homophones provided by the pinyin library; for synonyms, we use Chinese semantic elements and Cilin to generate synonyms or near-synonyms.
[0023] Step 2: Candidate word generation: For each word in each position, we generate replacement words to replace the original text one by one to form a candidate adversarial sample pool; combined Figure 2 , the replacement examples we generate are as follows: "This" is replaced by "this", "this"; "Movie" is replaced by "film", "film"; "very" is replaced by "extraordinary", "silver"; "Good-looking" was replaced with "beautiful", "decent", "expressive"; The left side represents the original word, and the right side represents the generated replacement word. Different replacement words are separated by ","; the replacement words we generated have good semantic preservation characteristics. Step 3: To further improve the quality of adversarial samples, we use semantic similarity filtering to remove adversarial samples below the threshold in the adversarial sample pool, thereby obtaining a high-quality candidate adversarial sample pool; for semantic similarity, we use the universal sentence encoder USE to obtain the embedded representation of the sentence in the semantic space, and then calculate the cosine value of the two vectors as the semantic similarity. The semantic similarity technology is: x: represents the original input sentence x′: represents the sentence of the adversarial sample we constructed A: represents the semantic vector of the original input sentence after being encoded by the universal sentence encoder B: represents the semantic vector of the adversarial sample we constructed after being encoded by the universal sentence encoder n: represents the dimension of the semantic vector
[0024] Using the obtained high-quality candidate adversarial examples, we perform beam search to explore the adversarial example space that may cause the model to make incorrect predictions. One iteration of beam search includes the following steps.
[0025] Step 1: Use the word replacement method to perturb the current sample and use semantic similarity filtering to generate a high-quality adversarial sample pool.
[0026] Step 2: Input all adversarial samples in the adversarial sample pool to the attacked pre-trained language model, and calculate the probability of each adversarial sample model outputting the correct label. Because our goal is to select the adversarial sample with the smallest probability of outputting the correct label, we use the negative value of the predicted probability of the correct label as the score of the adversarial sample: m: represents the number of model output labels z: represents the logit vector of the output prediction of the last layer of the model i: represents the position of the true label y in the logit vector z i : is the i-th element in vector z y: The correct label of the text before the attack score(·): represents the score of the adversarial sample. The higher the score, the better the attack effect.
[0027] Step 3: All adversarial samples in the candidate pool are sorted in descending order according to the scores of the adversarial samples.
[0028] Step 4: If the adversarial sample with the highest score in the candidate pool can cause the model to make an incorrect prediction, the algorithm stops; otherwise, b adversarial samples with the highest scores are selected from the adversarial sample pool to perform the next round of iteration.
[0029] The above disclosure is only a preferred embodiment of the present invention, which certainly cannot be used to limit the scope of rights of the present invention. A person skilled in the art can understand all or part of the above-mentioned implementation process, and the equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. The present invention discloses a Chinese text anti-attack method based on beam search, which mainly includes: Get the Chinese dataset for the attack; Use the word segmentation tool to segment Chinese text; replace words in the segmented text; and use the beam search algorithm to find adversarial samples.
2. According to claim 1, a Chinese text anti-attack method based on beam search is characterized in that: The word replacement of the text after word segmentation adopts Chinese character replacement, pinyin replacement and synonym replacement.
3. A Chinese text anti-attack method based on beam search according to claim 1 or claim 2, characterized in that: The pinyin replacement includes replacing the target word with a word with the same or similar pinyin.
4. The Chinese text anti-attack method based on beam search according to claim 1 is characterized in that: During each iteration, the beam search algorithm uses the negative value of the predicted probability of the correct label as the score of the adversarial sample: m: represents the number of model output labels z: represents the logit vector of the output prediction of the last layer of the model i: represents the position of the true label y in the logit vector z i : is the i-th element in vector z y: The correct label of the text before the attack score(·): represents the score of the adversarial sample. The higher the score, the better the attack effect.
5. The Chinese text anti-attack method based on beam search according to claim 1 is characterized in that: During each iteration, the beam search algorithm only adds the perturbation at one position, and sorts the generated adversarial samples from small to large according to the probability of the model predicting the correct label, retaining the first b candidate adversarial samples with the smallest probability; where b represents the beam size of the beam search, that is, the number of adversarial samples retained after each iteration.