Chinese text-oriented adversarial sample adaptive generation method and system

Through an adaptive generation method of adversarial sample for Chinese text, combined with deep learning and multiple replacement strategies, the problem of readability and semantic change in the generation of Chinese text adversarial sample is solved, and the attack effect with efficient and low modification rate is achieved, which improves the generation quality and attack success rate of Chinese text adversarial sample.

CN120337910APending Publication Date: 2025-07-18SHAANXI SCI CONTROL TECH IND RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236857.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generate Chinese text adversarial samples, which have problems such as poor readability, large semantic changes, high modification rate and low attack success rate. Most methods are based on English design and lack attack strategies for Chinese characteristics.

Method used

An adaptive generation method for adversarial sample oriented to Chinese text is designed. Through Jieba word segmentation, deep learning model training, keyword contribution calculation and a variety of replacement strategies, including synonym replacement, emoji replacement, homophone replacement, etc., the replacement strategy is adaptively selected to reduce the number of perturbations and improve attack efficiency.

Benefits of technology

The generated adversarial samples have a high attack success rate at low modification rate, the model classification accuracy decreases by more than 40%, the grammar and semantics are basically unchanged, it is difficult for human eyes to recognize, and it has a certain degree of universality and migration, which improves the generation effect of Chinese text adversarial samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337910A_ABST
    Figure CN120337910A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of adversarial sample self-adaptive generation, and discloses a Chinese text-oriented adversarial sample self-adaptive generation method and system.The Chinese text-oriented adversarial sample self-adaptive generation method comprises the steps of firstly, designing a new keyword positioning algorithm, secondly, designing various different replacement strategies according to Chinese features, and finally, carrying out Chinese text-oriented adversarial sample self-adaptive generation. And finally, adaptively selecting a replacement strategy through the model to reduce the disturbance frequency and improve the disturbance efficiency. By designing a new adversarial sample method, the problems existing in current Chinese adversarial sample research are solved, and the success rate and attack efficiency of adversarial sample attacks are improved. Through research on confrontation sample attacks, a text-based artificial intelligence model and potential safety hazards of a large model are deeply mined to design better defense measures, and the safety and controllability of the artificial intelligence model are improved. And a more effective Chinese confrontation sample can be generated. The self-adaptive attack is designed to ensure that the generated adversarial sample has a very small modification rate and has a very high attack effect at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence information security, and particularly relates to an adversarial sample adaptive generation method and system for Chinese texts. Background Art

[0002] Currently, with the rapid development and wide application of deep learning, artificial intelligence, and large models, the number of studies on the security of natural language processing (NLP) has been increasing continuously, and remarkable achievements have been made in fields such as sentiment analysis, machine translation, and question answering systems.

[0003] However, due to certain defects in the model itself, adding subtle perturbations that are not easily detected to the sample will cause the model's output to be incorrect. Such perturbed samples are called adversarial samples. The discovery of adversarial samples not only reveals the vulnerability of the model but also poses a security threat to application scenarios involving sensitive information or critical decisions. Improving the security and robustness of the model has stimulated the research enthusiasm of researchers. At the same time, with the rise of large language models, ensuring the security and reliability of the model in practical applications has also attracted more and more attention from scholars.

[0004] Compared with generating adversarial samples for images, generating adversarial samples for text still poses certain challenges. First, the biggest difference is that text is discrete, which is contrary to the continuous domain of images. This difference causes some gradient-based attacks to not be directly applicable to text. Second, when adding subtle perturbations to text, it is easy to turn words into out-of-vocabulary (OOV) words, thus affecting text understanding and reducing readability. Finally, perturbations in text easily have a great impact on semantics, while in images, the impact of perturbations on visual effects is small. In summary, it is difficult to directly use methods for images to generate text adversarial samples.

[0005] Currently, there are still some problems in some studies on text adversarial sample attacks and defenses. Text adversarial samples generated by character-level perturbations have a great impact on the readability of text, and some word correction methods can defend against character-level attacks to a certain extent. Word-level text adversarial samples have poor attack effects on certain datasets, a high modification rate, and also affect syntactic consistency. Sentence-level text adversarial samples have a low success rate, and the quality of the generated text is not easy to control. Facing these problems, a new method is needed to achieve effective attacks.

[0006] Through the above analysis, the problems and defects of the prior art are as follows:

[0007] (1) Discrete text data makes some gradient-based attacks difficult to apply in text, and perturbations such as becoming out-of-set words (OOV) make the perturbations easy to detect and the semantics change;

[0008] (2) Most of the existing Chinese adversarial perturbations are character-level with character insertion, resulting in weak readability of adversarial texts. At the same time, some text checking tools can defend against such attacks to a certain extent;

[0009] (3) The existing word-level attacks using synonym replacement have poor attack effects on certain datasets and have the problem of too high modification rates, making it difficult to maintain semantic consistency with the original text;

[0010] (4) The existing sentence-level adversarial attacks based on text addition have low success rates, and the quality of the generated text is affected by the attack cost, making it difficult to guarantee the generation quality;

[0011] (5) Most of the current text adversarial attacks are designed based on English adversarial samples and do not have adversarial perturbations designed according to the characteristics of Chinese adversarial samples. Summary of the Invention

[0012] In view of the problems existing in the prior art, the present invention provides an adaptive generation method and system for adversarial samples for Chinese texts.

[0013] The present invention is implemented as follows. An adaptive generation method and system for adversarial samples for Chinese texts includes:

[0014] Step 1: Text preprocessing;

[0015] (1) Use the Jieba tool to segment Chinese texts and perform part-of-speech tagging;

[0016] (2) Clean the text data, delete meaningless URLs, symbols, spaces, stop words, and various tags;

[0017] (3) Add corresponding digital tags to various types of data;

[0018] For sentiment classification samples, the positive sample label is set to 1 and the negative sample is set to 0; for multi-classification samples, label classification is performed starting from 0 according to the number of categories;

[0019] (4) Use the training dataset to construct word vectors, convert words into TOKENs, and normalize the text;

[0020] First, count the frequencies of all words in the training set and arrange them in descending order; count the sorted words starting from 3 and convert them into corresponding TOKEN labels. At the same time, use 0 to pad the length of the text data to ensure that the texts input to the model have the same length. Use 1 to represent the start of the text and append it at the beginning of the text. The unused 2 represents unknown words, that is, data not appearing in the training set;

[0021] Step 2: Train an effective deep learning model;

[0022] (1) Set the input matrix parameters, set the hyperparameters of the model structure, and construct the model framework using a deep learning model;

[0023] When constructing the model, set the length of the word vector according to different datasets, construct the word embedding matrix, and use random initialization to take the word embedding matrix as the first layer of the model. This step converts discrete words into continuous vector representations; input the continuous vector representations into the set deep learning model to obtain the vector output by the model, and finally, through the conversion of the linear layer and the Softmax layer, convert the output vector into the confidence score of the corresponding category;

[0024] (2) Input the preprocessed data into the model, and train and adjust the parameters of the model according to the deep learning method;

[0025] Send the data processed in step one into the model, optimize the model using the Adam optimizer, and continuously optimize the model parameters using the training set;

[0026] (3) Obtain the optimal parameters of the model, and solidify the model as a subsequent usage tool;

[0027] Save the trained model parameters and the trained model as a pkl file using the built-in saving function in pytorch for subsequent attack experiments;

[0028] The three models of LSTM / TEXTCNN / BERT are trained by adjusting the model parameters and data;

[0029] Step 3: Calculate the contribution of text keywords and locate keywords;

[0030] Step 4: Design a keyword replacement strategy and process the keywords;

[0031] Step 5: Generate adversarial examples.

[0032] Furthermore, the calculation of the contribution of text keywords and keyword location:

[0033] (1) Intercept the text to obtain the information of the context of the corresponding word;

[0034] 1) For each word in the text, remove the text after the word;

[0035] Express the i-th text in the dataset as x i ={w0, w1,..., w n-1 , w n}, where n is the fixed length of the input text. For the importance of the context information of the word w j , remove all the text after the j-th word to obtain Remove the j-th word to obtain

[0036] 2) Input the intercepted text into the model to obtain the confidence scores output by the model;

[0037] The obtained Input into the model to obtain confidence scores c = {c0, c1,..., c d} where d is the number of categories of the input text. The obtained Input into the model to obtain confidence scores c′ = {c'0, c′1,..., c' d};

[0038] 3) Calculate the change between the model scores and the corresponding labels, and use the change amount as the weight of the word related to the above text;

[0039] Suppose the i-th text expression is x i with the category k (k ∈ d), and obtain the corresponding confidence score change s1 = c k - c′ k as the above text information of the corresponding word;

[0040] (2) Intercept the text to obtain the following text information of the corresponding word;

[0041] 1) For each word in the text, remove the text before the word;

[0042] Express the i-th text in the dataset as x i = {w0, w1,..., w n-1 , w n}} where n is the fixed length of the input text. For the importance of the above text information of the word w j , remove all the text before the j-th word to obtain Then remove the j-th word to obtain

[0043] 2) Input the intercepted text into the model to obtain the confidence scores output by the model;

[0044] The obtained Input into the model to obtain confidence scores t = {t0, t1,..., t d}} where d is the number of categories of the input text. The obtained Input into the model to obtain confidence scores t′ = {t'0, t′1,..., t' d}};

[0045] (3) Calculate the change between the model score and the corresponding label, and use the change amount as the weight of the word related to the previous context;

[0046] Assume that the i-th text expression is x i The category is k (k ∈ d), and the corresponding confidence score change s2 = t k -t' k to serve as the following information of the corresponding word;

[0047] (3) Calculate the keyword contribution according to the context information of the corresponding word in the text;

[0048] Use the previous and following information corresponding to the word obtained in steps (1) and (2), and through to calculate the word contribution;

[0049] (4) Use the contribution to sort, locate and select keywords;

[0050] 1) Sort the words from largest to smallest according to the context information weight;

[0051] Calculate the context information contribution for each word in each text, record the position coordinates of the word, and sort them from largest to smallest;

[0052] 2) Select the words with high contribution as keywords for modification;

[0053] Select keywords in order from largest to smallest contribution.

[0054] Furthermore, design a keyword replacement strategy to process the keywords:

[0055] Note that all the following keyword replacement strategies should ensure that people can understand the meaning of the text after replacement, and the less noticeable the perturbation is, the better;

[0056] (1) Modify the word using a synonym;

[0057] 1) Use GloVe to calculate the word to obtain the corresponding word vector;

[0058] Use GloVe to construct a vector dictionary and convert the word into the corresponding vector;

[0059] 2) Find the closest word vector to the target word vector in the word vector space as the synonym replacement;

[0060] Find the word with the same POS as the keyword in the GloVe dictionary, calculate the cosine similarity with the keyword word vector, and select the word with the largest cosine similarity as the candidate for the synonym replacement of the current keyword;

[0061] (2) Modify words using emojis;

[0062] This modification strategy mainly targets emotional words and uses emojis to replace text for expression;

[0063] (1) Construct an emotional word library and annotate the words in the library;

[0064] Collect common emotional words, such as "happy", "sad", "sorrowful", "angry", etc.;

[0065] (2) Collect the corresponding emojis in the library and construct an emoji library;

[0066] (3) Determine whether the keyword is in the library. If it is, use an emoji for replacement; otherwise, use a synonym for replacement;

[0067] Use GloVe to construct a vector dictionary and convert words into corresponding vectors;

[0068] (3) Modify words using a dictionary;

[0069] (1) Use the nltk function to obtain the POS meaning of the word in the text

[0070] Use the part-of-speech judgment function in the nltk function library to obtain the part of speech POS of the keyword;

[0071] (2) Obtain the Chinese interpretation of the word in the dictionary and select the interpretation with the same POS as the replacement content for the keyword;

[0072] Use the dictionary to obtain the translation content of the keyword and select the translation with the same nature according to the part of speech POS of the keyword as the modified content of the keyword;

[0073] (4) Modify words using homophones;

[0074] (1) Construct a homophone dictionary;

[0075] (2) Use homophones to replace words;

[0076] (5) Modify words using homographs;

[0077] (1) Construct a homograph dictionary;

[0078] (2) Use homographs to replace words. If there are Chinese characters in the keyword that do not have homographs, use homophones for replacement;

[0079] (6) Modify words using Chinese character decomposition;

[0080] (1) Decompose commonly used Chinese characters with left-right structure and construct a Chinese character decomposition dictionary;

[0081] 2) Split the Chinese characters in the keyword. If there is a Chinese character in the keyword that is not in a left-right structure, use a homophone to replace it;

[0082] (7) Modify the word using word swapping;

[0083] Modify the order of the words

[0084] (8) Modify the word using character insertion;

[0085] Construct a set of meaningless characters. Each time a replacement is made, randomly select a character from the character set and randomly insert it at one of the three positions: "before", "in the middle", and "after" the word;

[0086] (9) Modify the word using flat and retroflex sounds;

[0087] (10) Modify the word using pinyin.

[0088] Furthermore, the modification of the word using flat and retroflex sounds is as follows:

[0089] 1) Count the Chinese characters with flat and retroflex sounds among the commonly used Chinese characters and construct a Chinese character flat and retroflex dictionary;

[0090] 2) Make the flat and retroflex sounds of the Chinese characters in the keyword. If there is a Chinese character in the keyword without a corresponding flat or retroflex sound, use a homophone to replace it.

[0091] Furthermore, the modification of the word using pinyin;

[0092] Replace the keyword with pinyin.

[0093] Furthermore, the generation of adversarial samples:

[0094] (1) Obtain the keywords sorted according to weights and process the keywords;

[0095] 1) Take out the sorted keywords one by one and obtain the replacement content by processing each word;

[0096] According to the keywords obtained in Step 3, select the keyword with the largest weight that has not been modified yet; according to the different replacement contents of the current keyword obtained in Step 4, replace the same text to obtain different adversarial samples of the same text;

[0097] 2) Make the corresponding modifications to the text to obtain a phased adversarial text, and input it into the model to obtain the model output score;

[0098] Input the different adversarial texts of the same sample obtained currently into the model to obtain the classification confidence of the model;

[0099] 3) If the model misclassifies, end step (1);

[0100] If the model has been misclassified currently, select a successfully attacked text as the modified adversarial sample, end the attack on the current text, otherwise proceed to the next step;

[0101] 4) Otherwise, select the modification scheme that has the greatest impact on the model output as the modification of the current keyword, and count the modification ratio. At the same time, continue to modify the remaining keywords and return to step 1);

[0102] For different samples of the same text obtained, select the adversarial sample that has the greatest impact on text classification among the output confidences of the model as the adversarial sample generated at this stage; use the currently generated adversarial sample as the text and continue to step 1), continue to select unmodified words for attack;

[0103] 5) When the number of modified keywords is greater than 20% of the total text length, end step (1);

[0104] (2) Obtain the modified adversarial sample and convert it into the original text according to the TOKEN;

[0105] For the obtained text adversarial sample sequence, convert it into the original text data according to the TOKEN marker and save it to a text document;

[0106] Another object of the present invention is to provide an adaptive adversarial sample generation system for Chinese text, including:

[0107] A text preprocessing module for text preprocessing; cleaning invalid data in the text to obtain text effective for classification;

[0108] A training module for training an effective deep learning model;

[0109] A calculation module for calculating the contribution of text keywords and keyword positioning;

[0110] A design module for designing a keyword replacement strategy to process keywords;

[0111] A generation module for generating adversarial samples.

[0112] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the adaptive adversarial sample generation method for Chinese text.

[0113] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the method for adaptively generating adversarial examples for Chinese text.

[0114] Another object of the present invention is to provide an information data processing terminal for implementing the system for adaptively generating adversarial examples for Chinese text.

[0115] Combined with the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solution to be protected by the present invention from the following aspects:

[0116] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving this problem, closely combined with the technical solution to be protected by the present invention and the results and data in the R & D process, etc., analyze in detail and deeply how the technical solution of the present invention solves the technical problems and the creative technical effects brought after solving the problems. The specific description is as follows:

[0117] The present invention aims to design an adaptive method for generating text adversarial examples for Chinese text features. First, a new keyword positioning algorithm is designed. Secondly, multiple different replacement strategies are designed according to Chinese features. Finally, the replacement strategy is adaptively selected by the model to reduce the number of perturbations and improve the perturbation efficiency. By designing a new adversarial example method, the problems existing in the current research on Chinese adversarial examples are solved, and the success rate and attack efficiency of adversarial example attacks are improved. Through the research on adversarial example attacks, the security risks of text-based artificial intelligence models and even large models are deeply explored to design better defense measures and promote the secure and controllable implementation of artificial intelligence models. More effective Chinese adversarial examples can be generated. An adaptive attack is designed to ensure that the generated adversarial examples have a very small modification rate and a quite high attack effect. At the same time, while ensuring that the grammar and semantics remain basically unchanged, it makes it impossible for the human eye to judge whether it is an adversarial example, but at the same time it can mislead the learning model and the attack is successful. Similarly, this method can also be extended to different languages and even large models to obtain the same effect.

[0118] (1) A new replacement strategy for adversarial examples for Chinese;

[0119] (2) Adaptation of keyword selection and keyword replacement algorithm;

[0120] (3) The present invention belongs to a hybrid attack combining sentence level and word level, and the model classification accuracy rate drops by more than 40%;

[0121] (4) It has a better attack success rate than other methods under a modification rate of less than 20%;

[0122] (5) The combined hybrid attack at the word and sentence levels makes the current defense strategy have limited defense effect;

[0123] (6) The adaptive strategy effectively reduces the modification rate and avoids the generation of invalid statements, and the average modification rate is less than 15%;

[0124] (7) The new keyword replacement strategy can have a small impact on semantics and reduce the possibility of human eye recognition, and the similarity with the original text is greater than 80%;

[0125] (8) It has certain generality, migration and scalability.

[0126] Second, by studying text adversarial sample attacks, the vulnerabilities and limitations of natural language processing (NLP) models can be revealed, thereby promoting model improvement and enhancing its robustness and generalization ability. For enterprises relying on NLP technology, this means more stable system performance and higher user satisfaction. The research on attacks can further explore the model principle and design more robust software to obtain benefits.

[0127] In practical systems such as spam detection, harmful text detection, and malware killing, the research on adversarial sample attacks helps to enhance the security of the system. By training the model to identify and resist adversarial samples, the impact of malicious attacks on the system can be reduced, protecting user data and privacy and reducing economic losses caused by such threats.

[0128] The present invention designs a method for text adversarial sample attack, which improves the efficiency of current Chinese text adversarial attacks by designing a new keyword contribution calculation method, a keyword replacement strategy that conforms to Chinese characteristics, and an adaptive keyword replacement strategy selection method, covering almost all keyword replacement strategies applicable to Chinese, and systematically studies Chinese adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0129] Figure 1 is a flowchart of an adaptive generation method for adversarial samples for Chinese text provided by an embodiment of the present invention.

[0130] Figure 2 is a block diagram of the structure of an adaptive generation system for adversarial samples for Chinese text provided by an embodiment of the present invention.

[0131] Figure 3 is a flowchart of adversarial sample generation provided by an embodiment of the present invention.

[0132] Figure 4 is an example diagram of a text replacement strategy provided by an embodiment of the present invention.

[0133] Figure 5 is a diagram showing the influence of the perturbation ratio on the classification accuracy provided by an embodiment of the present invention.

[0134] Figure 6 It is a sample diagram using the NetEase content detection API provided by an embodiment of the present invention; among them, (a) is a normal sample and (b) is an adversarial sample. Detailed implementation manners

[0135] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the following further describes the present invention in detail in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0136] This method first preprocesses the input Chinese text data, uses the Jiba tool to implement word segmentation and part-of-speech tagging, and identifies the key semantic units in the text. Then, invalid information such as website addresses, special symbols, spaces, stop words, and HTML tags is deleted through cleaning operations to ensure the purity of the text content. Subsequently, corresponding digital tags are assigned to the data according to the task type (such as sentiment classification or multi-classification), and the word frequency is statistically calculated using the training data set to construct a word vector mapping table. After converting the high-frequency words into TOKENs in sequence, the text is normalized to a unified length (using specific TOKENs to mark and fill the beginning and unknown words of the text), providing a standardized and structured data format for the subsequent model input.

[0137] The preprocessed text data is sent into a deep learning model for training. First, set the hyperparameters of the input matrix and the model structure, and use random initialization to construct a word embedding matrix to convert discrete TOKENs into continuous vectors. According to different task requirements, this method can select model architectures such as LSTM, TEXTCNN, or BERT. Through the linear layer and the Softmax layer, the output vector of the model is converted into the confidence score of the corresponding category. During the training process, the Adam optimizer is used to continuously adjust and optimize the model parameters, and finally the optimal model parameters are converged, and the trained model is solidified into a pkl file using the save function of pytorch, providing a reliable basis for subsequent adversarial attack experiments.

[0138] Based on the trained model, this method introduces a keyword contribution calculation mechanism. By analyzing the gradient information or attention weights of each input vocabulary in the prediction process of the model, the importance of each word in the classification decision is evaluated. Using this information, the system can locate the keywords that have the greatest impact on the model decision and identify which parts of the text play a key role in the final classification result. This step not only helps to reveal the internal decision logic of the model, but also provides a basis for subsequent design of keyword replacement strategies, making the generation of adversarial samples more targeted and controllable.

[0139] Based on the aforementioned keyword contribution calculation results, this method designs a series of keyword replacement strategies, including methods such as synonym replacement, part-of-speech conversion, and semantic reconstruction, to fine-tune and replace high-contribution keywords. On the premise of ensuring the overall semantic coherence of the text, by replacing local keywords, perturbed text samples are generated, so as to achieve the purpose of misleading the deep learning model to make misclassifications. The entire adversarial sample generation process realizes adaptive adjustment, ensuring that the generated samples maintain a certain degree of naturalness and readability while meeting the attack success rate, and finally realizing the adaptive generation of adversarial samples for Chinese texts.

[0140] As Figure 1 、 Figure 3 shown, a method for adaptively generating adversarial samples for Chinese texts provided by an embodiment of the present invention includes the following steps:

[0141] S101: Text preprocessing;

[0142] This step can clean the invalid data in the text and obtain text effective for classification;

[0143] (1) Use the Jieba tool to segment the Chinese text and perform part-of-speech tagging;

[0144] Since Chinese does not have natural separation by spaces like English, tools need to be used for word segmentation in order to perform better preprocessing and part-of-speech tagging, etc.;

[0145] (2) Clean the text data, delete meaningless URLs, symbols, spaces, stop words, and various tags;

[0146] The purpose of this step is to retain words meaningful for classification so that the model can extract text features more effectively and prevent overfitting of the model training caused by excessive meaningless information;

[0147] (3) Add corresponding digital labels to each category of data;

[0148] For sentiment classification samples, the positive sample label is set to 1, and the negative sample is set to 0; for multi-classification samples, label classification is performed starting from 0 according to the number of categories;

[0149] (4) Use the training data set to construct word vectors, convert words into TOKENs, and normalize the text;

[0150] First, count the frequencies of all words in the training set and sort them in descending order. Start counting the sorted words from 3 and convert them into corresponding TOKEN labels. At the same time, pad the text data with 0 to ensure that the text input into the model has the same length. Use 1 to represent the start of the text and append it to the beginning of the text. Use 2 for the unused words, that is, the data not appearing in the training set, which represents unknown words.

[0151] S102: Train an effective deep learning model;

[0152] (1) Set the input matrix parameters and the hyperparameters of the model structure, and build the model framework using a deep learning model (LSTM / TEXTCNN / BERT);

[0153] When building the model, set the word vector length according to different datasets, build the word embedding matrix, and use random initialization to make the word embedding matrix the first layer of the model. This step converts discrete words into continuous vector representations. Input the continuous vector representations into the set deep learning model to obtain the output vector of the model. Finally, through the conversion of the linear layer and the Softmax layer, convert the output vector into the confidence scores of the corresponding categories;

[0154] (2) Input the preprocessed data into the model and train and adjust the parameters of the model according to the deep learning method;

[0155] Send the data processed in S101 into the model, optimize the model through the Adam optimizer, and continuously optimize the model parameters using the training set;

[0156] (3) Obtain the optimal parameters of the model and solidify the model as a subsequent tool for use;

[0157] Save the trained model parameters and the trained model as a pkl file using the built-in saving function in pytorch for subsequent attack experiments;

[0158] The three models of LSTM / TEXTCNN / BERT are trained by adjusting the model parameters and data, and for different datasets, the classification accuracy of the three models all reaches more than 90%;

[0159] S103: Calculate the contribution of text keywords and locate the keywords;

[0160] (1) Intercept the text to obtain the information of the context of the corresponding word;

[0161] 1) For each word in the text, remove the text after the word;

[0162] Express the i-th text in the dataset as x i={w0, w1, ..., w n-1 , w n}, where n is the fixed length of the input text. For the importance of the context information of the word w j , remove all the text after the j-th word to get Then remove the j-th word to get

[0163] 2) Input the intercepted text into the model to obtain the confidence score output by the model;

[0164] Input the obtained into the model to get the confidence score c = {c0, c1, ..., c d}, where d is the number of categories of the input text. Input the obtained into the model to get the confidence score c′ = {c'0, c′1, ..., c' d};

[0165] 3) Calculate the change between the model score and the corresponding label, and use the change amount as the weight related to the word and its context;

[0166] Assume that the i-th text expression is x i with the category k (k ∈ d), and obtain the corresponding confidence score change s1 = c k - c′ k as the context information of the corresponding word;

[0167] (2) Intercept the text to obtain the context information of the corresponding word;

[0168] 1) For each word in the text, remove the text before the word;

[0169] Express the i-th text in the dataset as x i ={w0, w1, ..., w n-1 , w n}, where n is the fixed length of the input text. For the importance of the context information of the word w j , remove all the text before the j-th word to get Then remove the j-th word to get

[0170] 2) Input the intercepted text into the model to obtain the confidence score output by the model;

[0171] Input the obtained into the model to get the confidence score t = {t0, t1, ..., t d}, where d is the number of categories of the input text. Input the obtained The input model obtains a confidence score \(t'=\{t'_0,t'_1,\cdots,t' d \}\);

[0172] 3) Calculate the change between the model score and the corresponding label, and use the change amount as the weight related to the above text for the word;

[0173] Suppose the \(i\)-th text expression is \(x i with the category \(k(k\in d)\), and obtain the corresponding confidence score change \(s_2 = t k -t'\) k as the following context information for the corresponding word;

[0174] (3) Calculate the keyword contribution according to the context information of the corresponding word in the text;

[0175] Use the above and following context information corresponding to the word obtained in steps (1) and (2) to calculate the word contribution;

[0176] (4) Use the contribution to sort, locate, and select keywords;

[0177] 1) Sort the words from largest to smallest according to the context information weight;

[0178] Calculate the context information contribution for each word in each text, record the position coordinates of the word, and sort them from largest to smallest;

[0179] 2) Select the words with high contribution as keywords for modification;

[0180] Select keywords in order from largest to smallest contribution.

[0181] S104: Design a keyword replacement strategy to process the keywords, and its example is Figure 4 as shown.

[0182] Note that all the following keyword replacement strategies should ensure that people can understand the meaning of the text after replacement, and the less noticeable the perturbation, the better.

[0183] (1) Modify the word using a synonym;

[0184] 1) Use GloVe to calculate the word to obtain the corresponding word vector;

[0185] Use GloVe to construct a vector dictionary and convert the word into the corresponding vector;

[0186] 2) Find the word closest to the target word vector in the word vector space as the synonym replacement;

[0187] Search for words in the GloVe dictionary that are the same as the keyword POS, calculate the cosine similarity with the keyword word vector, and select the word with the largest cosine similarity as the candidate for synonym replacement of the current keyword.

[0188] (2) Modify the word using emojis;

[0189] This modification strategy mainly targets sentiment words and uses emojis to replace the text to represent the expression.

[0190] 1) Build a sentiment word library and label the words in the library;

[0191] Collect common sentiment words, such as "happy", "sad", "upset", "angry", etc.

[0192] 2) Collect the corresponding emojis in the library and build an emoji library;

[0193] 3) Determine whether the keyword is in the library. If it is in the library, use emojis for replacement; otherwise, use synonyms for replacement.

[0194] Use GloVe to build a vector dictionary and convert words into corresponding vectors;

[0195] (3) Modify the word using a dictionary;

[0196] 1) Use the nltk function to obtain the POS meaning of the word in the text

[0197] Use the part-of-speech judgment function in the nltk function library to obtain the part of speech POS of the keyword;

[0198] 2) Obtain the Chinese interpretation of the word in the dictionary and select the interpretation with the same POS as the keyword replacement content;

[0199] Use the dictionary to obtain the translation content of the keyword, and select the translation with the same nature according to the part of speech POS of the keyword as the modification content of the keyword.

[0200] (4) Modify the word using homophones;

[0201] 1) Build a homophone dictionary;

[0202] 2) Use homophones to replace the word;

[0203] (5) Modify the word using homographs;

[0204] 1) Build a homograph dictionary;

[0205] 2) Use homographs to replace the word. If there are no homographs for a Chinese character in the keyword, use homophones for replacement;

[0206] (6) Modify the words using Chinese character splitting;

[0207] 1) Split the commonly used Chinese characters with left-right structure, and construct a Chinese character splitting dictionary;

[0208] 2) Split the Chinese characters in the keyword. If there are Chinese characters in the keyword that are not in the left-right structure, use homophonic characters for replacement;

[0209] (7) Modify the words using word swapping;

[0210] Modify the order of the words

[0211] (8) Modify the words using character insertion;

[0212] Construct a set of meaningless characters. Each time a replacement is made, randomly select a character from the character set and randomly insert it at one of the three positions of "before", "in the middle" and "after" the word.

[0213] (9) Modify the words using flat and retroflex sounds;

[0214] 1) Count the Chinese characters with flat and retroflex sounds in the commonly used Chinese characters, and construct a Chinese character flat and retroflex dictionary;

[0215] 2) Make the flat and retroflex sounds of the Chinese characters in the keyword. If there are Chinese characters in the keyword that do not have corresponding flat and retroflex sounds, use homophonic characters for replacement;

[0216] (10) Modify the words using pinyin.

[0217] Replace the keyword with pinyin.

[0218] S105: Adversarial sample generation

[0219] (1) Obtain the keywords sorted according to weights, and process the keywords;

[0220] 1) Sequentially take out the sorted keywords, and process each word to obtain replacement content;

[0221] According to the keywords obtained in S103, select the keyword with the largest weight that has not been modified yet; according to the different replacement contents of the current keyword obtained in S104, replace the same text to obtain different adversarial samples of the same text;

[0222] 2) Make corresponding modifications to the text to obtain a phased adversarial text, and input it into the model to obtain the model output score;

[0223] Input the different adversarial texts of the same sample obtained currently into the model to obtain the classification confidence of the model;

[0224] 3) If the model misclassifies, end step (1);

[0225] If the model has been misclassified currently, select a successfully attacked text as the modified adversarial sample, end the attack on the current text, otherwise proceed to the next step;

[0226] 4) Otherwise, select the modification scheme that has the greatest impact on the model output as the current keyword modification scheme, and count the modification ratio. At the same time, continue to modify the remaining keywords and return to step 1);

[0227] For different samples of the same text obtained, select the adversarial sample that has the greatest impact on text classification among the output confidences of the model as the adversarial sample generated at the current stage; use the currently generated adversarial sample as the text and continue to step 1), continue to select unmodified words for attack.

[0228] 5) When the number of modified keywords is greater than 20% of the total text length, end step (1).

[0229] (2) Obtain the modified adversarial sample and convert it into the original text according to the TOKEN;

[0230] For the obtained text adversarial sample sequence, convert it into the original text data according to the TOKEN markers and save it to a text document.

[0231] As Figure 2 shown, an adversarial sample adaptive generation system for Chinese text provided by an embodiment of the present invention includes:

[0232] A text preprocessing module for text preprocessing; cleaning invalid data in the text to obtain text effective for classification;

[0233] A training module for training an effective deep learning model;

[0234] A calculation module for calculating text keyword contributions and keyword positioning;

[0235] A design module for designing a keyword replacement strategy and processing keywords;

[0236] A generation module for generating adversarial samples.

[0237] Another object of the present invention is to provide a computer device, where the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the adversarial sample adaptive generation method for Chinese text.

[0238] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the method for adaptively generating adversarial examples for Chinese texts.

[0239] Another object of the present invention is to provide an information data processing terminal for implementing the system for adaptively generating adversarial examples for Chinese texts.

[0240] Specific implementation of the present invention:

[0241] Table 1 Attack effects on different models and datasets

[0242]

[0243] It can be seen from the experimental results that the method proposed by the present invention has achieved better attack effects than WordHandling. It is worth noting that for the spam dataset trec06, the attack effect of the proposed method is much higher than that of WordHandling, because the content of spam is mostly commercial advertisements, pornographic marketing, sales fraud or phishing websites, etc., and the information features are relatively single and concentrated, while the content of normal emails is more extensive and diverse.

[0244] Table 2 Comparison of attack efficiencies

[0245]

[0246] The time costs required by PWWS and the proposed method are both closely related to the sentence length, but on each dataset, PWWS is 3 to 10 times slower than the proposed method. In summary, the proposed solution is far superior to PWWS in terms of time efficiency.

[0247] Specific application fields or related products of the present invention:

[0248] The present invention conducts research on text classification models, which are widely used in fields such as text content detection, sentiment analysis, intelligent recommendation, spam classification, etc., as Figure 5 shown.

[0249] In this paper, the generated adversarial texts are tested on the NetEase Yidun text detection API. The adversarial texts generated by the present invention can easily cause misclassification, as Figure 6 shown.

[0250] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.

[0251] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. An adversarial sample adaptive generation method for Chinese text, characterized in that It includes the following steps: Step 1. Text preprocessing step, including: (1) Use the Jiba tool to perform word segmentation and part-of-speech tagging on Chinese text; (2) Clean the text data, delete URLs, symbols, spaces, stop words, and various tags; (3) Add corresponding digital tags to various types of data; (4) Use the training dataset to construct word vectors, convert each word in the text into a TOKEN, and normalize the text to ensure that the input text has a unified length; Step 2. Deep learning model training step, including: (1) Set the input matrix and model structure hyperparameters, construct a word embedding matrix, and randomly initialize discrete TOKENs as continuous vectors; (2) Use deep learning models such as LSTM, TEXTCNN, or BERT to construct a classification model, and convert the model output into confidence scores for corresponding categories through a linear layer and a Softmax layer; (3) Use the Adam optimizer to train and adjust the parameters of the preprocessed data until the optimal model parameters are obtained, and solidify them into the model file for subsequent attack experiments; Step 3. Keyword contribution calculation and localization step, including: Based on the trained model, calculate the contribution of each input vocabulary to the model prediction through gradient information or attention weight analysis, and locate the keywords that have the greatest impact on the classification result; Step 4. Keyword replacement strategy design and adversarial sample generation step, including: According to the located keywords, design synonym replacement, part-of-speech conversion, or semantic reconstruction strategies, and appropriately perturb the high-contribution keywords, thereby generating perturbed adversarial samples to make the trained deep learning model produce misclassifications while maintaining the semantic coherence of the text.

2. The adversarial example adaptive generation method for Chinese text according to claim 1, wherein, The calculation of the contribution of text keywords and keyword localization: (1) Intercept the text to obtain the information above the corresponding word; 1) For each word in the text, remove the text after the word; Express the i-th text in the dataset as x i ={w0, w1,..., w n-1 , w n}, where n is the fixed length of the input text. For the importance of the context information of the word w j , remove all the text after the j-th word to get Then remove the j-th word to get 2) Input the intercepted text into the model to obtain the confidence score output by the model; The obtained input model yields confidence scores c = {c0, c1,..., c d}, where d is the number of classes of the input text. The obtained input model yields confidence scores c' = {c'0, c'1,..., c' d}; 3) Calculate the change between the model score and the corresponding label, and use the change amount as the weight related to the word and the above text; Suppose the $i$-th text expression is $x$. i Its category is $k$ ($k \in d$), and the corresponding confidence score change is obtained as $s_1 = c$. k $- c'$. k to be used as the upstream information of the corresponding word. (2) Intercept the text to obtain the information below the corresponding word; 1) For each word in the text, remove the text before the word; Express the i-th text in the dataset as x i ={w0, w1,..., w n-1 , w n}, where n is the fixed length of the input text. For the importance of the context information of the word w j , remove all the text before the j-th word to get Then remove the j-th word to get 2) Input the intercepted text into the model to obtain the confidence score output by the model; The obtained input model yields confidence scores \(t = \{t_0, t_1, \ldots, t\) d \}\), where \(d\) is the number of classes of the input text. The obtained input model yields confidence scores \(t'=\{t'_0, t'_1, \ldots, t'\) d \}; 3) Calculate the change between the model score and the corresponding label, and use the change amount as the weight related to the word and the above text; Suppose the i-th text expression is x i The category is k (k ∈ d), and the corresponding confidence score change s2 = t k -t' k to be used as the context information of the corresponding word; (3) Calculate the keyword contribution degree according to the context information of the corresponding word in the text; Using the above and below context information corresponding to the words obtained in step (1) and step (2), through to calculate the word contribution degree; (4) Use the contribution degree to sort, locate, and select keywords; 1) Sort the words from large to small according to the context information weight; Calculate the contribution degree of context information for each word in each text, record the position coordinates of the words, and sort them from large to small; 2) Select the words with high contribution degree as keywords for modification; Select keywords in order from large to small contribution degree in turn.

3. The adversarial example adaptive generation method for Chinese text according to claim 1, characterized in that The design of the keyword replacement strategy and the processing of keywords: Note that all the following keyword replacement strategies should ensure that people can understand the meaning of the text after replacement. At the same time, the less noticeable the perturbation is, the better. (1) Modify words using synonyms. 1) Use GloVe to calculate words and obtain corresponding word vectors. Use GloVe to construct a vector dictionary and convert words into corresponding vectors. 2) In the word vector space, find the word vector closest to the target word vector as a synonym replacement. In the GloVe dictionary, find words with the same POS as the keyword, calculate the cosine similarity with the keyword word vector, and select the word with the largest cosine similarity as the candidate for synonym replacement of the current keyword. (2) Modify words using emojis. This modification strategy mainly targets sentiment words and uses emoji replacement text to replace expressions. 1) Construct a sentiment word library and annotate the words in the library. Collect common sentiment words, such as "happy", "sad", "upset", "angry", etc. 2) Collect the corresponding emojis in the library and construct an emoji word library. 3) Determine whether the keyword is in the library. If it is in the library, use emojis for replacement; otherwise, use synonyms for replacement. Use GloVe to construct a vector dictionary and convert words into corresponding vectors. (3) Modify words using a dictionary. 1) Use the nltk function to obtain the POS meaning of the word in the text. Use the part-of-speech judgment function in the nltk function library to obtain the POS of the keyword. 2) Obtain the Chinese interpretation of the word in the dictionary and select the interpretation with the same POS as the keyword replacement content. Use the dictionary to obtain the translation content of the keyword, and select the translation with the same nature according to the POS of the keyword as the modification content of the keyword. (4) Modify words using homophones. 1) Construct a homophone dictionary. 2) Use homophones to replace words. (5) Modify words using homographs. 1) Construct a homograph dictionary. 2) Use homographs to replace words. If there are Chinese characters in the keyword that do not have homographs, use homophones for replacement. (6) Modify words using Chinese character splitting. 1) Split commonly used Chinese characters with left-right structures to construct a Chinese character splitting dictionary. 2) Split the Chinese characters in the keyword. If there are Chinese characters in the keyword that are not left-right structures, use homophones for replacement. (7) Modify words using word swapping. Modify the order of words. (8) Modify words using character insertion. Construct a set of meaningless characters. Each time, randomly select a character from the character set and randomly insert it in the "front", "middle", and "back" positions of the word. (9) Modify words using flat and curled tongue sounds. (10) Modify words using pinyin.

4. The adversarial example adaptive generation method for Chinese text according to claim 3, wherein The modification using flat and curled tongue sounds for words: 1) Count the Chinese characters with flat and curled tongue sounds in commonly used Chinese characters to construct a Chinese character flat and curled tongue dictionary. 2) Make the flat and curled tongue sounds of the Chinese characters in the keyword. If there are Chinese characters in the keyword that do not have corresponding flat and curled tongue sounds, use homophones for replacement.

5. The adversarial sample adaptive generation method for Chinese text according to claim 3, wherein The modification using pinyin for words: Replace the keyword with pinyin.

6. The adversarial sample adaptive generation method for Chinese text as described in claim 1, wherein The generation of adversarial examples: (1) Obtain the keywords sorted according to weights and process the keywords; 1) Sequentially take out the sorted keywords, and perform keyword processing on each word to obtain replacement content; According to the keywords obtained in Step 3, select the keyword with the largest weight that has not been modified currently; according to the different replacement contents of the current keyword obtained in Step 4, replace the same text to obtain different adversarial samples of the same text; 2) Make corresponding modifications to the text to obtain a phased adversarial text, and input it into the model to obtain the model output score; Input the different adversarial texts of the same sample obtained currently into the model to obtain the classification confidence of the model; 3) If the model misclassifies, end Step (1); If the model has been misclassified currently, then select a successfully attacked text as the modified adversarial sample, end the attack on the current text, otherwise proceed to the next step; 4) Otherwise, select the modification scheme that has the greatest impact on the model output as the modification of the current keyword, and count the modification ratio. At the same time, continue to modify the remaining keywords and return to Step 1); For different samples of the same text obtained, select the adversarial sample that has the greatest impact on the text classification among the output confidences of the model as the adversarial sample generated at the current stage; use the currently generated adversarial sample as the text and continue to go to Step 1), and continue to select unmodified words for attack; 5) When the number of keyword modifications is greater than 20% of the total text length, end Step (1); (2) Obtain the modified adversarial sample and convert it into the original text according to the TOKEN; For the obtained text adversarial sample sequence, convert it into the original text data according to the TOKEN marker and save it to a text document.

7. An adversarial example adaptive generation system for Chinese text that implements the adversarial example adaptive generation method for Chinese text according to any one of claims 1-6, characterized in that, The adversarial sample adaptive generation system for Chinese text includes: A text preprocessing module for text preprocessing; cleaning invalid data in the text to obtain text effective for classification; A training module for training an effective deep learning model; A calculation module for calculating the contribution of text keywords and keyword positioning; A design module for designing a keyword replacement strategy and processing keywords, and its example is shown in Figure 2; A generation module for generating adversarial samples.

8. A computer device, characterized in that, The computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the adversarial sample adaptive generation method for Chinese text as described in any one of Claims 1-6.

9. A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the adversarial sample adaptive generation method for Chinese text as described in any one of Claims 1-6.

10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the adversarial sample adaptive generation system for Chinese text as described in Claim 7.