A Text Adversarial Attack Method Based on Gaussian White Noise

By using Gaussian white noise and multiple iteration loops in text adversarial attacks, the adversarial samples are generated and the model is attacked, and the problems of low attack efficiency and insufficient model robustness in the existing technology are solved, and more efficient parameter updates and model generalization capabilities are achieved.

CN114528832BActive Publication Date: 2025-06-10SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210136853.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-06-10
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

The existing text-adversarial attack methods have low attack efficiency and single attack methods within the controllable range, resulting in poor model robustness and insufficient generalization capabilities.

Method used

The text adversarial attack method based on Gaussian white noise is adopted. By adding perturbation terms to the gradient of the word embedding vector, an adversarial sample is generated, and model attacks are carried out through multiple iteration loops to improve the effectiveness of parameter updates.

Benefits of technology

It improves the effectiveness of parameter updates in text-defense attacks, improves the robustness and generalization capabilities of the model, and solves the problems of low attack efficiency and single attack methods in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528832B_ABST
    Figure CN114528832B_ABST
Patent Text Reader

Abstract

The present invention provides a text adversarial attack method based on Gaussian white noise, belonging to the technical field of natural language processing. First, a natural language text is input into the model. After passing through the word embedding layer, the core layer, and the linear layer, the forward training loss is obtained, and the gradient is obtained through backpropagation. Then, several perturbation terms are added to the gradient of the word embedding vector to generate an adversarial sample, and the forward calculation of the loss value is continued. After multiple iterative trainings, the model can learn the optimal solution within a very small range of changes in the current vector. The perturbation terms prompt the update of the model parameters to move in the direction of increasing loss, achieving the purpose of attacking the model. In this process, the present invention makes the intensity of parameter update have a certain randomness within a set range by adding Gaussian white noise perturbation terms, so that the model has higher fault tolerance performance and improves the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a text adversarial attack method based on Gaussian white noise. Background Art

[0002] Deep learning technology has been increasingly deeply applied in various industries, and large-scale deep learning models have been deployed in application systems such as face recognition, vehicle recognition, and machine translation. The connection between natural language processing technology and real life has become increasingly close, and various applications have been successfully launched, such as public opinion detection, intelligent search, unstructured information extraction, etc., greatly enhancing the real-life experience. As an effective means to continuously improve the generalization ability of deep learning models, adversarial attacks have been favored by the industrial community.

[0003] Adversarial attack refers to the process of adding imperceptible noise to the original sample, which makes the classification model make a wrong classification judgment on the newly constructed sample. This process is called an adversarial attack, and the newly constructed sample is called an adversarial sample. Adversarial attacks originated from the research of image classification. Researchers found that for a high-performance panda image classifier, by adding the FGSM perturbation term to the original image, the classifier would determine it as a gibbon with a confidence of 99.3%. Therefore, researchers have shown great interest in the stability and security of deep learning models. Currently, in the field of natural language processing, the research on adversarial attacks is far less than that in the image field. The reason is that the features of images are continuously variable and perceivable, while the words in text have a certain discreteness, and the semantics contained are highly complex and non-quantifiable. Therefore, the research on adversarial attacks in the field of natural language processing is a technical challenge and also has very important research value.

[0004] In the field of text adversarial attacks, the research methods in the academic community mainly include the following three methods: gradient-based attacks, confidence-based attacks, and transfer-based attacks. Among them, the gradient-based attack method uses the internal structure information of the model, solves the corresponding gradient through the loss function, perturbs the parameter gradient, and then generates adversarial samples to guide the model to move in the direction of the wrong label; the confidence-based attack method only relies on the confidence information of the model's predicted label, and does not need to know the internal structure information of the model. In terms of text adversarial attacks, the confidence-based attack method attempts to traverse and modify each word until it can interfere with the output label of the model, and regards the entire modification process as a model attack. The confidence-based text adversarial attack method is limited by the length of the text, and the attack efficiency is low; the transfer-based attack method is a method of generating adversarial samples for a specific model, and placing them in other models with different structures and different training sets, taking advantage of the transferability of adversarial samples, but the transfer-based attack method relies on the same distribution of the training set, which is difficult to achieve in reality.

[0005] In summary, the gradient-based model attack method is currently a more commonly used adversarial attack method. Developing a more effective text adversarial attack method plays a key role in improving the model's anti-interference ability. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a text adversarial attack method based on Gaussian white noise, which constructs a model interference term using a gradient-based method, solves the problems of poor model robustness and single attack method in the existing technical solutions, gives the adversarial attack partial randomness, and realizes efficient text attack within a controllable range.

[0007] The technical solution of the present invention is:

[0008] A text adversarial attack method based on Gaussian white noise, first inputting natural language text into the model, after word embedding layer, core layer, linear layer, obtaining forward training loss, and obtaining gradient through back propagation; then adding perturbation term to the gradient of word embedding vector, generating adversarial sample, and continuing forward calculation of loss value;

[0009] After several iterations of training, the model learns the optimal solution for the current vector within a very small range of variation; the disturbance term causes the model parameter update to move in the direction of increasing loss, thereby achieving the purpose of attacking the model.

[0010] Furthermore,

[0011] The process is as follows:

[0012] 1) Process the given text;

[0013] 2) Forward model training;

[0014] 3) Set the number of iterations K, the gradient change range, the initial gradient, and the initial embedding vector;

[0015] 4) Calculate the Embedding gradient interference term;

[0016] 5) Determine whether the interference term exceeds the range; if yes, execute step 6); otherwise, execute step 7);

[0017] 6) Limit the interference term within a specific range;

[0018] 7) Use the interference term to generate adversarial samples;

[0019] 8) Determine whether the current number of iterations reaches K;

[0020] 9) Restore the gradient to the initial gradient;

[0021] 10) Forward model training;

[0022] 11) Output the model detection result.

[0023] Among them,

[0024] In step 8), if the current number of iterations does not reach K, then

[0025] 8.1) Normalize the gradient to a zero matrix;

[0026] 8.2) Forward model training;

[0027] 8.3) Increase the number of iterations by 1;

[0028] 8.4) Return to step 4).

[0029] Furthermore,

[0030] Specifically include:

[0031] S1. Process the given text; set the maximum character length in the text to L, and pad the insufficient part with "padding"; the dimension of the text tensor T is RL, where R represents the real number space;

[0032] S2. Input T into the Embedding layer to obtain the embedding vector matrix X;

[0033] S3. Forward model training, and then input X into the neural network Encoder layer and two fully connected layers in sequence. The activation function of the last fully connected layer is selected as Sigmoid to obtain the forward loss Loss.

[0034] Calculate the gradient g of Loss with respect to X, and the calculation formula is as follows

[0035]

[0036] S4. Set the number of iterations \(K\), the gradient change range \(\varepsilon\), and the initial gradient \(g\) 0 \(= g\), and the initial embedding vector \(X\) 0 \(= X\); Back up the gradient \(g\) and set it as \(g\) copy ; where \(K\) is a positive integer and \(\varepsilon\) is a positive number approaching 0.

[0037] S5. For the \(t\)-th iteration, calculate the interference term \(r\) based on the gradient \(g\) in the Embedding layer t Then, judge the value range of the interference term. If \(\|r\) adv \(\|\) adv \(\|\) 2 \(>\varepsilon\), execute S6, otherwise execute S7;

[0038] where \(\|\cdot\|\) 2 represents the 2-norm of the vector, and the number of iterations \(1\leq t\leq K\) and \(t\) is a positive integer.

[0039] S6. Limit the gradient within a controllable range and correct \(r\) according to the following formula. The corrected interference term can satisfy \(\|r\) adv \(\|\) adv \(\|\) 2 \(\leq\varepsilon\)

[0040]

[0041] Then, sequentially execute step S7;

[0042] S7. Use the gradient ascent method to perform a model attack, add the interference term to the embedding vector \(X\) t to generate an adversarial sample

[0043] \(X\) t \(= X\) t-1 \(+ r\) adv

[0044] S8. If the current iteration number \(t < K\) and the gradient of the current model is \(g\) t-1 , normalize the gradient to the zero matrix, then input the adversarial sample in S6 into S3 to obtain the gradient \(g\) t and the adversarial attack loss \(Loss\) t , and update the model parameters based on this; increase the current iteration number \(t\) by 1 and continue to execute step S5;

[0045] S9. If the current iteration number \(t = K\) and the gradient of the current model is \(g\) t-1 , restore the gradient to the initial gradient, that is, \(g\) t \(= g\) 0; Then, input the adversarial examples in S6 into S3 to obtain a new gradient g t ; Restore the embedding vector X t to the original embedding vector, i.e., X t = X 0 ; Based on X t , update the parameters according to g t .

[0046] S10. In the training phase, input the text, calculate the overall loss, and perform optimization training according to the deep learning training method; in the prediction phase, input the text and output the prediction results of binary classification for the text.

[0047] Furthermore

[0048] S5.1 The deterministic part of the interference term is obtained by normalizing the gradient g t at the current stage

[0049]

[0050] S5.2 The random part of the interference term is randomly sampled from a normal distribution with a mean of 0 and a variance of

[0051]

[0052] where N(.,.) represents the Gaussian normal distribution, and f(θ) represents the probability density function of the normal distribution

[0053] S5.3 Perform a moving average on S5.1 and S5.2 to obtain an adversarial attack interference term r adv with randomness, where the parameter ρ ∈ (0, 1) is a random real number within the range of (0, 1)

[0054] r adv = ρ * m 1 + (1 - ρ) * m 2

[0055] According to the above formula, it can be obtained that

[0056] The beneficial effects of the present invention are

[0057] It solves the problems of low attack efficiency and single attack method of traditional text adversarial attack methods within a controllable range. Gaussian white noise is added to the attack interference term, and the model is attacked through multiple iterative loops, improving the effectiveness of parameter update in text adversarial attacks and enhancing the robustness and generalization ability of the model. Description of the Drawings

[0058] ​Figure 1 It is the overall implementation flowchart of the technical solution;

[0059] Figure 2 It is the overall framework diagram of the model adopted in the technical solution;

[0060] Figure 3 It is the comparison diagram of the attack intensity in the gradient interference term. Specific implementation manners

[0061] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Given the text content, the present invention proposes a text adversarial attack method based on Gaussian white noise for the problems of poor model robustness and insufficient generalization ability in the process of text adversarial attack. The generation process of adversarial samples takes the text classification task as an example, and the processing flow of adversarial samples in other natural language processing tasks is similar to that of the text classification task, which will not be elaborated here. The text classification takes binary classification as an example.

[0063] A text adversarial method based on Gaussian white noise, comprising:

[0064] S1. Process the given text. Set the maximum character length in the text to L, and the insufficient part is completed by "padding". The dimension of the text tensor T is RL, where R represents the real number space.

[0065] S2. Input T into the Embedding layer to obtain the embedding vector matrix X.

[0066] S3. Train the forward model, and then input X into the neural network Encoder layer and two fully connected layers in sequence. The activation function of the last fully connected layer is selected as Sigmoid to obtain the forward loss Loss. Calculate the gradient g of Loss with respect to X, and the calculation formula is as follows

[0067]

[0068] S4. Set the number of iterations K, the gradient change range ε, and the initial gradient g 0 = g, and the initial embedding vector X 0 = X. Back up the gradient g in step S3, and set it as g copy . Where K is a positive integer and ε is a positive number approaching 0.

[0069] S5. For the \(t\)-th iteration, based on the gradient \(g\) in the Embedding layer t calculate the interference term \(r\) adv , then determine the value range of the interference term. If \(\|r\) adv \| 2 >\(\varepsilon\), execute step S6; otherwise, execute step S7. Here, \(\|\cdot\|\) 2 represents the 2-norm of the vector, and the iteration number \(1\leq t\leq K\) and \(t\) is a positive integer.

[0070] S6. Limit the gradient within a controllable range and correct \(r\) according to the following formula. The corrected interference term can satisfy \(\|r\) adv \|\leq\varepsilon adv \| 2

[0071]

[0072] Then, sequentially execute step S7.

[0073] S7. Use the gradient ascent method to perform a model attack, add the interference term to the embedding vector \(X\) t to generate an adversarial sample

[0074] \(X\) t = \(X\) t-1 + \(r\) adv

[0075] S8. If the current iteration number \(t < K\) and the gradient of the current model is \(g\) t-1 , normalize the gradient to the zero matrix, then input the adversarial sample in step S6 into step S3 to obtain the gradient \(g\) t and the adversarial attack loss \(Loss\) t . Based on this, update the parameters of the model. Increase the current iteration number \(t\) by 1 and continue to execute step S5.

[0076] S9. If the current iteration number \(t = K\) and the gradient of the current model is \(g\) t-1 , restore the gradient to the initial gradient, i.e., \(g\) t = \(g\) 0 . Then input the adversarial sample in step S6 into step S3 to obtain the new gradient \(g\) t . Restore the embedding vector \(X\) t to the original embedding vector, i.e., \(X\) t = \(X\) 0 . Based on \(X\) t , update the parameters according to \(g\) t . Then, sequentially execute step 10.

[0077] ​S10. In the training phase, input the text, calculate the overall loss, and perform optimization training according to the deep learning training method. In the prediction phase, input the text and output the prediction results of binary classification for the text.

[0078] In step S3 described above, the overall framework diagram of the model is shown in Figure 2 . Among them, the Encoder layer is a text sentence representation processing layer, and generally uses mature network structures such as LSTM, Bert, Roberta, etc. The model involved in the present invention can be regarded as a general model.

[0079] In step S5 described above, the specific method is as follows

[0080] S5.1 For the deterministic part of the interference term, it is obtained by normalizing the gradient g t of the current stage.

[0081]

[0082] S5.2 For the random part of the interference term, it is randomly sampled from a Gaussian distribution with a mean of 0 and a variance of .

[0083]

[0084] Among them N(.,.) represents the Gaussian normal distribution, and f(θ) represents the probability density function of the normal distribution.

[0085] S5.3 By performing a moving average on steps S5.1 and S5.2, an adversarial attack interference term r adv with randomness can be obtained, where the parameter ρ ∈ (0, 1) is a random real number within the range of (0, 1).

[0086] r adv = ρ * m 1 + (1 - ρ) * m 2

[0087] According to the above formula, it can be obtained that As Figure 3 shown, within the ε-ball, the traditional method moves with a step size of m 1 . In the present invention, a random step size depending on the current gradient is given to each movement, changing the intensity of each attack.

[0088] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A text adversarial attack method based on Gaussian white noise, characterized in that, first, input a natural language text into the model. After passing through the word embedding layer, the core layer, and the linear layer, obtain the forward training loss, and obtain the gradient through backpropagation; then add an interference term to the gradient of the word embedding vector to generate an adversarial sample, and continue to calculate the loss value forward; After several iterations of training, the model learns the optimal solution within the changing range of the current vector; the interference term prompts the update of the model parameters to move in the direction of increasing loss, achieving the purpose of attacking the model; The process steps are as follows: 1) Process the given text; 2) Forward model training; 3) Set the number of iterations K, the gradient change range, the initial gradient, and the initial embedding vector; 4) Calculate the Embedding gradient interference term; 5) Judge whether the interference term exceeds the range; if yes, execute step 6), otherwise execute step 7); 6) Limit the interference term within a specific range; 7) Use the interference term to generate an adversarial sample; 8) Judge whether the current number of iterations reaches K; 9) Restore the gradient to the initial gradient; 10) Forward model training; 11) Output the model detection result; In step 8), if the current number of iterations does not reach K, then perform 8.1) Normalize the gradient to a zero matrix; 8.2) Forward model training; 8.3) Increase the number of iterations by 1; 8.4) Return to step 4); Specifically include: S1. Process the given text; set the maximum character length in the text to L, and pad the insufficient part; the dimension of the text tensor T is R L , where R represents the real number space; S2. Input T into the Embedding layer to obtain the embedding vector matrix X; S3. Forward model training, and then input X into the neural network Encoder layer and two fully connected layers in sequence. The activation function of the last fully connected layer is selected as Sigmoid to obtain the forward loss Loss; Calculate the gradient g of Loss with respect to X, and the calculation formula is as follows S4. Set the number of iterations K, the gradient change range ε, and the initial gradient g 0 = g, and the initial embedding vector X 0 = X; Back up the gradient g and set it as g copy ; where K is a positive integer and ε is a positive number approaching 0; S5. For the t-th iteration, in the Embedding layer, based on the gradient g t calculate the interference term r adv , then judge the value range of the interference term. If ||r adv || 2 > ε, execute S6, otherwise execute S7; where ||.|| 2 represents the 2-norm of the vector, the number of iterations 1 ≤ t ≤ K and is a positive integer; S6. Limit the gradient within a controllable range, and perform the correction of r according to the following formula; the corrected interference term satisfies ||r adv ; the corrected interference term satisfies ||r adv || 2 ≤ε Sequentially execute S7; S7. Use the gradient ascent method to perform a model attack, add the interference term to the embedding vector X t to generate adversarial samples x t = x t-1 + r adv S8. If the current iteration number t < K and the gradient of the current model is g t-1 , normalize the gradient to a zero matrix, then input the adversarial example in S6 into S3 to obtain the gradient g t and the adversarial attack loss Loss t . Based on this, update the parameters of the model; increase the current iteration number t by 1 and continue to execute step S5; S9. If the current iteration number t = K and the gradient of the current model is g t-1 , restore the gradient to the initial gradient, i.e., g t = g 0 ; then input the adversarial sample in S6 into S3 to obtain a new gradient g t ; restore the embedding vector X t to the original embedding vector, i.e., X t = X 0 ; update the parameters based on g t on the basis of X t . S10. In the training stage, input the text, calculate the overall loss, and perform optimization training according to the deep learning training method; in the prediction stage, input the text and output the prediction result of binary classification of the text. S5.1 The deterministic part of the interference term is obtained by normalizing the gradient g at the current stage t and is obtained by normalization S5.2 The random part of the interference term is obtained by randomly sampling from a normal distribution with a mean of 0 and a variance of ​ where $N(\cdot,\cdot)$ represents the Gaussian normal distribution, and $f(\theta)$ represents the probability density function of the normal distribution S5.3 performs a moving average on S5.1 and S5.2 to obtain the adversarial attack interference term r with randomness, where the parameter ρ ∈ (0, 1) is a random real number within the range of (0, 1). adv ​ r adv = ρ * m 1 + (1 - ρ) * m 2 According to the above formula, we can obtain .

Citation Information

Patent Citations

  • Model training method and system

    CN111027717A

  • Text classification model training method and device, text classification method and device and storage medium

    CN112131366A

  • Model training method and device, equipment, storage medium and program product

    CN112580732A

  • Processing method and system of mathematical application question answering model and storage medium

    CN112784536A