An Adversarial Attack Method Based on Modified Boundary Attack

By analyzing the contribution degree of different pixel points in the image in a decision-based attack method and combining the successful and failed sampling information, an adversarial attack method based on corrected boundary attack is proposed, which solves the problem of failure to effectively utilize historical query information and failed sampling in the prior art, and achieves more efficient noise compression and adversarial sample generation.

CN111160400BActive Publication Date: 2025-05-27TIANJIN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201911245233.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-06
Publication Date
2025-05-27
Estimated Expiration
2039-12-06

AI Technical Summary

Technical Problem

Existing decision-based attack methods fail to effectively utilize the historical query information of the target model, resulting in a lack of targetedness in generating adversarial samples and failing to fully utilize the failed sampling information to optimize noise compression.

Method used

A confrontational attack method based on corrected boundary attack is proposed. By analyzing the contribution degree of different pixel points in the image to the wrong segmentation, attacking pixel points with larger contributions, and combining the sampling information of success and failure, a confrontation sample with stronger attack capabilities is constructed. Specific steps include constructing a noise set, initializing the perturbation spatial parameters, correcting the boundary attack, generating new adversarial samples, and attacking by adaptively adjusting the noise step size.

Benefits of technology

By correcting the boundary attack method, a higher noise compression amplitude is achieved under the same number of queries, which significantly improves the attack effect against samples, and has stronger noise compression capabilities compared with other boundary attack methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111160400B_ABST
    Figure CN111160400B_ABST
Patent Text Reader

Abstract

The present invention discloses an adversarial attack method based on modified boundary attack. Step 1: Collect image and label information to form <image, category> pairs and construct an image data set. Step 2: Take the original image x i , and then obtain the set x * composed of adversarial samples. Step 3: Construct a noise set z * , and construct and initialize a set of perturbation space parameters W. Step 4: By calculating the mean of the perturbation space parameters W, construct a perturbation space, randomly sample the perturbation in the perturbation space, and generate a set of vectors η in the tangential direction of the noise. Step 5: Perform modified boundary attack to construct a new adversarial sample x'. Step 6: Input the new adversarial sample x' into the target model to adjust the perturbation space parameters W. Step 7: Repeat Steps 4, 5, and 6 a total of B - 1 times to obtain the final adversarial sample x', and input the adversarial sample into the target model for classification to obtain the classification result F(x'). The present invention achieves the purpose of constructing adversarial samples with stronger attack capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning security, and particularly to a method for black-box adversarial decision-making attack for a depth image recognition system. Background Art

[0002] Decision-based attack is an important category in adversarial attack methods. Different from iterative or optimization-based attack methods, decision-based attacks do not require a large amount of computational resources to repeatedly differentiate the target model. Instead, they randomly walk within the input space of the original image and query the target model a certain number of times to achieve non-black-box attacks and compression of adversarial noise, and can generate adversarial samples with smaller noise amplitudes with higher efficiency and fewer restrictions. However, existing decision-based attacks, such as Boundary Attack, Evolutionary Attack, etc., do not model the sensitivity of the target model to the noise of each pixel of the image using historical queries of the target model. Decision-based attacks essentially sample within the neighborhood of the original image in the input space and search for the smallest possible noise amplitude change while ensuring misclassification. In fact, different pixel points in an image contribute differently to misclassification, and different models also have different sensitive regions for images. All this information can be obtained through historical queries of the target model. It can be considered that the results of historical queries are an unbiased approximation of the noise sensitivity of pixel points.

[0003] On the other hand, failed samples (i.e., samples falling into the correct category) in decision attacks actually contain the position information of the decision boundary. Although failed samples cannot directly compress the noise amplitude, they characterize the direction with a higher probability of crossing the decision boundary. Since the attacker hopes that as many samples as possible fall on the other side of the decision boundary relative to the correct category, the information of failed samples can make new samples avoid areas with a higher probability of failure as much as possible. However, current decision-based attack methods do not utilize this key information containing the decision boundary of the target model. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes an adversarial attack method based on modified boundary attack, which analyzes the contribution degree of different pixel points in an image to misclassification, attacks the pixel points with greater contribution, and combines successful and failed samples to achieve the purpose of constructing more powerful adversarial samples.

[0005] The adversarial attack method based on modified boundary attack of the present invention includes the following steps:

[0006] Step 1, collect image and label information, form <image, category> pairs, and construct an image data set;

[0007] Step 2: Take the original image x i , and add random Gaussian noise to x i to obtain such that the target classifier (DNN) outputs a classification result F(x i * ) ≠ y i , and then obtain the set x * ;

[0008] Step 3: Construct the noise set z * , and the expression is as follows:

[0009]

[0010] Construct and initialize the perturbation space parameter set W, and the expression is as follows:

[0011]

[0012] Step 4: By calculating the mean of the perturbation space parameter W, construct the perturbation space, and randomly sample the perturbation in the perturbation space to obtain the set η of the tangential direction vectors of the noise, and the expression is as follows:

[0013]

[0014] Step 5: Modify the boundary attack according to the following formula:

[0015]

[0016] where, represents the pixel point with the largest absolute value in z * , r represents the ratio of the number of pixel points included in the new sampling to the number of pixel points of the current noise, that is, the pixel retention rate;

[0017] The modified boundary attack operation selects the pixel point with the largest absolute value in the current noise according to the ratio of r and forms a mask T to filter out insensitive image regions; T constructs a screening mechanism for the image noise region while effectively compressing the sampling space.

[0018]

[0019] Thus, construct a new adversarial sample x':

[0020]

[0021] where, δ is the tangential step size of the added noise, and ε is the radial step size of the added noise, both of which are hyperparameters of this algorithm;

[0022] Step 6: First, input the new adversarial sample \(x'\) into the target model, denoted as \(F(\cdot)\). Then, use the adversarial sample construction method with an adaptively adjusted noise step size to attack the target model. Adjust \(x\) according to the result returned by the target model * and adjust the perturbation space parameter \(W\):

[0023] If \(F(x') \neq y\), that is, the output result of the model for the adversarial sample \(x'\) is inconsistent with its true class label, it indicates that the sampling is successful, which also means the attack is successful. At this time, further compress the noise, replace \(x\) with \(x'\) * and set the perturbation space parameter set \(W\) to an empty set

[0024] x * = x',

[0025] If \(F(x') = y\), it means the sampling fails. At this time, record the failed sampling and feedback it to \(x\) * , that is, update \(\eta\) to the perturbation space parameter set \(W\):

[0026] W = W ∪ \(\eta\);

[0027] Step 7: Repeat Step 4, Step 5, and Step 6 a total of \(B - 1\) times, where \(B\) is the maximum number of queries for each image. Obtain the final adversarial sample \(x'\), and input the adversarial sample into the target model for classification to obtain the classification result \(F(x')\);

[0028] The attack effect is measured by the noise compression amplitude \(\theta\) of the adversarial sample:

[0029]

[0030] where \(X\) represents the set of test images, \(x'\) represents the adversarial sample generated by the decision attack, \(x^*\) represents the initial adversarial sample, \(|X|\) represents the total number of elements in \(X\), and \(\theta \in (0, 1)\) is used to measure the noise compression ability of the decision attack

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] Compared with other adversarial attack methods of boundary attacks, an adversarial attack method based on modified boundary attack of the present invention only adjusts the pixels with a relatively large current noise amplitude during each attack, and at the same time combines the sampling information of both success and failure to guide new sampling. The modified boundary attack achieves the highest noise compression amplitude with the same number of queries on different target models BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the overall flowchart of an adversarial attack method based on modified boundary attack of the present invention

[0034] Figure 2 Comparison chart of different effects of adversarial samples generated by different attack methods;

[0035] Figure 3 Comparison chart of the change of θ with the change of B and r on Tiny-Imagenet. Detailed implementation manners

[0036] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it shall not be used as a basis for limiting the present invention.

[0037] As Figure 1 shown, it is the overall flowchart of an adversarial attack method based on modified boundary attack of the present invention.

[0038] Step 1: Form <image, category> pairs from the collected image and label information. There are a total of n categories for all images, and the category labels here are 0 to (n - 1);

[0039] Use a large-scale image classification dataset (ImageNet) to form an image set (Img):

[0040]

[0041] wherein, x i represents the RGB pixel value of the i-th image, and its dimension is W×H×C, which respectively represent the width, height, and number of channels of the image (here it is 3), and N d represents the total number of images in the image set (Img);

[0042] Construct an image description set (Label) corresponding to each image in the image set (IMG):

[0043]

[0044] wherein, y i represents the category number corresponding to the i-th image;

[0045] The final dataset is composed of the image set (Img) and the image description set (Label) corresponding to each image;

[0046] Step 2: Take the original image x i , add random Gaussian noise to x i to obtain such that the output classification result of the target classifier (DNN) F(x i * ) ≠ y i , and then obtain the set x * composed of adversarial samples:

[0047]

[0048] Among them, is x i the adversarial sample obtained by adding random Gaussian noise;

[0049] Step 3, construct and initialize the attack parameters:

[0050] Construct the noise z i * :

[0051]

[0052] The noise set z * The expression is as follows:

[0053]

[0054] Construct and initialize the perturbation space parameter set W i is an empty set

[0055]

[0056] The expression of the perturbation space parameter set W is as follows:

[0057]

[0058] Step 4, by calculating the mean value of the perturbation space parameter W, construct the perturbation space, randomly sample the perturbation in the perturbation space, and generate η i , representing the tangential direction vector of the sampled noise:

[0059]

[0060] The set η of the tangential direction vectors of the noise is expressed as follows:

[0061]

[0062] When that is, when the set W is an empty set,

[0063]

[0064]

[0065] Step 5, correct the boundary attack according to the following formula, and only adjust the pixels with a relatively large current noise amplitude. As shown in the following formula, that is, only change the values of the top r largest pixels:

[0066]

[0067] Among them, represents the pixel point with the largest absolute value in z * The ratio of the number of pixel points included in the new sampling to the number of pixel points of the current noise, that is, the pixel retention rate.

[0068] The modified boundary attack operation selects the pixel point with the largest absolute value in the current noise according to the ratio of r, and forms a mask T to filter out insensitive image regions. While effectively compressing the sampling space, T constructs a screening mechanism for the image noise region.

[0069]

[0070] Thus, a new adversarial sample x′ is constructed:

[0071]

[0072] Among them, δ is the tangential step size of the added noise, and ε is the radial step size of the added noise, both of which are hyperparameters of this algorithm.

[0073] Step 6, first input the new adversarial sample x′ into the target model. Here, the target model refers to the deep neural network model Inception-v3, which includes convolutional operations, pooling operations, etc., denoted as F(·). Then use the adversarial sample construction method with adaptive adjustment of the noise step size to attack the target model, and adjust x * and the perturbation space parameter W:

[0074] If F(x′)≠y, that is, the output result of the model for the adversarial sample x′ is inconsistent with its true class label, it means that the sampling is successful, which also means that the attack is successful. At this time, further compress the noise, replace x* with x′ and set the perturbation space parameter set W to an empty set

[0075] x* = x′,

[0076] If F(x′) = y, it means that the sampling fails. At this time, record the failed sampling and feedback it to x*, that is, update η to the perturbation space parameter set W:

[0077] W = W ∪ η

[0078] Step 7, repeat Step 4, Step 5, and Step 6 a total of B - 1 times, where B is the maximum number of queries for each image. Obtain the final adversarial sample x′, and input the adversarial sample into the target model for classification to obtain the classification result F(x′).

[0079] The attack effect is measured by the noise compression amplitude θ of the adversarial sample:

[0080]

[0081] Among them, X represents the set of test images, x' represents the adversarial example generated by the decision attack, x* represents the initial adversarial example, and |X| represents the total number of elements in X. θ ∈ (0, 1) is used to measure the noise compression ability of the decision attack. A higher θ indicates that the attack method can compress the adversarial noise to a lower level with the same number of queries.

[0082] As Figure 2 shown, it is a comparison chart of the different effects of adversarial examples generated by different attack methods. The leftmost of each row is the original image, comparing the C&W attack (Whey), the boundary attack (Boundary), the Bayesian boundary attack (BiasedBoundary), and the optimization attack (Evolutionary). The rightmost is the adversarial example generated by the adversarial example construction method of the modified boundary attack of the present invention. After adding the adversarial noise generated by the adversarial example construction method of the modified boundary attack, the classification results on the Inception-v3 model change from (waterbird, goldfish, hammerhead shark, red sea turtle, green mamba) to (redshank, starfish, lizard, hippopotamus, eel) from top to bottom. Since the modified boundary attack uses the current noise to correct the sampled normal distribution, it can be seen that the amplitude in the region with a higher noise amplitude has been significantly compressed.

[0083] As Figure 3 shown, it is the change of the compression amplitude θ with the changes of B and r on Tiny-Imagenet. Among them Figure 3 (a)(b)(c) represent the change of the compression amplitude θ of the algorithm proposed by the present invention with the change of the query number B after different attack algorithms. More query numbers B can provide more opportunities for the decision attack to compress the adversarial noise. It can be seen that the modified boundary attack has a noise compression amplitude exceeding other methods at all query numbers. Figure 3 (d) represents the change of the compression amplitude θ with the change of the pixel retention rate r. The pixel retention rate is related to the dimensional compression of the sampling space. There is a balance between exploration and exploitation for this parameter. The smaller r is, the more concentrated the sampling process is on the regions where the noise amplitude is already large. However, if r is too small, only a small number of pixels with the largest noise amplitude will be retained. Therefore, the selection of this parameter needs to balance the size of the search space and the noise compression efficiency.

[0084] Experiments show that, compared with the boundary attack, θ of the modified boundary attack can reach 2-3 times that of the boundary attack in some cases, which verifies the effectiveness of adjusting the sampling distribution according to the current noise and using the historical failed queries.

Claims

1. A counterattack method based on modified boundary attack, characterized in that, the method comprises the following steps: Step 1, collect image and label information, form <image, category> pairs, and construct an image data set; Step 2, obtain the original image x i , add random Gaussian noise to x i to obtain x i * . Input the adversarial sample x i * into the target model for classification to obtain the classification result such that the target classifier outputs the classification result and then obtain the set x composed of adversarial samples * , where y i represents the class number corresponding to the i-th image; Step 3, construct the noise set z * , and the expression is as follows: Construct and initialize the perturbation space parameter set W, and the expression is as follows: where N d represents the total number of images in the image set (Img); Step 4, by calculating the mean value of the perturbation space parameter W, construct a perturbation space, randomly sample the perturbation in the perturbation space, and obtain the set η of the tangential direction vectors of the noise, and the expression is as follows: Step 5, modify the boundary attack according to the following formula: Among them, represents the pixel point with the largest absolute value in z * r represents the ratio of the number of pixel points included in the new sampling to the number of pixel points of the current noise, that is, the pixel retention rate; The modified boundary attack operation selects the pixel points with the largest absolute value in the current noise according to the ratio of r, and forms a mask T to filter out insensitive image regions; T constructs a screening mechanism for the image noise region while effectively compressing the sampling space; Thus, a new adversarial sample x′ is constructed: where δ is the tangential step size of the added noise, and ε is the radial step size of the added noise, both of which are hyperparameters of this algorithm; Step 6: First, input the new adversarial sample \(x'\) into the target model, denoted as \(F(\cdot)\), and then use the adversarial sample construction method with an adaptively adjusted noise step size to attack the target model. Adjust \(x\) according to the result returned by the target model * and adjust the perturbation space parameter \(W\): If F(x′)≠y, that is, the output result of the target model for the adversarial sample x′ is inconsistent with its true class label, it indicates that the sampling is successful, which also means that the attack is successful. At this time, further compress the noise and replace x with x′ * And set the perturbation space parameter set W to an empty set If F(x′) = y, it indicates sampling failure. At this time, the failed sampling is recorded and fed back to x * , that is, η is updated to the perturbation space parameter set W: W = W ∪ η; Step 7, repeat Step 4, Step 5, and Step 6 for a total of B - 1 times, where B is the maximum number of queries for each image, obtain the final adversarial sample x′, and input the adversarial sample into the target model for classification to obtain the classification result F(x′); The attack effect is measured by the noise compression amplitude θ of the adversarial sample: Among them, \(X\) represents the set of test images, \(x'\) represents the adversarial sample generated by the decision attack, and \(x\) * represents the initial adversarial sample, \(|X|\) represents the total number of elements in \(X\), and \(\theta\in(0,1)\) is used to measure the noise compression ability of the decision attack.

Citation Information

Patent Citations

  • A step size self-adaptive attack resisting method based on model extraction

    CN109948663A