Black box adversarial sample generation method based on Bayesian evolutionary optimization

By using Bayesian evolutionary optimization and Gaussian process approximation neural network in the black box attack method, the problem of high computing resources consumption of existing black box attack methods is solved, and efficient adversarial sample generation and attack effects are achieved.

CN120014382APending Publication Date: 2025-05-16NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411982358.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing black box attack methods require a large number of query target models, which makes computing resources expensive and difficult to implement effective attacks in the real world.

Method used

Using Bayesian evolutionary optimization method, the Gaussian process is constructed to approximate the replacement of neural networks, reducing the number of neural network queries required for adversarial sample generation.

Benefits of technology

It significantly reduces the number of neural network queries required to generate adversarial samples, reduces the consumption of computing resources, and improves the attack effect of adversarial samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014382A_ABST
    Figure CN120014382A_ABST
Patent Text Reader

Abstract

The invention discloses a black box adversarial sample generation method based on Bayesian evolutionary optimization. The method comprises the following steps: determining a neural network classifier to be attacked; constructing a Gaussian process and an acquisition function; obtaining an original image, and setting an optimization variable; acquiring a training data set; training a Gaussian process by using the training data set to approximately replace a neural network classifier; optimizing the optimization variables by using a Gaussian process, and generating candidate adversarial samples based on the optimization variables; taking the candidate adversarial samples as the input of a neural network classifier, and outputting the candidate adversarial samples as adversarial samples in response to the fact that the output category is wrong; and in response to the fact that the output category is correct, performing image data sampling by using the acquisition function, adding the sampled image and the category thereof as training data into the training data set, and performing Gaussian process training again until an adversarial sample is obtained. According to the method, the number of times of neural network query required for generating the adversarial sample can be reduced, and consumption of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a black-box adversarial sample generation method based on Bayesian evolutionary optimization. Background Art

[0002] In recent years, with the continuous progress and development of machine learning theory and technology, especially the breakthroughs in computer vision and multimedia, machine learning has been widely used in technical fields such as autonomous driving, target detection, medical image processing, biological image recognition, and face recognition. However, the rapid development of machine learning has also brought many security issues. Studies have shown that the output of deep neural networks is highly sensitive to small changes in the input. For example, for image classification tasks, by adding some imperceptible perturbations to the original image, the deep neural network can make incorrect output results. The image with added perturbations is called an adversarial sample. From the perspective of the human eye, there is no difference between the adversarial sample and the corresponding original image.

[0003] The process of generating adversarial samples that cause deep neural networks to predict incorrect results is defined as an adversarial attack. Adversarial attacks can be divided into white-box attacks and black-box attacks. White-box attacks refer to known information such as the structure and parameters of the target model, and adversarial samples can be constructed by analyzing the structure of the target model to conduct adversarial attacks. Black-box attacks refer to the inability to obtain the internal structure and parameters of the target model, but the target model can be accessed and queried.

[0004] At present, black-box attacks mainly include transfer-based attack methods and query-based attack methods. In the transfer-based attack method, a local model is selected as the proxy model of the attacked target model, and then adversarial samples are generated for the local model, and the adversarial samples are used to attack the target model. However, due to the limited transferability of adversarial samples between two different models, the attack performance and effect of the transfer-based attack method are poor. In the query-based attack method, a slight perturbation is added to the input image, the output change of the target model is observed, and the gradient of the target model is roughly estimated through a series of queries. After obtaining the model gradient, adversarial samples are generated by using the gradient ascent technique, and then the adversarial samples are used to attack the target model. The query-based attack method has good attack performance and effect, but in order to generate the corresponding adversarial samples, the query-based attack method needs to perform a large number of queries on the target model, which consumes more computing resources, and a large number of queries are also prone to cause system monitoring, which is usually difficult to achieve in the real world. Summary of the invention

[0005] In order to solve some or all of the technical problems existing in the above-mentioned prior art, the present invention provides a black-box adversarial sample generation method based on Bayesian evolutionary optimization.

[0006] The technical solution of the present invention is as follows:

[0007] A black-box adversarial sample generation method based on Bayesian evolutionary optimization is provided, including:

[0008] Determine the neural network classifier to be attacked;

[0009] Construct Gaussian processes and acquisition functions;

[0010] Obtain the original image and its corresponding category output by the neural network classifier, and set the optimization variable and its initial value. The optimization variable is the weight of the pixel difference between the original image and the image in its neighborhood.

[0011] Obtaining a training data set, the training data including sample images and their corresponding categories output by a neural network classifier;

[0012] Using the training data set to train the Gaussian process to approximately replace the neural network classifier;

[0013] Use Gaussian process to optimize the optimization variables, and generate candidate adversarial samples based on the optimization variables and the original image;

[0014] The candidate adversarial sample is used as the input of the neural network classifier. In response to the output category being wrong, the candidate adversarial sample is output as the adversarial sample. In response to the output category being correct, the image data is sampled using the acquisition function, the sampled image and its corresponding category output by the neural network classifier are added to the training data set as training data, and the Gaussian process training, optimization variable optimization and candidate adversarial sample generation process are re-performed until the adversarial sample is obtained.

[0015] In some possible implementations, the acquisition function includes: any one of a PI function, an EI function, an LCB function, and a UC function.

[0016] In some possible implementations, candidate adversarial samples are generated based on the optimization variables and the original image in the following manner:

[0017] Calculate the pixel difference between the original image and its neighboring image;

[0018] Calculate the product of each pixel difference and the corresponding optimization variable respectively and superimpose them;

[0019] The superposition results are processed using the sign function to determine the disturbance direction;

[0020] According to the determined disturbance direction, a preset disturbance value is added to the original image to obtain a candidate adversarial sample.

[0021] In some possible implementations, the method further includes:

[0022] For each acquisition function, set the corresponding selection probability;

[0023] Before using the acquisition function to sample image data, a random number is generated, and the acquisition function is selected based on the relationship between the random number and the selection probability.

[0024] In some possible implementations, when the selection probability is a numerical range, selecting the acquisition function according to the relationship between the random number and the selection probability includes: selecting the acquisition function corresponding to the selection probability to which the random number belongs;

[0025] When the selection probability is a numerical value, selecting the acquisition function according to the relationship between the random number and the selection probability includes: selecting the acquisition function corresponding to the selection probability having the smallest difference with the random number.

[0026] In some possible implementations, the method further includes:

[0027] For each acquisition function, set the corresponding information entropy;

[0028] After each selection of the acquisition function, the information entropy is updated according to the selection result of the acquisition function, and the selection probability is updated according to the updated information entropy.

[0029] The main advantages of the technical solution of the present invention are as follows:

[0030] The black-box adversarial sample generation method based on Bayesian evolutionary optimization of the present invention adopts Gaussian process to approximate the replacement of neural network. In the optimization process of adversarial samples, the predicted output of Gaussian process is used to replace the output of neural network, which can significantly reduce the number of neural network queries required to generate adversarial samples and reduce the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0032] Figure 1 This is a flowchart of a black-box adversarial sample generation method based on Bayesian evolutionary optimization according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0034] The technical solution provided by the embodiments of the present invention is described in detail below with reference to the accompanying drawings.

[0035] refer to Figure 1 An embodiment of the present invention provides a black-box adversarial sample generation method based on Bayesian evolutionary optimization. The generated adversarial sample can be used to conduct adversarial attacks on a neural network classifier. The method includes the following steps S1-S7:

[0036] Step S1, determining the neural network classifier to be attacked.

[0037] Specifically, according to actual needs, the neural network classifier to be attacked is determined.

[0038] Step S2, constructing a Gaussian process and acquisition function.

[0039] The Gaussian process GP is an extension of the multivariate Gaussian distribution to infinite dimensions, and is a distribution of functions. The Gaussian process consists of a mean function and a covariance function, and is a probability model. Based on a given set of data, the Gaussian process can be used to predict the function value of an unobserved point. To this end, in one embodiment of the present invention, a neural network classifier to be attacked is constructed by constructing a Gaussian process agent.

[0040] Furthermore, the acquisition function includes: any one of a PI (probability of improvement) function, an EI (expected improvement) function, an LCB (lower confidence bound) function, and a UC (Upper Confidence Bound) function.

[0041] The PI function is used to evaluate the probability that the objective function value will exceed the currently known optimal value when evaluating at a candidate point. A high PI function point means that there is a higher probability of finding a better solution than the current optimal solution when evaluating at that point. Its mathematical expression is:

[0042] PI(x)=P(f(x)≥f(x + )+∈);

[0043] Among them, f(x) is the objective function, x is the candidate point, and x +is the currently known optimal solution, and ∈ is a small positive threshold.

[0044] The EI function not only considers the probability of exceeding the current optimal value, but also the degree of exceeding. It measures the expected improvement when evaluating at a candidate point. A high point in the EI function means that evaluating at this point may bring greater improvement. Its mathematical expression is:

[0045] EI(x) = E[max(0, f(x)-f(x) + ))];

[0046] Where E represents the expected value, f(x) is the objective function, x is the candidate point, and x + This is the best solution known so far.

[0047] The LCB function is a collection function based on confidence intervals. It encourages exploration of areas with high uncertainty. A point with a high LCB function means that the point may be underestimated and is worth further exploration. Its mathematical expression is:

[0048] LCB(x)=u(x)-κ·σ(x);

[0049] Among them, u(x) is the posterior mean of the objective function at point x, σ(x) is the posterior standard deviation, and κ is a hyperparameter set to control the trade-off between exploration and utilization.

[0050] The UC function, also known as UCB, is another acquisition function based on confidence intervals. It encourages the selection of points with high expected values ​​and high uncertainties. Its mathematical expression is:

[0051] UCB(x)=u(x)+κ·σ(x);

[0052] Among them, u(x) is the posterior mean of the objective function at point x, σ(x) is the posterior standard deviation, and κ is a hyperparameter set to control the trade-off between exploration and utilization.

[0053] Step S3, obtaining the original image and its corresponding category output by the neural network classifier, and setting the optimization variables and their initial values, where the optimization variables are the weights of the pixel differences between the original image and the images in its neighborhood.

[0054] In one embodiment of the present invention, an original image is obtained according to actual conditions, and the original image is input into a neural network classifier to be attacked to obtain a category corresponding to the original image output by the neural network classifier.

[0055] Specifically, assuming that x represents the original image and the original image x is used as the input of the neural network classifier, a multidimensional output vector M(x, θ) corresponding to the original image x output by the neural network classifier can be obtained. The category corresponding to the maximum value in the multidimensional output vector M(x, θ) is the category y corresponding to the original image x.

[0056] Furthermore, in one embodiment of the present invention, in order to generate adversarial samples, the weight of the pixel difference between the original image and the image in its neighborhood is set as an optimization variable, and the initial value of the optimization variable is set.

[0057] In the embodiment of the present invention, the initial value of the optimization variable is set according to the actual situation, for example, a random number within [-1, 1] is used.

[0058] In an embodiment of the present invention, the neighborhood of the original image is randomly selected from an image data set from which the original image is obtained.

[0059] Step S4, obtaining a training data set, the training data including sample images and their corresponding categories output by the neural network classifier.

[0060] In one embodiment of the present invention, a training data set is obtained through existing known image data.

[0061] Specifically, multiple sample images x are obtained from the existing image data. i , respectively, the sample image x i Input the neural network classifier and get the sample image x output by the neural network classifier i The corresponding category y i , for each sample image x i and its corresponding category y i As a training data (x i ,y i ), and obtain a training data set D including multiple training data = {(x1, y1), ..., (x i ,y i ),...,(x n ,y n )}.

[0062] Step S5, using the training data set to train the Gaussian process to approximately replace the neural network classifier.

[0063] In one embodiment of the present invention, a training data set is used to train and estimate the mean function and covariance function of a Gaussian process, so as to approximately replace a neural network classifier with a Gaussian process.

[0064] Gaussian process is a probability model used to model functions. A Gaussian process can be defined as an infinite-dimensional multivariate Gaussian distribution, and each finite-dimensional subset is a Gaussian distribution. A Gaussian process is defined by a mean function and a covariance function (kernel function). The Gaussian process can be specifically expressed as:

[0065] f(x)~GP(m(x),k(x,x′));

[0066] Among them, f(x) represents a potential function, m(x) represents the mean function, and k(x, x′) represents the covariance function, which is used to measure the correlation between the function values ​​corresponding to the input x and x′.

[0067] Furthermore, in an embodiment of the present invention, the Gaussian process is trained in the following manner:

[0068] Step 501, initially selecting a mean function and a covariance function to simulate the nonlinear characteristics of the neural network;

[0069] Step 502, calculating the prior distribution on the training data points according to the selected mean function and covariance function;

[0070] Step 503, updating the prior distribution using the training data set through Bayesian reasoning to obtain the posterior distribution;

[0071] Step 504: Based on the obtained posterior distribution, the mean function and the covariance function are optimized and updated using the maximum likelihood estimation (MLE) or the maximum a posteriori probability estimation (MAP).

[0072] In the embodiment of the present invention, considering that the covariance already contains sufficient information, the mean function is set to 0.

[0073] Step S6, optimizing the optimization variables using a Gaussian process, and generating candidate adversarial samples based on the optimization variables and the original image.

[0074] In one embodiment of the present invention, Gaussian process is used to optimize the optimization variables using differential evolution algorithm.

[0075] It should be noted that the difference algorithm is a conventional algorithm in this field and will not be described in detail here.

[0076] Furthermore, in one embodiment of the present invention, generating candidate adversarial samples based on the optimized variables and the original image includes the following steps:

[0077] Calculate the pixel difference between the original image and its neighboring image;

[0078] Calculate the product of each pixel difference and the corresponding optimization variable respectively and superimpose them;

[0079] The superposition results are processed using the sign function to determine the disturbance direction;

[0080] According to the determined disturbance direction, a preset disturbance value is added to the original image to obtain a candidate adversarial sample.

[0081] In the embodiment of the present invention, the sign function is the sgn function.

[0082] In the embodiment of the present invention, the interference value is set according to the actual situation, for example, it is set to a number in the range of [-1 to 1], such as -0.1, -0.05, 0.05, or 0.1.

[0083] Step S7, using the candidate adversarial sample as the input of the neural network classifier, in response to the output category being wrong, outputting the candidate adversarial sample as the adversarial sample; in response to the output category being correct, using the acquisition function to sample image data, adding the sampled image and its corresponding category output by the neural network classifier as training data to the training data set, and returning to step S5.

[0084] Specifically, the candidate adversarial sample obtained in step S6 is input into the neural network classifier to obtain the category corresponding to the candidate adversarial sample output by the neural network classifier, and it is determined whether the category corresponding to the candidate adversarial sample is different from the category corresponding to the corresponding original image. If so, the category corresponding to the output candidate adversarial sample is determined to be wrong, and the candidate adversarial sample is output as an adversarial sample; if not, the category corresponding to the output candidate adversarial sample is determined to be correct, and the acquisition function is used to sample the image data, and the sampled image and its corresponding category output by the neural network classifier are added to the training data set as training data, and the step S5 is returned to use the updated training data set to update the Gaussian process training until the adversarial sample is obtained.

[0085] Furthermore, in one embodiment of the present invention, using the acquisition function to sample image data specifically includes:

[0086] Calculate the acquisition function value corresponding to each image data, and select the image data corresponding to the lowest acquisition function value as the sampled image.

[0087] Furthermore, in order to reduce the approximation error of the Gaussian process, improve the prediction accuracy of the Gaussian process, and generate adversarial samples within a smaller number of queries, the method provided by an embodiment of the present invention further includes:

[0088] For each acquisition function, set the corresponding selection probability;

[0089] Before using the acquisition function to sample image data, a random number is generated, and the acquisition function is selected based on the relationship between the random number and the selection probability.

[0090] Specifically, in one embodiment of the present invention, for the PI function, EI function, LCB function, and UC function given above, the selection probability corresponding to each acquisition function is set. The selection probability can be a numerical range or a numerical value, which is specifically set according to actual needs.

[0091] When the selection probability is a numerical range, selecting the acquisition function according to the relationship between the random number and the selection probability further includes:

[0092] Select the acquisition function corresponding to the selection probability to which the random number belongs.

[0093] When the selection probability is a numerical value, selecting the acquisition function according to the relationship between the random number and the selection probability further includes:

[0094] Select the acquisition function that corresponds to the smallest selection probability of the difference with the random number.

[0095] Furthermore, in order to further improve the prediction accuracy of the Gaussian process, the method provided by an embodiment of the present invention further includes:

[0096] For each acquisition function, set the corresponding information entropy;

[0097] After each selection of the acquisition function, the information entropy is updated according to the selection result of the acquisition function, and the selection probability is updated according to the updated information entropy.

[0098] In the embodiment of the present invention, the information entropy of the acquisition function is used to represent the probability of the acquisition function being selected.

[0099] A black-box adversarial sample generation method based on Bayesian evolutionary optimization provided by an embodiment of the present invention adopts a Gaussian process to approximate a neural network. During the optimization process of the adversarial sample, the predicted output of the Gaussian process is used to replace the output of the neural network, which can significantly reduce the number of neural network queries required to generate adversarial samples and reduce the consumption of computing resources.

[0100] Moreover, on the basis of using Gaussian process to approximate the replacement of neural network, by setting the selection probability and information entropy to select the acquisition function and iteratively update the Gaussian process, the prediction accuracy of the Gaussian process can be improved, the number of neural network queries required can be further reduced, and the attack effect of adversarial samples can be improved.

[0101] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In addition, "front", "back", "left", "right", "upper" and "lower" in this article are all referenced to the placement state shown in the accompanying drawings.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A black-box adversarial sample generation method based on Bayesian evolutionary optimization, characterized in that: include: Determine the neural network classifier to be attacked; Construct Gaussian processes and acquisition functions; Obtain the original image and its corresponding category output by the neural network classifier, and set the optimization variable and its initial value. The optimization variable is the weight of the pixel difference between the original image and the image in its neighborhood. Obtaining a training data set, the training data including sample images and their corresponding categories output by a neural network classifier; Using the training data set to train the Gaussian process to approximately replace the neural network classifier; Use Gaussian process to optimize the optimization variables, and generate candidate adversarial samples based on the optimization variables and the original image; Taking the candidate adversarial sample as the input of the neural network classifier, and in response to the output category being wrong, outputting the candidate adversarial sample as the adversarial sample; In response to the output category being correct, the acquisition function is used to sample image data, and the sampled image and its corresponding category output by the neural network classifier are added to the training data set as training data. The Gaussian process training, optimization variables optimization and candidate adversarial sample generation process are re-performed until the adversarial sample is obtained.

2. The black-box adversarial sample generation method based on Bayesian evolutionary optimization according to claim 1 is characterized in that: The acquisition function includes: any one of the PI function, the EI function, the LCB function, and the UC function.

3. The black-box adversarial sample generation method based on Bayesian evolutionary optimization according to claim 1 is characterized in that: Candidate adversarial examples are generated based on the optimized variables and the original image in the following way: Calculate the pixel difference between the original image and its neighboring image; Calculate the product of each pixel difference and the corresponding optimization variable respectively and superimpose them; The superposition results are processed using the sign function to determine the disturbance direction; According to the determined disturbance direction, a preset disturbance value is added to the original image to obtain a candidate adversarial sample.

4. The black-box adversarial sample generation method based on Bayesian evolutionary optimization according to claim 2 is characterized in that: The method further comprises: For each acquisition function, set the corresponding selection probability; Before using the acquisition function to sample image data, a random number is generated, and the acquisition function is selected based on the relationship between the random number and the selection probability.

5. The black-box adversarial sample generation method based on Bayesian evolutionary optimization according to claim 4 is characterized in that: When the selection probability is a numerical range, selecting the acquisition function according to the relationship between the random number and the selection probability includes: selecting the acquisition function corresponding to the selection probability to which the random number belongs; When the selection probability is a numerical value, selecting the acquisition function according to the relationship between the random number and the selection probability includes: selecting the acquisition function corresponding to the selection probability having the smallest difference with the random number.

6. The black-box adversarial sample generation method based on Bayesian evolutionary optimization according to claim 4 is characterized in that: The method further comprises: For each acquisition function, set the corresponding information entropy; After each selection of the acquisition function, the information entropy is updated according to the selection result of the acquisition function, and the selection probability is updated according to the updated information entropy.

Citation Information

Cited By

  • Delay differential equation modeling method based on Bayesian optimization and neural network

    CN121920244A