An Adversarial Sample Detection Method and Device Based on Neuron Activation Map
Through the method based on neuron activation graph, a deep model is constructed and a noise sample data set is generated, the weighting of the graph is calculated and input into the detector, which solves the problem that deep learning models are vulnerable to adversarial sample attacks, and effectively detects diversity adversarial samples, improving the security and robustness of the model.
Patent Information
- Application Number
- CN202210976888.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-15
AI Technical Summary
Deep learning models are vulnerable to attacks from adversarial samples, resulting in model error recognition and affecting the safety and robustness of fields such as autonomous driving systems. Existing adversarial sample detection methods rely on prior knowledge of adversarial attacks and are difficult to transfer to unknown attack methods.
A method of detecting adversarial samples based on neuron activation graph is proposed. By constructing a depth model, generating a noise sample data set, calculating the weighting of the graph and inputting it into the detector, the detection of adversarial samples is realized. This method uses graph structure to explain data correlation, extracts spatial characteristics of interlayer activation neurons, and can detect diversity adversarial samples.
Effective detection of adversarial samples is realized, the security and robustness of deep learning models are improved, and the problem of low detection tolerance rate that only depends on specific attack methods is not high.
Smart Images

Figure QLYQS_2 
Figure QLYQS_3 
Figure QLYQS_6
Abstract
Description
Technical Field
[0001] This patent relates to the fields of artificial intelligence and its security, and image classification. Specifically, it relates to a method and device for detecting adversarial samples based on neuron activation maps. Background Art
[0002] With the improvement of hardware computing power, the support of large - data storage, and the improvement of theoretical frameworks, deep - learning technology has been applied to numerous fields by virtue of its powerful feature - extraction ability and fitting ability, including the fields of computer vision, natural language processing, bioinformatics, and so on. At the same time, deep - learning technology has gradually moved from the laboratory to industrialization, with autonomous driving applications being the most prominent. In an autonomous driving system, road - sign recognition, license - plate recognition, pedestrian recognition, road recognition, obstacle detection, etc. all involve computer - vision technology, while voice - command control involves speech - recognition technology. As deep - learning technology is further widely applied, the problems therein have gradually emerged.
[0003] As early as 2014, researchers found that deep models are vulnerable to adversarial samples, that is, adversarial attacks. Specifically, a trained deep model has a good recognition accuracy for benign samples in the test set. However, after adding tiny and carefully designed adversarial perturbations to the benign samples that could originally be correctly recognized, the resulting adversarial samples will be misrecognized by the deep model. Adversarial attacks expose the vulnerabilities in deep models, and such vulnerabilities will hinder the further development of deep - learning technology. Taking the autonomous driving system as an example, adversarial attacks will have a fatal impact on its safety. For example, if some small stickers are stuck on a "STOP" road sign, the road - sign recognition model in the autonomous driving system will recognize "STOP" as a speed limit of "40", which is very dangerous for drivers and pedestrians.
[0004] According to whether the attacker knows the internal details of the deep model, adversarial attacks can be divided into white - box attacks and black - box attacks; according to whether the attacker sets an attack target, adversarial attacks can be divided into targeted attacks and untargeted attacks; according to the scenario where the attack occurs, adversarial attacks can be divided into electronic - countermeasure attacks and physical - countermeasure attacks. The ultimate goal of studying adversarial attacks is to discover the vulnerabilities in deep models and improve the security and robustness of the models. Some existing methods for detecting adversarial samples mostly rely on prior knowledge of adversarial attacks, that is, detection based on attacks, such as adversarial - sample detection based on PGD adversarial attacks. Moreover, their detection performance is positively correlated with the number and diversity of adversarial examples. Therefore, they usually perform well on certain specific attacks but are difficult to transfer to unknown attack methods.
[0005] Many big data are presented in the form of large-scale graphs or networks. Many non-graph-structured big data are often converted into graph models for analysis. The graph data structure well expresses the correlation between data. Some past work has tried to understand and explain the internal mechanism of deep neural networks. One of the ways to achieve this goal includes representing the neural network as a graph structure and studying selected graph attributes such as clustering coefficient, path length, and modularity, etc. Some research work in recent years has also shown that some metrics of the graph have strong descriptive ability for the interpretable aspects of the model.
[0006] Based on the above considerations, this patent proposes an adversarial sample detection method based on neuron activation graphs to detect adversarial sample images. Summary of the Invention
[0007] The purpose of the present invention is to provide an adversarial sample detection method and device based on neuron activation graphs in view of the deficiencies of the prior art.
[0008] The purpose of the present invention is achieved through the following technical solutions: An adversarial sample detection method based on neuron activation graphs includes the following steps:
[0009] (1) Obtain an image data set and its corresponding set of true class labels; construct a deep model f, train the deep model f using the image data set and its corresponding set of true class labels, and the deep model f saves a checkpoint file; then generate a noise sample data set using the image data set, and construct a binary data set X from the image data set and the noise sample data set B and its corresponding set of true class labels Y B ;
[0010] (2) Read the checkpoint file saved by the deep model f, and generate a mapping weight set W corresponding to the binary data set X B ; B ;
[0011] (3) Map the mapping weight set W generated in step (2) B to the graph set G B ;
[0012] (4) Calculate the corresponding weighted degree set D of the graph set G B ; B ;
[0013] (5) Construct a detector, which consists of two fully connected layers and one classification layer. Each fully connected layer consists of the output of the fully connected layer plus a relu activation function layer, and the classification layer consists of two neurons;
[0014] Train the detector, and set the hyperparameters for training the detector: the number of training epochs of the detector is Epoch2 , the number of batch processing samples is M 2 , the gradient update rule is stochastic gradient descent; the training loss function of the detector is
[0015]
[0016] where D B,n represents the nth weighted degree of each round of batch processing, n = 1, 2…, n, …M 2 , D B,n is selected from the weighted degree set D B ; Y B,n represents the true class label of the weighted degree D B,n ; p(D B,n ) represents the predicted class label of D B,n ;
[0017] Until the training loss function converges, save and output the trained adversarial sample detector;
[0018] (7) Obtain the weighted degree of the image sample to be measured, and input the weighted degree of the image sample to be measured into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be measured. If the predicted class label of the image sample to be measured is 0, the image sample to be measured is a clean sample; if the predicted class label of the image sample to be measured is 1, the image sample to be measured is an adversarial sample.
[0019] Furthermore, the step (1) specifically includes the following sub-steps:
[0020] (1.1) Obtain an image dataset X containing a total of N sam image samples. The image dataset X is where x i is the i-th image sample of the image dataset; the true class label set Y of the image dataset X is Each image sample x i has a corresponding true class label y i , indicating that the image sample x i belongs to the y i th class, where y i = {0, 1, 2..., n - 1};
[0021] (1.2) Construct a deep model f: Use the Lenet-5 model as the deep model f; the Lenet-5 model consists of three convolutional layers and two fully connected layers, where the output of each convolutional layer consists of a batch processing layer and a relu activation layer;
[0022] (1.3) Train the deep model f: Set the hyperparameters of the training: the number of training epochs of the deep model f is Epoch 1 , the number of batch processing samples is M1 , the gradient update rule is stochastic gradient descent; the training loss function of the deep model f is:
[0023]
[0024] where x h represents the h-th image sample in each round of batch processing, h = 1, 2…, h,…M 1 , x h is selected from the image dataset X; y h represents the true class label of the image sample x h ; p(x h ) represents the predicted class label of the image sample x h ; ω represents the weight parameters of the deep model f;
[0025] After the training is completed, the deep model f saves the checkpoint file;
[0026] (1.4) Generate the noisy sample dataset X ** :
[0027] (1.4.1) Test whether the predicted class label of the image sample x i in the test image dataset X is consistent with the true class label. If not, then remove the image sample x i without adding noise. If consistent, then proceed to step (1.4.2);
[0028] (1.4.2) Randomly select an image sample x j from the image dataset X with a true class label of 0, where x j ≠x i ; Test whether the predicted class label of the image sample x j in the deep model f is consistent with the true class label. If not, then remove the image sample x j without adding noise. If consistent, then proceed to step (1.4.3);
[0029] (1.4.3) Perform high-pass filtering on the image sample x j , extract the high-frequency noise contained in the image sample x j . After normalizing the noise to [-5, 5], add it to the image sample x i to obtain the noisy sample with added noise
[0030] (1.4.4) Randomly select image samples with class labels of 1, 2,…, m - 2 or m - 1 from the image dataset X, and repeat steps (1.4.2) and (1.4.3) to obtain the noisy sample set of the image sample x i
[0031] (1.4.5) Input the noise samples in the noise sample set of the image sample x i one by one into the deep model f for testing. If the predicted class label of the noise sample in the deep model f is inconsistent with the true class label, it indicates that the noise sample is a successful noise sample, and then the noise sample is retained in the noise sample set; if they are consistent, it indicates that the attack on the noise sample fails, and then the noise sample is removed from the noise sample set;
[0032] (1.4.6) Repeat steps (1.4.1)-(1.4.5) for all image samples in the image data set X to obtain the noise sample sets of all image samples, and the noise samples in the noise sample sets of all samples constitute the noise sample set X * ;
[0033] (1.4.7) Randomly select N(1.4.7) Randomly select N * noise samples from the noise sample set X sam to form the noise sample data set X ** ;
[0034] (1.5) Construct a binary data set X ** from the image data set X and the noise sample data set X B , and the binary data set X B is X B ={X, X **}; and set the true class label of all noise samples in the noise sample data set X ** to 1; set the true class label of all image samples in the image data set X to 0 to obtain the true class label set Y B corresponding to the binary data set X B , and the true class label set Y B is Y B ={0, 1}.
[0035] Furthermore, the step (2) specifically includes the following sub-steps:
[0036] (2.1) Read the checkpoint of the deep model f: Read the checkpoint file saved by the deep model f generated in step (1) and save it to the matrix θ, where θ contains all the parameters of the model, and K c represents the kernel of the c-th layer of the model. For the convolutional layer or the fully connected layer, there are f c neurons; the number of model layers is deep_n, and the number of nodes is node_n;
[0037] (2.2) Take any sample X B from the binary data set X a and input it into the deep model f to obtain the outputs of each layer of the deep model f:
[0038] [O 1 ,O 2 ,...,O c ,...,O deep_n =f(X a ;ω);
[0039] Among them, O c represents the output of the c-th layer of the deep model f, where c = 1, 2…c,…deep_n;
[0040] (2.3) The output O of the c-th layer of the deep model f c corresponds to the weight coefficient ω c,c+1 , O c has an output dimension of 1*dim c ; where ω c,c+1 represents the weight connection between the c-th layer and the (c + 1)-th layer of the deep model f. If it is connected between a fully connected layer and a fully connected layer or between a fully connected layer and a convolutional layer, ω c,c+1 has a dimension of dim c+1 *dim c ; if it is connected between a convolutional layer and a convolutional layer or between a convolutional layer and a fully connected layer, ω c,c+1 has a dimension of dim c *dim c+1 *h*w, where h*w represents the convolutional kernel dimension of the convolutional layer; after expanding the dimension of ω c,c+1 and multiplying it by the weight coefficient, a new mapping weight is obtained:
[0041] ω c ′ ,c+1 =ω c,c+1 *expand(O c ,axis,dim c );
[0042] Among them, expand() represents the dimension expansion operation, axis represents the dimension expansion direction, and dim c-1 represents the number of dimensions to be expanded; expand(O c ,axis,dim c ) is the weight coefficient after expanding the dimension of ω c,c+1 ;
[0043] Repeat the above steps to obtain the mapping weight W of the sample X a ={ω′ a ,ω′ 1,2 ,...,ω′ 2,3 ,...,ω′ c,c+1 ,...,ω′ deep_n-1,deep_n};
[0044] (2.4) Repeat steps (2.2) and (2.3) for all samples in the binary dataset X B to obtain the mapping weight set W B ={W, W **}={W 1 , W 2 , …, W a , …}; where W is the mapping weight set of the image dataset X, and W ** is the mapping weight set of the noise sample dataset X ** .
[0045] Furthermore, step (3) specifically includes the following sub-steps:
[0046] (3.1) Initialize the graph G a :
[0047] The graph G a is an undirected weighted graph. Map the filters or neurons of the K c -th layer of the deep model f to nodes , that is, create f c nodes in the K c -th layer and represent them as Set the node attributes of the graph to be empty, and the weighted edges between nodes in adjacent layers are , that is, initialize the graph as
[0048] (3.2) Map the mapping weight W a to the graph G a ;
[0049] Take any mapping weight W B from the mapping weight set W a and map it to the graph G a , specifically expressed by the following formula:
[0050]
[0051] where ||·|| is the norm formula, is the b-th filter of the c-th layer kernel, is the d-th filter of the next layer kernel, is the norm of the parameters between the two filters; is the weighted edge between the b-th node of the c-th layer kernel and the d-th node of the c+1-th layer kernel; repeat the above steps until each layer parameter of the mapping weight W a is mapped to the graph G a ;
[0052] (3.3) Map the mapping weight set W BRepeat steps (3.1) and (3.2) for all mapping weights in, and obtain the graph set G B ={G, G **}={G 1 , G 2 , …, G a , …}; where G is the graph set of the image dataset X, and G ** is the graph set of the noise sample dataset X ** .
[0053] Furthermore, step (4) specifically includes the following sub-steps:
[0054] (4.1) Take any graph G B from the graph set G a , and calculate the weighted degree D a of the graph G a . The calculation formula is as follows:
[0055]
[0056] Until the weighted degree D a of each node is calculated. The weighted degree D a is a one-dimensional vector, and its length is the number of nodes node_n of the graph;
[0057] (4.2) Repeat step (4.1) for all graphs in the graph set G B , and obtain the weighted degree set D B ={D, D **}={D 1 , D 2 , …, D a , …}; where D is the weighted degree set of the image dataset X, and D ** is the weighted degree set of the noise sample dataset X ** .
[0058] Furthermore, step (7) specifically includes the following sub-steps:
[0059] (7.1) Input the image sample to be tested into the deep model f, and repeat steps (2.2) and (2.3) to obtain the mapping weights of the image sample to be tested;
[0060] (7.2) Repeat steps (3.1) and (3.2) to map the mapping weights of the image sample to be tested to the graph of the image sample to be tested;
[0061] (7.3) Repeat step (4.1) to calculate the weighted degree of the graph of the image sample to be tested;
[0062] (7.4) Subsequently, the weighted degree of the image sample to be tested is input into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be tested. If the predicted class label of the image sample to be tested is 0, then the image sample to be tested is a clean sample; if the predicted class label of the image sample to be tested is 1, then the image sample to be tested is an adversarial sample.
[0063] The present invention also provides an adversarial sample detection device based on a neuron activation map, including one or more processors for implementing the above-mentioned adversarial sample detection method based on a neuron activation map.
[0064] The present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned adversarial sample detection method based on a neuron activation map.
[0065] The beneficial effects of the present invention are:
[0066] (1) The adversarial sample detection method based on a neuron mapping graph proposed in this patent utilizes the ability of a graph to explain data correlation, associates model weights and neuron activations, and maps them to the graph level, further extracts the node attributes of the graph, and realizes the detection of adversarial samples.
[0067] (2) The adversarial sample detection method proposed in this patent is not based on certain specific adversarial attacks, but considers the common characteristics of adversarial attacks, and only uses normal samples to generate noise samples. The noise of benign samples and other category samples is superimposed, and the noise samples are used to simulate adversarial samples, which has the diversity of adversarial samples. Description of the Drawings
[0068] Figure 1 It is a schematic flow chart of an adversarial sample detection method based on a neuron activation map;
[0069] Figure 2 It is an overall framework diagram of an adversarial sample detection method based on a neuron activation map;
[0070] Figure 3 It is a schematic process diagram of generating a noise sample data set;
[0071] Figure 4 It is a schematic diagram of an adversarial sample detection device based on a neuron activation map. Detailed Embodiments
[0072] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] In the present invention, a clean sample refers to a sample from a public dataset without artificially added noise. An adversarial sample refers to a sample with artificially added specific noise that does not affect the observation of the picture but can cause a classification error in the DNN. A noise sample refers to a sample generated by the method of generating noise samples in the present invention to simulate an adversarial sample.
[0074] In view of the risk that the deep model may be misclassified by potential adversarial samples. The present invention proposes an adversarial sample detection method based on neuron activation maps. Previous research work has mentioned that graph structures can well express the correlation between data. In the present invention, the samples are activated in each layer of the model to form a graph, and the graph node features are extracted, which can better extract the spatial features of the activated neurons between layers. Therefore, this method has a better adversarial detection effect. As Figure 2 shown, the specific technical concept is as follows: 1. Use a general method of generating noise samples to generate noise samples, use the noise samples to simulate adversarial samples, and form a binary dataset with normal samples; 2. Input the samples into the deep model to obtain the output of each layer of the model; 3. Reflect the output of each layer of the model back to the weights of the model to obtain the mapped weights; 4. Map the mapped weights into a graph, extract the weighted degree on the graph, and then use the weighted degree as a feature to input into the detector to achieve the detection of adversarial samples. Among them, the noise sample set is generated by simulating adversarial samples with normal samples. It has diversity and the aggressiveness of adversarial samples, and also avoids the problem of low detection fault tolerance caused by only using adversarial samples of a certain specific attack. As Figure 3 shown, in order to generate effective noise samples, the high-frequency noise of each type of normal sample is extracted using a high-pass filter. The high-frequency noise will contain class features and noise features, so it is more directional and general than Gaussian noise. Superimpose the high-frequency noise of each class with the normal samples of other classes to obtain noise samples. Since the noise samples and normal samples need to form a binary classifier, and the number of noise samples is more than that of normal samples, balance the dataset for the noise samples, that is, randomly select a part of the dataset to have the same number as the normal sample set.
[0075] The deep model structure of the embodiment of the present invention uses VGG16, the dataset uses Cifar10, and the detector is composed of two fully connected networks.
[0076] Embodiment 1
[0077] As Figure 1 and Figure 2 shown, the present invention provides an adversarial sample detection method based on a neuron activation map, comprising the following steps:
[0078] (1) Obtain an image data set and its corresponding set of true labels; construct a deep model f, train the deep model f using the image data set and its corresponding set of true labels, and save a checkpoint file for the deep model f; then generate a noise sample data set using the image data set, and construct a binary data set X from the image data set and the noise sample data set B and its corresponding set of true labels Y B .
[0079] The step (1) specifically includes the following sub-steps:
[0080] (1.1) In this embodiment, obtain an image data set X containing a total of N sam = 42000 image samples from the Cifar10 data set; the image data set X is X = {x 1 , x 2 ,.., x i ,..., x 42000}, where x i is the i-th image sample of the image data set, and all image samples in the image data set X are clean samples; the Cifar10 data set has a total of m = 10 classes; the set of true labels Y of the image data set X is Y = {y 1 , y 2 ,..., y i ,..., y 42000}, and each image sample x i has a corresponding true label y i , indicating that the image sample x i belongs to the y i -th class, where y i = {0, 1, 2..., 9};
[0081] The Cifar10 data set is a commonly used data set for image classification models, and the size of each image is 3 * 32 * 32.
[0082] (1.2) Construct a deep model f: use the Lenet-5 model as the deep model f; the Lenet-5 model consists of three convolutional layers and two fully connected layers, and the output of each convolutional layer consists of a batch layer and a relu activation layer;
[0083] (1.3) Train the deep model f: set the hyperparameters for training: the number of training epochs for the deep model f is Epoch 1, the number of batch processing samples is M 1 , the gradient update rule is stochastic gradient descent; the training loss function of the deep model f is:
[0084]
[0085] where, x h represents the h-th image sample of each round of batch processing, h = 1, 2…, h,…M 1 , x h is selected from the image dataset X; y h represents the true class label of the image sample x h ; p(x h ) represents the predicted class label of the image sample x h ; ω represents the weight parameter of the deep model f;
[0086] After the training is completed, the deep model f saves the checkpoint file.
[0087] (1.4) Generate the noise sample dataset X ** , as Figure 3 shown:
[0088] (1.4.1) Test whether the predicted class label and the true class label of the image sample x i in the deep model f are consistent. If they are not consistent, then the image sample x i is not added with noise. If they are consistent, then go to step (1.4.2);
[0089] (1.4.2) Randomly take the image sample x j with the true class label of 0 from the image dataset X, where, x j ≠x i ; Test whether the predicted class label and the true class label of the image sample x j in the deep model f are consistent. If they are not consistent, then the image sample x j is not added with noise. If they are consistent, then go to step (1.4.3);
[0090] (1.4.3) Perform high-pass filtering on the image sample x j , extract the high-frequency noise contained in the image sample x j . After normalizing the noise to [-5, 5], it is superimposed on the image sample x i to obtain the noise sample with added noise
[0091] (1.4.4) Randomly take the image samples with true class labels of 1, 2, 3, 4, 5, 6, 7, 8, or 9 from the image dataset X, and repeat steps (1.4.2) and (1.4.3) to obtain the image sample x iThe noise sample set
[0092] (1.4.5) Input each noise sample in the noise sample set of the image sample x i The noise sample set into the deep model f for testing. If the predicted class label of the noise sample in the deep model f is inconsistent with the true class label, it indicates that the noise sample is a successful noise sample, and then the noise sample is retained in the noise sample set; if they are consistent, it indicates that the attack on the noise sample fails, and then the noise sample is removed from the noise sample set;
[0093] (1.4.6) Repeat steps (1.4.1) - (1.4.5) for all image samples in the image data set X to obtain the noise sample set of all image samples, and the noise samples in the noise sample sets of all samples constitute the noise sample set X * ;
[0094] (1.4.7) Randomly select N * = 42000 noise samples from the noise sample set X sam to form the noise sample data set X ** ;
[0095] All image samples in the noise sample data set X ** are noise samples;
[0096] Select N sam = 42000 noise samples to form the noise sample data set X ** for the purpose of balancing the data;
[0097] (1.5) Construct a binary data set X ** from the image data set X and the noise sample data set X B , and the binary data set X B is X B = {X, X **}; And set the true class label of all noise samples in the noise sample data set X ** to 1. Samples with a true class label of 1 are noise samples; set the true class label of all image samples in the image data set X to 0. Samples with a true class label of 0 are clean samples; obtain the true class label set Y B corresponding to the binary data set X B , and the true class label set Y B is Y B = {0, 1}.
[0098] (2) Read the checkpoint file saved by the deep model f generated in step (1), and generate the mapping weight set W B corresponding to the binary data set XB 。
[0099] Step (2) specifically includes the following sub-steps:
[0100] (2.1) Read the checkpoint file saved by the deep model f and save it to the matrix θ, where θ contains all the parameters of the model, and K c represents the kernel of the c-th layer of the deep model f. For the convolutional layer or the fully connected layer, there are f c neurons; the number of model layers is deep_n, and the number of nodes is node_n;
[0101] In this embodiment, deep_n is 5 and node_n is 239.
[0102] (2.2) Take any sample X B from the binary dataset X a and input it into the deep model f to obtain the outputs of each layer of the deep model f:
[0103] [O 1 , O 2 ,..., O c ,..., O deep_n = f(X a ; ω);
[0104] where O c represents the output of the c-th layer of the deep model f, c = 1, 2…c,…deep_n; the dimension of each layer's output is a one-dimensional vector, and the vector length is related to the model structure.
[0105] (2.3) The weight coefficient corresponding to the output O c of the c-th layer of the deep model f is ω c,c+1 , and the output dimension of O c is 1*dim c ; where ω c,c+1 represents the weight connection between the c-th layer and the c + 1-th layer of the deep model f. If it is connected between a fully connected layer and a fully connected layer or between a fully connected layer and a convolutional layer, the dimension of ω c,c+1 is dim c+1 *dim c ; if it is connected between a convolutional layer and a convolutional layer or between a convolutional layer and a fully connected layer, the dimension of ω c,c+1 is dim c *dim c+1 *h*w, where h*w represents the convolutional kernel dimension of the convolutional layer; after expanding the dimension of ω c,c+1 and multiplying it by the weight coefficient, a new mapping weight is obtained:
[0106] ω′ c,c+1 = ω c,c+1 *expand(Oc ,axis,dim c );
[0107] Among them, expand() represents the expansion dimension operation, axis represents the direction of the expansion dimension, dim c-1 Indicates the number of expanded dimensions; expand(O c ,axis,dim c ) is ω c,c+1 The weight coefficient after the expanded dimension;
[0108] Repeat the above steps to get sample X a The mapping weight W a ={ω′ 1,2 ,ω′ 2,3 ,...,ω′ c,c+1 ,...,ω′ deep_n-1,deep_n}.
[0109] (2.4) The binary data set X B Repeat steps (2.2) and (2.3) for all samples in to obtain the mapping weight set W B = {W,W **}={W 1 ,W 2 ,…,W a ,…}; where W is the mapping weight set of the image dataset X, W ** is the noise sample dataset X ** The set of mapping weights.
[0110] (3) The mapping weight set W generated in step (2) B Mapped to the graph set G B .
[0111] The step (3) specifically includes the following sub-steps:
[0112] (3.1) Initialize graph G a :
[0113] The figure G a is an undirected weighted graph, and the K c The filters or neurons of the layer Mapping to Node That is, in K c Layer creation c nodes and are represented as Assume that the node attributes of the graph are empty, and the weight edge between the nodes of two adjacent layers is That is, the initialization graph is
[0114] (3.2) The mapping weight Wa Map to graph G a ;
[0115] From the mapping weight set W B Take any mapping weight W a Map to graph G a , specifically expressed by the following formula:
[0116]
[0117] where ||·|| is the norm formula, is the b-th filter of the c-th layer kernel, is the d-th filter of the next layer kernel, is the norm of the parameter between the two filters; is the weighted edge between the b-th node of the c-th layer kernel and the d-th node of the (c + 1)-th layer kernel; Repeat the above steps until each layer parameter of the mapping weight W a is mapped to graph G a .
[0118] (3.3) Repeat steps (3.1) and (3.2) for all mapping weights in the mapping weight set W B to obtain the graph set G B ={G, G **}={G 1 , G 2 ,…, G a ,…}; where G is the graph set of the image data set X, and G ** is the graph set of the noise sample data set X ** .
[0119] (4) Calculate the corresponding weighted degree set D B for the graph set G B .
[0120] The said step (4) specifically includes the following sub-steps:
[0121] (4.1) Take any graph G B from the graph set G a , and calculate the weighted degree D a of the graph G a , and the calculation formula is as follows:
[0122]
[0123] until the weighted degree D a of each node is calculated. The said weighted degree D a is a one-dimensional vector, and its length is the number of nodes node_n of the graph.
[0124] (4.2) Repeat step (4.1) for all the graphs in graph set G B to obtain the weighted degree set D B ={D, D **}={D 1 , D 2 , …, D a , …}; where D is the weighted degree set of image data set X, and D ** is the weighted degree set of noise sample data set X ** .
[0125] (5) Construct a detector, which consists of two fully connected layers and one classification layer. Each fully connected layer is composed of the output of the fully connected layer plus a relu activation function layer, and the classification layer consists of two neurons.
[0126] (6) Train the detector, and set the hyperparameters for training the detector: the number of training epochs of the detector is Epoch 2 , the number of batch samples is M 2 , and the gradient update rule is stochastic gradient descent; the training loss function of the detector is
[0127]
[0128] where D B,n represents the nth weighted degree of each batch in each round, n = 1, 2, …, n, … M 2 , D B,n is selected from the weighted degree set D B ; Y B,n represents the true class label of the weighted degree D B,n ; p(D B,n ) represents the predicted class label of D B,n .
[0129] Until the training loss function converges, save and output the trained adversarial sample detector; the trained adversarial sample detector has sufficient performance.
[0130] Using this adversarial sample detector with sufficient performance, the transfer detection of different types of adversarial samples can be achieved, improving the security of the deep learning model.
[0131] (7) Obtain the weighted degree of the image sample to be tested, and input the weighted degree of the image sample to be tested into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be tested. If the predicted class label of the image sample to be tested is 0, then the image sample to be tested is a clean sample; if the predicted class label of the image sample to be tested is 1, then the image sample to be tested is an adversarial sample.
[0132] The specific steps of step (7) include the following sub-steps:
[0133] (7.1) Input the image sample to be tested into the deep model f, and repeat steps (2.2) and (2.3) to obtain the mapping weights of the image sample to be tested.
[0134] (7.2) Repeat steps (3.1) and (3.2) to map the mapping weights of the image sample to be tested to the graph of the image sample to be tested.
[0135] (7.3) Repeat step (4.1) to calculate the weighted degree of the graph of the image sample to be tested.
[0136] (7.4) Subsequently, input the weighted degree of the image sample to be tested into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be tested. If the predicted class label of the image sample to be tested is 0, the image sample to be tested is a clean sample; if the predicted class label of the image sample to be tested is 1, the image sample to be tested is an adversarial sample.
[0137] The image datasets use the MNIST handwritten dataset and the CIFAR-10 color dataset. The MNIST dataset consists of 60,000 grayscale images of 28*28*1, which contain handwritten digits from 0 to 9; the CIFAR-10 dataset consists of 60,000 color images of 32*32*3, which contain ten common types of items or organisms in the real world, such as airplanes, birds, dogs, etc. 10,000 test set images are taken from each of the two datasets for scheme verification. Four common adversarial attack methods are used to generate adversarial samples for these two test sets to verify the performance of the adversarial sample detector. The attack methods include three white-box attacks and one black-box attack. The white-box attacks include FGSM, CW, and PGD, where CW and PGD are iterative attack methods with stronger attack effects, and the CW attack has higher concealment; the black-box attack is Boundary. Multiple adversarial attack methods are used to illustrate the effectiveness of the adversarial sample detector generated by the present invention.
[0138] It can be seen from the results in Table 1 that the adversarial sample detector proposed by the present invention has good detection effects on various adversarial samples and has good performance in different datasets.
[0139] Table 1 Transfer detection effects of adversarial samples based on the Lenet-5 model
[0140] FGSM CW Boundary PGD MNIST 99.6% 99.4% 99.5% 99.5% CIFAR-10 98.9% 97.7% 98.0% 98.3%
[0141] Corresponding to the foregoing embodiments of the adversarial sample detection method based on the neuron activation graph, the present invention also provides embodiments of an adversarial sample detection device based on the neuron activation graph.
[0142] See Figure 4, An adversarial sample detection device based on a neuron activation map provided by an embodiment of the present invention includes one or more processors for implementing the adversarial sample detection method based on the neuron activation map in the above embodiment.
[0143] Embodiments of the adversarial sample detection device based on the neuron activation map of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically defined device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the adversarial sample detection device based on the neuron activation map of the present invention is located. In addition to Figure 4 the shown processor, memory, network interface, and non-volatile memory, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the device, which will not be elaborated here.
[0144] The specific implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.
[0145] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0146] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the adversarial sample detection method based on the neuron activation map in the above embodiment.
[0147] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0148] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An adversarial sample detection method based on neuron activation maps, characterized in that, it includes the following steps: (1) Obtain an image dataset and its corresponding set of true class labels; construct a deep model f, and use the image dataset and its corresponding set of true class labels to train the deep model f. The deep model f saves checkpoint files; Then, the image dataset is applied to generate a noise sample dataset, and a binary dataset X is constructed from the image dataset and the noise sample dataset B and its corresponding set of true class labels Y B ; Generate a noise sample dataset X ** : (1.4.1) The image sample x in the test image dataset X i Whether the predicted class label in the deep model f is consistent with the true class label. If not, the image sample x is removed. i Without adding noise. If consistent, proceed to step (1.4.2); (1.4.2) Randomly select an image sample x with a true class label of 0 from the image dataset X j , where x j ≠x i ; Test whether the predicted class label of the test image sample x j in the deep model f is consistent with the true class label. If not, remove the image sample x j without adding noise. If consistent, proceed to step (1.4.3); (1.4.3) For the image sample x j Perform high-pass filtering on it, and extract the high-frequency noise contained in the image sample x j After normalizing the noise to [-5, 5], superimpose it on the image sample x i to obtain the noise sample with added noise (1.4.4) Randomly select image samples with true class labels of 1, 2, …, m-2 or m-1 from the image dataset X, and repeat steps (1.4.2) and (1.4.3) to obtain the image sample x i 's noise sample set (1.4.5) The image sample x i The noise sample set The noise samples in the deep model f are input one by one for testing. If the predicted class label of the noise sample in the deep model f is inconsistent with the true class label, it indicates that the noise sample has successfully added noise, and the noise sample is retained in the noise sample set; if they are consistent, it indicates that the noise sample has failed to add noise, and the noise sample is removed from the noise sample set; (1.4.6) Repeat steps (1.4.1) - (1.4.5) for all image samples in the image dataset X to obtain a noise sample set for all image samples. The noise samples in the noise sample set for all image samples constitute the noise sample set X * ; (1.4.7) Randomly extract N * noise samples from the noise sample set X sam to form a noise sample data set X ** ; (2) Read the checkpoint file saved by the deep model f and generate the binary dataset X B The corresponding set of mapping weights W B ; (3) Map the mapping weight set W generated in step (2) B to the graph set G B ; The specific steps of step (3) include the following sub-steps: (3.1) Initialize graph G a : The graph G a is an undirected weighted graph. Map the filters or neurons of the K c th layer of the deep model f to nodes That is, create f c nodes in the K c th layer and represent them as Let the node attributes of the graph be empty, and the weight edges between the nodes of adjacent layers be That is, initialize the graph as (3.2) Map the mapping weight W a to graph G a ; From the set of mapping weights W B Take any mapping weight W a Map it to graph G a , which is specifically represented by the following formula: where, ||·|| is the norm formula, is the b-th filter of the c-th layer kernel, is the d-th filter of the next layer kernel, is the norm of the parameters between the two filters; is the weight edge between the b-th node of the c-th layer kernel and the d-th node of the (c + 1)-th layer kernel; Repeat the above steps until each layer parameter of the mapping weight W a is mapped to the graph G a ; (3.3) Repeat steps (3.1) and (3.2) for all mapping weights in the mapping weight set W B to obtain the graph set G B ={G, G **}={G 1 , G 2 , …, G a , …}; where G is the graph set of the image data set X, and G ** is the graph set of the noise sample data set X ** ; (4) Calculate the graph set G B The corresponding weighted degree set D B ; (5) Construct a detector, which consists of two fully connected layers and one classification layer. Each fully connected layer consists of the output of the fully connected layer plus a relu activation function layer, and the classification layer consists of two neurons; (6) Train the detector and set the hyperparameters for training the detector: the number of training epochs for the detector is Epoch 2 , the number of samples in a batch is M 2 , and the gradient update rule is stochastic gradient descent; the training loss function of the detector is Among them, D B,n represents the nth weighted degree of each round of batch processing, where n = 1, 2, …, n, …, M 2 , D B,n is selected from the weighted degree set D B ; Y B,n represents the true class label of the weighted degree D B,n ; p(D B,n ) represents the predicted class label of D B,n . Until the training loss function converges, save and output the trained adversarial sample detector; (7) Obtain the weighted degree of the image sample to be tested, and input the weighted degree of the image sample to be tested into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be tested. If the predicted class label of the image sample to be tested is 0, then the image sample to be tested is a clean sample; if the predicted class label of the image sample to be tested is 1, then the image sample to be tested is an adversarial sample.
2. The adversarial sample detection method based on neuron activation maps according to claim 1, characterized in that, the specific steps of step (1) include the following sub-steps: (1.1) Obtain an image dataset X containing N sam image samples, and the image dataset X is where x i is the i-th image sample in the image dataset; the true label set Y of the image dataset X is For each image sample x i there is a corresponding true label y i , indicating that the image sample x i belongs to the y i -th class, where y i = {0, 1, 2..., m - 1}; (1.2) Construct the deep model f: Use the Lenet-5 model as the deep model f; the Lenet-5 model consists of three convolutional layers and two fully connected layers, and the output of each convolutional layer consists of a batch processing layer and a relu activation layer; (1.3) Train the deep model f: Set the hyperparameters for training: The number of training epochs for the deep model f is Epoch 1 , and the number of batch samples is M 1 , and the gradient update rule is stochastic gradient descent; The training loss function of the deep model f is: where x h represents the h-th image sample in each round of batch processing, h = 1, 2, …, h, …, M 1 , x h is selected from the image dataset X; y h represents the true class label of the image sample x h ; p(x h ) represents the predicted class label of the image sample x h ; ω represents the weight parameter of the deep model f After the training ends, the deep model f saves checkpoint files; (1.4) Generate the noise sample dataset X ** ; (1.5) The image dataset X and the noise sample dataset X ** Construct a binary data set X B , the binary data set X B For X B ={X,X ** }; and the noise sample data set X ** The true class labels of all noise samples in the image dataset X are set to 1; the true class labels of all image samples in the image dataset X are set to 0, and the binary dataset X is obtained. B The corresponding true class label set Y B , the true class label set Y B Y B ={0,1}.
3. The adversarial sample detection method based on neuron activation maps according to claim 2, characterized in that, the specific steps of step (2) include the following sub-steps: (2.1) Read the checkpoint of the deep model f: Read the checkpoint file saved by the deep model f generated in step (1) and save it to the matrix θ, where θ contains all the parameters of the model, and K c represents the kernel of the c-th layer of the model. For the convolutional layer or the fully connected layer, there are f c neurons; The number of model layers is deep_n, and the number of nodes is node_n; (2.2) From the binary dataset X B Take any sample X a Input it into the deep model f to obtain the outputs of each layer of the deep model f: [O 1 ,O 2 ,...,O c ,...,O deep_n ) = f(X a ; ω); Among them, O c represents the output of the c-th layer of the deep model f, where c = 1, 2,..., c,..., deep_n; (2.3) Output O of the c-th layer of the deep model f c The corresponding weight coefficient is ω c,c+1 , O c The output dimension of O is 1 * dim c ; where ω c,c+1 represents the weight connection between the c-th layer and the (c + 1)-th layer of the deep model f. If it is connected between a fully connected layer and a fully connected layer or between a fully connected layer and a convolutional layer, the dimension of ω c,c+1 is dim c+1 * dim c ; If it is connected between a convolutional layer and a convolutional layer or between a convolutional layer and a fully connected layer, the dimension of ω c,c+1 is dim c * dim c+1 * h * w, where h * w represents the convolutional kernel dimension of the convolutional layer; After expanding the dimension of ω c,c+1 and multiplying it by the weight coefficient, a new mapping weight is obtained: ω′ c,c+1 = ω c,c+1 *expand(O c , axis, dim c ); Among them, expand() represents the operation of expanding dimensions, axis represents the direction of expanding dimensions, and dim c-1 represents the number of dimensions to be expanded; expand(O c , axis, dim c ) is ω c,c+1 the weight coefficient after expanding dimensions; Repeat the above steps to obtain sample X a The mapping weight W a ={ω′ 1,2 ,ω′ 2,3 ,...,ω′ c,c+1 ,...,ω′ deep_n-1,deep_n}; (2.4) Repeat steps (2.2) and (2.3) for all samples in the binary dataset X B to obtain the mapping weight set W B ={W, W **}={W 1 , W 2 , …, W a , …}; where W is the mapping weight set of the image dataset X, and W ** is the mapping weight set of the noise sample dataset X ** .
4. The adversarial sample detection method based on neuron activation maps according to claim 3, characterized in that, the specific steps of step (4) include the following sub-steps: (4.1) From the graph set G B Take any graph G a , calculate the weighted degree D a of graph G, and the calculation formula is as follows: a Until the weighted degree D of each node is calculated a , the weighted degree D a is a one-dimensional vector with a length equal to the number of nodes node_n in the graph; (4.2) Repeat step (4.1) for all the graphs in graph set G B to obtain the weighted degree set D B ={D, D **}={D 1 , D 2 , …, D a , …}; where D is the weighted degree set of image data set X, and D ** is the weighted degree set of noise sample data set X ** .
5. The adversarial sample detection method based on neuron activation maps according to claim 4, characterized in that, the specific steps of step (7) include the following sub-steps: (7.1) Input the image sample to be tested into the deep model f, and repeat steps (2.2) and (2.3) to obtain the mapping weights of the image sample to be tested; (7.2) Repeat steps (3.1) and (3.2) to map the mapping weights of the image sample to be tested to the graph of the image sample to be tested; (7.3) Repeat step (4.1) to calculate the weighted degree of the graph of the image sample to be tested; (7.4) Subsequently, input the weighted degree of the image sample to be tested into the trained adversarial sample detector. The adversarial sample detector outputs the predicted class label of the image sample to be tested. If the predicted class label of the image sample to be tested is 0, then the image sample to be tested is a clean sample; if the predicted class label of the image sample to be tested is 1, then the image sample to be tested is an adversarial sample.
6. An adversarial sample detection device based on neuron activation maps, characterized in that, Comprising one or more processors for implementing the adversarial sample detection method based on neuron activation maps according to any one of claims 1-5.
7. A computer-readable storage medium having a program stored thereon, wherein, when the program is executed by a processor, it is used to implement the adversarial sample detection method based on neuron activation maps according to any one of claims 1-5.