Deep neural network anti-attack method based on important neurons
Through the deep neural network adversarial attack method based on important neurons, the neuron importance score is calculated to generate a mask model, and the adversarial samples are updated through cross entropy loss and gradient, the problem of poor adversarial attack timeliness and migration in the existing technology is solved, and efficient adversarial attacks are achieved and the security of the model is improved.
Patent Information
- Application Number
- CN202510586370.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing network path-based adversarial attack methods have poor performance in terms of timeliness and migrations, and lack credible mechanism support, resulting in uncertainty in the adversarial attack effect.
A deep neural network adversarial attack method based on important neurons is adopted to generate a binary mask by calculating the importance score of the neuron, obtaining the mask model, and computing the cross entropy loss of the adversarial sample through the mask model and the source model, obtaining the first gradient and the second gradient, weighting to obtain the total gradient, and updating the iterative adversarial sample.
The migration of adversarial samples is improved, and the generated adversarial samples are strongly robust, which can better locate important features of image categories, improve the success rate of adversarial attacks, and provide trusted mechanism support for deep neural networks, improving the security of artificial intelligence models.
Smart Images

Figure CN120106233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of trusted artificial intelligence and computer vision technology, and in particular to a deep neural network counterattack method based on important neurons. Background Art
[0002] In recent years, the field of computer vision has focused on the problem of adversarial sample generation. Minor image perturbations can attack deep neural networks, which has aroused concern about the security of neural networks. Black-box transfer attacks, in which attackers do not know the target model information and attack the target model through the source model, have high transferability and are difficult to defend. They help to discover neural network defects in practical applications and guide network performance improvement. The adversarial sample phenomenon has prompted reflection on the reliability and security of artificial intelligence, and exploring neural network interpretability methods and analyzing adversarial sample mechanisms have become key research topics. As an important branch of understanding the decision-making process of neural networks, interpretable methods based on network path attribution have attracted much attention because they can intuitively display the path nodes that affect the decision. However, there is still a lack of a unified framework for attribution visualization and evaluation standards. There is still little related work on network path interpretability methods in adversarial attack problems, resulting in poor performance in balancing the timeliness and transferability of adversarial attacks.
[0003] Different from the new perspective based on network paths, mainstream black-box transfer attack schemes include data perspective and optimization perspective. The data perspective is inspired by data augmentation technology in the field of machine learning, and reduces the overfitting of adversarial samples to the source model through a series of established transformations such as scaling, cropping, and rotation. The optimization perspective relies on traditional numerical algorithms and uses a sophisticated iterative process to ensure that the algorithm seeks transferable adversarial samples, usually using a momentum-based gradient optimization attack method. Regardless of whether it is a data perspective or an optimization perspective, the corresponding adversarial attack schemes lack credible mechanism support, especially the decision-making process of the neural network is not clearly involved, which undoubtedly increases the uncertainty of the adversarial attack effect. Summary of the invention
[0004] In order to solve the technical problems that the existing network path-based counterattacks have poor timeliness and migration, the black box migration attack lacks a credible mechanism support, and the uncertainty of the counterattack effect is high, the purpose of the present invention is to provide a deep neural network counterattack method based on important neurons, and the technical scheme adopted is as follows: Prepare a data set and a source model in sequence, wherein the data set includes input samples and corresponding true labels, and set corresponding operating parameters for the source model, iterate the data set, and when the number of iterations reaches the preset operating parameters, calculate the importance scores of all neurons in the source model, generate a binary mask, and obtain a masked mask model; After the number of iterations reaches the preset running parameters, the adversarial sample is output, and the cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model and the source model respectively, and the first gradient and the second gradient of the cross entropy loss with respect to the adversarial sample are obtained accordingly, and the first gradient and the second gradient are weighted to obtain the total gradient; The adversarial samples are updated iteratively according to the total gradient until the iteration ends and the final adversarial sample is generated.
[0005] Preferably, the operating parameters include the total number of iterations, the maximum allowable distortion, the momentum factor, the mask ratio, the number of masks, the balance parameter, the step size and the update period, and the calculation formula corresponding to the update period is:
[0006] in, Indicates the update cycle; Indicates the total number of iterations; Indicates the number of masking times.
[0007] Preferably, when the number of iterations reaches a preset operating parameter, the importance scores of all neurons in the source model are calculated, a binary mask is generated, and a masked mask model is obtained, including: When the number of iterations reaches the update cycle, the importance scores of all neurons in the source model are obtained using the neuron importance calculation method; The neurons in each layer are sorted in descending order based on the importance score, and the binary mask is determined layer by layer through the mask ratio; The mask model after masking is obtained by combining the binary mask with the source model. The corresponding calculation formula is:
[0008] in, represents a mask model; Represents the source model; represents the element-wise product between a neuron and a binary mask; Represents a binary mask.
[0009] Preferably, a neuron importance calculation method is used to obtain the importance scores of all neurons in the source model, and the corresponding calculation formula is:
[0010] in, Indicates Layer The neurons are related to the input samples Importance score; Represents the input sample of the dataset Through the network output of the source model; Represents the input sample And the neuron activation response The network output when it approaches 0; Indicates Layer The activation response value of a neuron.
[0011] Preferably, the adversarial sample is output after the number of iterations reaches a preset running parameter, the cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model and the source model respectively, and the first gradient and the second gradient of the cross entropy loss with respect to the adversarial sample are obtained accordingly, and the first gradient and the second gradient are weighted to obtain the total gradient, including: When the number of iterations reaches the preset running parameters, the adversarial sample after the corresponding number of iterations is output. , and the corresponding true label ; The cross entropy loss between the adversarial sample and the corresponding true label is calculated by the mask model: ; The cross entropy loss between the adversarial sample and the corresponding true label is calculated by the source model: ; The first gradient and the second gradient are obtained in sequence through the back propagation method according to the cross entropy loss, and the total gradient is obtained by weighting the first gradient and the second gradient through the balance parameter.
[0012] Preferably, the first gradient and the second gradient are obtained in sequence through the back propagation method according to the cross entropy loss, and the corresponding calculation formula is:
[0013]
[0014] in, represents the first gradient; represents the second gradient; represents the gradient operator; The total gradient is obtained by weighting the first gradient and the second gradient by the balance parameter. The corresponding calculation formula is:
[0015] in, represents the total gradient; represents the equilibrium parameter.
[0016] Preferably, the adversarial samples are updated in sequence according to the total gradient until the iteration ends to generate a final adversarial sample, including: According to the total gradient, the adversarial sample is updated using momentum iteration to generate adversarial samples of adjacent iterations; Continue to iterate the adversarial samples of adjacent iterations until the total number of iterations is reached, the iteration ends, and the final adversarial sample is generated.
[0017] Preferably, according to the total gradient, momentum is used to iteratively update the adversarial sample to generate adversarial samples of adjacent iterations. The corresponding calculation formula is:
[0018]
[0019] in, Indicates the adjacent order, i.e. Adversarial examples of iterations; Indicates Adversarial examples of iterations; represents the truncation function used to satisfy the infinite norm constraint against noise; represents the maximum disturbance; Indicates the step size of each iteration; represents a symbolic function; Indicates Momentum in iterations; represents the momentum factor; Indicates Momentum in iterations; represents the total gradient; Represents a norm operation.
[0020] The present invention has the following beneficial effects: The adversarial attack method provided in the present application starts from the perspective of the interpretability method of neural network path attribution to improve the effect of adversarial sample migration; that is, analyze the data set, and when the number of iterations reaches the update cycle, calculate the importance score of important neurons through neuron contribution, determine the binary mask, and effectively generate a mask model that is conducive to migration attack, and then calculate the cross entropy loss of the adversarial sample and the corresponding true label through the mask model and the source model respectively, obtain the first gradient and the second gradient, and weighted to obtain the total gradient, that is, introduce the loss to ensure that the generated adversarial sample has strong robustness; finally, update and iterate the adversarial sample to generate the final version of the adversarial sample; compared with the existing method, it can better locate the important features of the corresponding image category and improve the success rate of adversarial attacks on adversarial samples facing different attack target models, that is, provide reliable mechanism support through deep neural networks, ensure the technical effect of deep neural networks used in adversarial attacks, and effectively issue security warnings about deep neural network models to researchers or users, which is helpful to reversely improve the security of artificial intelligence research and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 A flowchart of an implementation of a deep neural network anti-attack method based on important neurons provided by an embodiment of the present invention; Figure 2 An example flow chart of a single-iteration algorithm for adversarial samples of a deep neural network adversarial attack method based on important neurons provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a deep neural network anti-attack method based on important neurons proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0024] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0025] The following is a detailed description of a method for countering attacks on a deep neural network based on important neurons provided by the present invention in conjunction with the accompanying drawings.
[0026] See also Figure 1 , which shows a flowchart of the steps of an important neuron-based deep neural network anti-attack method provided by an embodiment of the present invention, the method comprising: Step S1: Prepare a data set and a source model in sequence. The data set includes input samples and corresponding true labels, and set corresponding operating parameters for the source model. Iterate the data set. When the number of iterations reaches the preset operating parameters, calculate the importance scores of all neurons in the source model, generate a binary mask, and obtain a masked mask model. Step S2: after the number of iterations reaches the preset running parameters, the adversarial sample is output, the cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model and the source model respectively, and the first gradient and the second gradient of the cross entropy loss with respect to the adversarial sample are obtained accordingly, and the first gradient and the second gradient are weighted to obtain the total gradient; Step S3: Update the adversarial sample in turn according to the total gradient until the iteration ends and generate the final adversarial sample.
[0027] To better illustrate, in deep learning models, when faced with inputs that have been slightly perturbed or maliciously constructed, namely adversarial samples, incorrect predictions are usually made, causing model errors. To solve this problem, adversarial samples are generated to simulate potential attackers' attack methods on the model, test and evaluate the robustness of deep learning models, so as to discover and strengthen the weaknesses of the model and improve the robustness and defense capabilities of the model. Generating adversarial samples can reveal whether the performance of the model will drop sharply when faced with specific types of perturbations. If the model is susceptible to specific adversarial samples, it means that its generalization ability is insufficient. That is, through the generation and attack of adversarial samples, the defense mechanism of the model can be tested and improved, so as to better respond to potential security threats in practical applications and promote the sustainable development of AI (Artificial Intelligence).
[0028] It can be understood that a data set and a source model are prepared in sequence, the data set includes input samples and corresponding true labels, and the source model is an initial model used to train the data set; in this embodiment, the input samples are images in an ImageNet-compatible data set, which include clean color images of different subjects and the same size taken by a camera, and the true labels are the names of the subjects in the corresponding clean color images, that is, csv files; preferably, 1,000 clean color images are selected from the data set as input samples in this embodiment; convolutional neural network architectures, namely DenseNet121 (a specific model in the DenseNet family), SqueezeNet1_0 (a lightweight convolutional neural network architecture), and MobileNetV3-L (version v3 in the MobileNet model) are used as source models for implementing adversarial attacks, and the operating parameters provided by PyTorch are used as the operating parameters of the source model.
[0029] Furthermore, in step S1, the operating parameters include the total number of iterations, the maximum allowable distortion, the momentum factor, the mask ratio, the number of masks, the balance parameter, the step length and the update period. The calculation formula corresponding to the update period is:
[0030] in, Indicates the update cycle; Indicates the total number of iterations; Indicates the number of masking times.
[0031] Preferably, in this embodiment, the total number of iterations is set ; Maximum allowable distortion ; Momentum Factor ; Mask times ; Balance parameters ; Step size, that is, the step size of each iteration of the source model ; Then determine the update cycle through the calculation formula ; It can be explained that the mask ratio is set differently for different source models, namely the mask ratio of DenseNet121, SqueezeNet1_0, and MobileNetV3-L The corresponding settings are 20%, 15%, and 15% respectively.
[0032] See also Figure 2 , which shows an example flow chart of a single-iteration algorithm for adversarial samples of a deep neural network adversarial attack method based on important neurons provided by an embodiment of the present invention.
[0033] Furthermore, in step S1, when the number of iterations reaches a preset operating parameter, the importance scores of all neurons in the source model are calculated, a binary mask is generated, and a mask model after masking is obtained, including: Step S11: When the number of iterations reaches the update cycle, the importance scores of all neurons in the source model are obtained using the neuron importance calculation method.
[0034] Specifically, when the number of iterations Reach the update cycle of the mask model When, based on the input sample Each time an image and its corresponding true label are input into different source models, and the corresponding true label is ; Use the neuron importance calculation method to obtain the importance scores of all neurons in the source model.
[0035] Furthermore, in step S11, a neuron importance calculation method is used to obtain the importance scores of all neurons in the source model, and the corresponding calculation formula is:
[0036] in, Indicates Layer The neurons are related to the input samples Importance score; Represents the input sample of the dataset Through the network output of the source model; Represents the input sample And the neuron activation response The network output when it approaches 0; Indicates Layer The activation response value of a neuron.
[0037] To explain, Represents the input sample And the neuron activation response The network output when it approaches 0; Indicates Layer The neurons are related to the input samples The importance score of the neuron for the input sample Contribution to decision making; The first-order Taylor approximation expansion of the importance score is presented. This setting not only improves computational efficiency but also helps to intuitively understand the contribution of each neuron to the decision of the source model.
[0038] Step S12: sort the neurons in each layer in descending order based on the importance score, and determine the binary mask layer by layer through the mask ratio.
[0039] It is explained that binary masking refers to binarizing the activation response value of each neuron, that is, according to the preset mask ratio, important neurons with higher importance scores retain their activation response values, while non-important neurons with lower importance scores have their activation response values set to 0, so as to highlight the important neurons that contribute more to the decision-making of the source model, and ignore the neurons with smaller contributions, which is conducive to more accurate positioning and utilization of key neurons in subsequent adversarial attacks.
[0040] Specifically, neurons are sorted in descending order layer by layer according to their importance scores, and binary masks are generated layer by layer through mask ratios. The corresponding logical expression is:
[0041] Among them, according to the different source models, the corresponding mask ratio Therefore, for different source models, the important neurons and non-important neurons determined are also different.
[0042] Step S13: Obtain a masked mask model based on the binary mask combined with the source model. The corresponding calculation formula is:
[0043] in, represents a mask model; Represents the source model; represents the element-wise product between a neuron and a binary mask; Represents a binary mask.
[0044] It is explained that the mask model refers to a model obtained by evaluating the importance scores of important neurons in the source model and, according to a preset mask ratio, retaining the activation response values of important neurons with higher importance scores and setting the activation response values of non-important neurons with lower importance scores to 0.
[0045] Furthermore, step S2 includes: Step S21: When the number of iterations reaches the preset operating parameters, output the adversarial sample after the corresponding number of iterations , and the corresponding true label .
[0046] It can be understood that when the number of iterations reaches the set operating parameters, that is, the mask model reaches the set update cycle, the importance scores of all neurons are obtained at this time, and the adversarial samples generated after the corresponding number of iterations are , which means The adversarial sample after iterations and the corresponding true label .
[0047] Step S22: The cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model: ; The cross entropy loss between the adversarial sample and the corresponding true label is calculated by the source model: .
[0048] It is explained that the cross entropy loss between the adversarial sample and the corresponding true label is calculated by the mask model and the source model respectively, which can evaluate the mask model and the source model in the adversarial sample. The performance difference between the two can be judged according to the cross entropy loss between the two. It can be judged whether the mask model successfully retains the key features of the source model and achieves the sparseness of important neurons while maintaining the model performance. If the cross entropy loss of the mask model is similar to or lower than that of the source model, it means that the mask model maintains good model performance while reducing the number of important neurons. This setting can reduce the complexity of the mask model.
[0049] Step S23: Obtain the first gradient and the second gradient in sequence through the back propagation method according to the cross entropy loss, and obtain the total gradient by weighting the first gradient and the second gradient through the balance parameter.
[0050] To illustrate, the first gradient refers to the cross entropy loss of the mask model with respect to the adversarial sample The gradient of reflects the mask model in the adversarial sample The gradient of the prediction error for the true label; similarly, the second gradient refers to the cross entropy loss of the source model with respect to the adversarial sample The gradient of the source model reflects the The total gradient obtained by weighting the two gradients by the balance parameter takes into account both the ability of the mask model in maintaining model performance and the retention of key features of the source model, so that the subsequently generated adversarial samples can maintain a good attack effect when attacking the mask model, and can retain the key features of the source model as much as possible, thereby achieving neuron sparsification without significantly reducing the performance of the model.
[0051] Preferably, a MIM (Momentum Iterative Fastest Descent) optimizer is provided based on gradient back-propagation, which can accelerate the generation process of adversarial samples and improve the efficiency of adversarial attacks. That is, by introducing a momentum term, accumulation can be performed in the gradient direction to achieve the purpose of accelerating the convergence speed, so that the adversarial samples can approach the optimal solution faster; and the MIM optimizer can also enhance the robustness of adversarial attacks, so that the generated adversarial samples are more stable and effective when facing the model defense mechanism.
[0052] Further, in step S23, the first gradient and the second gradient are obtained in sequence by back propagation method according to the cross entropy loss, and the corresponding calculation formula is:
[0053]
[0054] in, represents the first gradient; represents the second gradient; represents the gradient operator; The total gradient is obtained by weighting the first gradient and the second gradient by the balance parameter. The corresponding calculation formula is:
[0055] in, represents the total gradient; represents the equilibrium parameter.
[0056] Furthermore, step S3 includes: Step S31: According to the total gradient, use momentum iteration to update the adversarial sample and generate adversarial samples of adjacent iterations.
[0057] It is explained that iterative updates using the momentum iteration method can accelerate convergence and improve attack efficiency. In the process of momentum iteration, by accumulating the gradient information of previous iterations, the gradient direction can be smoothed, the oscillation can be reduced, and the adversarial sample can approach the optimal solution faster, which helps to jump out of the local optimal solution and explore a broader solution space, thereby enhancing the diversity and success rate of adversarial attacks.
[0058] Furthermore, in step S31, the corresponding calculation formula is:
[0059]
[0060] in, Indicates the adjacent order, i.e. Adversarial examples of iterations; Indicates Adversarial examples of iterations; represents the truncation function used to satisfy the infinite norm constraint against noise; represents the maximum disturbance; Indicates the step size of each iteration; represents a symbolic function; Indicates Momentum in iterations; represents the momentum factor; Indicates Momentum in iterations; represents the total gradient; Represents a norm operation.
[0061] Step S32: Continue to iterate the adversarial samples of adjacent iterations until the total number of iterations is reached, and the iteration ends to generate the final adversarial sample.
[0062] Specifically, until the total number of iterations is reached, in this embodiment, it means that the number of iterations reaches Then stop the iteration and generate the final adversarial sample, i.e. .
[0063] It can be understood that the adversarial attack method provided by the present application starts from the perspective of the interpretability method of neural network path attribution to improve the effect of adversarial sample migration; that is, analyze the data set, and when the number of iterations reaches the update cycle, calculate the importance score of important neurons through neuron contribution, determine the binary mask, and effectively generate a mask model that is conducive to migration attack, and then calculate the cross entropy loss of the adversarial sample and the corresponding true label through the mask model and the source model respectively, obtain the first gradient and the second gradient, and weight the total gradient, that is, introduce the loss to ensure that the generated adversarial sample has strong robustness; finally, update and iterate the adversarial sample to generate the final version of the adversarial sample; compared with the existing method, it can better locate the important features of the corresponding image category and improve the success rate of adversarial attacks on adversarial samples facing different attack target models, that is, provide reliable mechanism support through deep neural networks, ensure the technical effect of deep neural networks used in adversarial attacks, and effectively issue security warnings about deep neural network models to researchers or users, which helps to reversely improve the security of artificial intelligence research and application.
[0064] For better explanation, a migration attack experiment is used to intuitively compare the actual performance of using different source models to act on the adversarial attack method, so as to evaluate the effect of the adversarial attack method provided by the present application; preferably, in this embodiment, based on the source model, VGG16, ResNet-50 (Res-50), ResNet152 (Res-152), DN121, DenseNet161 (DN162), SqueezeNet1_1 (SN1_1) are selected as the target models of the attack, and the success rate of the migration attack is tested respectively, and the average success rate is statistically calculated; and the classic method of migration adversarial attack Y. Dong, et al. Boosting adversarial attacks with momentum. IEEE CVPR, 2018, pp. 9185–9193 is used as the benchmark method to be compared, as shown in Table 1, the migration attack success rate of different methods.
[0065] Table 1. The success rate of migration attack based on deep neural network adversarial attack methods and baseline methods
[0066] It is explained that the deep neural network adversarial attack method based on important neurons proposed in the present invention has a better migration attack success rate than the baseline method depending on the source model, which means that the adversarial attack method of the present invention can effectively improve the migration attack success rate of different models.
[0067] It can be explained that, through the ablation experiment, it is judged whether different mask ratios have an impact on the attack success rate of the counter-attack method proposed in the present invention, so as to determine the feasibility of the counter-attack method.
[0068] Table 2 Mask ratio Ablation experiment on DenseNet121
[0069] Table 3 Mask ratio Ablation experiment in SqueezeNet1_0
[0070] Table 4 Mask ratio Ablation experiment on MobileNetV3-L
[0071] It is explained that for different source models, with the increase of mask ratio, the attack success rate shows a trend of first rising and then falling. At a lower mask ratio, the attack success rate is relatively low. It may be because the mask ratio is too low, resulting in excessive blocking of key information, which indirectly affects the generation effect of adversarial samples; and with the increase of mask ratio, the attack success rate shows an upward trend, indicating that an appropriate mask ratio helps to improve the effect of adversarial attacks; but when the mask ratio is too high, the attack success rate shows a downward trend again. This may be because too much information is blocked, resulting in the adversarial sample losing the effective simulation of the initial adversarial sample, reducing the success rate of the attack; that is, for lightweight source models with fewer network parameters, the attack success rate is highest when the mask ratio is between 15% and 25%; for source models with larger network parameters, the attack success rate is highest when the mask ratio is between 20% and 30%; therefore, in practical applications, different mask ratios can be selected according to different source models selected according to needs to maximize the success rate of adversarial attacks.
[0072] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0073] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
Claims
1. A deep neural network counterattack method based on important neurons, characterized in that: The method comprises: Prepare a data set and a source model in sequence, wherein the data set includes input samples and corresponding true labels, and set corresponding operating parameters for the source model, iterate the data set, and when the number of iterations reaches the preset operating parameters, calculate the importance scores of all neurons in the source model, generate a binary mask, and obtain a masked mask model; After the number of iterations reaches the preset running parameters, the adversarial sample is output, and the cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model and the source model respectively, and the first gradient and the second gradient of the cross entropy loss with respect to the adversarial sample are obtained accordingly, and the first gradient and the second gradient are weighted to obtain the total gradient; The adversarial samples are updated iteratively according to the total gradient until the iteration ends and the final adversarial sample is generated.
2. According to claim 1, a deep neural network counterattack method based on important neurons is characterized in that: The operating parameters include the total number of iterations, the maximum allowable distortion, the momentum factor, the mask ratio, the number of masks, the balance parameter, the step size and the update period. The calculation formula corresponding to the update period is: ; in, Indicates the update cycle; Indicates the total number of iterations; Indicates the number of masking times.
3. A deep neural network counterattack method based on important neurons according to claim 2, characterized in that: When the number of iterations reaches the preset running parameters, the importance scores of all neurons in the source model are calculated, a binary mask is generated, and a masked mask model is obtained, including: When the number of iterations reaches the update cycle, the importance scores of all neurons in the source model are obtained using the neuron importance calculation method; The neurons in each layer are sorted in descending order based on the importance score, and the binary mask is determined layer by layer through the mask ratio; The mask model after masking is obtained by combining the binary mask with the source model. The corresponding calculation formula is: ; in, represents a mask model; Represents the source model; represents the element-wise product between a neuron and a binary mask; Represents a binary mask.
4. The method for countering attacks on a deep neural network based on important neurons according to claim 3 is characterized in that: The importance scores of all neurons in the source model are obtained by using the neuron importance calculation method. The corresponding calculation formula is: ; in, Indicates Layer The neurons are related to the input samples Importance score; Represents the input sample of the dataset Through the network output of the source model; Represents the input sample And the neuron activation response The network output when it approaches 0; Indicates Layer The activation response value of a neuron.
5. The method for countering attacks on a deep neural network based on important neurons according to claim 1, characterized in that: After the number of iterations reaches the preset running parameters, the adversarial sample is output, and the cross entropy loss between the adversarial sample and the corresponding true label is calculated through the mask model and the source model respectively, and the first gradient and the second gradient of the cross entropy loss with respect to the adversarial sample are obtained accordingly. The first gradient and the second gradient are weighted to obtain the total gradient, including: When the number of iterations reaches the preset running parameters, the adversarial sample after the corresponding number of iterations is output. , and the corresponding true label ; The cross entropy loss between the adversarial sample and the corresponding true label is calculated by the mask model: ; The cross entropy loss between the adversarial sample and the corresponding true label is calculated by the source model: ; The first gradient and the second gradient are obtained in sequence through the back propagation method according to the cross entropy loss, and the total gradient is obtained by weighting the first gradient and the second gradient through the balance parameter.
6. The method for countering attacks on a deep neural network based on important neurons according to claim 5, characterized in that: The first gradient and the second gradient are obtained in turn by back propagation method according to the cross entropy loss, and the corresponding calculation formula is: ; ; in, represents the first gradient; represents the second gradient; represents the gradient operator; The total gradient is obtained by weighting the first gradient and the second gradient by the balance parameter. The corresponding calculation formula is: ; in, represents the total gradient; represents the equilibrium parameter.
7. The method for countering attacks on a deep neural network based on important neurons according to claim 2, characterized in that: The adversarial samples are updated iteratively according to the total gradient until the iteration ends and the final adversarial sample is generated, including: According to the total gradient, the adversarial sample is updated using momentum iteration to generate adversarial samples of adjacent iterations; Continue to iterate the adversarial samples of adjacent iterations until the total number of iterations is reached, the iteration ends, and the final adversarial sample is generated.
8. The method for countering attacks on a deep neural network based on important neurons according to claim 7, characterized in that: According to the total gradient, momentum is used to iteratively update the adversarial sample to generate adversarial samples of adjacent iterations. The corresponding calculation formula is: ; ; in, Indicates the adjacent order, i.e. Adversarial examples of iterations; Indicates Adversarial examples of iterations; represents the truncation function used to satisfy the infinite norm constraint against noise; represents the maximum disturbance; Indicates the step size of each iteration; represents a symbolic function; Indicates Momentum in iterations; represents the momentum factor; Indicates Momentum in iterations; represents the total gradient; Represents a norm operation.
Citation Information
Patent Citations
Migratable attack countermeasure method based on feature importance perception
CN117556886A
Action recognition migration attack method based on adaptive gradient time sequence feature pruning
CN118587561A
Cited By
Counterfactual guided layer aggregation diverse neuron attribution attack method
CN122864691A