General adversarial disturbance generation method for adversarial samples
By screening the features of the intermediate layer of deep neural networks, designing the adversarial loss function and introducing momentum term optimization, a high-translocation universal adversarial perturbation was generated, and the existing methods depend on and overfitting a large number of labeled data was solved, and efficient general adversarial perturbation attack was achieved.
Patent Information
- Application Number
- CN202510267269.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing general anti-perturbation attack method requires a large number of annotated training samples, which can easily lead to overfitting specific models and poor migration.
By analyzing the characteristics of the intermediate layer of deep neural networks, screening important features and secondary features, designing a feature-based adversarial loss function, and introducing momentum term optimization general adversarial perturbation, suppressing important features to enhance secondary features, and optimizing general adversarial samples with high mobility.
Effective general adversarial perturbation is generated without large-scale annotation of data sets, which improves its mobility, avoids overfitting of specific models, and can achieve high attack success rates on unknown models and enhances the adversarial robustness of the model.
Smart Images

Figure CN120297359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of adversarial examples and deep neural networks, and particularly relates to a method for generating general adversarial perturbations of adversarial examples. Background Art
[0002] Artificial intelligence technology with deep neural networks as the core component has led the technological revolution in the world today, triggered profound social changes, and greatly promoted the levels of automation and intelligence. Typical applications include: face recognition technology is widely used in face verification access control and face payment, and object detection and object tracking technologies are widely used in fields such as video surveillance and vehicle detection. Although it brings great convenience to human society, artificial intelligence technology inevitably has security problems. Recently, some researchers have shown that by injecting subtle and invisible noise into the input samples to form new samples, the deep neural network classifier can be misled to output wrong results, and such samples are called adversarial examples.
[0003] From the construction method of adversarial examples, the key to adversarial examples lies in the adversarial perturbations added to clean samples. Existing attack methods all focus on how to generate efficient adversarial perturbations, and adversarial perturbations can be divided into two categories according to their practicality: sample-independent adversarial perturbations and general adversarial perturbations. Sample-independent adversarial perturbations mean that for each clean sample, the attack algorithm has to regenerate an adversarial perturbation. General adversarial perturbations mean that for a specific dataset and a specific model, the attack algorithm only needs to generate one adversarial perturbation, and this adversarial perturbation can mislead the model to output wrong recognition results on most samples of this dataset. Although the attack effect of sample-independent adversarial perturbations is stronger, their practicality is worse. On the contrary, general adversarial perturbations are slightly less aggressive than sample-independent adversarial attack algorithms, but their practicality is stronger because there is no need to re-execute the algorithm to calculate new adversarial perturbations during deployment.
[0004] Existing general adversarial perturbation attack methods require a large number of labeled training samples, while the general adversarial perturbation attacks trained with a small number of training samples have poor effects. The adversarial loss designed based on the model training loss is likely to cause the general adversarial perturbation to overfit to a specific model. For example, Moosavi-Dezfooli et al. first proposed the concept of general adversarial perturbation and used an iterative method to accumulate the adversarial perturbations that misclassify the training samples, thus constructing a general adversarial perturbation that can misclassify a large number of training samples. Mopuri et al. proposed to maximize the activation value of the intermediate network layer of the model. This method gets rid of the need for a large amount of labeled training data, but the attack effect is poor. Although many general adversarial perturbation attack methods have emerged in recent years, significantly reducing the requirements for training data and greatly improving the performance of adversarial attacks. However, the existing methods do not consider the transferability of general adversarial perturbations. In addition, maximizing the model training loss is likely to cause the problem that the trained general adversarial perturbation overfits to a specific model. Summary of the Invention
[0005] To solve some or all of the technical problems existing in the above-mentioned prior art, the present invention provides a method for generating general adversarial perturbations of adversarial samples, enabling the general adversarial perturbations to effectively capture the local and global information of the training data set, greatly improving the transferability of the general adversarial perturbations, thus eliminating the need for a large amount of labeled data sets, and the adversarial loss function used is not likely to cause the trained general adversarial perturbations to overfit to a specific model.
[0006] The technical solution of the present invention is as follows:
[0007] A method for generating general adversarial perturbations of adversarial samples is provided, including:
[0008] Sort all the feature maps in the channel dimension of any layer of the deep neural network according to the designed preset sorting method, split all the sorted feature maps into two groups according to a preset percentage, and label them as the first feature and the second feature respectively, and regard the first feature as the important feature and the second feature as the secondary feature;
[0009] Establish an adversarial loss function based on the feature maps, so that the adversarial loss function can suppress the important features and enhance the secondary features to obtain general adversarial samples;
[0010] Optimize the obtained general adversarial samples by using a momentum term to obtain general adversarial samples with attack transferability.
[0011] Further, in the above method for generating general adversarial perturbations of adversarial samples, the first feature includes the part from the first feature map to the preset percentage according to the preset sorting method after splitting all the sorted feature maps according to the preset percentage, and the second feature includes the part from the preset percentage to the last feature map in the sorting.
[0012] Further, in the above general adversarial perturbation generation method for adversarial samples, the adversarial loss function based on feature maps includes:
[0013]
[0014] Among them, represents the adversarial loss function based on features, log(·) represents the logarithmic function, |·| represents the absolute value function, and the ⊙ operation represents element-wise multiplication. represents the gradient corresponding to the important feature, i represents any channel dimension of any layer of the important feature in the deep neural network, S represents the important feature, k represents any channel dimension of any layer of the secondary feature in the deep neural network, and S ′ represents the secondary feature, x represents any adversarial sample, and δ represents the general adversarial perturbation. represents the gradient corresponding to the secondary feature. represents the feature map of the i-th channel of the l-th layer of the deep neural network f(x + δ) containing the general adversarial perturbation δ. represents the feature map of the k-th channel of the l-th layer of the deep neural network f(x + δ) containing the general adversarial perturbation δ.
[0015] Further, in the above general adversarial perturbation generation method for adversarial samples, the momentum term includes:
[0016]
[0017] Among them, g t+1 represents the momentum gradient of the (t + 1)-th iteration, t represents the t-th iteration, m represents the batch input size, and g t represents the momentum gradient of the t-th iteration. represents the derivative of x bs with respect to, represents the gradient of the data in the mini-batch, ||·||1 represents the L1 norm, and μ represents the attenuation degree of historical gradient information.
[0018] Further, in the above general adversarial perturbation generation method for adversarial samples, when using the momentum term to optimize the obtained general adversarial samples, the general adversarial perturbation δ needs to satisfy the following condition:
[0019] δ = δ + η · sign(g t+1 );
[0020] Among them, η represents the update step size during the iteration, and sign(·) represents the sign function.
[0021] Further, in the above general adversarial perturbation generation method for adversarial samples, when optimizing the obtained general adversarial samples using a momentum term, the general adversarial perturbation δ needs to satisfy the following conditions:
[0022] δ = min(max(-∈, δ), ∈);
[0023] It means that the maximum value of the general adversarial perturbation δ does not exceed ∈, and the minimum value is not less than -∈, where ∈ represents the maximum allowable modification amplitude of the general adversarial perturbation.
[0024] Further, in the above general adversarial perturbation generation method for adversarial samples, all the feature maps in the channel dimension of any layer in the deep neural network are sorted according to a preset sorting method, and all the sorted feature maps are split according to a preset percentage to obtain a first feature and a second feature, including:
[0025] Obtain the channel-level features of any layer in the deep neural network;
[0026] Calculate the mean value of the obtained channel-level features;
[0027] Calculate each feature map in the channel level to obtain a calculated value corresponding to each feature map;
[0028] Compare the calculated value corresponding to each feature map obtained with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value;
[0029] Arrange the feature maps greater than the mean value in descending order to obtain the sorted feature map indices;
[0030] Split the obtained feature map indices according to a preset percentage, and use the features before or within the preset percentage as important features, and the remaining features as secondary features.
[0031] Further, in the above general adversarial perturbation generation method for adversarial samples, comparing the calculated value corresponding to each feature map obtained with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value includes:
[0032]
[0033] where c i represents any value among the number of feature values of each feature map in the channel whose calculated value is greater than the channel feature mean value, represents the indicator function, when the expression in the function is true, otherwise it is 0; represents the feature value at (j, k) in the feature map, It represents calculating the mean of the obtained channel-level features. w represents the width of the feature map, h represents the height of the feature map, and i, j, and k respectively represent general feature index values, where k is the column at the j-th row of the feature map.
[0034] As an alternative implementation, in the above general adversarial perturbation generation method for adversarial samples, the preset percentage includes any value in the range of 0 - 100%.
[0035] As an alternative implementation, in the above general adversarial perturbation generation method for adversarial samples, the preset percentage includes any one of 10%, 20%, 30%, 40%, 50%, and 60%.
[0036] The main advantages of the technical solution of the present invention are as follows:
[0037] The general adversarial perturbation generation method for adversarial samples of the present invention analyzes the mechanism of features in the middle layer of the network. First, it proposes an important feature screening method, effectively separating important features from secondary features. Based on the separated important features and secondary features, a feature-based adversarial loss function is designed to guide the learning of general adversarial perturbations on how to suppress important features and enhance secondary features, so as to achieve the purpose of interfering with model features. In order to effectively capture the global information of the dataset and the local information of the current mini-batch during the iteration process, a momentum term is introduced to record the changes in historical gradients, and finally a general adversarial perturbation with high transferability is optimized. Therefore, it does not require a large number of labeled datasets, and the used adversarial loss function is not likely to cause the trained general adversarial perturbation to overfit a specific model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present invention. The illustrative embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0039] Figure 1 is a schematic flowchart of the general adversarial perturbation generation method for adversarial samples according to an embodiment of the present invention;
[0040] Figure 2 is a flowchart of the method for sorting all feature maps in the channel dimension of any layer in the deep neural network according to a preset sorting method and splitting the sorted all feature maps according to a preset percentage to obtain the first feature and the second feature in the general adversarial perturbation generation method for adversarial samples provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0042] The following will describe in detail the technical solutions provided by the embodiments of the present invention with reference to the appended Figure 1-2 figures as shown.
[0043] As shown in the appended Figure 1 figures, the embodiments of the present invention provide a general adversarial perturbation generation method for adversarial examples, and the method includes the following steps S1 to S3:
[0044] Step S1: Sort all the feature maps in the channel dimension of any layer of the deep neural network according to a preset sorting method, split all the sorted feature maps according to a preset percentage to obtain a first feature and a second feature, and use the first feature as the important feature and the second feature as the secondary feature.
[0045] As an example, the preset sorting of the above-mentioned feature maps includes: for example, the first convolutional layer outputs 64 feature maps, and the 64 feature maps are sorted according to the designed sorting method, and then different parts are selected as important features.
[0046] The general adversarial perturbation is an invisible adversarial perturbation that can mislead the deep neural network to output incorrect results on most data. Assume that f is a deep neural network trained on the dataset (X, Y). For any sample x ∈ X, there is f(x) = y, where y ∈ Y, f(x) represents the classification result predicted by the deep neural network, X represents the training sample dataset, y represents the true label corresponding to the input sample x, and Y represents the label set in the dataset. The general adversarial perturbation δ can be defined in the following form:
[0047] f(x + δ) ≠ f(x), for x ∼ X, s.t. ||δ|| p ≤ ε;
[0048] where, ||δ|| p represents the L p norm of the general adversarial perturbation δ. The present invention rewrites the above definition of the general adversarial perturbation, aiming to indirectly mislead the model decision result by interfering with the features in the middle layer of the model. Specifically, the general adversarial perturbation δ of the present invention is defined as follows:
[0049] f l (x + δ) ≠ f l(x), for x~X, s.t. ||δ|| p ≤ε;
[0050] Where represents the feature map of the i-th channel in the l-th layer of the deep neural network f. In the present invention, L p norm is the infinity norm, that is, L ∞ norm.
[0051] Specifically, as Figure 2 shown, all the feature maps in the channel dimension of any layer in the deep neural network are sorted in a preset sorting manner, and all the sorted feature maps are split according to a preset percentage, and the first feature and the second feature are obtained, including the following steps S100-step S105:
[0052] Step S100: Obtain the channel-level features of any layer in the deep neural network;
[0053] Step S101: Calculate the mean value of the obtained channel-level features;
[0054] Step S102: Calculate each feature map in the channel level to obtain the calculated value corresponding to each feature map;
[0055] Step S103: Compare the calculated value corresponding to each feature map with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value;
[0056] Step S104: Arrange the feature maps greater than the mean value in descending order to obtain the sorted feature map indices;
[0057] Step S105: Split the obtained feature map indices according to a preset percentage, and use the features before or within the preset percentage as important features, and the remaining features as secondary features.
[0058] As some optional implementation manners of this embodiment, the first feature includes the part from the first feature to the preset percentage according to the preset sorting manner after splitting all the sorted feature maps according to the preset percentage, and the second feature includes the part from the preset percentage to the last feature map.
[0059] Specifically, comparing the calculated value corresponding to each feature map with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value includes:
[0060]
[0061] Among them, c iRepresents any value among the number of eigenvalues where the calculated value of each feature map of the channel is greater than the mean value of the channel features. Represents the indicator function, when the expression in the function is true, Otherwise it is 0; Represents the eigenvalue at the position (j, k) in the feature map. Represents calculating the mean value of the obtained channel-level features, w represents the width of the feature map, h represents the height of the feature map, i, j, and k respectively represent the general feature index values, the k-th column of the j-th row of the feature map.
[0062] As an example, there are many redundant features in a deep neural network, so there must be important features with strong correlation to the model decision and secondary features with low correlation to the model decision. To screen out the important features, the present invention proposes to sort all the feature maps of the channel dimension of a certain network layer according to a certain strategy, and then select the first half of the sorted results as important features and the second half as secondary features. For the features in the l-th layer, there is Represents that the l-th layer has n feature maps with a resolution of width w and height h. To screen out the important features of the l-th layer, the present invention proposes an important feature screening strategy, and the specific process includes:
[0063] Input: Deep neural network model f, important feature ratio r;
[0064] Output: Important feature index S and secondary feature index S';
[0065] Obtain the channel-level features of the l-th layer
[0066] Calculate the mean value of the channel-level features
[0067] Calculate the number of eigenvalues c0, c1,..., c where each feature map of the channel is greater than the mean value of the channel features n , where where Represents the indicator function, when the expression in the function is true, Otherwise it is 0; Represents the eigenvalue at the position (j, k) in the feature map;
[0068] Sort c0, c1,..., c in descending order n and obtain the sorted indices s0, s1,..., s n ;
[0069] Select the first r% of the sorted indices as important features S and the remaining part as secondary features S'.
[0070] As some alternative implementation manners of this embodiment, the above preset percentage includes any value in the range of 0-100%.
[0071] As some alternative implementation manners of this embodiment, the above preset percentage includes any one of 10%, 20%, 30%, 40%, 50% and 60%.
[0072] As an example, the above preset percentage is set to 25%. Thus, the obtained feature map indexes can be split according to the ratio of 25%. The features within the preset percentage before the first 25% are regarded as important features, and the remaining 75% of the features are regarded as secondary features. That is, the first feature mentioned above includes that after splitting all the sorted feature maps according to the ratio of 25%, the feature maps from 0 to 25% are the first feature, that is, the important feature, and the features from 25% to 75% are the second feature, that is, the secondary feature.
[0073] Step S2: Establish an adversarial loss function based on the feature map, so that the adversarial loss function can suppress the important features and enhance the secondary features to obtain a universal adversarial sample;
[0074] To achieve the purpose of interfering with the features in the middle layer of the deep neural network, the present invention designs a feature-based adversarial loss function to guide the update direction of the universal adversarial perturbation. Specifically, the adversarial loss function involved in the present invention is as follows:
[0075]
[0076] Where log(·) represents the logarithmic function, and |·| represents the absolute value function. The purpose of the adversarial loss is to minimize the activation value of the important features and maximize the activation value of the secondary features. To further identify the important regions of the selected important features, the present invention introduces gradient information into the loss function and rewrites the feature-based loss function into the following form:
[0077]
[0078] Where, represents the feature-based adversarial loss function, log(·) represents the logarithmic function, |·| represents the absolute value function, the ⊙ operation represents element-wise multiplication, represents the gradient corresponding to the important feature, i represents any channel dimension of any layer of the important features in the deep neural network, S represents the important features, k represents any channel dimension of any layer of the secondary features in the deep neural network, S ′ represents the secondary features, x represents any adversarial sample, δ represents the universal adversarial perturbation, represents the gradient corresponding to the secondary feature, Denote the feature map of the \(i\)-th channel in the \(l\)-th layer of the deep neural network \(f(x + \delta)\) containing the universal adversarial perturbation \(\delta\). Denote the feature map of the \(k\)-th channel in the \(l\)-th layer of the deep neural network \(f(x + \delta)\) containing the universal adversarial perturbation \(\delta\).
[0079] Step S3: Optimize the obtained universal adversarial samples using the momentum term to obtain universal adversarial samples with attack transferability.
[0080] To effectively optimize the universal adversarial perturbation, the present invention introduces a momentum term during the optimization process to record the global information of the dataset and the local information of each iteration batch. Specifically, the present invention updates the universal adversarial perturbation using the following update strategy;
[0081] Specifically, the momentum term includes:
[0082]
[0083] where \(g\) t+1 denotes the momentum gradient at the \((t + 1)\)-th iteration, \(t\) denotes the \(t\)-th iteration, \(m\) denotes the batch input size, and \(g\) t denotes the momentum gradient at the \(t\)-th iteration, denotes the derivative with respect to \(x\) bs , denotes the gradient of the data in the mini-batch, \(\|\cdot\|_1\) denotes the \(L_1\) norm, and \(\mu\) denotes the historical gradient information decay rate.
[0084] When optimizing the obtained universal adversarial samples using the momentum term, the universal adversarial perturbation \(\delta\) needs to satisfy the following condition:
[0085] \(\delta=\delta+\eta\cdot sign(g\) t+1 );
[0086] where \(\eta\) denotes the update step size during the iteration process. This iteration process means that at the \((t + 1)\)-th iteration, the universal adversarial perturbation \(\delta\) is updated by \(\eta\) steps along the gradient direction determined by \(sign(g\) t+1 ), and \(sign(\cdot)\) denotes the sign function.
[0087] Specifically, when optimizing the obtained universal adversarial samples using the momentum term, the universal adversarial perturbation \(\delta\) needs to satisfy the following condition:
[0088] \(\delta=\min(\max(-\epsilon,\delta),\epsilon)\);
[0089] where \(\epsilon\) denotes the maximum allowable modification amplitude of the universal adversarial perturbation. The meaning of this formula is that the maximum value of the universal adversarial perturbation does not exceed \(\epsilon\), and the minimum value is not less than \(-\epsilon\).
[0090] Thus, through the above momentum term formula, \(L\) is achieved.∞ The norm truncates the universal adversarial perturbation to the minimum maximum modification amplitude allowed. Through the iteration of the above formula, the universal adversarial perturbation with high attack transferability is finally obtained.
[0091] In summary, the universal adversarial perturbation in the method for generating universal adversarial perturbations of adversarial samples of the present invention has at least the following effects:
[0092] 1) It can be deployed offline. Once the universal adversarial perturbation is generated, it can be added to most samples without relying on the model and data, so as to mislead the target model.
[0093] 2) It can achieve a higher attack success rate on unknown deep neural network models.
[0094] 3) It can be added to clean samples to construct adversarial samples, enhance the training data, and fine-tune the deep neural network to enhance its adversarial robustness.
[0095] Therefore, the method for generating universal adversarial perturbations of adversarial samples of the present invention analyzes the mechanism of the features in the middle layer of the network, proposes an important feature screening method, and effectively separates the important features from the secondary features. Based on the separated important features and secondary features, a feature-based adversarial loss function is designed to guide the universal adversarial perturbation to learn how to suppress important features and enhance secondary features, so as to achieve the purpose of interfering with the model features. In order to effectively capture the global information of the dataset and the local information of the current mini-batch during the iteration process, a momentum term is introduced to record the changes in the historical gradients, and finally a universal adversarial perturbation with high transferability is optimized.
[0096] Therefore, the adversarial perturbation of the present invention and the adversarial samples with adversarial perturbations can be applied to any image processing scenario.
[0097] As an example, when applying the method for generating universal adversarial perturbations of adversarial samples of the present invention to the scenarios of face verification access control and face payment technology, the adversarial samples applied to this environment can be deployed offline. Once the universal adversarial perturbation is generated, it can be added to most samples without relying on the model and data, so as to mislead the target model. At the same time, it can also achieve a higher attack success rate on unknown deep neural network models. It can also be added to clean samples to construct adversarial samples, enhance the training data, and fine-tune the deep neural network to enhance its adversarial robustness. The universal adversarial perturbation applied to this scenario can effectively capture the local and global information of the training dataset, greatly improving the transferability of the universal adversarial perturbation. Therefore, a large number of labeled datasets are not required, and the adversarial loss function used is not likely to cause the trained universal adversarial perturbation to overfit to a specific model.
[0098] As another example, when applying the general adversarial perturbation generation method for adversarial samples of the present invention to video surveillance and vehicle detection scenarios, the adversarial samples applied to this environment can be deployed offline. Once the general adversarial perturbation is generated, it does not need to rely on the model and data, and can be added to most samples to mislead the target model. At the same time, it can also achieve a higher attack success rate on unknown deep neural network models. It can also be added to clean samples to construct adversarial samples, enhance the training data, and fine-tune the deep neural network to enhance its adversarial robustness. The general adversarial perturbation applied to this scenario can effectively capture the local and global information of the training data set, greatly improving the transferability of the general adversarial perturbation. Therefore, a large number of labeled data sets are not required, and the used adversarial loss function is not likely to cause the trained general adversarial perturbation to overfit to a specific model.
[0099] In summary, the present invention provides a method for generating general adversarial perturbations for adversarial samples, enabling the general adversarial perturbations to effectively capture the local and global information of the training data set, greatly improving the transferability of the general adversarial perturbations. Therefore, a large number of labeled data sets are not required, and the used adversarial loss function is not likely to cause the trained general adversarial perturbations to overfit to a specific model.
[0100] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. In addition, in this article, "front", "rear", "left", "right", "upper", and "lower" are all referenced based on the placement state shown in the drawings.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A general adversarial perturbation generation method for adversarial examples, characterized in that, Including: Sort all the feature maps in the channel dimension of any layer of the deep neural network according to the designed preset sorting method, split all the sorted feature maps into two groups according to a preset percentage, and label them as the first feature and the second feature respectively. The first feature is regarded as the important feature, and the second feature is regarded as the secondary feature; Build an adversarial loss function based on the feature maps, so that the adversarial loss function can suppress the important feature and enhance the secondary feature to obtain a universal adversarial sample; Optimize the obtained universal adversarial sample using a momentum term to obtain a universal adversarial sample with attack transferability.
2. The general adversarial perturbation generation method for adversarial samples according to claim 1, wherein The first feature includes the part from the first feature map to the preset percentage according to the preset sorting method after splitting all the sorted feature maps according to the preset percentage. The second feature includes the part from the preset percentage to the last feature map in the sorting.
3. The general adversarial perturbation generation method for adversarial examples according to claim 1, characterized in that The adversarial loss function based on the feature maps includes: Among them, represents the feature-based adversarial loss function, log(·) represents the logarithmic function, |·| represents the absolute value function, and the ⊙ operation represents element-wise multiplication. represents the gradient corresponding to the important feature, i represents any channel dimension of any layer of the important feature in the deep neural network, S represents the important feature, k represents any channel dimension of any layer of the secondary feature in the deep neural network, and S ′ represents the secondary feature, x represents any adversarial sample, and δ represents the general adversarial perturbation. represents the gradient corresponding to the secondary feature. represents the feature map of the i-th channel in the l-th layer of the deep neural network f(x + δ) containing the general adversarial perturbation δ. represents the feature map of the k-th channel in the l-th layer of the deep neural network f(x + δ) containing the general adversarial perturbation δ.
4. The general adversarial perturbation generation method for adversarial samples according to claim 1, characterized in that, The momentum term includes: Among them, g t+1 represents the momentum gradient of the (t + 1)-th iteration, t represents the t-th iteration, m represents the batch input size, and g t represents the momentum gradient of the t-th iteration, represents the derivative with respect to x bs derivation, represents the gradient of the data in the mini-batch, ||·||1 represents the L1 norm, and μ represents the attenuation degree of historical gradient information.
5. The general adversarial perturbation generation method for adversarial examples according to claim 4, wherein When optimizing the obtained universal adversarial sample using a momentum term, the universal adversarial perturbation δ needs to satisfy the following conditions: δ = δ + η·sign(g t+1 ); where η represents the update step size in the iterative process, and sign(·) represents the sign function.
6. The general adversarial perturbation generation method for adversarial samples according to claim 4, characterized in that, When optimizing the obtained universal adversarial sample using a momentum term, the universal adversarial perturbation δ needs to satisfy the following conditions: δ = min(max(-∈, δ), ∈); It means that the maximum value of the universal adversarial perturbation δ does not exceed ∈, and the minimum value is not less than -∈, where ∈ represents the maximum allowable modification amplitude of the universal adversarial perturbation.
7. The general adversarial perturbation generation method for adversarial examples according to claim 1, wherein Sorting all the feature maps in the channel dimension of any layer in the deep neural network according to the preset sorting method and splitting all the sorted feature maps according to the preset percentage to obtain the first feature and the second feature includes: Obtain the channel-level features of any layer in the deep neural network; Calculate the mean value of the obtained channel-level features; Calculate for each feature map in the channel level to obtain the calculated value corresponding to each feature map; Compare the calculated value corresponding to each feature map with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value; Arrange the feature maps greater than the mean value in descending order to obtain the sorted feature map indices; Split the obtained feature map indices according to the preset percentage, and regard the features before or within the preset percentage as important features, and the remaining features as secondary features.
8. The general adversarial perturbation generation method for adversarial examples according to claim 7, wherein Comparing the calculated value corresponding to each feature map with the calculated mean value to obtain the number of calculated values corresponding to each feature map that are greater than the mean value includes: Among them, c i represents any value among the number of eigenvalues where the calculated value of each feature map of the channel is greater than the mean value of the channel features, represents the indicator function, when the expression in the function is true, otherwise it is 0; represents the eigenvalue at (j, k) in the feature map, represents calculating the mean value of the obtained channel-level features, w represents the width of the feature map, h represents the height of the feature map, i, j, and k respectively represent the general feature index values, the k-th column of the j-th row of the feature map.
9. The general adversarial perturbation generation method for adversarial examples according to claim 7, wherein The preset percentage includes any value in 0 - 100%; 10. The general adversarial perturbation generation method for adversarial samples according to claim 7, characterized in that, The preset percentage includes any one of 10%, 20%, 30%, 40%, 50% and 60%.
Citation Information
Cited By
Generalized adversarial sample group generation method based on easily confused category feature injection
CN121145982A