Method for realizing migratable directional attack through multi-branch disturbance optimization
Through the multi-branch perturbation optimization method, the deep neural network is used to extract multi-scale features and adaptively fusion, which solves the problem of underutilizing feature information in targeted attacks, improves the migrationability and attack success rate of the adversarial samples, and enhances the robustness of the defense mechanism.
Patent Information
- Application Number
- CN202510336423.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing adversarial sample defense methods do not fully consider the multi-scale feature information and key local features of the original sample in targeted attacks, resulting in limited migration and single optimization targets, limiting the misleading ability of the adversarial sample.
Through the multi-branch perturbation optimization method, a deep neural network is used to extract key local features and multi-scale features, and generate and adaptively fusion perturbation matrix, including key local feature extraction, deep feature fusion, multi-scale perturbation generation and weighted fusion, and gradually generate adversarial samples.
It improves the migrationability of adversarial samples and the success rate of targeted attacks, enhances the robustness of the defense mechanism, and the generated adversarial samples have better universality and attack effects among different models.
Smart Images

Figure CN120279282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a transferable directional attack method achieved through multi-branch perturbation optimization. Background Art
[0002] In the context of the rapid development of science and technology, the security issues of AI systems have become increasingly prominent. In particular, in practical applications, AI systems may encounter malicious attacks, resulting in system output errors or unpredictable behaviors. Adversarial sample attacks, as a type of attack against AI systems, aim to trick AI systems into making wrong decisions by imposing tiny perturbations on input data. This type of attack poses a significant threat to the security and reliability of AI systems.
[0003] In order to enhance the anti-attack capability of AI systems, researchers have conducted a large number of studies on adversarial sample defense methods, aiming to improve the robustness and stability of the system. However, in the field of adversarial sample targeted attacks, existing research is relatively scarce, and most of the work focuses on strategies based on feature utilization and data enhancement. Feature-based defense methods use sample features through different strategies to improve the transferability of adversarial samples, such as by measuring the similarity between feature maps or maximizing feature differences. Data enhancement-based methods explore methods to improve their transferability by applying multiple transformations to input samples, such as ODI transformation. Although these methods have shown certain potential in improving the transferability of adversarial samples, there are still some problems that need to be solved. First, the existing methods do not fully consider the feature information of the original samples at different scales in the study of targeted attacks, resulting in limited transferability. Second, in the process of adversarial sample generation, the existing methods have a relatively single optimization target and do not consider key local information, which limits the misleading ability of adversarial samples.
[0004] This patent application proposes a new method that aims to further improve the transferability of adversarial samples by deeply mining the key local features and multi-scale features of the original samples, thereby effectively improving the success rate of targeted attacks and enhancing the robustness of the defense mechanism. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method for achieving transferable directed attack through multi-branch perturbation optimization, by utilizing multi-faceted sample feature information, performing random weights in a deep neural network model, and generating and adaptively fusing multi-scale sample feature perturbations in iterations, and gradually generating adversarial samples, thereby achieving the purpose of improving the directed transferability of adversarial samples.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0007] A transferable directed attack method realized by multi-branch perturbation optimization, comprising the following steps:
[0008] S1. Apply a key local feature extraction module to extract the key local features of the sample;
[0009] S2. Store the key local features in a deep feature fusion module;
[0010] S3. In a deep neural network model, fuse the global features and key local features deeply to generate a perturbation matrix;
[0011] S4. Extract the multi-scale features of the sample and input them into a classifier to generate multi-scale perturbations;
[0012] S5. Adaptively fuse the multi-scale perturbations to obtain multi-scale perturbations;
[0013] S6. Weightedly fuse the perturbation matrix generated in S3 and the multi-scale perturbation matrix generated in S5 as the final perturbation matrix; Iteratively generate the final adversarial sample according to the final perturbation matrix.
[0014] A further improvement of the technical solution of the present invention is that: in S1, the original sample is input into the key local feature extraction module, and the key local features are obtained through network analysis to provide a basis for subsequent feature fusion;
[0015] Among them, the original sample features include original image features and key local features; the key local features are the image attention heat map regions obtained by network analysis of the original image.
[0016] A further improvement of the technical solution of the present invention is that: in S1, when the sample is input into the key local feature extraction module, the key local feature extraction module extracts the local key local features of the sample; the key local feature extraction module consists of a deep neural network analysis method Grad-CAM; when the sample passes through Grad-CAM analysis, the attention activation map region M is output; then the attention region M is redefined as R according to the mapping rule T for subsequent extraction of the sample key local features; the mapping rule T refers to defining the part with a contribution to the attention heat map greater than the threshold as 1, and the rest as 0;
[0017] The method for obtaining key local features according to the original sample is as follows:
[0018] f l = CAM(x)
[0019] Among them, CAM(·) is the Grad-cam function, x is the original sample, and f l is the generated key local feature information.
[0020] A further improvement of the technical solution of the present invention lies in that: S3 specifically includes the following steps:
[0021] S3.1 The optimization process of the source model adversarial sample is carried out in the deep deep feature fusion module; during the optimization process, first generate a random number within the range of 0.0 to 1.0, and if the random number is less than the preset probability value, enter the deep feature fusion module to perform subsequent feature optimization operations;
[0022] S3.2 Shuffle and reorganize the current feature, global feature, and key local feature of the same batch to obtain the feature as the input of the next network layer;
[0023] In the network layer that meets the optimization conditions, generate a set of random fusion weights, and weighted-fuse the three features according to the corresponding weights according to the probability to obtain the feature as the input of the next network layer; after the optimization of the current network layer is completed, continue to enter the next layer for optimization until the perturbation matrix p is finally generated a 。
[0024] A further improvement of the technical solution of the present invention lies in that: in S3.2, the calculation method for the fusion optimization of the current feature, global feature, and key local feature is as follows:
[0025] f si c =S(f i c )
[0026] f si l =S(f i l )
[0027]
[0028] f i ′=α i f i +(1-α i )f i p
[0029] Wherein, S(·) is a random shuffle function, i refers to the current level of the surrogate model, and the original image feature f i c is shuffled into f si c , similarly, the key local feature f i l is randomly shuffled into f si l , β i is the weight parameter of the key local feature; f ip is f si c and f si l the fusion result of, f i is the output feature of the previous layer, α i is f i and f i p the weight of linear interpolation fusion.
[0030] A further improvement of the technical solution of the present invention is that: S4 specifically includes the following steps:
[0031] S4.1 Before each iteration, input the current sample into the multi-scale feature extraction module to obtain sample features at different scales;
[0032] The method for generating multi-scale sample features according to the current sample is as follows:
[0033] x t Ti = D(x t , N)
[0034] where x t is the iterative sample at the current moment, N is the predefined downsampling step, D(·) is the downsampling function, and x t Ti is to generate sample features at different scales;
[0035] S4.2 Input the sample features at different scales into the classifier, calculate a one-step perturbation result according to the reverse gradient, and generate multi-scale perturbations.
[0036] A further improvement of the technical solution of the present invention is that: S5 specifically includes the following steps:
[0037] S5.1 First, upsample the perturbation results at different scales to the original sample size;
[0038] S5.2 According to the perturbations at different scales and the downsampled samples, calculate the contribution degree of the current-scale perturbation to the directional attack loss;
[0039] S5.3 Based on the perturbation contributions at all different scales, adaptively calculate the weight corresponding to each perturbation, and finally calculate the weighted sum to obtain the multi-scale perturbation p b .
[0040] A further improvement of the technical solution of the present invention is that: in S5.1, the method of upsampling the perturbation results at different scales to the original sample size is as follows:
[0041] p i u = U(p i, x)
[0042] Among them, p i is the one-step perturbation result generated by samples of different scales, U(·) is the upsampling function, and p i u is the perturbation matrix consistent with the size of the original sample.
[0043] A further improvement of the technical solution of the present invention lies in: In S5.2, the method for calculating the contribution degree of the current scale perturbation to the targeted attack loss is as follows:
[0044] L i = Logit(x i Ti + p i u , y')
[0045] Among them, Logit(·) represents the Logit loss function, y' is the target class, and L i is the contribution of the current scale perturbation to the targeted attack loss.
[0046] A further improvement of the technical solution of the present invention lies in: In S5.3, the adaptive fusion strategy is as follows:
[0047]
[0048] p b = ∑p i u × ω i
[0049] Among them, ω i is the weight parameter of the perturbation matrix at each scale, and p b is the result of the final multi-scale perturbation fusion.
[0050] Due to the adoption of the above technical solution, the technical progress obtained by the present invention is:
[0051] 1. Based on the idea of deep learning, the present invention constructs a deep feature fusion module and inserts it into a deeper layer of the source model; in traditional methods, the construction of perturbations often focuses on the shallow features of the model, making it difficult to fully mine and utilize the deep semantic features of the model, resulting in insufficient diversity of perturbations; by inserting the deep feature fusion module into a deeper layer, the present invention can perturb the semantic features at a high level, thus greatly enriching the perturbation diversity and making the generated adversarial samples more deceptive.
[0052] 2. The present invention introduces the idea of multi-scale in the targeted attack, considers the feature information of samples at different resolutions, and adaptively fuses the perturbation matrices at different scales. The attack method with a single scale cannot fully consider the feature differences of samples at different resolutions, which easily leads to limited attack effects. Especially when facing complex samples, the attack success rate is relatively low. By considering multi-scale feature information and adaptively fusing the perturbation matrices, the feature changes of samples can be captured more comprehensively, thereby improving the attack success rate and making the attack more targeted and effective.
[0053] 3. The present invention adopts a multi-branch perturbation construction and synthesis process. A single perturbation construction method is difficult to meet the diverse needs of adversarial samples in different scenarios, and in cross-model attacks, the transferability of adversarial samples is relatively poor, which limits its application scope. The multi-branch perturbation construction and synthesis process not only realizes transferability but also improves the ability of adversarial samples to mislead the target model, making the adversarial samples have better generality and attack effects between different models. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the transferable targeted attack method proposed by the present invention;
[0055] Figure 2 is a flowchart for generating key local features based on the current sample;
[0056] Figure 3 is a schematic diagram of the deep feature fusion module. DETAILED DESCRIPTION OF THE INVENTION
[0057] The following further describes the present invention in detail with reference to the drawings and embodiments:
[0058] In the embodiment of the present invention, as Figure 1 shown, for the transferable targeted attack method realized by multi-branch perturbation optimization, the used targeted attack method is mainly based on two aspects: the deep feature fusion module and the multi-scale perturbation fusion module. The main methods include the following steps:
[0059] S1. Apply the key local feature extraction module to extract the key local features of the sample;
[0060] The features include two parts: the original image features and the key local features. The key local features are the image attention heat map regions obtained by analyzing the original image through the network. The original image features and the key local features are randomly fused with the features from the previous layer to increase the diversity and richness of the perturbation.
[0061] First, input the original sample into the key local feature extraction module, and obtain the key local features through network analysis, providing a basis for subsequent feature fusion;
[0062] The following combines Figure 2 to elaborate on the acquisition of key local feature information in S1.
[0063] When a sample is input into the key local feature extraction module, the key local feature extraction module extracts the local key local features of the sample according to a certain strategy; the key local feature extraction module consists of a deep neural network analysis method Grad-CAM; when the sample is analyzed by Grad-CAM, the attention activation map region M is output; then the attention region M is redefined as R according to the mapping rule T for subsequent extraction of the sample's key local features; the mapping rule T refers to defining as 1 the part of the attention heatmap that contributes more than the threshold, and the rest as 0, and the value of the threshold ranges from 0 to 1, generally taking 0.5;
[0064] The method for obtaining key local features from the original sample is as follows:
[0065] f l = CAM(x)
[0066] where CAM(·) is the Grad-cam function, x is the original sample, and f l is the generated key local feature information.
[0067] S2. Store the key local features in the deep feature fusion module;
[0068] The deep feature fusion module is as shown in the DFF module in Figure 1 : As shown in the upper part, before feature fusion, the global feature of the original image and the key local features are stored in different deeper layers of the source model; this is because compared with perturbing low-level features, perturbing high-level features is more conducive to changing the region of the original image that has a large impact on the classification result and achieving higher transferability. Figure 1 The stored features are to save two related features to deeper layers of the source model so as to introduce perturbations at a higher semantic level, thereby enhancing the diversity of perturbations and the effect of targeted attacks.
[0069] The stored features are to save two related features to deeper layers of the source model so as to introduce perturbations at a higher semantic level, thereby enhancing the diversity of perturbations and the effect of targeted attacks.
[0070] S3. Deeply fuse the global features and key local features in the deep neural network model and generate a perturbation matrix;
[0071] As shown in Figure 3As shown, in the iterative optimization process, the images of the same batch are first subjected to input transformation processing, and then the global features and key features are randomly shuffled and recombined in the deep neural network model, which helps to improve the confusion of adversarial samples and the success rate of attacks. With a certain probability, the current features, global features, and key local features are combined according to the random fusion weights to form image features with rich perturbations and strong transferability, and finally a perturbation matrix is generated;
[0072] S3 specifically includes:
[0073] S3.1 The optimization process of the source model's adversarial samples is carried out in the deep feature fusion module; in the optimization process, a random number in the range of 0.0 to 1.0 is first generated. If the random number is less than the preset probability value, it enters the deep feature fusion module for subsequent feature optimization operations; this approach aims to avoid excessive interference with the original features and thus prevent the degradation of attack performance;
[0074] S3.2 Shuffle and recombine the current features, global features, and key local features of the same batch to obtain the features as the input of the next network layer;
[0075] In the network layer that meets the optimization conditions, the specific operation is to generate a set of random fusion weights, and the three features are weighted and fused according to the corresponding weights according to the probability to obtain the features as the input of the next network layer. After the optimization of the current network layer is completed, continue to enter the next layer for optimization until the perturbation matrix p is finally generated a ;
[0076] The calculation method for the fusion optimization of the current features, global features, and key local features is as follows:
[0077] f si c = S(f i c )
[0078] f si l = S(f i l )
[0079]
[0080] f i ' = α i f i + (1 - α i )f i p
[0081] Among them, S(·) is a random shuffling function, i refers to the current layer of the surrogate model, and the original image feature f ic is scrambled into f si c , similarly, the key local feature f i l is randomly scrambled into f si l , β i is the weight parameter of the key local feature. f i p is f si c and f si l 's fusion result, f i is the output feature of the previous layer, α i is f i and f i p The weight of linear interpolation fusion.
[0082] S4. Extract the multi-scale features of the sample and input them into the classifier to generate multi-scale perturbations;
[0083] As Figure 1 shown in the lower part, S4 specifically includes the following steps:
[0084] S4.1 Before each iteration, input the current sample into the multi-scale feature extraction module to obtain the sample features at different scales; the sample at different scales represents the performance and features of the image at different resolutions. By processing this multi-scale sample information, various features from details to the whole, from local to global can be captured, improving the attack performance;
[0085] The method for generating multi-scale sample features according to the current sample is as follows:
[0086] x t Ti = D(x t , N)
[0087] where x t is the iterative sample at the current moment, N is the predefined downsampling step, D(·) is the downsampling function, and x t Ti is to generate sample features at different scales;
[0088] S4.2 Input the sample features at different scales into the classifier, and calculate a one-step perturbation result according to the reverse gradient to generate multi-scale perturbations.
[0089] S5. Adaptive fusion of multi-scale perturbations to obtain multi-scale perturbations;
[0090] As Figure 1As shown in the lower part, the multi-scale perturbation results generated by S4 are adaptively fused; before fusion, to meet the requirement of dimension matching, the low-scale perturbation matrix needs to be extended and filled; finally, fusion is performed according to the designed adaptive fusion strategy;
[0091] S5 specifically includes the following steps:
[0092] S5.1 First, upsample the perturbation results of different scales to the original sample size;
[0093] The method of upsampling the perturbation results of different scales to the original sample size is as follows:
[0094] p i u =U(p i ,x)
[0095] where p i is the one-step perturbation result generated by different-scale samples in S4.2, U(·) is the upsampling function, and p i u is the perturbation matrix consistent with the original sample size;
[0096] S5.2 Calculate the contribution degree of the current-scale perturbation to the targeted attack loss based on the perturbations of different scales and the downsampled samples;
[0097] The method of calculating the contribution degree of the current-scale perturbation to the targeted attack loss is as follows:
[0098] L i =Logit(x i Ti +p i u ,y')
[0099] where Logit(·) represents the Logit loss function, y' is the target class, and L i is the contribution of the current-scale perturbation to the targeted attack loss;
[0100] S5.3 According to the contributions of all perturbations of different scales, adaptively calculate the weight corresponding to each perturbation, and finally calculate the weighted sum to obtain the multi-scale perturbation p b ;
[0101] The adaptive fusion strategy is as follows:
[0102]
[0103] p b =∑p i u ×ω i
[0104] Among them, ω i is the weight parameter of the perturbation matrix at each scale, and p b is the result of the final multi-scale perturbation fusion.
[0105] S6. Perform weighted fusion on the perturbation matrix generated in S3 and the multi-scale perturbation matrix generated in S5 to serve as the final perturbation matrix;
[0106] S7. Iteratively generate the final adversarial sample according to the final perturbation matrix.
[0107] In summary, in view of the problems of single perturbation in existing targeted attacks and the failure to consider the multi-scale sample characteristics under targeted attacks, the present invention proposes a method for transferable targeted attacks through multi-branch perturbation optimization. By obtaining and storing global clean features and key local features, and reasonably fusing them with a certain weight ratio during model inference optimization, sample features with rich perturbations are formed, and finally the perturbation pa is generated. Secondly, by utilizing the feature information of different scales of samples, multi-scale perturbations pb are constructed and adaptively fused. The two-branch perturbations are superimposed on the original sample to generate an adversarial sample, which has the advantages of perturbation diversity and strong transferability. Through the above optimization process, the problems of single perturbation in targeted attacks and the failure to consider multi-scale samples under targeted attacks are solved, and on this basis, the transferability of targeted attacks is greatly improved.
Claims
1. A transferable directed attack method achieved by multi-branch perturbation optimization, characterized in that: It includes the following steps: S1. Apply the key local feature extraction module to extract the key local features of the sample; S2. Store the key local features in the deep feature fusion module; S3. In the deep neural network model, fuse the global features, key local features, and deep features, and generate a perturbation matrix; S4. Extract the multi-scale features of the sample and input them into the classifier to generate multi-scale perturbations; S5. Adaptively fuse the multi-scale perturbations to obtain the multi-scale perturbations; S6. Weightedly fuse the perturbation matrix generated in S3 and the multi-scale perturbation matrix generated in S5 as the final perturbation matrix; Iteratively generate the final adversarial sample according to the final perturbation matrix.
2. The method for implementing transferable targeted attack through multi-branch perturbation optimization according to claim 1, characterized in that: In S1, the original sample is input into the key local feature extraction module, and the key local features are obtained through network analysis, providing a basis for subsequent feature fusion; Among them, the original sample features include the original image features and the key local features; the key local features are the image attention heat map regions obtained by analyzing the original image through the network.
3. The method for realizing transferable targeted attack through multi-branch perturbation optimization according to claim 2, wherein: In S1, when the sample is input into the key local feature extraction module, the key local feature extraction module extracts the local key local features of the sample; The key local feature extraction module consists of a deep neural network analysis method Grad-CAM; when the sample passes through the Grad-CAM analysis, the attention activation map region M is output; then the attention region M is redefined as R according to the mapping rule T for subsequent extraction of the sample key local features; The mapping rule T refers to defining the part with a contribution to the attention heat map greater than the threshold as 1, and the rest as 0; The method for obtaining the key local features according to the original sample is as follows: f l = CAM(x) Among them, CAM(·) is the Grad-cam function, x is the original sample, and f l is the generated key local feature information.
4. The method for realizing transferable targeted attack through multi-branch perturbation optimization according to claim 1, characterized in that: S3 specifically includes the following steps: S3.1 The optimization process of the source model adversarial sample is carried out in the deep deep feature fusion module; during the optimization process, first generate a random number within the range of 0.0 to 1.0, and if this random number is less than the preset probability value, enter the deep feature fusion module for subsequent feature optimization operations; S3.2 Shuffle and reorganize the current features, global features, and key local features of the same batch to obtain the features as the input of the next network layer; In the network layer that meets the optimization conditions, a set of random fusion weights is generated, and the three features are weighted and fused according to the corresponding weights according to the probability to obtain the features as the input of the next network layer; after the optimization of the current network layer is completed, continue to enter the next layer for optimization until the perturbation matrix p is finally generated a .
5. The method for realizing transferable targeted attacks through multi-branch perturbation optimization according to claim 4, characterized in that: In S3.2, the calculation method for the fusion and optimization of the current features, global features, and key local features is as follows: f si c = S(f i c ) f si l = S(f i l ) f i ′ = α i f i +(1 - α i )f i p Among them, S(·) is a random shuffling function, i refers to the current layer of the surrogate model, and the original image feature f i c is shuffled into f si c . Similarly, the key local feature f i l is randomly shuffled into f si l , and β i is the weight parameter of the key local feature; f i p is the fusion result of f si c and f si l . f i is the output feature of the previous layer, and α i is the weight of the linear interpolation fusion of f i and f i p .
6. The method for realizing transferable directed attack through multi-branch perturbation optimization according to claim 1, characterized in that: S4 specifically includes the following steps: S4.1 Before each iteration, input the current sample into the multi-scale feature extraction module to obtain the sample features at different scales; The method for generating multi-scale sample features according to the current sample is as follows: x t Ti = D(x t , N) where x t is the iterative sample at the current moment, N is the predefined downsampling step, D(·) is the downsampling function, and x t Ti is to generate sample features of different scales; S4.2 Input the sample features at different scales into the classifier, and calculate a one-step perturbation result according to the reverse gradient to generate multi-scale perturbations.
7. The method for realizing transferable targeted attacks through multi-branch perturbation optimization according to claim 1, characterized in that: S5 specifically includes the following steps: S5.1 First, upsample the perturbation results at different scales to the original sample size; S5.2 Calculate the contribution degree of the current scale perturbation to the directional attack loss according to the perturbations at different scales and the downsampled samples; S5.3 Calculate the weight corresponding to each perturbation adaptively according to the perturbation contributions at all different scales, and finally calculate the weighted sum to obtain the multi-scale perturbation p b .
8. The method for realizing transferable directed attack through multi-branch perturbation optimization according to claim 7, characterized in that: In S5.1, the method for upsampling the perturbation results at different scales to the original sample size is as follows: p i u = U(p i , x) Among them, p i is the one-step perturbation result of generating samples of different scales, U(·) is the upsampling function, and p i u is the perturbation matrix consistent with the size of the original sample.
9. The method for realizing transferable targeted attack through multi-branch perturbation optimization according to claim 7, characterized in that: In S5.2, the method for calculating the contribution degree of the current scale perturbation to the directional attack loss is as follows: L i = Logit(x i Ti + p i u , y') Among them, Logit(·) represents the Logit loss function, y' is the target class, and L i is for the current scale perturbation to the targeted attack loss.
10. The method for realizing transferable targeted attack through multi-branch perturbation optimization according to claim 7, characterized in that: In S5.3, the adaptive fusion strategy is as follows: p b = ∑p i u × ω i Among them, ω i is the weight parameter of the perturbation matrix at each scale, and p b is the result of the final multi-scale perturbation fusion.