Adversarial sample generation method based on DBINet and SMDANet

Through the salient feature detection and diffusion model combined with DBINet and SMDANet, the visual quality and attack effect problems of adversarial sample generation in the black box model are solved, and efficient and covert adversarial sample generation is achieved.

CN120747541APending Publication Date: 2025-10-03GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585344.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing adversarial sample generation methods rely on gradient information, making it difficult to generate adversarial samples with high visual quality and good attack effects in black-box models. In addition, the generation time is long and the image is easily distorted.

Method used

The method of combining DBINet and SMDANet is adopted to generate adversarial samples using salient feature detection and diffusion model. The key areas are quickly located through salient feature detection and the diffusion model is used to modify the image to generate high-quality adversarial samples.

Benefits of technology

Without relying on gradient information, efficient, high-visual-quality and concealed adversarial samples are generated, which expands the application scenarios and improves the accuracy and efficiency of attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747541A_ABST
    Figure CN120747541A_ABST
Patent Text Reader

Abstract

The invention discloses an adversarial sample generation method based on DBINet and SMDANet. The method comprises the following three steps: 1) generating a class activation thermodynamic diagram for a target model by using a self-designed double-branch saliency detection model DBINet, and generating an accurate weight distribution matrix in combination with an adaptive threshold processing technology so as to accurately identify a feature sensitive area in an image; 2) encoding an input image to be attacked into a submerged space for representation; and 3) in the denoising process of the semantic mask guided diffusion model SMDANet, strictly limiting the range of image modification by using a mask mechanism to ensure that adversarial disturbance only acts on a key feature region of the image, and meanwhile, applying adversarial condition constraints according to a target category guide vector in the later period of the whole process. Compared with a traditional method, the method has the advantages that the feature sensitive area is accurately positioned in a self-adaptive mode, the disturbance amplitude is effectively reduced, the attack effectiveness is maintained, meanwhile, the visual concealment of the adversarial sample is greatly improved, and the generation efficiency is improved. The innovation provides a brand-new and efficient evaluation tool for the safety test of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security, and specifically relates to a method for generating adversarial samples based on salient feature detection and diffusion model image editing technology. Background Art

[0002] Adversarial attacks, a crucial and challenging research area in machine learning, rely on carefully crafted, subtle perturbations to cause machine learning models to produce incorrect predictions when receiving these tampered input data. This attack poses a serious threat to model performance and reliability, particularly in critical application areas such as autonomous driving and military defense, where adversarial attacks can have catastrophic consequences and cause immeasurable losses. Therefore, in-depth research on how to effectively generate adversarial attack samples not only helps us better understand the workings of deep learning models but also provides a crucial basis for developing effective defense strategies against such threats.

[0003] Traditional adversarial attack methods rely heavily on gradient information, constructing perturbations by calculating the gradient of the input data with respect to the model's predictions. While this approach is theoretically highly effective, it suffers from significant limitations in practice. When working with black-box models or scenarios where gradient information is inaccessible, traditional gradient-based attack methods often struggle due to a lack of understanding of the target model's internal structure and parameters, significantly reducing their effectiveness. While current generative methods based on generative adversarial networks (GANs) and variational autoencoders (VAEs) can generate adversarial examples to a certain extent, they often suffer from severe image distortion. This results in adversarial examples being visually unnatural and easily discernible to human observers, thus reducing their practicality and stealth. How to efficiently and diversely generate adversarial examples with high visual quality and effective attack effectiveness without relying on gradient information has become a key issue that needs to be addressed in current adversarial attack research. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies and provide an adversarial sample generation method based on DBINet and SMDANet. This method uses salient feature information for prior guidance and a diffusion model to generate more realistic adversarial samples, thereby improving the stealth of attacks.

[0005] The technical solution for achieving the purpose of the present invention is:

[0006] A method for generating adversarial samples based on DBINet and SMDANet, characterized by comprising the following steps:

[0007] 1) Generate salient feature weights for the target model using the self-designed dual-branch saliency detection model DBINet, including the following steps:

[0008] Step 1.1: Build the DBINet training dataset based on the HKU-IS and ECSSD datasets. After preprocessing the image data, build the training set and test set. The specific process is as follows:

[0009] Step 1.1.1 Build the DBINet training dataset based on the HKU-IS and ECSSD datasets, with a total of 5447 pairs of labeled image data.

[0010] Step 1.1.2 processes the image data collected in step 1.1.1 into a uniform size of 224×224.

[0011] Step 1.1.3 divides the image data processed in step 1.1.2 into two groups, 1000 pairs of which are used as the test set and the remaining 4447 pairs as the training set.

[0012] Step 1.2: Based on the dataset created in step 1.1, train DBINet and obtain the trained parameters. The specific process is as follows:

[0013] Step 1.2.1 Load the training set constructed in step 1 to train the model.

[0014] Step 1.2.2 During the training process, the model output is calculated through forward propagation, and the loss value is calculated based on the true label. The loss function L total The formula is shown in Formula 3, which is divided into two parts:

[0015]

[0016] L total =L wbce +L iou (3)

[0017] Where N represents the number of pixels, y represents the label value of the corresponding pixel position in the image, α and β represent different weights, their sum is 1, and p represents the predicted value of the corresponding position.

[0018] Step 1.2.3: Update the model parameters using the backpropagation algorithm based on the loss value calculated in step 1.2.2, and record the training loss and performance indicators after each training batch.

[0019] In step 1.2.4, after each training round, use the test set to evaluate the model performance, monitor for overfitting, and finally save the parameters at the end of training.

[0020] Step 1.3 processes the original RGB input image into a 224×224 image. After DBINet loads the parameters obtained in step 1.2, the image data is simultaneously input into the global semantic encoding branch and the local detail encoding branch.

[0021] Step 1.4: In the global semantic branch, the image processed in step 1.3 is added with position encoding through the PEPM module to obtain a Token vector. The specific process is as follows:

[0022] Step 1.4.1: Cut the image input in step 1.3 into blocks through 4×4 convolution.

[0023] Step 1.4.2 flattens the block of step 1.4.1 and then adds the position code to obtain the final token.

[0024] Step 1.5: Input the token vector obtained in step 1.4 into the MSTB module, gradually extract hierarchical features, and obtain the global semantic feature map. The specific process is as follows:

[0025] Step 1.5.1 Input the Token vector obtained in step 1.4 into the three-stage Transformer module.

[0026] In step 1.5.2, a 3×3 convolution is used for downsampling between every two Transformer blocks.

[0027] Step 1.5.3 finally projects the features through a 1×1 convolution and outputs a feature map of size 256×14×14.

[0028] Step 1.6 In the local detail encoding branch, the image processed in step 1.3 is downsampled through the FDM module to obtain a smaller feature map. The specific process is as follows:

[0029] In step 1.6.1, the underlying features of size 64×224×224 are extracted through 3×3 convolution and batch normalized, and then the GELU activation function is used.

[0030] Step 1.6.2 performs progressive downsampling via three levels of depthwise separable convolution, where each level consists of:

[0031] a) 3×3 convolution with 1-pixel padding and 2-pixel stride for downsampling.

[0032] b) 1×1 convolution.

[0033] c) Batch Normalization.

[0034] d)GELU activation function.

[0035] Step 1.6.3 performs final downsampling through 4×4 convolution and outputs a feature map of size 512×14×14.

[0036] Step 1.7: Input the feature map obtained in step 1.6 into the MFEM module to perform multi-scale feature fusion and output a local detail feature map. The specific process is as follows:

[0037] Step 1.7.1 uses the multi-scale fusion feature module in the MFEM module to extract multi-scale features. Each layer contains a 3×3 regular convolution and three 3×3 dilated convolutions, with the corresponding receptive fields set to 7×7, 9×9, and 11×11 respectively. The four convolutions are connected through channel splicing 1×1 convolution and channel attention module.

[0038] In step 1.7.2, the multi-scale features are concatenated and fused through 1×1 convolution to output a feature map of size 256×14×14.

[0039] In step 1.7.3, the feature map obtained in step 1.7.2 is input into the spatial attention module to perform spatial domain reweighting on the feature map, highlight the responses of edge and texture areas, and output a feature map of size 256×14×14.

[0040] Step 1.8: Input the feature maps obtained in step 1.5 and step 1.7 into the FAM module for feature alignment, and output the aligned feature map of the global semantic feature map. The specific process is as follows:

[0041] Step 1.8.1 performs channel-domain adaptive weighting on the global feature map:

[0042] a) Compress the spatial dimension of the global semantic feature map obtained in step 1.5 by global average pooling with an output of 1.

[0043] b) After dimensionality reduction through 1×1 convolution, the network is input into the 1×1 convolution layer again through the ReLU activation function to generate 256 channel weights.

[0044] c) Sigmoid normalization is then used to multiply the original features channel by channel to obtain the true weight of each channel.

[0045] Step 1.8.2 applies the true weight of each channel obtained in step 1.8.1 to the global semantic feature map obtained in step 1.5 to obtain a global alignment feature map of size 256×14×14.

[0046] In step 1.8.3, the local detail feature map obtained in step 1.7 is subjected to 3×3 convolution and spatial domain alignment to obtain a local alignment feature map of size 256×14×14.

[0047] Step 1.8.4 adds the feature maps obtained in steps 1.8.2 and 1.8.3 element by element, and outputs an aligned feature map of size 256×14×14.

[0048] Step 1.9: Input the alignment feature map obtained in step 1.8 and the local detail feature map obtained in step 1.7 into the dynamic gate fusion module, fuse the feature maps, and output the fused feature map. The specific process is as follows:

[0049] Step 1.9.1: Concatenate the local detail feature map and alignment feature map generated in step 1.8 by channel.

[0050] In step 1.9.2, the feature map generated in step 1.9.1 is reduced in dimension through a 1×1 convolution, and then input into the 1×1 convolution layer again through the ReLU activation function to generate 2-channel weights.

[0051] Step 1.9.3: Pass the weights obtained in step 1.9.2 through the Softmax function to ensure that the sum of the weights is 1.

[0052] Step 1.9.4 performs weighted summation on the two feature maps according to the fusion weights obtained in step 1.9.3, and outputs a fused feature map of size 256×14×14.

[0053] Step 1.10: Input the fused feature map obtained in step 1.9 into the FRU module, align and refine it, enhance the quality of the fused features, and output the refined feature map. The specific process is as follows:

[0054] Step 1.10.1 reorganizes the local features of the fused feature map obtained in step 1.9 through 3×3 convolution.

[0055] Step 1.10.2 then uses group normalization to divide the 256 channels into 8 groups for independent normalization.

[0056] Step 1.10.3 then uses the GELU activation function to enhance the nonlinear expression capability.

[0057] Step 1.10.4 finally outputs the refined feature map with an output size of 256×14×14.

[0058] Step 1.11 uses the channel attention module to process the refined feature map obtained in step 1.10, dynamically calculates the importance weight of each channel, and outputs the corresponding feature map. The specific process is as follows:

[0059] Step 1.11.1 performs global average pooling on the input refined feature map to compress the 256×14×14 feature map into a 256×1×1 channel descriptor.

[0060] Step 1.11.2 generates channel weights through a two-level fully connected layer channel attention module, where:

[0061] Step 1.11.2.1 First layer: linear transformation (from 256 to 32 dimensions) followed by ReLU activation function.

[0062] Step 1.11.2.2 Second layer: Linear transformation (from 32 to 256 dimensions) followed by Sigmoid activation function.

[0063] Step 1.11.3 multiplies the generated 256-dimensional channel weights with the original feature map channel by channel, and outputs the enhanced 256×14×14 feature map.

[0064] Step 1.12 uses the decoder to decode the feature map obtained in step 1.11, normalizes it, and finally outputs a single-channel saliency feature weight map. The specific process is as follows:

[0065] Step 1.12.1 performs feature decoding after four levels of upsampling, outputting a 16×224×224 feature map. Each level includes:

[0066] a) Bilinear interpolation upsampling doubles the resolution.

[0067] b) 3×3 convolution with 1 pixel padding to maintain the size.

[0068] c) Batch Normalization.

[0069] d)GELU activation function.

[0070] In step 1.12.2, 1×1 convolution is used to compress the 16-channel features into a single channel, and then a sigmoid function is used to output a single-channel saliency feature weight map of size 1×224×224.

[0071] 2) Based on the adaptive threshold, the single-channel saliency feature weight map obtained in 1) is binarized to generate a mask, where the weight of the outer area of ​​the mask is set to 0, and the weight of the inner area is retained as the corresponding weight in the heat map, as shown in formula (4):

[0072]

[0073] Where M(x, y) is the generated mask, H(x, y) is the saliency weight map generated by 2), x and y represent the positions of corresponding pixels, and T represents the adaptive threshold.

[0074] Step 2.1 The adaptive threshold will take into account the size of the mask generated under the current threshold condition. According to the results of multiple experiments, the effect is better when the mask coverage area is less than 10% of the original image size, as shown in formula (5):

[0075] P(H(x, y)≥T)≤0.1*image area (5)

[0076] Where P(·) represents the number of pixels that currently meet the conditions, H(x, y) is the saliency weight map generated in step 2), x and y represent the positions of the corresponding pixels, and T represents the adaptive threshold.

[0077] 3) The input image is processed into a size of 224×224, encoded and noise is added to obtain a noisy image sequence of length 30, including:

[0078] Step 3.1 preprocesses the input data, processes the image into 224×224 size, and calculates the embedding vector using the original label and the attack target label respectively.

[0079] In step 3.2, the preprocessed image is gradually denoised using a self-designed denoising diffusion implicit model.

[0080] Step 3.3 During the noise addition process, record the noised images generated at each step to obtain a sequence Image containing 30 noised images. [0,1,2,...,29] .

[0081] 4) Combine the image obtained in 3) and the mask obtained in 2) and perform denoising to generate an intermediate image of the adversarial sample, including:

[0082] Step 4.1: The last image in the noise image sequence Image 29 As the initial image of the denoising process, the embedding vector of the original label is used for guidance to generate the first denoised image A 29 .

[0083] Step 4.2 In each step i (i=29,28,...,0) of the denoising process, the denoised image A generated in the previous step is generated according to the guide word i , use the diffusion model denoiser to generate a new denoised image A i-1 , where the first 20 steps use the embedding vector generated by the original label for guidance, and the last 10 steps use the embedding vector generated by the attack target label for guidance. The guidance process is the same as shown in formula (6).

[0084]

[0085] Among them, A i is the denoised image generated in the previous step, f θ (·,·) is the diffusion model, embedding 原始 The embedding vector of the original label of the image generated in step 3.1, embedding 目标The embedding vector of the attack target label generated in step 3.1.

[0086] Step 4.3 At each step i (i = 29, 28, ..., 0) in the denoising process, the denoised image A generated in step 4.2 is converted according to the weights in the mask generated in step 2). i-1 The noisy image Image generated in step 3.3 i Synthesize and get a new image as the input A for the next step of denoising i , the calculation method of image synthesis is shown in formula (7):

[0087] A i-) =M⊙A i +(1-M)⊙Image i (7)

[0088] Among them, A i The denoised image generated in step 4.2, Image i is the noisy image generated in step 3.2, M is the mask generated in step 2.2, and ⊙ represents the pixel-by-pixel multiplication operation.

[0089] 5) Use the decoder to decode the final denoised image A0 to obtain the adversarial sample.

[0090] This technical solution mainly solves the problems existing in existing adversarial sample generation methods, such as the need to obtain internal information of the target model, poor quality of the generated image, and too long attack time. In response to these problems, this technical solution provides an adversarial sample generation method based on DBINet and SMDANet. The innovation of this method is that it abandons the traditional method's reliance on gradient information and adopts a semantic-guided approach to achieve efficient attacks on black-box models. This design enables the present invention to generate highly aggressive adversarial samples even when it is unable to obtain internal information of the target model, greatly expanding the application scenarios of adversarial attacks. In response to the problem of long attack time in existing generation methods, this technical solution introduces significant feature detection technology. Significant feature detection can quickly locate the key areas in the image that have the greatest impact on model decision-making, thereby focusing the attack on these areas, avoiding the high computational overhead of global optimization of the entire image in traditional methods, and improving the accuracy and efficiency of the attack. To address the problem that existing generation methods are prone to image distortion, this technical solution introduces a diffusion model to modify the image. This improvement not only solves the problem that images generated by traditional methods are severely distorted and easily recognized by human observers, but also provides higher visual quality assurance for the generation of adversarial samples, significantly improving the concealment and practicality of adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1An overall diagram of the adversarial sample generation process.

[0092] Figure 2 Schematic diagram of the mask generation process.

[0093] Figure 3 Schematic diagram of the principle of generating saliency feature maps for DBINet.

[0094] Figure 4 This is a schematic diagram of the FDM module principle.

[0095] Figure 5 Schematic diagram of the MFEM module principle.

[0096] Figure 6 Schematic diagram of the global semantic encoder principle.

[0097] Figure 7 Schematic diagram of the FAM module principle.

[0098] Figure 8 This is a schematic diagram of the principles of dynamic gating fusion and FRU module.

[0099] Figure 9 Schematic diagram of the saliency map decoding module.

[0100] Figure 10 Schematic diagram of the principle of the noisy image sequence generation process.

[0101] Figure 11 Schematic diagram of the principle of generating adversarial samples based on SMDANet. DETAILED DESCRIPTION

[0102] In order to more clearly illustrate the purpose, technical solutions and advantages of the present invention, a detailed description will be given below with reference to specific examples.

[0103] A method for generating adversarial samples based on DBINet and SMDANet, such as Figure 1 As shown, the specific steps are as follows:

[0104] (1) Preprocess the data.

[0105] Step 1: Preprocess the image and crop it to 224×224.

[0106] Step 2: Use the original label and the attack target label to calculate the embedding vector respectively. The calculated embedding vector shape is 2×77×1024.

[0107] (2) Use DBINet to generate a saliency map. The generation process is as follows: Figure 3 As shown, the specific structure of DBINet can be found in Figures 4 to 9 .

[0108] Reasoning stage:

[0109] Step 1: Import the trained parameters of the DBINet model.

[0110] Step 2: Input the image with the shape of 224×224 into the global semantic encoder of the DBINet model. The specific structure is shown in Table 1 and Figure 6 The specific process is:

[0111] The input image is cut into blocks through a convolution with a convolution kernel of 4×4, 96 output channels, and a stride of 4. The blocks are then flattened and positional encoding is added to obtain the final Token. The resulting Token vector is then input into a three-stage Transformer block. Between each two Transformer blocks, a convolution kernel of 3×3, a stride of 2, and a padding of 1 is used to downsample the output and gradually extract hierarchical features. The hierarchical features are then reduced in dimension through a convolution operation with a convolution kernel of 1×1 and an output channel of 256 to obtain a global semantic feature map of 256×14×14.

[0112] Step 3: Input the image with the shape of 224×224 into the local detail encoder of the DBINet model. The specific structure is shown in Table 2 and Figure 4 and Figure 5 The specific process is:

[0113] The input image is convolved with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, and then the underlying features of 64×224×224 are extracted through batch normalization and GELU activation function; then progressive downsampling is performed through three levels of depth-wise separable convolution, each level including: downsampling through convolution kernel size of 3×3, padding of 1, and stride of 2; channel number adjustment through convolution kernel size of 1×1; batch normalization and GELU activation function are used for regularization and enhancement of nonlinear expression ability; finally, a feature map of 512×28×28 is obtained, and then a convolution operation with a convolution kernel size of 4×4, a stride of 2, and a padding of 1 is used for final downsampling, outputting a feature map of 512×14×14.

[0114] The feature map is then input into the three-stage multi-scale feature fusion block. Each multi-scale feature fusion block consists of four branches:

[0115] 1) Convolution with a kernel size of 3×3, a stride of 1, and padding of 1.

[0116] 2) Convolution with kernel size of 3×3, stride of 1, padding of 2, and expansion of 2.

[0117] 3) Convolution with kernel size of 3×3, stride of 1, padding of 3, and expansion of 3.

[0118] 4) Convolution with kernel size of 3×3, stride of 1, padding of 4, and expansion of 4.

[0119] The feature maps are input into these four branches respectively, and then the four feature maps are spliced ​​by channel, fused using 1×1 convolution, and finally sent to the channel attention module to output a feature map of size 512×14×14.

[0120] Then, a 1×1 convolution is used to adjust the feature map obtained in the previous stage to obtain a feature map of size 256×14×14. The feature map is input into the spatial attention module to perform spatial domain reweighting on the feature map, highlighting the response of the edge and texture areas, and outputting a local detail feature map of size 256×14×14.

[0121] Step 4: Input the global semantic feature map and local detail feature map obtained in the first two steps into the multi-level feature fusion module to obtain the refined feature map. The specific structure is shown in Table 3 and Figure 7 and Figure 8 The specific process is:

[0122] The global semantic feature map is compressed by global average pooling with an output of 1 in the spatial dimension. After dimensionality reduction through 1×1 convolution, it is input into 1×1 convolution again through ReLU activation function to generate 256 channel weights. Then, Sigmoid normalization is used and multiplied with the original features channel by channel to obtain the true weight of each channel. The obtained true weight of each channel is applied to the global semantic feature map to obtain a global alignment feature map of size 256×14×14. The local detail feature map is subjected to 3×3 convolution and spatial domain alignment to obtain a local alignment feature map of size 256×14×14. Finally, the two feature maps are added element by element to output an alignment feature map of size 256×14×14.

[0123] The local detail feature map and the alignment feature map are spliced ​​by channel, and then the dimensionality is reduced by 1×1 convolution. After the ReLU activation function, it is input into the 1×1 convolution layer again to generate 2-channel weights. The Softmax function is used to ensure that the sum of the weights is 1. Finally, the obtained fusion weight is used to perform weighted summation on the two feature maps, and the output is a fused feature map with a size of 256×14×14.

[0124] The fused feature map is subjected to local feature reorganization through convolution with a convolution kernel size of 3×3 and a padding of 1. Group normalization is used to divide the 256 channels into 8 groups for independent normalization. The nonlinear expression ability is then enhanced through the GELU activation function. Finally, the refined feature map with an output size of 256×14×14 is output.

[0125] Step 5: Use the saliency map decoding module to decode the refined feature map obtained in the previous step and output the saliency feature weight map predicted by the model. The specific structure is shown in Table 4 and Figure 9 As shown, the specific process is:

[0126] Use the channel attention module to process the refined feature map, dynamically calculate the importance weight of each channel, output the corresponding feature map, and use the upsampling block to enlarge the size of the feature map. The upsampling block includes:

[0127] 1) Bilinear interpolation upsampling to double the resolution.

[0128] 2) Convolution with a kernel size of 3×3 and padding of 1.

[0129] 3) Batch normalization.

[0130] 4) GELU activation function.

[0131] The feature map is decoded through a four-level upsampling block, and finally a 1×1 convolution is used to compress the 16-channel features into a single channel. The Sigmoid function is used to output a single-channel saliency weight map of size 1×224×224.

[0132] Training phase:

[0133] Step 1: Build the DBINet training dataset based on the HKU-IS and ECSSD datasets, with a total of 5447 pairs of labeled image data.

[0134] Step 2 processes the image data collected in step 1 into a uniform size of 224×224.

[0135] Step 3 divides the image data processed in step 2 into two groups, 1000 pairs of which are used as test sets and the remaining 4447 pairs as training sets.

[0136] Step 4: Based on the data set generated in step 3, DBINet is trained to obtain the trained parameters. The specific process is as follows:

[0137] Step 4.1 Input the image in the training set into DBINet, calculate the model output through forward propagation, obtain the corresponding single-channel saliency feature map, and then calculate the loss value based on the true label. The loss function L tootal The formula is shown in Formula 3, which is divided into two parts:

[0138]

[0139] L total =L wbce +Liou (3)

[0140] Where N represents the number of pixels, y represents the label value of the corresponding pixel position in the image, and p represents the predicted value of the corresponding position.

[0141] Step 4.2 uses the backpropagation algorithm to update the model parameters based on the loss value calculated in step 4.1, and records the training loss and performance indicators after each training batch.

[0142] Step 4.3: After each training round, use the test set to evaluate the model performance, monitor overfitting, and finally save the parameters when training is completed.

[0143] (3) Generate a weighted mask based on the saliency map. See Figure 2 .

[0144] Step 1: Initialize the threshold T to 0.5, calculate the area of ​​pixels larger than the threshold, and use an adaptive algorithm to ensure that the area of ​​pixels that meet the conditions is less than 10% of the total area.

[0145] Step 3: Set the weights less than or equal to the threshold to 0, and keep the weights greater than the threshold unchanged to generate a mask.

[0146] (4) Obtain a noisy image sequence, see Figure 10 .

[0147] Step 1: Add noise to the processed image. The whole process will generate thirty noisy images at different time steps.

[0148] Step 2: Save these noisy images and pass them to the next step.

[0149] (5) Generate adversarial samples based on SMDANet, see Figure 11 .

[0150] Step 1: Take the last image in the noise image sequence Image 29 As the initial image of the denoising process, the embedding vector of the original label obtained in step 1 is used for guidance to generate the first denoised image A 29 .

[0151] Step 2: At each step i (i=29,28,...,0) in the denoising process, the denoised image A generated in the previous step is generated based on the guide word i , use the diffusion model denoiser to generate a new denoised image A i-1 , where the first 2 / 3 steps are guided by the embedding vector generated by the original label, and the last 1 / 3 steps are guided by the embedding vector generated by the attack target label.

[0152] Step 3: At each step i (i=29,28,...,0) in the denoising process, the denoised image A generated in the previous step is converted according to the weights in the mask. i Image at the corresponding time step in the noisy image sequence i Synthesize and get a new image as the input A for the next step of denoising i-1 .

[0153] Step 4: When the denoising step is completed, the obtained image is decoded, and the decoded image is the adversarial sample.

[0154] It should be noted that although the above embodiments describe the specific embodiments of the present invention in detail, this is not intended to limit the present invention. The scope of protection of the present invention is not limited to the above specific embodiments. On the basis of following the core principles of the present invention, any other embodiments derived from the concepts of the present invention should be deemed to fall within the scope of protection of the present invention.

[0155]

[0156]

[0157] Table 1: Global semantic encoder structure

[0158]

[0159]

[0160] Table 2: Local detail encoder structure

[0161]

[0162]

[0163] Table 3: Multi-level feature fusion module structure

[0164]

[0165] Table 4: Saliency map decoding module structure.

Claims

1. A method for generating adversarial samples based on DBINet and SMDANet, characterized in that: The steps include: 1) Generate salient feature weights for the target model using the self-designed dual-branch saliency detection model DBINet, including the following steps: Step 1.1: Build the DBINet training dataset based on the HKU-IS and ECSSD datasets. After preprocessing the image data, build the training set and test set. The specific process is as follows: Step 1.1.1 Build the DBINet training dataset based on the HKU-IS and ECSSD datasets. A total of 5447 pairs of labeled image data. Step 1.1.2 processes the image data collected in step 1.1.1 into a uniform size of 224×224. Step 1.1.3 divides the image data processed in step 1.1.2 into two groups, 1000 pairs of which are used as the test set and the remaining 4447 pairs as the training set. Step 1.2: Based on the dataset created in step 1.1, train DBINet and obtain the trained parameters. The specific process is as follows: Step 1.2.1 Load the training set constructed in step 1 to train the model. Step 1.2.2 During the training process, the model output is calculated through forward propagation, and the loss value is calculated based on the true label. The loss function L total The formula is shown in Formula 3, which is divided into two parts: L total =L wbce +L iou (3) Where N represents the number of pixels, y represents the label value of the corresponding pixel position in the image, α and β represent different weights, their sum is 1, and p represents the predicted value of the corresponding position. Step 1.2.3: Update the model parameters using the backpropagation algorithm based on the loss value calculated in step 1.2.2, and record the training loss and performance indicators after each training batch. In step 1.2.4, after each training round, use the test set to evaluate the model performance, monitor for overfitting, and finally save the parameters at the end of training. Step 1.3 processes the original RGB input image into a 224×224 image. After DBINet loads the parameters obtained in step 1.2, the image data is simultaneously input into the global semantic encoding branch and the local detail encoding branch. Step 1.4: In the global semantic branch, the image processed in step 1.3 is added with position encoding through the PEPM module to obtain a Token vector. The specific process is as follows: Step 1.4.1: Cut the image input in step 1.3 into blocks through 4×4 convolution. Step 1.4.2 flattens the block of step 1.4.1 and then adds the position code to obtain the final token. Step 1.5: Input the token vector obtained in step 1.4 into the MSTB module, gradually extract hierarchical features, and obtain the global semantic feature map. The specific process is as follows: Step 1.5.1 Input the Token vector obtained in step 1.4 into the three-stage Transformer module. In step 1.5.2, a 3×3 convolution is used for downsampling between every two Transformer blocks. Step 1.5.3 finally projects the features through a 1×1 convolution and outputs a feature map of size 256×14×14. Step 1.6 In the local detail encoding branch, the image processed in step 1.3 is downsampled through the FDM module to obtain a smaller feature map. The specific process is as follows: In step 1.6.1, the underlying features of size 64×224×224 are extracted through 3×3 convolution and batch normalized, and then the GELU activation function is used. Step 1.6.2 performs progressive downsampling via three levels of depthwise separable convolution, where each level consists of: a) 3×3 convolution with 1-pixel padding and 2-pixel stride for downsampling. b) 1×1 convolution. c) Batch Normalization. d)GELU activation function. Step 1.6.3 performs final downsampling through 4×4 convolution and outputs a feature map of size 512×14×14. Step 1.7: Input the feature map obtained in step 1.6 into the MFEM module to perform multi-scale feature fusion and output a local detail feature map. The specific process is as follows: Step 1.7.1 uses the multi-scale fusion feature module in the MFEM module to extract multi-scale features. Each layer contains a 3×3 regular convolution and three 3×3 dilated convolutions, with the corresponding receptive fields set to 7×7, 9×9, and 11×11 respectively. The four convolutions are connected through channel splicing 1×1 convolution and channel attention module. In step 1.7.2, the multi-scale features are concatenated and fused through 1×1 convolution to output a feature map of size 256×14×14. In step 1.7.3, the feature map obtained in step 1.7.2 is input into the spatial attention module to perform spatial domain reweighting on the feature map, highlight the responses of edge and texture areas, and output a feature map of size 256×14×14. Step 1.8: Input the feature maps obtained in step 1.5 and step 1.7 into the FAM module for feature alignment, and output the aligned feature map of the global semantic feature map. The specific process is as follows: Step 1.8.1 performs channel-domain adaptive weighting on the global feature map: a) Compress the spatial dimension of the global semantic feature map obtained in step 1.5 by global average pooling with an output of 1. b) After dimensionality reduction through 1×1 convolution, the network is input into the 1×1 convolution layer again through the ReLU activation function to generate 256 channel weights. c) Sigmoid normalization is then used to multiply the original features channel by channel to obtain the true weight of each channel. Step 1.8.2 applies the true weight of each channel obtained in step 1.8.1 to the global semantic feature map obtained in step 1.5 to obtain a global alignment feature map of size 256×14×14. In step 1.8.3, the local detail feature map obtained in step 1.7 is subjected to 3×3 convolution and spatial domain alignment to obtain a local alignment feature map of size 256×14×14. Step 1.8.4 adds the feature maps obtained in steps 1.8.2 and 1.8.3 element by element, and outputs an aligned feature map of size 256×14×14. Step 1.9: Input the alignment feature map obtained in step 1.8 and the local detail feature map obtained in step 1.7 into the dynamic gate fusion module, fuse the feature maps, and output the fused feature map. The specific process is as follows: Step 1.9.1: Concatenate the local detail feature map and alignment feature map generated in step 1.8 by channel. In step 1.9.2, the feature map generated in step 1.9.1 is reduced in dimension through a 1×1 convolution, and then input into the 1×1 convolution layer again through the ReLU activation function to generate 2-channel weights. Step 1.9.3: Pass the weights obtained in step 1.9.2 through the Softmax function to ensure that the sum of the weights is 1. Step 1.9.4 performs weighted summation on the two feature maps according to the fusion weights obtained in step 1.9.3, and outputs a fused feature map of size 256×14×14. Step 1.10: Input the fused feature map obtained in step 1.9 into the FRU module, align and refine it, enhance the quality of the fused features, and output the refined feature map. The specific process is as follows: Step 1.10.1 reorganizes the local features of the fused feature map obtained in step 1.9 through 3×3 convolution. Step 1.10.2 then uses group normalization to divide the 256 channels into 8 groups for independent normalization. Step 1.10.3 then uses the GELU activation function to enhance the nonlinear expression capability. Step 1.10.4 finally outputs the refined feature map with an output size of 256×14×14. Step 1.11 uses the channel attention module to process the refined feature map obtained in step 1.10, dynamically calculates the importance weight of each channel, and outputs the corresponding feature map. The specific process is as follows: Step 1.11.1 performs global average pooling on the input refined feature map to compress the 256×14×14 feature map into a 256×1×1 channel descriptor. Step 1.11.2 generates channel weights through a two-level fully connected layer channel attention module, where: Step 1.11.2.1 First layer: linear transformation (from 256 to 32 dimensions) followed by ReLU activation function. Step 1.11.2.2 Second layer: Linear transformation (from 32 to 256 dimensions) followed by Sigmoid activation function. Step 1.11.3 multiplies the generated 256-dimensional channel weights with the original feature map channel by channel, and outputs the enhanced 256×14×14 feature map. Step 1.12 uses the decoder to decode the feature map obtained in step 1.11, normalizes it, and finally outputs a single-channel saliency feature weight map. The specific process is as follows: Step 1.12.1 performs feature decoding after four levels of upsampling, outputting a 16×224×224 feature map. Each level includes: a) Bilinear interpolation upsampling doubles the resolution. b) 3×3 convolution with 1 pixel padding to maintain the size. c) Batch Normalization. d)GELU activation function. In step 1.12.2, 1×1 convolution is used to compress the 16-channel features into a single channel, and then a sigmoid function is used to output a single-channel saliency feature weight map of size 1×224×224. 2) Based on the adaptive threshold, the single-channel saliency feature weight map obtained in 1) is binarized to generate a mask, where the weight of the outer area of ​​the mask is set to 0, and the weight of the inner area is retained as the corresponding weight in the heat map, as shown in formula (4): Where M(x, y) is the generated mask, H(x, y) is the saliency weight map generated by 1), x and y represent the positions of corresponding pixels, and T represents the adaptive threshold. Step 2.1 The adaptive threshold will take into account the size of the mask generated under the current threshold condition. According to the results of multiple experiments, the effect is better when the mask coverage area is less than 10% of the original image size, as shown in formula (5): P(H(x, y)≥T)≤0.1*image area (5) Where P(·) represents the number of pixels that currently meet the conditions, H(x, y) is the saliency weight map generated in step 2), x and y represent the positions of the corresponding pixels, and T represents the adaptive threshold. 3) The input image is processed into a size of 224×224, encoded and noise is added to obtain a noisy image sequence of length 30, including: Step 3.1 preprocesses the input data, processes the image into 224×224 size, and calculates the embedding vector using the original label and the attack target label respectively. In step 3.2, the preprocessed image is gradually denoised using a self-designed denoising diffusion implicit model. Step 3.3 During the noise addition process, record the noised images generated at each step to obtain a sequence Image containing 30 noised images. [0,1,2,...,29] . 4) Combine the image obtained in 3) and the mask obtained in 2) and perform denoising to generate an intermediate image of the adversarial sample, including: Step 4.1: The last image in the noise image sequence Image 29 As the initial image of the denoising process, the embedding vector of the original label is used for guidance to generate the first denoised image A 29 . Step 4.2 In each step i (i=29,28,...,0) of the denoising process, the denoised image A generated in the previous step is generated according to the guide word i , use the diffusion model denoiser to generate a new denoised image A i-1 , where the first 20 steps use the embedding vector generated by the original label for guidance, and the last 10 steps use the embedding vector generated by the attack target label for guidance. The guidance process is the same as shown in formula (6): Among them, A i is the denoised image generated in the previous step, f θ (·,·) is the diffusion model, embedding 原始 The embedding vector of the original label of the image generated in step 3.1, embedding 目标 The embedding vector of the attack target label generated in step 3.

1. Step 4.3 At each step i (i = 29, 28, ..., 0) in the denoising process, the denoised image A generated in step 4.2 is converted according to the weights in the mask generated in step 2). i-1 The noisy image Image generated in step 3.3 i Synthesize and get a new image as the input A for the next step of denoising i , the calculation method of image synthesis is shown in formula (7): A i-) =M⊙A i +(1-M)⊙Image i (7) Among them, A i The denoised image generated in step 4.2, Image i is the noisy image generated in step 3.2, M is the mask generated in step 2.2, and ⊙ represents the pixel-by-pixel multiplication operation. 5) Use the decoder to decode the final denoised image A0 to obtain the adversarial sample.