Multilayer semantic information region selection method, system and equipment and storage medium
By integrating the semantic information of multiple convolutional layers, a sparse semantic mask is generated, which solves the problems of inaccurate localization of small objects and poor perturbation control in existing technologies, and achieves more accurate and effective adversarial attacks.
Patent Information
- Application Number
- CN202511008815.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to accurately pinpoint the critical locations of smaller objects and cannot effectively utilize the discontinuous sparsity of shallow features to control the sparsity of perturbations, resulting in insufficient precision and effectiveness in adversarial attacks.
By employing a multi-layer semantic information region selection method, the semantic information of multiple convolutional layers is integrated to obtain gradient, normalization, upsampling, and binarization masks, generating sparse semantic masks. Only key regions in the image are perturbed, and multiple rounds of iterative optimization are performed in conjunction with the gradient direction.
The generated adversarial examples outperform existing methods in both white-box and black-box scenarios, exhibiting higher attack effectiveness, sparsity, and transferability, more accurate localization, and reduced noise pollution.
Smart Images

Figure CN120953578A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep learning adversarial attack technology, and particularly relates to a method, system, device and storage medium for selecting multi-layer semantic information regions. Background Technology
[0002] The anchorless object detector CenterNet extracts key points of objects using a fully convolutional network model. In addition to other feature information, its structure is simpler and computationally more efficient than previous anchor box object detectors. The convolutional layers in a fully convolutional neural network contain rich semantic information. Shallower layers retain abundant spatial information, leading to higher sparsity and discontinuity in the extracted features. However, deeper layers, with their larger receptive fields, contain more global information, resulting in more continuous and regular geometric features in the extracted features. By extracting features with different characteristics from the convolutional layers, it is possible to better locate the high semantic spatial position of the target identified by the network. If adversarial attacks are concentrated on these pixels, the attacks will become more concentrated. In previous work, using the key region of the last convolutional layer was a trend. This is because the last convolutional layer has the largest receptive field of current convolutional networks, which can better distinguish the semantic position and information of the target. Because these algorithms abandon the use of spatial semantic responses extracted from shallower convolutional layers, they struggle to accurately locate the important positions of smaller objects and cannot utilize the discontinuous and sparse characteristics of shallow features to control the sparsity of perturbations. Summary of the Invention
[0003] The purpose of this invention is to provide a method, system, device, and storage medium for selecting multi-layer semantic information regions, so as to solve the problem that existing technologies have difficulty in accurately locating the important positions of small objects and cannot utilize the discontinuous and sparse characteristics of shallow features to control the sparsity of perturbations.
[0004] The embodiments of this application implement a multi-layer semantic information region selection method, which includes the following steps:
[0005] For all convolutional layers in the set, first obtain the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ;
[0006] G i Normalize to the range [0,1] to obtain
[0007] Will Upsample to the input image size to obtain
[0008] Will The semantic mask m is output through the binary function δ(x). i ;
[0009] m of all obtained convolutional layers i The final mask M′ is obtained by performing element-wise product.
[0010] Optionally, in some embodiments of this application, the set of convolutional layers is:
[0011]
[0012] in
[0013]
[0014] In the formula, L is the set of convolutional layers; M is all the layers in the set of convolutional layers; l i Let i be the i-th convolutional layer containing K channels; K represents the j-th channel of the i-th convolutional layer; K is the number of channels; i is the convolutional layer label; and j is the channel number label.
[0015] Optionally, in some embodiments of this application, the channel is obtained. The gradient is as follows:
[0016]
[0017] In the formula, For channel The gradient of f(x); f(x) is the model function; Let be the j-th channel of the i-th convolutional layer; sum all gradients along the channel direction, expressed as:
[0018]
[0019] In the formula, For channel The gradient of G; i It is the sum of all gradients along the directions from channel 1 to k; K is the number of channels; i is the convolutional layer number; j is the channel number number.
[0020] Optionally, in some embodiments of this application, G is used. i Normalized to the range [0,1], it can be represented as:
[0021]
[0022] In the formula, norm() is the normalization function; G is the sum of all gradients along channels 1 to k after normalization. i It is the sum of all gradients along the directions from channel 1 to k.
[0023] Optionally, in some embodiments of this application, Upsampling to the input image size is represented as:
[0024]
[0025] In the formula, Ψ is the upsampling function; This is the sum of all gradients along channels 1 to k after normalization. This is the sum of gradients restored to the input image size.
[0026] Optionally, in some embodiments of this application, The semantic mask m is output through the binary function δ(x). i As shown in the following formula:
[0027]
[0028] In the formula, δ(x) is a binary function, and x is the independent variable of the function; t α The custom threshold is set to 0.5.
[0029] Optionally, in some embodiments of this application, the m values of all obtained convolutional layers are... i The final mask M′ is obtained by performing element-wise product, as shown in the following equation:
[0030]
[0031] In the formula, M′ is the final mask, m i is the mask for the i-th layer; i is the convolutional layer number; M is all the layers in the convolutional layer set.
[0032] Accordingly, embodiments of this application also provide a semantic centralized adversarial perturbation system, including: including:
[0033] Get G i The module, for all convolutional layers in the set of convolutional layers, first obtains the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ;
[0034] get Module, G i Normalize to the range [0,1] to obtain
[0035] get Module, will Upsample to the input image size to obtain
[0036] Semantic mask m iModule, will The semantic mask m is output through the binary function δ(x). i ;
[0037] The final mask module M′ will obtain m from all convolutional layers. i The final mask M′ is obtained by performing element-wise product.
[0038] Accordingly, embodiments of this application also provide a computer device, including a storage device and a processor, wherein the storage device stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0039] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0040] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0041] The multi-layer semantic information region selection method proposed in this application can identify important parts of a target and its surrounding spatial location. Then, by combining the intermediate results of multiple convolutional layers, it extracts the key semantic locations assigned to the target in the image, generating a semantically concentrated adversarial image. The generated semantically concentrated image contains adversarial attack pixels only in certain regions of the image, achieving more accurate localization compared to ordinary adversarial images.
[0042] Traditional methods often rely on the feature map of the last convolutional layer, which, while possessing strong global semantic capabilities, is insufficient in capturing spatial information of small objects, easily leading to localization errors. This invention integrates the semantic information of multiple convolutional layers (including shallow and deep layers), preserving the spatial details of the shallow layers while combining the global semantics of the deep layers, thereby improving the ability to identify semantically key regions of small objects.
[0043] Shallow convolutional features are typically sparse and discontinuous, yet contain more explicit edge and structural information. This invention employs a multi-layer gradient aggregation and masking mechanism to concentrate perturbations on key regions of the image that have high semantic value but are sparsely distributed. This approach allows perturbations to be applied over a smaller area, improving their concealment and attack efficiency, while avoiding noise pollution caused by perturbations across the entire image.
[0044] This invention designs a complete process from gradient acquisition, normalization, upsampling, binarization masking, and multi-layer mask fusion from convolutional layers, resulting in a semantic mask that can be accurately focused on the target semantic region. This mask, multiplied by a global perturbation, forms a sparse perturbation map, updating only necessary regions with perturbations, thus solving the problem of traditional methods being unable to control the perturbation location and density.
[0045] Within the PGD (Projected Gradient Descent) framework, this invention controls the perturbation application region through masking and performs multi-round iterative optimization in conjunction with gradient direction. This results in generated adversarial examples that possess both attack effectiveness and good sparsity, imperceptibility, and transferability. Experiments show that this method outperforms existing methods (such as DAG, FGSM, etc.) in both white-box and black-box scenarios. Attached Figure Description
[0046] Figure 1 This is a flowchart of the multi-layer semantic information region selection method of the present invention;
[0047] Figure 2 This is a comparison chart of the perceptuality of adversarial examples in this invention;
[0048] Figure 3 This is a visualization of clean and adversarial examples on the MS-COCO dataset of this invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0050] The technical solution of this application is as follows:
[0051] Firstly, please refer to Figure 1 This application provides a method for selecting multi-layer semantic information regions, including:
[0052] S01. For all convolutional layers in the set of convolutional structures, first obtain the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ;
[0053] S02, G i Normalize to the range [0,1] to obtain
[0054] S03, will Upsample to the input image size to obtain
[0055] S04, will The semantic mask m is output through the binary function δ(x). i ;
[0056] S05, obtain m from all convolutional layers i The final mask M′ is obtained by performing element-wise product.
[0057] Traditional methods often rely on the feature map of the last convolutional layer, which, while possessing strong global semantic capabilities, is insufficient in capturing spatial information of small objects, easily leading to localization errors. This invention integrates the semantic information of multiple convolutional layers (including shallow and deep layers), preserving the spatial details of the shallow layers while combining the global semantics of the deep layers, thereby improving the ability to identify semantically key regions of small objects.
[0058] Shallow convolutional features are typically sparse and discontinuous, yet contain more explicit edge and structural information. This invention employs a multi-layer gradient aggregation and masking mechanism to concentrate perturbations on key regions of the image that have high semantic value but are sparsely distributed. This approach allows perturbations to be applied over a smaller area, improving their concealment and attack efficiency, while avoiding noise pollution caused by perturbations across the entire image.
[0059] This invention designs a complete process from gradient acquisition, normalization, upsampling, binarization masking, and multi-layer mask fusion from convolutional layers, resulting in a semantic mask that can be accurately focused on the target semantic region. This mask, multiplied by a global perturbation, forms a sparse perturbation map, updating only necessary regions with perturbations, thus solving the problem of traditional methods being unable to control the perturbation location and density.
[0060] Within the PGD (Projected Gradient Descent) framework, this invention controls the perturbation application region through masking and performs multi-round iterative optimization in conjunction with gradient direction. This results in generated adversarial examples that possess both attack effectiveness and good sparsity, imperceptibility, and transferability. Experiments show that this method outperforms existing methods (such as DAG, FGSM, etc.) in both white-box and black-box scenarios.
[0061] In S01:
[0062] In some embodiments, the set of convolutional layers is:
[0063]
[0064] in
[0065]
[0066] In the formula, L is the set of convolutional layers; M is all the layers in the set of convolutional layers; l i Let i be the i-th convolutional layer containing K channels; K represents the j-th channel of the i-th convolutional layer; K is the number of channels; i is the convolutional layer label; and j is the channel number label.
[0067] In some embodiments, obtain channel The gradient is as follows:
[0068]
[0069] In the formula, For channel The gradient of f(x); f(x) is the model function; Let be the j-th channel of the i-th convolutional layer.
[0070] Furthermore, the gradients are summed along the channel direction, which can be expressed as:
[0071]
[0072] In the formula, For channel The gradient of G; i It is the sum of all gradients along the directions from channel 1 to k; K is the number of channels; i is the convolutional layer number; j is the channel number number.
[0073] In S02:
[0074] In some embodiments, G i Normalized to the range [0,1], it can be represented as:
[0075]
[0076] In the formula, norm() is the normalization function; G is the sum of all gradients along channels 1 to k after normalization. i It is the sum of all gradients along the directions from channel 1 to k.
[0077] In S03:
[0078] In some embodiments, Upsampling to the input image size is represented as:
[0079]
[0080] In the formula, Ψ is the upsampling function; This is the sum of all gradients along channels 1 to k after normalization. This is the sum of gradients restored to the input image size.
[0081] It is understandable that... Upsampling to the input image size aims to restore the obtained gradient to the input image size.
[0082] In S04:
[0083] In some embodiments, The semantic mask m is output through the binary function δ(x). i As shown in the following formula:
[0084]
[0085] In the formula, δ(x) is a binary function, and x is the independent variable of the function; t α The custom threshold is set to 0.5.
[0086] It is understandable that all values greater than t are included. α The output is 1, which is less than or equal to t. α The value of t is output as 0; α ∈(0,1) is a manually set threshold used to control the sparsity of the semantic mask; in this application, it is set to 0.5.
[0087] In S05:
[0088] In some embodiments, the m values of all convolutional layers obtained will be... i The final mask M′ is obtained by performing element-wise product, as shown in the following equation:
[0089]
[0090] In the formula, M′ is the final mask, m i is the mask for the i-th layer; i is the convolutional layer number; M is all the layers in the convolutional layer set.
[0091] After obtaining the final semantic mask M′, the final perturbation needs to be calculated.
[0092] The multi-layer semantic information region selection method proposed in this application can identify important parts of a target and its surrounding spatial location. Then, by combining the intermediate results of multiple convolutional layers, it extracts the key semantic locations assigned to the target in the image, generating a semantically concentrated adversarial image. The generated semantically concentrated image contains adversarial attack pixels only in certain regions of the image, achieving more accurate localization compared to ordinary adversarial images.
[0093] This application reveals the high linearization between the boundaries within the decision hyperplane of a deep neural network model. Research on such adversarial examples can provide an important theoretical foundation and practical support for deep neural networks to resist adversarial examples.
[0094] Using the steps in patent application number 202410879180.9, this application collects gradients in the convolutional structure of patent application number 202410879180.9, taking the target class of interest in a single input image as a unit.
[0095] After obtaining the final semantic mask M′ of this application, the steps in patent application number 202410879180.9 are adopted, and L is minimized. ∞ Project Gradient Decent (PGD) of the norm is used to generate the global perturbation x.(G) Then, the semantic mask M′ and the global perturbation x are obtained using the multi-layer semantic information region selection algorithm of this invention. (G) Multiplying them together yields the final perturbation.
[0096] Secondly, embodiments of this application provide a semantic centralized counter-perturbation system, including:
[0097] Get G i The module, for all convolutional layers in the set of convolutional layers, first obtains the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ;
[0098] get Module, G i Normalize to the range [0,1] to obtain
[0099] get Module, will Upsample to the input image size to obtain
[0100] Semantic mask m i Module, will The semantic mask m is output through the binary function δ(x). i ;
[0101] The final mask module M′ will obtain m from all convolutional layers. i The final mask M′ is obtained by performing element-wise product.
[0102] The acquisition of G i In the module:
[0103] In some embodiments, the set of convolutional layers is:
[0104]
[0105] in
[0106]
[0107] In the formula, L is the set of convolutional layers; M is all the layers in the set of convolutional layers; l i Let i be the i-th convolutional layer containing K channels; K represents the j-th channel of the i-th convolutional layer; K is the number of channels; i is the convolutional layer label; and j is the channel number label.
[0108] In some embodiments, obtain channel The gradient is as follows:
[0109]
[0110] In the formula, For channel The gradient of f(x); f(x) is the model function; Let be the j-th channel of the i-th convolutional layer.
[0111] Furthermore, the gradients are summed along the channel direction, which can be expressed as:
[0112]
[0113] In the formula, For channel The gradient of G; i It is the sum of all gradients along the directions from channel 1 to k; K is the number of channels; i is the convolutional layer number; j is the channel number number.
[0114] The acquisition In the module:
[0115] In some embodiments, G i Normalized to the range [0,1], it can be represented as:
[0116]
[0117] In the formula, norm() is the normalization function; G is the sum of all gradients along channels 1 to k after normalization. i It is the sum of all gradients along the directions from channel 1 to k.
[0118] The acquisition In the module:
[0119] In some embodiments, Upsampling to the input image size is represented as:
[0120]
[0121] In the formula, Ψ is the upsampling function; This is the sum of all gradients along channels 1 to k after normalization. This is the sum of gradients restored to the input image size.
[0122] It is understandable that... Upsampling to the input image size aims to restore the obtained gradient to the input image size.
[0123] The semantic mask m i In the module:
[0124] In some embodiments, The semantic mask m is output through the binary function δ(x). i As shown in the following formula:
[0125]
[0126] In the formula, δ(x) is a binary function, and x is the independent variable of the function; t α The custom threshold is set to 0.5.
[0127] It is understandable that all values greater than t are included. α The output is 1, which is less than or equal to t. α The value of t is output as 0; α ∈(0,1) is a manually set threshold used to control the sparsity of the semantic mask; in this application, it is set to 0.5.
[0128] In the final mask M′ module:
[0129] In some embodiments, the m values of all convolutional layers obtained will be... i The final mask M′ is obtained by performing element-wise product, as shown in the following equation:
[0130]
[0131] In the formula, M′ is the final mask, m i is the mask for the i-th layer; i is the convolutional layer number; M is all the layers in the convolutional layer set.
[0132] After obtaining the final semantic mask M′, the final perturbation needs to be calculated.
[0133] Thirdly, this application provides a computer device including a storage device and a processor, wherein the storage device stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the multi-layer semantic information region selection method described above.
[0134] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0135] The memory includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or 3D interface display memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device. Of course, the memory may include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is often used to store the operating system and various application software installed on the computer device, such as the program code of the multi-layer semantic information region selection method. In addition, the memory can also be used to temporarily store various types of data that have been output or will be output.
[0136] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data, such as running program code for a multi-layer semantic information region selection method.
[0137] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the multi-layer semantic information region selection method described above.
[0138] The computer-readable storage medium stores an interface display program, which can be executed by at least one processor to cause the at least one processor to perform the steps of the multi-layer semantic information region selection method described above.
[0139] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the multi-layer semantic information region selection method described in the embodiments of this application.
[0140] Application examples
[0141] FGSM and PGD are two typical examples of minimizing L ∞ Norm-based attack methods. FGSM is a single-step perturbation attack; the algorithm calculates the gradient of the input image using a loss function and uses the sign of the gradient as the direction and perturbation strength to generate adversarial perturbations. PGD is an iterative version of FGSM, which calculates the current gradient in multiple steps and adjusts the perturbation direction in each iteration. This algorithm can be represented as:
[0142]
[0143] In the perturbation visibility experiment of the object detection task on the PascalVOC dataset, since the DAG is a global attack, P L0 Almost reaching 1, such as Figure 2 The S-PGD proposed in this invention is a multi-layer semantic information region selection method, whose P L0 The size depends on the sparsity of the mask; in this experiment, S-PGD P L0 The lowest value was 0.26. (P) L2 The magnitude depends on the average intensity of the perturbed pixels, since S-PGD is based on minimizing L. ∞ Norm, its P L2 The value is slightly higher than that of DAG.
[0144] Figure 2 Medium blue represents P L0 Green represents P L2 .
[0145] This application refers to the semantic centralized adversarial perturbation generation method as S-PGD, the core idea of which is based on minimizing L ∞ Under the norm-based PGD (Projected Gradient Descent) framework, perturbation updates are performed only on the regions in the image with the richest semantic information, in order to achieve higher perturbation sparsity and concealment.
[0146] Please see Figure 3 The S-PGD proposed in this application is a multi-layer semantic information region selection method. Its perturbation range depends on the sparsity of the mask, and it is lower than that of the global adversarial algorithm in this experiment.
[0147] Figure 3 In the middle, first row, from left to right, the first image is the original image, and the second image is the normal result;
[0148] The second row, from left to right, shows the attack result in the first image and the corresponding S-PGD perturbation in the second image.
[0149] from Figure 3 It can be seen that the semantic mask proposed in this invention has the characteristic of sparsity. By combining image data to form semantic concentrated adversarial perturbation S-PGD, the perturbation range is low and the adversarial attack effect is good.
[0150] Conducting white-box adversarial attack experiments on the Pascal VOC dataset for object detection has significant research value and practical implications. As a standard benchmark dataset in computer vision, Pascal VOC covers 20 common object classes with moderate image complexity, making it suitable for evaluating the impact of attack algorithms on detection performance, especially providing good comparative and interpretative power regarding changes in mean accuracy (mAP). Compared to image classification, object detection is more complex, requiring simultaneous interference with object localization and classification branches. Therefore, it can more comprehensively reveal the destructive capabilities of attack algorithms on detection models (such as Faster-RCNN) in multi-object, multi-scale scenes. Simultaneously, this experiment provides a solid experimental foundation for analyzing the sensitivity of different backbone networks to attacks, exploring robustness differences, and subsequently building more robust detection models. Therefore, the Pascal VOC dataset in this study not only ensures the accuracy of adversarial attack evaluation but also promotes the in-depth development of adversarial robustness research in the detection field.
[0151] In white-box attack experiments on the PascalVOC dataset, this application evaluates the attack performance of different adversarial attack algorithms on various mainstream backbone networks. The metrics include the average precision before and after the attack (mAP_clean and mAP_att), the attack success rate (ASR), and the average time required to generate a single adversarial example (Inf_t), used to comprehensively measure the effectiveness and efficiency of the algorithms.
[0152] This application compares the proposed method with the typical white-box attack algorithm DAG (Dense Adversarial Generation). DAG is an attack algorithm specifically designed for the two-stage detector Faster-RCNN, which iteratively perturbs and optimizes the network based on a single candidate box. However, due to...
[0153] Faster-RCNN generates a large number of candidate boxes at once during the inference phase. The DAG must apply perturbations to each candidate box separately, which makes its computational cost extremely high, making it the algorithm with the longest inference time.
[0154] In contrast, S-FGSM (Single-step FGSM) has a significant efficiency advantage because it only requires one forward and backward propagation, resulting in the shortest inference time. However, due to the limitations of the single-step attack strategy, its attack perturbation cannot fully approximate the target gradient direction, leading to a lower overall attack success rate (ASR) and insufficient practical interference capability.
[0155] S-PGD (Gradient-aligned Projected Gradient Descent), as a multi-step iterative attack strategy, achieves a better balance between accuracy and stability by guiding the direction update of adversarial perturbations through continuous gradients. Experiments show that S-PGD achieves the highest attack success rate in a DLA34 backbone network configuration, far exceeding baseline methods such as DAG and S-FGSM. Specific results are shown in Table 1, validating the superior attack performance and good generalization ability of S-PGD in complex target detection scenarios. Overall, this experiment not only compares the efficiency and effectiveness of various attack methods but also reveals the differences in robustness of different architectures against adversarial examples, providing important references for future research.
[0156] Table 1. Quantitative Indicators of White-Box Attacks on the PascalVOC Dataset
[0157] Methods Backbones <![CDATA[mAP clean ]]> <![CDATA[mAP att ]]> ASR Inf_t S-FGSM Res18 0.71 0.26 0.63 0.2 S-FGSM Res101 0.77 0.34 0.56 0.4 S-FGSM DLA34 0.79 0.38 0.52 0.4 S-PGD Res18 0.71 0.07 0.90 1.7 S-PGD Res101 0.77 0.07 0.91 2.2 S-PGD DLA34 0.79 0.05 0.94 2.3 DAG FR 0.70 0.05 0.92 12.0
[0158] Conducting black-box adversarial attack experiments on the Pascal VOC dataset for object detection has significant theoretical and practical value. Compared to white-box attacks, which directly access model parameters and gradient information, black-box attacks are closer to real-world application scenarios. Attackers cannot know the structure and details of the target model and must rely on perturbations generated on the source model to achieve cross-model transfer. As a standard benchmark dataset in the field of object detection, Pascal VOC has clear categories, standardized annotations, and moderate scenario complexity. It can effectively evaluate the attack transfer rate of different algorithms under black-box settings and facilitate reproducible comparative analysis with existing research. Since object detection tasks themselves involve multi-target localization and classification, and have complex structures, black-box attacks face greater challenges. Therefore, conducting experiments on this dataset helps to deeply analyze the generalization ability of adversarial perturbations, the sensitivity of different backbone networks to perturbations, and promote the research of more robust detection model design and security defense strategies.
[0159] In black-box attack experiments on the Pascal VOC dataset, this application systematically evaluates the transfer attack capabilities of different attack algorithms under various model combinations. In the experimental setup, the columns of the matrix represent the source network (i.e., the original model) used to generate adversarial perturbations, and the rows represent the detection model (i.e., the target network) used as the attack target, thus constructing a comprehensive black-box attack evaluation framework. This application compares and analyzes the proposed method with the classic attack algorithm DAG (Dense Adversarial Generation).
[0160] As an attack algorithm targeting anchor-box detectors such as Faster R-CNN, DAG exhibits strong attack performance in white-box environments, especially when attacking the original trained model Faster R-CNN, where its average precision (mAP) drops significantly, indicating its effectiveness in perturbing candidate box generation. However, DAG performs poorly in cross-model transfer, i.e., black-box attacks, particularly when attacking target networks with architectures such as ResNet101 and DLA34, where its attack transfer rate (ATR) is only 0.08 and 0.06, respectively, revealing its high dependence on model structure and candidate box characteristics.
[0161] In contrast, the S-PGD (Structured-PGD) proposed in this application demonstrates a significant advantage in transferability. Both methods effectively enhance the cross-model generality of adversarial examples by introducing more structure-aware and gradient-consistent perturbation strategies. In experiments, the S-PGD perturbation generated based on the DLA34 source model achieved a transfer rate of 0.21 against Faster R-CNN, significantly higher than the results of DAG, demonstrating good cross-architecture transferability and black-box attack performance. This result not only proves the effectiveness of the proposed method but also shows that fully utilizing gradient information and feature structure adversarial performance is crucial when designing highly transferable adversarial perturbations (see Table 2).
[0162] Table 2. Quantitative Indicators of Black-Box Attacks on the PascalVOC Dataset
[0163]
[0164]
[0165] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for selecting multi-layer semantic information regions, characterized in that, Includes the following steps: For all convolutional layers in the set, first obtain the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ; G i Normalize to the range [0,1] to obtain Will Upsample to the input image size to obtain Will The semantic mask m is output through the binary function δ(x). i ; m of all obtained convolutional layers i The final mask M′ is obtained by performing element-wise product.
2. The multi-layer semantic information region selection method according to claim 1, characterized in that, The set of convolutional layers is as follows: in In the formula, L is the set of convolutional layers; M is all the layers in the set of convolutional layers; l i Let i be the i-th convolutional layer containing K channels; K represents the j-th channel of the i-th convolutional layer; K is the number of channels; i is the convolutional layer label; and j is the channel number label.
3. The multi-layer semantic information region selection method according to claim 2, characterized in that, Obtain the channel The gradient is as follows: In the formula, For channel The gradient of f(x); f(x) is the model function; Let be the j-th channel of the i-th convolutional layer; sum all gradients along the channel direction, expressed as: In the formula, For channel The gradient of G; i It is the sum of all gradients along the directions from channel 1 to k; K is the number of channels; i is the convolutional layer number; j is the channel number number.
4. The multi-layer semantic information region selection method according to claim 2, characterized in that, G i Normalized to the range [0,1], it can be represented as: In the formula, norm() is the normalization function; G is the sum of all gradients along channels 1 to k after normalization. i It is the sum of all gradients along the directions from channel 1 to k.
5. The multi-layer semantic information region selection method according to claim 4, characterized in that, Will Upsampling to the input image size is represented as: In the formula, Ψ is the upsampling function; This is the sum of all gradients along channels 1 to k after normalization. This is the sum of gradients restored to the input image size.
6. The multi-layer semantic information region selection method according to claim 5, characterized in that, Will The semantic mask m is output through the binary function δ(x). i As shown in the following formula: In the formula, δ(x) is a binary function, and x is the independent variable of the function; t α The custom threshold is set to 0.
5.
7. The multi-layer semantic information region selection method according to claim 6, characterized in that, m of all obtained convolutional layers i The final mask M′ is obtained by performing element-wise product, as shown in the following equation: In the formula, M′ is the final mask, m i is the mask for the i-th layer; i is the convolutional layer number; M is all the layers in the convolutional layer set.
8. A semantic centralized anti-disturbance system, characterized in that, include: Get G i The module, for all convolutional layers in the set of convolutional layers, first obtains the channels. The gradient is calculated, and then all gradients are summed along the channel direction to obtain G. i ; get Module, G i Normalize to the range [0,1] to obtain get Module, will Upsample to the input image size to obtain Semantic mask m i Module, will The semantic mask m is output through the binary function δ(x). i ; The final mask module M′ will obtain m from all convolutional layers. i The final mask M′ is obtained by performing element-wise product.
9. A computer device, characterized in that, It includes a storage device and a processor, the storage device storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The device stores a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Anchor-frame-free model class gradient global confrontation sample generation method, system and equipment
CN118608898A