Target detection confrontation attack method based on image restoration

Through the methods of image restoration and semantic anchor vector alignment, the problem of existing object detection adversarial attack methods relying on explicit annotation information is solved, and flexible specific area attacks and imperceptible adversarial sample generation are achieved.

CN120823458APending Publication Date: 2025-10-21RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510872332.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing target detection adversarial attack methods heavily rely on explicit spatial annotation information or detection information, making it difficult to perform specific attacks on multiple targets separately, and the generated perturbations are easily perceived or detected.

Method used

By converting the original image into a mask image, generating repair hints for image restoration, extracting the semantic information of the feature map, and using the semantic anchor vector to iteratively update the adversarial samples, the semantic alignment of the adversarial samples is achieved, and imperceptible target adversarial samples are generated.

Benefits of technology

It achieves targeted attacks on arbitrary image regions without relying on detector output or manual annotation, and is capable of performing category-specific attacks on different targets in the same image, with the generated perturbations being tiny and imperceptible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823458A_ABST
    Figure CN120823458A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection confrontation attack method based on image restoration. The method comprises the following steps: converting an original image into a mask image; generating a repair prompt based on the confrontation target area and the confrontation intention; repairing the mask image to generate a guide image; extracting a feature map from the guide image and the adversarial sample; combining the historical feature gradient information with the current feature information; extracting feature information to obtain a locally enhanced activation graph, and performing flattening processing to obtain a semantic anchor point vector; and carrying out semantic alignment on the semantic anchor point vectors of the adversarial sample and the guide image, and carrying out iterative updating until disturbance convergence to obtain a target adversarial sample. According to the method, the target attack and the non-target attack can be executed on any image area without depending on detector output or manual annotation, and flexible target detection of a specific area is realized to confront the attack; different targets in the same image can be subjected to specific category attacks, the targets are subjected to accurate disturbance, and the disturbance is small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of target detection countermeasure technology, and in particular to a target detection countermeasure attack method based on image restoration. Background Art

[0002] Deep neural networks are susceptible to adversarial attacks. Attackers create adversarial examples by adding carefully crafted, imperceptible perturbations to input images, attempting to deceive the model and cause it to produce erroneous outputs. Consequently, examples generated by adversarial attacks can lead to serious security issues and significantly impact the systems supported by object detection models. For example, attackers can design attack algorithms and craft adversarial examples to mislead autonomous driving systems, causing them to misidentify objects and pose serious safety risks. To further improve the security and accuracy of object detection models, researchers have proposed various attack algorithms. These algorithms aim to expose model vulnerabilities through active attacks, thereby achieving the goal of using attack to promote defense and enhance the security and accuracy of object detection systems. Current adversarial attack algorithms typically rely on data annotation or detection bounding box information. While they can effectively reduce detection accuracy, their attack performance is heavily dependent on the availability of annotations or predicted bounding boxes. Furthermore, due to their limited frameworks, they are difficult to flexibly generate attacks from scratch. In real-world scenarios, with the widespread adoption of IoT applications, research on model security and defense strategies against adversarial attacks are constantly evolving. Existing attack strategies need to be more flexible and imperceptible—undetectable by defense strategies—to successfully mislead the model and promote research on the security of object detection models.

[0003] To improve the flexibility and imperceptibility of adversarial attacks against object detection, existing attack methods typically manipulate the feature response maps of the backbone network to disrupt the semantic coherence of objects and achieve highly transferable attacks, or attempt to introduce category priors and spatial constraints of the guided reference image to generate adversarial perturbations that are independent of the detection box. However, these methods still suffer from issues such as being unable to control the spatial location and category of the attack target or being limited by the semantic compatibility between the guided image and the original scene.

[0004] Therefore, the target detection attack method in the related art has the following defects: 1. Existing adversarial attack methods for object detection rely heavily on explicit spatial annotations or detection information. Object detection models must localize and classify specific objects, which requires attack algorithms to also target specific objects, increasing the difficulty of attack. Various attack methods typically perturb samples using the classification or regression information obtained by the detector, or by perturbing samples using annotations of the object itself and its context. This makes the attack algorithm's performance heavily dependent on existing annotations or predicted bounding boxes in the image.

[0005] 2. Existing adversarial attack methods for object detection struggle to target multiple targets individually. Due to the varying semantic gaps between categories and varying attack framework designs, existing methods typically target all targets as belonging to the same category, or target surrounding areas with random category attacks based on the target's semantic information. This makes it difficult to target different targets within the same image individually. Furthermore, the perturbations generated by these attack methods can be visually noticeable, making them easily detected by humans or defense mechanisms.

[0006] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0007] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the Invention

[0008] The purpose of the embodiments of the present disclosure is to provide an object detection counterattack method based on image restoration, thereby overcoming one or more problems caused by the limitations and defects of related technologies to at least a certain extent.

[0009] The present disclosure first provides a method for object detection against attacks based on image restoration, the method comprising: Convert the original image into a mask image; Generate a repair hint based on the target area of ​​the confrontation and the confrontation intention, wherein the repair hint is used to modify the semantics of the target area to the semantics of the confrontation intention and to keep the semantics of the content in the area around the target area consistent with the original semantics; Inpainting the mask image using the inpainting hint to generate a guide image, wherein a target region of the guide image includes the inpainting hint, and a non-target region of the guide image retains original visual content and is embedded with a modification of the target region that is semantically consistent with the inpainting hint; Extracting a feature map from the guidance image and the initially generated adversarial sample; Combining historical feature gradient information with current feature information in the feature map to obtain semantic information; Extracting feature information of the feature map to obtain a locally enhanced activation map, and flattening the activation map to obtain a semantic anchor vector; Semantically aligning the adversarial sample with the semantic anchor vector of the guide image, iteratively updating the semantically aligned adversarial sample until the perturbation converges, and obtaining a target adversarial sample.

[0010] In one embodiment of the present disclosure, the step of converting the original image into the mask image includes: The pixels in the target area of ​​the original image are assigned a value of 1, and the pixels in the non-target area are assigned a value of 0.

[0011] In one embodiment of the present disclosure, the step of extracting a feature map from the guide image and the initially generated adversarial sample includes: Utilize feature extractors to extract multi-scale deep feature maps from hierarchical convolutional layers; The semantic loss obtained by iterative calculation of the feature map is fed back to the gradient information of the feature map.

[0012] In one embodiment of the present disclosure, the step of combining the historical feature gradient information in the feature map with the current feature information to obtain semantic information includes: By utilizing the progressive accumulation strategy of feature gradient information accumulation and adaptive weight function, historical feature gradient information is combined with current feature information to obtain semantic information.

[0013] In one embodiment of the present disclosure, the step of extracting feature information of the feature map, obtaining a locally enhanced activation map, and flattening the activation map to obtain a semantic anchor vector includes: Extracting feature information from the feature map using a sliding window-based maximum pooling operator to obtain a locally enhanced activation map; The activation map is reshaped by flattening and represented by a one-dimensional vector. The one-dimensional vector is used as a semantic anchor vector. The semantic anchor vector contains multi-dimensional information from low-level texture to high-level semantics.

[0014] In one embodiment of the present disclosure, the semantic loss function in the process of semantically aligning the semantic anchor vectors of the adversarial sample and the guide image is as follows:

[0015] in, is the semantic loss function, For adversarial samples, To guide the image, and Represent the first i The semantic anchor vector extracted by the feature layer.

[0016] The present disclosure further provides an object detection anti-attack system based on image restoration, the system comprising: A mask image generation module, used to convert the original image into a mask image; A repair hint generation module is used to generate a repair hint based on the target area of ​​the confrontation and the confrontation intention, wherein the repair hint is used to modify the semantics of the target area to the semantics of the confrontation intention and to keep the semantics of the content in the area around the target area consistent with the original semantics; a guide image generation module, configured to inpaint the mask image using the inpainting hint to generate a guide image, wherein a target region of the guide image includes the inpainting hint, and a non-target region of the guide image retains the original visual content and is embedded with a modification of the target region that is semantically consistent with the inpainting hint; A feature map extraction module, configured to extract a feature map from the guide image and the initially generated adversarial sample; A feature gradient accumulation module, configured to combine historical feature gradient information in the feature map with current feature information to obtain semantic information; An anchor vector acquisition module, used to extract feature information of the feature map, obtain a locally enhanced activation map, and flatten the activation map to obtain a semantic anchor vector; The semantic alignment module is used to semantically align the adversarial sample with the semantic anchor vector of the guide image, and iteratively update the semantically aligned adversarial sample until the perturbation converges to obtain a target adversarial sample.

[0017] In one embodiment of the present disclosure, the anchor vector acquisition module includes: An activation map acquisition unit is used to extract feature information from the feature map using a maximum pooling operator based on a sliding window to obtain a locally enhanced activation map; A flattening processing unit is used to reshape the activation map by flattening processing, express it with a one-dimensional vector, and use the one-dimensional vector as a semantic anchor vector, where the semantic anchor vector contains multidimensional information from low-level texture to high-level semantics.

[0018] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: In the disclosed embodiments, a method and system for target detection adversarial attack based on image restoration can perform targeted attacks and non-target attacks on any image area without relying on detector output or manual labeling, thereby realizing flexible target detection adversarial attack on specific areas. The present invention can perform specific category attacks on different targets in the same image, and accurately perturb the target through feature information, making the perturbation tiny and imperceptible. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0020] Figure 1 A flowchart of a method for countering attacks on target detection based on image restoration in an exemplary embodiment of the present disclosure is shown; Figure 2 A diagram showing an attack architecture structure based on image restoration in an exemplary embodiment of the present disclosure is shown; Figure 3 The overall flow chart of the target detection counterattack framework based on image restoration in an exemplary embodiment of the present disclosure is shown; Figure 4 A diagram showing image processing results of an object detection counterattack method based on image restoration in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0022] In addition, the accompanying drawings are merely schematic illustrations of embodiments of the present disclosure and are not necessarily drawn to scale. Like reference numerals in the figures represent like or similar parts, and thus repeated descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.

[0023] This example implementation first provides a method for object detection against attacks based on image restoration. Figure 1 The method may include: S101-S107. Specifically as follows: S101, converting the original image into a mask image; S102, generating a repair prompt based on the target area of ​​the confrontation and the confrontation intention, wherein the repair prompt is used to modify the semantics of the target area to the semantics of the confrontation intention and to keep the semantics of the content in the area around the target area consistent with the original semantics; S103, repairing the mask image using the repair hint to generate a guide image, wherein the target area of ​​the guide image includes the repair hint, and the non-target area of ​​the guide image retains the original visual content and is embedded with a modification of the target area that is semantically consistent with the repair hint; S104, extracting a feature map from the guide image and the initially generated adversarial sample; S105, combining the historical feature gradient information in the feature map with the current feature information to obtain semantic information; S106, extracting feature information of the feature map to obtain a locally enhanced activation map, and flattening the activation map to obtain a semantic anchor vector; S107 , semantically aligning the adversarial sample with the semantic anchor vector of the guide image, iteratively updating the semantically aligned adversarial sample until the perturbation converges, and obtaining a target adversarial sample.

[0024] The present invention can perform targeted and non-target attacks on any image area without relying on detector output or manual labeling, realizing flexible target detection adversarial attacks in specific areas; the present invention can perform specific category attacks on different targets in the same image, and accurately perturb the target through feature information, making the perturbation tiny and imperceptible.

[0025] The specific process of each step in the above embodiment is described below.

[0026] Please refer to Figure 2 and Figure 3 , in S101, for a given image and the area where the target is located ,This region can correspond to a foreground object, a background block, or any region of interest, and can flexibly adapt to real-world adversarial targets, such as target category replacement or object erasure tasks.

[0027] For the target area , which is converted into a binary mask region by encoding , as shown below:

[0028] Among them, H and W represent The pixel height and width of the region. The pixels inside the image are assigned the value 1 and the rest are assigned the value 0. This spatial encoding provides precise control over the modified area.

[0029] Mask image The definition is as follows:

[0030] in, Represents a channel-wise broadcasted multiplication that accurately masks the image.

[0031] After the image is masked, the process proceeds to S102, where a repair prompt is generated based on the target region and the intention of the confrontation. The repair prompt is used to modify the semantics of the target region to the semantics of the intention of the confrontation, and to keep the semantics of the content in the area around the target region consistent with the original semantics. t express, t Automatically generated based on the target region and optional user intent, inpainting hints guide the image inpainting model to generate contextually coherent adversarial content by modifying relevant vocabulary in the original image content (e.g., replacing "airplane" and "fly" with "horse" and "jump" to change the target object while preserving the background information). Specifically,

[0032] Then, the process proceeds to step S103 of generating a guide image. Specifically, the formula for synthesizing a guide image (real object or scene component) using the image restoration diffusion model is as follows:

[0033] in, To guide the image, Diffusion model for image restoration. Guided image Include the repair hint in the target area t , guide image The non-target area retains the visual content of the original image and embeds a modification of the target area that is semantically consistent with the repair hint.

[0034] It should be noted that when no explicit attack intent is provided (i.e. t as a general hint), the repair model will autonomously generate diverse semantic perturbations, including but not limited to: inducing the detector to misclassify the target object; generating disruptive visual patterns in the target area to undermine contextual reasoning, thereby being able to support both targeted and non-targeted attack scenarios.

[0035] In S104, since the disturbance obtained through explicit image information guidance may be more obvious or the attack performance is insufficient, it is necessary to use a detector to extract the feature information of the image and convert it into implicit key information to guide disturbance generation.

[0036] Specifically, S104 includes S1041 and S1042: S1041, extracting multi-scale deep feature maps from the hierarchical convolutional layers using a feature extractor; S1042: The semantic loss obtained by iteratively calculating the feature map is transmitted back to the gradient information of the feature map.

[0037] For the guidance image generated in the previous steps , from the guidance image and the adversarial samples generated by initialization These high-dimensional feature representations are then aligned in a shared embedding space through dual objectives to ensure semantic consistency, thereby guiding the adversarial perturbation generation process. Specifically, the backbone network of the pre-trained detector can be used to extract discriminative feature representations, and the and The consistency at the feature level achieves adversarial attack. This paper defines the feature extractor as a pre-trained model (e.g. VFNet, YOLOv3), extracting multi-scale deep feature maps from hierarchical convolutional layers. The features in these feature maps encode progressive representations of the input, covering multi-dimensional information from low-level textures to high-level semantics.

[0038] For each middle layer , feature map extracted by the detector It can be formally expressed as:

[0039] Indicates the number of attack iterations The set of intermediate feature maps extracted when is the total number of attack steps. Indicates the The feature tensor of the layer, with channel dimension and spatial resolution .

[0040] In each iteration, the semantic loss is calculated , pass the semantic loss back to the gradient information of each feature map , which is expressed as follows:

[0041] Through the above process, the extraction of multi-scale depth feature maps and the return of the gradient information of semantic loss to the feature maps are realized.

[0042] Next, in order to effectively obtain semantic information, the present invention adopts a method based on feature gradient information accumulation and adaptive weight function The gradual accumulation strategy of historical feature gradient information With current feature information Combined to further retain the key feature information in the historical process, the process can be expressed as:

[0043] in, Control the influence of the gradient in the early stage, so as to retain more key semantic information of the image in the early attack iteration process, is defined as follows:

[0044] Utilize the above. The weight adjustment strategy is combined with the implicit information acquisition method of feature gradient accumulation to ensure that the discriminative feature information in the early stage of attack generation can be better utilized, while gradually attenuating the influence of non-critical feature information in the later stage as the iteration deepens, effectively avoiding the local overfitting problem caused by over-reliance on specific instance features, and promoting the robust evolution of adversarial features through smooth gradient transitions. This application utilizes the gradient information accumulation of multi-time dimension features to calculate the feature information in the spatiotemporal dimension, amplifying the shared adversarial model with cross-dimensional consistency, and using a step-by-step adaptive update scheme to enhance the expression ability of generated perturbations between models.

[0045] In order to ensure that the generated adversarial samples are visually realistic and accurately aligned with the semantics of the target reference image, and to reduce irrelevant disturbances to the background area, the present invention proposes a perturbation optimization mechanism for dual-target semantic anchor alignment. During the adversarial attack process, the existing technology usually uses mean square error loss to measure the deviation between the original area of ​​the image and the target category. This may not fully utilize the key semantic information of the image, resulting in insufficient expression ability of the generated perturbation and poor attack effect. The present application uses a perturbation optimization mechanism to measure the deep semantic association between the original area and the target category through semantic similarity, and realizes accurate semantic mapping in the feature space.

[0046] Specifically, the present application adopts a dual constraint strategy: on the one hand, it establishes a connection through semantic relevance to ensure that the adversarial modification is highly consistent with the target category features; on the other hand, it uses spatial context consistency constraints to maintain a natural transition between the disturbed area and the surrounding environment, and accurately attacks the target area. This semantic guidance mechanism based on the latent feature space can not only achieve fine-grained attacks on any area of ​​the image, but also effectively improve the visual concealment of the adversarial disturbance. Combined with the feature map obtained in the above steps, in order to extract the key semantic anchor points for subsequent alignment, the present invention uses the feature map to extract the key semantic anchor points for subsequent alignment. Apply the sliding window based maximum pooling operator Extract important feature information from the feature map to obtain a locally enhanced activation map , this process can be expressed as follows:

[0047] Then, the flattening operation Reshape into a one-dimensional vector representation , while retaining the localized semantic response, it is convenient to integrate feature information into the alignment loss. .

[0048] The vector As a key semantic anchor, it captures the The key feature information of the layer feature map contains multi-dimensional information from low-level texture to high-level semantics.

[0049] The dual-target semantic alignment mechanism guides adversarial perturbations in meaningful multi-dimensional semantic directions while ensuring that the loss is as imperceptible as possible. The dual-target semantic alignment strategy proposed in this paper forces alignment of key semantic information in the adversarial image with the inpainted reference image at the feature level, while minimizing feature perturbations in irrelevant target areas.

[0050] Loss Function Encouraging adversarial images With guided images Alignment in high-dimensional feature space preserves object semantics and contextual coherence by aligning key information in specific target areas and keeping features of irrelevant areas consistent.

[0051] set up and Respectively represent No. The flattened semantic anchor vector extracted by the feature layer. The dual-objective semantic loss function is defined as follows:

[0052] in, Represents the cosine similarity calculation. This semantic loss function promotes the overall semantic consistency between images across multiple feature layers by aligning key semantic anchors, ensuring that the adversarial perturbation not only meets the specific target requirements of the attack but also retains the natural visual effect. During the attack iteration, the adversarial sample By minimizing , using a momentum-based iterative attack method to continuously iterate and update until the optimized perturbation converges, and finally obtains a flexible target adversarial (attack) sample with imperceptible perturbations.

[0053] In this example embodiment, a target detection adversarial attack system based on image restoration is first provided, which includes: a mask image generation module, a restoration hint generation module, a guide image generation module, a feature map extraction module, a feature gradient accumulation module, an anchor vector acquisition module and a semantic alignment module.

[0054] Specifically, the mask image generation module is used to convert the original image into a mask image; the repair hint generation module is used to generate a repair hint based on the adversarial target area and the adversarial intention, wherein the repair hint is used to modify the semantics of the target area to the semantics of the adversarial intention and to keep the semantics of the content of the area around the target area consistent with the original semantics; the guide image generation module is used to use the repair hint to repair the mask image to generate a guide image, wherein the target area of ​​the guide image contains the repair hint, and the non-target area of ​​the guide image retains the original visual content and is embedded with a modification of the target area that is semantically consistent with the repair hint; the feature map extraction module is used to extract feature maps from the guide image and the initialized adversarial sample; the feature gradient accumulation module is used to combine the historical feature gradient information in the feature map with the current feature information to obtain semantic information; the anchor vector acquisition module is used to extract feature information of the feature map, obtain a locally enhanced activation map, and flatten the activation map to obtain a semantic anchor vector; the semantic alignment module is used to semantically align the adversarial sample and the semantic anchor vector of the guide image, iteratively update the adversarial sample after semantic alignment until the perturbation converges, and obtain the target adversarial sample.

[0055] In addition, the anchor vector acquisition module includes: an activation map acquisition unit and a flattening processing unit.

[0056] Specifically, the activation map acquisition unit is used to extract feature information from the feature map using a maximum pooling operator based on a sliding window to obtain a locally enhanced activation map; the flattening processing unit is used to reshape the activation map through flattening processing, represent it with a one-dimensional vector, and use the one-dimensional vector as a semantic anchor vector, which contains multidimensional information from low-level texture to high-level semantics.

[0057] Please refer to Figure 2 , Figure 2The following is a diagram of the attack architecture based on image restoration in the present invention. For the input clean image dataset, a substitution model (such as YOLOv3) is used, along with an image-to-text content extraction model and an image restoration diffusion model. Specifically, for the input clean image, an image content extraction model is used in combination with a word replacement strategy to obtain guiding restoration hints. The restoration hints and the mask of the target area are combined to guide the image restoration diffusion model to repair the key area (i.e., the adversarial target area), thereby obtaining a target guiding image. The feature extractor of the substitution model is then used to extract multi-scale features of the target reference image and the adversarial image obtained by adding initialization noise to the clean image. With the help of a feature gradient accumulation module and a maximum pooling operator, the key semantic vectors in the image are obtained from the feature space. When the perturbation of the adversarial sample is updated, the semantic similarity of the obtained semantic vector is calculated using cosine similarity. A loss function based on cosine similarity is used to align the key semantic information of the reference image and the adversarial image in the feature space, thereby optimizing and updating the perturbation through an iterative attack algorithm based on momentum.

[0058] Please refer to Figure 3 , Figure 3 This is the overall flow chart of the target detection counter-attack framework based on image restoration of the present invention. For the input clean image, the image content extraction model is used to obtain the key information of the target area, and then the information about the target such as the subject and predicate is replaced by words to modify the information of the target area of ​​the image and use it as a guiding restoration prompt. The restoration prompt and the mask of the target area are combined to guide the image restoration diffusion model to repair the key area (target area), and finally the target reference image and the adversarial image with the initial disturbance added to the clean image are obtained. For the obtained target reference image and clean image, the multi-scale features of the two are first extracted with the help of the feature extractor of the substitution model, and then the feature gradient information generated by the backpropagation of the feature map of the previous iteration round is obtained and passed through the symbolic function This process accumulates historical feature gradient information on the feature map of the current iteration, effectively preserving the key perturbation adversarial patterns in the feature map. After feature gradient accumulation, multiple sets of feature maps (adversarial samples, guided samples) can be obtained.

[0059] Multiple sets of features are pooled using the maximum pooling operator, where the pooling operator is set to the step size , the size is This downsampling operation helps better capture key semantic information from the features. For each pooled set of key semantic vectors (adversarial feature vectors and guided feature vectors), cosine similarity is calculated to determine the multi-dimensional semantic association between the adversarial image and the reference image. A constructed loss function is then used to perform dual-target semantic alignment, effectively perturbing specific target regions while minimizing the impact on irrelevant areas. The perturbation generation process is implemented using a momentum-based iterative attack algorithm. The loss function is backpropagated to determine the gradient of the input sample—the direction in which the model is sensitive to input changes. The algorithm then introduces momentum, combining the current gradient with historical gradient directions. This direction is then updated smoothly through a weighted average to prevent the perturbation from deviating from the target due to local optima or noise. In each iteration, the algorithm adds a small perturbation to the adversarial example along the direction of the momentum-accumulated gradient, ensuring that the perturbation amplitude does not exceed a preset threshold. This process is repeated until the adversarial example successfully misleads the model or the maximum number of iterations is reached.

[0060] The following is a detailed description of the method for countering attacks on target detection based on image restoration and its beneficial effects.

[0061] Test steps: ① Input a picture , the size of the picture is , extract the image content and get the text , and generate the corresponding mask according to the specified target area ; ② For the extracted text Perform word replacement strategies (such as replacing the subject and predicate) to modify the text into a semantically reasonable text with the original target changed to a target of a specified attack category. ; ③Text and mask As a guiding condition input to the image restoration diffusion model Repair and obtain the repaired guide image ; ④ Image Add initial perturbation to get the initial adversarial image ; ⑤The guide image and the initial adversarial image Input to the feature extractor to obtain multi-scale depth feature map and ; ⑥ Backpropagate the gradient of the previous iteration to the adversarial sample feature map The gradient map is weighted and accumulated into the adversarial sample feature map of the current iteration round On the top, we get the accumulated adversarial sample feature map ; ⑦ Feature map of adversarial samples after accumulation and guided feature maps Apply the sliding window based max pooling operator , and through flattening operation, we get the key semantic vectors and ; ⑧Through the semantic alignment loss function, the key semantic vector and The loss is calculated and then the initial adversarial image is attacked by the momentum-based iterative attack algorithm. Add noise; ⑨ After n rounds of iteration (e.g. n=10), an adversarial image with a high attack rate can be obtained .

[0062] Experimental results: As shown in Table 1, our image inpainting-based adversarial attack method for object detection demonstrates significant advantages in target attack tasks targeting specific regions. Compared to existing methods (TOG, a real-time attack method for target gradient attacks; CAI, a contextual attack method for object detection), our method achieves higher attack success rates in both foreground and background regions. Specifically, in background object implantation tasks (such as synthesizing pedestrians on an empty road), our image inpainting-driven attack constructs a semantic alignment mechanism guided by deep semantic features and a feature space decoupling optimization strategy. This allows for direct modulation of the local feature response of the target region without relying on prior detection boxes, enabling feature-level semantic reconstruction. Experimental results confirm that our image inpainting-driven attack framework maintains a stable attack success rate even in complex dynamic scenes, validating its robustness in representing multi-scale semantic features.

[0063] Table 1: Comparison of the success rate of adversarial attacks on target and background regions under multiple detectors / datasets ( )

[0064] like Figure 4As shown, the image restoration-driven attack method of the present application demonstrates the typical effect of multi-target attack, and is able to perform independent category attribute manipulation on different targets in the image. In real scenarios, the original image contains multiple objects to be detected (such as pedestrians, children, dogs), and the adversarial samples generated by the image restoration-driven attack framework can independently misclassify these targets into preset error categories (such as trucks, apples, bears). These error categories show a high confidence level in the detector and the original target object category cannot be detected. Even if these categories have a huge semantic gap (such as attacking a person as a car), the model detection is still guided to output a preset target label for each object. This multi-target independent attack capability has an important impact in safety-critical scenarios such as autonomous driving. For example, pedestrians and traffic signs on the road can be hidden at the same time, or multiple obstacles can be falsely generated to promote the further improvement of existing defense systems.

[0065] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A method for object detection against attacks based on image restoration, characterized in that: The method comprises: Convert the original image into a mask image; Generate a repair hint based on the target area of ​​the confrontation and the confrontation intention, wherein the repair hint is used to modify the semantics of the target area to the semantics of the confrontation intention and to keep the semantics of the content in the area around the target area consistent with the original semantics; Inpainting the mask image using the inpainting hint to generate a guide image, wherein a target region of the guide image includes the inpainting hint, and a non-target region of the guide image retains original visual content and is embedded with a modification of the target region that is semantically consistent with the inpainting hint; Extracting a feature map from the guidance image and the initially generated adversarial sample; Combining historical feature gradient information with current feature information in the feature map to obtain semantic information; Extracting feature information of the feature map to obtain a locally enhanced activation map, and flattening the activation map to obtain a semantic anchor vector; Semantically aligning the adversarial sample with the semantic anchor vector of the guide image, iteratively updating the semantically aligned adversarial sample until the perturbation converges, and obtaining a target adversarial sample.

2. The method for object detection against attacks based on image restoration according to claim 1, characterized in that: The step of converting the original image into a mask image comprises: The pixels in the target area of ​​the original image are assigned a value of 1, and the pixels in the non-target area are assigned a value of 0.

3. The method for object detection against attacks based on image restoration according to claim 1, characterized in that: The step of extracting a feature map from the guide image and the initially generated adversarial sample comprises: Utilize feature extractors to extract multi-scale deep feature maps from hierarchical convolutional layers; The semantic loss obtained by iterative calculation of the feature map is fed back to the gradient information of the feature map.

4. The method for object detection against attacks based on image restoration according to claim 3, characterized in that: The step of combining the historical feature gradient information in the feature map with the current feature information to obtain semantic information includes: By utilizing the progressive accumulation strategy of feature gradient information accumulation and adaptive weight function, historical feature gradient information is combined with current feature information to obtain semantic information.

5. The method for object detection against attacks based on image restoration according to claim 4, characterized in that: The step of extracting feature information of the feature map, obtaining a locally enhanced activation map, and flattening the activation map to obtain a semantic anchor vector includes: Extracting feature information from the feature map using a sliding window-based maximum pooling operator to obtain a locally enhanced activation map; The activation map is reshaped by flattening and represented by a one-dimensional vector. The one-dimensional vector is used as a semantic anchor vector. The semantic anchor vector contains multi-dimensional information from low-level texture to high-level semantics.

6. The method for object detection against attacks based on image restoration according to claim 5, characterized in that: The semantic loss function in the process of semantically aligning the semantic anchor vectors of the adversarial sample and the guide image is as follows: in, is the semantic loss function, For adversarial samples, To guide the image, and denote the semantic anchor vectors extracted from the i-th feature layer of the adversarial sample and the guidance image, respectively.

7. An object detection counterattack system based on image restoration, characterized in that: The system comprises: A mask image generation module, used to convert the original image into a mask image; A repair hint generation module is used to generate a repair hint based on the target area of ​​the confrontation and the confrontation intention, wherein the repair hint is used to modify the semantics of the target area to the semantics of the confrontation intention and to keep the semantics of the content in the area around the target area consistent with the original semantics; a guide image generation module, configured to inpaint the mask image using the inpainting hint to generate a guide image, wherein a target region of the guide image includes the inpainting hint, and a non-target region of the guide image retains the original visual content and is embedded with a modification of the target region that is semantically consistent with the inpainting hint; A feature map extraction module, configured to extract a feature map from the guide image and the initially generated adversarial sample; A feature gradient accumulation module, configured to combine historical feature gradient information in the feature map with current feature information to obtain semantic information; An anchor vector acquisition module, used to extract feature information of the feature map, obtain a locally enhanced activation map, and flatten the activation map to obtain a semantic anchor vector; The semantic alignment module is used to semantically align the adversarial sample with the semantic anchor vector of the guide image, and iteratively update the semantically aligned adversarial sample until the perturbation converges to obtain a target adversarial sample.

8. The object detection counterattack system based on image restoration according to claim 7, characterized in that: The anchor vector acquisition module includes: An activation map acquisition unit is used to extract feature information from the feature map using a maximum pooling operator based on a sliding window to obtain a locally enhanced activation map; A flattening processing unit is used to reshape the activation map by flattening processing, express it with a one-dimensional vector, and use the one-dimensional vector as a semantic anchor vector, where the semantic anchor vector contains multidimensional information from low-level texture to high-level semantics.