Confrontation method and device for defending attack of physical world on neural network, electronic equipment and medium

By building high-quality adversarial samples, using the globally-wide adversarial perturbation invisible to human eyes and adversarial patches of small areas, the target detection model is trained, and the problem of difficulty in preventing the physical world from attacking neural networks is solved in the existing technology, and effective defense against physically implementable attacks is achieved.

CN120071044APending Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120734.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-12
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively defend against attacks on neural networks by the physical world, especially when facing physically achievable attacks, the existing defense methods are not effective.

Method used

By building high-quality adversarial samples, using globally-wide adversarial perturbations invisible to human eyes and adversarial patches for small areas, the target detection model is trained, thereby enhancing its defense against physically achievable attacks.

Benefits of technology

This method can significantly improve the defense capabilities of the target detection model against physically achievable attacks, while maintaining detection performance on clean images, ensuring the robustness of the model when facing different types of attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071044A_ABST
    Figure CN120071044A_ABST
Patent Text Reader

Abstract

The invention relates to a confrontation method and device for defending attack of a physical world on a neural network, electronic equipment and a medium, and the method comprises the steps: obtaining a target detection task for a target visual file, the target detection task being used for indicating at least one target object needing to be detected; inputting the target visual file into a target detection model, and calculating to obtain a classification result and position information of each target object; the target detection model is obtained by training antagonistic samples, each antagonistic sample is obtained by adding antagonistic disturbance and antagonistic patches to corresponding original samples, the antagonistic disturbance comprises global noise set for the original samples, and the antagonistic patches comprise local patches set for objects in the original samples. According to the method and the device, high-quality antagonistic samples can be constructed based on antagonistic disturbance invisible to human eyes in a global range and small-region antagonistic patches, and the antagonistic samples are applied to training to defend various physical realizable attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to an adversarial method and apparatus, an electronic device, and a medium for defending against attacks on neural networks in the physical world. Background Art

[0002] Object detection is a fundamental task in computer vision. Object detection refers to detecting the position and size of an object in a given image or video, usually marked with a bounding box. At the same time, this task requires classifying all objects in the image. In recent years, neural networks have been widely used in fields such as computer vision, and the development of object detection methods has greatly benefited from the application of neural networks. However, neural networks are vulnerable to adversarial samples, which brings huge security risks to security-sensitive fields. Therefore, exploring strategies to enhance the defense against attacks on neural networks in the physical world has important practical significance. Summary of the Invention

[0003] In view of this, the present disclosure provides an adversarial method and apparatus, an electronic device, and a medium for defending against attacks on neural networks in the physical world, which can construct high-quality adversarial samples based on global-scale imperceptible adversarial perturbations and small-region adversarial patches, and apply these adversarial samples in training to defend against various physically realizable attacks.

[0004] According to one aspect of the present disclosure, there is provided an adversarial method for defending against attacks on neural networks in the physical world, including: obtaining an object detection task for a target visual file, where the object detection task is used to indicate at least one target object to be detected; inputting the target visual file into a target detection model for calculation to obtain classification results and position information of each target object; where the target detection model is trained using adversarial samples, and each adversarial sample is obtained by adding an adversarial perturbation and an adversarial patch to a corresponding original sample, the adversarial perturbation includes global noise set for the original sample, and the adversarial patch includes a local patch set for an object in the original sample.

[0005] In this way, this adversarial method can construct high-quality adversarial samples based on global-scale imperceptible adversarial perturbations and small-region adversarial patches, and apply these adversarial samples in training, which can enhance the defense ability of the target detection model against physically realizable attacks, and at the same time will not reduce the detection performance on clean images, so as to obtain the classification results (i.e., the category to which the target object belongs) and position information (i.e., the position of the target object in the target image) of the target object.

[0006] In a possible implementation, the method further includes: obtaining a plurality of original samples, each of the original samples including an original image and object information of each object in the original image, and each of the object information including category information and location information; determining adversarial samples required for the current training based on each of the original samples; performing model training based on the current adversarial samples until the model converges to obtain the target detection model.

[0007] In this way, the target detection model in this adversarial method will learn the features of adversarial samples during training, thereby enhancing the recognition ability of adversarial samples and improving the defense ability of the target detector.

[0008] In a possible implementation, determining adversarial samples required for the current training based on each of the original samples includes: selecting a target original image required for the current training from the original images of each of the original samples; adding an adversarial perturbation to the target original image, and adding an adversarial patch to the target original image; using the target original image added with the adversarial perturbation and the adversarial patch as the training image for the current training, and obtaining the adversarial samples required for the current training according to the training image and the object information corresponding to the training image.

[0009] In a possible implementation, adding an adversarial patch to the target original image includes: adding a corresponding adversarial patch for at least one first object among a plurality of objects in the target original image.

[0010] In this way, by adding adversarial patches to some objects in the target original image, this adversarial method can balance the adversarial ability and recognition ability of the target detection model.

[0011] In a possible implementation, adding a corresponding adversarial patch for at least one first object among a plurality of objects in the target original image includes: determining an object region corresponding to each of the first objects in the target original image according to the location information of each of the first objects in the target original image; randomly sampling each of the object regions to determine a patch region in each of the object regions; dividing each of the patch regions into a plurality of sub-regions, and calculating the image gradient of each of the sub-regions to obtain the gradient of each of the sub-regions; selecting a target sub-region of each of the patch regions from all the sub-regions according to the gradients of all the sub-regions, and adding an adversarial patch corresponding to the first object in the target sub-region.

[0012] In this way, when generating adversarial patches, this adversarial method uses two strategies: random sampling and gradient-based partitioning, enabling the adversarial samples to attack the weakest parts of the image as much as possible. While introducing more adversarial patches, it maximally preserves the feature information of the image itself, enabling the object detection model to learn more information about the image itself during the model training process and ensuring the recognition accuracy of the object detection model.

[0013] In a possible implementation, the method further includes: calculating a target gradient value based on historical perturbations, the target original image, and the class information in the original sample, where the historical perturbation is a preset initial perturbation when the current training is the first training, and the historical perturbation is determined based on the adversarial perturbation and adversarial patch used in the previous training when the current training is not the first training; determining global noise based on the target gradient value, a preset global perturbation intensity, and historical global noise, and using the global noise as the adversarial perturbation added to the target original image, where the historical global noise is a preset initial noise when the current training is the first training, and the historical global noise is the global noise used in the previous training when the current training is not the first training, and determining a local patch based on the target gradient value, a preset perturbation patch step size, and historical local patches, and using the local patch as the adversarial patch added to the target original image, where the historical local patch is a preset initial patch when the current training is the first training, and the historical local patch is the local patch used in the previous training when the current training is not the first training.

[0014] In this way, this adversarial method adds a total perturbation formed by adversarial perturbation and adversarial patches to the target original image, enabling the simulated attack to more comprehensively cover the potential attack space, enhancing the effect of adversarial training, improving the robustness of the object detection model against different types of attacks, enabling the model to effectively resist multiple attacks, and at the same time updating the perturbation using the previously calculated gradient information, which can reduce the additional training cost brought by internal maximization and improve the training efficiency.

[0015] In a possible implementation, the method further includes: modifying the global noise when the determined global noise is outside a preset perturbation range, where the preset perturbation range is determined based on the global perturbation intensity; and modifying the local patch when the determined local patch is outside a preset patch range, where the preset patch range is determined based on a preset patch perturbation intensity.

[0016] According to another aspect of the present disclosure, there is provided an adversarial device for defending neural network attacks in the physical world, including: an acquisition module configured to acquire a target detection task for a target visual file, the target detection task being used to indicate at least one target object to be detected, wherein the type of the target visual file is an image or a video; a calculation module configured to input the target visual file into a target detection model for calculation to obtain classification results and location information of each of the target objects; wherein the target detection model is trained using adversarial samples, and each of the adversarial samples is obtained by adding an adversarial perturbation and an adversarial patch to a corresponding original sample, the adversarial perturbation includes global noise set for the original sample, and the adversarial patch includes local patches set for objects in the original sample.

[0017] In this way, the present adversarial device can construct high-quality adversarial samples based on the globally imperceptible adversarial perturbation to the human eye and the adversarial patches in small regions, and apply these adversarial samples in training, which can enhance the defense ability of the target detection model against physically achievable attacks, and at the same time will not reduce the detection performance on clean images, so as to obtain the classification results (i.e., the categories to which the target objects belong) and location information (i.e., the positions of the target objects in the target image) of the target objects.

[0018] In a possible implementation manner, the device further includes a training module configured to: acquire a plurality of original samples, each of the original samples including an original image and object information of each object in the original image, and each of the object information including category information and location information; determine adversarial samples required for the current training based on each of the original samples; perform model training based on the current adversarial samples until the model converges to obtain the target detection model.

[0019] In this way, the target detection model in the present adversarial device will learn the features of the adversarial samples during the training process, thereby enhancing the recognition ability of the adversarial samples to improve the defense ability of the target detector.

[0020] In a possible implementation manner, determining the adversarial samples required for the current training based on each of the original samples includes: selecting a target original image required for the current training from the original images of each of the original samples; adding an adversarial perturbation to the target original image, and adding an adversarial patch to the target original image; using the target original image added with the adversarial perturbation and the adversarial patch as the training image for the current training, and obtaining the adversarial samples required for the current training according to the training image and the object information corresponding to the training image.

[0021] In a possible implementation, adding an adversarial patch to the target original image includes: adding a corresponding adversarial patch for at least one first object among multiple objects in the target original image.

[0022] In this way, by adding adversarial patches to some objects in the target original image, this adversarial device can balance the adversarial ability and recognition ability of the target detection model.

[0023] In a possible implementation, adding a corresponding adversarial patch for at least one first object among multiple objects in the target original image includes: determining object regions corresponding to the first objects in the target original image according to the position information of the first objects in the target original image; randomly sampling each of the object regions to determine patch regions in each of the object regions; dividing each of the patch regions into multiple sub-regions, and calculating image gradients for each of the sub-regions to obtain gradients of each of the sub-regions; selecting target sub-regions of each of the patch regions from all the sub-regions according to the gradients of all the sub-regions, and adding an adversarial patch corresponding to the first object in the target sub-region.

[0024] In this way, when generating adversarial patches, this adversarial device uses two strategies, namely random sampling and division according to gradients, which can make the adversarial samples attack the weakest parts of the image as much as possible. While introducing more adversarial patches, it can retain the feature information of the image itself to the greatest extent, enabling the target detection model to learn more information about the image itself during the model training process and ensuring the recognition accuracy of the target detection model.

[0025] In a possible implementation, the device further includes a determination module, configured to: calculate a target gradient value according to historical perturbations, the target original image, and class information in the original samples, where the historical perturbation is a preset initial perturbation when the current training is the first training, and the historical perturbation is determined according to the adversarial perturbation and adversarial patch used in the previous training when the current training is not the first training; determine a global noise according to the target gradient value, a preset global perturbation intensity, and historical global noise, and use the global noise as the adversarial perturbation added to the target original image, where the historical global noise is a preset initial noise when the current training is the first training, and the historical global noise is the global noise used in the previous training when the current training is not the first training; and determine a local patch according to the target gradient value, a preset perturbation patch step size, and historical local patches, and use the local patch as the adversarial patch added to the target original image, where the historical local patch is a preset initial patch when the current training is the first training, and the historical local patch is the local patch used in the previous training when the current training is not the first training.

[0026] In this way, by adding the total perturbation formed by the adversarial perturbation and adversarial patch to the target original image, the proposed adversarial device enables the simulated attack to more comprehensively cover the potential attack space, enhances the effect of adversarial training, improves the robustness of the target detection model against different types of attacks, enables the model to effectively resist various attacks, and updates the perturbation using the previously calculated gradient information, which can reduce the additional training cost brought by the internal maximization and improve the training efficiency.

[0027] In a possible implementation, the device further includes a modification module, configured to: modify the global noise when the determined global noise is outside a preset perturbation range, where the preset perturbation range is determined according to the global perturbation intensity; and modify the local patch when the determined local patch is outside a preset patch range, where the preset patch range is determined according to a preset patch perturbation intensity.

[0028] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing processor-executable instructions, where the processor is configured to implement the above method when executing the instructions stored in the memory.

[0029] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, where the computer program instructions implement the above method when executed by a processor.

[0030] According to another aspect of the present disclosure, there is provided a computer program product including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0031] The present disclosure obtains a target detection task for a target visual file, where the target detection task is used to indicate at least one target object to be detected, inputs the target visual file into a target detection model for calculation, and obtains classification results and location information of each target object. Among them, the target detection model is trained using adversarial samples, and each adversarial sample is obtained by adding an adversarial perturbation and an adversarial patch to a corresponding original sample. The adversarial perturbation includes global noise set for the original sample, and the adversarial patch includes a local patch set for an object in the original sample. In this way, by combining the globally imperceptible adversarial perturbation with the small-area adversarial patch, high-quality adversarial samples can be constructed, and these adversarial samples are used in training to defend against various physically realizable attacks, solving the current situation that existing defense methods are difficult to resist physically realizable attacks.

[0032] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings included in and constituting a part of this specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0034] Figure 1a Schematic diagrams showing a clean image, an adversarial perturbation, and an adversarial sample.

[0035] Figure 1b Schematic diagrams showing an adversarial patch and a texture.

[0036] Figures 2 to 3 Schematic diagram showing an adversarial method provided by an embodiment of the present disclosure for defending against attacks on a neural network in the physical world.

[0037] Figure 4 Schematic diagram showing a patch area provided by an embodiment of the present disclosure.

[0038] Figure 5 Block diagram showing an adversarial device provided by an embodiment of the present disclosure for defending against attacks on a neural network in the physical world. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0040] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0041] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0042] To facilitate the understanding of the technical solutions provided by the embodiments of the present disclosure by those skilled in the art, the technical environment for implementing the technical solutions will be described first below.

[0043] Object detection requires classifying and localizing all objects in an image simultaneously. Recently, the development of object detection methods has greatly benefited from the use of deep neural networks (DNNs), which are neural networks with a deep structure and are widely used in fields such as computer vision.

[0044] However, deep neural networks are vulnerable to adversarial samples. Adversarial samples are images generated by adding carefully constructed perturbations to clean images (i.e., "original images"). As Figure 1a shown, the original image showing a panda is classified as "panda" by the DNN, while the adversarial sample formed by adding an adversarial perturbation to the original image showing a panda is classified as "gibbon" by the DNN. Although the adversarial perturbations cannot be observed by the human eye, they can deceive the classification or detection tasks of the DNN, resulting in incorrect results. These adversarial samples exist not only in the digital world but also in the real world. Currently, some works have added patterns, i.e., adversarial patches, and textures to pictures, for example Figure 1bThe shown Projected Gradient Descent Attack (PGD) patches, adversarial patches, adversarial textures, and adversarial camouflage textures can be printed in the real world to directly attack object detectors in the real world, where the object detectors can be formed based on object detection models. This poses a huge security risk to many security-sensitive fields such as autonomous driving, security monitoring, and biometric systems. Therefore, exploring strategies to enhance the defense of object detectors against physical attacks has important practical significance.

[0045] Current solutions for defending against physical attacks can be divided into Adversarial Training (AT) methods and non-AT methods. Non-AT methods mainly include preprocessing-based methods, abnormal feature filtering-based methods, and methods of adding defense boxes. Preprocessing-based methods mainly process images before sending them into object detectors. They usually detect the positions where adversarial patterns that may be aggressive exist through certain means, remove these patterns, and then send the processed images into object detectors for object detection. Abnormal feature filtering-based methods introduce filters in object detection models to mitigate abnormal internal features caused by adding adversarial patterns. The filters will find regions with abnormal features in the images and smooth and modify these abnormal regions. Methods of adding defense boxes enhance the robustness of object detectors by training a defensive box to stick around the image. Non-AT methods are often threatened by powerful adaptive attacks. Adaptive attacks refer to the situation where the attacker knows the specific parameters of the neural network and will use these parameters to generate adversarial samples targeted at the object detection model. The adversarial samples generated in this way will be more aggressive towards the object detection model, making the effects of these non-AT defense methods worse or even ineffective.

[0046] AT methods improve the robustness of neural networks by adding adversarial samples to the training of neural networks. Such methods use specific algorithms to generate adversarial samples, combine the generated adversarial samples with the original training samples to form an extended training set, and use the extended training set to train the detector. Although AT can well defend against adaptive attacks, existing AT methods are basically trained under tiny perturbation attacks that are difficult to detect by the human eye, and they can only defend against tiny perturbation attacks. When facing physical realizable large-area adversarial texture attacks, these AT methods often do not have good effects.

[0047] To solve the above technical problems, the embodiments of the present disclosure provide an adversarial method for defending against attacks on neural networks in the physical world, using a composite AT based on perturbation and patches to solve the current situation that existing defense methods are difficult to resist physically realizable attacks. Now in combination withFigures 2 to 4 A schematic illustration of the adversarial method provided by the embodiments of the present disclosure for defending neural networks against attacks in the physical world is given. As Figure 2 shown, the adversarial method may include the following steps S101 to S102.

[0048] Step S101: Obtain a target detection task for a target visual file.

[0049] The type of the target visual file is an image or a video. The target detection task is used to indicate at least one target object to be detected. For example, from a certain target image (i.e., the target visual file) to be detected, two target objects, namely a cat and a cup, need to be detected.

[0050] Step S102: Input the target visual file into the target detection model for calculation to obtain the classification results and location information of each target object.

[0051] The target detection model may be formed based on DNN, and a target detector may be formed based on the target detection model. The target detection model is trained using adversarial samples. Each adversarial sample is obtained by adding an adversarial perturbation and an adversarial patch to the corresponding original sample. The adversarial perturbation is an invisible interference to the human eye added to the global range of the original image in the original sample, and the adversarial perturbation includes global noise set for the original sample. The adversarial patch is an interference added to the local area of the original image in the original sample, and the adversarial patch includes local patches set for the objects in the original sample.

[0052] In this way, the present adversarial method can construct high-quality adversarial samples based on the globally invisible adversarial perturbation to the human eye and small-area adversarial patches, and apply these adversarial samples in training, which can enhance the defense ability of the target detection model against physically achievable attacks, and at the same time will not reduce the detection performance on clean images, so as to obtain the classification results (i.e., the categories to which the target objects belong) and location information (i.e., the locations of the target objects in the target image) of the target objects.

[0053] This adversarial method may also include the training process of the target detection model: obtaining a plurality of original samples, each original sample including an original image and object information of each object in the original image, and each object information including category information and location information; determining adversarial samples required for the current training based on each original sample; performing model training based on the current adversarial samples until the model converges to obtain the target detection model. Among them, determining the adversarial samples required for the current training based on each original sample may include: selecting the target original image required for the current training from the original images of each original sample; adding an adversarial perturbation and an adversarial patch to the target original image; using the target original image added with the adversarial perturbation and the adversarial patch as the training image for the current training, and obtaining the adversarial samples required for the current training according to the training image and the object information corresponding to the training image.

[0054] In one example, batch training is used for model training, that is, the entire dataset is divided into several batches, and the number of samples in each batch is called the batch size. A complete forward calculation and backpropagation are performed on each batch to update the model parameters. The dataset in this example includes 100 original samples, which are divided into 10 batches, and the batch size in each batch is 10. The first batch is used as an example for illustration. The 10 original samples in the first batch are denoted as {x 1 ,x 2 ,…,x 10}. Select one of the original samples from {x 1 ,x 2 ,…,x 10} as the basis for determining the adversarial samples required for the current training. For example, the original sample is x 1 , that is, the original image of x 1 (denoted as p 1 ) is used as the target original image required for the current training. Add an adversarial perturbation and an adversarial patch to p 1 to obtain a new image (denoted as d 1 ), so that d 1 is used as the training image for the current training, and according to d 1 and the object information corresponding to d 1 (that is, the object information corresponding to x 1 ), obtain the adversarial samples required for the current training (denoted as a 1 ), and thus use a 1 for adversarial training. Among them, the order of adding the adversarial perturbation and the adversarial patch to p 1 is uncertain and can be flexibly set according to the actual situation. The processing process of the original samples in the remaining batches is the same as that of x 1, for the sake of brevity, it will not be elaborated here. In this way, the object detection model in this adversarial method will learn the features of adversarial samples during the training process, thereby enhancing the ability to recognize adversarial samples and improving the defense ability of the object detector.

[0055] Before training the model with adversarial samples, pre-trained weights can be used, that is, load a weight file trained on a clean sample set and then perform adversarial training. Among them, the clean sample set includes each original sample without added adversarial perturbations and adversarial patches.

[0056] Different from past AT methods, the AT provided by the embodiments of the present disclosure not only has adversarial perturbations for global range perturbations, but also has adversarial patches for local regions. The training objective for adding adversarial patches in adversarial training can be expressed by Equation 1 below:

[0057]

[0058] In Equation 1, θ represents the object detection model parameters, δ p represents the adversarial patch, m p represents the binary mask used to limit the patch perturbation area, δ p ⊙m p represents adding the adversarial patch to the weakest area described later, β represents the maximum perturbation intensity of the adversarial patch, L d is the loss function of the object detector f θ (including classification loss and regression loss), x represents the original object image, y represents the label information corresponding to x (i.e., class information), argmin represents the parameter value that obtains the minimum value in the function domain (i.e., the value of the independent variable), represents taking the expectation of the image distribution of the input original image x, that is, calculating the average effect of the loss function L d over the entire data distribution. In this way, this adversarial method can greatly improve the defense ability of the object detection model against patch attacks and defend against attacks that can be realized in the physical world through such a training objective.

[0059] During the training process of the above-mentioned object detection model, adding adversarial patches to the original object image may include: adding corresponding adversarial patches to at least one first object among multiple objects in the original object image. The first object refers to the object selected to add adversarial patches. For example, adversarial patches can be added to half of the objects in the multiple objects of the original object image. For instance, if there are four objects, namely a cat, a cup, a sofa, and a coffee table, in a certain original object image, adversarial patches can be added only to the cat and the cup, and the cat and the cup are the first objects. The number of first objects can be flexibly set according to actual needs, and the embodiments of the present disclosure do not limit this. In this way, by adding adversarial patches to some objects in the original object image, this adversarial method can balance the adversarial ability and recognition ability of the object detection model.

[0060] To enable the object detection model to pay more attention to the original information of the image, in the embodiments of the present disclosure, two strategies, namely random sampling and gradient-based partitioning, are used when generating adversarial patches, which can make the adversarial samples attack the weakest parts of the image as much as possible. Generally speaking, first, a larger region is selected. This larger region can be, but is not limited to, square. The embodiments of the present disclosure do not limit this. Then, this larger region is divided into multiple smaller regions. Finally, the weakest region (i.e., some smaller regions) is selected from this larger region and an adversarial patch is added. In this way, while introducing more adversarial patches, the characteristic information of the image itself can be retained to the greatest extent, enabling the object detection model to learn more information about the image itself during the model training process and ensuring the recognition accuracy of the object detection model.

[0061] Specifically, adding corresponding adversarial patches to at least one first object among multiple objects in the original object image may include: determining the object regions corresponding to each first object in the original object image according to the position information of each first object in the original object image; performing random sampling on each object region respectively to determine the patch regions in each object region; dividing each patch region into multiple sub-regions and calculating the image gradients of each sub-region to obtain the gradients of each sub-region; selecting the target sub-regions of each patch region from all sub-regions according to the gradients of all sub-regions, and adding adversarial patches corresponding to the first object in the target sub-regions (i.e., the above-mentioned weakest regions). Taking the object (denoted as o 1 ) in the original object image p 1 as an example, in p 1 , the object information of o 1 is marked by a recognition bounding box, and the region where this recognition bounding box is located in p 1 is the object region of o 1 . In this way, in p 1Randomly sample the central positions of patch regions (abbreviated as "patches") within the object region of o. These central positions follow a normal distribution, and the side length of the patch is proportional to the bounding box size, thereby obtaining o 1 patch regions, and then divide the patch regions into n 2 sub-regions, calculate the gradients of each sub-region, and optionally select the sub-regions where the gradient norm corresponding to the gradient is greater than a preset threshold as target sub-regions and add adversarial patches corresponding to o 1 to all target sub-regions, or directly select the top regions with larger gradient norms as target sub-regions and only add adversarial patches corresponding to o 1 to these target sub-regions, thereby forming gradient-guided adversarial patches on p 1 . Among them, n is a positive integer. For the first object in the target original image p 1 where other adversarial patches need to be added, it is the same as o 1 , and the way of adding adversarial patches to the remaining target original images is the same as p 1 . For the sake of brevity, this article will not elaborate further. In this way, this adversarial method adds adversarial patches through random sampling and gradient selection, ensuring that the regions that have the greatest impact on the output of the target detection model are effectively added with adversarial patches, improving the efficiency and effectiveness of adversarial training.

[0062] This adversarial method may also include the determination process of the adversarial perturbations and adversarial patches added to the target original image: calculate the target gradient value according to the historical perturbations, the target original image, and the class information in the original samples. Among them, the historical perturbation is the preset initial perturbation when this training is the first training, and the historical perturbation is determined according to the adversarial perturbation and adversarial patch used in the previous training when this training is not the first training; determine the global noise according to the target gradient value, the preset global perturbation intensity, and the historical global noise, and use the global noise as the adversarial perturbation added to the target original image in this training. Among them, the historical global noise is the preset initial noise when this training is the first training, and the historical global noise is the global noise used in the previous training when this training is not the first training. And, determine the local patch according to the target gradient value, the preset perturbation patch step size, and the historical local patch, and use the local patch as the adversarial patch added to the target original image in this training. Among them, the historical local patch is the preset initial patch when this training is the first training, and the historical local patch is the local patch used in the previous training when this training is not the first training.

[0063] For example, the target gradient value can be determined according to Equation 2 below:

[0064]

[0065] In Equation 2, g adv represents the target gradient value, represents taking the expectation of the image distribution of the input target original image x, that is, calculating the loss function L θ of the target detector f d for the average effect over the entire data distribution B, represents taking the gradient of x with respect to the calculated L d δ represents the historical perturbation, and y represents the class information corresponding to x.

[0066] For example, the global noise can be determined according to Equation 3 below:

[0067] δ g ← δ g + ∈·sign(g adv ) Equation 3

[0068] In Equation 3, δ on the left side of ← g represents the global noise (i.e., the adversarial perturbation added to the target original image in the current training), and δ on the right side of ← g represents the historical global noise, ∈ represents the preset global perturbation intensity, and g adv represents the target gradient value, and sign(g adv ) represents that when g adv is positive, the sign function returns 1, when g adv is zero, it returns 0, and when g adv is negative, it returns -1.

[0069] For example, the local patch can be determined according to Equation 4 below:

[0070] δ g ← δ p + α·sign(g adv ) Equation 4

[0071] In Equation 4, δ on the left side of ← p represents the local patch (i.e., the adversarial patch added to the target original image in the current training), and δ on the right side of ← p represents the historical local patch, α represents the preset perturbation patch step size, and for the other parameters, refer to Equation 3 and will not be elaborated here.

[0072] For example, the historical perturbation can be determined according to Equation 5 below:

[0073] δ = δ p ⊙ m p + δ g Equation 5

[0074] In Equation 5, δ represents the historical perturbation (i.e., the total perturbation added to the target original image), and for the remaining parameters, refer to the previous text and will not be elaborated here. In this way, this adversarial method adds a total perturbation formed by adversarial perturbation and adversarial patch to the target original image, enabling the simulated attack to more comprehensively cover the potential attack space, enhancing the effect of adversarial training, and improving the robustness of the target detection model against different types of attacks.

[0075] For example, this adversarial method can update the parameters of the target detection model using the model gradient determined by the following Equation 6 in combination with an optimizer:

[0076]

[0077] In Equation 6, g θ represents the model gradient, x + δ represents the target original image added with adversarial perturbation and adversarial patch, y represents the class information of the target original image, and for the remaining parameters, refer to the previous text and will not be elaborated here.

[0078] This adversarial method can also include a correction process for adversarial perturbation and adversarial patch: in the case where the determined global noise is outside the preset perturbation range, modify the global noise to be within the preset perturbation range, where the preset perturbation range is determined according to the global perturbation intensity, and this process can be expressed as δ g ← clip(δ g , -∈, ∈), clip(δ g , -∈, ∈) means modifying the global noise δ g to be within the preset perturbation range [-∈, ∈], and ∈ refers to the previous text; and in the case where the determined local patch is outside the preset patch range, modify the local patch to be within the preset patch range, where the preset patch range is determined according to the preset patch perturbation intensity, and this process can be expressed as δ p ← clip(δ p , -β, β), clip(δ p , -∈, ∈) means modifying the local patch δ p to be within the preset patch range [-β, β], and β refers to the previous text.

[0079] Now, an example is given to illustrate the determination process of the adversarial perturbation and adversarial patch added to the target original image in combination with the above Equations 1 to 6.

[0080] In this example, the initial perturbation, initial noise, and initial patch are all initialized to zero, and the global perturbation intensity and perturbation patch step size are preset according to actual needs. For the first training, select the target original image p 1 to form the adversarial sample (denoted as a 1 ) required for the first training, and for the second training, select the target original image p2 to form the required adversarial example (denoted as a 2 ), in this example, the object detection model of AT is simply referred to as the model.

[0081] In the case where it is determined that the current training is the first training, an initial perturbation is added to p 1 to form a new image (denoted as p 1 ’), and p 1 ’ is input into the model (denoted as M 0 ) for object detection, and L 1 for p d ’ can be obtained as (f θ (x + δ), y). Then, combined with Equation 2 above, the target gradient value g adv1 can be calculated. Among them, x + δ substituted into Equation 2 is p 1 ’ (x is p 1 , δ is the initial perturbation), and y is the category information of p 1 ; then, substituting g adv1 , the global perturbation intensity, and the initial noise into Equation 3 above, the global noise δ g1 can be obtained, and substituting g adv1 , the perturbation patch step size, and the initial patch into Equation 4 above, the local patch δ p1 can be obtained. And in the case where it is determined that δ g1 is outside the preset perturbation range, δ g1 is corrected, and in the case where it is determined that δ p1 is outside the preset patch range, δ p1 is corrected; substituting the finally determined δ g1 and δ p1 into Equation 5 above, the total perturbation δ 1 can be obtained. Among them, m p substituted into Equation 5 is determined based on the target sub-region determined from p 1 according to two strategies: random sampling and gradient division. As Figure 3 shows, adding δ 1 to p 1 , that is, adding δ g1 to all regions of p 1 and adding δ p1 to the target sub-regions of each first object in p 1 . For example, Figure 4 shows that the area corresponding to the white dashed box is the patch area with the adversarial patch added, so as to obtain the training image (denoted as d 1 ) used for the first training. According to the object information corresponding to d 1 and d 1 (that is, the object information of p 1 ), a1 ; Finally, a can be utilized 1 to perform adversarial training on M 0 . Specifically, the model gradient g can be obtained by combining the above formula (6), θ1 where x + δ in formula (6) is substituted with d 1 , y is the class information of p 1 , and the target detection model parameters θ of M θ1 are updated using g 0 to meet the training objective shown in the above formula (1), and the model after the first training (denoted as M 1 ) is obtained.

[0082] When it is determined that the current training is the second training, historical perturbations (i.e., the total perturbation δ 2 used in the first training) are added to p 1 to form a new image (denoted as p 2 ’). Inputting p 2 ’ into M 1 for target detection can obtain L 2 (f d (x + δ), y) for p θ ’. Then, combining the above formula (2), the target gradient value g adv2 can be calculated, where x + δ in formula (2) is substituted with p 2 ’ (x is p 2 , δ is the historical perturbation δ 1 ), and y is the class information of p 2 ; Next, substituting g adv2 , the global perturbation intensity, and δ g1 used in the first training (i.e., the historical global noise) into the above formula (3) can obtain the global noise δ g2 , and substituting g adv2 , the perturbation patch step size, and δ p1 used in the first training (i.e., the historical local patch) into the above formula (4) can obtain the local patch δ p2 . The correction steps for δ g2 and δ p2 are the same as those in the first training and will not be elaborated here; Substituting the finally determined δ g2 and δ p2 into the above formula (5) can obtain the total perturbation δ 2 , where m p substituted in formula (5) is determined based on the target sub-regions determined from p 2 according to two strategies: random sampling and gradient division. Adding δ 2 to p 2 is the same as the process of adding δ 1 to p 1In (which will not be elaborated here again), the training images for the second training are obtained (denoted as d 2 ), according to d 2 and the object information corresponding to d 2 (that is, the object information of p 2 ), a 2 is obtained; finally, a 2 can be used to perform adversarial training on M 1 . Specifically, the model gradient g θ2 can be obtained by combining the above formula 6. Among them, x + δ in formula 6 is substituted with d 2 , y is the class information of p 2 , and the target detection model parameters θ of M θ2 are updated using g 1 until the training objective shown in the above formula 1 is satisfied, and the model after the second training is obtained (denoted as M 2 ).

[0083] The subsequent training process is the same as the second training process. For the sake of brevity, it will not be elaborated here. In addition, the basic settings of model training such as the number of training times and batch size are specifically set flexibly according to the actual situation, and the embodiments of the present disclosure do not make limitations on this.

[0084] In this way, this adversarial method adds adversarial patches guided by small-region gradients and human-eye-invisible adversarial perturbations in the global range during the model training process, enabling the model to effectively resist various attacks. Moreover, by updating the perturbations using the previously calculated gradient information, the additional training cost brought by internal maximization can be reduced, and the training efficiency is improved.

[0085] The target detection model provided by the embodiments of the present disclosure not only theoretically has strong adversarial robustness but also demonstrates excellent performance in practical applications. Through reasonable perturbation strategies and efficient calculation methods, the target detection model can maintain a relatively high average precision (AP) and stability when facing various actual attacks. The embodiments of the present disclosure first propose an adversarial method that can effectively defend against various physically realizable attacks, and the effect of this method on the target detector when facing powerful adaptive physically realizable attacks is much better than various existing defense methods.

[0086] In the embodiments of the present disclosure, the proposed adversarial method was tested on a detector formed by a Faster Region-based Convolutional Neural Networks (Faster RCNN). Three attacks, namely adversarial patches, adversarial textures, and adversarial camouflages, were used to test the stability of the detector. Meanwhile, a comparison was made with the basic model without AT, and the performance of the detector on two clean images was also tested. For the dataset, the detector was trained on the MS-COCO dataset, clean samples 1 and adversarial patches were generated on the Inria dataset, and adversarial textures and adversarial camouflages were generated on the Synthetic dataset. The test metric was represented by AP, and the relevant test results are shown in Table 1.

[0087] Table 1 Defense effects of the embodiments of the present disclosure under clean samples and various physically realizable attacks

[0088]

[0089] It can be found from the results shown in Table 1 that the adversarial method provided by the embodiments of the present disclosure can enhance the defense ability of the target detector against physically realizable attacks, while the performance degradation on clean images is minimal. The adversarial method of the embodiments of the present disclosure has better generality. Users neither need to make any modifications to the model nor perform additional processing on the images. The model after AT can be directly used in various practical application scenarios. After dividing and selecting image regions based on image gradients, the proposed adversarial method mixes adversarial patches and globally adversarial noises that are hardly noticeable to the human eye for AT, making the target detection model more robust and capable of withstanding various physically realizable attacks, including physically realizable adversarial patches and textures.

[0090] The embodiments of the present disclosure also provide an adversarial device for defending neural network attacks in the physical world, including: an acquisition module configured to acquire a target detection task for a target visual file, where the target detection task is used to indicate at least one target object to be detected, and the type of the target visual file is an image or a video; a calculation module configured to input the target visual file into a target detection model for calculation to obtain classification results and location information of the target objects; where the target detection model is trained using adversarial samples, and each adversarial sample is obtained by adding an adversarial perturbation and an adversarial patch to a corresponding original sample, the adversarial perturbation includes global noise set for the original sample, and the adversarial patch includes local patches set for objects in the original sample.

[0091] In this way, the adversarial device can construct high-quality adversarial samples based on global-range human-eye-invisible adversarial perturbations and small-region adversarial patches, and apply these adversarial samples to training, which can enhance the defense ability of the object detection model against physically realizable attacks, while not degrading the detection performance on clean images, so as to obtain the classification result (i.e., the category to which the target object belongs) and location information (i.e., the location of the target object in the target image) of the target object.

[0092] In a possible implementation, the device further includes a training module, configured to: obtain a plurality of original samples, each of the original samples including an original image and object information of each object in the original image, and each of the object information including category information and location information; determine adversarial samples required for the current training based on each of the original samples; perform model training based on the current adversarial samples until the model converges, so as to obtain the object detection model.

[0093] In this way, the object detection model in this adversarial device will learn the features of the adversarial samples during the training process, thereby enhancing the recognition ability of the adversarial samples to improve the defense ability of the object detector.

[0094] In a possible implementation, determining the adversarial samples required for the current training based on each of the original samples includes: selecting a target original image required for the current training from the original images of each of the original samples; adding an adversarial perturbation to the target original image, and adding an adversarial patch to the target original image; using the target original image added with the adversarial perturbation and the adversarial patch as the training image for the current training, and obtaining the adversarial samples required for the current training according to the training image and the object information corresponding to the training image.

[0095] In a possible implementation, adding an adversarial patch to the target original image includes: adding a corresponding adversarial patch for at least one first object among a plurality of objects in the target original image.

[0096] In this way, by adding adversarial patches to some objects in the target original image, this adversarial device can balance the adversarial ability and recognition ability of the object detection model.

[0097] In a possible implementation, for at least one first object among multiple objects in the target original image, adding a corresponding adversarial patch includes: determining, according to the position information of each of the first objects in the target original image, an object region corresponding to each of the first objects in the target original image; randomly sampling each of the object regions to determine a patch region in each of the object regions; dividing each of the patch regions into multiple sub-regions, and calculating the image gradient for each of the sub-regions to obtain the gradient of each of the sub-regions; selecting a target sub-region of each of the patch regions from all the sub-regions according to the gradients of all the sub-regions, and adding an adversarial patch corresponding to the first object in the target sub-region.

[0098] In this way, when generating an adversarial patch, this adversarial device uses two strategies, namely random sampling and division according to the gradient, which can make the adversarial sample attack the weakest part of the image as much as possible. While introducing more adversarial patches, it maximally retains the feature information of the image itself, enabling the target detection model to learn more information about the image itself during the model training process and ensuring the recognition accuracy of the target detection model.

[0099] In a possible implementation, the device further includes a determination module for: calculating a target gradient value according to the historical perturbation, the target original image, and the class information in the original sample, where the historical perturbation is a preset initial perturbation when the current training is the first training, and the historical perturbation is determined according to the adversarial perturbation and the adversarial patch used in the previous training when the current training is not the first training; determining a global noise according to the target gradient value, a preset global perturbation intensity, and the historical global noise, and using the global noise as the adversarial perturbation added to the target original image, where the historical global noise is a preset initial noise when the current training is the first training, and the historical global noise is the global noise used in the previous training when the current training is not the first training, and determining a local patch according to the target gradient value, a preset perturbation patch step size, and the historical local patch, and using the local patch as the adversarial patch added to the target original image, where the historical local patch is a preset initial patch when the current training is the first training, and the historical local patch is the local patch used in the previous training when the current training is not the first training.

[0100] In this way, the present adversarial device adds a total perturbation formed by an adversarial perturbation and an adversarial patch to the target original image, enabling the simulated attack to more comprehensively cover the potential attack space, enhancing the effect of adversarial training, improving the robustness of the target detection model against different types of attacks, enabling the model to effectively resist multiple attacks, and at the same time updating the perturbation using the previously calculated gradient information, which can reduce the additional training cost brought by the internal maximization and improve the training efficiency.

[0101] In a possible implementation manner, the device further includes a modification module configured to: modify the global noise when the determined global noise is outside a preset perturbation range, where the preset perturbation range is determined according to the global perturbation intensity; and modify the local patch when the determined local patch is outside a preset patch range, where the preset patch range is determined according to a preset patch perturbation intensity.

[0102] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0103] The embodiments of the present disclosure also propose a computer-readable storage medium storing computer program instructions thereon, and the computer program instructions, when executed by a processor, implement the above method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0104] The embodiments of the present disclosure also propose an electronic device including: a processor; and a memory for storing processor-executable instructions, where the processor is configured to implement the above method when executing the instructions stored in the memory.

[0105] The embodiments of the present disclosure also provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0106] Figure 5 The block diagram of the adversarial device provided by the embodiments of the present disclosure for defending against attacks on neural networks in the physical world is shown. For example, the device 1900 can be provided as a server or a terminal device. Refer to Figure 5, the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0107] The apparatus 1900 may also include a power component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0108] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method.

[0109] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0110] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0111] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0112] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0113] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0114] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create an apparatus that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0115] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0116] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.

[0117] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for defending a neural network from attacks by the physical world, characterized in that: include: Acquire a target detection task for a target visual file, where the target detection task is used to indicate at least one target object to be detected; Inputting the target visual file into the target detection model for calculation to obtain the classification result and position information of each target object; Among them, the target detection model is obtained by training using adversarial samples, and each adversarial sample is obtained by adding adversarial perturbations and adversarial patches to the corresponding original samples. The adversarial perturbations include global noise set for the original samples, and the adversarial patches include local patches set for the objects in the original samples.

2. The method according to claim 1, characterized in that The method further comprises: Acquire multiple original samples, each of the original samples includes an original image and object information of each object in the original image, and each of the object information includes category information and position information; Based on the original samples, determining adversarial samples required for current training; The model is trained based on the current adversarial sample until the model converges to obtain the target detection model.

3. The method according to claim 2, characterized in that Based on the original samples, adversarial samples required for the current training are determined, including: Selecting a target original image required for current training from the original images of the original samples; adding an adversarial perturbation to the target original image, and adding an adversarial patch to the target original image; The target original image to which the adversarial perturbation and the adversarial patch are added is used as a training image for current training, and adversarial samples required for current training are obtained according to the training image and object information corresponding to the training image.

4. The method according to claim 3, characterized in that Adding an adversarial patch to the target original image includes: For at least one first object among a plurality of objects in the target original image, a corresponding adversarial patch is added.

5. The method according to claim 3, characterized in that: For at least one first object among the plurality of objects in the target original image, adding a corresponding adversarial patch comprises: Determining, according to the position information of each of the first objects in the target original image, an object area corresponding to each of the first objects in the target original image; Randomly sampling each of the object regions to determine a patch region in each of the object regions; Dividing each of the patch regions into a plurality of sub-regions, and performing image gradient calculation on each of the sub-regions to obtain a gradient of each of the sub-regions; A target sub-region of each patch region is selected from all sub-regions according to the gradients of all sub-regions, and an adversarial patch corresponding to the first object is added to the target sub-region.

6. The method according to any one of claims 3 to 5, characterized in that The method further comprises: Calculating a target gradient value according to the historical disturbance, the target original image, and the category information in the original sample, wherein the historical disturbance is a preset initial disturbance when the current training is the first training, and the historical disturbance is determined according to the adversarial disturbance and adversarial patch used in the previous training when the current training is not the first training; Determine the global noise according to the target gradient value, the preset global perturbation intensity, and the historical global noise, and use the global noise as the adversarial perturbation added to the target original image, wherein the historical global noise is the preset initial noise when the current training is the first training, and the historical global noise is the global noise used in the last training when the current training is not the first training, and A local patch is determined according to the target gradient value, a preset perturbation patch step size, and a historical local patch, and the local patch is used as an adversarial patch added to the target original image, wherein the historical local patch is a preset initial patch when the current training is the first training, and the historical local patch is a local patch used in the previous training when the current training is not the first training.

7. The method according to claim 6, characterized in that The method further comprises: In a case where the determined global noise is outside a preset disturbance range, modifying the global noise, wherein the preset disturbance range is determined according to the global disturbance intensity; and In the case that the determined local patch is outside a preset patch range, the local patch is modified, wherein the preset patch range is determined according to a preset patch disturbance intensity.

8. A countermeasure device for defending a neural network from attacks by the physical world, characterized in that: include: An acquisition module, used for acquiring a target detection task for a target visual file, wherein the target detection task is used for indicating at least one target object to be detected, wherein the type of the target visual file is an image or a video; A calculation module, used for inputting the target visual file into the target detection model for calculation, and obtaining the classification result and position information of each target object; Among them, the target detection model is obtained by training using adversarial samples, and each adversarial sample is obtained by adding adversarial perturbations and adversarial patches to the corresponding original samples. The adversarial perturbations include global noise set for the original samples, and the adversarial patches include local patches set for the objects in the original samples.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method described in any one of claims 1 to 7 when executing the instructions stored in the memory.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.