Optimization method and device of adversarial patch, equipment and storage medium
By combining the optimization loss function and the detection model, the generated adversarial patch solves the problem of insufficient concealment, achieves high naturalness and high attack in the real environment, and improves the application effect of adversarial patches.
Patent Information
- Application Number
- CN202510162870.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-04
AI Technical Summary
The existing adversarial patching methods are not concealed in the real environment, and the attack effect is affected by the external environment, resulting in limited application.
By obtaining the preset detection model, stable diffusion model and data set, setting the text alignment loss function, latent space alignment loss function and target category confidence score, generating an optimization loss function for the adversarial patch, and obtaining the optimized adversarial patch through the stabilization diffusion model, and verifying it with the detection model.
The generated adversarial patches are highly consistent with the text description in detail, improving nature, concealment and practicality, making them more effective and difficult to detect in real environments.
Smart Images

Figure CN120255918A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of adversarial machine learning technology. Specifically, it relates to an optimization method, device, equipment, and storage medium for adversarial patches. Background Art
[0002] An adversarial patch is a common physical attack method. By pasting a carefully designed adversarial patch onto a target, the detector cannot correctly identify the target.
[0003] Existing adversarial patch methods ignore the concealment of adversarial textures in the real application environment, making adversarial patches easily detectable by the naked eye, which restricts the application of related technologies. The current existing methods for improving the concealment of adversarial patches mainly introduce specific targets (such as cats, dogs, flowers) into the adversarial texture, making the adversarial pattern possess the visual features of the introduced target. Although the finally formed texture improves the naturalness compared to the unrestricted attack pattern, there is still an inconsistent phenomenon between the pattern and the environmental background. For example, an adversarial patch containing a dog appears in an oasis or desert. In addition, in actual applications, the attack effect of adversarial patches will be affected by the external environment (light, imaging distance, background), resulting in a decline in attack performance.
[0004] Therefore, how to design adversarial patches with high concealment and high aggressiveness is the key to the practical application of adversarial patches. Summary of the Invention
[0005] According to one aspect of the present application, an optimization method for adversarial patches is provided, including: obtaining a preset detection model, a stable diffusion model, and a dataset; setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimization loss function for adversarial patches; obtaining an optimized adversarial patch through the stable diffusion model based on the optimization loss function and the dataset; and verifying the optimized adversarial patch through the detection model.
[0006] According to some embodiments, setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimization loss function for adversarial patches includes: obtaining a preset input text; obtaining a first latent space variable generated by the stable diffusion model; generating a cross-attention map between the first latent space variable and the input text; and setting the text alignment loss function according to the cross-attention map.
[0007] According to some embodiments, setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimization loss function for adversarial patches includes: obtaining a second latent space variable corresponding to the adversarial patch through the stable diffusion model; and setting the latent space alignment loss function according to the second latent space variable.
[0008] According to some embodiments, a text alignment loss function, a latent space alignment loss function, and a target class confidence score are set to generate an optimized loss function for adversarial patches, including: obtaining the target class confidence score through a detection model.
[0009] According to some embodiments, based on the optimized loss function and a dataset, an optimized adversarial patch is obtained through a StableDiffusion model, including: setting a fusion method for the adversarial patch based on the dataset; iteratively optimizing the latent space of the StableDiffusion model through the detection model and the StableDiffusion model based on the optimized loss function and the fusion method; and obtaining the optimized adversarial patch through the latent space when the number of iterations for iteratively optimizing the latent space of the StableDiffusion model reaches a preset number of iterations.
[0010] According to some embodiments, setting a fusion method for the adversarial patch based on the dataset includes: obtaining target position information in the images of the dataset; obtaining an intermediate adversarial patch; generating an adversarial patch image based on the intermediate adversarial patch according to the target position information and a preset adversarial patch size ratio; and pixel-wise fusing the adversarial patch image with the images of the dataset.
[0011] According to some embodiments, verifying the optimized adversarial patch through a detection model includes: generating a target image containing the optimized adversarial patch based on the dataset; and detecting the target image through the detection model to verify the effect of the optimized adversarial patch.
[0012] According to one aspect of the present application, an optimization device for adversarial patches is provided, including: a preparation module for obtaining a preset detection model, a StableDiffusion model, and a dataset; a configuration module for setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for adversarial patches; an optimization module for obtaining an optimized adversarial patch through the StableDiffusion model based on the optimized loss function and the dataset; and a verification module for verifying the optimized adversarial patch through the detection model.
[0013] According to one aspect of the present application, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0014] According to one aspect of the present application, a computer-readable storage medium is provided, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the method as described above is implemented.
[0015] According to an embodiment of the present application, text-guided adversarial patch generation is achieved by attacking the StableDiffusion model, and the generation process of the adversarial patch is further refined through the latent space alignment module loss function, which not only ensures that the generated adversarial patch is highly consistent with the text description in details, but also improves the naturalness, concealment and practicality of the adversarial patch in the real environment, making it more effective and imperceptible in practical applications.
[0016] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application.
[0018] Figure 1 The flowchart showing an optimization method of an adversarial patch according to an exemplary embodiment of the present application.
[0019] Figure 2 The block diagram showing an optimization device of an adversarial patch according to an exemplary embodiment of the present application.
[0020] Figure 3 The block diagram showing an electronic device according to an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals refer to like or similar parts throughout the figures, and thus their repeated description will be omitted.
[0022] The described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of these specific details, or can be implemented in other ways, components, materials, devices or operations, etc. In these cases, well-known structures, methods, devices, implementations, materials or operations will not be shown or described in detail.
[0023] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0024] The terms "first", "second", etc. in the description and claims of this application and the above-mentioned accompanying drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0025] This application provides an optimization method, device, equipment, and storage medium for adversarial patches, which can improve the naturalness, concealment, and practicality of adversarial patches in a real environment.
[0026] Next, with reference to the accompanying drawings, a detailed description will be given of an optimization method, device, equipment, and storage medium for adversarial patches according to an embodiment of this application.
[0027] Figure 1 A flowchart showing an optimization method for adversarial patches according to an exemplary embodiment of this application is shown.
[0028] As Figure 1 shown, in step S100, a preset detection model, a stable diffusion model, and a data set are obtained.
[0029] For example, in step S100, the optimization device respectively obtains a preset detection model, a stable diffusion model, and a data set for the generation and optimization of adversarial patches.
[0030] The optimization device obtains a preset detection model.
[0031] According to some embodiments, the preset detection model can adopt a pre-trained YOLOv5 object detection model.
[0032] The optimization device obtains a preset stable diffusion model to encode the text, decode the latent space representation generated by the stable diffusion model, and gradually denoise the Gaussian noise in the latent space.
[0033] The optimization device obtains a preset data set.
[0034] According to some embodiments, the preset dataset may adopt the INRIA pedestrian dataset, which is divided into a training set and a test set. Among them, the training set is used as the base map for pasting adversarial patches to train the adversarial patches, including 614 pictures. The test set is used to evaluate the attack effect of the adversarial patches, including 288 pictures. The size of each picture is 416×416×3.
[0035] In step S200, a text alignment loss function, a latent space alignment loss function, and a target class confidence score are set to generate an optimized loss function for the adversarial patch.
[0036] For example, in step S200, the optimization device generates an optimized loss function for the adversarial patch based on the set text alignment loss function, latent space alignment loss function, and target class confidence score.
[0037] The optimization device first obtains the preset input text to be used as the prompt for the input stable diffusion model. For example, the input text is "a picture full of leaf-like green color".
[0038] The optimization device obtains the first latent space variable generated by the stable diffusion model and generates a cross-attention map between the first latent space variable and the input text.
[0039] According to some embodiments, the optimization device gradually denoises Gaussian noise in the latent space through the stable diffusion model and generates the first latent space variable during the process of gradual denoising. For example
[0040] Furthermore, the optimization device generates and saves a matrix map composed of the cross-attention between the first latent space variable and the input text, that is, the cross-attention map between the first latent space variable and the input text. Among them, the cross-attention between the first latent space variable and the input text reflects the correlation between the latent space variable and the text, and the larger the value, the stronger the correlation.
[0041] The optimization device sets a text alignment loss function according to the cross-attention map between the first latent space variable and the input text so that the generated adversarial patch conforms to the description of the input text.
[0042] According to some embodiments, the cross-attention map between the first latent space variable and the input text can ensure that the generated adversarial patch contains the latent features of the input text. The optimization device aligns the cross-attention maps continuously generated between the first latent space variable and the input text with the initial cross-attention map and obtains the text alignment loss function As follows.
[0043]
[0044] Among them, represents the j-th cross-attention map of each first latent space variable , and represents the corresponding initial cross-attention map of the first latent space variable .
[0045] The optimization device obtains the second latent space variable corresponding to the adversarial patch through the stable diffusion model, and sets the latent space alignment loss function based on this.
[0046] According to some embodiments, the optimization device can obtain the second latent space variable when generating the adversarial patch through the stable diffusion model. Among them, the optimization device obtains the initial second latent space variable when generating the adversarial patch for the first time and then obtains the second latent space variable every time the adversarial patch is generated The optimization device uses the initial second latent space variable to constrain the latent space and obtain the latent space alignment loss function to generate a relatively natural adversarial patch. The latent space alignment loss function can be expressed by the following formula.
[0047]
[0048] Among them, the latent space alignment loss function represents the loss of the latent space.
[0049] The optimization device also obtains the target class confidence score through the detection model.
[0050] According to some embodiments, the optimization device can detect the image pasted with the adversarial patch through the detection model, and output the target class probability, target bounding box information, and the confidence score F obj( I A ) in the image. The optimization device sets the target class confidence score according to F obj (I A ) as follows. As follows.
[0051]
[0052] According to the set text alignment loss function, latent space alignment loss function, and detection confidence score, the optimization device generates the optimization loss function of the adversarial patch.
[0053] According to some embodiments, the optimization loss function of the adversarial patch can be expressed by the following formula.
[0054]
[0055] Among them, α, β, and γ respectively represent the weights of the text alignment loss function, the latent space alignment loss function, and the weight of the target class confidence score.
[0056] According to some embodiments, the optimized loss function of the adversarial patch is constrained by multiple aspects of the prompt, cross-attention mechanism, latent space, and detection model, ensuring the naturality and semantic information of the adversarial patch while taking into account the aggressiveness of the adversarial patch.
[0057] In step S300, based on the optimized loss function and the dataset, an optimized adversarial patch is obtained through the StableDiffusion model.
[0058] For example, in step S300, the optimization device sets the fusion method of the adversarial patch based on a preset dataset, and iteratively optimizes the latent space of the StableDiffusion model according to the optimized loss function and the fusion method of the adversarial patch to obtain an optimized adversarial patch.
[0059] Based on a preset dataset, the optimization device sets the fusion method of the adversarial patch.
[0060] According to some embodiments, the optimization device first obtains the target position information of each image in the dataset and a preset size ratio of the adversarial patch.
[0061] Furthermore, the optimization device obtains an intermediate adversarial patch x' through the latent space of the StableDiffusion model, and rotates and scales the intermediate adversarial patch x' through an affine transformation matrix according to the target position information and the size ratio of the adversarial patch to generate an adversarial patch image I with an adversarial patch. M .
[0062] According to some embodiments, the adversarial patch image I M can be expressed by the following formula.
[0063] I M = S × R(θ) × x', (5)
[0064] Among them, S represents a scaling matrix for scaling the adversarial patch. R(θ) represents a rotation matrix for rotating the adversarial patch.
[0065] The scaling matrix S can be expressed by the following formula.
[0066]
[0067] Among them, S x and S y respectively represent the scaling coefficients corresponding to the x and y axes.
[0068] The rotation matrix R(θ) can be expressed by the following formula.
[0069]
[0070] The optimization device fuses the generated adversarial patch image with the images in the dataset pixel by pixel and obtains the input data for the detection model.
[0071] According to some embodiments, the optimization device takes the adversarial patch image I M and the image I r in the training set of the dataset and fuses them pixel by pixel at positions i, j to obtain the adversarial input data I A , which is used as the input data for the detection model. The adversarial input data I A can be expressed by the following formula.
[0072]
[0073] According to the optimization loss function and fusion method of the adversarial patch, the optimization device combines the detection model and the StableDiffusion model and performs iterative optimization with the latent space of the StableDiffusion model as the optimization target.
[0074] According to some embodiments, the optimization device generates an adversarial patch through the latent space of the StableDiffusion model, fuses the image containing the adversarial patch with the images in the dataset to obtain the input data for the detection model. Furthermore, the optimization device obtains the output result of the detection model based on the input data and adjusts the generation of the adversarial patch accordingly, that is, performs iterative optimization on the latent space of the StableDiffusion model.
[0075] When the number of iterations for iterative optimization of the latent space of the StableDiffusion model reaches the preset number of iterations (e.g., 200 times), the optimization device stops optimizing the latent space of the StableDiffusion model. At this time, the optimization device generates an adversarial patch through the latent space of the StableDiffusion model and uses it as the final optimized adversarial patch.
[0076] In step S400, the optimized adversarial patch is verified by the detection model.
[0077] For example, in step S400, the optimization device generates a target image containing the optimized adversarial patch and detects the target image through the detection model.
[0078] Based on the preset dataset, the optimization device generates a target image containing the optimized adversarial patch.
[0079] According to some embodiments, the optimization device fuses the image containing the optimized adversarial patch with the images in the dataset to generate a target image and pastes the adversarial patch image at the corresponding target position in the dataset image.
[0080] The optimization device detects images in the dataset of the pasted adversarial patch images through the detection model to verify the effect of the optimized adversarial patch.
[0081] According to some embodiments, the effects of the optimized adversarial patch include the unity of the adversarial patch and the target environment and the detection evasion effect of the adversarial patch.
[0082] According to the embodiments of the present application, it can be ensured that the generated adversarial patch is highly consistent with the text description in details, improving the naturalness, concealment and practicality of the adversarial patch in the real environment.
[0083] Figure 2 The block diagram of an optimization device for an adversarial patch according to an exemplary embodiment of the present application is shown.
[0084] As Figure 2 shown, the optimization device 100 includes a preparation module 110, a configuration module 120, an optimization module 130 and a verification module 140.
[0085] The preparation module 110 respectively obtains a preset detection model, a stable diffusion model and a dataset.
[0086] The configuration module 120 obtains a preset input text to be used as a prompt for the input stable diffusion model.
[0087] The configuration module 120 obtains the first latent space variable generated by the stable diffusion model and generates a cross-attention map between the first latent space variable and the input text.
[0088] The configuration module 120 sets a text alignment loss function according to the cross-attention map between the first latent space variable and the input text so that the generated adversarial patch conforms to the description of the input text.
[0089] The configuration module 120 obtains the second latent space variable corresponding to the adversarial patch through the stable diffusion model and sets a latent space alignment loss function based on this.
[0090] The configuration module 120 obtains the target class confidence score through the detection model.
[0091] According to the set text alignment loss function, latent space alignment loss function and detection confidence score, the configuration module 120 generates an optimization loss function for the adversarial patch.
[0092] Based on the preset dataset, the optimization module 130 sets the fusion method of the adversarial patch.
[0093] According to the optimization loss function and fusion method of the adversarial patch, the optimization module 130 combines the detection model and the stable diffusion model and performs iterative optimization with the latent space of the stable diffusion model as the optimization target.
[0094] When the number of iterations for the iterative optimization of the latent space of the Stable Diffusion model reaches the preset number of iterations, the optimization module 130 stops optimizing the latent space of the Stable Diffusion model and generates an adversarial patch through the latent space of the Stable Diffusion model, which is used as the final optimized adversarial patch.
[0095] Based on a preset data set, the verification module 140 generates a target image containing the optimized adversarial patch.
[0096] The verification module 140 detects the images in the data set of the images with the pasted adversarial patch through the detection model to verify the effect of the optimized adversarial patch.
[0097] Figure 3 A block diagram of an electronic device according to an exemplary embodiment of the present application is shown.
[0098] As Figure 3 shown, the electronic device 600 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0099] As Figure 3 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc. Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the methods according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 610 can execute the method as Figure 1 shown in
[0100] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0101] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0102] The bus 630 can represent one or more of several types of bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, an Accelerated Graphics Port, a processing unit, or a local bus using any of the various bus structures.
[0103] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0104] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. The technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0105] The software product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0106] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0107] The program code for performing the operations of the present application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0108] The above computer-readable medium bears one or more programs, which, when executed by a device, cause the computer-readable medium to implement the foregoing functions.
[0109] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from the embodiments. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0110] The above has introduced the embodiments of the present application in detail. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, those skilled in the art, based on the idea of the present application, the changes or deformations made in the specific implementation manner and application scope of the present application all belong to the protection scope of the present application. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An optimization method for anti-patches, characterized in that, Including: Obtain a preset detection model, a stable diffusion model, and a dataset; Set a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for the adversarial patch; Based on the optimized loss function and the dataset, obtain an optimized adversarial patch through the stable diffusion model; Verify the optimized adversarial patch through the detection model.
2. The method according to claim 1, wherein Setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for the adversarial patch includes: Obtain a preset input text; Obtain a first latent space variable generated by the stable diffusion model; Generate a cross-attention map between the first latent space variable and the input text; Set the text alignment loss function according to the cross-attention map.
3. The method according to claim 1, characterized in that, Setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for the adversarial patch includes: Obtain a second latent space variable corresponding to the adversarial patch through the stable diffusion model; Set the latent space alignment loss function according to the second latent space variable.
4. The method according to claim 1, characterized in that Setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for the adversarial patch includes: Obtain the target class confidence score through the detection model.
5. The method according to claim 1, wherein Based on the optimized loss function and the dataset, obtaining an optimized adversarial patch through the stable diffusion model includes: Based on the dataset, set the fusion method of the adversarial patch; Based on the optimized loss function and the fusion method, iteratively optimize the latent space of the stable diffusion model through the detection model and the stable diffusion model; When the number of times of iteratively optimizing the latent space of the stable diffusion model reaches a preset number of iterations, obtain the optimized adversarial patch through the latent space.
6. The method according to claim 5, characterized in that, Based on the dataset, setting the fusion method of the adversarial patch includes: Obtain the target position information in the pictures of the dataset; Obtain an intermediate adversarial patch; Based on the intermediate adversarial patch, generate an adversarial patch picture according to the target position information and a preset adversarial patch size ratio; Perform pixel fusion on the adversarial patch picture and the pictures of the dataset.
7. The method according to claim 1, wherein Verifying the optimized adversarial patch through the detection model includes: Based on the dataset, generate a target picture containing the optimized adversarial patch; Detect the target picture through the detection model to verify the effect of the optimized adversarial patch.
8. An optimization device for anti-patch, characterized in that, Including: A preparation module for obtaining a preset detection model, a stable diffusion model, and a dataset; A configuration module for setting a text alignment loss function, a latent space alignment loss function, and a target class confidence score to generate an optimized loss function for the adversarial patch; An optimization module for obtaining an optimized adversarial patch through the stable diffusion model based on the optimized loss function and the dataset; A verification module for verifying the optimized adversarial patch through the detection model.
9. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions, when executed by a processor, implement the method according to any one of claims 1-7.