Adversarial patch generation method and device, equipment and storage medium

By generating and optimizing adversarial patches, the problem of insufficient stealth and multi-angle attack capabilities in existing technologies is solved, achieving high stealth and strong attack effects in different background environments.

CN120852645APending Publication Date: 2025-10-28启元实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510859999.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing anti-patch techniques are insufficient in terms of stealth and multi-angle attack capabilities, are easily identified, and their effectiveness is affected by external environmental factors.

Method used

By acquiring a reference image of the target environment, generating a feature vector corresponding to the target environment, generating an adversarial patch using a stable diffusion model, rendering it onto a 3D model, and optimizing it in conjunction with a detection model, the adversarial patch is ensured to have high concealment and strong attack power in different background environments.

Benefits of technology

It improves the stealth and attack effectiveness of anti-patch in real-world environments, adapts to different perspectives and backgrounds, and enhances robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852645A_ABST
    Figure CN120852645A_ABST
Patent Text Reader

Abstract

The invention provides an adversarial patch generation method and device, equipment and a storage medium, and relates to the technical field of adversarial machine learning. The adversarial patch generation method comprises the following steps: acquiring a target environment reference image; generating a first feature vector corresponding to the target environment according to the target environment reference image; generating an adversarial patch through a preset stable diffusion model based on the first feature vector; rendering the adversarial patch to a preset 3D model to obtain adversarial samples of different background environments; and based on the adversarial sample, optimizing the adversarial patch through a preset detection model. According to the embodiment of the invention, the concealment and practicability of the adversarial patch in a real environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of adversarial machine learning technology, and more specifically, to a method, apparatus, device, and storage medium for generating adversarial patches. Background Technology

[0002] Counter-patching is a common physical attack technique that involves attaching carefully designed patches to a target, making it impossible for target detectors to correctly identify the target.

[0003] Existing adversarial patching techniques have significant shortcomings in terms of stealth and multi-angle attack capabilities. First, poor stealth makes adversarial patches easily detectable by the human eye in real-world environments. Second, existing methods are typically only effective at specific angles; their effectiveness decreases significantly as the viewing angle changes. In practical applications, the effectiveness of adversarial patches is also affected by external environmental factors, further weakening their stealth and attack performance.

[0004] Therefore, developing a highly covert, highly offensive, and multi-faceted countermeasure patch is the key to realizing the practical application of countermeasure patch technology. Summary of the Invention

[0005] According to one aspect of this application, a method for generating an adversarial patch is provided, comprising: acquiring a target environment reference image; generating a first feature vector corresponding to the target environment based on the target environment reference image; generating an adversarial patch based on the first feature vector using a preset stable diffusion model; rendering the adversarial patch onto a preset 3D model to obtain adversarial samples of different background environments; and optimizing the adversarial patch based on the adversarial samples using a preset detection model.

[0006] According to some embodiments, obtaining a target environment reference image includes: obtaining image datasets corresponding to different background environments; and determining the target environment reference image in the image dataset.

[0007] According to some embodiments, generating a first feature vector corresponding to a target environment based on a target environment reference image includes: obtaining a preset pre-trained model; encoding the target environment reference image into a visual feature space through the visual encoder of the pre-trained model to generate the first feature vector.

[0008] According to some embodiments, an adversarial patch is generated based on a first feature vector using a preset stable diffusion model, including: obtaining latent space variables based on the first feature vector using a stable diffusion model; and decoding the latent space variables into pixel space using a stable diffusion model to obtain the adversarial patch.

[0009] According to some embodiments, generating an adversarial patch based on a first feature vector using a preset stable diffusion model further includes: obtaining a preset pre-trained model; encoding the adversarial patch into a visual feature space using a visual encoder of the pre-trained model to generate a second feature vector; and obtaining an alignment loss function between the first feature vector and the second feature vector.

[0010] According to some embodiments, the 3D model includes a human body model and a shirt and / or pants model; rendering adversarial patches onto a preset 3D model to obtain adversarial samples for different background environments includes: setting texture information for the shirt and / or pants model based on the adversarial patches; fusing the human body model and the shirt and / or pants model with set texture information to obtain a complete 3D model; and obtaining adversarial samples based on the complete 3D model.

[0011] According to some embodiments, obtaining adversarial examples based on a complete 3D model includes: obtaining preset mapping parameters for mapping the 3D model to a 2D plane; projecting the complete 3D model onto the 2D plane according to the preset mapping parameters to obtain a 2D image corresponding to the complete 3D model; and generating adversarial examples based on image datasets corresponding to different background environments and the 2D image corresponding to the complete 3D model.

[0012] According to some embodiments, adversarial patches are optimized based on adversarial examples using a preset detection model, including: inputting adversarial examples into the detection model to obtain the target detection confidence level corresponding to the adversarial examples; setting a confidence loss function based on the target detection confidence level corresponding to the adversarial examples; obtaining the total loss function corresponding to the adversarial patches based on the alignment loss function and the confidence loss function; and optimizing the adversarial patches using the total loss function.

[0013] According to one aspect of this application, an adversarial patch generation apparatus is provided, comprising: a first execution module for acquiring a target environment reference image; a second execution module for generating a first feature vector corresponding to the target environment based on the target environment reference image; a third execution module for generating an adversarial patch based on the first feature vector using a preset stable diffusion model; a fourth execution module for rendering the adversarial patch onto a preset 3D model to obtain adversarial samples of different background environments; and a fifth execution module for optimizing the adversarial patch based on the adversarial samples using a preset detection model.

[0014] According to one aspect of this application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to perform the method as described above.

[0015] According to one aspect of this application, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the method as described above.

[0016] According to the embodiments of this application, one or more of the following beneficial effects can be achieved:

[0017] 1. By generating adversarial patches through a stable diffusion model and performing progressive denoising in the latent space, natural and subtle adversarial textures are generated. This avoids the problem of overly bright colors in traditional adversarial patches and improves the concealment and practicality of adversarial patches in real-world environments.

[0018] 2. Rendering the adversarial patch onto a 3D model allows the adversarial patch to adapt to different perspectives and background environments, improving the attack effectiveness and robustness of the adversarial patch in real-world scenarios.

[0019] 3. Align the feature vectors of the adversarial patch with the feature vectors of the target environment reference image in the visual space to ensure that the generated adversarial patch is highly similar to the environment, and ensure that the adversarial patch has strong offensive capabilities by minimizing the confidence score of the detection model.

[0020] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application.

[0022] Figure 1 A flowchart illustrating a method for generating an adversarial patch according to an example embodiment of this application is shown.

[0023] Figure 2 A schematic diagram of a natural environment image according to an example embodiment of this application is shown.

[0024] Figure 3 A schematic diagram of an urban environment image according to an example embodiment of this application is shown.

[0025] Figure 4 A schematic diagram of a target environment reference image according to an example embodiment of this application is shown.

[0026] Figure 5 A schematic diagram of a 2D image formed on a 2D plane of a complete 3D model according to an example embodiment of this application is shown.

[0027] Figure 6 A schematic diagram of an adversarial sample according to an example embodiment of this application is shown.

[0028] Figure 7 A schematic diagram of an apparatus for generating adversarial patches according to an example embodiment of this application is shown.

[0029] Figure 8 A block diagram of an electronic device according to an example embodiment of this application is shown. Detailed Implementation

[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. Like reference numerals in the drawings represent like or similar parts, and thus repetitive description thereof will be omitted.

[0031] The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of these specific details, or other methods, components, materials, apparatus, or operations may be employed. In these cases, well-known structures, methods, apparatuses, implementations, materials, or operations will not be shown or described in detail.

[0032] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0033] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0034] This application provides a method, apparatus, device, and storage medium for generating anti-patterns, which can improve the stealth and practicality of anti-patterns in real-world environments.

[0035] The following is a detailed description of a method, apparatus, device, and storage medium for generating anti-patch according to embodiments of this application, with reference to the accompanying drawings.

[0036] Figure 1 A flowchart illustrating a method for generating an adversarial patch according to an example embodiment of this application is shown.

[0037] like Figure 1 As shown, in step S100, a reference image of the target environment is obtained.

[0038] For example, in step S100, the generating device obtains a target environment reference image through a preset background environment image dataset.

[0039] The generating device acquires a preset detection model for detecting targets in different background environments.

[0040] According to some embodiments, the detection model may employ any of the pre-trained YOLOv2, YOLOv3, YOLOv4, YOLOv5, and Faster-RCNN detection models.

[0041] The generation device acquires a preset stable diffusion model, which is used to decode the latent space variables generated by the stable diffusion model and to progressively denoise the Gaussian noise in the latent space.

[0042] The generating device acquires a preset background environment image dataset, which includes image datasets corresponding to various different background environments.

[0043] According to some embodiments, the image dataset includes, for example: Figure 2 The natural environment images shown, and such as Figure 3 The image shown is of an urban environment. The generation device divides the image dataset into a training set and a test set. The number of images in the training set and the test set can be adjusted according to actual needs. For example, the training set may include 379 images, and the test set may include 130 images.

[0044] The generating device determines a reference image of the target environment from the image dataset.

[0045] According to some embodiments, the generation device randomly selects an image from an image dataset as a target environment reference image. Different detection models can correspond to different target environment reference images. For example, ... Figure 4 The target environment reference images shown correspond to the Faster-RCNN, YOLOv2, YOLOv3, YOLOv4, and YOLOv5 detection models from left to right.

[0046] According to some embodiments, the dimensions of the target environment reference image are 3×224×224, where 3 represents the number of channels and 224 represents the length and width of the image.

[0047] In step S200, a first feature vector corresponding to the target environment is generated based on the target environment reference image.

[0048] For example, in step S200, based on the acquired target environment reference image, the generating device generates a first feature vector corresponding to the target environment through a preset pre-trained model.

[0049] The generation device acquires a pre-trained model.

[0050] According to some embodiments, the pre-trained model may employ a CLIP (Contrastive Language-Image Pre-training) model, which includes a visual encoder and a text encoder. After passing through the visual encoder of the CLIP model, an image is mapped to a fixed-dimensional feature vector, representing the image's semantic representation. The space comprised of feature vectors generated from a large number of images in the CLIP model's visual encoder is called the visual feature space, used to capture the semantic similarity between images.

[0051] The generation device encodes the target environment reference image into the visual feature space through the visual encoder of the pre-trained model to generate the first feature vector corresponding to the target environment.

[0052] According to some embodiments, the first feature vector f r The dimension of f is 1×768, where 1 represents the first feature vector f. r The length of the sequence, 768, represents the first feature vector f. r Embedded dimensions.

[0053] In step S300, an adversarial patch is generated based on the first feature vector using a preset stable diffusion model.

[0054] For example, in step S300, the generating device uses the first feature vector as a condition and performs stepwise denoising of Gaussian noise in the latent space through a preset stable diffusion model to generate an adversarial patch.

[0055] Based on the first feature vector, the generating device obtains the latent space variables through a stable diffusion model.

[0056] According to some embodiments, based on the first feature vector f r The generator uses a stable diffusion model to handle Gaussian noise x in the latent space. T Stepwise denoising is performed to obtain the latent space variable x0. Where x... T The dimensions of x0 are 64×64×4.

[0057] According to some embodiments, denoising can be performed using a preset DDIM (Denoising diffusion implicit models), and its one-step denoising formula is as follows.

[0058]

[0059] in, The cumulative mean coefficient representing the time step from 0 to t is calculated using the following formula: α s It is the mean coefficient of the normal distribution in each denoising step. Its value from time 0 to T is obtained by stepwise linear interpolation from 0.9999 to 0.98. θ (x t ,t,f r σ represents the residual predicted by the stable diffusion model at time t. t is the variance of the normal distribution at time t, which is usually set to 0. z is random noise that follows a normal distribution with a mean of 0 and a variance of 1, and its dimensions are 64×64×4.

[0060] According to some embodiments, x can be t Generate x t-1 The one-step denoising process is defined as x t-1 =Denoise(x t Then from Gaussian noise x T The entire process of gradually denoising down to the latent space variable x0 can be expressed by the following formula.

[0061]

[0062] The generator decodes latent space variables into pixel space using a stable diffusion model to obtain adversarial patches.

[0063] According to some embodiments, the generation device maps the latent space variable x0 back to the pixel space from the latent space using a decoder of a stable diffusion model to obtain an image of the adversarial patch x′. Here, x′ has dimensions of 512×512×3.

[0064] Furthermore, the generation device encodes the adversarial patch into the visual feature space through the visual encoder of the pre-trained model to generate the second feature vector corresponding to the adversarial patch, and obtains the alignment loss function of the first feature vector and the second feature vector.

[0065] According to some embodiments, the generation device inputs the adversarial patch x′ into the visual decoder of the pre-trained model and obtains a second feature vector f corresponding to the adversarial patch x′. x′ The generation device will generate the first feature vector f rSecond eigenvector f x′ Alignment is performed in the visual feature space to obtain the first feature vector f. r Second eigenvector f x′ An alignment loss function is used to increase the similarity between the adversarial patch and the target environment reference image. Here, the first feature vector f... r Second eigenvector f x′ Alignment loss function It can be expressed by the following formula.

[0066]

[0067] In step S400, the adversarial patch is rendered onto a preset 3D model to obtain adversarial samples with different background environments.

[0068] For example, in step S400, the generating device renders the adversarial patch onto a preset 3D model to form a complete 3D model covered with the adversarial patch texture, so as to obtain adversarial samples with different background environments.

[0069] According to some embodiments, the 3D model includes a human body model as well as a top and / or pants model.

[0070] Based on adversarial patches, the generator sets the texture information of the top and / or pants models in a preset 3D model.

[0071] According to some embodiments, 3D model files, including human body models as well as clothing and / or pants models, can be loaded using the PyTorch3D tool. The model files include vertex, face, and texture information.

[0072] According to some embodiments, a textured t-shirt model is set. tex and trouser model texture tex The formula for applying the texture information of the adversarial patch x′ to the top and pants models is as follows.

[0073] tshirt tex =TexturesUV(x′, tshirt) faces tshirt uv (4)

[0074] trouser tex =TexturesUV(x′, trouser) faces trouser uv (5)

[0075] Among them, tshirt faces and trouser facesThese represent the triangular face information of the top and pants models in 3D space, respectively. uv and trouser uv These represent the vertex information of the top and pants models in 3D space, respectively.

[0076] The generation device merges the human body model with the top and / or pants model with pre-set texture information to obtain a complete 3D model.

[0077] According to some embodiments, the generation device merges a human body model, a shirt model, and / or a pants model into a unified 3D model to obtain a complete 3D model. The fusion formula for the complete 3D model is as follows.

[0078] mesh join =join meshes (mesh man mesh tshirt mesh trouser (6)

[0079] Among them, mesh man mesh tshirt and mesh trouser These are the mesh information of the human body model, the top clothing model, and the pants model in 3D space.

[0080] Furthermore, the generating device acquires preset mapping parameters for mapping the 3D model to a 2D plane.

[0081] According to some embodiments, the preset mapping parameters include camera parameters, lighting parameters, and position adjustment parameters.

[0082] According to some embodiments, camera parameters include the distance from the target to the camera (dis), the pitch angle of the target relative to the camera (elev), and the horizontal viewing angle of the target relative to the camera (azim). These parameters collectively define the camera's viewing angle and position for projecting the 3D model onto a 2D plane. Wherein, dis ~ Uniform(d - d + ), elev~Uniform(z - , z + ), azim~Uniform(a - a + ), d - and d + The values ​​can be 2.0m and 2.5m respectively, z - and z + The values ​​can be 0 degrees and 9 degrees respectively, a - and a + The values ​​can be 0 degrees and 360 degrees respectively.

[0083] According to some embodiments, lighting parameters can be set by randomly sampling from three light sources: uniform light, directional light, and point light, to simulate different lighting conditions and enhance the realism of the image.

[0084] According to some embodiments, the position adjustment parameters are used to randomly locate the 3D model in the background environment, and adapt to various external environmental factors by adjusting the position of the model in the background.

[0085] According to preset mapping parameters, the generating device projects the complete 3D model onto a 2D plane to obtain a 2D image corresponding to the complete 3D model.

[0086] According to some embodiments, the generation device can continuously adjust the position and pose of a complete 3D model in 3D space to attack it from different angles, and map the adjusted complete 3D model onto image datasets corresponding to different background environments. The 2D images formed by the projection of the complete 3D model onto a 2D plane at different angles are shown below. Figure 5 As shown.

[0087] Based on image datasets corresponding to different background environments and 2D images corresponding to complete 3D models, the generation device generates adversarial examples corresponding to different background environments.

[0088] According to some embodiments, the generating apparatus acquires a 2D image I formed by the projection of a complete 3D model onto a 2D plane. M And combined with images I from image datasets corresponding to different background environments b Generate adversarial examples I A Adversarial Sample I A The formula for generating is as follows.

[0089]

[0090] According to some embodiments, such as Figure 6 The adversarial examples shown correspond to different background environments, from left to right, to the Faster-RCNN, YOLOv2, YOLOv3, YOLOv4, and YOLOv5 detection models, respectively.

[0091] In step S500, the adversarial patch is optimized based on the adversarial sample using a preset detection model.

[0092] For example, in step S500, the generation device inputs the adversarial sample into a preset detection model to obtain the total loss function of the adversarial patch, and optimizes the adversarial patch accordingly.

[0093] The generation device inputs adversarial examples into the detection model to obtain the target detection confidence corresponding to the adversarial examples, and sets the confidence loss function accordingly.

[0094] According to some embodiments, the generation device will generate adversarial sample I A Input into the detection model and obtain adversarial example I. A Target detection confidence F for all target regions obj (I A The generation device acquires the target detection confidence F. obj (I A The maximum confidence score in the given data is used as the confidence loss function, and this minimum score is minimized. Confidence loss function It can be expressed by the following formula.

[0095]

[0096] Based on the alignment loss function and the confidence loss function, the generator obtains the total loss function corresponding to the adversarial patch.

[0097] According to some embodiments, the generation device will align the loss function. and confidence loss function Perform a weighted combination to obtain the total loss function corresponding to the adversarial patch. Total loss function It can be expressed by the following formula.

[0098]

[0099] Here, α and β are preset parameters, which can take values ​​of 1 and 0.01, respectively.

[0100] The generator optimizes the adversarial patch using the total loss function.

[0101] According to some embodiments, the generating device uses a total loss function. Gradient backpropagation is performed to account for Gaussian noise x in the latent space. T Perform iterative optimization.

[0102] According to some embodiments, the generation device repeatedly executes steps S300 to S500 until the iteration optimization round reaches a preset optimization round or the adversarial patch achieves a preset attack effect. The generation device obtains the iteratively optimized adversarial patch and uses it as the final version of the adversarial patch.

[0103] According to the embodiments of this application, the problem of traditional adversarial patches being too brightly colored can be avoided, the concealment and practicality of adversarial patches in real environments can be improved, and the adversarial patches can adapt to different perspectives and background environments, thereby improving the attack effect and robustness of adversarial patches in real scenarios.

[0104] Figure 7A schematic diagram of an apparatus for generating adversarial patches according to an example embodiment of this application is shown.

[0105] like Figure 7 As shown, the generating device 100 includes a first execution module 110, a second execution module 120, a third execution module 130, a fourth execution module 140, and a fifth execution module 150.

[0106] The first execution module 110 acquires a preset detection model for detecting targets in different background environments.

[0107] The first execution module 110 acquires a preset stable diffusion model for decoding the latent space variables generated by the stable diffusion model and for progressively denoising the Gaussian noise in the latent space.

[0108] The first execution module 110 obtains a preset background environment image dataset, which includes image datasets corresponding to various background environments.

[0109] The first execution module 110 determines the target environment reference image from the image dataset.

[0110] The second execution module 120 acquires the preset pre-trained model.

[0111] The second execution module 120 encodes the target environment reference image into the visual feature space through the visual encoder of the pre-trained model to generate the first feature vector corresponding to the target environment.

[0112] Based on the first feature vector, the third execution module 130 obtains the latent space variables through a stable diffusion model.

[0113] The third execution module 130 decodes the latent space variables into the pixel space using a stable diffusion model to obtain an adversarial patch.

[0114] The third execution module 130 encodes the adversarial patch into the visual feature space through the visual encoder of the pre-trained model to generate the second feature vector corresponding to the adversarial patch, and obtains the alignment loss function of the first feature vector and the second feature vector.

[0115] Based on the adversarial patch, the fourth execution module 140 sets the texture information of the top and / or pants models in the preset 3D model.

[0116] The fourth execution module 140 merges the human body model with the top and / or pants model with pre-set texture information to obtain a complete 3D model.

[0117] The fourth execution module 140 obtains the preset mapping parameters for mapping the 3D model to the 2D plane.

[0118] According to the preset mapping parameters, the fourth execution module 140 projects the complete 3D model onto the 2D plane to obtain the 2D image corresponding to the complete 3D model.

[0119] Based on the image datasets corresponding to different background environments and the 2D images corresponding to the complete 3D model, the fourth execution module 140 generates adversarial examples corresponding to different background environments.

[0120] The fifth execution module 150 inputs the adversarial sample into the detection model to obtain the target detection confidence corresponding to the adversarial sample, and sets the confidence loss function accordingly.

[0121] Based on the alignment loss function and the confidence loss function, the fifth execution module 150 obtains the total loss function corresponding to the adversarial patch.

[0122] The fifth execution module 150 optimizes the adversarial patch through the total loss function.

[0123] Figure 8 A block diagram of an electronic device according to an example embodiment of this application is shown.

[0124] like Figure 8 As shown, the electronic device 600 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0125] like Figure 8 As shown, the electronic device 600 is manifested in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc. The storage unit stores program code, which can be executed by the processing unit 610, causing the processing unit 610 to perform the methods described in this specification according to the various exemplary embodiments of this application. For example, the processing unit 610 can perform, for example... Figure 1 The method shown.

[0126] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.

[0127] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0128] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0129] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0130] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. The technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0131] Software products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0133] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0134] The aforementioned computer-readable medium carries one or more programs, which, when executed by a device, cause the computer-readable medium to perform the aforementioned functions.

[0135] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified accordingly and placed in one or more devices that are unique to this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0136] The embodiments of this application have been described in detail above. These descriptions are solely for the purpose of helping to understand the method and core ideas of this application. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of this application, its specific implementation methods, and its application scope, are all within the scope of protection of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating anti-patch, characterized in that, include: Obtain a reference image of the target environment; Based on the target environment reference image, a first feature vector corresponding to the target environment is generated; Based on the first feature vector, an adversarial patch is generated using a preset stable diffusion model; The adversarial patch is rendered onto a preset 3D model to obtain adversarial samples in different background environments; Based on the adversarial sample, the adversarial patch is optimized using a preset detection model.

2. The method according to claim 1, characterized in that, Obtain a reference image of the target environment, including: Obtain the image datasets corresponding to the different background environments; The target environment reference image is determined from the image dataset.

3. The method according to claim 1, characterized in that, Based on the target environment reference image, a first feature vector corresponding to the target environment is generated, including: Obtain the preset pre-trained model; The target environment reference image is encoded into the visual feature space by the visual encoder of the pre-trained model to generate the first feature vector.

4. The method according to claim 1, characterized in that, Based on the first feature vector, an adversarial patch is generated using a preset stable diffusion model, including: Based on the first feature vector, latent space variables are obtained through the stable diffusion model; The latent space variables are decoded into the pixel space using the stable diffusion model to obtain the adversarial patch.

5. The method according to claim 4, characterized in that, Based on the first feature vector, an adversarial patch is generated using a preset stable diffusion model, which further includes: Obtain the preset pre-trained model; The adversarial patch is encoded into the visual feature space by the visual encoder of the pre-trained model to generate a second feature vector; Obtain the alignment loss function for the first feature vector and the second feature vector.

6. The method according to claim 1, characterized in that, The 3D model includes a human body model as well as a top and / or pants model; The adversarial patch is rendered onto a preset 3D model to obtain adversarial samples in different background environments, including: Based on the adversarial patch, set the texture information of the top model and / or the pants model; The human body model and the top and / or pants model with textured information are merged to obtain a complete 3D model; The adversarial sample is obtained based on the complete 3D model.

7. The method according to claim 6, characterized in that, The adversarial example is obtained based on the complete 3D model, including: Obtain the preset mapping parameters for mapping a 3D model to a 2D plane; According to the preset mapping parameters, the complete 3D model is projected onto a 2D plane to obtain a 2D image corresponding to the complete 3D model; The adversarial examples are generated based on the image datasets corresponding to the different background environments and the 2D images corresponding to the complete 3D model.

8. The method according to claim 5, characterized in that, Based on the adversarial sample, the adversarial patch is optimized using a preset detection model, including: The adversarial sample is input into the detection model to obtain the target detection confidence level corresponding to the adversarial sample; Based on the target detection confidence corresponding to the adversarial example, a confidence loss function is set; Based on the alignment loss function and the confidence loss function, obtain the total loss function corresponding to the adversarial patch; The adversarial patch is optimized using the total loss function.

9. An apparatus for generating anti-patterns, characterized in that, include: The first execution module is used to acquire a reference image of the target environment; The second execution module is used to generate a first feature vector corresponding to the target environment based on the target environment reference image; The third execution module is used to generate an adversarial patch based on the first feature vector using a preset stable diffusion model. The fourth execution module is used to render the adversarial patch onto a preset 3D model to obtain adversarial samples with different background environments; The fifth execution module is used to optimize the adversarial patch based on the adversarial sample using a preset detection model.

10. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the method as described in any one of claims 1-8.