Diffusion model-based adversarial sample generation method and device for HR-Net detector

By using a diffusion-based adversarial example generation method, and optimizing the perturbation with a variational autoencoder and a multi-scale feature-aware loss function, the stealth and transferability issues of the HR-Net detector are addressed, thus improving its robustness.

CN120877016APending Publication Date: 2025-10-31CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717015.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing adversarial example generation methods have poor concealment and low transferability against the HR-Net detector, making it difficult to effectively improve its robustness.

Method used

We employ a diffusion-based adversarial example generation method. By encoding images into a latent space through a variational autoencoder, we optimize perturbations using PGD gradient attack, self-attention, and cross-attention, and introduce a multi-scale feature-aware perturbation loss function to generate high-quality adversarial examples.

Benefits of technology

The generated adversarial examples are highly concealed and transferable, significantly improving the robustness of the HR-Net detector and effectively resisting adversarial attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877016A_ABST
    Figure CN120877016A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a diffusion model-based adversarial sample generation method and device for an HR-Net detector. The method uses a variational auto-encoder to encode an image into a potential space to generate a potential representation; and generating disturbance in a potential space: in each iteration, generating disturbance by a PGD gradient attack method, and after the disturbance is optimized through self-attention and cross attention, balancing and fusing the optimized disturbance by using self-adaptive weight. Besides, a multi-scale feature perception disturbance loss function is introduced into each iterative calculation, so that countermeasure disturbance pertinently interfering the HR-Net multi-resolution feature extraction and fusion process is generated, and the disturbance effect is enhanced. And finally, decoding the potential representation added with the current disturbance back to an image space by using a variational self-decoder to obtain a high-quality adversarial sample with good concealment and high transferability, and performing robustness training by using the adversarial sample to significantly improve the security of the HR-Net detector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method and apparatus for generating adversarial examples based on a diffusion model for the HR-Net detector. Background Technology

[0002] With the rapid development of artificial intelligence technology, deep learning has made significant progress in the field of image generation and processing. In particular, the rise of deepfake technology in recent years has made it increasingly easy to create highly realistic forged images and videos. These technologies can seamlessly replace one person's facial features onto another's face, or generate completely non-existent facial images, with a realism sufficient to deceive human observers. While these technologies have positive applications in entertainment, artistic creation, and other fields, they have also brought about serious social problems, such as identity fraud, the spread of misinformation, and privacy violations.

[0003] To address the challenges posed by deepfake technology, researchers have developed various deepfake detection techniques. Among them, detectors based on High-Resolution Networks (HR-Net) have performed exceptionally well in deepfake detection tasks due to their ability to maintain high-resolution representations and capture multi-scale features. HR-Net effectively captures subtle forgery traces in images by connecting multiple sub-networks of different resolutions in parallel and repeatedly performing multi-resolution fusion, thus gaining widespread application in the field of deepfake detection. However, like other deep learning models, the HR-Net detector also faces the threat of adversarial attacks. Adversarial attacks refer to techniques that add minute perturbations, imperceptible to humans, to the input image, causing the model to make incorrect judgments. For deepfake detectors, adversarial attacks can cause forged images to be mistakenly identified as genuine images, thereby bypassing detection. This type of attack poses a serious threat to the security of deepfake detection systems, especially in critical application scenarios such as electronic forensics and news media review.

[0004] While adversarial attacks pose a threat to deepfake detectors, researching adversarial example generation methods also has a positive side. By generating high-quality adversarial examples and incorporating them into training data, the robustness of detectors can be improved, enabling them to resist various adversarial attacks. This method, known as adversarial training or robust training, is an effective means of enhancing the security of deep learning models.

[0005] However, existing adversarial example generation methods have some limitations. Traditional gradient-based methods (such as FGSM and PGD) can effectively generate adversarial examples, but these examples often contain obvious noise and artifacts, resulting in poor stealth. Optimization-based methods (such as C&W) can generate more stealthy adversarial examples, but they are computationally expensive and have low transferability (the success rate of attacks on other models). Generative model-based methods (such as AdvGAN) have improved in stealth and transferability, but they still struggle to generate adversarial examples that are both highly stealthy and have good transferability. Summary of the Invention

[0006] This application provides a diffusion-based adversarial example generation method for the HR-Net detector, which addresses the problem that while generating high-quality adversarial examples can improve the robustness of the HR-Net detector, existing adversarial example generation methods suffer from poor concealment and low transferability.

[0007] Accordingly, this application also provides an adversarial example generation device based on a diffusion model for the HR-Net detector, which ensures the implementation and application of the above method.

[0008] To address the aforementioned technical problems, this application discloses a diffusion-based adversarial example generation method for the HR-Net detector, the method comprising:

[0009] The image is encoded into the latent space using a variational autoencoder to generate a latent representation.

[0010] In the latent space, for each iteration, a perturbation is generated by the PGD gradient attack method, and after the perturbation is optimized by self-attention and cross-attention, adaptive weights are used to balance and fuse the perturbation optimized by self-attention and cross-attention to generate the current perturbation.

[0011] The variational autodecoder is used to decode the latent representation with the current perturbation back into the image space to obtain the current adversarial example.

[0012] In each iteration, the output of the HR-Net detector for the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a pre-defined multi-scale feature-aware perturbation loss function.

[0013] This application also discloses a diffusion-based adversarial example generation device for the HR-Net detector, the device comprising:

[0014] The encoding module is used to encode the image into the latent space using a variational autoencoder to generate a latent representation.

[0015] The perturbation generation module is used to generate perturbations in the latent space for each iteration by the PGD gradient attack method, and after optimizing the perturbations by self-attention and cross-attention, it uses adaptive weights to balance and fuse the perturbations optimized by self-attention and cross-attention to generate the current perturbation.

[0016] The decoding module is used to decode the latent representation with the added perturbation back into the image space using a variational autodecoder to obtain the current adversarial example;

[0017] In each iteration, the output of the HR-Net detector for the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a pre-defined multi-scale feature-aware perturbation loss function.

[0018] In this application, for the multi-resolution parallel structure of the HR-Net detector, a variational autoencoder is used to encode images into a latent space to generate a latent representation. Perturbations are generated in the latent space to improve the stealth and transferability of adversarial examples. In each iteration of the diffusion model, perturbations are generated by the PGD gradient attack method. Self-attention and cross-attention are used to optimize the perturbations, capturing the dependencies within and between perturbations. Adaptive weights are then used to balance and fuse the optimized perturbations, adapting to the needs of different images and different attack stages, adaptively generating the optimal current adversarial example. Furthermore, a multi-scale feature-aware perturbation loss function is introduced in each iteration to generate adversarial perturbations that specifically interfere with the multi-resolution feature extraction and fusion process of HR-Net, thereby enhancing the perturbation effect. Finally, a variational autodecoder is used to decode the latent representation with added perturbations back into the image space, resulting in high-quality adversarial examples with good stealth and high transferability. Using the adversarial examples generated in this application for robust training can significantly improve the security of the HR-Net detector.

[0019] Additional aspects and advantages of this application will be set forth in the following description, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description

[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0021] Figure 1 A flowchart illustrating the diffusion-based adversarial example generation method for the HR-Net detector provided in this application embodiment;

[0022] Figure 2 This is a schematic diagram of the structure of the adversarial example generation device based on the diffusion model for the HR-Net detector provided in an embodiment of this application. Detailed Implementation

[0023] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0024] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0025] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0026] The solutions provided in this application can be executed by any electronic device, such as a terminal device or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. Regarding the technical problems existing in the prior art, the adversarial example generation method and apparatus based on a diffusion model for the HR-Net detector provided in this application aim to solve at least one of the technical problems in the prior art.

[0027] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0028] In recent years, with the successful application of High-Resolution Networks (HR-Net) in computer vision, researchers have also begun to apply HR-Net to deep forgery detection tasks. Because HR-Net can maintain high-resolution representations and capture multi-scale features, it performs exceptionally well in deep forgery detection, becoming one of the most advanced detection methods currently available.

[0029] HR-Net is a neural network architecture specifically designed to preserve high-resolution representations. Unlike traditional convolutional neural networks, HR-Net effectively captures detailed information in images by connecting multiple subnetworks of different resolutions in parallel and repeatedly performing multi-resolution fusion.

[0030] The core idea of ​​HR-Net is to maintain high-resolution representation throughout the entire network. It starts with a high-resolution sub-network and then gradually adds low-resolution sub-networks to form a multi-resolution parallel sub-network. At each stage of the network, HR-Net exchanges feature maps of different resolutions, allowing high-resolution features to acquire global information from low-resolution features, while low-resolution features can also acquire detailed information from high-resolution features.

[0031] Specifically, HR-Net consists of four stages:

[0032] The first stage consists of a single high-resolution sub-network that processes the input image through a series of convolutional layers and residual blocks.

[0033] Second stage: Based on the first stage, add a low-resolution sub-network and perform multi-resolution fusion.

[0034] The third stage involves adding a lower-resolution sub-network based on the second stage and performing multi-resolution fusion.

[0035] The fourth stage involves adding a sub-network with the lowest resolution based on the third stage and performing multi-resolution fusion.

[0036] At each stage, HR-Net swaps and fuses feature maps of different resolutions. Specifically, for swapping from high resolution to low resolution, HR-Net uses convolutional layers with a stride of 2 for downsampling; for swapping from low resolution to high resolution, HR-Net uses bilinear interpolation for upsampling. This multi-resolution fusion mechanism allows HR-Net to simultaneously capture both local details and global semantic information of the image.

[0037] Because HR-Net can maintain high-resolution representations and capture multi-scale features, it performs exceptionally well in deepfake detection tasks. In particular, HR-Net can effectively capture subtle artifacts introduced during deepfake processing, such as blurred facial edges and texture inconsistencies, thus achieving high-precision deepfake detection.

[0038] Adversarial examples are samples that introduce subtle, imperceptible perturbations into the input data, causing deep learning models to make incorrect judgments. Generating high-quality adversarial examples and incorporating them into the training data can improve the robustness of the HR-Net detector, enabling it to resist various adversarial attacks. Following the successful application of diffusion models in image generation, researchers have begun exploring their application to adversarial example generation. For example, by introducing adversarial targets during the diffusion process, samples that are both natural and adversarial can be generated. However, these methods typically operate in the original image space, resulting in low computational efficiency and difficulty in controlling the size and distribution of perturbations.

[0039] Diffusion models are a type of generative model based on Markov chains that generate samples by progressively adding noise to the data (forward process) and progressively removing noise (backward process). The core idea of ​​diffusion models is that by learning how to reverse the noise addition process, high-quality samples can be generated from pure noise.

[0040] The forward process of the diffusion model is a fixed Markov chain that progressively adds Gaussian noise to the data until the data becomes pure noise. Specifically, given a data sample x0 (the original data sample), the forward process is defined as:

[0041]

[0042] Where q(x) t |x t-1 The transition probability from time step t-1 to time step t is defined as follows: β t These are predefined noise scheduling parameters. Let represent a normal distribution, and Ι be the identity matrix.

[0043] The reverse process of the diffusion model is a learned Markov chain that progressively recovers data from pure noise. Specifically, given a pure noise sample x... T (Final noise state), the reverse process is defined as:

[0044]

[0045] Where, p θ (x t-1 |x t The transition probability from time step t to time step t-1 is defined as follows: μ θ and Σ θ These are the mean and covariance functions parameterized by the neural network, where θ represents the parameters of the neural network.

[0046] The training objective of the diffusion model is to maximize the log-likelihood of the data, which is equivalent to minimizing the reconstruction error of a series of denoising autoencoders. Specifically, the training objective can be expressed as:

[0047]

[0048] Where ∈ is standard Gaussian noise, ∈ θ It is noise in the neural network prediction, x t This is data with t-step noise added. It expresses expectation.

[0049] While diffusion models can generate high-quality samples, they typically require multiple iterations to generate a single sample, resulting in low computational efficiency. Latent Diffusion Models (LDMs), on the other hand, perform the diffusion process in a compressed latent space, improving computational efficiency while maintaining high-quality generated images.

[0050] The latent diffusion model first uses an autoencoder to compress the image into a latent space, then applies the diffusion model in the latent space, and finally uses a decoder to transform the generated latent representation back into the image space. This approach not only improves computational efficiency but also enables the diffusion model to better capture the semantic information of the image, generating more natural and diverse samples.

[0051] Based on the above, this application provides a diffusion-based adversarial example generation method for the HR-Net detector. This method can be executed by any electronic device, and optionally, it can be executed on a server or a terminal device.

[0052] like Figure 1 As shown, the method may include the following steps:

[0053] Step 101: Encode the image into the latent space using a variational autoencoder (VAE encoder) to generate a latent representation;

[0054] Step 102: In the latent space, for each iteration, a perturbation is generated by the PGD gradient attack method, and after the perturbation is optimized by self-attention and cross-attention, adaptive weights are used to balance and fuse the perturbation optimized by self-attention and cross-attention to generate the current perturbation.

[0055] Step 103: Use the variational autodecoder (VAE decoder) to decode the latent representation with the added perturbation back into the image space to obtain the current adversarial example;

[0056] In each iteration, the output of the HR-Net detector for the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a pre-defined multi-scale feature-aware perturbation loss function.

[0057] The method proposed in this application is specifically designed for the parallel multi-resolution architecture of the HR-Net detector, aiming to generate adversarial examples that are both highly concealed and have good transferability, thereby improving the robustness of the HR-Net detector against adversarial attacks. The core idea of ​​this method is to combine the traditional PGD gradient attack method with a latent diffusion model, operate in the latent space of the diffusion model, and utilize self-attention and cross-attention mechanisms to dynamically guide the generation and fusion of perturbations.

[0058] Specifically, given a forged image x (input image) and an HR-Net detector f (target model), the goal is to generate an adversarial sample x′ (adversarial sample) such that f(x′) is misclassified as a real image (i.e., f(x′) = 0), while maintaining the visual similarity between x′ and x. To achieve this goal, in this embodiment, a pre-trained VAE encoder is first used to encode the image x into a latent space to obtain a latent representation z; then, an adversarial perturbation δ is generated in the latent space; finally, a VAE decoder is used to decode the perturbated latent representation z+δ back into the image space to obtain the adversarial sample x′.

[0059] Compared to generating adversarial perturbations directly in the image space, the latent space typically has a lower dimension, making the generation of adversarial perturbations more efficient. After being transformed nonlinearly by the VAE decoder, the perturbations in the latent space produce more natural and semantic changes in the image space, thereby improving the stealth of adversarial examples. The perturbations in the latent space can leverage the generative capabilities of the diffusion model to generate more diverse and effective adversarial examples, thereby improving transferability.

[0060] Regarding perturbation generation in the latent space, this application proposes a multi-scale feature-aware perturbation generation method tailored to the multi-resolution parallel structure of the HR-Net detector. Unlike traditional adversarial attack methods that typically consider only single-scale features and are therefore insufficient to effectively attack multi-scale feature fusion networks like HR-Net, the method in this application introduces a multi-scale feature-aware perturbation loss function. This allows for the simultaneous consideration of feature information at different scales in the latent space, generating adversarial perturbations that specifically interfere with the multi-resolution feature extraction and fusion process of HR-Net, thereby enhancing the perturbation's effectiveness. Furthermore, this application incorporates self-attention and cross-attention mechanisms to dynamically guide the generation and fusion of perturbations, resulting in adversarial examples with better stealth and transferability.

[0061] In this embodiment, for the multi-resolution parallel structure of the HR-Net detector, a variational autoencoder is used to encode the image into the latent space to generate a latent representation. Perturbations are generated in the latent space to improve the stealth and transferability of adversarial examples. In each iteration of the diffusion model, perturbations are generated by the PGD gradient attack method. Self-attention and cross-attention are used to optimize the perturbations, capturing the dependencies within and between perturbations. Adaptive weights are then used to balance and fuse the optimized perturbations, adapting to the needs of different images and different attack stages, and adaptively generating the optimal current adversarial example. Furthermore, a multi-scale feature-aware perturbation loss function is introduced in each iteration to generate adversarial perturbations that specifically interfere with the HR-Net multi-resolution feature extraction and fusion process, thereby enhancing the perturbation effect. Finally, a variational autodecoder is used to decode the latent representation with the added perturbation back into the image space, resulting in high-quality adversarial examples with good stealth and high transferability. Using the adversarial examples generated in this embodiment for robust training can significantly improve the security of the HR-Net detector.

[0062] In an optional embodiment, the image x is encoded into the latent space using a pre-trained VAE encoder to obtain a latent representation z, the expression of which is as follows:

[0063] z = Encoder(x) (4)

[0064] Among them, Encoder (VAE encoder) is a pre-trained VAE encoder.

[0065] In an optional embodiment, in the latent space, for each iteration, a perturbation is generated by the PGD gradient attack method, and after optimizing the perturbation through self-attention and cross-attention, adaptive weights are used to balance and fuse the perturbation optimized by self-attention and cross-attention to generate the current perturbation, including:

[0066] Initialize the perturbations in the potential space;

[0067] For each iteration, the output of the HR-Net detector to the adversarial examples generated in the previous iteration is calculated, and the multi-scale feature-aware perturbation loss function is calculated to obtain the multi-scale feature-aware loss.

[0068] Under the influence of multi-scale feature perception loss, the PGD gradient attack method is used to generate PGD perturbations that fuse feature information at different scales.

[0069] By optimizing the PGD perturbation through self-attention and cross-attention, self-attention optimized perturbation and cross-attention optimized perturbation are generated respectively;

[0070] Adaptive weights are used to balance and fuse self-attention optimization perturbations and cross-attention optimization perturbations to generate the current adversarial example.

[0071] The key to generating adversarial perturbations in the latent space is how to effectively utilize gradient information and the generative capabilities of the diffusion model. The specific implementation method for generating the current perturbation in this application embodiment is as follows:

[0072] Step 1. Initialize the perturbation δ in the latent space as a zero vector:

[0073] δ0=0 (5)

[0074] Where δ0 is the initial perturbation.

[0075] Step 2. For each iteration t = 1, 2, ..., T, perform the following steps:

[0076] Step 2.1. Use the VAE decoder to decode the current latent representation back to the image space:

[0077] x t =Decoder(z+δ) t-1 (6)

[0078] Where Decoder is a pre-trained VAE decoder, x t This is a current adversarial example.

[0079] Step 2.2. Calculate the HR-Net detector for x t The output is calculated, and the multi-scale feature-aware perturbation loss function is obtained to obtain the multi-scale feature-aware loss L. multi-scale .

[0080] Step 2.3. Calculate the gradient of the multi-scale feature-aware perturbation loss function with respect to the latent representation:

[0081]

[0082] Among them, g t(Gradient) is the gradient of the loss function with respect to the latent representation. This represents the gradient operator.

[0083] Step 2.4. Update the perturbation using the PGD gradient attack method:

[0084]

[0085] in, It is the PGD perturbation updated using the PGD gradient attack method, Clip [-ε,ε] This means that the perturbation is clipped to the range [-ε, ε], α t is the step size of the t-th iteration, sign indicates the sign of the gradient, and ε is the upper bound of the perturbation.

[0086] Step 2.5. Apply self-attention and cross-attention to optimize the PGD perturbation, and generate self-attention optimization perturbations respectively. and cross-attention optimization perturbation

[0087] Step 2.6. Use adaptive weights for dynamic fusion.

[0088] In an optional embodiment, the multi-scale feature-aware perturbation loss function is:

[0089]

[0090] Among them, L multi-scale It is a multi-scale feature-aware loss, where S is the number of scales of the HR-Net detector, and w i It is the weight of the i-th scale, L i It is the loss function at the i-th scale, f i It is the feature extractor of the HR-Net detector at the i-th scale, y target It is the target label.

[0091] By employing the aforementioned multi-scale feature-aware perturbation loss function, this embodiment of the application can simultaneously consider the feature information of HR-Net at different resolutions, generating adversarial perturbations that can effectively interfere with the multi-scale feature fusion process of HR-Net. Compared with traditional methods that only consider single-scale features, the method in this embodiment of the application, which enhances the perturbation effect by introducing a multi-scale feature-aware perturbation loss function, can generate more effective adversarial examples and achieve a higher success rate in attacking the HR-Net detector.

[0092] In an optional embodiment, the PGD perturbation is optimized through self-attention and cross-attention, generating self-attention optimized perturbations and cross-attention optimized perturbations respectively, including:

[0093] Generate self-attention optimized perturbations by optimizing PGD perturbations.

[0094]

[0095] Among them, SelfAttention is a self-attention mechanism;

[0096] Generate diffusion model perturbation δ in the latent space using a diffusion model. t Diff :

[0097]

[0098] Among them, DiffusionModel is a pre-trained diffusion model;

[0099] The diffusion model is perturbed using cross-attention. and the self-attention optimization perturbation Fusion, generating cross-attention optimization perturbation

[0100]

[0101] CrossAttention is a cross-attention mechanism.

[0102] In this embodiment, self-attention and cross-attention mechanisms are introduced to generate and fuse perturbations more effectively. The self-attention mechanism is used to capture the dependencies within the PDG perturbation, while the cross-attention mechanism is used to capture the dependencies between the PGD perturbation and the diffusion model perturbation.

[0103] In an optional embodiment, self-attention optimization perturbation is used to generate self-attention optimization perturbation. include:

[0104] Reshape the PGD perturbation into a PDG perturbation sequence:

[0105]

[0106] Where, δ seq This is the reshaped PDG perturbation sequence, where Reshape is the reshaping function, N is the sequence length (equal to the spatial dimension of the latent representation), and C is the number of channels. Let represent the N×C dimensional real space.

[0107] Using the PDG perturbation sequence as the query, key, and value, the self-attention weights are calculated as follows:

[0108] Calculate the query (Q), key (K), and value (V):

[0109] Q = δ seq (14)

[0110] K = δ seq (15)

[0111] V = δ seq (16)

[0112] Calculate the self-attention weights:

[0113]

[0114] Where A is the self-attention weight, Softmax is the Softmax function, and K... T It is the transpose of K. It is the scaling factor.

[0115] Apply self-attention weights to obtain self-attention output:

[0116] δ self =AV (18)

[0117] Where, δ self It is the self-attention output after applying attention weights.

[0118] Reshape the self-attention output back to its original shape to obtain the self-attention optimized perturbation:

[0119]

[0120] in, It is a self-attention optimization perturbation obtained after reshaping back to the original shape.

[0121] In this embodiment of the application, a self-attention mechanism is used to process PGD perturbations. Through the self-attention mechanism, each element in the PGD perturbation will consider the information of other elements, thereby generating a more structured and semantic perturbation.

[0122] In an optional embodiment, cross-attention is used to fuse the diffusion model perturbation and the self-attention optimization perturbation to generate a cross-attention optimization perturbation, including:

[0123] The diffusion model perturbation and the self-attention optimization perturbation are reshaped into the corresponding diffusion model perturbation sequence and self-attention optimization perturbation sequence:

[0124]

[0125] Where, δ self_seq and δ diff_seq These are the reshaped self-attention optimization perturbation sequence and the diffusion model perturbation sequence, respectively.

[0126] Using the self-attention optimization perturbation sequence as the query and the diffusion model perturbation sequence as the key and value, the cross-attention weights are calculated as follows:

[0127] Calculate the query (Q), key (K), and value (V):

[0128] Q = δ self_seq (twenty two)

[0129] K = δ diff_seq (twenty three)

[0130] V = δ diff_seq (twenty four)

[0131] Calculate the cross-attention weights:

[0132]

[0133] Apply cross-attention weights to obtain the cross-attention output:

[0134] δ cross =AV (26)

[0135] Where, δ cross It is the cross-attention output after applying cross-attention weights;

[0136] Reshape the cross-attention output back to its original shape to obtain the cross-attention optimization perturbation:

[0137]

[0138] in, It is the cross-attention optimization perturbation obtained after reshaping back to the original shape.

[0139] In this embodiment of the application, the PGD perturbation is optimized by self-attention to obtain a self-attention optimized perturbation. The cross-attention mechanism is used to fuse the self-attention optimized perturbation and the diffusion model perturbation. Through the cross-attention mechanism, each element in the PGD perturbation will consider the information of the corresponding element in the diffusion model perturbation, so that the final perturbation can simultaneously possess the effectiveness of the PGD perturbation and the naturalness of the diffusion model perturbation.

[0140] In an optional embodiment, adaptive weights are used to balance and fuse the self-attention optimization perturbation and the cross-attention optimization perturbation to generate the current adversarial example, including:

[0141] The fusion weights are dynamically adjusted based on the current attack success rate and perturbation characteristics.

[0142] The current adversarial example is generated by fusing self-attention optimization perturbations and cross-attention optimization perturbations using fusion weights.

[0143] In this embodiment, to further improve the quality of adversarial examples, a dynamic fusion and adaptive weight adjustment mechanism is proposed. Traditional adversarial example generation methods typically use fixed weights to fuse different types of perturbations, which is difficult to adapt to the needs of different images and different attack stages. However, the method in this embodiment dynamically adjusts the weights of self-attention optimization perturbation and cross-attention optimization, adaptively generating optimal adversarial examples based on the success rate of the current attack and the characteristics of the perturbations.

[0144] Specifically, the dynamic fusion function in this application embodiment is defined as:

[0145]

[0146] Where, λ t This is the fusion weight for the t-th iteration, which is dynamically adjusted based on the success rate of the current attack and the characteristics of the perturbation.

[0147] λ t =σ(α·ASR) t +β·SSIM t +γ·t / T) (29)

[0148] Where σ is the Sigmoid function, ASR t This is the current attack success rate, SSIM. t t / T is the structural similarity between the current adversarial sample and the original image, t / T is the number of normalization iterations, and α, β and γ are weight coefficients.

[0149] In this way, the method in this embodiment can dynamically adjust the weights of self-attention optimization perturbation and cross-attention optimization based on the progress of the current attack and the quality of the adversarial example, generating adversarial examples that are both effective and stealthy. Compared with traditional methods using fixed weights, the dynamic fusion and adaptive weight adjustment mechanism in this embodiment can generate higher-quality adversarial examples, significantly improving attack success rate, stealth, and transferability.

[0150] Based on the above method, the complete algorithm flow is given, as shown in Algorithm 1:

[0151]

[0152]

[0153] Through the above algorithm, the embodiments of this application can generate adversarial examples that are both highly concealed and have good transferability, effectively attacking the HR-Net detector, while providing high-quality adversarial data for the robust training of the HR-Net detector.

[0154] Based on the same principles as the methods provided in the embodiments of this application, the embodiments of this application also provide an adversarial example generation device based on a diffusion model for the HR-Net detector, such as... Figure 2 As shown, the device includes:

[0155] Encoding module 201 is used to encode the image into the latent space using a variational autoencoder to generate a latent representation;

[0156] The perturbation generation module 202 is used to generate perturbations in the latent space for each iteration by the PGD gradient attack method, and after optimizing the perturbations by self-attention and cross-attention, use adaptive weights to balance and fuse the perturbations optimized by self-attention and cross-attention to generate the current perturbation.

[0157] Decoding module 203 is used to decode the latent representation with the added current perturbation back into the image space using a variational autodecoder to obtain the current adversarial example;

[0158] In each iteration, the output of the HR-Net detector for the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a pre-defined multi-scale feature-aware perturbation loss function.

[0159] In this embodiment, for the multi-resolution parallel structure of the HR-Net detector, a variational autoencoder is used to encode the image into the latent space to generate a latent representation. Perturbations are generated in the latent space to improve the stealth and transferability of adversarial examples. In each iteration of the diffusion model, perturbations are generated by the PGD gradient attack method. Self-attention and cross-attention are used to optimize the perturbations, capturing the dependencies within and between perturbations. Adaptive weights are then used to balance and fuse the optimized perturbations, adapting to the needs of different images and different attack stages, and adaptively generating the optimal current adversarial example. Furthermore, a multi-scale feature-aware perturbation loss function is introduced in each iteration to generate adversarial perturbations that specifically interfere with the HR-Net multi-resolution feature extraction and fusion process, thereby enhancing the perturbation effect. Finally, a variational autodecoder is used to decode the latent representation with the added perturbation back into the image space, resulting in high-quality adversarial examples with good stealth and high transferability. Using the adversarial examples generated in this embodiment for robust training can significantly improve the security of the HR-Net detector.

[0160] The adversarial example generation device based on the diffusion model for the HR-Net detector provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.

[0161] The diffusion-based adversarial example generation device for the HR-Net detector in this application embodiment can execute the diffusion-based adversarial example generation method for the HR-Net detector provided in this application embodiment. The implementation principles are similar. The actions performed by each module and unit in the diffusion-based adversarial example generation device for the HR-Net detector in each embodiment of this application correspond to the steps in the diffusion-based adversarial example generation method for the HR-Net detector in each embodiment of this application. For detailed functional descriptions of each module of the diffusion-based adversarial example generation device for the HR-Net detector, please refer to the descriptions in the corresponding diffusion-based adversarial example generation method for the HR-Net detector shown above, which will not be repeated here.

[0162] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A diffusion-based adversarial example generation method for the HR-Net detector, characterized in that, The method includes: The image is encoded into the latent space using a variational autoencoder to generate a latent representation. In the latent space, for each iteration, a perturbation is generated by the PGD gradient attack method, and after the perturbation is optimized by self-attention and cross-attention, adaptive weights are used to balance and fuse the perturbation optimized by self-attention and cross-attention to generate the current perturbation. The variational autodecoder is used to decode the latent representation with the current perturbation back into the image space to obtain the current adversarial example. In each iteration, the output of the HR-Net detector to the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a preset multi-scale feature-aware perturbation loss function.

2. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 1, characterized in that, In the latent space, for each iteration, a perturbation is generated by the PGD gradient attack method, and after optimizing the perturbation through self-attention and cross-attention, adaptive weights are used to balance and fuse the perturbations optimized by self-attention and cross-attention to generate the current perturbation, including: Initialize the perturbations in the potential space; For each iteration, the output of the HR-Net detector to the adversarial examples generated in the previous iteration is calculated, and the multi-scale feature-aware perturbation loss function is calculated to obtain the multi-scale feature-aware loss. Under the influence of multi-scale feature perception loss, the PGD gradient attack method is used to generate PGD perturbations that fuse feature information at different scales. The PGD perturbation is optimized by self-attention and cross-attention, respectively generating self-attention optimized perturbation and cross-attention optimized perturbation; Adaptive weights are used to balance and fuse the self-attention optimization perturbation and the cross-attention optimization perturbation to generate the current adversarial example.

3. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 2, characterized in that, The multi-scale feature-sensing perturbation loss function is: Among them, L multi-scale It is a multi-scale feature-aware loss, where S is the number of scales of the HR-Net detector, and w i It is the weight of the i-th scale, L i It is the loss function at the i-th scale, f i It is the feature extractor of the HR-Net detector at the i-th scale, y target It is the target label.

4. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 2, characterized in that, The step of optimizing the PGD perturbation through self-attention and cross-attention, generating self-attention optimized perturbations and cross-attention optimized perturbations respectively, includes: The PGD perturbation is optimized using self-attention to generate a self-attention optimized perturbation; Diffusion model perturbations are generated in the latent space using a diffusion model; The cross-attention optimization perturbation is generated by fusing the diffusion model perturbation and the self-attention optimization perturbation using cross-attention.

5. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 4, characterized in that, The step of using self-attention to optimize the PGD perturbation and generating a self-attention optimized perturbation includes: The PGD perturbation is reshaped into a PDG perturbation sequence; The PDG perturbation sequence is used as the query, key, and value to calculate the self-attention weights; By applying the self-attention weights, a self-attention output is obtained; The self-attention output is reshaped back to its original shape to obtain the self-attention optimized perturbation.

6. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 4, characterized in that, The step of fusing the diffusion model perturbation and the self-attention optimization perturbation using cross-attention to generate the cross-attention optimization perturbation includes: The diffusion model perturbation and the self-attention optimization perturbation are reshaped into corresponding diffusion model perturbation sequences and self-attention optimization perturbation sequences; The self-attention optimization perturbation sequence is used as the query, and the diffusion model perturbation sequence is used as the key and value to calculate the cross-attention weight; Apply the aforementioned cross-attention weights to obtain the cross-attention output; The cross-attention output is reshaped back to its original shape to obtain the cross-attention optimized perturbation.

7. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 2, characterized in that, The step of using adaptive weights to balance and fuse the self-attention optimization perturbation and the cross-attention optimization perturbation to generate the current adversarial example includes: The fusion weights are dynamically adjusted based on the current attack success rate and perturbation characteristics. The self-attention optimization perturbation and the cross-attention optimization perturbation are fused using the fusion weights to generate the current adversarial example.

8. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 7, characterized in that, The calculation formula for dynamically adjusting the fusion weight based on the current attack success rate and perturbation characteristics is as follows: l t =σ(α·ASR t +β·SSIM t +γ·t / T) Where σ is the Sigmoid function, ASR t This is the current attack success rate, SSIM. t t / T is the structural similarity between the current adversarial sample and the original image, t / T is the number of normalization iterations, and α, β and γ are weight coefficients.

9. The adversarial example generation method based on a diffusion model for the HR-Net detector according to claim 8, characterized in that, The formula for calculating the current adversarial example by fusing the self-attention optimization perturbation and the cross-attention optimization perturbation using the fusion weights is as follows: in, It is a self-attention optimization perturbation. It is a cross-attention optimization perturbation.

10. A diffusion-based adversarial example generation device for the HR-Net detector, characterized in that, The device includes: The encoding module is used to encode the image into the latent space using a variational autoencoder to generate a latent representation. The perturbation generation module is used to generate perturbations in the latent space for each iteration by the PGD gradient attack method, and after optimizing the perturbations by self-attention and cross-attention, it uses adaptive weights to balance and fuse the perturbations optimized by self-attention and cross-attention to generate the current perturbation. The decoding module is used to decode the latent representation with the added perturbation back into the image space using a variational autodecoder to obtain the current adversarial example; In each iteration, the output of the HR-Net detector to the current adversarial example is calculated, and the effect of the perturbation is enhanced by introducing a preset multi-scale feature-aware perturbation loss function.