An adversarial attack method based on diffusion model and structure perception partition reconstruction
Patent Information
- Application Number
- CN202611047897.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-15
AI Technical Summary
尽管其识别精度极高,但由于深度神经网络固有的脆弱性,该系统极容易遭受对抗样本的攻击
本发明提供的一种基于扩散模型与结构感知分区重构的对抗攻击方法,方法调用人脸解析模型对源面部图像进行语义分割,并根据各面部区域的视觉敏感度分配差异化权重,得到潜在空间语义分割权重图,使得对抗扰动能够根据区域的重要性进行差异化分配。然后通过扩散模型进行逆向去噪,基于模型对当前潜在变量的去噪估计值,分别计算身份攻击梯度和结构保留梯度,再利用潜在空间语义分割权重图对身份攻击梯度和结构保留梯度进行加权融合,得到融合引导梯度,通过将语义解析技术与扩散模型的梯度引导相结合,实现了对抗扰动与结构保护在空间上的自适应协调,从而使得利用融合引导梯度进行迭代更新的对抗样本在获得高攻击成功率的同时,能够保持与源图像高度一致的纹理细节和面部结构保真度,可以兼顾攻击成功率与视觉隐蔽性。
Smart Images

Figure CN122551413B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to an adversarial attack method based on diffusion model and structure-aware partition reconstruction. Background Technology
[0002] In recent years, facial recognition systems have been widely deployed in security, authentication, and surveillance. Despite their extremely high recognition accuracy, the inherent vulnerability of deep neural networks makes these systems highly susceptible to adversarial attacks. In black-box attack scenarios, attackers can effectively deceive the unknown internal structure of the recognition model by injecting minute perturbations into the input image, causing it to make incorrect judgments. This poses a security threat to facial recognition systems.
[0003] Existing adversarial facial attack methods mainly face the following technical bottlenecks: The first category is gradient-based pixel perturbation methods, which often generate high-frequency noise distortion in sensitive facial areas, failing to maintain the structural consistency of facial features and resulting in poor visual naturalness. The second category is patch-based attack methods, whose local patterns are visually too abrupt, lacking adaptive control over different facial regions, and exhibiting extremely low concealment. The third category is generative model-based methods (especially attacks based on latent diffusion models). While these methods can generate realistic textures, they often apply blind global perturbations or large-scale semantic modifications during the attack generation process. Due to the lack of explicit constraints on image structure, the models tend to overfit the target identity features during optimization, resulting in a significant deviation from the original geometric layout and fine structure of the source face, producing visually identifiable facial structural distortions and identity drift.
[0004] Therefore, there is an urgent need in this field for a new adversarial attack method that can break through the global perturbation framework and has the ability to finely reshape space, so as to systematically improve the success rate and concealment in black-box attack scenarios. Summary of the Invention
[0005] In view of this, the present invention provides an adversarial attack method based on diffusion model and structure-aware partition reconstruction to improve the success rate and stealth of black-box attacks.
[0006] The technical solution adopted in this invention is: This invention provides an adversarial attack method based on a diffusion model and structure-aware partition reconstruction, comprising: S1 acquires the source facial image and the target identity image, and calls the pre-trained face parsing model, diffusion model, and target face recognition model; The source facial image is semantically segmented using a face parsing model, and weights are assigned based on the visual sensitivity of each facial region to obtain a latent spatial semantic segmentation weight map. The source facial image is mapped to the latent space of the diffusion model to obtain the source latent variables, and an initial latent variable is sampled from Gaussian noise as the starting point for inverse denoising. Inverse denoising is performed starting from the initial latent variables using a diffusion model. During the inverse denoising process, based on the denoised estimate of the current latent variables by the diffusion model, the identity features of the image corresponding to the denoised estimate and the identity features of the target identity image are extracted using the target face recognition model. The identity attack gradient is calculated with the aim of reducing the distance between the two identity features in the feature space. At the same time, the structure preservation gradient is calculated with the aim of maintaining the structural similarity between the image corresponding to the denoised estimate and the source face image. The identity attack gradient and the structure preservation gradient are weighted and fused using the latent space semantic segmentation weight graph to obtain the fused guided gradient. The fusion-guided gradient is injected into the inverse denoising update trajectory of the current step to correct the potential variables of the next step; The reverse denoising process is performed iteratively. After all denoising steps are completed, the final latent variables are decoded and output to obtain the adversarial example.
[0007] Furthermore, the step of using a face parsing model to perform semantic segmentation on the source facial image and assigning weights based on the visual sensitivity of each facial region to obtain a latent spatial semantic segmentation weight map includes: The source facial image is parsed into multiple facial semantic regions using a face parsing model; Based on the pixel categories of facial semantic regions, differentiated weight values are assigned to different semantic regions to obtain a semantic map; among them, regions with high visual sensitivity are given high weight values, and regions with low visual sensitivity are given low weight values. The semantic segmentation weights after weight allocation are downsampled and interpolated to make the spatial resolution of the semantic segmentation weights consistent with the latent space dimension of the diffusion model, and the latent space semantic segmentation weight map in tensor form is output.
[0008] Furthermore, the calculation of the identity attack gradient includes: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; The reconstructed image is input into the face recognition model to extract the identity embedding features of the reconstructed image, and the target identity image is input into the face recognition model to extract the identity embedding features of the target identity image. Calculate the first cosine similarity between the identity embedding features of the reconstructed image and the identity embedding features of the target identity image, and calculate the gradient of the first cosine similarity with respect to the latent variable estimate, as the identity attack gradient.
[0009] Furthermore, the computational structure preserves gradients, including: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; Calculate the structural similarity index between the reconstructed image and the source facial image; The gradient of the structural similarity index with respect to the estimated values of the latent variables is calculated and used as the structural preservation gradient.
[0010] Furthermore, the denoised latent variable estimate at the current time step is specifically a noise-free latent variable estimated from the noisy latent variables and network prediction noise at the current time step using a closed-form solution.
[0011] Furthermore, the step of using the latent space semantic segmentation weight graph to weight and fuse the identity attack gradient and the structure preservation gradient to obtain the fused guided gradient includes: Set a fusion coefficient to balance attack strength and structural fidelity; For each location in the latent space, the weight value of the latent space semantic segmentation weight map is multiplied by the structure preservation gradient to obtain the structure component; the weight value of the latent space semantic segmentation weight map is subtracted from 1 and then multiplied by the identity attack gradient to obtain the attack component. The fusion guidance gradient is obtained by weighted summation of the structural component and the attack component using the fusion coefficient.
[0012] Furthermore, injecting the fused guiding gradient into the inverse denoising update trajectory of the current step to correct the potential variables for the next step includes: In the inverse denoising process of the diffusion model, the fusion guiding gradient is multiplied by the noise scale of the current time step, and the product is then superimposed with the latent variable update calculated based on the predicted denoising tensor and random noise to obtain the latent variables for the next step.
[0013] In summary, the beneficial effects of the present invention are as follows: This invention provides an adversarial attack method based on a diffusion model and structure-aware partitioning reconstruction. The method calls a face parsing model to perform semantic segmentation on the source facial image and assigns differentiated weights based on the visual sensitivity of each facial region, obtaining a latent spatial semantic segmentation weight map. This allows adversarial perturbations to be differentiated according to the importance of each region. Then, inverse denoising is performed using the diffusion model. Based on the model's denoised estimates of the current latent variables, the identity attack gradient and structure preservation gradient are calculated separately. The latent spatial semantic segmentation weight map is then used to weightedly fuse the identity attack gradient and structure preservation gradient to obtain a fused guided gradient. By combining semantic parsing technology with gradient guidance from the diffusion model, adaptive coordination between adversarial perturbations and structure protection in space is achieved. This allows adversarial examples that use the fused guided gradient for iterative updates to achieve a high attack success rate while maintaining highly consistent texture details and facial structure fidelity with the source image, thus balancing attack success rate and visual concealment. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.
[0015] Figure 1 This is a flowchart of an adversarial attack method based on diffusion model and structure-aware partition reconstruction according to the present invention. Figure 2 The overall framework flowchart of the present invention based on diffusion model and structure-aware partition reconstruction. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Unless otherwise specified, the present invention and the various features in the embodiments can be combined with each other, all of which are within the protection scope of the present invention.
[0017] Reference Figure 1 and Figure 2 As shown in the figure, an adversarial attack method based on diffusion model and structure-aware partition reconstruction provided by an embodiment of the present invention includes: S10. Obtain the source facial image and the target identity image, and call the pre-trained face parsing model, diffusion model and target face recognition model; S20. Use a face parsing model to perform semantic segmentation on the source face image, and assign weights according to the visual sensitivity of each face region to obtain a latent space semantic segmentation weight map. S30. Map the source face image to the latent space of the diffusion model to obtain the source latent variables, and sample an initial latent variable from the Gaussian noise as the starting point for inverse denoising; S40. Perform inverse denoising starting from the initial latent variables using a diffusion model. During the inverse denoising process, based on the denoised estimate of the current latent variables using the diffusion model, extract the identity features of the image corresponding to the denoised estimate and the identity features of the target identity image using the target face recognition model. Calculate the identity attack gradient with the aim of reducing the distance between the two identity features in the feature space. Simultaneously, calculate the structure preservation gradient with the aim of maintaining the structural similarity between the image corresponding to the denoised estimate and the source face image. S50. The identity attack gradient and the structure preservation gradient are weighted and fused using the latent space semantic segmentation weight graph to obtain the fused guided gradient. S60. Inject the fused guiding gradient into the inverse denoising update trajectory of the current step to correct the potential variables of the next step; S70. Iteratively execute the reverse denoising. After all denoising steps are completed, decode and output the final latent variables to obtain the adversarial example.
[0018] The adversarial attack method of this invention combines the generation prior of the latent diffusion model with the dual gradient fusion mechanism based on face parsing, aiming to significantly improve the attack success rate and structural fidelity of adversarial samples in various scenarios.
[0019] It should be noted that in step S10, the source facial image serves as the spatial and textural reference source for structure preservation, while the target identity image serves as the feature direction target for identity attacks. Both are selected as image data sources by using high-quality sets of source facial images and sets of target identity images to be forged. The diffusion model uses an existing latent diffusion model, specifically the LDM version of Stable Diffusion, to provide a bidirectional mapping from pixel space to latent space. StableDiffusion is an open-source AI painting model developed by Stability AI, capable of quickly generating high-quality images based on text descriptions. The diffusion model includes a variational autoencoder (VAE), a U-Net denoising network, and a text encoder. The VAE is responsible for the bidirectional mapping between pixel space and latent space, and consists of an encoder and a decoder. U-Net is the core model of the diffusion process, performing iterative denoising tasks in the latent space. The text encoder converts the input text prompts into semantic embedding vectors, providing conditional control signals for the U-Net denoising process.
[0020] Simultaneously, this invention also incorporates a target face recognition model for implementing the supervision signal. This target face recognition model employs existing recognition models, such as IRSE50, IR152, and FaceNet, as well as the pre-trained face semantic parsing model EHANet. This step provides the data and model foundation for subsequent gradient extraction and reconstruction.
[0021] During pre-training, face parsing models are typically trained under supervision on large-scale face semantic segmentation datasets, such as the CelebAMask-HQ dataset (containing approximately 30,000 high-resolution face images and their corresponding fine semantic annotations). Model parameters are optimized using the cross-entropy loss function or a variant loss function with boundary weights. A fully trained model can achieve accuracy exceeding 78% on metrics such as mIoU (mean intersection-over-union ratio) and can process high-resolution images at speeds exceeding 55 FPS on a single consumer-grade GPU, meeting the demands of real-time applications.
[0022] The diffusion model is pre-trained on large-scale image and text datasets (such as LAION-5B), possessing strong generative prior knowledge, which can provide rich facial generation capabilities as constraints for subsequent adversarial attacks.
[0023] Target face recognition models are typically trained under supervision on large-scale face recognition datasets (such as MS-Celeb-1M, CASIA-WebFace, etc.). The widely used loss function in training is ArcFace (Additive Angular Margin Loss), which imposes strict classification margin constraints in the feature angular space, causing features of the same identity to be tightly aggregated in the high-dimensional space, while features of different identities are kept far apart.
[0024] Specifically, in step S20, the source facial image is semantically segmented at the pixel level using a face parsing model, and differentiated weights are assigned based on the visual perception sensitivity of each facial region, ultimately obtaining a semantic weight map aligned with the latent space dimension of the diffusion model. This semantic weight map, as a spatially modulated constraint matrix, provides a spatial basis for subsequent adaptive fusion of dual gradients. Traditional adversarial attack methods typically apply a uniform perturbation intensity constraint to the entire image, ignoring the significant differences in visual perception among different facial regions, resulting in a lack of targeted perturbation allocation. This step, based on the current source facial image, uses a parsing network to extract independent semantic regions. Unlike the globally uniform processing logic of traditional methods, this invention formulates a dynamic weighting rule based on the differences in human visual adversarial perturbation sensitivity. By applying high constraint weights to sensitive regions (such as facial features) and low constraint weights to non-sensitive regions (such as the background), a semantic weight map matrix aligned to the latent space is finally constructed, allowing the perturbation intensity to be adaptively allocated according to the importance of the region, thus laying the foundation for spatial coordination between aggression and fidelity in subsequent steps.
[0025] In this invention, after acquiring the source facial image and the target identity image, the method calls a face parsing model to perform semantic segmentation on the source facial image and assigns differentiated weights based on the visual sensitivity of each facial region, obtaining a latent space semantic segmentation weight map. This allows adversarial perturbations to be differentiated according to the importance of each region. Then, inverse denoising is performed using a diffusion model. Based on the model's denoised estimates of the current latent variables, the identity attack gradient and structure preservation gradient are calculated separately. The latent space semantic segmentation weight map is then used to weightedly fuse the identity attack gradient and structure preservation gradient to obtain a fused guided gradient. By combining semantic parsing technology with gradient guidance from the diffusion model, adaptive coordination between adversarial perturbations and structure protection in space is achieved. After determining the optimal gradient direction after nonlinear modulation (i.e., the fused guided gradient), the fused guided gradient is forcibly injected into the inverse denoising trajectory of the diffusion model (DDIM). Through repeated guidance within a set number of iterations, the final decoded output is an adversarial example with both extremely high visual concealment and powerful target deception effectiveness.
[0026] As a further refinement of the present invention, refer to Figure 2 As shown, in S20, a face parsing model is used to perform semantic segmentation on the source facial image, and weights are assigned according to the visual sensitivity of each facial region to obtain a latent spatial semantic segmentation weight map, including the following process: By using a face parsing model to parse the source facial image into multiple facial semantic regions, the segmented pixel categories can be identified. Assign corresponding weights Differential weight values are assigned to different semantic regions based on the pixel categories of facial semantic regions to obtain a semantic map; regions with high visual sensitivity are given high weight values, and regions with low visual sensitivity are given low weight values.
[0027] Specifically, this invention first uses a deep face analysis model to process the source input sample (i.e., the source facial image). It outputs a multi-channel semantic mask. To precisely control the perturbation boundary, this invention pre-establishes a strict visual sensitivity mapping table. For any extracted pixel... Its weight The allocation logic is strictly quantified, as shown in the following formula: ; This invention assigns extremely high weights of 0.95 and 0.80 to "Eyes" and "Mouth," which determine facial structural features, respectively; and intermediate weights of 0.70, 0.55, 0.40, and 0.30 to "Skin / Cheeks," "Nose / Eyebrows," "Ears," and "Neck," respectively; while extremely low weights of only 0.10 and 0.05 are assigned to "Hair" and "Background," which have extremely high visual tolerance.
[0028] The semantic segmentation weights, after being assigned weights, are downsampled and interpolated to ensure that the spatial resolution of the semantic segmentation weights matches the latent space dimension of the diffusion model, outputting a latent space semantic segmentation weight map in tensor form. Since the generation process of the diffusion model occurs in a highly compressed latent space, dimensionality adaptation of the aforementioned pixel-level mapping map is necessary. This invention employs a bilinear interpolation and downsampling algorithm to compress the high-resolution mask into a tensor with the same dimension as the latent variables (e.g., ...). ), constructing the final latent space semantic weight graph Finally, the This will serve as the core spatial adjustment matrix, running through all subsequent gradient fusion steps. This mechanism ensures that subsequent gradient modulation can play a pixel-level guiding role in the underlying feature space.
[0029] After performing the latent space mapping and denoising initialization steps, step S30 of this invention maps the source facial image to the latent space through the VAE encoder of the diffusion model to obtain source latent variables, and samples an initial latent variable from Gaussian noise as the starting point for inverse denoising. Thus, the starting point for image generation is placed in the compressed latent space of the diffusion model, rather than the original pixel space, enabling subsequent gradient guidance and image editing to be performed efficiently in a low-dimensional space. The rich facial generation priors learned during the pre-training of the diffusion model provide strong generation constraints for the gradual evolution of the latent variables.
[0030] As a further refinement of the present invention, calculating the identity attack gradient in S40 includes: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; The reconstructed image is input into the target face recognition model to extract the identity embedding features of the reconstructed image, and the target identity image is input into the target face recognition model to extract the identity embedding features of the target identity image. Calculate the first cosine similarity between the identity embedding features of the reconstructed image and the identity embedding features of the target identity image, and calculate the gradient of the first cosine similarity with respect to the latent variable estimate, as the identity attack gradient.
[0031] As a further refinement of the present invention, the gradient calculation of the structure in S40 is preserved, including: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; Calculate the structural similarity index between the reconstructed image and the source facial image; The gradient of the structural similarity index with respect to the estimated values of the latent variables is calculated and used as the structural preservation gradient.
[0032] As a further refinement of the present invention, the estimated latent variables after denoising at the current time step in S40 are specifically the noiseless latent variables estimated from the noisy latent variables and network prediction noise at the current time step using a closed-form solution. Based on the specific implementation process of S40 described above, the present invention, in the process of latent space inverse reasoning, not only calculates the identity attack gradient aimed at narrowing the distance to the target identity features, but also simultaneously introduces a structure-preserving gradient for anchoring the geometric layout of the original image. Using the semantic weight graph constructed above as a spatial mask, element-wise multiplicative modulation is performed on these two gradient streams with completely different physical meanings, realizing the nonlinear harmonious reshaping of structural protection and adversarial perturbation in space. This allows the conflict between attack strength and visual fidelity to be finely reconciled in different facial regions, overcoming the defect of blind global perturbation in existing generative attacks.
[0033] As a further refinement of the present invention, in S50, the identity attack gradient and the structure preservation gradient are weighted and fused using the latent space semantic segmentation weight graph to obtain the fused guided gradient, including: Set a fusion coefficient to balance attack strength and structural fidelity; For each location in the latent space, the weight value of the latent space semantic segmentation weight map is multiplied by the structure preservation gradient to obtain the structure component; the weight value of the latent space semantic segmentation weight map is subtracted from 1 and then multiplied by the identity attack gradient to obtain the attack component. The fusion guidance gradient is obtained by weighted summation of the structural component and the attack component using the fusion coefficient.
[0034] Based on the above calculation process of the fusion-guided gradient, it can be seen that in positions where the semantic weight value approaches 1 (such as the eye region), the structural component is close to the original structure-preserving gradient, while the attack component is significantly suppressed; in positions where the semantic weight value approaches 0 (such as the background region), the attack component is close to the original identity attack gradient, while the structural component is significantly suppressed; in the transition region where the weight value is between 0 and 1, the two types of gradients participate in the fusion proportionally. The fusion coefficient provides a global-level balance adjustment, allowing the attack strength and structural fidelity to be coordinated overall.
[0035] In a specific embodiment, the actual execution process of steps S40 and S50 is as follows: First, in order to obtain the deviation between the current generated state and the target, a closed-form solution is used from the current time step. noise latent variables and network prediction noise Estimating noiseless latent variables That is, the latent variable estimate after denoising at the current time step.
[0036] Then through the decoder Decoding the noiseless latent variables and projecting them onto the image domain yields the corresponding reconstructed image. The target face recognition model can then be used to extract the embedding features of the reconstructed image. Similarly, the target face recognition model can be used to extract the identity embedding features of the target identity image. The identity attack gradient needs to be calculated. Defined as a reconstruction feature With target identity characteristics Cosine similarity between The partial derivative of . The direction of this gradient represents the steepest ascent path that "disguises" the image as the target's identity.
[0037] Then, calculate the identity attack gradient. Its target guides the generated image's features to converge towards the target's identity. At time step... Estimates of clean latent variables were predicted based on diffusion networks. Through the decoder The reconstructed embedding function of the target face recognition model is input. In the middle, maximize its resemblance to the target identity image. Maximize the cosine similarity between the reconstructed image and the target identity image in the feature space of the target face recognition model: ; in, These are the clean latent variable predictions calculated algebraically at each time step. For decoder, For the embedding function of the target face recognition model, Image for target identity.
[0038] Secondly, to counteract the structural anomalies that may occur during the above process, a structure-preserving gradient is introduced. It is defined as reconstructing an image. Compared with the original source image The Structural Similarity Index (SSIM) between them is used for The partial derivative of is used to constrain the spatial and texture consistency between the generated image and the source facial image. The corresponding calculation process is shown in the following formula: ; in, This is the structural similarity function.
[0039] Finally, nonlinear spatial fusion is performed, utilizing the latent spatial semantic segmentation weight graph. The two gradients mentioned above are adaptively fused to obtain the fused guided gradient, as shown in the following equation: ; in, This represents element-wise multiplication. The preset fusion coefficient is used to balance the strength of the resistance and the structural fidelity, and is set to... =0.5. To guide the fusion gradient.
[0040] Under the constraints of the above formula, In the region where the value approaches 1, the first term on the right-hand side of the equation (identity attack) is severely suppressed, while the second term (structure preservation) dominates; and in... In regions where the value approaches 0, the situation is completely reversed. This multiplicative modulation enables a delicate adaptive cutting of the gradient signal in the image space.
[0041] As a further refinement of the present invention, in S60, the fused guiding gradient is injected into the inverse denoising update trajectory of the current step to correct the potential variables of the next step, including: In the inverse denoising process of the diffusion model, the fusion guiding gradient is multiplied by the noise scale of the current time step. The product is then superimposed with the latent variable update calculated based on the predicted denoising tensor and random noise to obtain the latent variables for the next step. The noise scale dynamically changes as the denoising time step progresses, thereby automatically adjusting the gradient injection amplitude.
[0042] Specifically, the fusion-guided gradient of this invention is superimposed on the standard denoising update trajectory as an additional adversarial perception correction term. After multiplying by the noise scale, the magnitude of the gradient injection can adaptively match the noise level at each time step. In the early stages of denoising, when the noise level is high and the noise scale is large, the gradient injection magnitude is correspondingly larger to quickly establish identity feature tendencies at a coarse-grained level. In the later stages of denoising, when the noise level is low and the noise scale approaches zero, the gradient injection magnitude is correspondingly smaller to facilitate fine-tuning of the structure at a fine-grained level.
[0043] Specifically, this invention determines the fusion guiding gradient. Then, it needs to be directly coupled into the underlying logic of the diffusion model DDIM iterative update. Its iterative update process consists of three parts: a baseline term based on the predicted clean latent variables, a drift term based on the predicted noise, and the core... The driving adversarial guiding term. Finally, the fused guiding gradient... Injected into the inverse sampling trajectory of the Diffusion Model (DDIM). The next latent tensor with adversarial perception constraints. The iterative update formula is: ; in, The cumulative noise progress factor. For step-size dependent noise scales, The denoised tensor predicted by the U-Net network. It is random Gaussian noise.
[0044] Secondly, a total of T iterative sampling steps are performed along the above trajectory, so that the model can continuously absorb the adversarial features assigned across space during each denoising step.
[0045] Finally, after completing T-step sampling, the refined final latent tensor is... Input to VAE decoder, output final adversarial example This sample represents the output that balances stealth and strong attack power.
[0046] The entire reverse denoising process starts from the pure Gaussian noise distribution Initially, after an optimized T=200-step iterative loop (experiments show that T=200 achieves the highest ASR gain while maintaining extremely low FID), the system outputs implicit latent variables containing all adversarial features after completing all T-step denoising and gradient injection. .
[0047] Finally, through the decoder Perform a complete upsampling decoding to output a high-fidelity, highly transferable final adversarial image. Referring to the ablation experiment results in Table 1, the combined impact of the dual gradient mechanism of the proposed method (DPS-attack) on the final attack success rate and FID metric is revealed. Experimental data for ASR (Automatic Speech Recognition), SSIM (Structural Similarity Index), and FID (French Initial Distance) show that the adversarial examples generated by this invention not only achieve a state-of-the-art attack success rate exceeding 85% on both the CelebA-HQ and FFHQ datasets, but also perfectly preserve the original hair-level texture details, greatly enhancing the model's application potential in black-box penetration scenarios.
[0048] Table 1 Ablation Experiment Results ; Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for adversarial attack based on diffusion model and structure-aware partition reconstruction, characterized in that, include: Obtain the source facial image and the target identity image, and call the pre-trained face parsing model, diffusion model and target face recognition model; The source facial image is semantically segmented using a face parsing model, and weights are assigned based on the visual sensitivity of each facial region to obtain a latent spatial semantic segmentation weight map. The source facial image is mapped to the latent space of the diffusion model to obtain the source latent variables, and an initial latent variable is sampled from Gaussian noise as the starting point for inverse denoising. Inverse denoising is performed starting from the initial latent variables using a diffusion model. During the inverse denoising process, based on the denoised estimate of the current latent variables by the diffusion model, the identity features of the image corresponding to the denoised estimate and the identity features of the target identity image are extracted using the target face recognition model. The identity attack gradient is calculated with the aim of reducing the distance between the two identity features in the feature space. At the same time, the structure preservation gradient is calculated with the aim of maintaining the structural similarity between the image corresponding to the denoised estimate and the source face image. The identity attack gradient and the structure preservation gradient are weighted and fused using the latent space semantic segmentation weight graph to obtain the fused guided gradient. The fusion-guided gradient is injected into the inverse denoising update trajectory of the current step to correct the potential variables of the next step; The reverse denoising process is performed iteratively. After all denoising steps are completed, the final latent variables are decoded and output to obtain the adversarial example.
2. The adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 1, characterized in that, The step of using a face parsing model to perform semantic segmentation on the source facial image and assigning weights based on the visual sensitivity of each facial region to obtain a latent spatial semantic segmentation weight map includes: The source facial image is parsed into multiple facial semantic regions using a face parsing model; Based on the pixel categories of facial semantic regions, differentiated weight values are assigned to different semantic regions to obtain a semantic map; among them, regions with high visual sensitivity are given high weight values, and regions with low visual sensitivity are given low weight values. The semantic segmentation weights after weight allocation are downsampled and interpolated to make the spatial resolution of the semantic segmentation weights consistent with the latent space dimension of the diffusion model, and the latent space semantic segmentation weight map in tensor form is output.
3. The adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 1, characterized in that, The calculation of the identity attack gradient includes: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; The reconstructed image is input into the face recognition model to extract the identity embedding features of the reconstructed image, and the target identity image is input into the face recognition model to extract the identity embedding features of the target identity image. Calculate the first cosine similarity between the identity embedding features of the reconstructed image and the identity embedding features of the target identity image, and calculate the gradient of the first cosine similarity with respect to the latent variable estimate, as the identity attack gradient.
4. The adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 1, characterized in that, The computational structure preserves gradients, including: Based on the latent variable estimates after denoising at the current time step, the latent variable estimates are decoded by the decoder of the diffusion model to obtain the reconstructed image; Calculate the structural similarity index between the reconstructed image and the source facial image; The gradient of the structural similarity index with respect to the estimated values of the latent variables is calculated and used as the structural preservation gradient.
5. An adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 3 or 4, characterized in that, Specifically, the denoised latent variable estimate at the current time step is the noise-free latent variable estimated from the noisy latent variables and network prediction noise at the current time step using a closed-form solution.
6. The adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 1, characterized in that, The step of using the latent space semantic segmentation weight graph to weight and fuse the identity attack gradient and the structure preservation gradient to obtain the fused guided gradient includes: Set a fusion coefficient to balance attack strength and structural fidelity; For each location in the latent space, the weight value of the latent space semantic segmentation weight map is multiplied by the structure preservation gradient to obtain the structure component; the weight value of the latent space semantic segmentation weight map is subtracted from 1 and then multiplied by the identity attack gradient to obtain the attack component. The fusion guidance gradient is obtained by weighted summation of the structural component and the attack component using the fusion coefficient.
7. The adversarial attack method based on diffusion model and structure-aware partition reconstruction according to claim 1, characterized in that, The fused guiding gradient is injected into the inverse denoising update trajectory of the current step to correct the potential variables for the next step, including: In the inverse denoising process of the diffusion model, the fusion guiding gradient is multiplied by the noise scale of the current time step, and the product is then superimposed with the latent variable update calculated based on the predicted denoising tensor and random noise to obtain the latent variables for the next step.
Citation Information
Patent Citations
Diffusion model customized privacy protection method and system based on mask attention mechanism elimination
CN120337271A
Semantically constrained visual recommendation system diffusion attack method and system
CN121235772A