Adversarial disturbance generation method for image tampering prevention

Through global feature spatial correlation damage and local face attribute distortion, an anti-perturbation generation method for image tamper prevention is generated, which solves the problems of weak resistance to concept migration and insufficient toxicity in the prior art, and achieves a stronger image tamper prevention effect.

CN120543367APending Publication Date: 2025-08-26BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510721161.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing anti-customization method fails to fully utilize deep facial image features, resulting in weak resistance to concept migration and insufficient toxicity, which makes it impossible to effectively protect personal privacy and copyright.

Method used

Through global feature spatial correlation corruption and local face attribute distortion, an adversarial perturbation generation method is generated for image tampering prevention, including global feature correlation corruption losses and local face attribute distortion losses, building GoodAC losses, and combining feature loss and LDM denoising losses to generate adversarial perturbation.

Benefits of technology

It enhances the conceptual migration resistance and semantic theft ability to combat customized samples, and improves the effectiveness and image quality of image tampering prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543367A_ABST
    Figure CN120543367A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of anti-customization of artificial intelligence generated contents, and particularly discloses an adversarial disturbance generation method for image tampering prevention, which comprises the following steps of: destroying spatial correlation of deep learning model perception features on a global level, enhancing the robustness of a protected image to concept transfer, and generating a confrontation disturbance of the protected image; obtaining global feature correlation damage loss; a personalized precise face attribute attack strategy is generated on a local level, attacks are concentrated on an individual image structure, the toxicity of resisting disturbance is enhanced, and local face attribute distortion loss is obtained; constructing a GoodAC loss based on the global feature correlation damage loss and the local face attribute distortion loss; and furthermore, an adversarial disturbance for image tampering prevention is generated. According to the method, the problems that the concept migration resistance is weak and the toxicity is insufficient due to the fact that deep face image features, namely image contents and face attributes, cannot be fully utilized in an existing anti-customization method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of anti-customization of artificial intelligence-generated content, and specifically relates to a method for generating adversarial disturbances to prevent image tampering. Background Art

[0002] Currently, artificial intelligence generated content (AIGC) has formed a technological framework covering multimodal scenarios. In particular, image generation technology based on diffusion models (DM) has become a key representative of AIGC and a research hotspot. Diffusion models have achieved significant success in personalized image synthesis, but due to unauthorized misuse, they pose significant privacy risks to society.

[0003] Anti-customization technology aims to protect personal privacy and copyright by processing the original image so that when used to train a customized model, the generated results do not accurately reflect the characteristics of the original object. Currently, several anti-customization methods have been proposed. For example, adversarial examples are used to add tiny perturbations to the original image that are imperceptible to the human eye, preventing the model from correctly learning and reproducing image features during training or generation. Other methods interfere with the model's learning process by adding watermarks, noise, and style transfer.

[0004] However, these methods have several limitations. First, most of them ignore the intrinsic properties and generative mechanisms of the diffusion model and employ direct adversarial design, which compromises optimization effectiveness and results in low attack success rates (referred to as customization, where the de-customized sample can learn the features of the original image, for example, the FDFR face detection failure rate is a metric). These methods require significant computational resources and time. Furthermore, these solutions often lack attention to the key elements of the protected content and targeted design of the operating mechanisms of the target model, resulting in unsatisfactory optimization effectiveness and efficiency. Second, these methods lack the ability to resist concept transfer (i.e., transferring a learned new concept to a new context renders it invulnerable to concept manipulation) and generalization (i.e., weak resistance to various customization methods). Adversarial perturbations are often only effective against specific model architectures or training settings, and their effectiveness significantly degrades when applied to different or updated models. However, the architecture and training of stealing models in real-world scenarios are diverse and complex. Furthermore, when processing high-resolution images, these methods can lead to degraded image quality, impacting user experience and practical applications. This poses challenges for practical applications and fails to fully address the risk of diffusion model abuse.

[0005] In summary, existing AC methods fail to fully utilize deep facial image features, namely image content and facial attributes, resulting in weak protection against concept transfer and insufficient toxicity, resulting in unsatisfactory protection. Therefore, how to leverage the intrinsic properties and generative mechanisms of diffusion models to develop anti-customization technologies with high resistance to concept transfer and semantic theft has become a pressing scientific issue. Summary of the Invention

[0006] The purpose of this invention is to solve the problem that existing anti-customization methods fail to fully utilize deep facial image features, namely image content and facial attributes, resulting in weak resistance to concept transfer and insufficient toxicity. A method for generating adversarial perturbations to prevent image tampering is proposed.

[0007] The technical solution of the present invention is: a method for generating an adversarial disturbance for preventing image tampering, comprising the following steps: Perform global feature space correlation destruction: By destroying the spatial correlation of the deep learning model's perceptual features at the global level, the robustness of the protected image to concept transfer is enhanced, and the global feature correlation destruction loss is obtained; Perform local facial attribute distortion: By generating personalized and precise facial attribute attack strategies at the local level, the attack is focused on the individual image structure, enhancing the toxicity of the adversarial perturbation and obtaining the local facial attribute distortion loss; Based on the global feature correlation destruction loss and the local facial attribute distortion loss, the GoodAC loss is constructed; Obtain feature loss based on the feature map of the perturbed image and the feature map of the clean image; Obtain LDM denoising loss based on latent diffusion model; GoodAC loss, feature loss and LDM denoising loss are combined to generate adversarial perturbations for preventing image tampering.

[0008] Preferably, the global feature space association destruction specifically includes the following steps: Input the protected image into the feature extraction module of the deep learning model for feature extraction, and output a feature map; Performing block angle transformation on the feature map; the block angle transformation includes dividing the feature map into blocks and rearranging them, and performing angle transformation on the feature map; The difference of feature maps before and after transformation is calculated to construct the global feature correlation destruction loss.

[0009] Preferably, the expression formula for performing block angle transformation on the feature map is:

[0010] in, Indicates that after transformation from layer l The feature maps of all channels extracted, Indicates the layer before transformation l The feature maps of all channels extracted, Indicates that the feature map is divided into blocks and rearranged. Indicates the angle transformation of the feature map. Indicates that the feature map is the last upsampling block from UNet and the first downsampling block Extracted from.

[0011] Preferably, the global feature correlation destruction loss is expressed as:

[0012] in, represents the global feature correlation destruction loss, Represents the layer of UNet from the LDM model after transformation l The feature maps of all channels extracted, Represents the UNet layer from the LDM model before transformation l The feature maps of all channels extracted, Represents the layer of UNet in the LDM model l The weight coefficient of Indicates that the feature map is the last upsampling block from UNet and the first downsampling block Extracted from Represents the Frobenius norm, and the input of UNet is the image to be protected.

[0013] Preferably, the performing of local facial attribute distortion specifically includes the following steps: For images that need protection , generate semantic hints to describe user attributes; the semantic hints are:

[0014] in, represents the semantic cue generation function; Select the semantic hint combination corresponding to the top 15 attributes with the highest confidence to get the hint , and then generate image features with accurate facial attributes. The specific formula is:

[0015] in, Represents clean image features with accurate facial attributes, represents the perturbed image features with accurate facial attributes, represents the perturbation image that needs to be protected, Represents the layer of UNet in the LDM model ; Application Detector To identify important local features, the specific formula is:

[0016] in, represents the edge map of the clean image, A feature map representing the perturbed image; based on and , constructing local facial attribute warping loss.

[0017] Preferably, the calculation formula of the local facial attribute distortion loss is:

[0018] in, represents the local facial attribute distortion loss, represents the Frobenius norm.

[0019] Preferably, the calculation formula of the GoodAC loss is:

[0020] in, Indicates GoodAC loss, represents the global feature correlation destruction loss, represents the local facial attribute distortion loss, Represents the weight factor that controls the distortion loss of local facial attributes.

[0021] Preferably, the calculation formula of the feature loss is:

[0022] in, represents the feature loss, express The corresponding expectations, Indicates that at a specific time step and UNet layers The feature map of the perturbed image, Indicates that at a specific time step and layer UNet The feature map of the clean image, represents the Frobenius norm.

[0023] Preferably, the calculation formula of the LDM denoising loss is:

[0024] in, represents the LDM denoising loss, represents the parameters of the diffusion model, An image representing the LDM denoising loss metric, Express expectations, represents the initial noise, represents the noise prediction network, Indicates the t +1 step image, represents the diffusion model "time step" parameter, represents the semantic condition of the diffusion model, represents a Gaussian distribution, represents the Frobenius norm.

[0025] Preferably, the generation formula of the adversarial perturbation for preventing image tampering is:

[0026] in, represents the adversarial perturbation against image tampering prevention, i.e., anti-customized adversarial perturbation, represents the learning rate, represents the symbolic function, express right The gradient, represents the LDM denoising loss, Indicates GoodAC loss, represents the feature loss, represents the weight factor that controls feature loss, represents the perturbation image that needs to be protected, represents the infinite norm, represents the adversarial perturbation cost.

[0027] The beneficial effects of the present invention are: 1. The present invention destroys the spatial correlation in the perceptual features through global feature space association destruction, so that the model loses some stable information related to the concept during the generation process, thereby enhancing the resistance of anti-customized samples to such concept migration.

[0028] 2. The present invention accurately locates and distorts local facial attributes through local facial attribute distortion, thereby destroying the ability of LDM to learn fine-grained semantic information such as facial features, and ultimately improving the effectiveness of anti-customization. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1The figure shows a flow chart of a method for generating an adversarial disturbance for preventing image tampering provided in Example 1 of the present invention.

[0030] Figure 2 The figure shows a schematic diagram of the global-local collaborative anti-customization framework structure provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0031] The exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are merely exemplary and are intended to illustrate the principles and spirit of the present invention, rather than to limit the scope of the present invention.

[0032] Definitions of Abbreviations and Key Terms Text-generated image: refers to giving the input model a text description, and the model can output an image that matches the content of the text description.

[0033] Latent Diffusion Model (LDM): A specific technical school (model) for implementing text-based graph technology, it is currently the de facto standard for text-based graph models. There are also two other schools: autoregressive models (e.g., the DALL-E model) and generative adversarial models (e.g., GANs).

[0034] Customization technology: refers to the technology that enables the basic model of the cultural graph (latent diffusion models LDMs in this article) to learn new specific concepts or styles.

[0035] Adversarial perturbation: refers to a pixel-level perturbation image generated by an adversarial algorithm (occupying about 1% to 5% of the original image, depending on different settings).

[0036] Anti-customized samples: These are new images created by adding adversarial perturbations to the original image (e.g., a selfie or other facial image). These new images prevent certain customization techniques based on text-based graph models from learning or stealing specific concepts. In this embodiment of the present invention, these specifically refer to facial features.

[0037] Define the problem: Assume that the initial image set that the user needs to protect is , define the perturbation function , mapping the initial image set to the protected image set The attacker may obtain an initial image set and use it to fine-tune the noise prediction network of LDM (latent diffusion model) under the DreamBooth paradigm. , so that the latent diffusion model can generate images containing user information in different scenarios (i.e., concept transfer), and the generated images have clear user facial features and peculiar details (i.e., semantic stealing). The goal of this invention is to design an optimal perturbation function , to minimize the training parameters The customization capability of the potential diffusion model is as follows:

[0038] in, represents minimization, represents an evaluation metric used to measure the fidelity of the image generated by the latent diffusion model with the initial image, represents the loss function of DreamBooth, denotes the parameters of the potential diffusion model, represents the constraint against disturbance, represents the infinity norm.

[0039] Example 1: like Figure 1 As shown, a method for generating adversarial perturbations for preventing image tampering includes the following steps: S1. Perform global feature space correlation destruction: By destroying the spatial correlation of the deep learning model's perceptual features at the global level, the robustness of the protected image to concept transfer is enhanced, and the global feature correlation destruction loss is obtained; S2. Perform local facial attribute distortion: By generating personalized and precise facial attribute attack strategies at the local level, the attack is focused on the individual image structure, enhancing the toxicity of the adversarial perturbation and obtaining the local facial attribute distortion loss; S3. Construct GoodAC loss based on global feature correlation destruction loss and local facial attribute distortion loss; S4. Obtain feature loss based on the feature map of the perturbed image and the feature map of the clean image; S5. Obtain LDM denoising loss based on latent diffusion model; S6. Combine GoodAC loss, feature loss and LDM denoising loss to generate adversarial perturbations for preventing image tampering.

[0040] In this embodiment, a global-local collaborative anti-customization (GoodAC) framework is proposed to generate adversarial disguises with strong resistance to concept transfer and semantic theft. Figure 2As shown in the figure, this method aims to generate strong adversarial perturbations by collaboratively destroying global perceptual correlations and distorting local facial attributes. The CrossAttnUpBlock2D module in the figure is an upsampling module that combines a cross-attention mechanism with upsampling. When it comes to enhancing resistance to concept transfer, perceptual features are fundamental to image generation, and latent diffusion models rely heavily on the spatial consistency of perceptual features to maintain identity during concept transfer. Therefore, the robustness of protected images to concept transfer is enhanced by globally destroying spatial correlations within the perceptual feature space. Specifically, feature maps are extracted from the first downsampling block and the last upsampling block of the UNet, and block-wise transformation is used to destroy the spatial connectivity within these feature maps. Based on this, embodiments of the present invention design a loss function to optimize adversarial perturbations, effectively preventing concepts in protected images from being successfully transferred to new scenarios. When it comes to defending against semantic theft, given that fine-grained semantic features such as facial features are crucial for protection, accurately identifying these details is crucial. Therefore, it is proposed to apply stronger perturbations locally in key image regions, thereby weakening the latent diffusion model's ability to extract detailed user-specific information from the image. Specifically, a text inversion module is used to obtain precise semantic descriptions as semantic guidance, and edge detection techniques are combined to accurately extract facial details. As a result, the perturbations are effectively concentrated in these key areas, limiting the model's ability to capture personalized and precise facial attributes, thereby increasing the difficulty of accurately reproducing details from the original image.

[0041] In this embodiment, perceptual features refer to the key characteristic representations extracted from specific objects by deep learning models during image processing. These features typically appear in intermediate layers of the model, such as convolutional layers and residual blocks. Existing research has shown that perceptual features are fundamental to both recognition and generation tasks. During customization, the Latent Diffusion Model (LDM) leverages the spatial correlation in these perceptual features to effectively transfer personalized subject concepts to new scenarios for image generation. Inspired by this mechanism, this paper proposes to enhance the resistance of anti-customization samples to this concept transfer by destroying the spatial correlation in perceptual features.

[0042] Specifically, when LDM transfers a concept associated with the identifier "a specific person" to a new context (e.g., "a specific person looking in a mirror"), it relies on a set of spatially ordered perceptual features to maintain the integrity of the concept. If the spatial associations of these features are disrupted, the model loses some of the stable information associated with the concept during generation, making it difficult to preserve the original characteristics of "a specific person," thereby achieving anti-customization.

[0043] In order to destroy the spatial correlation between perceptual features, this embodiment introduces a feature map transformation called block angle transformation. The global feature space correlation destruction is specifically as follows: Input the protected image into the feature extraction module of the deep learning model for feature extraction, and output a feature map; Performing block angle transformation on the feature map; the block angle transformation includes dividing the feature map into blocks and rearranging them, and performing angle transformation on the feature map; The difference of feature maps before and after transformation is calculated to construct the global feature correlation destruction loss.

[0044] Specifically, the feature map is output. On the one hand, the feature map The process of dividing into blocks and rearranging them destroys the internal continuity and spatial correlation of the feature map, making it difficult for the model to capture the correlation pattern. This process is defined as On the other hand, for the feature map Applying angular transformations (such as rotations) changes its global spatial arrangement, thereby weakening the model's understanding of orientation and shape, a process defined as , the whole process can be expressed as:

[0045] in, Indicates that after transformation from layer l The feature maps of all channels extracted, Indicates the layer before transformation l The feature maps of all channels extracted, Indicates that the feature map is the last upsampling block from UNet and the first downsampling block By calculating the difference between the feature maps before and after the transformation, a global feature correlation destruction loss is constructed to more effectively detect adversarial perturbations and interfere with perceptual features. The global feature correlation destruction loss is expressed as:

[0046] in, represents the global feature correlation destruction loss, Indicates that after transformation from layer l The feature maps of all channels extracted, Indicates the layer before transformation l The feature maps of all channels extracted, Presentation Layer l The weight coefficient of Indicates that the feature map is the last upsampling block from UNet and the first downsampling block Extracted from Represents the Frobenius norm. It measures the difference between the transformed feature map and the untransformed feature map. Based on the block-angle transformation strategy, the difference between the transformed feature map and the untransformed feature map is calculated to generate adversarial camouflage, effectively enhancing the anti-customization effect globally.

[0047] In this embodiment, the local face attribute distortion is specifically performed as follows: To enhance the resistance of anti-customization examples to semantic theft, existing AC research has primarily focused on gradient-based optimization methods, exploring the role of high-frequency image features in accelerating the optimization process. However, these methods are unable to effectively segment or identify semantically meaningful information in images for targeted optimization, and therefore have limited ability to mitigate semantic theft. Therefore, our goal is to precisely locate and distort local facial attributes, thereby disrupting the LDM's ability to learn fine-grained semantic information such as facial features, ultimately improving the effectiveness of anti-customization.

[0048] This paper proposes local facial attribute warping, which uses adaptive semantic guidance and edge detection to locate key features in anti-custom perturbations. It includes two key steps: one is to generate a more detailed description of the protected image to guide LDM to encode richer internal features of fine-grained semantic attributes; the other is to apply edge detection between the feature maps of the adversarial disguise and the clean image during the perturbation process, and optimize the perturbation by minimizing the L2 loss between the detected edges.

[0049] The embodiment of the present invention first uses a text reversal module (through<https: / / imageprompt.org / image-to-prompt> Implementing a text reversal module) generates fine-grained image descriptions, which serve as semantic guides to enrich the internal features of LDM. , generating semantic hints ,in, represents the semantic cue generation function constructed by the text reversal module, which describes the user attributes as a set of values ​​(e.g., ). Select the top 15 attributes with the highest confidence and combine them into a refined prompt , guiding LDM to generate image features with accurate facial attributes and improve the accuracy of local interference. The specific formula is:

[0050] in, Represents clean image features with accurate facial attributes, represents the perturbed image features with accurate facial attributes, Represents the perturbation image that needs to be protected.

[0051] Then, features are extracted from the last upsampling block of UNet and the detector is applied To identify important local features, in this embodiment, the detector Canny edge detection is used for recognition, and the specific formula is:

[0052] in, represents the edge map of the clean image, Represents the feature map of the perturbed image. and , construct the local facial attribute distortion loss, which is expressed as:

[0053] in, represents the local facial attribute distortion loss, represents the Frobenius norm.

[0054] In summary, our proposed method leverages precise semantic properties and edge-focused optimization to target key image features, effectively weakening the model's ability to extract sensitive facial information, thereby providing strong protection.

[0055] In this embodiment, the calculation formula of the GoodAC loss is:

[0056] in, Indicates GoodAC loss, represents the global feature correlation destruction loss, represents the local facial attribute distortion loss, Represents the weight factor that controls the distortion loss of local facial attributes.

[0057] In this embodiment, the calculation formula of the feature loss is:

[0058] in, represents the feature loss, Express expectations, Indicates that at a specific time step and layer The feature map of the perturbed image, Indicates that at a specific time step and layer The feature map of the clean image, represents the Frobenius norm.

[0059] In this embodiment, the calculation formula of the LDM denoising loss is:

[0060] in, represents the LDM denoising loss, represents the parameters of the diffusion model, An image representing the LDM denoising loss metric, Express expectations, represents the initial noise, represents the noise prediction network, Indicates the t +1 step image, represents the diffusion model "time step" parameter, represents the semantic condition of the diffusion model, represents a Gaussian distribution, represents the Frobenius norm.

[0061] In this embodiment, the generation formula of the adversarial perturbation for preventing image tampering is:

[0062] in, represents the adversarial perturbation against image tampering prevention, i.e., anti-customized adversarial perturbation, represents the learning rate, represents the symbolic function, express right The gradient, represents the LDM denoising loss, Indicates GoodAC loss, represents the feature loss, represents the weight factor that controls feature loss, represents the perturbation image that needs to be protected, represents the infinite norm, represents the adversarial perturbation cost.

[0063] Example 2: Based on Example 1, this example uses the CelebA-HQ and VGGFace2 datasets to evaluate the performance of anti-customization samples generated using the method provided by this invention in resisting information theft. The evaluation is conducted from two perspectives: toxicity and concept transfer.

[0064] Experimental setup: First, GoodAC and various baseline model attacks are performed on SDv-2.1 to generate adversarial samples. Then, SDv-2.1 is fine-tuned using the Dreambooth method and reasoning is performed using four text prompts according to SimAC. In the sks individual, the training terms are synonymous with the inference terms. The other three examples involve concept transfer, in which the concept of "sks person" is transferred to other contexts. For each prompt, 32 generated images are randomly selected to calculate each indicator and report their average values. As shown in Tables 1 and 2, Tables 1 and 2 are comparisons of the method proposed in the present invention with other open source anti-customization methods on the CelebA-HQ dataset and the VGGFace2 dataset, respectively. The embodiment of the present invention evaluates the performance under four different prompts during the customization process, and the underline represents the baseline method.

[0065] The present invention uses four evaluation metrics to measure the effectiveness of anti-customization. Specifically, they include: (1) Face Detection Failure Rate (FDFR), which measures whether a face can be detected in the generated image; (2) Identity Score Matching (ISM), which measures the identity similarity between the generated image and the original image; (3) SER-FQA, a facial image quality assessment metric; and (4) BRISQUE, a widely used general-purpose no-reference image quality assessment metric.

[0066] According to the experimental results, the GoodAC of the present invention outperforms other state-of-the-art methods, and the following conclusions are drawn: 1. When the training item is also the inference item, our adversarial disguise achieves significantly better performance on both datasets (maximum ISM reduction exceeds 50% and FDFR success rate exceeds 95% on CelebA-HQ, and maximum ISM reduction exceeds 50% and FDFR success rate exceeds 90% on VGGFace2).

[0067] Previous studies have generally performed poorly in resisting concept transfer. The present invention hypothesizes that this is due to insufficient consideration of the importance of spatially aware features in learning visual representations and their fundamental role in the task of generating diffuse models. In contrast, our GoodAC method exploits this by disrupting a set of spatially ordered perceptual features, thereby disrupting the integrity of sks-based concepts during model concept transfer, making them difficult to transfer (e.g., the Eiffel Tower). This method achieves significant results, with ISM reductions exceeding 50%.

[0068] 3. In actual experiments, SimAC performs weaker than CelebA-HQ on VGGFace2. Detailed analysis shows that SimAC performs significantly worse on some faces (e.g., n00005 in VGGFace2). This is likely due to the unique facial attributes used in the model's recognition and generation processes, which SimAC and previous studies have not fully accounted for.

[0069] Table 1 Results of the CelebA-HQ dataset

[0070] Table 2 VGGFace2 dataset results

[0071] AdvDM (Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples), published at ICML 2023, is designed to protect human-created artwork from being illegally imitated and generated by AI-for-Art applications based on diffusion models (DMs). It is effective against text inversion (a customized technique). Its core idea is to generate adversarial examples to interfere with the feature extraction and generation process of the diffusion model, thereby preventing unauthorized artwork from being learned, imitated, or copied by the diffusion model.

[0072] AdvDM suffers from the following issues: The texture, semantic, and fusion modes it provides require manual configuration of weight parameters, lacking dynamic adaptability. Experiments show that its fusion mode can easily lead to local over-perturbations (e.g., grid-like artifacts) in complex background images. AdvDM's defenses have only been validated against textual inversion and image-to-image attacks, and do not consider combined attack scenarios (e.g., combining DreamBooth or Lora fine-tuning with style transfer). Unfortunately, fine-tuning techniques based on DreamBooth and Lora are the most important customization methods for conceptual facial manipulation. Furthermore, grid-like artifacts can have a significant visual impact.

[0073] Anti-DB (Anti-DreamBooth: Protecting Users from Personalized Text-to-Image Synthesis) was published at ICCV 2023. By adding imperceptible adversarial noise to user images, the anti-customized samples significantly degrade the quality of the generated images when DreamBooth is used for semantic stealing, thereby protecting users from harm. The algorithm consists of two steps: alternating training: In each iteration, a proxy model is first fine-tuned using clean images, and then this model is used to optimize the adversarial noise. Model update: The proxy model is updated with the optimized noise, and this process is repeated to better simulate the fine-tuning process of malicious users.

[0074] The core flaw of Anti-DB is that it indiscriminately applies limited perturbations to every time step of the diffusion model. Unfortunately, the perturbations generated at some time steps are effective in improving the final anti-customization capabilities, and fail to explore the unique information of the LDM model structure. This makes Anti-DB less capable of defending against semantic theft.

[0075] SimAC (SimAC: A Simple Anti-Customization Method for Protecting FacePrivacy against Text-to-Image Synthesis of Diffusion Models), published at CVPR 2024, aims to protect user privacy and prevent the malicious customization of diffusion models (DMs) during text-to-image generation. SimAC addresses potential vulnerabilities in diffusion models by introducing adversarial noise and optimization strategies to disrupt the model's ability to reconstruct user input images, thereby protecting user privacy.

[0076] The main disadvantage of SimAC is that it focuses too much on the gradient information of the model itself and uses it for perturbation generation. Existing research shows that focusing only on gradient information perturbation will lead to localization and poor resistance to concept transfer.

[0077] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A method for generating adversarial perturbations for preventing image tampering, characterized in that: The following steps are involved: Perform global feature space correlation destruction: By destroying the spatial correlation of the deep learning model's perceptual features at the global level, the robustness of the protected image to concept transfer is enhanced, and the global feature correlation destruction loss is obtained; Perform local facial attribute distortion: By generating personalized and precise facial attribute attack strategies at the local level, the attack is focused on the individual image structure, enhancing the toxicity of the adversarial perturbation and obtaining the local facial attribute distortion loss; Based on the global feature correlation destruction loss and the local facial attribute distortion loss, the GoodAC loss is constructed; Obtain feature loss based on the feature map of the perturbed image and the feature map of the clean image; Obtain LDM denoising loss based on latent diffusion model; Combine GoodAC loss, feature loss and LDM denoising loss to generate adversarial perturbations for preventing image tampering.

2. The method for generating adversarial disturbances for preventing image tampering according to claim 1, characterized in that: The global feature space association destruction specifically includes the following steps: Input the protected image into the feature extraction module of the deep learning model for feature extraction, and output a feature map; Performing block angle transformation on the feature map; the block angle transformation includes dividing the feature map into blocks and rearranging them, and performing angle transformation on the feature map; The difference of feature maps before and after transformation is calculated to construct the global feature correlation destruction loss.

3. The method for generating an adversarial disturbance for preventing image tampering according to claim 2, wherein: The expression formula for performing block angle transformation on the feature map is: in, represents the feature maps of all channels extracted from layer l after transformation, F l represents the feature maps of all channels extracted from layer l before transformation, T block Indicates that the feature map is divided into blocks and rearranged, T rot Indicates the angle transformation of the feature map, l∈{l down ,l up } indicates that the feature map is from the last upsampling block l of UNet up and the first downsampling block l down Extracted from.

4. The method for generating an adversarial disturbance for preventing image tampering according to claim 2, wherein: The global feature correlation destruction loss is expressed as: Among them, L perceptual represents the global feature correlation destruction loss, represents the feature maps of all channels extracted from layer l of UNet in the LDM model after transformation, F l represents the feature maps of all channels extracted from layer l of UNet in the LDM model before transformation, λ l Represents the weight coefficient of layer l of UNet in the LDM model, l∈{l down ,l up } indicates that the feature map is from the last upsampling block l of UNet up and the first downsampling block l down , ||·||2 represents the Frobenius norm, and the input of UNet is the image to be protected.

5. The method for generating adversarial disturbances for preventing image tampering according to claim 1, wherein: The local face attribute distortion specifically includes the following steps: For the image x that needs to be protected, a semantic hint is generated to describe the user attributes; the semantic hint is: P = TextInversion(x) Among them, TextInversion represents the semantic hint generation function; The semantic clue combination corresponding to the top 15 attributes with the highest confidence is selected to obtain the clue ξ, and then generate image features with accurate facial attributes. The specific formula is: F clean =l(x,ξ),F adv =l(x adv ,x) Among them, F clean represents the clean image features with accurate facial attributes, F adv represents the perturbed image features with accurate facial attributes, x adv represents the perturbation image to be protected, l(·) represents the layer l of UNet in the LDM model; Apply detector C to identify important local features. The specific formula is: E clean =C(F clean ),E adv =C(F adv ) Among them, E clean Represents the edge map of the clean image, E adv A feature map representing the perturbed image; Based on E clean and E adv , constructing local facial attribute warping loss.

6. The method for generating adversarial disturbances for preventing image tampering according to claim 5, wherein: The calculation formula of the local facial attribute distortion loss is: THE edge =||And adv -AND clean ||2 Among them, L edge represents the local facial attribute distortion loss, and ||·||2 represents the Frobenius norm.

7. The method for generating adversarial disturbances for preventing image tampering according to claim 1, wherein: The calculation formula of the GoodAC loss is: L GoodAC =L perceptual +β·L edge Among them, L GoodAC Indicates GoodAC loss, L perceptual represents the global feature correlation destruction loss, L edge represents the local facial attribute distortion loss, and β represents the weight factor controlling the local facial attribute distortion loss.

8. The method for generating adversarial disturbances for preventing image tampering according to claim 1, wherein: The calculation formula of the feature loss is: L feat =E||F l,t,adv -F l,t,clean ||2 Among them, L feat represents feature loss, E represents || F l,t,adv -F l,t,clean ||2 corresponding expectation, F l,t,adv represents the feature map of the perturbed image at a specific time step t and layer l of UNet, F l,t,clean represents the feature map of the clean image at a specific time step t and layer l of UNet, and ||·||2 represents the Frobenius norm.

9. The method for generating adversarial disturbances for preventing image tampering according to claim 1, wherein: The calculation formula of the LDM denoising loss is: Among them, L cond (θ, x0) represents the LDM denoising loss, θ represents the parameters of the diffusion model, and x0 represents the image metric of the LDM denoising loss. represents the expectation, ∈ represents the initial noise, ∈ θ represents the noise prediction network, x t+1 represents the image at step t+1, t represents the diffusion model "time step" parameter, c represents the diffusion model semantic condition, N(0,1) represents the Gaussian distribution, and ||·||2 represents the Frobenius norm.

10. The method for generating adversarial disturbances for preventing image tampering according to claim 1, wherein: The generation formula of the adversarial perturbation for preventing image tampering is: Among them, δ represents the adversarial perturbation for preventing image tampering, that is, anti-customized adversarial perturbation, α represents the learning rate, and sign represents the sign function. Indicates L cond +L GoodAC +γ·L feat x adv The gradient of L cond (θ, x0) represents the LDM denoising loss, L GoodAC Indicates GoodAC loss, L feat represents feature loss, γ represents the weight factor that controls feature loss, x adv represents the perturbation image to be protected, ||·|| ∞ represents the infinity norm, and η represents the adversarial perturbation cost.