Quality-guaranteeing text image joint concept erasing method for diffusion model

By constructing visual semantic sets and performing attention matching methods, the contradiction between erasing effectiveness and universality of concept erasing methods in diffusion models is solved, and more complete and accurate concept erasing and model universality are achieved.

CN120198540APending Publication Date: 2025-06-24UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224101.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-24

Smart Images

  • Figure CN120198540A_ABST
    Figure CN120198540A_ABST
Patent Text Reader

Abstract

The invention discloses a quality guarantee text image joint concept erasing method for a diffusion model, which comprises the following steps of: constructing a visual semantic set, and taking an image containing a target concept as a sample in the visual semantic set; projecting images in the visual semantic set into a text space to obtain image embedding; visual features are extracted from image embedding, and attention matching is carried out on an attention layer of the diffusion model by adopting corresponding text embedding and the visual features; training the diffusion model; and using the trained diffusion model to generate an image according to the text cue word. According to the method disclosed by the invention, the target concept can be erased more completely and accurately, and the non-target concept appearing in the image is ensured not to be erased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for jointly erasing quality-preserving text and images for diffusion models, belonging to the technical field of image recognition. Background Art

[0002] Image generation technology, especially the use of diffusion models, can generate high-quality images with consistent semantics according to text prompts. Despite the great success of diffusion models, they also bring many potential security risks. One of them is that they may be used to generate inappropriate content, such as pornographic images.

[0003] Concept erasure methods, that is, making diffusion models avoid generating content related to a certain concept. In the prior art, words are used to describe the target concept to be erased, and by reducing the generation probability of the target concept or anchoring the target concept to other benign concepts, for example, the methods described in the literature Rohit Gandikota, Joanna Materzynska, Jaden Fiotto Kaufman, and David Bau. Erasing concepts from diffusion models. In ICCV, pages 2426–2436, 2023 and the literature Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu. Ablating concepts in text-to-image diffusion models. In ICCV, pages 22691–22702, 2023, can make the model avoid generating corresponding content when facing text descriptions related to the target concept. However, the existing prompt learning methods can optimize in the text embedding space to find out the text that can still induce the model to generate the target concept, thus destroying the erasing effectiveness of the above methods.

[0004] In addition, there is also a framework using adversarial training in the prior art, which perturbs the words describing the target concept and uses a broader semantics to represent the target concept. For example, the methods described in the literature Jing Wu and Mehrtash Harandi. Scissorhands: Scrub data influence via connection sensitivity in networks. In ECCV, 2024 or the literature Wu Jing, Le Trung, Hayat Munawar, and Harandi Mehrtash. Erasediff: Erasing data influence in diffusion models. arXiv, 2024.

[0005] However, such methods may cause other benign concepts to be wrongly erased, reducing the quality of normal generated content, thus undermining the generality of the diffusion model.

[0006] Therefore, it is necessary to conduct a more in-depth study on the existing concept erasure methods, meeting the two requirements of erasure effectiveness and model generality, and ensuring the normal use of the model is not affected while erasing the target concept. Summary of the Invention

[0007] In order to overcome the above problems, an in-depth study has been carried out, and a quality-preserving text-image joint concept erasure method for diffusion models is proposed, including the following steps:

[0008] S1. Construct a visual semantic set, in which images containing the target concept are used as samples;

[0009] S2. Project the images in the visual semantic set into the text space to obtain image embeddings;

[0010] S3. Extract visual features from the image embeddings, and in the attention layer of the diffusion model, perform attention matching using the corresponding text embeddings and visual features;

[0011] S4. Train the diffusion model;

[0012] S5. Use the trained diffusion model to generate images according to text prompts.

[0013] In a preferred embodiment, in S1, in addition to the real visual semantic set, a supplementary visual semantic set is also constructed through the diffusion model.

[0014] In a preferred embodiment, by giving a target concept, a text prompt is constructed using a template, and an image containing the target concept is generated through the original diffusion model as a supplementary visual semantic set.

[0015] In a preferred embodiment, S2 includes the following sub-steps:

[0016] S21: Encode the text description corresponding to the target concept in the visual semantic set;

[0017] S22: Encode the images in the visual semantic set;

[0018] S23: Use an image adapter to project the image encoding into the same space as the text description encoding.

[0019] In a preferred embodiment, in S3, the attention layer is a cross-attention layer of a text attention layer and an image attention layer

[0020] In a preferred embodiment, S3 includes the following sub-steps:

[0021] S31: Extract text-guided visual features from the image embedding based on the text description encoding;

[0022] S32: Perform text attention matching with the text description encoding as the guiding information and perform image attention matching with the visual features as the guiding information.

[0023] In a preferred embodiment, the attention layer query matrix is shared during the text attention matching and image attention matching processes.

[0024] In a preferred embodiment, in S4, the loss function during the training process is set to:

[0025]

[0026] where, || ||2 represents the L2 norm, t represents the time step of the current denoising process, x t represents the latent space variable at the t-th time step, c is the target concept condition, represents the noise prediction result of the trained model at the time step t under the condition c, ∈ θ (x t , t) represents the noise prediction result of the model before training at the time step t without conditions, and η is the erasure strength.

[0027] The present invention also provides an electronic device, including:

[0028] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the above methods.

[0029] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above.

[0030] The beneficial effects of the present invention include:

[0031] (1) Using text-image pairs to jointly represent the target concept, so as to more completely and accurately erase the target concept;

[0032] (2) Ensuring that non-target concepts appearing in the image are not erased, and maintaining the generation quality of normal concepts. Description of the Drawings

[0033] Figure 1 It is a schematic flow diagram of a quality-preserving text-image joint concept erasing method for a diffusion model according to a preferred embodiment of the present invention. Detailed Embodiments

[0034] The present invention will be further described in detail below with reference to the drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more definite.

[0035] The special term "exemplary" here means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, unless otherwise specified, the drawings do not have to be drawn to scale.

[0036] According to a quality-preserving text-image joint concept erasing method for a diffusion model provided by the present invention, as Figure 1 shown, it includes the following steps:

[0037] S1. Construct a visual semantic set, in which an image containing the target concept is used as a sample;

[0038] S2. Project the images in the visual semantic set into the text space to obtain image embeddings;

[0039] S3. Extract visual features from the image embeddings, and in the attention layer of the diffusion model, perform attention matching between the corresponding text embeddings and the visual features;

[0040] S4. Train the diffusion model;

[0041] S5. Use the trained diffusion model to generate an image according to the text prompt.

[0042] In the present invention, there is no limitation on the construction method of the supplementary visual semantic set, and those skilled in the art can freely establish it according to actual needs.

[0043] Preferably, in addition to the real visual semantic set, a supplementary visual semantic set is constructed through the diffusion model.

[0044] In the existing erasure methods, users can optimize in the text embedding space through the prompt learning method to find the text that can still induce the model to generate the target concept, thereby destroying the effectiveness of the existing methods. In the present invention, through the diffusion model, all the concept texts that may be generated by prompt learning are incorporated into the supplementary visual semantic set, so as to effectively supplement the real visual semantic set, thereby improving the erasure performance.

[0045] Preferably, in S1, by giving the target concept, use a template to construct a text prompt, and generate a picture containing the target concept through the diffusion model as the supplementary visual semantic set.

[0046] The diffusion model refers to any existing diffusion model, such as the classic text-to-image model StableDiffusion.

[0047] More preferably, in S1, it includes the following steps:

[0048] S11. According to the target concept to be erased, find its corresponding descriptive words, for example:

[0049] Church – “church”, fish – “tench”, nudity – “nudity”, Van Gogh – “Van Gogh”.

[0050] S12. Obtain a picture of the target concept;

[0051] Preferably, set a text prompt for generating a picture related to the target concept, for example:

[0052] A photo of a church – “a photo of the church”,

[0053] A drawing in the style of Van Gogh – “a drawing by Van Gogh”.

[0054] Input the text prompt into the diffusion model, and the diffusion model generates a picture.

[0055] S13. Detect the generated picture, and retain the image related to the target concept as the supplementary visual semantic completion concept representation.

[0056] Preferably, for the generated image, use the CLIP classifier to determine whether the image is relevant to the target concept. If the prediction score of the CLIP classifier exceeds the threshold, it is considered that the generated image is relevant to the target concept; otherwise, the image is deleted.

[0057] In traditional erasure methods, only text words are used to describe the target concept to be erased. By reducing the generation probability of the target concept or anchoring the target concept to other benign concepts, the model is made to avoid generating corresponding content when faced with text descriptions related to the target concept.

[0058] According to the present invention, in S2, by projecting the visual features into the same space as the text features, the text and the image can simultaneously guide the generation process, thereby improving the erasure performance.

[0059] Preferably, S2 includes the following sub-steps:

[0060] S21. Encode the text description corresponding to the target concept in the visual semantic set;

[0061] Preferably, use the text encoder of CLIP to encode the text description corresponding to the target concept, denoted as c text 。

[0062] S22. Encode the images in the visual semantic set;

[0063] Preferably, use the image encoder of CLIP to encode the images, denoted as

[0064] S23. Use an image adapter to project the encoded image into the same space as the encoded text description, and denote the projected image embedding as

[0065] Any existing image adapter can be used for the image adapter, which is not limited in the present invention.

[0066] In the present invention, in S3, the attention layer is a cross-attention layer of a text attention layer and an image attention layer. Through S3, the model focuses on the visual content related to the target concept, avoids affecting non-target concepts, and ensures that non-target concepts appearing in the image are not erased.

[0067] Preferably, S3 includes the following sub-steps:

[0068] S31. Extract text-guided visual features from the image embedding based on the encoded text description, denoted as:

[0069]

[0070] Among them, represents the visual feature, Softmax() is the Softmax function, and d r is the dimension of the text description encoding, is the query matrix, and K r is the key matrix, and V r is the value matrix, and K r = V r = c text , that is, the image embedding is used as the query matrix, and the text description encoding is used as the key matrix and the value matrix.

[0071] S32. Use the text description encoding as the guiding information for text attention matching, and use the visual feature as the guiding information for image attention matching.

[0072] Preferably, the attention layer query matrix is shared during the text attention matching and the image attention matching processes.

[0073] Through this step, through the attention operation of the text feature and the visual feature, the part in the visual feature with a high semantic similarity to the target concept is obtained. This mechanism can enable the model to focus on the visual content related to the target concept and avoid affecting the non-target concepts.

[0074] Furthermore, text embedding is used for text attention matching, which is expressed as:

[0075]

[0076] Among them, Z′ t represents the final hidden variable obtained by the processing of the text attention layer. The text attention layer query matrix is Q = z t W q , the text attention layer key matrix is K = z t W k , the text attention layer value matrix is V = z t W v , z t is the hidden space variable of the diffusion model at time step t, and W q , W k , W v are the weight matrices of the text attention module.

[0077] Use the visual feature for image attention matching, which is expressed as:

[0078]

[0079] Among them, Z″ t represents the final hidden variable obtained by the processing of the image attention layer. The image attention layer query matrix is Q = z t Wq , the key matrix of the image attention layer is K′ = z t W′ k , the key matrix of the image attention layer is V = z t W′ v , W q , W′ k , W′ v is the weight matrix of the image attention module.

[0080] According to the present invention, the latent variable Z finally obtained after the attention layer processes is t as follows:

[0081] Z t = Z′ t + Z″ t

[0082] In the present invention, through the above process, text and visual interactions are carried out between the attention layer and the latent variable of the diffusion model to guide the conditional generation process, thereby achieving a more complete concept representation.

[0083] S4 includes the following sub-steps:

[0084] S41. Set the form of the diffusion model probability suppression parameter;

[0085] S42. Reparameterize the probability suppression parameter and obtain the loss function based on the reparameterized result;

[0086] S43. Use the original diffusion model parameters as the initial parameters and update the parameters through the obtained loss function.

[0087] In S41, set the diffusion model probability suppression parameter θ * to satisfy:

[0088]

[0089] where x represents the generation result, P θ (x) represents the result distribution generated by the model before training, represents the result distribution generated by the model after training, P(c|x) represents the probability that the result generated by the model before training belongs to category c, c is the target concept condition, η is the erasure intensity.

[0090] Furthermore, according to Bayes' theorem, there is:

[0091]

[0092] Then the diffusion model probability suppression parameter θ * satisfies:

[0093]

[0094] In S42, the reparameterization is expressed as:

[0095]

[0096] where t represents the time step of the current denoising process, and x t represents the latent space variable at the t-th time step, represents the noise prediction result of the trained model at time step t under the condition c, and ∈ θ (x t , t) represents the noise prediction result of the model before training at time step t unconditionally.

[0097] The loss function is expressed as:

[0098]

[0099] where ||||2 represents the L2 norm.

[0100] In S5, the usage method of the trained diffusion model is the same as that of the original diffusion model. Since the parameter θ * has erased the content related to the target concept during the training process, the model will no longer generate results related to c at this time.

[0101] In the present invention, various embodiments of the methods described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.

[0103] Embodiment

[0104] Example 1

[0105] An erasure method experiment was conducted using a certain dataset, including the following steps:

[0106] S1. Construct a visual semantic set, in which images containing the target concept are used as samples;

[0107] S2. Project the images in the visual semantic set into the text space to obtain image embeddings;

[0108] S3. Extract visual features from the image embeddings, and in the attention layer of the diffusion model, perform attention matching between the corresponding text embeddings and visual features;

[0109] S4. Train the diffusion model;

[0110] S5. Use the trained diffusion model to generate images according to text prompts.

[0111] In S1, the following steps are included:

[0112] S11. According to the target concept to be erased, find its corresponding descriptive words;

[0113] S12. Obtain pictures of the target concept, where the diffusion model uses Stable Diffusion;

[0114] S13. Detect the generated pictures, and retain the images related to the target concept as supplementary visual semantic completion concept representations.

[0115] In S2, the following sub-steps are included:

[0116] S21. Encode the text description corresponding to the target concept in the visual semantic set; use the text encoder of CLIP to encode the text description corresponding to the target concept, denoted as c text .

[0117] S22. Encode the images in the visual semantic set; use the image encoder of CLIP to encode the images, denoted as

[0118] S23. Use an image adapter to project the image encoding into the same space as the text description encoding, and denote the projected image embedding as

[0119] In S3, the following sub-steps are included:

[0120] S31. Based on the text description encoding, extract text-guided visual features from the image embeddings, denoted as:

[0121]

[0122] S32. Perform text attention matching using the text description encoding as the guiding information, and perform image attention matching using the visual features as the guiding information.

[0123] In the process of text attention matching and image attention matching, share the attention layer query matrix, and use text embedding for text attention matching, which is expressed as:

[0124]

[0125] Use visual features for image attention matching, which is expressed as:

[0126]

[0127] The latent variable Z finally obtained after the attention layer processing t is:

[0128] Z t = Z′ t + Z″ t

[0129] S4 includes the following sub-steps:

[0130] S41. Set the form of the diffusion model probability suppression parameter to satisfy:

[0131]

[0132] S42. Reparameterize the probability suppression parameter, and obtain the loss function based on the reparameterized result, which is expressed as:

[0133]

[0134] S43. Use the original diffusion model parameters as the initial parameters, and update the parameters through the obtained loss function.

[0135] Comparative Example 1

[0136] Perform the same experiment as in Example 1, except that the FMN method, UCE method, ESD method, SH method, SalUn method, and AdvUnlearn method are used respectively.

[0137] Among them, the FMN method can be found in the literature [9] Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. Forget-me-not: Learning to forget in text-to-image diffusion models. In CVPR Workshops, pages 1755–1764, 2024.

[0138] The UCE method can be found in the literature [5] Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzynska, and David Bau. Unified concept editing in diffusion models. In WACV, pages 5111–5120, 2024.

[0139] The ESD method can be found in the literature [4] Rohit Gandikota, Joanna Materzynska, Jaden Fiotto Kaufman, and David Bau. Erasing concepts from diffusion models. In ICCV, pages 2426–2436, 2023.

[0140] The SH method can be found in the literature

[11] Jing Wu and Mehrtash Harandi. Scissorhands: Scrub data influence via connection sensitivity in networks. In ECCV, 2024.

[0141] The SalUn method can be found in the literature

[13] Chongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong, Dennis Wei, and Sijia Liu. Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation. In ICLR, 2024.

[0142] The AdvUnlearn method can be found in the literature

[10] Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, and Sijia Liu. Defensive unlearning with adversarial training for robust concept erasure in diffusion models. NeurIPS, 2024.

[0143] Concepts "nude", "Van Gogh", and "parachute" were erased respectively. The commonly used evaluation metrics pre-ASR, ASR, FID, and CLIP were used to evaluate the erasure effects in Example 1 and Comparative Example 1. The results are shown in Tables 1 - 3.

[0144] Table 1 Performance comparison of erasing "nude"

[0145]

[0146] Table 2 Performance comparison of erasing "Van Gogh"

[0147]

[0148] Table 3 Performance comparison of erasing "parachute"

[0149]

[0150] It can be seen from Tables 1 - 3 that under the four evaluation metrics, the method in Example 1 can achieve good performance compared with various methods in Comparative Example 1, and its comprehensive performance is especially better than various existing methods.

[0151] The present invention has been described in combination with preferred embodiments above, but these embodiments are only exemplary and only serve an illustrative purpose. On this basis, various substitutions and improvements can be made to the present invention, and all of these fall within the protection scope of the present invention.

Claims

1. A quality-preserving text image joint concept erasing method for diffusion model, characterized in that: The following steps are involved: S1, constructing a visual semantic set, in which images containing target concepts are used as samples; S2, project the image in the visual semantics set into the text space to obtain image embedding; S3, extract visual features from image embedding, and use the corresponding text embedding to perform attention matching with the visual features in the attention layer of the diffusion model; S4, training the diffusion model; S5. Generate an image based on the text prompt words using the trained diffusion model.

2. The quality-preserving text image joint concept erasing method for diffusion model according to claim 1 is characterized in that: In S1, in addition to the real visual semantic set, a supplementary visual semantic set is constructed through the diffusion model.

3. The quality-preserving text image joint concept erasing method for diffusion model according to claim 2 is characterized in that: By giving a target concept, a template is used to construct text prompt words, and a diffusion model is used to generate pictures containing the target concept as a supplementary visual semantic set.

4. The quality-preserving text image joint concept erasing method for diffusion model according to claim 1 is characterized in that: S2 includes the following sub-steps: S21, encoding the text description corresponding to the target concept in the visual semantics set; S22, encode the image in the visual semantic set; S23. Use an image adapter to project the image code into the same space as the text description code.

5. The quality-preserving text image joint concept erasing method for diffusion model according to claim 1 is characterized in that: In S3, the attention layer is a cross attention layer of a text attention layer and an image attention layer.

6. The quality-preserving text image joint concept erasing method for diffusion model according to claim 1 is characterized in that: S3 includes the following sub-steps: S31, extracting text-guided visual features from image embeddings based on text description encoding; S32. Use the text description encoding as guiding information to perform text attention matching, and use the visual features as guiding information to perform image attention matching.

7. The quality-preserving text image joint concept erasing method for diffusion model according to claim 6 is characterized in that: The attention layer query matrix is ​​shared in the text attention matching and image attention matching processes.

8. The quality-preserving text image joint concept erasing method for diffusion model according to claim 1 is characterized in that: In S4, the loss function during training Set to: Among them, ||||2 represents the L2 norm, t represents the time step of the current denoising process, and x t represents the latent space variable at the tth time step, c is the target concept condition, represents the noise prediction result of the trained model under condition c at time step t, ∈ θ (x t , t) represents the noise prediction result of the pre-training model at time step t without any condition, and η is the erasing strength.

9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.