Text-to-Image Diffusion Model Concept Erasure Method, System, Device and Medium
By fine-tuning training of the text-to-image diffusion model, combining concept localization and self-attention distortion loss function, the target concept in the model was successfully erased, solving the problem that the deep visual concept cannot be effectively removed in the existing technology, and achieving the safety and reliability of the model.
Patent Information
- Application Number
- CN202510331508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Existing text-to-image diffusion models may contain bad or sensitive content that is not suitable for public dissemination during the generation process. The existing concept erasure method cannot effectively deal with the deep visual concepts that the model has learned, and rely on specific prompt words to truly erase visual concepts from the model.
Through fine-tuning of the model, the target concept is erased from the model. The specific methods include obtaining text to the image diffusion model and the target concept image collection, noise processing on the target concept image and inputting it into the model, adjusting the target concept generation probability to construct the concept localized erase loss function, and fine-tuning the model for the model combining the self-attention distortion loss function to obtain the model for erasing the target concept.
The real forgetting of the target concept is achieved, the probability of generating bad content is reduced, the robustness and accuracy of concept erasing is improved, and the safety and reliability of the model is ensured.
Smart Images

Figure CN119850773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of concept erasure for text-to-image diffusion models, and particularly to a method, system, device and medium for concept erasure of text-to-image diffusion models. Background Art
[0002] In recent years, generative models, especially text-to-image diffusion model technology, have developed rapidly. Due to their ability to generate high-quality images, such models have been widely used in many fields such as art creation, content generation, and advertising design. These models can generate high-quality and detailed images through natural language prompts.
[0003] However, since these models are usually trained based on large-scale datasets, the quality and source of the training data are often difficult to control, resulting in the possibility that the content generated by the models may include inappropriate or sensitive content that is not suitable for public dissemination (collectively referred to as inappropriate content hereinafter).
[0004] Some existing methods achieve the prevention and control of inappropriate content by filtering training data. However, the filtering of inappropriate data is limited. Especially when faced with large-scale unprocessed network data, it is difficult to ensure that all inappropriate content is eliminated. In addition, data filtering methods cannot effectively deal with the deep visual concepts that the models have learned. Therefore, some researchers have proposed using a simpler and more efficient concept erasure technology to erase or edit the learned concepts by fine-tuning the pre-trained models. However, existing concept erasure methods often rely on specific prompts and can only erase the visual content corresponding to specific prompts, rather than truly erasing the visual concepts from the models. If users use adversarial or paraphrased prompts, the models may still be induced to generate inappropriate content.
[0005] In view of this, the present invention is specifically proposed. Summary of the Invention
[0006] The object of the present invention is to provide a method, system, device and medium for concept erasure of text-to-image diffusion models, which erase the target concept from the model by means of model fine-tuning to obtain a safer and more reliable diffusion model.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] A method for concept erasure of text-to-image diffusion models includes:
[0009] Obtain a given text-to-image diffusion model and a set of target concept images;
[0010] Add noise to the set of target concept images and input it into the text-to-image diffusion model; obtain the objective function by adjusting the target concept generation probability, and localize the corresponding target concept image by generating a target concept segmentation mask, and construct a concept localization erasure loss function in combination with the objective function; construct a self-attention distortion loss function in combination with the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; fine-tune and train the text-to-image diffusion model in combination with the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model that erases the target concept.
[0011] A text-to-image diffusion model concept erasure system for implementing the foregoing method, the system comprising:
[0012] A model and image set acquisition unit for acquiring a given text-to-image diffusion model and a set of target concept images;
[0013] A fine-tuning training unit for adding noise to the set of target concept images and inputting it into the text-to-image diffusion model; obtaining the objective function by adjusting the target concept generation probability, and localizing the corresponding target concept image by generating a target concept segmentation mask, and constructing a concept localization erasure loss function in combination with the objective function; constructing a self-attention distortion loss function in combination with the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; fine-tuning and training the text-to-image diffusion model in combination with the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model that erases the target concept.
[0014] A processing device comprising: one or more processors; a memory for storing one or more programs;
[0015] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.
[0016] A readable storage medium storing a computer program that implements the foregoing method when executed by a processor.
[0017] As can be seen from the technical solutions provided by the present invention above, on the one hand, from the perspective of the generation probability of concepts, the generation probability of the target concept is reduced as a whole, truly realizing the forgetting of the target concept. At the same time, in order to avoid the influence of irrelevant concepts in the image, a concept localization method is proposed, introducing a target concept segmentation mask to prevent the introduction of irrelevant visual information, ensuring the accuracy and efficiency of concept erasure; on the other hand, starting from the mechanism of the model itself, by disturbing the self-attention response of the model, the generation of the target concept is further disrupted at the model level, improving the robustness of concept erasure. Generally speaking, the present invention considers from both the probability distribution level and the model level, truly realizing robust and efficient concept erasure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a method for concept erasure of a text-to-image diffusion model provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic diagram of the overall framework of a method for concept erasure of a text-to-image diffusion model provided by an embodiment of the present invention;
[0021] Figure 3 It is a schematic diagram of the visualization comparison result of the erasure effect between the method of the present invention and the existing concept erasure method provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of a system for concept erasure of a text-to-image diffusion model provided by an embodiment of the present invention;
[0023] Figure 5 It is a schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0025] First, the following explanations are given for the terms that may be used in this article:
[0026] Descriptions with semantic meanings such as "comprising", "including", "containing", "having" or other similar ones shall be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be construed as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.
[0027] The term "consisting of" means excluding any technical feature element not explicitly listed. If this term is used in a claim, this term will make the claim a closed type, so that it does not include technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements explicitly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.
[0028] The following provides a detailed description of a method, system, device and medium for concept erasure of a text-to-image diffusion model provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. For those conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the instruments used in the embodiments of the present invention without indicating the manufacturer, they are all conventional products that can be obtained through commercial purchase.
[0029] Embodiment 1
[0030] A method for concept erasure of a text-to-image diffusion model according to an embodiment of the present invention, as Figure 1 shown, mainly includes the following steps:
[0031] Step 1, obtain a given text-to-image diffusion model and a set of target concept images.
[0032] In the embodiment of the present invention, the given text-to-image diffusion model is an existing model, specifically a text-to-image diffusion model that needs to erase the target concept, and the set of target concept images also belongs to an existing image data set. The above two types of information (i.e., the model and the image set) can be obtained through conventional means.
[0033] Step 2, construct a loss function considering from the probability distribution level and the model level, and perform fine-tuning training.
[0034] In the embodiments of the present invention, the target concept image set is subjected to noise addition processing and input into the text-to-image diffusion model; the target function is obtained by adjusting the target concept generation probability, and the corresponding target concept image is localized by generating a target concept segmentation mask, and the concept localization erasure loss function is constructed in combination with the target function; the self-attention distortion loss function is constructed in combination with the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; the text-to-image diffusion model is fine-tuned and trained in combination with the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model that erases the target concept.
[0035] The preferred implementation of this step is as follows:
[0036] (1) Construct the concept localization erasure loss function.
[0037] (1.1) Obtain the target function by adjusting the target concept generation probability.
[0038] Denote the target concept image set as , where is a single target concept image, i = 1, 2, …, n, and n is the number of target concept images; after adding noise to the target concept image set, a single noise-added target concept image is denoted as , and based on the variational lower bound, we get:
[0039] ;
[0040] where represents the mathematical expectation of adding noise to the target concept image at time step t, is the noise predicted by the text-to-image diffusion model (abbreviated as the predicted noise), is the proportional symbol, exp is the exponential function with the natural constant e as the base, is the probability of the text-to-image diffusion model generating the target concept, is the noise, and t is the time step of the noise addition process.
[0041] By minimizing the probability of generating the target concept, the target function is determined as:
[0042] ;
[0043] where is the target function.
[0044] (1.2) Localize the corresponding target concept image by generating a target concept segmentation mask, and the steps are as follows: Introduce the Segment Anything model; for each target concept image, use the Segment Anything model for segmentation to obtain the target concept segmentation mask; through the target concept segmentation mask, eliminate the visual information irrelevant to the target concept.
[0045] (1.3) Construct a concept localization erasure loss function.
[0046] The concept localization erasure loss function is expressed as:
[0047] ;
[0048] where, is the concept localization erasure loss function, M is the target concept segmentation mask, is the symbol of Hadamard product, representing the operation of multiplying the corresponding elements of vectors.
[0049] (2) Construct a self-attention distortion loss function.
[0050] (2.1) Aggregate all self-attention maps, and through singular value decomposition, select the vector corresponding to the largest singular value.
[0051] Specifically: Denote the number of attention maps as K, upsample each self-attention map to a set dimension, and then aggregate them to obtain the aggregated self-attention map ; Through singular value decomposition, obtain:
[0052] ;
[0053] where, U and V are the unitary matrices obtained by decomposition; S is a diagonal matrix, SVD represents singular value decomposition, is 's mean matrix, select the vector corresponding to the largest singular value, that is, the diagonal elements of the diagonal matrix are singular values, and the row where the largest singular value is located is denoted as Y, and the row vector of the Y-th row in the unitary matrix V is vector.
[0054] (2.2) Use the vector corresponding to the largest singular value to construct a self-attention distortion loss function.
[0055] (3) Construct the final loss function and perform fine-tuning training.
[0056] (3.1) Combine the concept localization erasure loss function and the self-attention distortion loss function to construct the final loss function L, which is expressed as:
[0057] ;
[0058] Among them, and correspondingly represent the concept localization erasure loss function and the self-attention distortion loss function, and are two weight coefficients.
[0059] (3.2) Use the final loss function L to fine-tune and train the text-to-image diffusion model.
[0060] The present invention can erase the target concept (i.e., the specific concept to be erased) in the text-to-image diffusion model, so as to be used in scenarios such as privacy protection and model security prevention and control. In implementation, it can be installed on devices such as computers and mobile phones in the form of software to provide real-time target concept erasure; it can also be installed on a server to provide a large number of target concept erasures.
[0061] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the method provided by the embodiments of the present invention.
[0062] I. Overall overview of the solution.
[0063] The solution provided by the embodiments of the present invention can erase the high-probability generation area in the text-to-image diffusion model. Different from the existing erasure method based on prompts, the present invention does not rely on specific prompts for optimization, but considers from both the probability distribution level and the model level to achieve the erasure effect. The solution of the present invention includes: generation probability adjustment, concept localization, and self-attention distortion. Among them, the generation probability adjustment derives the objective function through the variational lower bound (ELBO) to maximize the generation error of the model for the target concept. The concept localization performs localization processing on the concept image by generating a target concept segmentation mask, so that the model only focuses on the target concept area during the fine-tuning process. The self-attention distortion disrupts the self-attention response of the model and destroys the model's ability to understand the structure and shape of the target concept. Thanks to the above improvements, the present invention significantly improves the performance of concept erasure and reaches an advanced level on the evaluation dataset.
[0064] II. Detailed introduction of the solution.
[0065] As Figure 2 shown, the solution of the present invention is mainly divided into generation probability adjustment, concept localization, and self-attention distortion. The following will introduce each part above separately.
[0066] 1. Generation probability adjustment.
[0067] Given the set of target concept images and the pre-trained text-to-image diffusion model , first add noise to each target concept image to obtain a noisy image From the derivation of the evidence lower bound (ELBO), the following conclusions can be drawn:
[0068] ;
[0069] where denotes the mathematical expectation of adding noise to the target concept image at time step t, is the proportionality symbol, exp is the exponential function with the natural constant e as the base, is the model to generate the probability of the target visual concept, is the added noise, and t is the time step of adding noise. Therefore, the optimization objective of concept erasure can be defined as minimizing the generation probability of the target visual concept, that is:
[0070] ;
[0071] From the above derivation of the evidence lower bound, the objective function can be defined as:
[0072] .
[0073] The above objective function defines concept erasure from the perspective of probability distribution, avoiding the introduction of specific prompt words and achieving true visual concept erasure.
[0074] 2. Concept localization.
[0075] To achieve more accurate concept erasure and avoid the influence of irrelevant concept information in the image, the present invention proposes concept localization. Given any image, the Segment Anything Model (SAM) can segment different instances. Therefore, use SAM to segment the target concept image and select the mask M of the target concept (target concept segmentation mask). During the concept erasure process, omit the visual information irrelevant to the target concept. Combining the aforementioned objective function, the concept localization erasure loss function can be defined as:
[0076] .
[0077] 3. Self-attention distortion.
[0078] To further erase the target concept at the model level, the present invention proposes self-attention distortion. In the text-to-image diffusion model, the self-attention mechanism reflects the model's understanding of visual concept shapes, structures, etc. Concept erasure effects at the model level can be achieved by disrupting the model's self-attention.
[0079] In the diffusion model, K self-attention maps are generated, which contain multiple different resolutions. For example, the U-Net in the text-to-image diffusion model generates 16 self-attention maps, including multiple resolutions, such as: 64×64×64×64, 32×32×32×32, 16×16×16×16, 8×8×8×8.
[0080] In the embodiments of the present invention, by upsampling the set dimensions (e.g., the last two dimensions) of all self-attention maps to a specified resolution (e.g., the highest resolution), and then aggregating, an aggregation result is obtained. 。
[0081] For the th self-attention map The processing is expressed as:
[0082] ;
[0083] Among them, Upsampling is the upsampling operation. For example, it can be implemented through the Bilinear interpolation operation. is the result after performing the upsampling operation on the th self-attention map ; is the symbol of the set of real numbers, h and w are the height and width of the self-attention map respectively, and 64×64 is an example of the resolution.
[0084] The aggregation process is expressed as:
[0085] ;
[0086] Among them, represents the first two dimensions, , represents the reference position during the calculation of the self-attention map, and the subsequent colon means taking all elements. The above formula shows the solution process of each value of the first two dimensions, and the comprehensive result is the aggregation result ; is , W is an element in the highest resolution. For example, W = 64.
[0087] Performing singular value decomposition on results in:
[0088] ;
[0089] Among them, is The mean matrix, where U and V are unitary matrices obtained by decomposition; S is a diagonal matrix, and the elements on the diagonal of the diagonal matrix are singular values. The row where the largest singular value is located is denoted as Y, and the row vector of the Y-th row in V is taken as , and a self-attention distortion loss function is introduced as follows:
[0090] ;
[0091] where, where is the mean vector. Through this self-attention distortion loss function, the erasing effect of visual concepts can be achieved at the model level, enabling the model to forget information such as the shape and structure of the target concept.
[0092] Combining the above three parts, the final loss function for fine-tuning the model during concept erasure is written as:
[0093] ;
[0094] Based on the above loss function L, the model is fine-tuned and trained; taking Stable Diffusion v1.4 (abbreviated as SDv1.4) as the benchmark model, three to five collected target concept images are used as the fine-tuning dataset (i.e., the target concept image set). During the fine-tuning process, is set to 1, is set to 500, the Adam (Adaptive Moment Estimation) optimizer is used, the learning rate is set to 0.0001, and the total number of fine-tuning rounds is set to 10 rounds. The CLIP score is used as the evaluation index for the concept erasure effect, and the smaller its value, the better the erasure effect. Considering that the relevant fine-tuning training process can be achieved through conventional methods, it will not be elaborated here.
[0095] Among them, Stable Diffusion v1.4 is the 1.4 version of the stable diffusion model, and CLIP is Contrastive Language-Image Pretraining, which is a multi-modal model. Both are conventional models.
[0096] III. Effect description.
[0097] Existing concept erasure methods often rely on specific prompts. For example, to erase the visual concept of "Mickey Mouse", existing methods can only prevent the generation of this concept when the user gives the prompt "Mickey Mouse". However, when the user gives a paraphrased text as the prompt, such as "Disney classic character", the visual concept of "Mickey Mouse" will still be generated, resulting in the failure of the erasure method. This is because existing methods only change the mapping relationship from text to image in the generation model, but do not truly erase the model's generation ability for the target visual concept. Therefore, existing concept erasure methods have poor robustness and can be easily bypassed by adversarial prompts or paraphrased prompts.
[0098] In view of the above problems, the above solution provided by the embodiments of the present invention starts from the perspective of the concept generation probability, defines the goal of concept erasure as minimizing the generation probability, provides a theoretical basis for the entire erasure process, is more interpretable, and based on the variational lower bound, derives a concept erasure loss function based on the mean square error of the predicted noise, integrally reducing the generation probability of the target concept and truly achieving the forgetting of the target concept. To avoid the influence of irrelevant concepts in the image, a concept localization method is proposed. By introducing a mask of the target concept in the fine-tuning stage, the introduction of irrelevant visual information is prevented, ensuring the accuracy and efficiency of concept erasure. In addition, starting from the mechanism of the diffusion model itself, the present invention points out the relationship between the self-attention layer and the model's visual understanding ability, and proposes a self-attention distortion loss function, further destroying the generation of the target concept at the model level and improving the robustness of concept erasure.
[0099] Thanks to the above improvements, the present invention has significantly improved the performance of concept erasure and reached an advanced level on the evaluation dataset. Specifically: The evaluation dataset used selects and collects paraphrased prompt words corresponding to three concepts, namely "Pikachu", "Mickey", and "Snoopy", with 15 paraphrased prompt words for each concept. The CLIP score of the present invention on this dataset is 61.91, achieving an improvement of 1.59 compared to the existing best method UCE. As Figure 3 shown, a visual comparison of the erasure effects between the method of the present invention and existing concept erasure methods is provided, indicating that the present method shows better robustness when dealing with paraphrased prompt words and the erasure effect is more thorough; Figure 3 Among them, ESD, FMN, and UCE are all existing concept erasure schemes. ESD (Erased Stable Diffusion) is stable diffusion erasure, FMN (Forget-Me-Not) is a forgetting method for diffusion models, and UCE (Unified Concept Editing) is unified concept editing. It should be noted that the text-to-image diffusion model involved in the present invention can be applicable to multiple languages. Figure 3 Only an example in the English language is provided. In actual applications, users can input text in the same language according to their needs.
[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented through software, or can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0101] Embodiment 2
[0102] The present invention also provides a text-to-image diffusion model concept erasure system, which is mainly used to implement the method provided in the foregoing embodiments, such as Figure 4 shown, the system mainly includes:
[0103] A model and image set acquisition unit for acquiring a given text-to-image diffusion model and a target concept image set;
[0104] A fine-tuning training unit for performing noise addition processing on the target concept image set and inputting it into the text-to-image diffusion model; obtaining an objective function by adjusting the target concept generation probability, and localizing the corresponding target concept image by generating a target concept segmentation mask, and constructing a concept localization erasure loss function in combination with the objective function; constructing a self-attention distortion loss function by combining the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; fine-tuning and training the text-to-image diffusion model by combining the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model that erases the target concept.
[0105] Considering that the processing details of each part involved in this system have been introduced in detail in the previous embodiments, they will not be elaborated here.
[0106] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
[0107] Embodiment 3
[0108] The present invention also provides a processing device, such as Figure 5As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0109] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.
[0110] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:
[0111] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;
[0112] The output device can be a display terminal;
[0113] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.
[0114] Embodiment 4
[0115] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.
[0116] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.
[0117] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.
Claims
1. A text-to-image diffusion model concept erasing method, characterized in that: include: Get a given text-to-image diffusion model and a target concept image set; The target concept image set is subjected to noise processing and input into the text-to-image diffusion model; the target function is obtained by adjusting the target concept generation probability, and the corresponding target concept image is localized by generating a target concept segmentation mask, and a concept localization erasure loss function is constructed in combination with the target function; a self-attention distortion loss function is constructed in combination with the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; the text-to-image diffusion model is fine-tuned and trained in combination with the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model with the target concept erased; The self-attention distortion loss function constructed by combining the self-attention map generated by the text-to-image diffusion model includes: aggregating all the self-attention maps, and selecting the vector corresponding to the maximum singular value through singular value decomposition; using the vector corresponding to the maximum singular value to construct the self-attention distortion loss function; The step of aggregating all self-attention graphs and selecting the vector corresponding to the maximum singular value through singular value decomposition includes: The number of attention maps is recorded as K, and the dimension of each self-attention map is set for upsampling, and then aggregated to obtain the aggregated self-attention map ; Through singular value decomposition, we get: ; Among them, U and V are unitary matrices obtained by decomposition; S is a diagonal matrix, and SVD represents singular value decomposition. for The mean matrix of Vector, that is, the diagonal elements in the diagonal matrix are singular values, and the row with the largest singular value is denoted as Y. The row vector of the Yth row in the unitary matrix V is vector; Self-attention distortion loss function It is expressed as: ; in, for The mean vector of .
2. The text-to-image diffusion model concept erasing method according to claim 1, characterized in that: The objective function obtained by adjusting the target concept generation probability includes: Set of target concept images Recorded as ,in, is a single target concept image, i=1,2,…,n, n is the number of target concept images; After the target concept image set is subjected to noise processing, a single target concept image after noise processing is recorded as , based on the variational lower bound, we get: ; in, Represents the target concept image at time step t Adding Noise The mathematical expectation after is the noise predicted by the text-to-image diffusion model, is proportional to the sign, exp is an exponential function with the natural constant e as the base, Text-to-image diffusion model The probability of generating the target concept; By minimizing the probability of generating the target concept , determine the objective function as: ; in, is the objective function.
3. A text-to-image diffusion model concept erasure method according to claim 1 or 2, characterized in that: The localization processing of the corresponding target concept image by generating a target concept segmentation mask comprises: Introducing the Split Everything model; For each target concept image, the segmentation model is used to segment it and obtain the target concept segmentation mask; The target concept segmentation mask is used to remove visual information irrelevant to the target concept.
4. The text-to-image diffusion model concept erasing method according to claim 2, characterized in that: The concept localization erasure loss function is expressed as: ; in, is the concept localization erasure loss function, M is the target concept segmentation mask, is the Hadamard product symbol.
5. The text-to-image diffusion model concept erasing method according to claim 1, characterized in that: The fine-tuning training of the text-to-image diffusion model by combining the concept localization erasure loss function with the self-attention distortion loss function includes: Combining the concept localization erasure loss function and the self-attention distortion loss function, the final loss function L is constructed, which is expressed as: ; in, , The corresponding representation concept localization erasure loss function and self-attention distortion loss function are and are two weight coefficients; The text-to-image diffusion model is fine-tuned using the final loss function L.
6. A text-to-image diffusion model concept erasure system, characterized in that: For implementing the method described in any one of claims 1 to 5, the system comprises: A model and image set acquisition unit, used to acquire a given text-to-image diffusion model and a target concept image set; A fine-tuning training unit is used to perform noise processing on the target concept image set and input it into the text-to-image diffusion model; obtain the target function by adjusting the target concept generation probability, and localize the corresponding target concept image by generating a target concept segmentation mask, and construct a concept localization erasure loss function in combination with the target function; construct a self-attention distortion loss function in combination with the self-attention map generated by the text-to-image diffusion model to adjust the self-attention response; fine-tune the text-to-image diffusion model in combination with the concept localization erasure loss function and the self-attention distortion loss function to obtain a text-to-image diffusion model with the target concept erased.
7. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Ink erase concept scroll
CA2591514A1
Personalized text-to-image generation method and system based on diffusion model
CN118429485A