Image Generation Content Suppression Method and System Based on Text-to-Image Diffusion Model
By constructing the word embedding matrix of the target prompt word and performing singular value decomposition and optimization, the attention map is evaluated using cross attention and alignment loss, the problem of specific subject generation in the diffusion model is solved, and a more efficient image generation and editing effect is achieved.
Patent Information
- Application Number
- CN202310935657.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-07-27
AI Technical Summary
The existing diffusion model is difficult to effectively suppress the generation of specific subjects in the input prompt words in image generation, resulting in the occurrence of unwanted specific visual and semantic properties in the generated image.
By constructing the word embedding matrix of the target prompt words, perform singular value decomposition, soft-weighted regularization and inference optimization are introduced, and attention maps are evaluated using cross-attention and alignment loss, and the generation of specific subjects is suppressed.
Without the need to fine-tune the model, the ability to suppress the generation of specific subjects is significantly improved, the editing effect of image generation is enhanced, and the quality and accuracy of generated images are improved.
Smart Images

Figure CN117251589B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image generation technologies, and particularly to an image generation content suppression method and system based on a text-to-image diffusion model. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] Diffusion models have recently achieved remarkable results. However, despite their great success in the field of image generation currently, they may still be unable to suppress the generation of specific subjects in the input prompt; Stable Diffusion (SD), as a powerful generation model, in many practical applications of image generation, especially in specific contexts, needs to have the ability to suppress specific semantic information while having significant generation ability, and to remove specific context information from a given real image. However, there is a key semantic misrepresentation problem in the current model, that is, the SD model may not be able to suppress the generation of specific subjects in the input prompt. For example, when the input prompt is "a man without glasses", the entity "glasses" will still be generated in the image, as Figure 1 . Another example, given a real image, such as Figure 1 , the user expects to edit the image under the condition that the prompt is "Yoshua Bengio without beard". However, the current SD-based image editing method cannot remove the beard information, as Figure 1 . Removing a specific subject is more challenging than replacement because it requires filling reasonable content in the area where the specific subject is removed in the image. The specific visual and semantic attributes of the images generated by the SD model are determined by the input prompt. A simple strategy is to remove the target text (i.e., "glasses"). However, as Figure 1As shown, the glasses still exist. This is because many of the human images collected in the training set contain glasses, but often do not contain the "glasses" label; the current solution to such problems is to directly delete the target text embeddings from the text encoder. However, this still generates the target subject. Experiments have confirmed that the End of Text (EOT) embedding appended at the end of the prompt contains meaningful, redundant, and repetitive semantic information. Some concurrent studies fine-tune the SD model, which however leads to "catastrophic forgetting". Take an example, consider the input prompt "a man without glasses". The SD model is fine-tuned to remove "glasses". However, when given "a man with glasses", the fine-tuned SD model usually fails to generate "glasses". And it is usually not easy to design a suitable semantic constraint, and a simple implementation leads to unexpected side effects, and the output image may have additional suppression for non-target prompts. Summary of the Invention
[0004] To solve the above problems, the present disclosure proposes an image generation content suppression method and system based on a text-to-image diffusion model, constructs a word embedding matrix of a target prompt, and proposes a method for regularizing the target prompt information, and further suppresses the generation of the subject in the target prompt during inference, and determines the text embedding of specific visual attributes of the generated image.
[0005] According to some embodiments, the present disclosure adopts the following technical solutions:
[0006] An image generation content suppression method based on a text-to-image diffusion model, comprising:
[0007] Obtain the text input prompt of the given image to be generated and map it to a text embedding;
[0008] Divide the text embedding into two parts: the embedding expected to be suppressed and the embedding encouraged to be retained, construct a target text embedding matrix, perform singular value decomposition on the matrix part composed of the embedding expected to be suppressed and [EOT] in the target text embedding matrix, and extract the suppressed semantic information;
[0009] Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, output the corresponding attention maps of the features expected to be suppressed and the attention maps of the features encouraged to be retained through cross-attention, and propose two attention losses to evaluate the attention maps; introduce an alignment loss to align the attention maps of the features encouraged to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the subject expected to be suppressed, and finally generate an image after removing the entity expected to be suppressed.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions:
[0011] An image generation content suppression system based on a text-to-image diffusion model, comprising:
[0012] A data acquisition module, configured to acquire a text input target prompt for a given image to be generated and map it to a text embedding;
[0013] A suppression module, configured to divide the text embedding into two parts: an embedding expected to be suppressed and an embedding encouraged to be retained, construct a target text embedding matrix, perform singular value decomposition on the matrix part composed of the embedding expected to be suppressed and [EOT] in the target text embedding matrix, and extract the suppressed semantic information;
[0014] Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, output corresponding attention maps of features expected to be suppressed and features encouraged to be retained through cross-attention, propose two attention losses to evaluate the attention maps; introduce an alignment loss to align the attention maps of features encouraged to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the entity expected to be suppressed, and finally generate an image after removing the entity expected to be suppressed.
[0015] According to some embodiments, the present disclosure adopts the following technical solutions:
[0016] A non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the image generation content suppression method based on the text-to-image diffusion model as described above.
[0017] According to some embodiments, the present disclosure adopts the following technical solutions:
[0018] An electronic device, comprising: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the image generation content suppression method based on the text-to-image diffusion model as described above.
[0019] Compared with the prior art, the beneficial effects of the present disclosure are:
[0020] Without fine-tuning the image generator, the present disclosure proposes an image generation content suppression method based on the text-to-image diffusion model, and uses two methods: soft weighted regularization and optimization during inference. In the former, a word embedding matrix of the target prompt is constructed, and a method for regularizing the target prompt information is proposed. Optimization during inference aims to further suppress the generation of the subject in the target prompt and encourage the generation that is expected to be retained. The present disclosure removes the word embedding of the target text and can extract the corresponding target information from the [EOT] embedding. Without further fine-tuning the SD model, the ability of the SD model to generate the expected subject and suppress the unwanted subject is greatly improved. The method of the present disclosure is quantitatively and qualitatively evaluated through several experiments, its effectiveness is verified, the ability to edit real images is achieved, and its versatility is further demonstrated through the context editing task of given real images. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.
[0022] Figure 1 It is a failure example of the Stable Diffusion (SD) model of the present disclosure;
[0023] Figure 2 It is the analysis process of the text embedding of the present disclosure;
[0024] Among them, Figure 2 (a) in represents the subject of the generated representation, Figure 2 (b) in represents the restored experimental result, Figure 2 (c) in represents that the generated image has aligned semantic information,
[0025] Figure 3 It is the model construction diagram of the image generation content suppression method of the embodiment of the present disclosure;
[0026] Figure 4 It is the effect of singular value decomposition of the embodiment of the present disclosure;
[0027] Figure 5 It is the comparison between the embodiment of the present disclosure and P2P in suppressing the target prompt;
[0028] Figure 6 It is the comparison between the embodiment of the present disclosure and the baseline method;
[0029] Figure 7 It is the ablation study on soft weighted regularization and optimization during inference of the embodiment of the present disclosure;
[0030] Figure 8 It is the schematic diagram of enhancing the generation of the subject of the embodiment of the present disclosure;
[0031] Figure 9 This is a schematic diagram of a failure case of an embodiment of the present disclosure. Specific embodiments
[0032] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0033] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0035] Embodiment 1
[0036] In an embodiment of the present disclosure, an image generation content suppression method based on a text-to-image diffusion model is provided, including:
[0037] Step 1: Obtain the text input target prompt of the given image to be generated and map it to a text embedding;
[0038] Step 2: Divide the text embedding into two parts: the embedding expected to be suppressed and the embedding encouraged to be retained, construct a target text embedding matrix, perform singular value decomposition on the matrix part composed of the embedding expected to be suppressed and [EOT] in the target text embedding matrix, and extract the suppressed semantic information;
[0039] Step 3: Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, output the corresponding attention map of the feature expected to be suppressed and the attention map of the feature encouraged to be retained through cross-attention, propose two attention losses to evaluate the attention map; introduce an alignment loss to align the attention map of the feature encouraged to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the entity expected to be suppressed, and finally generate an image after removing the entity expected to be suppressed.
[0040] The goal of the present disclosure is to suppress the generation of a specific subject in the input prompt. Given an existing pre-trained SD model as the diffusion model, the purpose is to remove a specific subject that the user does not want to appear in the output image without further training. The SD model trains a denoising network based on UNet ∈θ To predict the noise ∈, the following objective function is followed:
[0041]
[0042] Among them, the encoded text embedding c is extracted by the text encoder Γ of CLIP. Given the prompt p: c = Γ(p). z t is a noise sample according to the timestamp t ∼ [1, T], and T is the number of time steps. The SD model trains the encoder E and the decoder D. Then the diffusion process is carried out in the latent space. Here, the encoder maps the image x to the latent representation z0 = E(x), and the decoder D reverses the latent representation z0 into an image The sampling process is as follows:
[0043]
[0044] Among them, α t is a scalar function.
[0045] Cross-attention is adopted in this model. The SD model interacts with the text prompt through the cross-attention layer. Given the text embedding c and the image feature representation f, the key matrix K = Ψ K , Ψ V , Ψ Q , can be generated, and the value matrix V = Ψ K (c), the value matrix V = Ψ V (c) and the query matrix Q = Ψ Q (f). Then the attention map is calculated by the following formula:
[0046]
[0047] Among them, d is the projection dimension of the key and the query.
[0048] The deterministic DDIM model
[33] is used for image inversion. This process is as follows:
[0049]
[0050] The DDIM inversion generates the latent noise of the real image, which can produce an approximate reconstruction of the real image when input into the diffusion process. In this disclosure, Null-text
[24] is used to invert the real image because the noise inverted based on DDIM under the classifier-free guidance mechanism cannot accurately reconstruct the real image.
[0051] As an embodiment, the specific implementation manner of the image generation content suppression method based on the text-to-image diffusion model of this disclosure is as follows:
[0052] First, obtain the text input target prompt for the image to be generated and map it to a text embedding;
[0053] The text encoder Γ maps the text input target prompt p to a text embedding, (In the SD model, M = 768, N = 77). This is achieved by adding a text start ([SOT]) symbol in front of the text input target prompt p and N - |p| - 1 text end ([EOT]) embedding padding symbols at the end, resulting in a total of N symbols. Define the text embedding:
[0054]
[0055] Observing the text embedding, it is found that the [EOT] embedding carries important semantic information. For example, when generating an image with the prompt "aman without glasses", it generates a subject that matches "glasses", such as Figure 2 shown in (a). When the prompt embedding vector of "glasses" is set to zero, it is observed that the model still generates "glasses" and does not suppress the generation of "glasses". Replacing all [EOT] embeddings with 0 still generates the "glasses" subject. Finally, setting both the prompt embeddings of "glasses" and [EOT] to zero successfully removes "glasses" from the generated image. The results show that the [EOT] embedding contains important information of the input prompt. However, simply setting them to zero often leads to changes in information unrelated to the suppressed content.
[0056] It is observed that the [EOT] embedding has a low-rank property, indicating that it contains redundant semantic information. For example, when generating an image with the prompt "whiteand black long coated puppy", |p| = 6, the [EOT] embedding is Construct the matrix Then perform weighted nuclear norm minimization (WNNM):
[0057]
[0058] where Ψ = UΣV T is the singular value decomposition (SVD) of Ψ, is the generalized soft threshold operator with the weighted vector w, i.e., The singular values σ0 ≥ … ≥ σ 69 and the weights satisfy 0 ≤ w0 ≤ … ≤ w 69 . The purpose of this operation is to retain large singular values as much as possible and set small singular values to zero.
[0059] Restore further Use it to replace the original c EOT to generate an image. Use to represent the rank. The results of this experiment are as shown in Figure 2 (b). Two metrics: PSNR and SSIM are used to evaluate the difference between the reconstructed image and the output image of the SD model. When is selected, it means that all [EOT] embeddings are set to zero, and the generated image has similar semantic information to the original image when using all [EOT] embeddings. When increasing the value, the generated image approaches the output image of the SD model. Visually, when , the generated image is similar to the output image of the SD model. When , acceptable metric values (PSNR = 40.288, SSIM = 0.994) can be obtained. These results indicate that [EOT] embeddings have low-rank characteristics and contain redundant semantic information.
[0060] It is observed that each [EOT] embedding is semantically aligned, that is, their semantic information has a high degree of similarity. This phenomenon is demonstrated both qualitatively and quantitatively in Figure 2 (c). For example, the input prompt is "A man with a beard wearing glasses and a beanie in blue shirt". Randomly select an [EOT] embedding to replace all the prompt word embeddings, as shown in Figure 2 (c). The generated image has aligned semantic information, as shown in Figure 2 (c), which is also proven by the distance of each [EOT] embedding ( Figure 2 (c)). Most [EOT] embeddings have small distances, indicating that they are semantically aligned.
[0061] Furthermore, based on the previous analysis, it is necessary to suppress the target information from the [EOT] embeddings and the target prompt word embeddings. To achieve this goal, two methods are introduced, called soft weighted regularization and optimization during inference. For the former, a target prompt word embedding matrix is designed, and a method is proposed to regularize the target prompt word embeddings. Optimization during inference aims to further suppress the generation of the subject corresponding to the target prompt word and encourage the generation that is expected to be retained.
[0062] Divide the text embeddings into two parts: embeddings expected to be suppressed and embeddings encouraged to be retained, construct a target text embedding matrix, perform singular value decomposition on the target text embedding matrix, and extract the suppressed semantic information;
[0063] Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, and output the corresponding attention map of the feature to be suppressed and the attention map of the feature to be retained through cross-attention, and propose two attention losses to evaluate the attention map; introduce an alignment loss to align the attention map of the feature to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the subject to be suppressed, and finally generate an image after removing the entity to be suppressed.
[0064] Specifically, construct the target text embedding matrix, including: the text embedding of the text encoder is Divide into two parts of text embeddings: c SE and c PE ; c SE is the embedding to be suppressed, and c PE is the embedding to be retained. Therefore, there is:
[0065]
[0066] Construct the target text embedding matrix as The target text embedding matrix contains the information to be suppressed.
[0067] Perform singular value decomposition on the target text embedding matrix. When performing singular value decomposition, χ is obtained to guide the extraction of the semantic information to be suppressed from the embedding, including:
[0068]
[0069] where Singular value n0 = min(M, N - |p| - 1). U is the left singular vector matrix; Σ is the singular value matrix; V is the right singular vector matrix;
[0070] Intuitively, the embedding matrix mainly contains the information to be suppressed, because c EOT has redundant information ( Figure 2 (b)(c)). When performing SVD, χ is obtained to guide the extraction of the semantic information to be suppressed from the embedding c EOT . In particular, assume that the main singular values correspond to the information to be suppressed. To suppress the target information, introduce soft weighted regularization for each singular value:
[0071]
[0072] Then restore the word embedding matrix Here Note that the restored result is And Among them, σ is the original singular value. Since σ is arranged from large to small, the larger σ corresponds to the smaller e -σ , e -σ ·σ means suppressing the larger singular values mainly;
[0073] Consider a special case, that is, setting the largest K or the smallest K singular values to zero. When setting the largest K (here, K = 2) singular values to zero, the target prompt (such as glasses or a beard) can be removed. When setting the smallest K singular values to zero (here, K = 70), the target prompt information is retained. This shows that the main singular values correspond to the target information to be suppressed.
[0074] Furthermore, during inference optimization, for a specific time step t, in the order of the diffusion process T→1, the output of the diffusion network is obtained: and the corresponding attention map: where the attention map corresponds to c PE , while corresponds to the c that is expected to be suppressed SE . After soft weighted regularization, there is a new text embedding Similarly, the attention map can be obtained: Since the goal is to further suppress the generation of the main body of the target prompt word and encourage the retention of the expected generation. The present disclosure proposes two attention losses to evaluate the attention map and modify the text embedding to guide the attention map to focus on the specific region corresponding to the prompt word that wants to be retained. The alignment loss is introduced:
[0075]
[0076] That is to say, this loss attempts to align the attention map of the text prompt word that is expected to be retained at the time step. To further suppress the generation of the target main body that is expected to be removed, the diversity loss is proposed:
[0077]
[0078] Then the complete objective function of the model is:
[0079]
[0080] Among them, λ1 and λ2 are balance parameters, λ1 = 1, λ2 = 0.5, which are used to balance the effects of retention and suppression.
[0081] Use this loss to update the text embedding
[0082] For real - image editing, first, use Null - Text to perfectly invert the given real image. Because when performing content suppression on the images generated by the SD model, the latent representation that conforms to the Gaussian distribution and the text prompt are input into the SD model; while when performing content suppression on real images, the real image needs to be perfectly inverted first, and then content suppression is performed through the SD; then use the proposed method to suppress the generation of the subject in the input prompt.
[0083] Experimental settings
[0084] Training datasets and details. For image generation, randomly select 100 captions provided in the COCO validation set as input to the SD model. For real - image editing, randomly select 100 images and their corresponding text prompts from Unsplash(https: / / unsplash.com / ) and the COCO dataset.
[0085] Evaluation metrics. CLIPscore is a metric for evaluating the semantic similarity between the prompt and the edited image. The Fréchet Inception Distance (FID) is also used for evaluation. To evaluate the suppression degree of the target prompt information after editing, a reverse FID (IFID) is used, which measures the similarity between two sets. In this metric, the larger the better.
[0086] Baseline methods. P2P introduces an attention re - weighting, which scales the attention map by artificially specifying parameters to enhance or weaken the influence degree of the target prompt. This disclosure is compared with the baseline methods of the same period. Concept - ablation introduces an anchor concept, which can match the desired generated image. As shown in Table 1, compared with the baseline methods using two metrics, the method proposed in this disclosure achieves better performance.
[0087] Table 1 Comparison of two metrics with baseline methods
[0088]
[0089] Compared with the baseline method of this disclosure, this disclosure can suppress the target prompt without further fine - tuning the SD model.
[0090] Remove the subject from the real image. Figure 5 Shows the comparison between the P2P and the method of this disclosure. It is found that P2P has challenges in suppressing the target prompt, and in most cases, there is no obvious change in the edited image. The method of this disclosure successfully removes the objects of specific targets and generates high - quality images, indicating that the proposed method has more accurate suppression ability. The method of this disclosure usually uses the surrounding scene to fill the area where the target subject is suppressed, seeFigure 5 Note that the present disclosure does not qualitatively and quantitatively compare with ESD and concept-ablation because they are only applicable to image generation tasks rather than real image editing, and erasing each object requires a specific model.
[0091] The present disclosure evaluates the performance of the proposed method on the collected dataset. As shown in Table 1, the proposed method achieves the best scores on both CLIP-score and IFID metrics, indicating that the present disclosure has superior ability to suppress target prompt information. For example, on the IFID score, the present disclosure has a significant advantage (P2P vs Ours: 92.53 vs 166.3). This is achieved without any modification to the SD parameters.
[0092] Subject suppression in SD. In Figure 6 , the present disclosure provides a qualitative comparison of different baseline methods. As Figure 6 (top) shows, the SD model fails to generate an output that semantically matches the input prompt. The present disclosure finds that P2P is not very effective in suppressing the content of the target prompt, and the target subject remains retained. However, by incorporating the method of the present disclosure into the SD model (Ours), the subject generation can be accurately suppressed. The method of the present disclosure can precisely control different prompt words. In Figure 6 , there is no qualitative and quantitative comparison with ESD and concept-ablation because they require fine-tuning of the SD model for each target prompt word. Figure 6 It is shown that it can remove the painting styles of Tyler Edlin and Van Gogh, as well as the local object car, just like ESD and concept-ablation.
[0093] As shown in Table 1, the present disclosure evaluates the performance when removing certain painter styles and local objects. P2P has poor performance. ESD and concept-ablation have better scores. For the style of Tyler Edlin, ESD achieves the best score. However, both ESD and concept-ablation require fine-tuning of the SD model. The present disclosure wins on both Clipscore and IFID metrics (except for the paintings of Tyler Edlin) without further fine-tuning of the SD model. These results demonstrate the effectiveness of the proposed method in suppressing or removing target content.
[0094] As shown in Table 1, the present disclosure conducted a user questionnaire survey. The present disclosure asked users to select an image in which a target subject (e.g., a car) is more accurately suppressed. The present disclosure conducted a questionnaire survey on 24 users (four options, single selection required), and each user completed 20 multiple-choice questions. The results show that the method of the present disclosure is superior to other methods in suppressing target prompts.
[0095] The method of the present disclosure can also be used to enhance the generation of the subject. Instead of extracting the target embedding, the present disclosure enhances the added prompt. As Figure 7 shown, when compared with GLIGEN, competitive results can be obtained.
[0096] Ablation analysis. As Figure 7 shown, only using soft weighted regularization can achieve less accurate suppression. When combined with optimization during inference, more reasonable results are obtained.
[0097] The present disclosure found that the SD model may not be able to suppress the generation of the subject in the input prompt. The present disclosure explored the text embedding of the prompt words input to the diffusion model. The present disclosure found that [EOT] contains meaningful, redundant, and repetitive semantic information. To suppress the generation of the subject, the present disclosure provides two methods: soft weighted regularization and optimization during inference. For the former, the present disclosure designed a word embedding matrix for the target prompt word and proposed a method to remove the target prompt word information. Optimization during inference encourages retaining the desired prompt word information while further removing the target prompt word information.
[0098] Example 2
[0099] In one embodiment of the present disclosure, an image generation content suppression system based on a text-to-image diffusion model is provided, including:
[0100] A data acquisition module for acquiring the text input target prompt word of the given image to be generated and mapping it to a text embedding;
[0101] A suppression module for dividing the text embedding into two parts: the embedding expected to be suppressed and the embedding encouraged to be retained, constructing a target text embedding matrix, performing singular value decomposition on the matrix part composed of the embedding expected to be suppressed and [EOT] in the target text embedding matrix, and extracting the suppressed semantic information;
[0102] Soft weighted regularization is introduced for each singular value to restore the target text embedding matrix; the target text embedding matrix is input into the diffusion model, and through cross-attention, the corresponding attention maps of the features to be suppressed and the features to be retained are output, and two attention losses are proposed to evaluate the attention maps; an alignment loss is introduced to align the attention maps of the features to be retained within a certain number of time steps; a diversity loss is proposed to suppress the generation of the entities to be suppressed, and finally an image after removing the entities to be suppressed is generated.
[0103] Embodiment 3
[0104] In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the image generation content suppression method based on the text-to-image diffusion model is implemented.
[0105] Embodiment 4
[0106] In an embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the image generation content suppression method based on the text-to-image diffusion model.
[0107] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0109] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the scope of protection of the present disclosure.
Claims
1. An image generation content suppression method based on a text-to-image diffusion model, characterized in that Including: Obtain a text input target prompt for a given image to be generated, and map it to a text embedding; Divide the text embedding into two parts: the embedding expected to be suppressed and the embedding encouraged to be retained, construct a target text embedding matrix, perform singular value decomposition on the matrix part composed of the embedding expected to be suppressed and EOT in the target text embedding matrix, and extract the suppressed semantic information; the EOT represents the text end embedding symbol; Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, output the corresponding attention maps of the features expected to be suppressed and the features encouraged to be retained through cross-attention, propose two attention losses to evaluate the attention maps; introduce an alignment loss to align the attention maps of the features encouraged to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the subject expected to be suppressed, and finally generate an image after removing the entity expected to be suppressed.
2. The image generation content suppression method based on the text-to-image diffusion model according to claim 1, wherein The step of obtaining the text input target prompt for the image to be generated maps the text input target prompt to a text embedding through a text encoder, and defines the text embedding by adding a text start SOT symbol in front of the text input target prompt and embedding a certain number of text end EOT embedding symbols at the end.
3. The method for suppressing image generation content based on a text-to-image diffusion model according to claim 2, wherein Construct the target text embedding matrix, including: the text embedding of the text encoder is , split into two parts of text embeddings: and is the embedding to be suppressed, is the embedding encouraged to be retained, is the target prompt word for the text input, so there is: , construct the target text embedding matrix as , the target text embedding matrix contains information that is expected to be suppressed.
4. The method for suppressing image generation content based on the text-to-image diffusion model according to claim 3, wherein, Perform singular value decomposition on the target text embedding matrix. When performing singular value decomposition, obtain To guide the extraction of semantic information that needs to be suppressed from the embedding, including: Among them , singular value , U is the left singular vector matrix; Σ is the singular value matrix; V is the right singular vector matrix.
5. The method for suppressing image generation content based on the text-to-image diffusion model according to claim 4, wherein Introduce soft weighted regularization for each singular value to restore the target text embedding matrix, including: Among them, σ is the original singular value. Since σ is arranged from large to small, the larger σ corresponds to the smaller , which means suppressing the larger singular values mainly; Then restore the word embedding matrix , , the result of restoration is , and .
6. The method for suppressing image generation content based on the text-to-image diffusion model according to claim 5, wherein After soft weighted regularization, there is a new text embedding , obtain an attention map through the cross-attention of the diffusion model: , the attention map corresponds to , while corresponds to the one expected to be suppressed ; propose two attention losses to evaluate the attention map and modify the text embedding to guide the attention map to focus on specific regions corresponding to the prompt words that are encouraged to be retained.
7. The method for suppressing image generation content based on the text-to-image diffusion model according to claim 6, characterized in that, The alignment loss aligns the attention maps of the text prompts expected to be retained within a time step, including: 。 8. An image generation content suppression system based on a text-to-image diffusion model, characterized in that, Including: A data acquisition module for obtaining a text input target prompt for a given image to be generated and mapping it to a text embedding; A suppression module for dividing the text embedding into two parts: the embedding expected to be suppressed and the embedding encouraged to be retained, constructing a target text embedding matrix, performing singular value decomposition on the matrix part composed of the embedding expected to be suppressed and EOT in the target text embedding matrix, and extracting the suppressed semantic information; the EOT represents the text end embedding symbol; Introduce soft weighted regularization for each singular value to restore the target text embedding matrix; input the target text embedding matrix into the diffusion model, output the corresponding attention maps of the features expected to be suppressed and the features encouraged to be retained through cross-attention, propose two attention losses to evaluate the attention maps; introduce an alignment loss to align the attention maps of the features encouraged to be retained within a certain number of time steps; propose a diversity loss to suppress the generation of the subject expected to be suppressed, and finally generate an image after removing the entity expected to be suppressed.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method for suppressing image generation content based on a text-to-image diffusion model as described in any one of claims 1-7 is implemented.
10. An electronic device, characterized in that, Including: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device implements the method for suppressing image generation content based on a text-to-image diffusion model as described in any one of claims 1-7.
Citation Information
Patent Citations
Text summarization method and system based on deep learning combined with accumulated attention mechanism
CN109635284A
Figure clothing conversion method and system
CN111476241A