Text-driven image editing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202310086738.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-01-16
AI Technical Summary
[0005]本发明提供一种文本驱动图像编辑方法、装置、电子设备及存储介质,用以解决现有技术中限制在特定图像领域且依赖于人为指定特定图像的可编辑区域后才能编辑图像所导致的基于文本描述编辑图像的操作复杂性和应用范围受限的缺陷,实现无需人为标识掩码及无需预先训练对抗网络的情况下基于文本描述驱动图像编辑的目的,大幅提高了基于文本描述编辑图像的便捷高效性和可靠高质性,同时也大幅提高了基于文本驱动编辑图像的适用范围
[0036]本发明提供的文本驱动图像编辑方法、装置、电子设备及存储介质,其中文本驱动图像编辑方法,终端设备首先获取原始图像、定位文本和目标文本;由于定位文本用于定位原始图像中的待编辑对象,目标文本用于生成新的对象,因此基于原始图像、定位文本和目标文本确定的前向加噪潜像、方向增强潜像和定位文本在原始图像中的空间掩码进行图像重建,生成将原始图像中的待编辑对象替换为新的对象后的目标图像。以此实现无需人为标识掩码及无需预先训练对抗网络的情况下基于文本描述驱动图像编辑的目的,大幅提高了基于文本描述编辑图像的便捷高效性和可靠高质性,同时也大幅提高了基于文本驱动编辑图像的适用范围。
Smart Images

Figure CN116109733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a text-driven image editing method, apparatus, electronic device, and storage medium. Background Technology
[0002] As is well known, text-based image editing is a technique that edits the original image based on given text, that is, replacing specific objects in the original image with objects described in the given text. Therefore, it has received much attention in recent years and has also made great progress.
[0003] In related technologies, image editing based on text descriptions typically uses a pre-trained generative adversarial network (GAN). This network generates an image that depicts the same person as in the original image but with altered clothing. The original image is from a specific set of person images, and the altered clothing is from another specific set of clothing images. For example, if the text description is "replace the pants in person image A (column 4 of the person image set) with the skirt (column 2 of the clothing image set)," the text description and a manually marked editable region mask in person image A are input into the GAN model. The GAN then performs the pants-to-skirt replacement operation on the editable region mask of person image A, outputting a person image B that depicts the same person as in person image A but with the pants changed to a skirt.
[0004] However, existing text-based image editing is limited to specific human figures and is not applicable to arbitrary natural images. Furthermore, it relies on manually specifying the editable area of a particular image before human figures can be edited, which increases the complexity of text-based image editing. Summary of the Invention
[0005] This invention provides a text-driven image editing method, apparatus, electronic device, and storage medium to address the shortcomings of existing technologies that limit image editing to specific image domains and require manual specification of editable areas, resulting in operational complexity and limited application scope. It achieves text-driven image editing without the need for manual masking or pre-trained adversarial networks, significantly improving the convenience, efficiency, reliability, and quality of text-driven image editing, while also greatly expanding its applicability.
[0006] This invention provides a text-driven image editing method, comprising:
[0007] Acquire an original image, location text, and target text, wherein the location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object;
[0008] Based on the original image, the localized text, and the target text, determine the forward noise latent image, the directional enhancement latent image, and the spatial mask of the localized text in the original image;
[0009] Image reconstruction is performed based on the forward-noise latent image, the directional enhancement latent image, and the spatial mask to generate a target image, which is the image generated after the object to be edited in the original image is replaced with the new object.
[0010] According to a text-driven image editing method provided by the present invention, the step of determining a forward-noise latent image, a directional enhancement latent image, and a spatial mask of the localized text in the original image based on the original image, the located text, and the target text includes:
[0011] Based on the original image, a first forward-noise latent image and a second forward-noise latent image are determined; the first forward-noise latent image is used for image reconstruction.
[0012] Based on the second forward-noise latent image and the target text, the direction-enhanced latent image is determined;
[0013] Based on the original image and the location text, determine the spatial mask of the location text in the original image.
[0014] According to a text-driven image editing method provided by the present invention, determining the orientation-enhanced latent image based on the second forward-noising latent image and the target text includes:
[0015] Denoising and decoding are performed based on the second forward-noising latent image and the target text to determine a noise-free image of the text.
[0016] Decoding is performed based on the second forward-noise latent image to determine that there is no text in the image;
[0017] Based on the text-free image and the text-free image, directional enhancement guidance and encoding are performed to determine the directional enhancement latent image.
[0018] According to a text-driven image editing method provided by the present invention, determining the spatial mask of the positioned text in the original image based on the original image and the positioned text includes:
[0019] Semantic alignment is performed based on the original image and the located text. The semantic alignment includes binary segmentation using a decoding unit and a CLIP visual transformation unit that match the number of channels in the original image, thereby determining the spatial mask of the located text in the original image.
[0020] According to a text-driven image editing method provided by the present invention, determining a first forward-noise latent image and a second forward-noise latent image based on the original image includes:
[0021] The original image is encoded to determine the latent image;
[0022] The latent image is subjected to forward noise addition for a first preset number of steps to determine the first forward-noiseed latent image;
[0023] Perform a second preset number of forward noise additions on the first forward-noise latent image to determine the second forward-noise latent image.
[0024] According to a text-driven image editing method provided by the present invention, the step of reconstructing an image based on the forward-noising latent image, the directional enhancement latent image, and the spatial mask to generate a target image includes:
[0025] The inverse latent image result is determined by superimposing the first forward-noise latent image, the directional enhancement latent image, and the spatial mask; the first forward-noise latent image is the image after forward-noiseing the original image for a first preset number of steps.
[0026] Based on the superimposed inverse latent image results, the initial text-driven image editing results are determined;
[0027] Based on the initial text-driven image editing results, the original image, and the spatial mask, determine the loss in preserving non-editable regions in the original image;
[0028] The target image is generated based on the initial text-driven image editing results and the loss in preserving the non-editable regions.
[0029] The present invention also provides a text-driven image editing device, comprising:
[0030] The acquisition module is used to acquire an original image, location text, and target text. The location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0031] A determination module is used to determine, based on the original image, the localized text, and the target text, a forward-noise latent image, a directional enhancement latent image, and a spatial mask of the localized text in the original image;
[0032] The editing module is used to perform image reconstruction based on the forward-noise latent image, the directional enhancement latent image, and the spatial mask to generate a target image, wherein the target image is the image generated after the object to be edited in the original image is replaced with the new object.
[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the text-driven image editing method as described above.
[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the text-driven image editing method as described above.
[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the text-driven image editing method as described above.
[0036] This invention provides a text-driven image editing method, apparatus, electronic device, and storage medium. In the text-driven image editing method, the terminal device first acquires an original image, location text, and target text. Since the location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object, image reconstruction is performed based on the forward-noise latent image, the orientation-enhancing latent image, and the spatial mask of the location text in the original image, determined by the original image, location text, and target text. This generates a target image after replacing the object to be edited in the original image with the new object. This achieves text-based image editing without the need for manually identifying masks or pre-training adversarial networks, significantly improving the convenience, efficiency, reliability, and high quality of text-based image editing, and also greatly expanding the applicability of text-driven image editing. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0038] Figure 1 This is a flowchart illustrating the text-driven image editing method provided by the present invention;
[0039] Figure 2 This is a schematic diagram of the overall model structure of the text-driven image editing method provided by the present invention;
[0040] Figure 3 This is a schematic diagram showing the result of the text-driven image editing method provided by the present invention;
[0041] Figure 4 This is a schematic diagram of the structure of the text-driven image editing device provided by the present invention;
[0042] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0044] The following is combined with Figures 1-5 This invention describes a text-driven image editing method, apparatus, electronic device, and storage medium. The execution subject of the text-driven image editing method can be a terminal device or a server. The terminal device can be a personal computer (PC), portable device, laptop computer, smartphone, tablet computer, portable wearable device, or other electronic device. The server can refer to a single server, or a server cluster composed of multiple servers, a cloud computing center, etc. This invention does not limit the specific form of the terminal device or server. The following method embodiments use a terminal device as an example for illustration.
[0045] Figure 1 This is a flowchart illustrating the text-driven image editing method provided by the present invention, as shown below. Figure 1 As shown, this text-driven image editing method includes the following steps:
[0046] Step 110: Obtain the original image, location text, and target text. The location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0047] The original image can be any natural image and can be an RGB image, where RGB represents the three color channels: red (R), green (G), and blue (B). The object to be edited, located by the positioning text, can be an existing instance in the original image, while the new object generated by the target text is the new instance that needs to replace the existing instance. For example, if the object to be edited located by the positioning text is a dog in the original image, the new object generated by the target text could be a cat. In this case, the dog is an existing instance in the original image, and the cat is the new instance that needs to replace the dog. Furthermore, there can be one or more original images; the number of positioning texts and target texts can also be one or more each. The specific number of original images, positioning texts, and target texts is not specifically limited here.
[0048] Specifically, the terminal device acquires the original image, location text, and target text. This can be achieved by the user inputting these elements into the terminal device. The input methods can include, but are not limited to, input on the terminal device itself, input through other device applications, or uploading via photo capture. For example, the user can manually input the original image, location text, and target text onto the terminal device; the user can also manually input these elements into an application on another device connected to the terminal device; or the terminal device can first upload a captured original image along with image information containing the location and target text, and then recognize the image information. No specific limitations are placed on the method of acquiring the original image, location text, and target text.
[0049] Step 120: Based on the original image, the localized text, and the target text, determine the spatial masks of the forward noise latent image, the orientation enhancement latent image, and the localized text in the original image;
[0050] Specifically, based on the acquired original image, localized text, and target text, the terminal device can determine the forward-noise latent image, the directional enhancement latent image, and the spatial mask of the localized text in the original image. The forward-noise latent image can reflect the latent space or potential space in the original image, facilitating the rapid and accurate determination of editable and non-editable regions. The directional enhancement latent image can reflect the degree of proximity between the diffusion result of the target text in the original image and the original image, thereby protecting the remaining regions outside the editable regions from being constrained by the generation process or affected by the diffusion process. The spatial mask of the localized text in the original image can accurately reflect the editable region of the object to be edited in the original image.
[0051] Step 130: Reconstruct the image based on the forward-noise latent image, the orientation-enhanced latent image, and the spatial mask to generate the target image. The target image is the image generated after the object to be edited in the original image is replaced with a new object.
[0052] Specifically, the terminal device performs image reconstruction based on the forward-noise latent image, the directional enhancement latent image, and the spatial mask. In other words, it performs the operation of replacing the object to be edited in the original image with a new object, thereby generating the target image.
[0053] The text-driven image editing method provided by this invention involves a terminal device first acquiring an original image, location text, and target text. Since the location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object, image reconstruction is performed based on the forward-noise latent image, the orientation-enhancing latent image, and the spatial mask of the location text in the original image, determined by the original image, location text, and target text. This generates a target image after replacing the object to be edited in the original image with the new object. This achieves text-based image editing without the need for manually identifying masks or pre-training adversarial networks, significantly improving the convenience, efficiency, reliability, and quality of text-based image editing, and also greatly expanding its applicability.
[0054] Optionally, the specific implementation process of step 120 may include:
[0055] First, based on the original image, a first forward-noise latent image and a second forward-noise latent image are determined; the first forward-noise latent image is used for image reconstruction; then, based on the second forward-noise latent image and the target text, a direction-enhancing latent image is determined; then, based on the original image and the localized text, the spatial mask of the localized text in the original image is determined.
[0056] Specifically, the terminal device performs initial diffusion by adding Gaussian noise with different step counts to the data distribution of the latent image corresponding to the original image, thereby determining the first forward-noising latent image and the second forward-noising latent image. The number of Gaussian noise addition steps for the first and second forward-noising latent images are different, and the number of Gaussian noise addition steps for the second forward-noising latent image is greater than that for the first forward-noising latent image. At this time, the terminal device can use the first forward-noising latent image for the image reconstruction operation in step 130. At the same time, the terminal device can use the second forward-noising latent image as the enhancement direction to guide dense diffusion, and perform dense diffusion based on the second forward-noising latent image and the target text, as well as dense diffusion based on the second forward-noising latent image, to ensure that the difference between the two dense diffusion results is minimized when determining the direction enhancement latent image. Furthermore, the terminal device determines the spatial mask of the located text in the original image by semantically aligning the original image and the located text.
[0057] The text-driven image editing method provided by this invention involves a terminal device determining a first forward-noising latent image and a second forward-noising latent image for image reconstruction based on the original image. Then, based on the second forward-noising latent image and the target text, a directional enhancement latent image is determined. Simultaneously, the spatial mask of the located text in the original image is determined based on the original image and the located text. This improves the reliability and accuracy of determining different forward-noising latent images, directional enhancement latent images, and spatial masks, providing sufficient data support for subsequent text-driven image editing.
[0058] Optionally, based on the second forward-noise latent image and the target text, a direction-enhanced latent image is determined, the specific implementation process of which may include:
[0059] First, denoising and decoding are performed based on the second forward-noising latent image and the target text to determine the text-free image; then, decoding is performed based on the second forward-noising latent image to determine the text-free image; finally, orientation enhancement guidance and encoding are performed based on the text-free image and the text-free image to determine the orientation enhancement latent image.
[0060] Specifically, the terminal device can use a pre-trained denoising U-Net network to denoise the second forward-noising latent image z. T Denoise the target text t2 and determine the denoised second forward latent image z′. T-1 For the denoised second forward latent image z′ T-1 Decode the text to determine the noise-free image x′ T-1 Noise-free image x′ T-1 It is an RGB image; simultaneously, the terminal device can target the second forward-noising latent image z. T Decode the second forward RGB image x. T For the second forward RGB image x T After applying text-free constraints, determine the text-free image x. T-1 No text image x T-1 It is also an RGB image. At this point, based on the text-free image x′ T-1 and textless image x T-1 Enhanced directional guidance is performed using a textless image x. T-1 To enhance directionality, guide the text to a noise-free image x′ T-1 and textless image x T-1 The difference between them is infinitesimally small, thus determining the direction of RGB image x′ that corresponds to the minimum difference. T-1 , for directional enhancement of RGB image x′ T-1 Encode the directional enhancement latent image z″. T-1 .
[0061] It should be noted that, in order to enable the editable regions in the original image to be edited according to the target text t2, a pre-trained ViT-L / 14 CLIP model is used to perform text-driven image content editing operations. That is, during the diffusion process, the ViT-L / 14 CLIP model will edit the second forward RGB image x. T The cosine distance between the CLIP embedding of t1 and the CLIP embedding of the target text t2 is used to specify the CLIP-based loss. The target text t2 is embedded into the embedding space defined as E. L(t2), used for the second forward-noising latent image z T A time-varying image encoder can be defined as E′ I Based on this, we can use equation (1) to measure the similarity between embeddings. The text guidance function is defined using cosine distance:
[0062]
[0063] In equation (1), ⊙ represents the dot product, m is the spatial mask of the located text in the original image, and D(z) T For the second forward-noising latent image z T Decode, x T Let be the second forward RGB image, and t be the time variable. Equation (1) shows that the diffusion process here is not limited by any additional non-editable regions and is evaluated within the editable region of the original image; and the diffusion process here can be the process of how the target text works.
[0064] Furthermore, it should be noted that, to enhance the consistency between the edited text-driven image and the content of the target text, this invention proposes an enhanced direction-guided diffusion module to strengthen cross-modal guidance. This enhanced direction-guided diffusion module is a strategy for guiding the diffusion model and does not require training a separate classifier model. To provide classifier-free guidance, the class-conditional diffusion model ∈ […]. θ (x T The category label y in |y) is replaced with an empty label. Obtain the class-label-free conditional diffusion model During the sampling process, the output of the category conditional diffusion model is in ∈ θ (x T Extend further in the direction of |y) and away from Prediction of conditional diffusion models without classification guidance The expression is given by equation (3):
[0065]
[0066] Wherein, the guidance scale s = 5. Conditional diffusion model prediction. This is then used to guide the diffusion direction toward the target text t2, that is, the category label y in the above formula is replaced with the target text t2: thus obtaining the directional enhancement latent image z″ T-1 That is, the result obtained from equation (4)
[0067]
[0068] The text-driven image editing method provided by this invention allows a terminal device to determine a text-free image by denoising and decoding a second forward-noising latent image and the target text, to determine a text-free image by decoding the second forward-noising latent image, and to determine a directional enhancement latent image by directional enhancement guidance and encoding the text-free image and the text-free image. This achieves the purpose of directional enhancement guided by the target text, thereby improving the accuracy and reliability of dense diffusion and ensuring the high quality of subsequent image editing results.
[0069] Optionally, based on the original image and the located text, a spatial mask of the located text in the original image is determined. The specific implementation process may include:
[0070] Semantic alignment is performed based on the original image and the located text. The semantic alignment includes binary segmentation using a decoding unit and a CLIP visual transformation unit that match the number of channels in the original image, thereby determining the spatial mask of the located text in the original image.
[0071] Specifically, the terminal device performs semantic alignment on the original image x0 and the localized text t1. This can be achieved by inputting the original image and the localized text t1 into a pre-constructed cross-modal entity-level calibration model for semantic alignment. This model includes decoding units and CLIP visual transformation units that match the number of channels in the original image. Based on this, when the original image and the localized text t1 are input into the cross-modal entity-level calibration model, the localized text t1 is fed into each CLIP visual transformation unit to obtain a conditional vector. Each conditional vector is used to adjust the input activation of the corresponding decoding unit, enabling the decoding unit to correlate the activation within the CLIP with the output segmentation and informing each decoding unit of the segmentation target. The input image x0 then passes through each CLIP visual transformation unit to obtain the corresponding... Subsequently, the activations extracted from different decoding units S are added to the internal activations of the decoding unit of embedding size P before each CLIP visual transform unit, which generates binary segmentation by linearly projecting the tokens of the last layer of its CLIP visual transform unit:
[0072]
[0073] In equation (5), S = [3, 7, 9], D = 64, P = 16; To link CLIP's capabilities with the segmentation results, this invention employs a universal binary prediction setting, namely, setting a threshold for the binary segmentation of the spatial mask m of the located text in the original image. The threshold K ranges from 0 to 255, typically set to 150. This yields the spatial mask of the located text in the original image, which accurately reflects the editable area of the object to be edited in the original image.
[0074] The text-driven image editing method provided by this invention involves a terminal device semantically aligning the original image and the located text using a decoding unit and a CLIP visual transformation unit that match the number of channels in the original image, thereby determining the spatial mask of the located text in the original image. This achieves cross-modal instance-level calibration of the original image and the located text, improving the efficiency and accuracy of automatically determining the spatial mask of the located text in the original image.
[0075] Optionally, based on the original image, a first forward-noising latent image and a second forward-noising latent image are determined. The specific implementation process may include:
[0076] The original image is encoded to determine the latent image; the latent image is subjected to forward noise addition for a first preset number of steps to determine the first forward-noised latent image; the first forward-noised latent image is subjected to forward noise addition for a second preset number of steps to determine the second forward-noised latent image.
[0077] Specifically, the terminal device first encodes the original image x0 to determine the latent image z of the original image x0, z = ε(x0), where ε(x0) is the encoding of the original image x0; then, for the forward process of the latent image z, it performs T-1 steps of Gaussian noise addition to determine the first forward-noised latent image z. T-1 Then, for the first forward-noisy latent image z T-1 The forward process involves a single step of Gaussian noise addition to determine the second forward-noised latent image z. T That is, the second forward-noising latent image z T This refers to the noisy latent image obtained by adding Gaussian noise in T steps during the forward process of the latent image z. The first preset number of steps is T-1 steps, and the second preset number of steps is 1 step.
[0078] The text-driven image editing method provided by this invention involves a terminal device first encoding the original image, then performing a first preset number of forward noise additions, and finally performing a second preset number of forward noise additions to determine a first and a second forward noise-added latent image. This method achieves diffusion of the original image through encoding and forward noise addition, improving the reliability and efficiency of determining different forward noise-added latent images.
[0079] Optionally, the specific implementation process of step 130 may include:
[0080] First, the inverse latent image is determined by superimposing the first forward-noising latent image, the directional enhancement latent image, and the spatial mask. The first forward-noising latent image is the image after forward-noising the original image for a first preset number of steps. Then, the initial text-driven image editing result is determined based on the superimposed inverse latent image result. Further, the loss for preserving uneditable regions in the original image is determined based on the initial text-driven image editing result, the original image, and the spatial mask. Finally, the target image is generated based on the initial text-driven image editing result and the loss for preserving uneditable regions.
[0081] Specifically, the terminal device superimposes the first forward-noise latent image, the directional enhancement latent image, and the spatial mask to determine the superimposed reverse latent image result. Then regarding the results of the reverse latent image Perform a reverse process, which is the opposite of the aforementioned forward process, to obtain the target latent image result. And the results of the latent image of the target After decoding, the initial text-driven image editing result is obtained. At this point, the initial text-driven image editing results are used. Given the original image x0 and the spatial mask m, determine the loss for preserving non-editable regions in the original image. The calculation formula is shown in equation (6).
[0082]
[0083] In equation (6), x1 = x0⊙(1-m), ⊙ represents dot product. Based on this, the terminal device will use the initial text-driven image editing result. Loss of preservation of non-editable areas The result obtained after addition is determined as the target image.
[0084] It should be noted that even if the loss based on CLIP is Editing occurs within the editable area of the original image, but it still affects the entire original image. To improve this issue, this invention proposes a non-editable region protection module that incorporates a spatial mask m into the diffusion process. The latent image in subsequent latent diffusion steps is obtained by mixing the first forward-noising latent image z with the adjusted spatial mask m. T-1 and directional enhancement latent image z″ T-1 These two results produce z″ T-1 ⊙m latent +z T-1 ⊙(1-m latent Since the entire latent image is modified in each denoising step, but subsequent blending forces m... latentThe areas outside the mask remain unchanged. At this stage, the background is strictly preserved by replacing the entire area outside the mask with a comparable region from the original image. Subsequent latent image denoising ensures consistency, even if the resulting blended latent images are not always coherent. After the latent image diffusion process is complete, a decoder is used to decode the resulting latent image into the output image. Furthermore, a loss is applied to preserve non-editable regions outside the editable regions. This guides the diffusion outside the mask, causing the area surrounding the editable region to expand in the direction of the original image x0.
[0085] The text-driven image editing method provided by this invention involves the terminal device first determining the inverse latent image result after superimposing a first forward-noise latent image, a directional enhancement latent image, and a spatial mask. Then, based on the inverse latent image result, the original image, and the spatial mask, the loss for preserving uneditable regions in the original image is determined. Finally, the target image is generated based on the inverse latent image result and the loss for preserving uneditable regions, thereby improving the accuracy and reliability of the text-driven image editing result.
[0086] Reference Figure 2 This is a schematic diagram of the overall model structure of the text-driven image editing method provided by the present invention, as shown below. Figure 2 As shown, the first forward-noising latent image z is determined based on the original image x0. T-1 Second forward-noise latent image z T And based on the second forward-noising latent image z T Determine the noise-free image x′ of the text and the target text t2. T-1 , and based on the second forward-noising latent image z T Determine if there is no text in the image x T-1 Based on the text-free image x′ T-1 and textless image x T-1 After directional enhancement guidance, the directional enhancement latent image z″ is determined. T-1 Simultaneously, after semantic alignment of the original image x0 and the located text t1, the spatial mask m of the located text in the original image is determined; finally, the first forward-noising latent image z is then... T-1 , Directional enhancement latent image z″ T-1 The target image is generated by first superimposing the spatial mask m, and then decoding it. Here, ε is the encoder. For decoder, This enhances directional guidance; and the specific process involved can be referred to in the aforementioned embodiments. It will not be repeated here. Thus, the following can be obtained: Figure 3 The diagram showing the results is in Figure 3In the diagram, Input is the original image, and Output is the generated target image. "Flower" represents the location text as "flower," "Chrysanthemum" represents the target text as "chrysanthemum," "A horse" represents the location text as "a horse," and "Azebra" represents the target text as "a zebra." "Plate" represents the location text as "plate," and "Dinner" represents the target text as "dinner." "A dog" represents the location text as "a dog," and "A cat" represents the target text as "a cat." "T-shirt" represents the location text as "t-shirt," and "Polo shirt" represents the target text as "open-necked short-sleeved shirt." "Flower" represents the location text as "flower," and "Hydrangea" represents the target text as "hydrangea." "Cornfield" represents the location text as "cornfield," and "Grass" represents the target text as "lawn." "A fish" represents the location text as "a fish," and "A goldfish" represents the target text as "a goldfish."
[0087] The text-driven image editing apparatus provided by the present invention will be described below. The text-driven image editing apparatus described below can be referred to in correspondence with the text-driven image editing method described above.
[0088] Reference Figure 4 This is a schematic diagram of the structure of the text-driven image editing device provided by the present invention, as shown below. Figure 4 As shown, the text-driven image editing device 400 includes:
[0089] The acquisition module 410 is used to acquire the original image, the positioning text, and the target text. The positioning text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0090] The determination module 420 is used to determine the spatial mask of the forward noise latent image, the orientation enhancement latent image, and the localized text in the original image based on the original image, the localized text, and the target text.
[0091] The editing module 430 is used to reconstruct an image based on a forward-noise latent image, a directional enhancement latent image, and a spatial mask, and generate a target image. The target image is the image generated after the object to be edited in the original image is replaced with a new object.
[0092] Optionally, the determining module 420 can be used to determine a first forward-noise latent image and a second forward-noise latent image based on the original image; the first forward-noise latent image is used for image reconstruction; based on the second forward-noise latent image and the target text, a direction-enhancing latent image is determined; based on the original image and the located text, a spatial mask of the located text in the original image is determined.
[0093] Optionally, the determination module 420 can also be used to perform denoising and decoding based on the second forward-noising latent image and the target text to determine a text-free image; to perform decoding based on the second forward-noising latent image to determine a text-free image; and to perform orientation enhancement guidance and encoding based on the text-free image and the text-free image to determine an orientation enhancement latent image.
[0094] Optionally, the determination module 420 can also be used to perform semantic alignment based on the original image and the located text. The semantic alignment includes binary segmentation using a decoding unit and a CLIP visual transformation unit that match the number of channels of the original image, thereby determining the spatial mask of the located text in the original image.
[0095] Optionally, the determining module 420 can also be used to encode the original image to determine the latent image; to perform forward noise addition on the latent image for a first preset number of steps to determine the first forward-noised latent image; and to perform forward noise addition on the first forward-noised latent image for a second preset number of steps to determine the second forward-noised latent image.
[0096] Optionally, the editing module 430 can be used to determine the superimposed inverse latent image result based on the superposition of the first forward-noising latent image, the orientation-enhanced latent image, and the spatial mask; the first forward-noising latent image is the image after forward-noising the original image for a first preset number of steps; based on the superimposed inverse latent image result, the initial text-driven image editing result is determined; based on the initial text-driven image editing result, the original image, and the spatial mask, the loss for preserving uneditable regions in the original image is determined; based on the initial text-driven image editing result and the loss for preserving uneditable regions, the target image is generated.
[0097] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device 500 may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a text-driven image editing method, which includes:
[0098] The process involves acquiring the original image, the location text, and the target text. The location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0099] Based on the original image, the localized text, and the target text, determine the spatial masks of the forward noise latent image, the orientation enhancement latent image, and the localized text in the original image;
[0100] Image reconstruction is performed based on forward-noise latent image, directional enhancement latent image and spatial mask to generate target image. The target image is the image generated after the object to be edited in the original image is replaced with a new object.
[0101] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0102] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the text-driven image editing method provided by the above methods, the method comprising:
[0103] The process involves acquiring the original image, the location text, and the target text. The location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0104] Based on the original image, the localized text, and the target text, determine the spatial masks of the forward noise latent image, the orientation enhancement latent image, and the localized text in the original image;
[0105] Image reconstruction is performed based on forward-noise latent image, directional enhancement latent image and spatial mask to generate target image. The target image is the image generated after the object to be edited in the original image is replaced with a new object.
[0106] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the text-driven image editing methods provided by the methods described above, the method comprising:
[0107] The process involves acquiring the original image, the location text, and the target text. The location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object.
[0108] Based on the original image, the localized text, and the target text, determine the spatial masks of the forward noise latent image, the orientation enhancement latent image, and the localized text in the original image;
[0109] Image reconstruction is performed based on forward-noise latent image, directional enhancement latent image and spatial mask to generate target image. The target image is the image generated after the object to be edited in the original image is replaced with a new object.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text-driven image editing method, characterized in that, include: Acquire an original image, location text, and target text, wherein the location text is used to locate the object to be edited in the original image, and the target text is used to generate a new object; Based on the original image, the localized text, and the target text, determine the forward noise latent image, the directional enhancement latent image, and the spatial mask of the localized text in the original image; Image reconstruction is performed based on the forward-noise latent image, the directional enhancement latent image, and the spatial mask to generate a target image. The target image is the image generated after the object to be edited in the original image is replaced with the new object. The step of determining the spatial mask of the forward-noise latent image, the directional enhancement latent image, and the localized text in the original image based on the original image, the localized text, and the target text includes: Based on the original image, a first forward-noise latent image and a second forward-noise latent image are determined; the first forward-noise latent image is used for image reconstruction; the number of Gaussian noise addition steps corresponding to the first forward-noise latent image and the second forward-noise latent image is different, and the number of Gaussian noise addition steps corresponding to the second forward-noise latent image is greater than the number of Gaussian noise addition steps corresponding to the first forward-noise latent image. Based on the second forward-noise latent image and the target text, the direction-enhanced latent image is determined; Based on the original image and the location text, determine the spatial mask of the location text in the original image; The step of determining the directional enhancement latent image based on the second forward-noising latent image and the target text includes: Denoising and decoding are performed based on the second forward-noising latent image and the target text to determine a noise-free image of the text. Decoding is performed based on the second forward-noise latent image to determine that there is no text in the image; Based on the text-free image and the text-free image, directional enhancement guidance and encoding are performed to determine the directional enhancement latent image.
2. The text-driven image editing method according to claim 1, characterized in that, The step of determining the spatial mask of the location text in the original image based on the original image and the location text includes: Semantic alignment is performed based on the original image and the located text. The semantic alignment includes binary segmentation using a decoding unit and a CLIP visual transformation unit that match the number of channels in the original image, thereby determining the spatial mask of the located text in the original image.
3. The text-driven image editing method according to claim 1, characterized in that, The step of determining the first forward-noising latent image and the second forward-noising latent image based on the original image includes: The original image is encoded to determine the latent image; The latent image is subjected to forward noise addition for a first preset number of steps to determine the first forward-noiseed latent image; Perform a second preset number of forward noise additions on the first forward-noise latent image to determine the second forward-noise latent image.
4. The text-driven image editing method according to any one of claims 1 to 3, characterized in that, The step of reconstructing the target image based on the forward-noising latent image, the directional enhancement latent image, and the spatial mask includes: The inverse latent image result is determined by superimposing the first forward-noise latent image, the directional enhancement latent image, and the spatial mask; the first forward-noise latent image is the image after forward-noiseing the original image for a first preset number of steps. Based on the superimposed inverse latent image results, the initial text-driven image editing results are determined; Based on the initial text-driven image editing results, the original image, and the spatial mask, determine the loss in preserving non-editable regions in the original image; The target image is generated based on the initial text-driven image editing results and the loss in preserving the non-editable regions.
5. A text-driven image editing device, characterized in that, The apparatus is used to implement the text-driven image editing method as described in any one of claims 1 to 4.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the text-driven image editing method as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text-driven image editing method as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the text-driven image editing method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Image processing method, model training method and related device
CN114943789A