Image editing method and device and storage medium
By utilizing the potential diffusion model and fine-tuning model, combined with the text information input by the user, an extended area that meets the similarity conditions is generated, which solves the problems of low degree of image editing, insufficient accuracy and inability to retain the visual attributes of the original image in the prior art, and achieves efficient and accurate image extended editing.
Patent Information
- Application Number
- CN202311526868.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has low degree of automation in image editing, insufficient accuracy, and cannot perform extended editing based on the content of the original image. The editing effect is poor, making it difficult to retain the visual attributes of the original image.
By obtaining the text information of the image to be edited and the user input, the pre-trained potential diffusion model (LDM) and fine-tuning model are used to generate the target image, and the similarity conditions are met between the extended area and the area to be edited, thereby realizing extended editing of the image.
It realizes efficient extended editing of edited images, retains the visual attributes of the original image, improves the user experience, and improves the automation and accuracy of image editing.
Smart Images

Figure CN120013837A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of image recognition, image editing and related artificial intelligence technologies, and in particular to an image editing method, device and storage medium. Background Art
[0002] Among the related technologies, image editing and generation technology is an important research direction in the field of computer vision. Its main purpose is to edit and generate images through computer algorithms to achieve certain goals. At present, the research status of image editing and generation technology mainly includes the following aspects:
[0003] (1) Image restoration: Image restoration refers to the use of computer algorithms to repair defects, noise, distortion, etc. in images to improve the quality and clarity of the image. Currently, image restoration technologies mainly include methods based on interpolation, texture synthesis, and deep learning.
[0004] (2) Image enhancement: Image enhancement refers to enhancing an image through computer algorithms to improve the brightness, contrast, clarity, etc. At present, image enhancement technologies mainly include methods based on histogram equalization, filtering, and deep learning.
[0005] (3) Image synthesis: Image synthesis refers to the synthesis of multiple images through computer algorithms to generate a new image. At present, image synthesis technology mainly includes methods based on image fusion, image splicing, and deep learning.
[0006] (4) Image generation: Image generation refers to the process of generating a new image through a computer algorithm to meet certain requirements. Currently, image generation technologies mainly include methods based on generative adversarial networks (GANs), variational autoencoders (VAEs), and deep learning.
[0007] In general, the research status of image editing and generation technology mainly includes image restoration, image enhancement, image synthesis and image generation, among which deep learning technology is being used more and more widely in the field of image editing and generation.
[0008] In addition, the application scenarios of image editing and text image technology are very wide. The following are some common application scenarios:
[0009] (1) Image design applications: For example, image editing and text-based graphics technology can be used in advertising design, such as making posters, flyers, billboards, etc.; film and video production, such as special effects production, scene synthesis, character design, etc.; image design in game development, such as character design, scene design, special effects production, etc.; artistic creation, such as digital art, virtual reality art, interactive art, etc.
[0010] (2) Photo beautification on mobile phones: Image editing and text-based image processing techniques on mobile devices can be used to beautify photos, such as applying filters, retouching, and graffiti to photos to make them more beautiful. They can also be used on social media, such as cropping, compositing, and adding text to photos to make them more interesting.
[0011] (3) Image editing and text-based image processing technology on mobile devices can be used on e-commerce platforms, such as modifying, compositing, and adding tags to product photos to make them more attractive. They can also be used in education and training, such as producing teaching materials, designing courseware, and producing learning materials to make education more vivid.
[0012] In general, the application scenarios of image editing and image processing technology are very wide and can be applied to various fields, bringing convenience and innovation to people's lives and work. Summary of the invention
[0013] In order to overcome the problems existing in the related art, the present disclosure provides an image editing method, device and storage medium.
[0014] According to a first aspect of an embodiment of the present disclosure, there is provided an image editing method, comprising: acquiring an image to be edited and text information input by a user, wherein the text information is descriptive information that the user expects to perform extended editing on the image to be edited; based on the image to be edited and the text information, generating a target image, wherein the target image comprises an extended area after extended editing of the image to be edited, wherein a similarity condition is satisfied between the extended area and the area to be edited in the image to be edited.
[0015] In one embodiment, generating a target image based on the image to be edited and the text information includes: determining a region to be edited in the image to be edited; obtaining a first text information marker based on the image to be edited, the text information, the region to be edited, and a pre-trained latent diffusion model LDM, wherein the first text information marker is used to mark an extended region that has a similarity condition with the region to be edited; calling a fine-tuned latent diffusion model to iteratively denoise the first text information marker to obtain a target image.
[0016] In one embodiment, the first text information marker is obtained based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM, including: obtaining a first text feature between the image to be edited and the text information based on the image to be edited, the text information, and a pre-trained latent diffusion model LDM; performing mask text inversion on the area to be edited in the image to be edited, and fusing it with the first text feature to obtain a first text information marker.
[0017] In one embodiment, the first text feature between the image to be edited and the text information is obtained based on the image to be edited, the text information and a pre-trained latent diffusion model LDM, including: performing word segmentation processing on the text information and extracting a second text feature of the text after the word segmentation processing; adding random noise to the image to be edited and encoding the image to be edited with added random noise to obtain a third text feature; inputting the second text feature and the third text feature into the pre-trained LDM to obtain a first text feature.
[0018] In one embodiment, determining the area to be edited in the image to be edited includes: in response to a user selecting a target area in the image to be edited, using the target area as the area to be edited; or in response to the user not selecting a target area in the image to be edited, using the entire area of the image to be edited as the area to be edited.
[0019] In one embodiment, the method further includes: adding random noise to the image to be edited, and based on the pre-trained LDM, iteratively looping the image to be edited with the random noise added to obtain a fourth text feature; calling the fine-tuned latent diffusion model to iteratively denoise the first text information marker to obtain a target image, including: fusing the second text information marker corresponding to the fourth text feature with the first text information marker; calling the fine-tuned latent diffusion model to iteratively denoise the fused text marker to obtain the target image.
[0020] In one embodiment, the calling of the fine-tuned latent diffusion model to iteratively denoise the first text information marker to obtain a target image includes: determining the attention channel value of the text information marker based on the key projection and query projection of the multi-head attention mechanism between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model; and multiplying the attention channel value with the image to be edited to obtain the target image.
[0021] According to a second aspect of an embodiment of the present disclosure, there is provided an image editing device, comprising: an acquisition unit, for acquiring an image to be edited and text information input by a user, wherein the text information is descriptive information that the user expects to perform extended editing on the image to be edited; and an execution unit, for generating a target image based on the image to be edited and the text information, wherein the target image comprises an extended area after extended editing of the image to be edited, wherein a similarity condition is satisfied between the extended area and the area to be edited in the image to be edited.
[0022] In one embodiment, the execution unit generates a target image based on the image to be edited and the text information in the following manner: determining a region to be edited in the image to be edited; obtaining a first text information marker based on the image to be edited, the text information, the region to be edited, and a pre-trained latent diffusion model LDM, wherein the first text information marker is used to mark an extended region that has a similarity condition with the region to be edited; calling a fine-tuned latent diffusion model to iteratively denoise the first text information marker to obtain a target image.
[0023] In one embodiment, the execution unit obtains a first text information label based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM in the following manner: based on the image to be edited, the text information, and the pre-trained latent diffusion model LDM, a first text feature between the image to be edited and the text information is obtained; mask text inversion is performed on the area to be edited in the image to be edited, and the mask text is fused with the first text feature to obtain a first text information label.
[0024] In one embodiment, the execution unit obtains a first text feature between the image to be edited and the text information based on the image to be edited, the text information and a pre-trained latent diffusion model LDM in the following manner: performing word segmentation processing on the text information and extracting a second text feature of the text after the word segmentation processing; adding random noise to the image to be edited and performing encoding conversion on the image to be edited with added random noise to obtain a third text feature; inputting the second text feature and the third text feature into the pre-trained LDM to obtain a first text feature.
[0025] In one embodiment, the execution unit determines the area to be edited in the image to be edited in the following manner: in response to a user selecting a target area in the image to be edited, the target area is used as the area to be edited; or in response to the user not selecting a target area in the image to be edited, the entire area of the image to be edited is used as the area to be edited.
[0026] In one embodiment, the execution unit is also used to: add random noise to the image to be edited, and based on the pre-trained LDM, iteratively loop process the image to be edited with the random noise added to obtain a fourth text feature; the execution unit calls the fine-tuned latent diffusion model in the following manner to iteratively denoise the first text information marker to obtain a target image: fuse the second text information marker corresponding to the fourth text feature with the first text information marker; call the fine-tuned latent diffusion model to iteratively denoise the fused text marker to obtain a target image.
[0027] In one embodiment, the execution unit calls the fine-tuned latent diffusion model in the following manner to iteratively denoise the first text information marker to obtain a target image: between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model, based on the key projection and query projection of the multi-head attention mechanism, the attention channel value of the text information marker is determined; and the attention channel value is multiplied by the image to be edited to obtain the target image.
[0028] According to a third aspect of an embodiment of the present disclosure, there is provided an image editing device, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to: execute the image editing method described in the first aspect or any one of the embodiments of the first aspect.
[0029] According to the fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to execute the image editing method described in the first aspect or any one of the embodiments of the first aspect.
[0030] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining an image to be edited and description information for extended editing of the image to be edited input by a user, and generating a target image after extended editing of the image to be edited based on the image to be edited and the description information, wherein the style of the extended area is similar to that of the area to be edited in the image to be edited.
[0031] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0033] Figure 1 is a schematic diagram of an image editing method according to an exemplary embodiment.
[0034] Figure 2 The figure is a flowchart of an image editing method according to an exemplary embodiment.
[0035] Figure 3 is a schematic diagram of an image editing method according to an exemplary embodiment.
[0036] Figure 4 The present invention is a flowchart showing a method of generating a target image based on an image to be edited and text information according to an exemplary embodiment.
[0037] Figure 5 It is a flowchart showing a method of obtaining a first text information label based on an image to be edited, text information, a region to be edited, and a pre-trained latent diffusion model LDM according to an exemplary embodiment.
[0038] Figure 6 The present invention is a flowchart showing a method of obtaining a first text feature between an image to be edited and text information based on an image to be edited, text information and a pre-trained latent diffusion model LDM according to an exemplary embodiment.
[0039] Figure 7 The present invention is a flowchart showing a method for determining a to-be-edited area in an to-be-edited image according to an exemplary embodiment.
[0040] Figure 8 The figure is a flowchart of an image editing method according to an exemplary embodiment.
[0041] Fig. 9 The invention is a block diagram of an image editing device according to an exemplary embodiment.
[0042] Fig.10 The invention is a block diagram of an image editing device according to an exemplary embodiment.
[0043] Fig.11 The invention is a block diagram of an image editing device according to an exemplary embodiment. DETAILED DESCRIPTION
[0044] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure.
[0045] In the related art, there are still some problems with image editing, such as:
[0046] (1) The degree of automation is not high enough: Although image editing and image processing technology can automatically complete some tasks, in some cases, human intervention and adjustment are still required.
[0047] (2) Insufficient accuracy: Although image editing and text-based mapping techniques can process large amounts of data, in some cases, such as when processing complex images or text, their accuracy may not be high enough.
[0048] (3) It is impossible to perform extended editing based on the original image content. Usually, the difficulty of extended editing lies in preserving the visual attributes of the original image, which will be an important reason hindering the progress of related technologies.
[0049] (4) Poor editing effect. Most current technologies are based on methods similar to image editing (PhotoShop, PS), and have a high threshold. The edited image effect cannot fit the original image, nor can it be expanded to a larger image.
[0050] Based on this, the embodiment of the present disclosure provides an image editing method, the essence of which is to establish new image content relationships by relying on the understanding of the objects contained in the current image, and at the same time, it can be infinitely expanded, retaining the inherent visual properties of the image.
[0051] In an exemplary embodiment, the image editing method provided by the embodiment of the present disclosure can realize modeling and reasoning on PC and mobile terminals, and can be embedded in App in the future.
[0052] In the embodiments of the present disclosure, the key theoretical basis that needs to be used is the latent space diffusion model. Regarding the latent diffusion model (Latent Diffusion Model, LDM), in order to facilitate the technical description below, its technical principle is briefly explained.
[0053] LDM consists of two key stages, in the first stage, the autoencoder maps the image to a latent space z0=E(I) and the decoder maps it back to an image D(E(I))≈I. In the second stage, a diffusion model ∈θ is trained to denoise the noisy latent, where αt is a factor that determines the noise level at each time step t and ∈~N(0,1) is Gaussian noise. The diffusion model is then trained to predict the added Gaussian noise with the LDM loss:
[0054]
[0055] Among them, y is the input text condition and cθ is the text encoder.
[0056] Figure 1 is a schematic diagram of an image editing method according to an exemplary embodiment. Figure 1 As shown, based on Figure 1 As well as the above-mentioned LDM modeling method, the embodiments of the present disclosure are described in detail below:
[0057] Figure 2 is a flowchart of an image editing method according to an exemplary embodiment. Figure 2 As shown, the following steps are included.
[0058] In step S11, the image to be edited and the text information input by the user are obtained.
[0059] In the embodiment of the present disclosure, the text information is description information that the user expects to perform extended editing on the image to be edited.
[0060] In the embodiment of the present disclosure, the user inputs an image to be edited and inputs content that the user wants to expand on the image to be edited.
[0061] In step S12, a target image is generated based on the image to be edited and the text information.
[0062] In the embodiment of the present disclosure, the target image includes an extended region after the image to be edited is extended and edited, and the extended region satisfies a similarity condition with the region to be edited in the image to be edited.
[0063] Figure 3 is a schematic diagram of an image editing method according to an exemplary embodiment. Figure 3 As shown. M1, M2 and M3 in the image A to be edited are the areas selected by the user, that is, the editing and expansion are performed according to the style details of the areas selected by the user. The user also needs to input how to expand the image to be edited, for example: regenerate a cat. In this way, based on the areas M1, M2 and M3 selected by the user and the user's extended description of the image to be edited, the target image B is generated. Among them, B is a randomly generated image, that is, the image generated each time is not necessarily the same, but it has the same style as the image to be edited and is expanded and edited based on the user's description.
[0064] In the disclosed embodiment, based on the image to be edited input by the user and the user's extended description of the image to be edited, a target image similar in style to the image to be edited and based on the user's description can be generated, allowing the user to perform extended editing on any image, thereby improving the user experience.
[0065] Figure 4 is a flowchart showing a method of generating a target image based on an image to be edited and text information according to an exemplary embodiment. Figure 4 As shown, the following steps are included.
[0066] In step S21, a region to be edited is determined in the image to be edited.
[0067] In the disclosed embodiment, the user determines the area to be edited in the image to be edited. For example, the user may randomly select a few places in the image to be edited, and expand based on the style of these places. Alternatively, the user may not select the image to be edited, and directly expand based on the overall style of the image to be edited. Compared with directly using the image to be edited, randomly selecting a few places in the image to be edited can make the style and details of the generated target image closer to the original image, and the effect is better.
[0068] In step S22, a first text information tag is obtained based on the image to be edited, the text information, the region to be edited, and the pre-trained latent diffusion model LDM.
[0069] In the embodiment of the present disclosure, the first text information mark is used to mark the extended area that has a similarity condition with the area to be edited.
[0070] In the disclosed embodiment, the image to be edited input by the user is added with random noise N(0,1) to obtain zt, and the obtained zt is converted into a feature representation of the latent space after being encoded by the encoder. At the same time, the text information input by the user is segmented to obtain a[1], a[2], a[3], and a[1], a[2], a[3] are marked to obtain tokens[a1, a2, a3], and tokens[a1, a2, a3] are input into the Transformer model to obtain the text feature representation, i.e., F(f). That is, the feature representations corresponding to the image to be edited and the user's text information are obtained, and the feature representations corresponding to the two are input into the pre-trained latent diffusion model (LatentDF) to obtain the fused L(f). The area to be edited in the image to be edited is obtained. This process can select the area to be edited for the user, or when the user does not select it, the entire image to be edited is selected. In the area to be edited, the area of each object is represented by masks Mk1, Mk2,..., MkN, i.e., M1, M2, M3. Multiply L(f) by M1, M2, and M3 respectively to obtain a new object y*=[a1*, a2*, a3*], that is, to obtain the first text information tag.
[0071] In step S23, the fine-tuned latent diffusion model is called to iteratively denoise the first text information tag to obtain a target image.
[0072] In the disclosed embodiment, the first text information marker is input into the fine-tuned latent diffusion model, and the first text information marker generated above is iteratively denoised to obtain a target image, that is, an expanded and edited image.
[0073] In the disclosed embodiment, based on the image to be edited and the text information input by the user, the target image expected by the user is generated. When the user is not satisfied with the content of a certain image, the user can give full play to his imagination and can obtain the expected effect image based on the image to be edited through text description without being proficient in painting. The position of the objects in the image to be edited can be rearranged to adapt to the target layout, and the visual attributes of the image to be edited are retained.
[0074] Figure 5 is a flowchart of obtaining a first text information tag based on an image to be edited, text information, a region to be edited, and a pre-trained latent diffusion model LDM according to an exemplary embodiment, as shown in FIG. Figure 5 As shown, the following steps are included.
[0075] In step S31, based on the image to be edited, the text information and the pre-trained latent diffusion model LDM, a first text feature between the image to be edited and the text information is obtained.
[0076] In the disclosed embodiment, random noise N(0,1) is added to the image to be edited input by the user to obtain zt, and the obtained zt is converted into a feature representation of the latent space after being encoded by the encoder. At the same time, the text information input by the user is segmented to obtain a[1], a[2], a[3], and a[1], a[2], a[3] are marked to obtain tokens[a1, a2, a3], and tokens[a1, a2, a3] are input into the Transformer model to obtain the text feature representation, i.e., F(f). That is, the feature representations corresponding to the image to be edited and the user text information are obtained, and the feature representations corresponding to the two are respectively input into the pre-trained latent diffusion model (LatentDF) to obtain the fused L(f), i.e., the first text feature.
[0077] In step S32, the mask text of the area to be edited in the image to be edited is inverted and fused with the first text feature to obtain a first text information mark.
[0078] In the disclosed embodiment, the region to be edited in the image to be edited is obtained. This process can be multiple objects selected by the user or the entire image to be edited when the user does not select. The region of each object is represented by a mask Mk1, Mk2, ..., MkN, i.e., M1, M2, M3. The L(f) obtained above is multiplied by M1, M2, M3 respectively to obtain the first text information mark y*=[a1*, a2*, a3*].
[0079] That is, through mask text inversion, the concepts of multiple objects are learned from a single input image to be edited I to text tags a1, a2, ..., aN, where the area of each object in the area to be edited is represented by a mask Mk1, Mk2, ..., MkN, and the details of the object are further understood by fine-tuning the diffusion model ∈θ and optimizing the additional text tags a[1], ..., a[L]. Figure 1 As described on the upper left, the region mask of each object is jointly used to generate a new object representation y*. In this process, LatentDF also uses the text tokens [a1, a2, a3, ...] from the previous stage. The essence of this iterative process is to obtain y* optimized by the model. Therefore, after the first stage, the optimized text tag and fine-tuned model ∈θ* of the object y* = [a1*, ..., aN*] are obtained.
[0080] In the disclosed embodiment, a masked text method is used to learn the segmentation of multiple to-be-edited regions in a single to-be-edited image, thereby achieving continuous extended editing of different contents.
[0081] Figure 6 is a flowchart showing a method of obtaining a first text feature between an image to be edited and text information based on an image to be edited, text information and a pre-trained latent diffusion model LDM according to an exemplary embodiment, such as Figure 6 As shown, the following steps are included.
[0082] In step S41, the text information is segmented and a second text feature of the text after the segmentation is extracted.
[0083] In the disclosed embodiment, the text information input by the user is segmented to obtain a[1], a[2], a[3], a[1], a[2], a[3] are marked to obtain tokens[a1, a2, a3], and tokens[a1, a2, a3] are input into the Transformer model to obtain the text feature representation F(f), that is, the second text feature.
[0084] In step S42, random noise is added to the image to be edited, and encoding conversion is performed on the image to be edited with the random noise added, so as to obtain a third text feature.
[0085] In the disclosed embodiment, random noise N(0,1) is added to the image to be edited input by the user to obtain zt, and the obtained zt is converted into a feature representation of the latent space after being encoded by an encoder, that is, the third text feature is obtained.
[0086] In step S43, the second text feature and the third text feature are input into the pre-trained LDM to obtain the first text feature.
[0087] In the disclosed embodiment, the second text feature and the third text feature obtained above are input into a pre-trained latent diffusion model (LatentDF) to obtain the fused L(f), ie, the first text feature.
[0088] In the disclosed embodiment, the text information input by the user and the image to be edited input by the user are respectively converted into corresponding feature representations, and the feature representations of the two are fused.
[0089] In the disclosed embodiment, in response to the user selecting a target area in the image to be edited, the target area is used as the area to be edited. Alternatively, in response to the user not selecting a target area in the image to be edited, the entire area of the image to be edited is used as the area to be edited.
[0090] In the disclosed embodiment, the region to be edited can be a target region selected by the user in the image to be edited, or the user does not select the region to be edited, that is, the entire image to be edited is used as the region to be edited. Compared with directly using the image to be edited, randomly selecting several locations in the image to be edited can make the style and details of the generated target image closer to the original image, and the effect is better.
[0091] Figure 7 is a flow chart showing a method of determining a region to be edited in an image to be edited according to an exemplary embodiment. Figure 7 As shown, the following steps are included.
[0092] In step S51, random noise is added to the image to be edited, and based on the pre-trained LDM, the image to be edited with the random noise added is iteratively processed to obtain a fourth text feature.
[0093] In the disclosed embodiment, random noise is added to the image to be edited to obtain zt, which is then input into the pre-trained latent diffusion model LatentDF for iterative loop processing to obtain the generated L(f), i.e., the fourth text feature. It only contains the features of the image to be edited, and its purpose is to enhance the style of the image.
[0094] In step S52, the second text information mark corresponding to the fourth text feature is merged with the first text information mark.
[0095] In the disclosed embodiment, the fourth text feature obtained through the iterative cycle, i.e., the second text information mark [a1**, a2**, a3**] corresponding to the enhanced feature of the image to be edited (the style of the image to be edited) is fused with the first text information mark y*=[a1*, a2*, a3*] obtained in the first stage to obtain the fused F(f**).
[0096] In step S53, the fine-tuned latent diffusion model is called to iteratively denoise the fused text tags to obtain a target image.
[0097] In the disclosed embodiment, a fine-tuned latent diffusion model is called to iteratively denoise the fused text mark F(f**) to obtain a target image.
[0098] In the disclosed embodiment, the first text information tag is merged with the second text information tag containing only the image to be edited, so that the style of the original image can be enhanced, that is, the detail processing of the generated image will be closer to the image to be edited, so that the generated target image is more in line with user expectations.
[0099] Figure 8 is a flowchart of an image editing method according to an exemplary embodiment. Figure 8 As shown, the following steps are included.
[0100] In step S61, the attention channel value of the text information tag is determined based on the key projection and query projection of the multi-head attention mechanism between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model.
[0101] In the disclosed embodiment, for fine-tuning the latent diffusion model, the main focus is on the key and value projections in the cross-attention layer of the denoising network to determine the attention channel values of the text information tags.
[0102] In step S62, the attention channel value is multiplied by the image to be edited to obtain the target image.
[0103] In the disclosed embodiment, based on the obtained first text information tag, and in order to avoid overfitting, a second text information tag that increases the style of the image to be edited is further obtained, and the previous preservation loss is applied. In the disclosed embodiment, a training-free method is proposed to control the layout to avoid dataset collection. The cross attention in the denoising network of the text-to-image diffusion model can reflect the position of each generated object specified by the corresponding text tag, and its calculation formula is:
[0104] Atten = δ(Q l (zt)K l (y) T )
[0105] Where Atten is the cross attention of the denoising network layer l, Q l , K l are the query and key projections, is the softmax operation along the y dimension of the text embedding, and zt is the latent intermediate feature of the image. The size of the calculated attention Al is hl×wl×d, where hl and wl are the spatial dimensions of the feature zt, and d is the length of the input text token. More specifically, for each text token, we can obtain an attention map of size hl×wl, which reflects the relevance to its concept. For example, in the attention map with the text "cat", the position within the region containing the cat should have a larger value than other positions. Therefore, zt can be optimized to achieve the goal of having a larger value in the target region. First, the layer with a resolution of 16×16 contains the most meaningful semantic information. Therefore, we choose l as the layer with hl=wl=16, and for each layer l, each channel of the cross attention Atten represents the spatial relevance to the corresponding text token. For example, if you want to optimize the position of the object represented by v2*, you can extract the corresponding channel Atten2 and multiply it with the target region mask of v2*.
[0106] In the disclosed embodiment, the multi-head attention mechanism is applied between the Transformers of the fine-tuned latent diffusion model, and only the query and key parts of the multi-head attention mechanism are used. In addition, based on the subsequent multiplication with the image that has not been processed by the multi-head attention mechanism, the amount of calculation will be greatly reduced, so that the detail generation effect of the more refined image is better.
[0107] An exemplary embodiment of the present disclosure is described by taking the case where the text information input by the user is divided into three word segments and the input image to be edited is divided into three groups of objects as an example:
[0108] The disclosed embodiment is mainly divided into two stages. In the first stage, random noise N(0,1) is added to the image to be edited input by the user to obtain zt, and the obtained zt is converted into a feature representation of the latent space after being encoded by the encoder. At the same time, the text information input by the user is segmented to obtain a[1], a[2], a[3], and a[1], a[2], and a[3] are marked to obtain tokens[a1, a2, a3], and tokens[a1, a2, a3] are input into the Transformer model to obtain the text feature representation, i.e., F(f). The text feature representation F(f) and the feature representation of the latent space of zt are input into the pre-trained LatentDF to obtain L(f) after the fusion of the two. The area to be edited in the image to be edited is obtained. This process can select the entire image to be edited for multiple objects selected by the user or when the user does not select. The area of each object is represented by masks Mk1, Mk2,..., MkN, i.e., M1, M2, and M3 in the figure. Multiply L(f) by M1, M2, and M3 respectively to obtain a new object y*=[a1*, a2*, a3*]. That is, the image to be edited can capture three groups of objects M1, M2, and M3, and establish an understanding of different objects based on text tags (tokens) and the image encoder (Encoder), that is, the updated object y*=[a1*,..., aN*] can be obtained in the first stage.
[0109] That is, through mask text inversion, the concepts of multiple objects are learned from a single input image I to be edited into text tags a1, a2, ..., aN, where the region of each object is represented by a mask Mk1, Mk2, ..., MkN, and the details of the object are further understood by fine-tuning the diffusion model ∈θ and optimizing the additional text tags a[1], ..., a[L]. Figure 1As described on the upper left, the region mask of each object is used together to generate a new object representation y*. In this process, LatentDF also uses the text tokens [a1, a2, a3, ...] of the previous stage. For this iterative process, its essence is to obtain y* optimized by the model. Therefore, after the first stage, the optimized text tag and fine-tuned model ∈θ* of the object y* = [a1*, ..., aN*] are obtained. Then, in the second stage, the position of the object is determined according to the layout map IL specified by the user by rearranging the training-free optimized layout editing method:
[0110] I edit =out′-C(I,c θ (y * ),ε θ* ,I L )
[0111] Among them, Iedit is the newly edited image, C is our training target, and out' is the iterative noise based on the image to be edited and the objects contained in the image. The whole process is similar to the iterative denoising of LDM.
[0112] Then in the second stage, the above-obtained zt is input into the trainable LatentDF, and an iterative loop is performed to obtain the generated L(f), which only contains the features of the image to be edited. In order to enhance the style of the image, [a1**, a2**, a3**] is finally obtained, which is fused with y*=[a1*, a2*, a3*] to obtain F(f**), so that y*=[a1*, a2*, a3*] is closer to the style of the image to be edited.
[0113] Through the generated F(f**) and zt (which is the mean of Cross Attention in engineering), we iteratively denoise to generate the object feature y**, and send it to the trainable model LatentDF for iterative denoising until our target image is generated.
[0114] Among them, in the first stage, to rearrange the layout of the input image, it is first necessary to extract the concepts of multiple objects in a single input image to best preserve their visual features, such as shape, color, and texture. In the disclosed embodiment, masked text inversion is used to learn the concept of each individual object and embed it into a unique text tag. Then, the diffusion model is fine-tuned to better grasp the detailed texture of the learned object. The original technology on text inversion only supports learning the concept of a single object from a set of images (usually 3-5). However, in the disclosed embodiment, multiple concepts need to be learned from a single image. The latent vector z0=E(I) encoded from the input image using an autoencoder has local properties in the spatial dimension, and the performance of the encoder is similar to that of a downsampler. Therefore, the concepts of different objects can be sorted out by simply applying a spatial mask. It should be noted here that the entire potential loss is not calculated in the disclosed embodiment, but the loss is only propagated within the object area to update the corresponding text tag, which can be simply described as:
[0115]
[0116] Where Mk is the mask of the kth object (k = 1, ..., N is the index of N objects), and the input text condition y includes text tags [a1, ..., aN]. In fact, this mask can be roughly generated manually or automatically generated using CLIP segmentation. This process is repeated independently for each N object to optimize the text tag of each object. To avoid overfitting, each optimization is run for only 200 steps, which is much less than the original text inversion of 3000-5000 steps. Since a single text tag can only store limited information about an object, this may cause obvious distortion or artifacts during the sampling process. Therefore, the denoising network ∈ θ is further fine-tuned to better grasp the detailed texture of the object.
[0117] For the fine-tuning model, the main focus is on the query and key projection in the cross-attention layer of the denoising network, which is also the most effective way to achieve the fine-tuning goal. Therefore, in the present disclosed embodiment, by setting the input text condition to an optimized tag, and in order to avoid overfitting, an additional L trainable tags are further appended to the end of the text condition, and the previous preservation loss is applied. But in general, the traditional method is to rearrange the positions of the objects to edit the layout by learning the concepts of multiple objects a1*,...,aN* and fine-tuning the model ∈θ*. A direct way to control the layout is to add new layout conditions to a stable diffusion model, however, this method requires further fine-tuning using additional datasets. In contrast, a training-free method is proposed in the present disclosed embodiment to control the layout to avoid dataset collection. The cross-attention in the denoising network of the text-to-image diffusion model can reflect the position of each generated object specified by the corresponding text tag, and its calculation formula is:
[0118] Atten = δ(Q l (zt)K l (y) T )
[0119] Where Atten is the cross attention of the denoising network layer l, Ql, Kl are the query and key projections, is the softmax operation along the y dimension of the text embedding, and zt is the latent intermediate feature of the image. The size of the calculated attention Al is hl×wl×d, where hl and wl are the spatial dimensions of the feature zt, and d is the length of the input text token. More specifically, for each text token, we can obtain an attention map of size hl×wl, which reflects the relevance to its concept. For example, in the attention map with the text "cat", the position within the region containing the cat should have a larger value than other positions. Therefore, zt can be optimized to achieve the goal of having a larger value in the target region. First, the layer with a resolution of 16×16 contains the most meaningful semantic information. Therefore, we choose l as the layer with hl=wl=16, and for each layer l, each channel of the cross attention Atten represents the spatial relevance to the corresponding text token. For example, if you want to optimize the position of the object represented by v2*, you can extract the corresponding channel Atten2 and multiply it with the target region mask of v2*.
[0120] Finally, by optimizing the potential zt with the loss Loss, during training, only at large time steps, i.e., t>=0.5, it is enough to fix the layout of the generated image. Iterative optimization with maximum steps and early stopping is applied for t=1.0, 0.8, 0.6. Although the model has memorized the background in the process mentioned above, background distortion is still introduced during the layout control optimization process. In order to preserve the original background, it starts to blend with the original input image in the area without objects, with time step t>=0.7.
[0121] For model training: First, Masked text is used to optimize each token of different objects, and the optimization is performed for 200 steps with a batch size of 4. The learning rate is set to 0.001, and it takes about 35 seconds per token on a single Nvidia 4090GPU. Next, all optimized tokens are concatenated with additional rare tokens for 1000 steps of fine-tuning with a batch size of 4. The learning rate is set to 0.0001, which takes about 5 minutes on an Nvidia 4090GPU. To sample the image, LDM uses DDIM here to sample for 50 steps. For layout control, by optimizing zt, the learning rate is reduced to one-tenth, and the t value ranges from 1.0 to 0.5. In the final inference stage, it takes about 18 seconds to generate an image with a new layout on the PC side. It takes 32 seconds to generate a new image using the CPU on the mobile phone side.
[0122] In the embodiments of the present disclosure, the focus is on how to generate new objects during extended editing based on original objects.
[0123] In the embodiments of the present disclosure, the mask text inversion, shield text inversion and mask text inversion involved can be replaced and are not limited to a certain text inversion form.
[0124] In the related art, due to the sparsity and ambiguity of text descriptions, it is difficult to accurately control the layout of the generated image through the pre-trained text-to-image diffusion model, so the first consideration for its improvement is the need to fine-tune the pre-trained text-to-image model, thereby adding layout guidance as an additional condition in addition to the text. Secondly, consider it to perform local denoising on different areas of each object, and then globally fuse the results after each denoising step. Compared with calculating multiple denoising directions for each object, it will cause artifacts and discontinuities at the boundaries of objects, but the disclosed embodiment avoids the problem of local boundaries by directly denoising the entire image, and optimizes the potential layout control image to avoid gaps and discontinuities between objects.
[0125] In the disclosed embodiment, a framework for continuous layout extension editing of a single image is proposed, which retains the visual attributes of the image to be edited and realizes diffuse continuous editing of a single image. It is mainly implemented through two main modules: First, in order to retain the features of multiple objects in the image, it is necessary to decompose different objects (or targets) and embed them into separate text features using masked text inversion. As the model is trained, it is possible to use the learned concepts of different objects to regenerate images, align them with the user-specified layout, or expand them to generate new objects. Secondly, in order to retain the visual attributes of the input image, it is necessary to learn the concepts of multiple objects in a single image, and use the learned concepts of different objects to regenerate new images under different layouts.
[0126] In the disclosed embodiments, a framework for continuous layout extension editing based on a single image is implemented. Through the continuous extension editing capability of the image, it can be applied to multiple application scenarios such as tablet drawing software, image editing functions, text painting creation, style design, etc. A masked text method is also proposed to learn the splitting of multiple objects in a single image, thereby realizing continuous extension editing of different contents. It is also an optimization method that does not require training, and performs connectivity control of layout diffusion based on multiple potential diffusion models. It can realize terminal-side deployment methods, including central processing unit (CPU) and neural network processor (NPU) reasoning.
[0127] Through the present disclosure, the positions of objects in the input image can be rearranged to fit the target layout without changing their visual properties. A key component of learning objects from a single image is masked text inversion, which decomposes multiple concepts into different tags, and can enhance the input image to different sizes and angles, and even repair missing parts of the object caused by occlusion before applying masked text inversion. It can be expanded in technical fields such as continuous image editing, image restoration, and drawing. Extended editing of any image is achieved, and the original visual properties of the image are retained. Based on the image to be edited input by the user and the user's extended description of the image to be edited, a target image similar in style to the image to be edited and based on the user's description can be generated, so that the user can perform extended editing on any image, improving the user experience.
[0128] Based on the same concept, an embodiment of the present disclosure also provides an image editing device.
[0129] It is understandable that in order to realize the above functions, the image editing device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.
[0130] Fig. 9 FIG. 1 is a block diagram of an image editing device 100 according to an exemplary embodiment. Fig. 9 The device 100 includes an acquisition unit 101 and an execution unit 102.
[0131] The acquisition unit 101 is used to acquire the image to be edited and the text information input by the user, where the text information is the description information that the user expects to perform extended editing on the image to be edited; the execution unit 102 is used to generate a target image based on the image to be edited and the text information, where the target image includes an extended area after the image to be edited is extended and edited, and the extended area satisfies a similarity condition with the area to be edited in the image to be edited.
[0132] In one embodiment, the execution unit 102 generates a target image based on the image to be edited and the text information in the following manner: determining the area to be edited in the image to be edited; obtaining a first text information marker based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM, wherein the first text information marker is used to mark an extended area that has a similarity condition with the area to be edited; calling a fine-tuned latent diffusion model to iteratively denoise the first text information marker to obtain a target image.
[0133] In one embodiment, the execution unit 102 obtains a first text information marker based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM in the following manner: based on the image to be edited, the text information, and the pre-trained latent diffusion model LDM, a first text feature between the image to be edited and the text information is obtained; mask text inversion is performed on the area to be edited in the image to be edited, and the mask text is fused with the first text feature to obtain a first text information marker.
[0134] In one embodiment, the execution unit 102 obtains a first text feature between the image to be edited and the text information based on the image to be edited, the text information and a pre-trained latent diffusion model LDM in the following manner: performing word segmentation on the text information and extracting a second text feature of the text after the word segmentation; adding random noise to the image to be edited and performing encoding conversion on the image to be edited with the random noise added to obtain a third text feature; inputting the second text feature and the third text feature into the pre-trained LDM to obtain a first text feature.
[0135] In one embodiment, the execution unit 102 determines the area to be edited in the image to be edited in the following manner: in response to the user selecting a target area in the image to be edited, the target area is used as the area to be edited; or in response to the user not selecting a target area in the image to be edited, the entire area of the image to be edited is used as the area to be edited.
[0136] In one embodiment, the execution unit 102 is also used to: add random noise to the image to be edited, and based on the pre-trained LDM, iteratively loop process the image to be edited with the random noise added to obtain a fourth text feature; the execution unit calls the fine-tuned latent diffusion model in the following manner to iteratively denoise the first text information marker to obtain a target image: fuse the second text information marker corresponding to the fourth text feature with the first text information marker; call the fine-tuned latent diffusion model to iteratively denoise the fused text marker to obtain the target image.
[0137] In one embodiment, the execution unit 102 calls the fine-tuned latent diffusion model to iteratively denoise the first text information tag to obtain a target image: between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model, based on the key projection and query projection of the multi-head attention mechanism, the attention channel value of the text information tag is determined; the attention channel value is multiplied by the image to be edited to obtain the target image.
[0138] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0139] Fig.10 1 is a block diagram of a device 200 for image editing according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0140] Reference Fig.10, the device 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .
[0141] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.
[0142] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0143] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.
[0144] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0145] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), and when the device 200 is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 204 or sent via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.
[0146] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0147] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor assembly 214 can also detect the position change of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200 and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.
[0148] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0149] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0150] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the instructions can be executed by the processor 220 of the device 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0151] Fig.11 is a block diagram of a device 300 for image editing according to an exemplary embodiment. For example, the device 300 may be provided as a server. Fig.11 , the apparatus 300 includes a processing component 322, which further includes one or more processors, and a memory resource represented by a memory 332 for storing instructions, such as an application, that can be executed by the processing component 322. The application stored in the memory 332 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 322 is configured to execute the instructions to perform the above method.
[0152] The device 300 may also include a power supply component 326 configured to perform power management of the device 300, a wired or wireless network interface 350 configured to connect the device 300 to a network, and an input / output (I / O) interface 358. The device 300 may operate based on an operating system stored in the memory 332, such as Windows Server™, Mac OSX™, Unix™, Linux™, FreeBSD™, or the like.
[0153] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 332 including instructions, which can be executed by the processing component 322 of the device 300 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0154] It is to be understood that in the present disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings.
[0155] It is further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or degree of importance. In fact, the expressions "first", "second", etc. can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0156] It will be further understood that the terms “center”, “longitudinal”, “lateral”, “front”, “back”, “up”, “down”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0157] It can be further understood that, unless otherwise specified, “connection” includes a direct connection without other components between the two, and also includes an indirect connection with other components between the two.
[0158] It is further understood that, although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order or serial order shown, or requiring the execution of all the operations shown to obtain the desired results. In certain environments, multitasking and parallel processing may be advantageous.
[0159] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modifications, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.
[0160] It should be understood that the present disclosure is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.
Claims
1. An image editing method, characterized in that: include: Acquire an image to be edited and text information input by a user, wherein the text information is description information that the user expects to perform extended editing on the image to be edited; A target image is generated based on the image to be edited and the text information. The target image includes an extended area after the image to be edited is extended and edited. The extended area satisfies a similarity condition with the area to be edited in the image to be edited.
2. The method according to claim 1, characterized in that The step of generating a target image based on the image to be edited and the text information includes: Determining a to-be-edited area in the to-be-edited image; Based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM, a first text information tag is obtained, where the first text information tag is used to mark an extended area having a similarity condition with the area to be edited; The fine-tuned latent diffusion model is called to iteratively denoise the first text information tag to obtain a target image.
3. The method according to claim 2, characterized in that The obtaining of a first text information tag based on the image to be edited, the text information, the region to be edited, and a pre-trained latent diffusion model LDM includes: Based on the image to be edited, the text information and a pre-trained latent diffusion model LDM, obtaining a first text feature between the image to be edited and the text information; The mask text is inverted on the to-be-edited area in the to-be-edited image, and is fused with the first text feature to obtain a first text information mark.
4. The method according to claim 3, characterized in that The obtaining, based on the image to be edited, the text information and a pre-trained latent diffusion model LDM, a first text feature between the image to be edited and the text information comprises: Performing word segmentation processing on the text information, and extracting a second text feature of the text after the word segmentation processing; adding random noise to the image to be edited, and performing encoding conversion on the image to be edited with the random noise added, to obtain a third text feature; The second text feature and the third text feature are input into the pre-trained LDM to obtain the first text feature.
5. The method according to any one of claims 2 to 4, characterized in that The step of determining the area to be edited in the image to be edited comprises: In response to the user selecting a target area in the image to be edited, taking the target area as the area to be edited; or In response to the user not selecting a target area in the image to be edited, the entire area of the image to be edited is used as the area to be edited.
6. The method according to claim 2, characterized in that The method further comprises: Adding random noise to the image to be edited, and performing iterative loop processing on the image to be edited with the random noise added based on the pre-trained LDM to obtain a fourth text feature; The calling and fine-tuning the latent diffusion model to iteratively denoise the first text information mark to obtain a target image includes: Merging the second text information mark corresponding to the fourth text feature with the first text information mark; The fine-tuned latent diffusion model is called to iteratively denoise the fused text tags to obtain the target image.
7. The method according to claim 2, characterized in that The calling and fine-tuning the latent diffusion model to iteratively denoise the first text information mark to obtain a target image includes: Determining the attention channel value of the text information tag based on the key projection and query projection of the multi-head attention mechanism between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model; The attention channel value is multiplied by the image to be edited to obtain a target image.
8. An image editing device, characterized in that: include: An acquisition unit, used to acquire the image to be edited and text information input by the user, wherein the text information is description information that the user expects to perform extended editing on the image to be edited; An execution unit is used to generate a target image based on the image to be edited and the text information, wherein the target image includes an extended area after the image to be edited is extended and edited, and a similarity condition is satisfied between the extended area and the area to be edited in the image to be edited.
9. The device according to claim 8, characterized in that The execution unit generates a target image based on the image to be edited and the text information in the following manner: Determining a to-be-edited area in the to-be-edited image; Based on the image to be edited, the text information, the area to be edited, and a pre-trained latent diffusion model LDM, a first text information tag is obtained, where the first text information tag is used to mark an extended area having a similarity condition with the area to be edited; The fine-tuned latent diffusion model is called to iteratively denoise the first text information tag to obtain a target image.
10. The device according to claim 9, characterized in that The execution unit obtains a first text information tag based on the image to be edited, the text information, the region to be edited, and a pre-trained latent diffusion model LDM in the following manner: Based on the image to be edited, the text information and a pre-trained latent diffusion model LDM, obtaining a first text feature between the image to be edited and the text information; The mask text is inverted on the to-be-edited area in the to-be-edited image, and is fused with the first text feature to obtain a first text information mark.
11. The device according to claim 10, characterized in that The execution unit obtains a first text feature between the image to be edited and the text information based on the image to be edited, the text information and a pre-trained latent diffusion model LDM in the following manner: Performing word segmentation processing on the text information, and extracting a second text feature of the text after the word segmentation processing; adding random noise to the image to be edited, and performing encoding conversion on the image to be edited with the random noise added, to obtain a third text feature; The second text feature and the third text feature are input into the pre-trained LDM to obtain the first text feature.
12. The device according to any one of claims 9 to 11, characterized in that The execution unit determines the area to be edited in the image to be edited in the following manner: In response to the user selecting a target area in the image to be edited, taking the target area as the area to be edited; or In response to the user not selecting a target area in the image to be edited, the entire area of the image to be edited is used as the area to be edited.
13. The device according to claim 9, characterized in that The execution unit is also used for: Adding random noise to the image to be edited, and performing iterative loop processing on the image to be edited with the random noise added based on the pre-trained LDM to obtain a fourth text feature; The execution unit calls the fine-tuned latent diffusion model in the following manner to iteratively denoise the first text information tag to obtain a target image: Merging the second text information mark corresponding to the fourth text feature with the first text information mark; The fine-tuned latent diffusion model is called to iteratively denoise the fused text tags to obtain the target image.
14. The device according to claim 9, characterized in that The execution unit calls the fine-tuned latent diffusion model in the following manner to iteratively denoise the first text information tag to obtain a target image: Determining the attention channel value of the text information tag based on the key projection and query projection of the multi-head attention mechanism between the respective attention mechanism neural network layers in the fine-tuned latent diffusion model; The attention channel value is multiplied by the image to be edited to obtain a target image.
15. An image editing device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the image editing method described in any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute the image editing method according to any one of claims 1 to 7.
Citation Information
Cited By
Image editing processing method and device, electronic equipment and storage medium
CN120298545A