Image editing method and electronic equipment
By dividing into multiple stages in the chain generation process of diffusion network, the problem of stiff image effects and unnatural object interaction in the existing AI image editing technology is solved, and natural object interaction and harmonious tone and light and shadow effects are achieved.
Patent Information
- Application Number
- CN202311693131.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
When existing AI image editing technology performs strong editing operations, it is easy to cause the final image effect to be stiff, the interaction between objects is unnatural, and the tones and light and shadow are not harmonious.
By dividing into multiple stages in the chain generation process of the diffusion network, focusing on generating image content and object details, attributes, and light, style, and tones of the whole picture, the natural object interaction relationship and harmonious tone and light and shadow effects are achieved.
When ensuring that the image content is not changed or basically not changed, the interaction relationship between the objects in the generated final image is natural, without the sense of splicing and splitting, and the tone and light and shadow are harmonious.
Smart Images

Figure CN120147473A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing, and in particular, to an image editing method and an electronic device. Background Art
[0002] Image editing has wide applications in design creation, photography, and social media dissemination. Since traditional image editing technologies require extremely high professional technical capabilities and a large amount of creative time consumption, it is difficult to achieve popular and efficient creation; therefore, artificial intelligence (AI) image editing technologies have emerged. AI image editing technologies can use computers to automatically and efficiently implement image editing, greatly reducing the creation and application thresholds.
[0003] For example, AI image editing technologies can be used to perform editing operations on images such as adding filters, stylizing, adjusting lighting effects, adding objects, modifying objects, and deleting objects. However, using existing AI image editing technologies to perform strong editing operations such as adding objects, modifying objects, and deleting objects on images is likely to cause problems such as a rigid effect in the finally edited image, unnatural interaction between objects, and disharmonious tone and lighting. Summary of the Invention
[0004] In view of this, the present application provides an image editing method and an electronic device. This method can make the interaction relationships between various objects in the finally edited image natural, without a sense of splicing and fragmentation, and the tone and lighting are harmonious.
[0005] In a first aspect, an embodiment of the present application provides an image editing method, which includes: First, obtain a first editing task, a first image, and input text; where the first editing task is used to indicate editing of the content of the first image, and the input text includes a first text, and the first text includes text for describing a target object targeted by the first editing task; then, loop through the following steps until i is equal to H, the initial value of i is 2, and i and H are positive integers: Input the first text and the (i - 1)th edited image into a diffusion network, and perform the i-th generation process by the diffusion network to obtain the i-th first intermediate image, where the first edited image is generated according to the first editing task, the first image, the first text, and the first first intermediate image, and the first first intermediate image is generated by the diffusion network performing the first generation process based on the first text and a preset noise image; generate the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image; increment i by 1; then, input the input text and the H-th edited image into the diffusion network, and perform N generation processes by the diffusion network to obtain a second image; where N is a positive integer.
[0006] Among them, the first H generation processes can be regarded as the first stage of the chain generation process of the diffusion network, and the subsequent N generation processes can be regarded as the subsequent stages of the chain generation process of the diffusion network; that is, the present application divides the chain generation process of the diffusion network into multiple stages, and the processing operations performed in different stages are different. Correspondingly, the tasks emphasized in different stages are different: that is, the first stage emphasizes generating image content, and the subsequent stages emphasize generating details, attributes of each object in the image, and lighting, style, tone, etc. of the entire image; in this way, without changing or basically not changing the content of the first image or retaining the ID features of the specified object in the first image, the interaction relationships of each object in the finally obtained second image are natural, without a sense of splicing and fragmentation, and the tone and lighting are harmonious.
[0007] Exemplarily, the subsequent N generation processes can correspond to at least one subsequent stage, and each stage can include at least one generation process; that is, the chain generation process of the diffusion network in the present application can include at least two stages.
[0008] Exemplarily, the first editing task can also be called a strong editing task, and can include but is not limited to: adding a target object to the first image, modifying an object in the first image (for example, replacing the object in the first image with the target object), replacing the target object with an object in the first image, and deleting the target object in the first image.
[0009] For example, if the first editing task is to add a target object to the first image, the first image is a "photo of a flower", and the first text is "add a bee", then the target object for the first editing task is a "bee", and the second image is an "image with a bee added to the flower".
[0010] For example, if the first editing task is to replace the object in the first image with the target object, the first image is a "photo of a dog wearing a Christmas hat in front of a gift box and a Christmas tree", and the first text is "a cat", then the target object for the first editing task is a "cat", and the second image is an "image of a cat wearing a Christmas hat in front of a gift box and a Christmas tree".
[0011] For example, if the first editing task is to replace the target object with an object in the first image, the first image is a "photo of an orange cat", and the first text is "a black cat wearing a Christmas hat in front of a gift box and a Christmas tree", then the target object for the first editing task is a "black cat", and the second image is an "orange cat wearing a Christmas hat in front of a gift box and a Christmas tree". Among them, the orange cat in the second image retains the ID features of the orange cat in the first image.
[0012] For example, if the first editing task is to delete a target object in the first image, the first image is a "photo of a cat and a dog", and the first text is "delete a dog", then the target object of the first editing task is "dog", and the second image is a "photo of a cat".
[0013] Exemplarily, the first editing task can be determined according to the user's operation, that is, selected by the user.
[0014] Exemplarily, the first image can be an image to be edited; there are various ways to obtain the first image. For example, it can be captured by the camera of the terminal device, downloaded from a social platform, or obtained by taking a screenshot. This application does not limit this.
[0015] Exemplarily, the input text can be input by the user to the terminal device and includes text for describing the content of the editing task (including the first editing task). Among them, the input text can be composed of at least one of words, phrases, and sentences.
[0016] Exemplarily, the first text can also be composed of at least one of words, phrases, and sentences. For example, "add a bee", "bee", "add a butterfly", etc.
[0017] Exemplarily, the diffusion network can also be called a diffusion model. For example, the Latent Diffusion Model (LDM), the Cascade Diffusion Model (CDM), etc. This application does not limit this.
[0018] Exemplarily, in the step of inputting the input text and the H-th edited image into the diffusion network and performing N generation processes by the diffusion network to obtain the second image, the input of the diffusion network at the next time node includes: the input text and the output of the diffusion network at the previous time node. That is to say, inputting the input text and the H-th edited image into the diffusion network and performing N generation processes by the diffusion network to obtain the second image can be understood as: the input text and the H-th edited image can be input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 1)-th generation process) to output the (H + 1)-th intermediate image. Then, the input text and the (H + 1)-th intermediate image are input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 2)-th generation process) to output the (H + 2)-th intermediate image; and so on, until the input text and the (H + N - 1)-th intermediate image are input into the diffusion network, and the diffusion network performs the (H + N)-th generation process to obtain the second image. That is to say, when N is equal to 1, the (H + 1)-th intermediate image is the second image; when N is equal to 2, the (H + 2)-th intermediate image is the second image; and so on, which will not be elaborated here.
[0019] Exemplarily, the size of the second image is the same as that of the first image.
[0020] Exemplarily, the target object can be any entity such as a person, an animal, a plant, a building, etc. (things that objectively exist and can be distinguished from each other), and the present application places no restrictions thereon.
[0021] According to the first aspect, the input text and the H-th edited image are input into the diffusion network, and the diffusion network performs N generation processes to obtain the second image, including: inputting the first text and the H-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image.
[0022] In this case, the input text only includes the first text. The chain generation process of the diffusion network includes two stages. The first stage includes H generation processes, and the second stage includes N generation processes. The second stage of the chain generation process of the diffusion network is used to generate details, attributes, and light, shadow, style, color tone, etc. of the entire image for each object. At this time, the image editing method of the present application can be used to execute an editing task.
[0023] According to the first aspect, or any one of the above implementation manners of the first aspect, the input text further includes a second text, and the second text includes the content of the second editing task, and the second editing task is used to instruct to perform editing on the first image other than the content.
[0024] Exemplarily, the second text can be composed of at least one of words, phrases, and sentences.
[0025] Exemplarily, the second editing task can be used to perform editing on the first image other than the content. The second editing task includes but is not limited to: adding features, adding filters, adding / modifying light and shadow effects, adjusting colors, etc. to the objects (including at least one of the target object or the original objects in the first image) in the H-th edited image. For example, the first editing task is: adding a target object to the first image, the first text is "add a bee", the second editing task is: adding features to the target object, and the second text is "yellow". Another example, the first editing task is: adding a target object to the first image, the first text is "add a bee", the second editing task is: adding a filter, and the second text is "fresh and natural".
[0026] Exemplarily, there can be one second editing task. Correspondingly, the second text can be a group (each group of the second text can be composed of at least one of words, phrases, and sentences).
[0027] In one possible way, the chain generation process of the diffusion network includes two stages. Among them, the first stage includes H generation processes, and the second stage includes N generation processes. In this case, the second stage of the chain generation process of the diffusion network is used to perform the second editing task, and generate the details and attributes of each object, as well as the lighting, style, color tone, etc. of the whole image.
[0028] In one possible way, the chain generation process of the diffusion network includes three stages; among them, the first stage includes H generation processes, the second stage includes N1 generation processes, the third stage includes N2 generation processes, N1 + N2 = N, and N1 and N2 are positive integers. In this case, the second stage of the chain generation process of the diffusion network is used to perform the second editing task, and the third stage of the chain generation process of the diffusion network is used to generate the details and attributes of each object, as well as the lighting, style, color tone, etc. of the whole image.
[0029] At this time, the image editing method of the present application can be used to perform two editing tasks.
[0030] It should be noted that whether the chain generation process of the diffusion network includes two stages or three stages, one stage or two stages of the subsequent stages are to input the first text, the second text, and the i-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image.
[0031] According to the first aspect, or any implementation manner of the above first aspect, the second text is in M groups, and each group of the second text includes the content of a second editing task, where M is an integer greater than 1; inputting the input text and the H-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image, including: inputting the first text, the first group of the second text, and the H-th edited image into the diffusion network, and the diffusion network performs R 1 generation processes to obtain the first second intermediate image; looping through the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: inputting the first text, the first group of the second text to the j-th group of the second text, and the (j - 1)-th second intermediate image into the diffusion network, and the diffusion network performs R j generation processes to obtain the j-th second intermediate image; adding 1 to j; using the M-th second intermediate image as the second image; where the number of generation processes performed in the M - 1 loops and R 1 sum up to N.
[0032] In this case, each stage from the second stage to the M + 1 stage is used to perform a second editing task, and each stage or some stages from the second stage to the M + 1 stage are used to generate the details and attributes of each object, as well as the lighting, style, color tone, etc. of the whole image.
[0033] At this time, the image editing method of the present application can be used to execute three or more editing tasks.
[0034] According to the first aspect, or any implementation manner of the above first aspect, the second text is in M groups, and each group of the second text includes the content of a second editing task, where M is an integer greater than 1; input the input text and the Hth editing image into the diffusion network, and the diffusion network performs N generation processes to obtain a second image, including: input the first text, the first group of second text, and the Hth editing image into the diffusion network, and the diffusion network performs R 1 generation processes to obtain the first second intermediate image; loop and execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of second text to the jth group of second text, and the (j - 1)th second intermediate image into the diffusion network, and the diffusion network performs R j generation processes to obtain the jth second intermediate image; increment j by 1; input the first text, the first group of second text to the Mth group of second text, and the Mth second intermediate image into the diffusion network, and the diffusion network performs P generation processes to obtain the second image; where the sum of the number of generation processes executed in M - 1 loops, R 1 and P is equal to N, and P is a positive integer.
[0035] In this case, each stage from the second stage to the (M + 1)th stage is used to execute a second editing task, and the stages after the (M + 1)th stage are used to generate details, attributes, and lighting, style, tone, etc. of each object.
[0036] At this time, the image editing method of the present application can be used to execute three or more editing tasks.
[0037] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: obtaining a first region, where the first region is a region specified by the user; inputting the first text and the (i - 1)th editing image into the diffusion network, and the diffusion network performs the ith generation process to obtain the ith first intermediate image, including: inputting the first image, the first text, the first region, and the (i - 1)th editing image into the diffusion network, and the diffusion network performs the ith generation process to obtain the ith first intermediate image. In this case, the user can customize and select the addition region of the target object, which can meet the personalized requirements of the user.
[0038] Exemplarily, when the first editing task is to add a target object to the first image, the first region may be a region specified by the user for adding the target object. When the first editing task is to replace an object in the first image with a target object, the first region may be a region of the object to be replaced specified by the user. When the first editing task is to delete a target object in the first image, the first region may be a region of the object to be deleted specified by the user.
[0039] It should be noted that this application may also not require the user to specify a region for adding a target object, which can reduce the difficulty of the user editing the image.
[0040] According to the first aspect, or any implementation manner of the above first aspect, the first image, the first text, the first region, and the (i - 1)-th edited image are input into a diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image, including: generating a first mask map according to the first region; wherein, the size of the first mask map is the same as the size of the first image, the pixel values of the pixel points in the first region in the first mask map are 0, and the pixel values of the pixel points in other regions in the first mask map are 1; performing mask processing on the first image pair by using the first mask map to obtain the masked first image; wherein, the pixel values of the pixel points in the first region in the masked first image are 0; inputting the first mask map, the masked first image, the first text, and the (i - 1)-th edited image into the diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image.
[0041] It should be understood that in a possible case, the pixel values of the pixel points in the first region in the first mask map are 1, and the pixel values of the pixel points in other regions in the first mask map are 0; then, after taking the difference between 1 and the first mask map, the difference between 1 and the first mask map is used to perform mask processing on the first image pair to obtain the masked first image; this application does not limit this.
[0042] Exemplarily, performing mask processing on the first image pair by using the first mask map to obtain the masked first image may specifically be multiplying the first mask map by the first image to obtain the masked first image.
[0043] According to the first aspect, or any implementation of the above first aspect, generating the i-th edited image based on the first editing task, the first image, the first text, and the i-th first intermediate image includes: generating the i-th second mask image according to the first text; wherein, the pixel values of the pixel points in the second region of the i-th second mask image are 1, and the pixel values of the pixel points in other regions of the i-th second mask image are 0, and the second region is the region for adding the target object; fusing the first image and the i-th first intermediate image according to the i-th second mask image to obtain the i-th edited image.
[0044] It should be understood that in a possible case, the pixel values of the pixel points in the second region of the second mask image are 0, and the pixel values of the pixel points in other regions of the i-th second mask image are 1.
[0045] It should be noted that relative to the first region, the second region is a more refined region for adding the target object. The second region can be adaptively determined by the system.
[0046] It should be noted that when i is different, the i-th second mask image is also different; as the value of i continuously increases, the second region in the i-th second mask image is also more refined.
[0047] According to the first aspect, or any implementation of the above first aspect, fusing the first image and the i-th first intermediate image according to the second mask image to obtain the i-th edited image includes: performing the i-th noise addition process on the first image to obtain the i-th first image after noise addition; fusing the i-th first image after noise addition and the i-th first intermediate image according to the i-th second mask image to obtain the i-th edited image.
[0048] Exemplarily, based on a sampler (which can be understood as a forward denoising algorithm), the i-th reverse denoising can be performed on the first image to obtain the i-th first image after noise addition.
[0049] It should be noted that the process of performing the i-th reverse denoising on the first image and the process of performing the i-th generation process (that is, the i-th forward denoising) by the diffusion network are inverse processes to each other.
[0050] Exemplarily, the sampler can be, for example, Denoising Diffusion Implicit Models (DDIM), Denoising Diffusion Probabilistic Models (DDPM), etc., and the present application does not limit this.
[0051] According to the first aspect, or any implementation of the above first aspect, fusing the i-th noise-added first image and the i-th first intermediate image according to the i-th second mask image to obtain the i-th edited image includes: generating the i-th third mask image according to the i-th second mask image, where the pixel values of the pixel points in the second region of the i-th third mask image are 0, and the pixel values of the pixel points in other regions of the i-th third mask image are 1; multiplying the i-th second mask image by the i-th first intermediate image to obtain the i-th first fused image; multiplying the i-th third mask image by the i-th noise-added first image to obtain the i-th second fused image; adding the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
[0052] This way of fusing the first image and the i-th first intermediate image can be applied to the scenario where the first editing task is to add a target object to the first image.
[0053] That is to say, selecting the content of the regions other than the second region in the first image and the content of the second region in the first intermediate image for fusion can ensure that the content of the regions other than the second region in the first image is not damaged.
[0054] Exemplarily, when the pixel values of the pixel points in the second region of the second mask image are 0 and the pixel values of the pixel points in other regions of the i-th second mask image are 1, the pixel values of the pixel points in the second region of the third mask image are 1, and the pixel values of the pixel points in other regions of the i-th second mask image are 0; at this time, the i-th third mask image can be multiplied by the i-th first intermediate image to obtain the i-th first fused image; multiplying the i-th second mask image by the i-th noise-added first image to obtain the i-th second fused image.
[0055] Exemplarily, when the first editing task is to replace a target object with an object in the first image, fusing the i-th noise-added first image and the i-th first intermediate image according to the i-th second mask image to obtain the i-th edited image includes: obtaining the transformation relationship between the pixels of the first image and the pixels in the i-th first intermediate image according to the i-th second mask image and the first image; projecting the pixels in the second region of the i-th second mask image onto the i-th first intermediate image according to the transformation relationship; setting the pixel values of the pixels obtained by projection in the i-th first intermediate image to the pixel values of the corresponding position pixels in the i-th noise-added first image.
[0056] Exemplarily, when the first editing task is to replace the object in the first image with a target object, the i-th edited image is obtained by fusing the i-th noise-added first image and the i-th first intermediate image according to the i-th second mask image, including: obtaining the transformation relationship between the pixels of the first image and the pixels in the i-th first intermediate image according to the i-th second mask image and the first image; projecting the pixels in the second region of the i-th second mask image onto the i-th noise-added first image according to the transformation relationship; and setting the pixel value of the projected pixels in the i-th noise-added first image to the pixel value of the corresponding position pixels in the i-th first intermediate image.
[0057] According to the first aspect, or any one of the implementation manners of the above first aspect, the diffusion network includes a cross-attention module. The i-th intermediate feature map output by the g-th network layer in the cross-attention module includes: the weight of each word in the first text relative to each pixel point in the feature map input to the cross-attention module; generating the i-th second mask image according to the first text, including: obtaining the i-th first attention feature map according to the i-th intermediate feature map; wherein, the i-th first attention feature map includes the weight of the keyword of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module; performing smoothing processing on the i-th first attention feature map to obtain the i-th second attention feature map; and performing binarization processing on the i-th second attention feature map to obtain the i-th second mask image.
[0058] Exemplarily, the size of the i-th intermediate feature map is less than or equal to the size of the first image.
[0059] Exemplarily, the size of the i-th second mask image is the same as the size of the i-th intermediate feature map. When the size of the i-th intermediate feature map is less than the size of the first image, the i-th second mask image can be scaled so that the size of the i-th second mask image is the same as the size of the first image. Then, the i-th third mask image is determined by using the scaled i-th second mask image; and the scaled i-th second mask image is multiplied by the i-th first intermediate image to obtain the i-th first fused image.
[0060] Wherein, the feature map input to the cross-attention module is obtained by processing the first image through other modules of the diffusion network, and the spatial relationship between the first image and the feature map input to the cross-attention module is basically corresponding.
[0061] In this way, cross-attention module performs cross-modal alignment between text and image, automatically locates the accurate and reasonable position (i.e., the second region) for adding (or replacing) the target object through the keywords of the target object in the first text (such as bees, butterflies, etc.), making the relationship between the target object and the original object in the first image more natural.
[0062] Exemplarily, there are various ways of smoothing processing. For example, softmax function can be used for smoothing processing, and the present application does not limit this.
[0063] Exemplarily, the cross-attention module may include a softmax layer, and the g-th network layer may refer to the softmax layer (the softmax layer of the penultimate layer).
[0064] In a second aspect, an image editing device provided by an embodiment of the present application includes: an acquisition module, configured to acquire a first editing task, a first image, and input text; wherein, the first editing task is used to indicate editing the content of the first image, and the input text includes a first text, and the first text includes text for describing a target object targeted by the first editing task; an image generation module, configured to repeatedly execute the following steps until i is equal to H, the initial value of i is 2, and i and H are positive integers: input the first text and the (i - 1)-th edited image into a diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image, wherein the first edited image is generated according to the first editing task, the first image, the first text, and the first first intermediate image, and the first first intermediate image is generated by the diffusion network performing the first generation process based on the first text and a preset noise image; generate the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image; increment i by 1; when i is equal to H, input the input text and the H-th edited image into the diffusion network, and the diffusion network performs N times of generation processes to obtain a second image; wherein, N is a positive integer.
[0065] According to the second aspect, the second text is in M groups, each group of second text includes the content of a second editing task, and M is an integer greater than 1; the image generation module is specifically configured to input the first text, the first group of second text, and the H-th edited image into the diffusion network, and the diffusion network performs R 1 times of generation processes to obtain the first second intermediate image; repeatedly execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of second text to the j-th group of second text, and the (j - 1)-th second intermediate image into the diffusion network, and the diffusion network performs R j times of generation processes to obtain the j-th second intermediate image; increment j by 1; use the M-th second intermediate image as the second image; wherein, the number of generation processes performed in the M - 1 times of loops is the same as R 1The sum is equal to N.
[0066] According to a second aspect, or any implementation manner of the above second aspect, the second text is in M groups, and each group of the second text includes the content of a second editing task, where M is an integer greater than 1; the image generation module is specifically configured to input the first text, the first group of the second text, and the Hth edited image into a diffusion network, and the diffusion network performs R 1 generative processes to obtain the first second intermediate image; loop and execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of the second text to the jth group of the second text, and the (j - 1)th second intermediate image into the diffusion network, and the diffusion network performs R j generative processes to obtain the jth second intermediate image; increment j by 1; input the first text, the first group of the second text to the Mth group of the second text, and the Mth second intermediate image into the diffusion network, and the diffusion network performs P generative processes to obtain a second image; where the number of generative processes executed in the M - 1 loops, R 1 and P sum up to N, and P is a positive integer.
[0067] According to a second aspect, or any implementation manner of the above second aspect, the image generation module is specifically configured to generate an ith second mask image according to the first text; where the pixel values of the pixel points in the second region of the ith second mask image are 1, and the pixel values of the pixel points in other regions of the ith second mask image are 0, and the second region is the region for adding a target object; fuse the first image and the ith first intermediate image according to the ith second mask image to obtain the ith edited image.
[0068] According to a second aspect, or any implementation manner of the above second aspect, the image generation module is specifically configured to perform an ith noise addition process on the first image to obtain an ith noise-added first image; fuse the ith noise-added first image and the ith first intermediate image according to the ith second mask image to obtain the ith edited image.
[0069] According to a second aspect, or any implementation manner of the above second aspect, the image generation module is specifically configured to generate an ith third mask image according to the ith second mask image, where the pixel values of the pixel points in the second region of the ith third mask image are 0, and the pixel values of the pixel points in other regions of the ith third mask image are 1; multiply the ith second mask image by the ith first intermediate image to obtain an ith first fused image; multiply the ith third mask image by the ith noise-added first image to obtain an ith second fused image; add the ith first fused image and the ith second fused image to obtain the ith edited image.
[0070] According to a second aspect, or any implementation manner of the above second aspect, the diffusion network includes a cross-attention module. The i-th intermediate feature map output by the g-th network layer in the cross-attention module includes: the weights of each word in the first text relative to each pixel point in the feature map input to the cross-attention module; an image generation module, specifically configured to obtain the i-th first attention feature map according to the i-th intermediate feature map; wherein, the i-th first attention feature map includes the weights of the keyword of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module; perform smoothing processing on the i-th first attention feature map to obtain the i-th second attention feature map; perform binarization processing on the i-th second attention feature map to obtain the i-th second mask map.
[0071] According to a second aspect, or any implementation manner of the above second aspect, the first editing task includes any one of the following: adding a target object to the first image, replacing the object in the first image with the target object, replacing the target object with the object in the first image, or deleting the target object in the first image.
[0072] It should be understood that the above image editing device can be used to execute the image editing method in the first aspect or any possible implementation manner of the first aspect.
[0073] The second aspect and any implementation manner of the second aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the second aspect and any implementation manner of the second aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0074] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, the memory is coupled to the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the image editing method in the first aspect or any possible implementation manner of the first aspect.
[0075] The third aspect and any implementation manner of the third aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the third aspect and any implementation manner of the third aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0076] Fourth aspect, an embodiment of the present application provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the image editing method in the first aspect or any possible implementation manner of the first aspect are executed.
[0077] The fourth aspect and any implementation manner of the fourth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fourth aspect and any implementation manner of the fourth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0078] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, the computer or the processor is caused to execute the image editing method in the first aspect or any possible implementation manner of the first aspect.
[0079] The fifth aspect and any implementation manner of the fifth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fifth aspect and any implementation manner of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0080] Sixth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer or the processor is caused to execute the image editing method in the first aspect or any possible implementation manner of the first aspect.
[0081] The sixth aspect and any implementation manner of the sixth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the sixth aspect and any implementation manner of the sixth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here. Description of the Drawings
[0082] Figure 1A It is a schematic diagram of a mobile phone interface shown by way of example;
[0083] Figure 1B It is a schematic diagram of a mobile phone interface shown by way of example;
[0084] Figure 1C It is a schematic diagram of a mobile phone interface shown by way of example;
[0085] Figure 1DSchematic diagram of a mobile phone interface shown exemplarily;
[0086] Figure 1E Schematic diagram of a mobile phone interface shown exemplarily;
[0087] Figure 1F Schematic diagram of a mobile phone interface shown exemplarily;
[0088] Figure 1G Schematic diagram of a mobile phone interface shown exemplarily;
[0089] Figure 1H Schematic diagram of a mobile phone interface shown exemplarily;
[0090] Figure 1I Schematic diagram of a mobile phone interface shown exemplarily;
[0091] Figure 1J Schematic diagram of a mobile phone interface shown exemplarily;
[0092] Figure 1K Schematic diagram of a mobile phone interface shown exemplarily;
[0093] Figure 1L Schematic diagram of a mobile phone interface shown exemplarily;
[0094] Figure 2 Schematic diagram of a system framework shown exemplarily;
[0095] Figure 3A Schematic diagram of the editing process of an image shown exemplarily;
[0096] Figure 3B Schematic diagram of the editing process of an image shown exemplarily;
[0097] Figure 3C Schematic diagram of the editing process of an image shown exemplarily;
[0098] Figure 4A Schematic diagram of the editing process of an image shown exemplarily;
[0099] Figure 4B Schematic diagram of the editing process of an image shown exemplarily;
[0100] Figure 5 Schematic diagram of the editing process of an image shown exemplarily;
[0101] Figure 6A Comparison diagram of editing effects shown exemplarily;
[0102] Figure 6B Comparison diagram of editing effects shown exemplarily;
[0103] Figure 6CExemplary comparison diagram of editing effects;
[0104] Figure 7A Schematic diagram of the editing process of an exemplary image;
[0105] Figure 7B Schematic diagram of the fusion process shown exemplarily;
[0106] Figure 8A Schematic diagram of the editing process of an exemplary image;
[0107] Figure 8B Schematic diagram of the fusion process shown exemplarily;
[0108] Figure 9 Schematic diagram of an exemplary image editing device;
[0109] Figure 10 Schematic diagram of the structure of an exemplary device. Detailed implementation manners
[0110] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0111] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations.
[0112] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of the target objects.
[0113] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0114] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.
[0115] Figures 1A to 1L It is a schematic diagram of a mobile phone interface shown exemplarily.
[0116] Referring to Figure 1A , exemplarily, 101 is the main interface of the mobile phone. The main interface of the mobile phone includes one or more controls, including but not limited to: application icons (for example, the application icon of the Huawei Video application, the application icon of the browser application, the application icon 102 of the camera application), network identifier, battery identifier, etc.
[0117] Continuing to refer to Figure 1A , in a possible situation, when the user clicks on the application icon 102 of the camera application, the mobile phone responds to the user's operation behavior and displays the application interface 103 of the camera application, as Figure 1B shown. Exemplarily, the application interface 103 of the camera application includes one or more controls, including but not limited to: camera switching option, flash option, AI shooting option, settings option, night scene option, portrait option, photo-taking option, photo-taking confirmation option 104, and image preview option 105, etc. In addition, the application interface 103 of the camera application may further include a preview frame, and the preview frame displays the image captured by the camera.
[0118] Referring to Figure 1B , exemplarily, when the user wants to take a photo, they can click on the photo-taking confirmation option 104. The mobile phone responds to the user's operation behavior, instructs the camera to take a photo, and displays a thumbnail of the image captured by the camera in the image preview option 105, as Figure 1C shown.
[0119] Referring to Figure 1C , exemplarily, when the user needs to view or edit the captured image, they can click on the image preview option 105. The mobile phone responds to the user's operation behavior and displays the image preview interface 106, as Figure 1D shown. Exemplarily, the image preview interface 106 may include one or more controls, including but not limited to: share option, favorite option, cloud camera option 107, delete option, etc.
[0120] It should be noted that after the user clicks on an image from the mobile phone's gallery (or album), the mobile phone can also respond to the user's operation behavior and display the image preview interface 106.
[0121] Referring to Figure 1D, Exemplarily, when the user needs to edit an image, the user can click on the cloud camera option 107; in response to the user's operation, the mobile phone displays the cloud camera main interface 108, as Figure 1E shown. Exemplarily, the cloud camera main interface 108 may include one or more controls, including but not limited to: the add object option 109, the delete object option, the modify object option, etc.
[0122] Referring to Figure 1E , Exemplarily, when the user needs to add an object to the image, the user can click on the add object option 109; when the user needs to delete an object in the image, the user can click on the delete object option; when the user needs to replace an object in the first image with a target object, the user can click on the modify object option. It should be understood that the cloud camera main interface 108 may also include a replace option ( Figure 1E not shown), when the user needs to replace a target object with an object in the first image, the user can click on the replace option. The following takes the user clicking on the add object option 109 as an example to illustrate the interaction process between the mobile phone and the user.
[0123] Continuing to refer to Figure 1E , Exemplarily, the user clicks on the add object option 109, and in response to the user's operation, the mobile phone displays the add object interface 110, as Figure 1F or Figure 1G shown. Exemplarily, the add object interface 110 may include one or more controls, including but not limited to: an edit box, a confirm option, a cancel option, a finish option, etc. Exemplarily, the user can enter text in the edit box of the add object interface 110 to describe the target object, such as "Add a bee" as Figure 1F shown, or, "Bee, yellow, fresh" as Figure 1G shown.
[0124] Referring to Figure 1F , Exemplarily, the user clicks on the confirm option, and in response to the user's operation, the mobile phone executes the image editing method involved in the present application, generates an edited image and displays it as Figure 1H shown (or, the mobile phone sends an edit request to the server, and the server executes the image editing method involved in the present application, generates an edited image and returns it to the mobile phone, and the mobile phone displays the edited image). When the user confirms that the current edit is completed, the user can click on the Figure 1H finish option in, and in response to the user's operation, the mobile phone displays the image preview interface 106 or the cloud camera main interface 108. When the user is not satisfied with the current edit result, the user can click on the cancel option, and the mobile phone can, in response to the user's operation, display the add object interface 110.
[0125] Referring again to Figure 1E, Exemplarily, the user clicks the add object option 109, and the mobile phone displays the add object interface 111 in response to the user's operation behavior, as Figure 1I shown. Exemplarily, the add object interface 111 may include one or more controls, including but not limited to: an edit box, an OK option, a cancel option, a finish option, a feature 1 option for adding an object, a feature Y (where Y is an integer greater than 1) option for adding an object, an add filter option, etc.
[0126] Referring to Figure 1I , Exemplarily, the user can enter text associated with the target object in the edit box of the add object interface 111, such as "Add a bee" as Figure 1J shown.
[0127] Referring to Figure 1J , Exemplarily, when the user needs to add a feature description of the target object, the user can click the feature 1 option for adding an object, and the mobile phone adds a new edit box 112 to the add object interface 111 in response to the user's operation behavior, as Figure 1K shown. Exemplarily, the user can enter text for describing the feature of the target object in the edit box 112, such as "yellow" as Figure 1K shown.
[0128] It should be understood that when the user still needs to continue adding feature descriptions of the target object, the user can click feature 2 of the add object,......, feature Y of the add object, which will not be elaborated here.
[0129] Continuing to refer to Figure 1K , Exemplarily, the user can also perform editing on the image other than the content, such as adding a filter. The user can click the add filter option, and the mobile phone adds another edit box 113 to the add object interface 111 in response to the user's operation behavior, as Figure 1L shown. Exemplarily, the user can enter text for describing the filter in the edit box 113, such as "fresh and natural" as Figure 1L shown.
[0130] It should be understood that this application does not limit whether the user clicks the feature 1 option for adding an object first or the add filter option first.
[0131] It should be understood that the cloud camera option can also be an application independent of the camera, that is, the cloud camera application. The user can use the cloud camera application to edit the image, and this application does not limit this.
[0132] Figure 2 It is a schematic diagram of the system framework shown exemplarily. Figure 2 The shown system includes a terminal device and a server.
[0133] Among them, the terminal device includes but is not limited to: mobile phones, tablet computers, laptop computers, wearable devices, etc., and this application does not limit this. The specific implementation form of the server can be a cloud server, a physical (independent) server, a cluster server, etc., and this application does not limit this.
[0134] Referring to Figure 2 , exemplarily, the user uses a mobile phone to edit an image. After the user clicks the Figure 1F OK option in it, the mobile phone responds to the user's operation behavior and sends an edit request to the server, and the server executes the image editing method involved in this application. Among them, the edit request may include inputting the text "Add a bee" and the first image (i.e., the image to be edited). After the server completes the editing of the first image (such as adding a bee to the first image) and obtains the second image (with the bee added), the second image can be returned to the mobile phone, and the mobile phone can store and display the second image.
[0135] Continuing to refer to Figure 2 , exemplarily, the user uses a laptop computer to edit an image. The laptop computer can respond to the user's operation behavior and send an edit request to the server, and the server executes the image editing method involved in this application. Among them, the edit request may include inputting the text "Add a butterfly, fresh and natural" and the first image. After the server completes the editing of the first image (such as adding a butterfly to the first image and adding a "fresh and natural" filter), and obtains the second image (with the butterfly added and the fresh and natural filter), the second image can be returned to the laptop computer, and the laptop computer can store and display the second image.
[0136] Continuing to refer to Figure 2 , exemplarily, the user uses a tablet to edit an image. The tablet can respond to the user's operation behavior and send an edit request to the server, and the server executes the image editing method involved in this application. Among them, the edit request includes inputting the text "Bee, yellow" and the first image. After the server completes the editing of the first image (such as adding a yellow bee to the first image) and obtains the second image (with the yellow bee added), the second image can be returned to the tablet, and the tablet can store and display the second image.
[0137] It should be understood that the image editing method involved in this application can also be executed by the terminal device, and this application does not limit this.
[0138] Figure 3A It is a schematic diagram of the exemplary image editing process.
[0139] S301, obtain the first editing task, the first image, and the input text.
[0140] Exemplarily, the first editing task can be used to indicate editing the content of the first image. For example, the first editing task can be at least one of adding a target object to the first image, modifying an object in the first image (such as replacing the object in the first image with a target object), replacing a target object with an object in the first image, and deleting a target object in the first image.
[0141] Exemplarily, the first editing task can be determined according to the user's operation, that is, selected by the user. For example, if the user clicks the add object option 109 in Figure 1E , the first editing task can be adding a target object to the first image; if the user clicks the delete object option in Figure 1E , the first editing task can be deleting the target object in the first image; if the user clicks the modify object option in Figure 1E , the first editing task can be replacing the object in the first image with a target object.
[0142] For example, if the first editing task is to add a target object to the first image, the first image is a "photo of a flower", and the first text is "add a bee", then the target object for the first editing task is "bee", and the second image is an "image of a flower with a bee added".
[0143] For example, if the first editing task is to replace the object in the first image with a target object, the first image is a "photo of a dog wearing a Christmas hat in front of a gift box and a Christmas tree", and the first text is "a cat", then the target object for the first editing task is "cat", and the second image is an "image of a cat wearing a Christmas hat in front of a gift box and a Christmas tree".
[0144] For example, if the first editing task is to replace a target object with an object in the first image, the first image is a "photo of an orange cat", and the first text is "a black cat wearing a Christmas hat in front of a gift box and a Christmas tree", then the target object for the first editing task is "black cat", and the second image is an "orange cat wearing a Christmas hat in front of a gift box and a Christmas tree".
[0145] For example, if the first editing task is to delete the target object in the first image, the first image is a "photo of a cat and a dog", and the first text is "delete a dog", then the target object for the first editing task is "dog", and the second image is a "photo of a cat".
[0146] Exemplarily, the first image can be an image to be edited; the first image acquisition method can include various ways. For example, it can be captured by the camera of a terminal device, downloaded from a social platform, or obtained by taking a screenshot. This application does not limit this.
[0147] Exemplarily, the input text can be input by a user to a terminal device and include text for describing the content of an editing task (including the first editing task). Among them, the input text can be composed of at least one of words, phrases, and sentences.
[0148] Exemplarily, the input text can include a first text, and the first text can include text for describing the target object targeted by the first editing task; among them, the first text can also be composed of at least one of words, phrases, and sentences. For example, "add a bee", "bee", "add a butterfly", and so on.
[0149] Next, the following S302 - S308 can be executed to implement the editing of the first image based on the first editing task and the input text.
[0150] Figure 3B It is a schematic diagram of the editing process of the exemplary image shown.
[0151] Referring to Figure 3B , exemplarily, the editing process of the image involves a diffusion network (which can also be called a diffusion model, for example, LDM, CDM, etc., and this application does not limit it) and an editing processing module.
[0152] Exemplarily, the chain - type generation process of the diffusion network can correspond to T (T is a positive integer) + 1 states, and the T + 1 states correspond to T + 1 time nodes: X T (the T - th state) corresponds to the time node t = T, X T-1 (the (T - 1)-th state) corresponds to the time node t = T - 1,......, X T-H+1 (the (T - H + 1)-th state) corresponds to the time node t = T - H + 1, X T-H (the (T - H)-th state) corresponds to the time node t = T - H, X T-H-1 (the (T - H - 1)-th state) corresponds to the time node t = T - H - 1,......, X1 (the 1 - st state) corresponds to the time node t = 1, X 0 (the 0 - th state) corresponds to the time node t = 0. The chain - type generation process of the diffusion network can include T times of generation processing. Input the current time node and the input information corresponding to the current time node into the diffusion network, and the diffusion network performs one - time generation processing to obtain the state corresponding to the next time node; that is to say, each generation processing is located between two adjacent time nodes.
[0153] Exemplarily, the specific operations performed by the editing processing module correspond to the first editing task.
[0154] Exemplarily, the present application can implement the editing of the first image based on a diffusion network. Specifically, in the chain generation process of the diffusion network, different tasks are emphasized in different stages. Among them, in the first stage of the chain generation process of the diffusion network (from time node t = T to time node t = T - H), the objects in the second image can be generated (which can include: some original objects in the first image (the corresponding first editing task is: deleting the target object in the first image), or some original objects and the target object in the first image (the corresponding first editing task is: replacing the object in the first image with the target object), or all the original objects in the first image and the target object to be added to the first image (the corresponding first editing task is: adding the target object to the first image), or some original objects in the first image and other objects described in the first text (the corresponding first editing task is: replacing the target object with the object in the first image)); that is, a process from scratch, which can realize the generation of image content. In the subsequent stages of the chain generation process of the diffusion network, the details, attributes of each object in the second image, and the light, shadow, style, color tone, etc. of the whole image can be generated; in this way, without changing or basically not changing the image content, the interaction relationship of each object in the finally obtained second image is natural, without the sense of splicing and splitting, and the color tone and light and shadow are harmonious.
[0155] It should be noted that the subsequent stages of the chain generation process of the diffusion network can include at least one stage, and each stage in the subsequent stages of the chain generation process of the diffusion network can include at least one generation process.
[0156] S302, input the first text and the preset noise image into the diffusion network, and perform the first generation process by the diffusion network to obtain the first first intermediate image.
[0157] Refer to Figure 3B , exemplarily, the time node t = T, the first text, and the preset noise image can be input into the diffusion network, and the first generation process is performed by the diffusion network to obtain the first intermediate image (for the convenience of description, the intermediate images generated by the H generation processes in the first stage are called the first intermediate images).
[0158] Exemplarily, the preset noise image can be obtained by sampling Gaussian noise, and the size of the preset noise image can be the same as the size of the first image.
[0159] S303, generate the first edited image according to the first editing task, the first image, the first text, and the first first intermediate image.
[0160] In one possible way, each first editing task corresponds to an editing processing module; in this way, the editing processing module corresponding to the first editing task can be used to generate the first editing image according to the first image, the first text, and the first intermediate image.
[0161] In one possible way, each first editing task corresponds to a sub-module in the editing processing module; in this way, the sub-module corresponding to the first editing task can be used to generate the first editing image according to the first image, the first text, and the first intermediate image.
[0162] S304, input the first text and the (i - 1)-th editing image into the diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image; the initial value of i is 2, and i is a positive integer.
[0163] After that, during the second to the H-th generation processes of the diffusion network, the input of the diffusion network includes the first text and the (i - 1)-th editing image, and the output is the i-th first intermediate image. As Figure 3B shown.
[0164] S305, generate the i-th editing image according to the first editing task, the first image, the first text, and the i-th first intermediate image.
[0165] Exemplarily, S305 can refer to the description of S303 above and will not be elaborated here.
[0166] Refer to Figure 3B , when i = 2, the time node t = T - 1, the first text, and the first editing image can be input into the diffusion network to obtain the second first intermediate image; then, the editing processing module corresponding to the first editing task generates the second editing image according to the first text, the first image, and the second first intermediate image.
[0167] Refer to Figure 3B , when i = H, the time node t = T - H + 1, the first text, and the (H - 1)-th editing image can be input into the diffusion network to obtain the H-th first intermediate image; then, the editing processing module corresponding to the first editing task generates the H-th editing image according to the first text, the first image, and the H-th first intermediate image.
[0168] S306, determine whether i is equal to H.
[0169] Exemplarily, after S305 is executed, it can be determined whether i is equal to H; if i is equal to H, the first stage of the diffusion chain generation process ends, and S308 is executed, that is, enter the subsequent stage of the diffusion chain generation process; if i is not equal to H, then S307 is executed.
[0170] S307, increment i by 1.
[0171] Exemplarily, after executing S307, it is possible to return to execute S304.
[0172] S308, input the input text and the H-th edited image into the diffusion network, and let the diffusion network perform N generation processes to obtain a second image, where N is a positive integer.
[0173] Exemplarily, in S308, the input for the next time node of the diffusion network includes: the input text and the output of the diffusion network at the previous time node.
[0174] Exemplarily, N = T - H.
[0175] Refer to Figure 3B , exemplarily, it is possible to input the time node t = T - H, the input text, and the H-th edited image into the diffusion network, and let the diffusion network perform 1 generation process (i.e., the (H + 1)-th generation process) to output the (H + 1)-th intermediate image. Then, input the time node t = T - H - 1, the input text, and the (H + 1)-th intermediate image into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 2)-th generation process) to output the (H + 2)-th intermediate image; and so on, until the time node t = 1, input the time node t = 1, the input text, and the (H + N - 1)-th intermediate image into the diffusion network, and let the diffusion network perform the (H + N)-th generation process to obtain the second image.
[0176] That is to say, when N is equal to 1, the (H + 1)-th intermediate image is the second image; when N is equal to 2, the (H + 2)-th intermediate image is the second image; and so on, which will not be elaborated here.
[0177] In this way, the present application divides the chain generation process of the diffusion network into different stages, and the tasks emphasized in different stages are different: the first stage emphasizes generating image content, and the subsequent stages emphasize generating details, attributes of each object in the image, and lighting, style, color tone, etc. of the whole image; in this way, it is possible to ensure that the interaction relationship between each object in the finally obtained second image is natural, without a sense of splicing and fragmentation, and the color tone and lighting are harmonious without changing or basically not changing the image content.
[0178] In a possible way, the input text only includes the first text. At this time, the chain generation process of the diffusion network includes two stages. The first stage corresponds to S302 - S307, and the second stage corresponds to S308; among them, the first stage includes H generation processes, and the second stage includes N generation processes. In this case, the second stage of the chain generation process of the diffusion network is used to generate details, attributes of each object, and lighting, style, color tone, etc. of the whole image.
[0179] Exemplarily, a specific implementation manner of S308 may be to input the first text and the i-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image. Exemplarily, the time node t = T - H, the first text, and the H-th edited image may be input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 1)-th generation process) to output the (H + 1)-th intermediate image. Then, the time node t = T - H - 1, the first text, and the (H + 1)-th intermediate image are input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 2)-th generation process) to output the (H + 2)-th intermediate image; and so on. Until the time node t = 1, the time node t = 1, the first text, and the (H + N - 1)-th intermediate image are input into the diffusion network, and the diffusion network performs the (H + N)-th generation process to obtain the second image.
[0180] In a possible manner, the input text includes a first text and a second text; wherein, the second text includes the content of the second editing task, and the second text may be composed of at least one of words, phrases, and sentences. Wherein, the second editing task may be used to perform editing on the first image other than the content, and the second editing task includes but is not limited to: adding features to each object (including at least one of the target object or the original object in the first image) in the H-th edited image, adding filters, adding / modifying lighting effects, adjusting colors, etc. For example, the first editing task is: adding a target object in the first image, the first text is "add a bee", the second editing task is: adding features to the target object, and the second text is "yellow". Another example, the first editing task is: adding a target object in the first image, the first text is "add a bee", the second editing task is: adding a filter, and the second text is "fresh and natural".
[0181] Exemplarily, the second editing task may be one, correspondingly, the second text may be a group (each group of the second text may be composed of at least one of words, phrases, and sentences). In a possible manner, the chain generation process of the diffusion network includes two stages, the first stage corresponds to S302 - S307, and the second stage corresponds to S308. Wherein, the first stage includes H generation processes, and the second stage includes N generation processes. In this case, the second stage of the chain generation process of the diffusion network is used to execute the second editing task, and generate details, attributes of each object, and lighting, style, color tone, etc. of the whole image.
[0182] Exemplarily, there can be one second editing task, and correspondingly, the second text can be a group. In one possible way, the chain generation process of the diffusion network includes three stages. The first stage corresponds to S302 - S307, and S308 corresponds to the second and third stages. Among them, the first stage includes H generation processes, the second stage includes N1 generation processes, the third stage includes N2 generation processes, N1 + N2 = N, and N1 and N2 are positive integers. In this case, the second stage of the chain generation process of the diffusion network is used to perform the second editing task, and the third stage of the chain generation process of the diffusion network is used to generate details, attributes of each object, and lighting, style, hue, etc. of the whole picture.
[0183] Exemplarily, whether the chain generation process of the diffusion network includes two stages or three stages, the specific implementation of S308 can be to input the first text, the second text, and the i-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image. Exemplarily, the time node t = T - H, the first text, the second text, and the H-th edited image can be input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 1)-th generation process) to output the (H + 1)-th intermediate image. Then, the time node t = T - H - 1, the first text, the second text, and the (H + 1)-th intermediate image are input into the diffusion network, and the diffusion network performs 1 generation process (i.e., the (H + 2)-th generation process) to output the (H + 2)-th intermediate image; and so on, until the time node t = 1, when the time node t = 1, the first text, the second text, and the (H + N - 1)-th intermediate image are input into the diffusion network, and the diffusion network performs the (H + N)-th generation process to obtain the second image.
[0184] It should be noted that when the chain generation process of the diffusion network includes two stages, the N generation processes in the S308 process are used to perform the second editing task and generate details, attributes of each object, and lighting, style, hue, etc. of the whole picture. When the chain generation process of the diffusion network includes three stages, the first N1 generation processes in the S308 process are used to perform the second editing task; the last N2 generation processes in the S308 process are used to generate details, attributes of each object, and lighting, style, hue, etc. of the whole picture.
[0185] Exemplarily, there can be multiple second editing tasks (represented by M, where M is an integer greater than 1), and correspondingly, the second text can be M groups.
[0186] In one case, the chain generation process of the diffusion network can include M + 1 stages. The first stage corresponds to S302 - S307, and S308 corresponds to the second stage to the (M + 1)-th stage; among them, the first stage includes H generation processes, and the j-th stage among the second stage to the (M + 1)-th stage can include R jThe secondary generation process, where j is an integer less than or equal to M. At this time, each stage from the second stage to the (M + 1)-th stage is used to perform a second editing task, and each stage or partial stage from the second stage to the (M + 1)-th stage is used to generate details, attributes, and lighting, style, color tone, etc. of each object.
[0187] In one case, the chain generation process of the diffusion network may include G (G > M) + 1 stages. The first stage corresponds to S302 to S307, and S308 corresponds to the second stage to the (G + 1)-th stage. Among them, the first stage includes H times of generation processing. The j-th stage from the second stage to the (M + 1)-th stage may include R j times of generation processing, where j is an integer less than or equal to M; the (M + 2)-th stage to the (G + 1)-th stage may include P (P is a positive integer) times of generation processing. At this time, each stage from the second stage to the (M + 1)-th stage is used to perform a second editing task, and the (M + 2)-th stage to the (G + 1)-th stage is used to generate details, attributes, and lighting, style, color tone, etc. of each object.
[0188] Figure 3C It is a schematic diagram of the editing process of the exemplary image. In Figure 3C it shows the implementation manner of S308 when the chain generation process of the diffusion network includes M + 1 stages.
[0189] Referring to Figure 3C , exemplarily, S308 may include the following steps: S3081 to S3084:
[0190] S3081, input the first text, the first group of second texts, and the H-th edited image into the diffusion network, and the diffusion network performs R 1 times of generation processing to obtain the first second intermediate image.
[0191] Exemplarily, the time node t = T - H, the first text, the first group of second texts, and the H-th edited image can be input into the diffusion network, and the diffusion network performs 1 time of generation processing (i.e., the (H + 1)-th generation processing) to obtain the (H + 1)-th intermediate image (for the sake of distinction, the intermediate images generated by each generation processing in the second stage are called third intermediate images, that is, the first third intermediate image is obtained). Then, the time node t = T - H - 1, the first text, the first group of second texts, and the first third intermediate image are input into the diffusion network, and the diffusion network performs 1 time of generation processing (i.e., the (H + 2)-th generation processing) to obtain the second third intermediate image. And so on. After the (H + R 1 - 1)-th generation processing of the diffusion network, the time node t = T - H - R 1 + 1, the first text, the first group of second texts, and the R 1- One third intermediate image is input into the diffusion network, and one generation process (i.e., the (H + R)th generation process) is performed by the diffusion network to obtain the Rth third intermediate image (for the sake of distinction, the intermediate image generated by the last generation process in the second stage is called the first second intermediate image). 1 That is, when R is equal to 1, the (H + 1)th third intermediate image is the first second intermediate image; when R is equal to 2, the (H + 2)th third intermediate image is the first second intermediate image; and so on, which will not be elaborated here. 1 After that, the following S3082 to S3084 can be looped until j is equal to G; where the initial value of j is 2 and j is a positive integer.
[0192] It should be noted that S3082 to S3084 correspond to the third stage to the (G + 1)th stage in the diffusion chain generation process. 1 S3082: Input the first text, the first group of second texts to the jth group of second texts, and the (j - 1)th second intermediate image into the diffusion network, and perform R generation processes by the diffusion network to obtain the jth second intermediate image.
[0193] Exemplarily, when j = 2, S3082 can be as follows:
[0194] Exemplarily, input the time node t = T - H - R, the first text, the first group of second texts to the second group of second texts, and the first second intermediate image into the diffusion network, and perform one generation process (i.e., the (H + R + 1)th generation process) by the diffusion network to obtain the first fourth intermediate image (for the sake of distinction, the intermediate image generated by each generation process in the third stage to the (G + 1)th stage is called the fourth intermediate image, that is, obtain the first fourth intermediate image). Among them, the first group of second texts to the second group of second texts includes 2 groups of second texts, namely the first group of second texts and the second group of second texts.
[0195] Then, input the time node t = T - H - R - 1, the first text, the first group of second texts to the second group of second texts, and the first fourth intermediate image into the diffusion network, and perform one generation process (i.e., the (H + R + 2)th generation process) by the diffusion network to obtain the second fourth intermediate image. And so on, input the time node t = T - H - R - R j That is, when R is equal to 1, the (H + 1)th third intermediate image is the first second intermediate image; when R is equal to 2, the (H + 2)th third intermediate image is the first second intermediate image; and so on, which will not be elaborated here.
[0196] Exemplarily, when j = 2, S3082 can be as follows:
[0197] Exemplarily, input the time node t = T - H - R, the first text, the first group of second texts to the second group of second texts, and the first second intermediate image into the diffusion network, and perform one generation process (i.e., the (H + R + 1)th generation process) by the diffusion network to obtain the first fourth intermediate image (for the sake of distinction, the intermediate image generated by each generation process in the third stage to the (G + 1)th stage is called the fourth intermediate image, that is, obtain the first fourth intermediate image). Among them, the first group of second texts to the second group of second texts includes 2 groups of second texts, namely the first group of second texts and the second group of second texts. 1 That is, when R is equal to 1, the (H + 1)th third intermediate image is the first second intermediate image; when R is equal to 2, the (H + 2)th third intermediate image is the first second intermediate image; and so on, which will not be elaborated here. 1 + 1 generation process) to obtain the first fourth intermediate image (for the sake of distinction, the intermediate image generated by each generation process in the third stage to the (G + 1)th stage is called the fourth intermediate image, that is, obtain the first fourth intermediate image). Among them, the first group of second texts to the second group of second texts includes 2 groups of second texts, namely the first group of second texts and the second group of second texts.
[0198] Next, input the time node t = T - H - R 1 - 1, the first text, the first group of second texts to the second group of second texts, and the first fourth intermediate image into the diffusion network, and perform one generation process (i.e., the (H + R 1 + 2 generation process) to obtain the second fourth intermediate image. And so on, input the time node t = T - H - R 1 - R 2+1, the first text, the second text from the 1st group to the 2nd group of the second text, and the R 2 -1 fourth intermediate images are input into the diffusion network and are subjected to 1 generation process (i.e., the H+R 1 +R 2 th generation process) by the diffusion network to obtain the R 2 th fourth intermediate image (i.e., the 2nd second intermediate image).
[0199] When j is equal to other values, it can be deduced by analogy and will not be elaborated here.
[0200] That is to say, when R 2 is equal to 1, the R 1 +1 fourth intermediate image is the 2nd second intermediate image; when R 2 is equal to 2, the R 1 +2 fourth intermediate images are the 2nd second intermediate image; and so on, which will not be elaborated here.
[0201] S3083, determine whether j is equal to M.
[0202] Exemplarily, determine whether j is equal to M; when j is equal to M, execute S3085; when j is less than M, S3084 can be executed.
[0203] S3084, increment j by 1.
[0204] Exemplarily, after S3084 is executed, S3082 can be executed again.
[0205] S3085, use the Mth second intermediate image as the second image.
[0206] It should be noted that when the chain generation process of the diffusion network includes G+1 stages, S308 includes the above S3081 to S3085, and S3085 can be replaced by inputting the first text, the second text from the 1st group to the Mth group of the second text, and the Mth second intermediate image into the diffusion network, and the diffusion network performs P generation processes to obtain the second image; among them, the sum of the number of generation processes executed in the M-1 loops, R 1 and P is equal to N.
[0207] Refer again to Figure 1G or Figure 1L , exemplarily, the user clicks the confirmation option, the mobile phone responds to the user's operation behavior, and the mobile phone executes the image editing method involved in the present application, generates the edited image and displays it as Figure 1Has shown (alternatively, the mobile phone sends an editing request to the server, and the server executes the image editing method involved in this application to generate an edited image and return it to the mobile phone, and the mobile phone displays the edited image). When the user confirms the completion of this editing, they can click Figure 1H the completion option in, and the mobile phone responds to the user's operation behavior and displays an image preview interface or the main interface of the cloud camera. When the user is not satisfied with the result of this editing, they can click the cancel option, and the mobile phone can respond to the user's operation behavior and display the object addition interface 110.
[0208] It should be noted that, referring to Figure 1G or Figure 1L , when the mobile phone finishes adding a bee to the image, it can display the image after adding the bee (or a thumbnail of the image after adding the bee) on the mobile phone interface; when the mobile phone finishes the image with a yellow bee added, it can display the image after adding the yellow bee (or a thumbnail of the image after adding the yellow bee) on the mobile phone interface; when the mobile phone finishes adjusting the filter of the image with a yellow bee added to a fresh style, it can display the image with a yellow bee added and the filter in a fresh style (or a thumbnail of the image with a yellow bee added and the filter in a fresh style) on the mobile phone interface. That is to say, after each stage of the diffusion network of this application is executed, it can display the image generated in this stage or a thumbnail of the image.
[0209] The following takes the first editing task of adding a target object to the first image as an example to illustrate S304 and S305 above.
[0210] Figure 4A It is a schematic diagram of the image editing process shown exemplarily. Figure 4A In the embodiment, the first editing task is to add a target object to the first image; during the Figure 4A image editing process, there is no need for the user to specify the area for adding the target object, which can reduce the difficulty of the user editing the image.
[0211] S401, obtain the first editing task, the first image, and the input text.
[0212] S402, input the first text and the preset noise image into the diffusion network, and the diffusion network performs the first generation process to obtain the first intermediate image of the first one.
[0213] S403, generate the first edited image according to the first editing task, the first image, the first text, and the first intermediate image of the first one.
[0214] Among them, the process of generating the first edited image and the subsequent generation of the i-th edited image is similar, and the description of the subsequent generation of the i-th edited image can be referred to, which will not be elaborated here.
[0215] S404. Input the first text and the (i - 1)-th edited image into the diffusion network, and perform the i-th generation process by the diffusion network to obtain the i-th first intermediate image. The initial value of i is 2, and i is a positive integer.
[0216] Exemplarily, S401 - S404 can refer to the descriptions of S301 - S304 above and will not be elaborated here.
[0217] Exemplarily, in S305 above, generating the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image may include the following steps S405 - S406:
[0218] S405. Generate the i-th second mask map according to the first text.
[0219] Exemplarily, the image editing process also involves a region localization module. Among them, the region localization module can be used to locate the precise region for adding the target object.
[0220] Exemplarily, the region localization module can generate the i-th second mask map according to the first text. In one possible way, the pixel values of the pixel points in the second region of the i-th second mask map are 1, and the pixel values of the pixel points in other regions of the i-th second mask map are 0. In one possible way, the pixel values of the pixel points in the second region of the i-th second mask map are 0, and the pixel values of the pixel points in other regions of the i-th second mask map are 1. Among them, the second region is the region for adding the target object.
[0221] Figure 4B It is a schematic diagram of the exemplary image editing process.
[0222] In Figure 4B The diffusion network may further include a Cross Attention (CA) module. Among them, the CA module includes multiple network layers. Among them, the feature output from the g-th (g is a positive integer) network layer of the CA module to the (g + 1)-th network layer (subsequently referred to as the i-th intermediate feature map) includes: the weight of each word in the first text relative to each pixel point in the feature map input to the CA module (that is, the feature map obtained by processing the first image by some network layers in the diffusion network, and the spatial relationship between the feature map input to the CA module and the first image is basically corresponding). Among them, the size of the i-th intermediate feature map is less than or equal to the size of the first image.
[0223] For example, assume that the first text is "Add a bee", then the first text includes 3 words: word 1 ("Add"), word 2 ("a"), and word 3 ("bee"); the size of the i-th intermediate feature map is 10*10, that is, the intermediate feature map includes 100 pixel points. The pixel value of any pixel point (such as pixel point A) among these 100 pixel points = the feature vector of word 1 * w1 + the feature vector of word 2 * w2 + the feature vector of word 3 * w3; where. w1 is the weight of word 1 relative to pixel point A, which can be used to represent the probability that pixel point A is the semantics described by word 1; w2 is the weight of word 2 relative to pixel point A, which can be used to represent the probability that pixel point A is the semantics described by word 2; w3 is the weight of word 3 relative to pixel point A, and the probability w3 of word 3 can be used to represent the probability that pixel point A is the semantics described by word 3. Among them, the greater the weight of a word relative to a pixel point, the greater the probability that the pixel point is the semantics described by the word.
[0224] Exemplarily, S405 may include S4051 to S4053, where S4051 to S4053 may be executed by the region localization module:
[0225] S4051, according to the i-th intermediate feature map, obtain the i-th first attention feature map.
[0226] Exemplarily, the region localization module may obtain the i-th first attention feature map according to the i-th intermediate feature map output by the g-th network layer in the CA module in the diffusion network. Among them, the size of the i-th first attention feature map is the same as the size of the i-th intermediate feature map, and the i-th first attention feature map includes the weights of the keyword of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module. The pixel value of a pixel point in the i-th first attention feature map is the weight of the keyword of the target object relative to this pixel point in the feature map input to the cross-attention module.
[0227] Among them, the keyword of the target object may refer to the name of the target object. For example, if the first text is "Add a bee", the keyword of the target object may be "bee".
[0228] S4052, perform smoothing processing on the i-th first attention feature map to obtain the i-th second attention feature map.
[0229] Exemplarily, the objective of S4052 is to reduce the difference in pixel values between two adjacent pixel points in the i-th first attention feature map.
[0230] Exemplarily, there are various ways of smoothing processing, and this application does not limit this; for example, use the softmax function to perform smoothing processing on the first attention feature map to obtain the second attention feature map.
[0231] S4053, perform binarization processing on the i-th second attention feature map to obtain the i-th second mask map.
[0232] Exemplarily, the pixel value of each pixel point in the i-th second attention feature map can be compared with a threshold to generate the i-th second mask map. For example, set the pixel value of the pixel point with a pixel value greater than the threshold to 1, and set the pixel value of the pixel point with a pixel value less than or equal to the threshold to 0; in this way, the i-th second mask map can be obtained. In this case, the area with a pixel value of 1 in the i-th second mask map can be the second area. For example, set the pixel value of the pixel point with a pixel value greater than the threshold to 0, and set the pixel value of the pixel point with a pixel value less than or equal to the threshold to 1; in this way, the i-th second mask map can be obtained. In this case, the area with a pixel value of 0 in the i-th second mask map can be the second area. This method can be used to determine the precise position for adding the target object within the entire image range.
[0233] That is to say, the second mask maps determined at each time node may be different; and as time goes by, the later the time node, the finer the second area in the determined second mask map.
[0234] In this way, the present application can locate the precise addition position of the target object in the first image, and can ensure that while adding the target object in the first image, the content of other areas in the original image except the second area is not damaged.
[0235] Exemplarily, the size of the i-th second mask map is the same as the size of the i-th intermediate feature map. When the size of the i-th intermediate feature map is smaller than the size of the first image, the i-th second mask map can be scaled so that the size of the i-th second mask map is the same as the size of the first image. The i-th second mask map in subsequent S407 - S409 is the scaled i-th second mask map.
[0236] Exemplarily, after obtaining the i-th second mask map, the first image and the i-th first intermediate image can be fused according to the i-th second mask map to obtain the i-th edited image; refer to S406 - S410 below.
[0237] Exemplarily, in Figure 4B the embodiment of, the editing processing module may include a fusion module and a noise addition module. The fusion module can be used for fusion and can execute S407 - S410, and the noise addition module can be used for noise addition processing and can be used to execute S406.
[0238] S406, perform the i-th noise addition processing on the first image to obtain the i-th first image after noise addition processing.
[0239] Exemplarily, based on a sampler (which can be understood as a forward denoising algorithm), the first image can be reversely denoised for the i-th time to obtain the i-th first image processed by denoising.
[0240] It should be noted that the process of reversely denoising the first image for the i-th time and the process of the diffusion network performing the i-th generation process (i.e., the i-th forward denoising) are inverse processes of each other.
[0241] Exemplarily, the sampler can be, for example, the Denoising Diffusion Implicit Models (DDIM), the Denoising Diffusion Probabilistic Models (DDPM), etc. The present application does not limit this.
[0242] Next, according to the i-th second mask image, the i-th first image processed by denoising and the i-th first intermediate image can be fused to obtain the i-th edited image. The following steps S407 to S410 can be referred to:
[0243] S407: Generate the i-th third mask image according to the i-th second mask image.
[0244] Exemplarily, 1 can be subtracted from the pixel value of each pixel point in the i-th second mask image to obtain the i-th third mask image.
[0245] Exemplarily, when the pixel value of the pixel points in the second region of the i-th second mask image is 1 and the pixel value of the pixel points in other regions of the i-th second mask image is 0, the pixel value of the pixel points in the second region of the i-th third mask image is 0, and the pixel value of the pixel points in other regions of the i-th second mask image is 1.
[0246] S408: Multiply the i-th second mask image by the i-th first intermediate image to obtain the i-th first fused image.
[0247] S409: Multiply the i-th third mask image by the i-th first image processed by denoising to obtain the i-th second fused image.
[0248] S410: Add the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
[0249] Exemplarily, the following formula (1) can be referred to for fusion:
[0250] z i =x i Mi +y i (1 - M i ) (1)
[0251] Among them, in formula (1), z i is the i-th edited image, M i is the i-th second mask image, (1 - Mi) is the i-th third mask image, x i is the i-th first intermediate image, y i is the i-th first image processed by adding noise.
[0252] It should be noted that when the pixel value of the pixel points in the second region of the i-th second mask image is 0 and the pixel value of the pixel points in other regions of the i-th second mask image is 1, the pixel value of the pixel points in the second region of the i-th third mask image is 1, and the pixel value of the pixel points in other regions of the i-th third mask image is 0. In this case, the i-th third mask image can be multiplied by the i-th first intermediate image to obtain the i-th first fused image; and the i-th second mask image can be multiplied by the i-th first image processed by adding noise to obtain the i-th second fused image.
[0253] S411, determine whether i is equal to H.
[0254] Exemplarily, after S410 is executed, it can be determined whether i is equal to H; if i is equal to H, then S413 is executed, that is, enter the subsequent stage of the diffusion chain generation process; if i is not equal to H, then S412 is executed.
[0255] S412, increment i by 1.
[0256] Exemplarily, after S412 is executed, S404 can be returned to for execution.
[0257] S413, input the input text and the i-th edited image into the diffusion network, and the diffusion network performs N times of generation processing to obtain a second image, where N is a positive integer.
[0258] Exemplarily, S413 can refer to the description of the above embodiments and will not be elaborated here.
[0259] Figure 5 It is a schematic diagram of the editing process of the exemplary image shown. Figure 5 In the embodiment, the first editing task is to add a target object to the first image; during the Figure 5 editing process of the image, the area for adding the target object can be specified by the user.
[0260] S501, obtain the first editing task, the first image, and the input text.
[0261] S502. Obtain a first region, which is the region specified by the user for adding a target object.
[0262] Exemplarily, the user can specify a region for adding a target object (subsequently referred to as the first region).
[0263] Exemplarily, the user can Figures 1F to 1L in any mobile phone interface, frame the first region in the first image.
[0264] It should be noted that the first region specified by the user is a rough region for adding a target object.
[0265] S503. Input the first text, a preset noise image, and the first region into a diffusion network, and the diffusion network performs the first generation process to obtain the first first intermediate image.
[0266] Exemplarily, the implementation process of S503 can refer to the implementation process of S505 described below and will not be elaborated here.
[0267] S504. Generate the first edited image according to the first editing task, the first image, the first text, and the first first intermediate image.
[0268] Exemplarily, S501, S503, and S504 can refer to the descriptions of S301 to S303 above and will not be elaborated here.
[0269] The difference between S503 and S504 and S302 and S303 is that the input of the diffusion network in S503 and S504 newly adds the first region.
[0270] S505. Input the first image, the first text, the (i - 1)-th edited image, and the first region into the diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image; the initial value of i is 2, and i is a positive integer.
[0271] Exemplarily, S505 may include the following S5051 to S5053:
[0272] S5051. Determine a first mask map according to the first region.
[0273] In a possible way, the pixel value of the pixel points in the first region of the first mask map is 0, and the pixel value of the pixel points in other regions of the first mask map is 1.
[0274] In a possible way, the pixel value of the pixel points in the first region of the first mask map is 1, and the pixel value of the pixel points in other regions of the first mask map is 0.
[0275] For S5052, the first image pair is masked using the first mask image to obtain the masked first image.
[0276] Exemplarily, when the pixel value of the pixel points in the first region of the first mask image is 0 and the pixel value of the pixel points in other regions of the first mask image is 1, the pixel value of the pixel points in the first mask image can be multiplied by the pixel value of the pixel points at the corresponding position in the first image to obtain the masked first image.
[0277] Exemplarily, when the pixel value of the pixel points in the first region of the first mask image is 1 and the pixel value of the pixel points in other regions of the first mask image is 0, the difference between the pixel value of the pixel points in the first mask image and 1 can be multiplied by the pixel value of the pixel points at the corresponding position in the first image to obtain the masked first image.
[0278] That is to say, masking can erase the content in the first region of the first image. In this way, the pixel values of the pixel points in the first region of the masked first image are all 0.
[0279] For S5053, the first mask image, the masked first image, the first text, and the (i - 1)-th edited image are input into the diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image.
[0280] Exemplarily, in the above S305, generating the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image may include the following steps S506 to S507:
[0281] For S506, the i-th second mask image is generated according to the first text.
[0282] Exemplarily, S506 may include S5061 to S5063, where S5061 to S5063 may be executed by the region localization module:
[0283] For S5061, the i-th first attention feature map is obtained according to the i-th intermediate feature map.
[0284] For S5062, the i-th first attention feature map is smoothed to obtain the i-th second attention feature map.
[0285] For S5063, the i-th second attention feature map is binarized to obtain the i-th second mask image.
[0286] In a possible way, Figure 5 S5063 of the embodiment may refer to Figure 4A the description of S4053 in the embodiment, which will not be elaborated here.
[0287] In one possible way, the pixel value of each pixel point in the first region of the i-th second attention feature map can be compared with a threshold to generate the i-th second mask map. For example, the pixel value of a pixel point with a pixel value greater than the threshold in the first region is set to 1, the pixel value of a pixel point with a pixel value less than or equal to the threshold is set to 0, and the pixel value of a pixel point outside the first region is set to 0; in this way, the i-th second mask map can be obtained. In this case, the region with a pixel value of 1 in the i-th second mask map can be the second region. For example, the pixel value of a pixel point with a pixel value greater than the threshold is set to 0, the pixel value of a pixel point with a pixel value less than or equal to the threshold is set to 1, and the pixel value of a pixel point outside the first region is set to 1; in this way, the i-th second mask map can be obtained.
[0288] Exemplarily, the size of the i-th second mask map is the same as the size of the i-th intermediate feature map. When the size of the i-th intermediate feature map is smaller than the size of the first image, the i-th second mask map can be scaled so that the size of the i-th second mask map is the same as the size of the first image. The i-th second mask map in subsequent S507 to S509 is the scaled i-th second mask map.
[0289] In this case, the region with a pixel value of 0 in the i-th second mask map can be the second region. This way can, on the premise of meeting the personalized needs of users, determine the precise position (i.e., the second region) for adding the target object within the first region.
[0290] It should be noted that, relative to the first region, the second region is a more refined region for adding the target object.
[0291] Exemplarily, the size of the i-th second mask map is the same as the size of the i-th intermediate feature map. When the size of the i-th intermediate feature map is smaller than the size of the first image, the i-th second mask map can be scaled so that the size of the i-th second mask map is the same as the size of the first image.
[0292] Then, according to the i-th second mask map, the first image and the i-th first intermediate image can be fused to obtain the i-th edited image; specifically, refer to the following steps S507 to S511:
[0293] S507, perform the i-th noise addition process on the first image to obtain the i-th noise-added first image.
[0294] After that, according to the i-th second mask image, the i-th first image after noise addition and the i-th first intermediate image can be fused to obtain the i-th edited image. Specifically, the following steps S508 to S5011 can be referred to:
[0295] S508, generate the i-th third mask image according to the i-th second mask image.
[0296] S509, multiply the i-th second mask image by the i-th first intermediate image to obtain the i-th first fused image.
[0297] S510, multiply the i-th third mask image by the i-th first image after noise addition to obtain the i-th second fused image.
[0298] S511, add the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
[0299] Exemplarily, S507 to S511 can refer to the descriptions of the above S406 to S410, and will not be elaborated here.
[0300] S512, determine whether i is equal to H.
[0301] Exemplarily, after S511 is executed, it can be determined whether i is equal to H; if i is equal to H, then S514 is executed, that is, enter the subsequent stage of the diffusion chain generation process; if i is not equal to H, then S513 is executed.
[0302] S513, increment i by 1.
[0303] Exemplarily, after S513 is executed, it can return to execute S505.
[0304] S514, input the first image, the input text, the i-th edited image, and the first region into the diffusion network, and the diffusion network performs N times of generation processing to obtain the second image, where N is a positive integer.
[0305] Exemplarily, S512 to S514 can refer to the descriptions of the above S406 to S413, and will not be elaborated here.
[0306] It should be noted that during the execution of S514, first, according to the first region, the first mask image needs to be determined; then, the first image pair is masked by the first mask image to obtain the first image after masking; after that, the first mask image, the first image after masking, the input text, and the i-th edited image are input into the diffusion network, and the diffusion network performs N times of generation processing to obtain the second image.
[0307] Figures 6A to 6CThe figure is a comparison diagram of exemplary editing effects.
[0308] Figures 6A to 6C (1), (2), and (3) are respectively: the first image, the second image generated by the method of the prior art, and the second image generated by the method of the present application. Among them, in the second image generated by the method of the prior art, except for the area where the target object is added, other areas are damaged, and it is impossible to maintain the original content of other areas except for the area where the target object is added. While in the second image generated by the method of the present application, the original content of other areas except for the area where the target object is added is maintained.
[0309] The test results obtained by testing the method of the prior art and the method of the present application with different target objects can be shown in Table 1 as follows:
[0310] Table 1
[0311] Category Prior Art This Application Bee 74% 91.16% Butterfly 48.8% 88.07% Ladybug 84.6% 97.43%
[0312] Referring to Table 1, when adding a bee to the first image by the method of the prior art, the success rate of the obtained second image is 74%; when adding a bee to the first image by the method of the present application, the success rate of the obtained second image is 91.16%; among them, the higher the success rate, the better the image effect. By comparison, compared with the prior art, the absolute improvement ratio of the present application is 17.16%, and the relative improvement ratio is 23.18%.
[0313] Referring to Table 1, when adding a butterfly to the first image by the method of the prior art, the success rate of the obtained second image is 48.8%; when adding a butterfly to the first image by the method of the present application, the success rate of the obtained second image is 88.07%; among them, the higher the success rate, the better the image effect. By comparison, compared with the prior art, the absolute improvement ratio of the present application is 39.27%, and the relative improvement ratio is 80.47%.
[0314] Referring to Table 1, when adding a ladybug to the first image by the method of the prior art, the success rate of the obtained second image is 84.6%; when adding a ladybug to the first image by the method of the present application, the success rate of the obtained second image is 97.43%; among them, the higher the success rate, the better the image effect. By comparison, compared with the prior art, the present application has an absolute improvement of 12.83% and a relative improvement ratio of 15.13%.
[0315] The following takes the first editing task as an example of replacing the target object with the object in the first image to illustrate S304 and S305 above.
[0316] Figure 7A The figure is a schematic diagram of the editing process of the image shown for exemplary purposes. Figure 7AIn the embodiment, the first editing task is to modify the object in the first image.
[0317] S701, obtain the first editing task, the first image, and the input text.
[0318] For example, the first editing task is to replace the target object with the object in the first image, the first image is "a photo of an orange cat"; the first text is "a black cat wearing a Christmas hat in front of a gift box and a Christmas tree".
[0319] S702, input the first text and the preset noise image into the diffusion network, and perform the first generation process by the diffusion network to obtain the first intermediate image of the first one.
[0320] S703, generate the first edited image according to the first editing task, the first image, the first text, and the first intermediate image of the first one.
[0321] S704, input the first text and the (i - 1)th edited image into the diffusion network, and perform the ith generation process by the diffusion network to obtain the ith first intermediate image; the initial value of i is 2, and i is a positive integer.
[0322] Exemplarily, S701 to S704 can refer to the descriptions of the above S301 to S304, and will not be elaborated here.
[0323] Exemplarily, in S305 above, generating the ith edited image according to the first editing task, the first image, the first text, and the ith first intermediate image may include the following steps S705 to S711:
[0324] S705, generate the ith second mask image according to the first text.
[0325] Exemplarily, S705 may include S7051 to S7053, where S7051 to S7053 can be executed by the region localization module:
[0326] S7051, obtain the ith first attention feature map according to the ith intermediate feature map.
[0327] S7052, perform smoothing processing on the ith first attention feature map to obtain the ith second attention feature map.
[0328] S7053, perform binarization processing on the ith second attention feature map to obtain the ith second mask image.
[0329] Exemplarily, S7051 to S7053 can refer to the descriptions of the above S4051 to S4053, and will not be elaborated here.
[0330] Among them, the i-th second mask image in subsequent S707 to S708 is the scaled i-th second mask image.
[0331] Next, according to the i-th second mask image and the first image, the i-th first intermediate image can be fused to obtain the i-th edited image; reference can be made to the descriptions in S706 to S711 below.
[0332] Exemplarily, according to the i-th second mask image and the first image, the transformation relationship between the pixels of the first image and the pixels in the i-th first intermediate image can be obtained; specifically, reference can be made to S706 to S708:
[0333] S706, determine the segmentation map of the object to be replaced in the first image and the detection frame of the object to be replaced in the first image.
[0334] Figure 7B For example, a schematic diagram of the fusion process is shown. Among them, Figure 7B is a single fusion process shown on the basis of Figure 4B .
[0335] Referring to Figure 7B , exemplarily, the editing processing module may include an open-domain segmentation module, an open-domain detection module, and a post-processing module; among them, the open-domain segmentation module can perform segmentation processing on the first image to obtain the segmentation map of the object to be replaced in the first image (which can be represented by Mask_source); the open-domain detection module can perform detection processing on the first image to obtain the detection frame of the object to be replaced in the first image (which can be represented by Box_source).
[0336] It should be noted that the open-domain detection module in the editing processing module can also be replaced by a post-processing module, and the post-processing module can perform post-processing on the segmentation map Mask_source of the object to be replaced in the first image to obtain the detection frame Box_source of the object to be replaced in the first image. That is to say, the open-domain detection module is an optional module.
[0337] In a possible way, the pixel values of the pixels in the area where the object to be replaced is located (subsequently referred to as the third area) in the segmentation map Mask_source of the object to be replaced in the first image are 1, and the pixel values of the pixels in other areas of the segmentation map Mask_source of the object to be replaced in the first image are 0; as Figure 7B shown.
[0338] In a possible way, the pixel value of the pixels in the third region of the segmentation map Mask_source of the object to be replaced in the first image is 0, and the pixel value of the pixels in other regions of the segmentation map Mask_source of the object to be replaced in the first image is 1.
[0339] Exemplarily, the detection box Box_source of the object to be replaced in the first image can be the smallest box containing the object to be replaced in the first image.
[0340] S707, determine the detection box of the i-th second mask map.
[0341] Exemplarily, the post-processing module can also perform post-processing on the i-th second mask map to obtain the detection box of the i-th second mask map (which can be represented by Box_target).
[0342] Exemplarily, the detection box Box_target of the i-th second mask map can be the smallest box containing the second region in the i-th second mask map.
[0343] S708, determine the transformation relationship between the pixels of the first image and the pixels of the i-th first intermediate image according to the detection box of the i-th second mask map and the detection box of the object to be replaced in the first image.
[0344] Exemplarily, the editing processing module can also include a transformation module, and the transformation module can be used to analyze the four vertices of the detection box Box_target of the i-th second mask map and the four vertices of the detection box Box_source of the object to be replaced in the first image to determine the transformation relationship between the pixels of the first image and the pixels of the i-th first intermediate image (i.e., the affine transformation relationship).
[0345] Among them, the transformation relationship between the pixels of the first image and the pixels of the i-th first intermediate image can indicate the transformation relationship between the pixels of the object to be replaced in the first image and the pixels of the target object in the i-th first intermediate image.
[0346] S709, perform the i-th noise addition processing on the first image to obtain the i-th first image after noise addition processing.
[0347] Exemplarily, S709 can refer to the description of S406 above and will not be elaborated here. S709 can be executed by the noise addition module included in the editing processing module.
[0348] S710, project the pixel point P1 k =(X1 k , Y1 k ) in the third region onto the i-th first intermediate image to obtain P2k = (X2 k , Y2 k ).
[0349] Where k is a positive integer.
[0350] S711. Set the pixel value of the pixel point with coordinates (X2 k , Y2 k ) in the i-th first intermediate image to the pixel value of the pixel point with coordinates (X1 k , Y1 k ) in the i-th first image after noise addition processing, to obtain the i-th edited image.
[0351] Exemplarily, the editing processing module may further include a fusion module, and S711 may be executed by the fusion module.
[0352] S712. Determine whether i is equal to H.
[0353] Exemplarily, after S711 is executed, it can be determined whether i is equal to H; if i is equal to H, then execute S714, that is, enter the subsequent stage of the diffusion chain generation process; if i is not equal to H, then execute S713.
[0354] S713. Increment i by 1.
[0355] Exemplarily, after S713 is executed, it can return to execute S704.
[0356] On the one hand, as the number of diffusion steps gradually approaches 0 (that is, the time node t gradually approaches 0), the image noise becomes smaller and the content becomes clearer, and the automatic positioning of the area where the target object is located will be more accurate, which helps to obtain a more accurate replacement relationship, that is, the transformation relationship from the object to be replaced in the first image to the corresponding target object in the diffusion network generation graph, so as to achieve precise replacement; on the other hand, the (i - 1)-th edited image is used as the input for the i-th generation process of the diffusion network, and the information of the object to be replaced in the first image can be input in a timely manner, which helps the algorithm in the generation process to generate adjustments, making the posture and body shape of the object to be replaced in the first image and the object to be replaced in the generation graph more matching, and the interaction with other contents and objects in the generation graph more harmonious and natural. In this way, the specified object (personalized object, such as one's own cats and dogs and other personal unique items) in the user input image can be replaced into the generation graph, realizing placing the user-specified object into various scene pictures and achieving "AI photo shooting".
[0357] S714. Input the input text and the i-th edited image into the diffusion network, and the diffusion network performs N times of generation processing to obtain the second image, where N is a positive integer.
[0358] Exemplarily, S714 can refer to the description of the above embodiments and will not be elaborated here.
[0359] Taking the first editing task of replacing the object in the first image with a target object as an example, the above S304 and S305 will be described below.
[0360] Figure 8A It is a schematic diagram of the editing process of the exemplary image shown. Figure 8A In the embodiment, the first editing task is to modify the object in the first image.
[0361] S801, obtain the first editing task, the first image, and the input text.
[0362] For example, the first editing task is to replace the object in the first image with a target object. The first image can be "an image of a dog wearing a Christmas hat in front of a gift box and a Christmas tree", and the first text is "a cat".
[0363] S802, input the first text and the preset noise image into the diffusion network, and the diffusion network performs the first generation process to obtain the first intermediate image of the first one.
[0364] S803, generate the first editing image according to the first editing task, the first image, the first text, and the first intermediate image of the first one.
[0365] S804, input the first text and the (i - 1)th editing image into the diffusion network, and the diffusion network performs the ith generation process to obtain the ith first intermediate image; the initial value of i is 2, and i is a positive integer.
[0366] Exemplarily, S801 to S804 can refer to the description of the above S301 to S304 and will not be elaborated here.
[0367] Exemplarily, in the above S305, generating the ith editing image according to the first editing task, the first image, the first text, and the ith first intermediate image may include the following steps S805 to S811:
[0368] S805, generate the ith second mask image according to the first text.
[0369] Exemplarily, S805 may include S8051 to S8053, where S8051 to S8053 may be executed by the region localization module:
[0370] S8051, obtain the ith first attention feature map according to the ith intermediate feature map.
[0371] S8052, perform smoothing processing on the ith first attention feature map to obtain the ith second attention feature map.
[0372] S8053 binarizes the i-th second attention feature map to obtain the i-th second mask map.
[0373] Exemplarily, S8051 to S8053 may refer to the description of S4051 to S4053 above, and will not be elaborated here.
[0374] Among them, the i-th second mask map in subsequent S807 to S808 and S810 is the scaled i-th second mask map.
[0375] Next, the i-th first intermediate image can be fused with the first image according to the i-th second mask map to obtain the i-th edited image; reference can be made to the description of S806 to S811 below.
[0376] Exemplarily, the transformation relationship between the pixels of the first image and the pixels in the i-th first intermediate image can be obtained according to the i-th second mask map and the first image; specifically, reference can be made to S806 to S808:
[0377] S806 determines the segmentation map of the object to be replaced in the first image and the detection frame of the object to be replaced in the first image.
[0378] Figure 8B For exemplary illustration of the fusion process diagram. Among them, Figure 8B is a fusion process shown on the basis of Figure 4B One fusion process is shown on the basis of
[0379] Refer to Figure 8B , exemplarily, the editing processing module may include an open-domain segmentation module, an open-domain detection module, and a post-processing module; among them, the open-domain segmentation module can perform segmentation processing on the first image to obtain the segmentation map of the object to be replaced in the first image (which can be represented by Mask_source); the open-domain detection module can perform detection processing on the first image to obtain the detection frame of the object to be replaced in the first image (which can be represented by Box_source).
[0380] It should be noted that the open-domain detection module in the editing processing module can also be replaced by a post-processing module, and the post-processing module can perform post-processing on the segmentation map Mask_source of the object to be replaced in the first image to obtain the detection frame Box_source of the object to be replaced in the first image. That is to say, the open-domain detection module is an optional module.
[0381] In one possible way, the pixel value of the pixels in the third region (subsequently referred to as the third region) where the object to be replaced is located in the segmentation map Mask_source of the object to be replaced in the first image is 1, and the pixel values of the pixels in other regions in the segmentation map Mask_source of the object to be replaced in the first image are 0; as shown in 8B.
[0382] In one possible way, the pixel value of the pixels in the third region in the segmentation map Mask_source of the object to be replaced in the first image is 0, and the pixel values of the pixels in other regions in the segmentation map Mask_source of the object to be replaced in the first image are 1.
[0383] Exemplarily, the detection box Box_source of the object to be replaced in the first image can be the smallest box containing the object to be replaced in the first image.
[0384] S807, determine the detection box of the i-th second mask map.
[0385] Exemplarily, the post-processing module can also perform post-processing on the i-th second mask map to obtain the detection box of the i-th second mask map (which can be represented by Box_target).
[0386] Exemplarily, the detection box Box_target of the i-th second mask map can be the smallest box containing the second region in the i-th second mask map.
[0387] S808, determine the transformation relationship between the pixels of the first image and the pixels of the i-th first intermediate image according to the detection box of the i-th second mask map and the detection box of the object to be replaced in the first image.
[0388] Exemplarily, S808 can refer to the description of S708 above and will not be elaborated here.
[0389] S809, perform the i-th noise addition process on the first image to obtain the i-th first image after noise addition processing.
[0390] Exemplarily, S809 can refer to the description of S406 above and will not be elaborated here.
[0391] S810, project the pixel point P1 k =(X1 k , Y1 k ) in the second region of the i-th second mask map onto the i-th first image after noise addition processing to obtain P2 k =(X2 k , Y2 k ).
[0392] Exemplarily, k is a positive integer.
[0393] S811. Set the pixel value of the pixel point with coordinates (X2 k , Y2 k ) in the i-th first image after noise addition as the pixel value at coordinates (X1 k , Y1 k ) in the i-th first intermediate image to obtain the i-th edited image.
[0394] S812. Determine whether i is equal to H.
[0395] Exemplarily, after S811 is executed, it can be determined whether i is equal to H. If i is equal to H, then execute S814, that is, enter the subsequent stage of the diffusion chain generation process. If i is not equal to H, then execute S813.
[0396] S813. Increment i by 1.
[0397] Exemplarily, after S813 is executed, it can return to execute S804.
[0398] On the one hand, as the number of diffusion steps gradually approaches 0 (that is, the time node t gradually approaches 0), the image noise becomes smaller and the content becomes clearer. The automatic positioning of the area where the target object is located will be more accurate, which helps to obtain a more accurate replacement relationship, that is, the transformation relationship from the object to be replaced in the first image to the corresponding target object in the diffusion network generation graph, so as to achieve precise replacement. On the other hand, the (i - 1)-th edited image as the input for the i-th generation process of the diffusion network can obtain more information inputs of other original objects (original objects other than the object to be replaced) in the first image, which helps to maintain other original objects in the first image in the final output result, and also helps the algorithm generation adjustment during the generation process to make the interaction between other original objects and the target object in the first image more harmonious and natural. In this way, the user-specified category object can be replaced into the user input image (i.e., the first image), and the user-specified category object can be placed in the user-specified scene picture.
[0399] S814. Input the input text and the i-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain the second image, where N is a positive integer.
[0400] Exemplarily, S814 can refer to the description of the above embodiment and will not be elaborated here.
[0401] Figure 9 It is a schematic diagram of an image editing device shown exemplarily. The image editing device can be used to execute the method of the foregoing embodiment. Therefore, the beneficial effects it can achieve can refer to the beneficial effects in the corresponding method provided above and will not be elaborated here.
[0402] Reference Figure 9 , exemplarily, an image editing device may include:
[0403] An acquisition module 901, configured to acquire a first editing task, a first image, and input text; wherein, the first editing task is used to indicate editing of the content of the first image, the input text includes a first text, and the first text includes text for describing a target object targeted by the first editing task;
[0404] An image generation module 902, configured to repeatedly execute the following steps until i is equal to H, the initial value of i is 2, and i and H are positive integers: input the first text and the (i - 1)-th edited image into a diffusion network, and perform an i-th generation process by the diffusion network to obtain the i-th first intermediate image, wherein the 1st edited image is generated according to the first editing task, the first image, the first text, and the 1st first intermediate image, and the 1st first intermediate image is generated by the diffusion network through a 1st generation process based on the first text and a preset noise image; generate the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image; increment i by 1; when i is equal to H, input the input text and the H-th edited image into the diffusion network, and perform N generation processes by the diffusion network to obtain a second image; wherein, N is a positive integer.
[0405] It should be noted that the image generation module 902 may include the above-mentioned editing processing module and region positioning module.
[0406] Exemplarily, the image generation module 902 is specifically configured to input the first text and the H-th edited image into the diffusion network, and perform N generation processes by the diffusion network to obtain a second image.
[0407] Exemplarily, the input text further includes a second text, and the second text includes the content of a second editing task, and the second editing task is used to indicate editing of the first image other than the content.
[0408] Exemplarily, the second text is in M groups, each group of second text includes the content of a second editing task, and M is an integer greater than 1; the image generation module 902 is specifically configured to input the first text, the first group of second text, and the H-th edited image into the diffusion network, and perform R 1 generation processes to obtain the 1st second intermediate image; repeatedly execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of second text to the j-th group of second text, and the (j - 1)-th second intermediate image into the diffusion network, and perform R jPerform the j-th generation process to obtain the j-th second intermediate image; increment j by 1; use the M-th second intermediate image as the second image; where the sum of the number of generation processes performed in M - 1 loops and R 1 is equal to N.
[0409] Exemplarily, the second text is in M groups, each group of second text includes the content of a second editing task, and M is an integer greater than 1; the image generation module is specifically configured to input the first text, the first group of second text, and the H-th editing image into the diffusion network, and perform R 1 generation processes to obtain the first second intermediate image; loop through the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of second text to the j-th group of second text, and the j - 1-th second intermediate image into the diffusion network, and perform R j generation processes to obtain the j-th second intermediate image; increment j by 1; input the first text, the first group of second text to the M-th group of second text, and the M-th second intermediate image into the diffusion network, and perform P generation processes by the diffusion network to obtain the second image; where the sum of the number of generation processes performed in M - 1 loops, R 1 and P is equal to N, and P is a positive integer.
[0410] Exemplarily, the image generation module 902 is further configured to obtain a first region, where the first region is a region specified by the user; the image generation module 902 is specifically configured to input the first image, the first text, the first region, and the i - 1-th editing image into the diffusion network, and perform the i-th generation process by the diffusion network to obtain the i-th first intermediate image.
[0411] Exemplarily, the image generation module 902 is specifically configured to generate a first mask map according to the first region; where the size of the first mask map is the same as the size of the first image, the pixel values of the pixel points in the first region of the first mask map are 0, and the pixel values of the pixel points in other regions of the first mask map are 1; perform mask processing on the first image pair using the first mask map to obtain the masked first image; where the pixel values of the pixel points in the first region of the masked first image are 0; input the first mask map, the masked first image, the first text, and the i - 1-th editing image into the diffusion network, and perform the i-th generation process by the diffusion network to obtain the i-th first intermediate image.
[0412] Exemplarily, the second text is in M groups, each group of second text includes the content of a second editing task, and M is an integer greater than 1; the image generation module 902 is specifically configured to input the first text, the first group of second text, and the H-th editing image into the diffusion network, and perform R 1The first generation process is performed to obtain the first second intermediate image; the following steps are repeatedly executed until j is equal to G, where the initial value of j is 2, j is a positive integer, and G is an integer greater than or equal to M: The first text, the second texts of the first group to the j-th group, and the (j - 1)-th second intermediate image are input into the diffusion network, and the diffusion network performs R j generation processes to obtain the j-th second intermediate image; j is incremented by 1; where the sum of the number of generation processes performed in the G - 1 loops is equal to N 1 and R
[0413] Exemplarily, the image generation module 902 is specifically configured to generate the i-th second mask image according to the first text; wherein, the pixel values of the pixel points in the second region of the i-th second mask image are 1, and the pixel values of the pixel points in other regions of the i-th second mask image are 0, and the second region is the region for adding the target object; according to the i-th second mask image, the first image and the i-th first intermediate image are fused to obtain the i-th edited image.
[0414] Exemplarily, the image generation module 902 is specifically configured to perform the i-th noise addition process on the first image to obtain the i-th noise-added first image; according to the i-th second mask image, the i-th noise-added first image and the i-th first intermediate image are fused to obtain the i-th edited image.
[0415] Exemplarily, the image generation module 902 is specifically configured to generate the i-th third mask image according to the i-th second mask image, where the pixel values of the pixel points in the second region of the i-th third mask image are 0, and the pixel values of the pixel points in other regions of the i-th third mask image are 1; multiply the i-th second mask image by the i-th first intermediate image to obtain the i-th first fused image; multiply the i-th third mask image by the i-th noise-added first image to obtain the i-th second fused image; add the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
[0416] Exemplarily, the diffusion network includes a cross-attention module, and the i-th intermediate feature map output by the g-th network layer in the cross-attention module includes: the weights of each word in the first text with respect to each pixel point in the feature map input into the cross-attention module;
[0417] The image generation module 902 is specifically configured to obtain the i-th first attention feature map according to the i-th intermediate feature map; wherein, the i-th first attention feature map includes the weights of the keywords of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module; smooth the i-th first attention feature map to obtain the i-th second attention feature map; and perform binarization processing on the i-th second attention feature map to obtain the i-th second mask map.
[0418] Exemplarily, the first editing task includes any one of the following: adding a target object to the first image, replacing an object in the first image with the target object, replacing the target object with an object in the first image, or deleting the target object in the first image.
[0419] In one example, Figure 10 FIG. shows a schematic block diagram of a device 1000 according to an embodiment of the present application. The device 1000 may include: a processor 1001 and a transceiver / transceiver pin 1002. Optionally, it further includes a memory 1003.
[0420] Each component of the device 1000 is coupled together through a bus 1004. Among them, the bus 1004 includes not only a data bus but also a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus 1004 in the figure.
[0421] Optionally, the memory 1003 may be used to store the instructions in the foregoing method embodiments. The processor 1001 may be configured to execute the instructions in the memory 1003, control the receiving pin to receive signals, and control the transmitting pin to transmit signals.
[0422] The device 1000 may be the electronic device or the chip of the electronic device in the foregoing method embodiments.
[0423] Exemplarily, the electronic device may be a server or a terminal device.
[0424] Wherein, all the relevant contents of each step involved in the foregoing method embodiments may be cited in the function descriptions of the corresponding functional modules, and will not be elaborated herein.
[0425] The embodiment of the present application further provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits. When the one or more processors execute computer instructions, the steps of the relevant method steps described above are executed to implement the method in the foregoing embodiments. Among them, the interface circuit is the transceiver / transceiver pin 1002.
[0426] This embodiment also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device is caused to execute the above-related method steps to implement the method in the above embodiment.
[0427] This embodiment also provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer is caused to execute the above-related steps to implement the method in the above embodiment.
[0428] In addition, an embodiment of the present application also provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other. Among them, the memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory so that the chip executes the methods in the above method embodiments.
[0429] Among them, the electronic device, the computer-readable storage medium, the computer program product or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0430] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0431] In the several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0432] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0433] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0434] Any content of each embodiment of the present application, as well as any content of the same embodiment, can be freely combined. Any combination of the above content is within the scope of the present application.
[0435] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0436] The steps of the methods or algorithms described in combination with the disclosed content of the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules. The software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, compact discs read-only (CD-ROM), or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0437] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes a computer-readable storage medium and a communication medium, where the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0438] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. An image editing method, characterized in that, the method includes: obtaining a first editing task, a first image, and input text; wherein, the first editing task is used to indicate editing of the content of the first image, the input text includes a first text, and the first text includes text for describing a target object targeted by the first editing task; repeatedly execute the following steps until i is equal to H, the initial value of i is 2, and i and H are positive integers: input the first text and the (i - 1)-th edited image into a diffusion network, and perform the i-th generation process by the diffusion network to obtain the i-th first intermediate image, wherein the 1st edited image is generated according to the first editing task, the first image, the first text, and the 1st first intermediate image, and the 1st first intermediate image is generated by the diffusion network through the 1st generation process based on the first text and a preset noise image; generate the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image; increment i by 1; input the input text and the H-th edited image into the diffusion network, and perform N generation processes by the diffusion network to obtain a second image; wherein, N is a positive integer.
2. The method according to claim 1, characterized in that, the step of inputting the input text and the H-th edited image into the diffusion network, and performing N generation processes by the diffusion network to obtain a second image includes: input the first text and the H-th edited image into the diffusion network, and perform N generation processes by the diffusion network to obtain the second image.
3. The method according to claim 1, characterized in that, the input text further includes a second text, and the second text includes the content of a second editing task, and the second editing task is used to indicate editing of the first image other than the content.
4. The method according to claim 3, characterized in that, the second text is in M groups, and each group of second text includes the content of a second editing task, and M is an integer greater than 1; the step of inputting the input text and the H-th edited image into the diffusion network, and performing N generation processes by the diffusion network to obtain a second image includes: Input the first text, the first group of second texts, and the H-th edited image into the diffusion network, and perform R 1 times of generation processing by the diffusion network to obtain the first second intermediate image; repeatedly execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: Input the first text, the second texts of the 1st group to the jth group, and the (j - 1)th second intermediate image into the diffusion network, and perform R j times of generation processing by the diffusion network to obtain the jth second intermediate image; increment j by 1; use the M-th second intermediate image as the second image; Among them, the sum of the number of generation processes executed in the M-1 times of loops and R 1 is equal to N.
5. The method according to claim 3, characterized in that, the second text is in M groups, and each group of second text includes the content of a second editing task, and M is an integer greater than 1; the step of inputting the input text and the H-th edited image into the diffusion network, and performing N generation processes by the diffusion network to obtain a second image includes: Input the first text, the first group of second texts, and the H-th edited image into the diffusion network, and perform R 1 times of generation processing by the diffusion network to obtain the first second intermediate image; repeatedly execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: Input the first text, the second texts of the 1st group to the jth group, and the (j - 1)th second intermediate image into the diffusion network, and perform R j times of generation processing by the diffusion network to obtain the jth second intermediate image; increment j by 1; input the first text, the first group of second text to the M-th group of second text, and the M-th second intermediate image into the diffusion network, and perform P generation processes by the diffusion network to obtain the second image; Among them, the number of generation processes executed in the (M - 1) - th loop, R 1 and the sum of P is equal to N, where P is a positive integer.
6. The method according to any one of claims 1 to 5, wherein, the method further comprises: obtaining a first region, where the first region is a region specified by the user; the step of inputting the first text and the (i-1)th edited image into the diffusion network, and performing the ith generation process by the diffusion network to obtain the ith first intermediate image, includes: inputting the first image, the first text, the first region, and the (i-1)th edited image into the diffusion network, and performing the ith generation process by the diffusion network to obtain the ith first intermediate image.
7. The method according to claim 6, wherein, the step of inputting the first image, the first text, the first region, and the (i-1)th edited image into the diffusion network, and performing the ith generation process by the diffusion network to obtain the ith first intermediate image, includes: generating a first mask map according to the first region; wherein, the size of the first mask map is the same as the size of the first image, the pixel values of the pixel points in the first region of the first mask map are 0, and the pixel values of the pixel points in other regions of the first mask map are 1; performing mask processing on the first image pair by using the first mask map to obtain a masked first image; wherein, the pixel values of the pixel points in the first region of the masked first image are 0; inputting the first mask map, the masked first image, the first text, and the (i-1)th edited image into the diffusion network, and performing the ith generation process by the diffusion network to obtain the ith first intermediate image.
8. The method according to any one of claims 1 to 7, wherein, the step of generating the ith edited image according to the first editing task, the first image, the first text, and the ith first intermediate image, includes: generating the ith second mask map according to the first text; wherein, the pixel values of the pixel points in the second region of the ith second mask map are 1, the pixel values of the pixel points in other regions of the ith second mask map are 0, and the second region is the region for adding the target object; fusing the first image and the ith first intermediate image according to the ith second mask map to obtain the ith edited image.
9. The method according to claim 8, wherein, the step of fusing the first image and the ith first intermediate image according to the second mask map to obtain the ith edited image, includes: performing the ith noise addition process on the first image to obtain the ith noise-added first image; fusing the ith noise-added first image and the ith first intermediate image according to the ith second mask map to obtain the ith edited image.
10. The method according to claim 9, wherein, Fusing the i-th denoised first image and the i-th first intermediate image according to the i-th second mask image to obtain the i-th edited image includes: Generating an i-th third mask image according to the i-th second mask image, where the pixel values of the pixel points in the second region of the i-th third mask image are 0, and the pixel values of the pixel points in other regions of the i-th third mask image are 1; Multiplying the i-th second mask image by the i-th first intermediate image to obtain an i-th first fused image; Multiplying the i-th third mask image by the i-th denoised first image to obtain an i-th second fused image; Adding the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
11. The method according to any one of claims 8 to 10, wherein, the diffusion network includes a cross-attention module, and the i-th intermediate feature map output by the g-th network layer in the cross-attention module includes: the weights of each word in the first text relative to each pixel point in the feature map input to the cross-attention module; Generating the i-th second mask image according to the first text includes: Obtaining an i-th first attention feature map according to the i-th intermediate feature map; wherein, the i-th first attention feature map includes the weights of the keyword of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module; Smoothing the i-th first attention feature map to obtain an i-th second attention feature map; Performing binarization processing on the i-th second attention feature map to obtain the i-th second mask image.
12. The method according to any one of claims 1 to 11, wherein, the first editing task includes any one of the following: adding the target object to the first image, replacing the object in the first image with the target object, replacing the target object with the object in the first image, or deleting the target object in the first image.
13. An image editing device, wherein, the device includes: An acquisition module for acquiring a first editing task, a first image, and an input text; wherein, the first editing task is used to indicate editing of the content of the first image, the input text includes a first text, and the first text includes text for describing the target object targeted by the first editing task; An image generation module is used to repeatedly execute the following steps until i is equal to H. The initial value of i is 2, and i and H are positive integers: Input the first text and the (i - 1)-th edited image into a diffusion network, and the diffusion network performs the i-th generation process to obtain the i-th first intermediate image. Among them, the first edited image is generated according to the first editing task, the first image, the first text, and the first first intermediate image, and the first first intermediate image is generated by the diffusion network through the first generation process based on the first text and a preset noise image; Generate the i-th edited image according to the first editing task, the first image, the first text, and the i-th first intermediate image; Increment i by 1; When i is equal to H, input the input text and the H-th edited image into the diffusion network, and the diffusion network performs N generation processes to obtain a second image; where N is a positive integer.
14. The apparatus according to claim 13, wherein, the second text is in M groups, and each group of the second text includes the content of a second editing task, and M is an integer greater than 1; The image generation module is specifically configured to input the first text, the first group of second texts, and the H-th edited image into the diffusion network, and perform R 1 times of generation processing by the diffusion network to obtain the first second intermediate image; loop and execute the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first group of second texts to the j-th group of second texts, and the (j - 1)-th second intermediate image into the diffusion network, and perform R j times of generation processing by the diffusion network to obtain the j-th second intermediate image; Increment j by 1; use the M-th second intermediate image as the second image; where the sum of the number of generation processes executed in M - 1 loops and R 1 is equal to N.
15. The apparatus according to claim 13, wherein, the second text is in M groups, and each group of the second text includes the content of a second editing task, and M is an integer greater than 1; The image generation module is specifically configured to input the first text, the first set of second texts, and the Hth edited image into the diffusion network, and perform R 1 times of generation processing by the diffusion network to obtain the first second intermediate image; loop through the following steps until j is equal to M, the initial value of j is 2, and j is a positive integer: input the first text, the first set of second texts to the jth set of second texts, and the (j - 1)th second intermediate image into the diffusion network, and perform R j times of generation processing by the diffusion network to obtain the jth second intermediate image; Increment j by 1; input the first text, the second texts of the first group to the M-th group, and the M-th second intermediate image into the diffusion network, and perform P generation processes by the diffusion network to obtain the second image; wherein, the sum of the number of generation processes executed in M-1 loops, R 1 and P is equal to N, and P is a positive integer.
16. The apparatus according to any one of claims 13 to 15, wherein, the image generation module is specifically configured to generate the i-th second mask image according to the first text; wherein, the pixel values of the pixel points in the second region of the i-th second mask image are 1, and the pixel values of the pixel points in other regions of the i-th second mask image are 0, and the second region is the region for adding the target object; According to the i-th second mask image, fuse the first image and the i-th first intermediate image to obtain the i-th edited image.
17. The apparatus according to claim 16, wherein, the image generation module is specifically configured to perform the i-th noise addition process on the first image to obtain the i-th noise-added first image; According to the i-th second mask image, fuse the i-th noise-added first image and the i-th first intermediate image to obtain the i-th edited image.
18. The apparatus according to claim 17, wherein, the image generation module is specifically configured to generate the i-th third mask image according to the i-th second mask image. The pixel values of the pixel points in the second region of the i-th third mask image are 0, and the pixel values of the pixel points in other regions of the i-th third mask image are 1; Multiply the i-th second mask image by the i-th first intermediate image to obtain the i-th first fused image; Multiply the i-th third mask image by the i-th noise-added first image to obtain the i-th second fused image; Add the i-th first fused image and the i-th second fused image to obtain the i-th edited image.
19. The device according to any one of claims 16 to 18, wherein, the diffusion network includes a cross-attention module, and the i-th intermediate feature map output by the g-th network layer in the cross-attention module includes: the weight of each word in the first text relative to each pixel point in the feature map input to the cross-attention module; the image generation module is specifically configured to obtain the i-th first attention feature map according to the i-th intermediate feature map; wherein, the i-th first attention feature map includes the weight of the keyword of the target object in the first text relative to each pixel point in the feature map input to the cross-attention module; smooth the i-th first attention feature map to obtain the i-th second attention feature map; perform binarization processing on the i-th second attention feature map to obtain the i-th second mask map.
20. The device according to any one of claims 13 to 19, wherein, the first editing task includes any one of the following: adding the target object to the first image, replacing the object in the first image with the target object, replacing the target object with the object in the first image, or deleting the target object in the first image.
21. An electronic device, wherein, it includes: a memory and a processor, the memory is coupled with the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device is enabled to execute the image editing method according to any one of claims 1 to 12.
22. A chip, wherein, it includes one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method according to any one of claims 1 to 12 are executed.
23. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program runs on a computer or a processor, the computer or the processor is enabled to execute the image editing method according to any one of claims 1 to 12.
24. A computer program product, wherein, the computer program product contains computer instructions, and when the computer instructions are executed by a computer or a processor, the steps of the method according to any one of claims 1 to 12 are executed.