Image editing method and apparatus, device, and storage medium
By using target mask maps and noise prediction processing in the image editing model, precise editing of specified areas is achieved, solving the problem that image editing in existing technologies cannot make precise modifications, and improving the quality of image generation and the consistency between images and text.
Patent Information
- Application Number
- PCT/CN2025/088422
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-15
- Filing Date
- 2025-04-11
- Publication Date
- 2025-10-23
AI Technical Summary
Existing technologies cannot accurately modify and edit specified areas in image editing based on text descriptions, resulting in discrepancies between the generated image and the text description, thus affecting the quality of the generated image.
By acquiring the region to be edited and the target prompt, an image editing model is used to add preset noise to the region to be edited based on the target mask image. Noise prediction and image generation processing are then performed to achieve precise editing of local regions.
It improves the accuracy of image editing and the matching degree between images and text, reduces the difficulty of image processing, and improves the quality of generated images.
Smart Images

Figure CN2025088422_23102025_PF_FP_ABST
Abstract
Description
Image editing method, device, equipment and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410455156.2, filed on April 15, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to an image editing method, device, equipment and storage medium. BACKGROUND
[0003] With the development of computer technology, there is an increasing demand for editing local regions in original images according to text descriptions.
[0004] Currently, in the image editing task according to the text description, the semantic image editing is used to process the entire image without distinction, which cannot accurately modify and edit the specified region, resulting in the generated image being different from the text description, and affecting the image generation quality. SUMMARY
[0005] The present disclosure provides an image editing method, device, equipment and storage medium, which can accurately edit the specified region, and improve the matching degree of text and image and the image generation quality.
[0006] In a first aspect, embodiments of the present disclosure provide an image editing method, comprising:
[0007] obtaining a to-be-edited region in an original image and a target prompt word corresponding to the to-be-edited region, wherein the target prompt word is text information used to describe the expected effect of image editing;
[0008] determining a target mask image according to the to-be-edited region, adding a preset noise to the to-be-edited region based on the target mask image through an image editing model, and obtaining a local noise image;
[0009] performing noise prediction processing and image generation processing on the to-be-edited region of the local noise image based on the target prompt word through the image editing model, and outputting a target image according to the noise prediction result and the image generation result.
[0010] In a second aspect, embodiments of the present disclosure also provide an image editing device, which comprises:
[0011] an obtaining module configured to obtain a to-be-edited region in an original image and a target prompt word corresponding to the to-be-edited region, wherein the target prompt word is text information used to describe the expected effect of image editing;
[0012] add a preset noise to the to-be-edited region based on the target mask map through the image editing model, to obtain a local noise image;
[0013] generate a target image according to the noise prediction result and the image generation result.
[0014] In a third aspect, the present disclosure also provides an electronic device, which comprises:
[0015] one or more processors;
[0016] a storage device configured to store one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the image editing method according to any embodiment of the present disclosure.
[0018] In a fourth aspect, the present disclosure also provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are configured to perform the image editing method according to any embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0020] FIG. 1 is a flowchart of an image editing method according to an embodiment of the present disclosure;
[0021] FIG. 2 is a schematic diagram of an editing interface according to an embodiment of the present disclosure;
[0022] FIG. 3 is a flowchart of another image editing method according to an embodiment of the present disclosure;
[0023] FIG. 4 is a schematic diagram of training of an image editing model according to an embodiment of the present disclosure;
[0024] FIG. 5 is a schematic diagram of an image editing device according to an embodiment of the present disclosure; and
[0025] FIG. 6 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so as to enable a more thorough and complete understanding of the present disclosure. It is understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0028] The term “comprising” and variations thereof as used herein are open-ended, that is “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related terms are defined in the following description.
[0029] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0030] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly indicated in the context.
[0031] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0032] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0033] For example, when responding to the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware, such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0034] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner in which the prompt information may be presented in a text manner. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0035] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0036] It can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0037] It should be noted that in the embodiments of the present disclosure, some existing industry solutions such as software, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solution of the present disclosure, but it does not mean that the applicant has or will necessarily use the solution.
[0038] FIG. 1 is a flowchart of an image editing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to a local editing scenario. For example, a user inputs a region selection operation for an original image to determine a to-be-edited region in the original image, and inputs a target prompt word for the to-be-edited region. An image editing model generates a target object in the to-be-edited region based on the original image, the to-be-edited region, and the target prompt word, obtains a target image with the target object, and outputs the target image. The target object represents a to-be-edited region object accurately generated based on the target prompt word. The method can be executed by an image editing device, which can be implemented in the form of software and / or hardware. Optionally, the image editing device is implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server, etc.
[0039] As shown in FIG. 1, the method includes:
[0040] S110, obtaining a to-be-edited region in an original image and a target prompt word corresponding to the to-be-edited region.
[0041] In the embodiments of the present disclosure, the original image represents a picture or a video frame of a video to be subjected to an image editing operation. The original image can include a picture or a video frame stored in an electronic device. Alternatively, the original image can include a picture or a video frame downloaded from a server. Alternatively, the original image can also include a picture or a video frame taken by a user. If an image editing operation of a local region is needed, the original image can be displayed through an editing interface. The editing interface can be an interactive interface displayed by a local editing control in a client. The client can include an application program, a mini-program or a web client, etc. The local editing control is used to display the editing interface.
[0042] The region to be edited represents an image region to be subjected to an image editing operation. The region to be edited can be determined through a region selection operation on the original image in the editing interface. The region selection operation can include a brushing operation, a framing operation or a sliding operation, etc. For example, a brushing operation on the original image in the editing interface is obtained, and the region to be edited is determined according to the brushed region. Alternatively, a framing operation on the original image in the editing interface is obtained, and the region to be edited is determined according to the framed region. Alternatively, a sliding operation on the original image in the editing interface is obtained, and the region to be edited is determined according to the sliding track.
[0043] The target prompt word is text information used to describe an expected effect of image editing. Specifically, the target prompt word represents a description text of a target object expected to be generated in the region to be edited. For example, the target prompt word can represent a car, indicating that the expected effect of image editing is to replace the object in the region to be edited with a car.
[0044] For example, the original image is displayed in the editing interface; a region selection operation on the original image is obtained, and a region to be edited in the original image is determined according to the region selection operation; a text input operation on the region to be edited is obtained, and a target prompt word corresponding to the region to be edited is determined according to the text input operation.
[0045] FIG. 2 is a schematic diagram of an editing interface provided by an embodiment of the present disclosure. As shown in FIG. 2, the editing interface 200 includes an editing region 210 and a result display region 220, etc. The editing region 210 includes an image loading control 230 and a prompt word input control 240. In response to a triggering event of the image loading control 230, an original image 250 is obtained, and the original image 250 is displayed at a position corresponding to the image loading control 230. In response to a brushing operation on the original image 250, a brushed region in the original image 250 is determined as a region to be edited. In response to a text input operation on the prompt word input control 240, a target prompt word corresponding to the region to be edited is obtained. The prompt word input control 240 includes a text box control.
[0046] S120, determine a target mask image according to the region to be edited, and add preset noise to the region to be edited based on the target mask image through an image editing model to obtain a local noise image.
[0047] The target mask image is used to cover the non-region-to-be-edited in the original image to prompt the image editing model to perform image editing on the region to be edited. For example, an original mask image consistent with the size of the original image is generated, a target object in the region to be edited of the original image is identified, the region to be edited in the original mask image is determined according to the coordinates of the target object, the region to be edited in the original mask image is filled with white, and the non-region-to-be-edited is filled with black to obtain the target mask image. Alternatively, the target object in the region to be edited of the original image can also be dilated to expand the boundary of the target object, and then the target mask image is determined based on the dilated target object.
[0048] The image editing model can be a diffusion model trained based on a sample image set, a sample mask image set, and a description text set. The sample mask image set is determined based on a sample dilated image corresponding to a sample image in the sample image set. The sample dilated image represents an image after performing a dilation operation on a target region corresponding to the sample image. The description text set is determined based on image content of the target region. The target region can represent a target object carrying semantics segmented by a panoramic segmentation model in the sample image.
[0049] The diffusion model includes an encoder, a noise prediction network, and a decoder. The input of the encoder includes the original image, which is used to compress the original image to a low-dimensional space to obtain a latent feature map. The decoder is used to restore the low-dimensional image after completing the image editing task to the size of the original image to obtain a target image. The noise prediction network is used to add noise to the region to be edited in the latent feature map under the constraint of the target mask image, and keep the latent feature unchanged for the non-region-to-be-edited to obtain a local noise image. In addition, under the constraint of the target prompt word, the noise prediction network is used to perform noise prediction processing and image generation processing on the region to be edited in the local noise image zT to obtain the prediction noise corresponding to the time step t=T and the image-text correlation corresponding to the time step t=T. Then, based on the image-text correlation corresponding to the time step t=T and the local noise image zT corresponding to the time step t=T, a local editing image corresponding to the time step t=T is generated. The prediction noise corresponding to the time step t=T is subtracted from the local editing image corresponding to the time step t=T to obtain a local noise image zT-1 corresponding to the time step t=T-1. The local noise image zT-1 corresponding to the time step t=T-1 is input into the noise prediction network to generate a local noise image zT-2 corresponding to the time step t=i-2, and so on, until a local noise image z0 corresponding to the time step t=0 is generated, i.e., a low-dimensional feature map of the target image is obtained.
[0050] Optionally, the noise prediction network can be an Unet network, including a convolutional layer, a down-sampling layer, a down-sampling layer based on a multi-head attention mechanism, an intermediate network, an up-sampling layer based on a multi-head attention mechanism, and an up-sampling layer. By introducing the multi-head attention mechanism, the text and the image are associated. The down-sampling layer includes a plurality of residual modules. The down-sampling layer based on the multi-head attention mechanism can include a residual module and a Transformer module, and the diffusion model can include at least one down-sampling layer based on the multi-head attention mechanism. If there are a plurality of down-sampling layers based on the multi-head attention mechanism, the different down-sampling layers based on the multi-head attention mechanism include different numbers of Transformer modules. The intermediate network includes a residual module and a Transformer module. The up-sampling layer based on the multi-head attention mechanism can include a residual module and a Transformer module, and the diffusion model can include at least one up-sampling layer based on the multi-head attention mechanism. If there are a plurality of up-sampling layers based on the multi-head attention mechanism, the different up-sampling layers based on the multi-head attention mechanism include different numbers of Transformer modules. The up-sampling layer includes a plurality of residual modules.
[0051] The preset noise can be noise that meets a preset requirement in terms of distribution attribute in the noise feature map. For example, the preset noise can be random noise that meets a Gaussian distribution. The noise feature map can be a noise image that is consistent in size with the original image. The noise feature region can be determined by superimposing the target mask image and the noise feature map. The noise feature region represents a specific region in the noise feature map corresponding to the region to be edited. Since the region to be edited in the target mask image is white and the remaining regions are black, superimposing the target mask image on the noise feature map can cover the non-edited region in the noise feature map, thereby determining the noise feature region corresponding to the region to be edited.
[0052] The local noise image represents an image obtained by adding the preset noise to the region to be edited of the original image. Since the image after adding noise only has the region to be edited in a noise state, the image after adding noise is taken as the local noise image. Optionally, in order to reduce the amount of calculation, the original image is compressed to a low-dimensional latent space by an encoder to obtain a latent feature map of the original image. The preset noise is added to the region to be edited of the latent feature map to obtain the local noise image.
[0053] Exemplarily, the target mask image is determined according to the region to be edited, and a preset noise is added to the region to be edited based on the target mask image through the image editing model to obtain a local noise image, including: determining a target mask image according to the foreground content in the region to be edited; generating a latent feature map of the original image through the image editing model, and determining the region to be edited in the latent feature map in combination with the target mask image and the latent feature map; and performing a preset number of noise addition operations based on the region to be edited in the latent feature map to obtain a local noise image.
[0054] For example, if a local region in the original image is painted, the foreground content in the painted region of the original image is identified to obtain a target object. A target mask image is generated based on the target object in the original image. The original image is compressed through an encoder in the image editing model to obtain a latent feature map of the original image. Since the region to be edited in the target mask image is white and the other regions are black, the target mask image is superimposed on the latent feature map, which can cover the non-edited region in the latent feature map, so that the region to be edited in the latent feature map is determined. A preset number of noise addition operations are performed based on the region to be edited in the latent feature map through the image editing model to obtain a local noise image.
[0055] In some embodiments, the latent feature map is represented as z0, the region to be edited in z0 is determined under the constraint of the target mask image through the image editing model, random noise is added to the region to be edited in z0 to obtain a local noise image z1. Then, random noise is added to the region to be edited in the local noise image z1 through the image editing model to obtain a local noise image z2. The noise addition step is iteratively performed until a local noise image zT is obtained.
[0056] Optionally, the preset number of noise addition operations based on the region to be edited in the latent feature map to obtain a local noise image includes: obtaining a noise feature map, wherein the noise feature map represents a preset noise; determining a noise feature region in combination with the target mask image and the noise feature map; and performing a preset number of noise addition operations based on the region to be edited and the noise feature region to obtain a local noise image.
[0057] For example, the latent feature map is represented as z0, and the noise feature map is represented as S. z0, S, and the target mask map are input into the Unet network of the image editing model. The Unet network determines the region to be edited in z0 under the constraints of the target mask map. The Unet network also determines the noise feature region under the constraints of the target mask map. The noise features in the noise feature region are superimposed on the pixel features in the region to be edited in z0 to obtain the local noise map z1. Then, using the same method, random noise is added to the region to be edited in the local noise map z1 through the Unet network to obtain the local noise map z2. The noise addition step is iterated until the local noise map zT is obtained.
[0058] S130 , performing noise prediction processing and image generation processing on the to-be-edited area of the local noise image based on the target prompt word through the image editing model, and outputting a target image according to the noise prediction result and the image generation result.
[0059] Exemplarily, the image editing model determines predicted noise based on image features in the to-be-edited area of the local noise image; the image editing model generates text features based on the target prompt word, and determines the correlation between the text features and the local noise image; a local edited image is generated based on the correlation and the local noise image, and a denoising operation is performed on the local edited image based on the predicted noise, and a target image is output.
[0060] For example, the image editing model includes a text mapping module for mapping the target prompt word into a text vector as a text feature. The text features, the local noise image ZT, and the time step t=T are input into the Unet network of the image editing model. The Unet network outputs the predicted noise corresponding to time step t=T, calculates the correlation between the text features and the local noise image ZT, and adjusts the pixel distribution of the local noise image ZT based on the correlation to obtain the local edited image corresponding to time step t=T. The predicted noise corresponding to time step t=T is subtracted from the local edited image corresponding to time step t=T to obtain the local noise image at time step t=T-1. The text features, the local noise image ZT-1, and the time step t=T-1 are input into the Unet network of the image editing model. Using a similar method, the local noise image ZT-2 at time step t=T-2 can be obtained. Similarly, after performing T rounds of denoising and image generation, the local noise image Z0 at time step t=0 is obtained. The local noise image Z0 is input into the decoder of the image editing model, and the local noise image Z0 after denoising and image generation is decompressed by the decoder to restore it to the size of the original image to obtain the target image.
[0061] Optionally, the method further comprises: obtaining mask setting information, the mask setting information comprising a blur level. determining, by the image editing model, predicted noise according to image features in the to-be-edited region of the local noise image; determining, by the image editing model, a text feature based on the target prompt word, and determining relevance between the text feature and the local noise image; determining a filling degree of the generated object to the to-be-edited region based on the blur level; and generating a local edited image according to the relevance, the filling degree, and the local noise image, wherein the to-be-edited region in the local edited image comprises the generated object. performing a denoising operation on the local edited image based on the predicted noise, and outputting a target image.
[0062] The filling degree represents distribution information of the generated object in the to-be-edited region. Under different blur levels, the filling degree of the generated object to the to-be-edited region is different. For a case with a higher filling degree, the generated object is close to the edge of the to-be-edited region. For a case with a lower filling degree, the distance between the generated object and the edge of the to-be-edited region is increased.
[0063] For example, the editing interface further comprises a mask setting control for inputting mask setting information. If a triggering operation on the mask setting control is detected, the mask setting information is obtained, and the mask setting information is parsed to obtain the blur level. After the text feature is generated based on the target prompt word by the image editing model, and the relevance between the text feature and the local noise image is determined, the filling degree of the generated object to the to-be-edited region is determined based on the blur level. A local edited image is generated according to the relevance, the filling degree, and the local noise image, wherein the to-be-edited region in the local edited image comprises the generated object. A denoising operation is performed on the local edited image based on the predicted noise, and a target image is output.
[0064] The technical scheme of the embodiment of the present disclosure accurately describes the expected editing effect of the to-be-edited region by the target prompt word for the to-be-edited region, determines the target mask based on the to-be-edited region, adds the preset noise to the to-be-edited region based on the target mask by the image editing model, obtains the local noise image, and then performs denoising processing and text generation processing on the to-be-edited region of the local noise image based on the target prompt word by the image editing model, to generate the target object in the to-be-edited region that meets the expected editing effect, accurately edit the specified region, and improve the image-text matching degree and the image generation quality. Since only the to-be-edited region is subjected to the noise adding processing and the denoising processing, the image processing difficulty and the color difference between the target object and the original image are reduced. The embodiment of the present disclosure solves the problem that the image editing method in the related art cannot accurately modify and edit the local region, and improves the image-text consistency and the image generation quality.
[0065] FIG. 3 is a flowchart of another image editing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of training an image editing model. The embodiment of the present disclosure is based on the above-mentioned embodiments and additionally defines the training process of the image editing model.
[0066] As shown in FIG. 3, the method comprises the following steps.
[0067] S310, obtaining a sample image set.
[0068] The sample image set is a collection of sample images. The sample images can include countable instance objects, such as people, vehicles, animals, and the like. The sample images can also include regions without fixed shapes, such as sky, grass, snow, trees, and the like.
[0069] S320, for a sample image in the sample image set, determining a target region in the sample image, performing an inflation operation on the target region to obtain a sample inflation image, and determining a sample mask image set according to the sample inflation image corresponding to the sample image set.
[0070] The target region represents an instance object and / or a region without a fixed shape in the sample image that carries semantic information. The inflation operation is a morphological operation in image processing, which is used to make the target region larger and the boundary rougher.
[0071] The panoramic segmentation model is used to segment the sample image to obtain the instance object and / or the region without a fixed shape in the sample image. The inflation operation is performed on the instance object and / or the region without a fixed shape in the sample image to obtain the sample inflation image. The embodiment of the present disclosure makes the actual input target mask image in the model prediction process exceed the target region boundary by randomly inflating the sample image, completely solves the complete filling problem of the model, makes the generated object of the to-be-edited region in the target image more real and natural, and improves the image generation quality.
[0072] The sample mask image set is determined according to the sample inflation image corresponding to each sample image in the sample image set. For example, for the sample inflation image corresponding to the sample image set, a reference mask image of the sample image is generated according to the target region in the sample inflation image; a Gaussian blur process is performed on the reference mask image to obtain at least two sample mask images corresponding to the sample image; and the sample mask image set is determined according to the at least two sample mask images corresponding to the sample images in the sample image set, wherein the at least two sample mask images are associated with the description text.
[0073] In the embodiments of the present disclosure, a reference mask image of a sample image is generated based on a target region in a sample expansion image, wherein a region to be edited of the reference mask image of the sample image is filled with white color and a non-region to be edited is filled with black color. The reference mask image includes an instance mask image and a semantic mask image, etc.
[0074] The reference mask image is subjected to Gaussian blur processing by using Gaussian kernels of different sizes, to obtain at least two sample mask images of different roughness levels corresponding to the sample image. The sample mask image set is constituted according to the at least two sample mask images corresponding to each sample image in the sample image set, to meet the image editing requirements of different precisions. For the case that the target region represents an instance target, the image editing model is trained by using the at least two sample mask images of different roughness levels, so that the image editing model can be input with a fine instance target boundary and a rough rectangle as a mask image, and output a target image with a higher image-text matching degree. For the case that the target region represents a region without a fixed shape, the image editing model is trained by using the at least two sample mask images of different roughness levels, so that the image editing model can be input with a fine instance target boundary and a rough rectangle as a mask image, and output a target image in which a generated object fills the entire region to be edited.
[0075] In S330, image content recognition is performed on the target region, to obtain a description text corresponding to the target region, and a description text set is determined according to the description texts of the target regions corresponding to the sample image set.
[0076] For example, the image content recognition and understanding of the target region are performed by using a preset image description generation model, to output a description text corresponding to the target region. The description text set is constituted according to the description texts of the target regions of each sample image in the sample image set. The association between the description text of the sample image and the sample mask image is established, to obtain a graph-text data pair.
[0077] In S340, a preset editing model is trained according to the sample image set, the sample mask image set and the description text set, to obtain an image editing model.
[0078] The preset editing model can include an encoder, a noise prediction network and a decoder. Optionally, the noise prediction network can be a Unet network, including a convolutional layer, a down-sampling layer, a down-sampling layer based on a multi-head attention mechanism, an intermediate network, an up-sampling layer based on a multi-head attention mechanism, and an up-sampling layer. By introducing the multi-head attention mechanism, the text and the image are associated. The down-sampling layer includes a plurality of residual modules. The down-sampling layer based on the multi-head attention mechanism can include a residual module and a Transformer module, and the diffusion model can include at least one down-sampling layer based on the multi-head attention mechanism. If there are a plurality of down-sampling layers based on the multi-head attention mechanism, the different down-sampling layers based on the multi-head attention mechanism include different numbers of Transformer modules. The noise prediction network includes a residual module and a Transformer module. The up-sampling layer based on the multi-head attention mechanism can include a residual module and a Transformer module, and the diffusion model can include at least one up-sampling layer based on the multi-head attention mechanism. If there are a plurality of up-sampling layers based on the multi-head attention mechanism, the different up-sampling layers based on the multi-head attention mechanism include different numbers of Transformer modules. The up-sampling layer includes a plurality of residual modules.
[0079] FIG. 4 is a schematic diagram of training of an image editing model according to an embodiment of the present disclosure. As shown in FIG. 4, the input of the encoder 410 includes a sample image 420, which is used to compress the sample image 420 to a low-dimensional space to obtain a latent feature map 430. The decoder 440 is used to restore the low-dimensional image 450 after completing the image editing task to the size of the sample image to obtain an edited result image 460. The noise prediction network 470 is used to iteratively perform T times of adding noise operation on the to-be-edited region 490 in the latent feature map 430 under the constraint of the sample mask image 480, and keep the latent features unchanged for the non-to-be-edited region, to obtain a local noise image 4100. In addition, under the constraint of the time step and the description text, the noise prediction processing and the image generation processing are performed on the to-be-edited region in the local noise image zT, to obtain the prediction noise corresponding to the time step t=T and the correlation between the image and the text. Then, the local edited image corresponding to the time step t=T is generated based on the correlation between the image and the text corresponding to the time step t=T and the local noise image zT. The prediction noise corresponding to the time step t=T is subtracted from the local edited image corresponding to the time step t=T to obtain the local noise image zT-1 corresponding to the time step t=T-1. The local noise image zT-1 corresponding to the time step t=T-1 is input into the noise prediction network to generate the local noise image zT-2 corresponding to the time step t=T-2, and so on, until the local noise image z0 corresponding to the time step t=0 is generated, i.e., the low-dimensional image 450 of the edited result image 460 is obtained.
[0080] The low-dimensional image 450 is input into the decoder 440, and the decoder 440 decompresses the low-dimensional image 450 to obtain an edited result image 460 having the same image size as the sample image 420. A loss value of the generated object in the to-be-edited region and the real instance target or the region without a fixed shape in the edited result image 460 is calculated, the model parameters are updated in the process of back propagation according to the loss value, and finally the trained image editing model is obtained.
[0081] Optionally, after the reference mask image is subjected to Gaussian blur processing to obtain at least two sample mask images corresponding to the sample image, the method further includes: determining a blur level of the sample mask image according to a size of a Gaussian kernel used in the Gaussian blur processing, and marking the corresponding sample mask image according to the blur level. Since the reference mask image is subjected to Gaussian blur processing using Gaussian kernels of different sizes, the sample mask images obtained have different roughness levels. The sample mask images can be marked with blur levels based on the size of the Gaussian kernel used in the Gaussian blur processing. For example, a fine sample mask image corresponds to a blur level of 0, and as the roughness level of the sample mask image increases, the blur level corresponding to the sample mask image increases.
[0082] The training of the preset editing model according to the sample image set, the sample mask image set, and the description text set to obtain the image editing model includes: training the preset editing model according to the sample image set, the sample mask image set, the blur level corresponding to the sample mask image, and the description text set to obtain the image editing model. Specifically, the blur level is mapped to a blur level vector, and the blur level vector is injected into the noise prediction network, so that the model learns the corresponding relationship between different blur levels and sample mask images of different roughness levels, and then controls the filling degree of the generated object to the to-be-edited region.
[0083] The technical scheme of the embodiment of the present disclosure adds noise to the to-be-edited region and removes the noise, increases the local correlation, improves the matching degree of the edited result image and the description text, improves the generation stability of the model, avoids the generation error due to the fact that the description text describes the whole image during the process of adding noise to the whole image and removing the noise, and causes the edited result image to fail to meet the expectation. For the non-to-be-edited region, the latent features of the sample image are used as input in the adding noise and removing noise operations corresponding to all time steps of the model training, the reconstruction difficulty of the model in the image editing task is reduced, the color difference of the generated object is reduced, and the consistency of the model is improved.
[0084] FIG. 5 is a structural schematic diagram of an image editing device provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server.
[0085] As shown in FIG. 5, the apparatus includes an acquisition module 510, a noise adding module 520, and an image generation module 530.
[0086] The acquisition module 510 is configured to acquire a to-be-edited region in an original image and a target prompt word corresponding to the to-be-edited region, wherein the target prompt word is text information used to describe an expected effect of image editing.
[0087] The noise adding module 520 is configured to determine a target mask based on the to-be-edited region, add preset noise to the to-be-edited region based on the target mask through an image editing model, and obtain a local noise image.
[0088] The image generation module 530 is configured to perform noise prediction processing and image generation processing on the to-be-edited region of the local noise image based on the target prompt word through the image editing model, and output a target image according to a noise prediction result and an image generation result.
[0089] Optionally, the training manner of the image editing model includes:
[0090] acquiring a sample image set;
[0091] For a sample image in the sample image set, a target region in the sample image is determined, an expansion operation is performed on the target region to obtain a sample expansion image, and a sample mask set is determined according to the sample expansion image corresponding to the sample image set.
[0092] Image content recognition is performed on the target region to obtain description text corresponding to the target region, and a description text set is determined according to the description text of the target region corresponding to the sample image set.
[0093] A preset editing model is trained according to the sample image set, the sample mask set, and the description text set to obtain an image editing model.
[0094] Further, the sample mask set is determined according to the sample expansion image corresponding to the sample image set, and includes:
[0095] For the sample expansion image corresponding to the sample image set, a reference mask of the sample image is generated according to the target region in the sample expansion image.
[0096] The reference mask is subjected to Gaussian blur processing to obtain at least two sample masks corresponding to the sample image.
[0097] The sample mask set is determined according to the at least two sample masks corresponding to the sample image in the sample image set.
[0098] Optionally, after the reference mask image is Gaussian blurred to obtain at least two sample mask images corresponding to the sample image, the method further includes:
[0099] According to the size of the Gaussian kernel used in the Gaussian blurring, a blur level of the sample mask image is determined, and the corresponding sample mask image is labeled according to the blur level;
[0100] The training of the preset editing model according to the sample image set, the sample mask image set, and the description text set includes:
[0101] The training of the preset editing model according to the sample image set, the sample mask image set, and the description text set includes:
[0102] Optionally, the obtaining module 510 is specifically configured to:
[0103] The original image is displayed in the editing interface;
[0104] The region selection operation for the original image is obtained, and the to-be-edited region in the original image is determined according to the region selection operation;
[0105] The text input operation for the to-be-edited region is obtained, and the target prompt word corresponding to the to-be-edited region is determined according to the text input operation.
[0106] Optionally, the noise adding module 520 is specifically configured to:
[0107] The target mask image is determined according to the foreground content in the to-be-edited region;
[0108] The latent feature map of the original image is generated by the image editing model, and the to-be-edited region in the latent feature map is determined in combination with the target mask image and the latent feature map;
[0109] The noise adding operation is performed a set number of times based on the to-be-edited region in the latent feature map, and a local noise image is obtained.
[0110] Further, the noise adding operation is performed a set number of times based on the to-be-edited region in the latent feature map, and a local noise image is obtained, including:
[0111] The noise feature map is obtained, wherein the noise feature map represents a preset noise;
[0112] The noise feature region is determined in combination with the target mask image and the noise feature map;
[0113] The noise adding operation is performed a set number of times based on the to-be-edited region and the noise feature region, and a local noise image is obtained.
[0114] Optionally, the image generation module 530 is specifically configured to:
[0115] determine, by the image editing model, predicted noise according to image features in the to-be-edited region of the local noise image;
[0116] generate, by the image editing model, text features based on the target prompt word, and determine relevance of the text features to the local noise image;
[0117] generate a local edited image according to the relevance and the local noise image, perform a denoising operation on the local edited image based on the predicted noise, and output a target image.
[0118] Optionally, the image generation module 530 further comprises:
[0119] a grade setting module configured to obtain mask setting information, the mask setting information comprising a blur grade;
[0120] Further, the generating a local edited image according to the relevance and the local noise image comprises:
[0121] determining a filling degree of a generated object to the to-be-edited region based on the blur grade;
[0122] generating a local edited image according to the relevance, the filling degree and the local noise image, wherein the to-be-edited region in the local edited image comprises the generated object.
[0123] The image editing apparatus provided by the embodiments of the present disclosure can perform the image editing method provided by any of the embodiments of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0124] It should be noted that each unit and module included in the above apparatus is only divided according to the function logic, and is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for convenient mutual distinction, and is not used to limit the protection scope of the embodiments of the present disclosure.
[0125] FIG. 6 is a structural diagram of an electronic device according to an embodiment of the disclosure. Below, referring to FIG. 6, a structural diagram of an electronic device (e.g., a terminal device or a server in FIG. 6) 600 suitable for implementing an embodiment of the disclosure is illustrated. The terminal device in an embodiment of the disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device illustrated in FIG. 6 is merely an example, and should not impose any limitation on the functions and use range of an embodiment of the disclosure.
[0126] As illustrated in FIG. 6, the electronic device 600 can include a processing device (e.g., a central processing unit, a graphic processing unit, or the like) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0127] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, or the like; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, or the like; a storage device 608 including, for example, a magnetic tape, a hard disk, or the like; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or via a wire to exchange data. Although FIG. 6 illustrates the electronic device 600 having various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0128] In particular, according to an embodiment of the disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the disclosure includes a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-described functions defined in the methods of an embodiment of the disclosure are performed.
[0129] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0130] The electronic device provided by the embodiments of the present disclosure and the image editing method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.
[0131] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the image editing method provided by the above embodiments.
[0132] It should be noted that the computer readable medium of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0133] In some embodiments, the client, server, or both can communicate using any known or future developed network protocols, such as the HyperText Transfer Protocol (HTTP), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current or future developed networks.
[0134] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and can be accessed via the electronic device.
[0135] The computer-readable medium described above carries one or more programs that, when executed by the electronic device, cause the electronic device to perform:
[0136] Obtain a to-be-edited region in an original image and a target prompt word corresponding to the to-be-edited region, wherein the target prompt word is text information used to describe an expected effect of image editing;
[0137] Determine a target mask image according to the to-be-edited region, add a preset noise to the to-be-edited region based on the target mask image through an image editing model, and obtain a local noise image;
[0138] Perform noise prediction processing and image generation processing on the to-be-edited region of the local noise image based on the target prompt word through the image editing model, and output a target image according to a noise prediction result and an image generation result.
[0139] Computer program code for carrying out operations of the present disclosure can be written in any one or combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, application specific circuitry, or field programmable gate array ("FPGA") circuitry can execute the program code.
[0140] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The computer program product of the first aspect can include a non-transitory computer-readable medium storing code that, when executed, causes a computer to perform operations for the first aspect.
[0141] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0142] The functions described in this document can be implemented in hardware, software, or any combination thereof. In some embodiments, the functions described in this document can be implemented in one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0143] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0144] The above description merely illustrates the preferred embodiments of the disclosure and a principle for applying the technologies. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0145] Further, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included for the purpose of providing a thorough disclosure, these should not be construed as limitations on the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0146] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An image editing method, comprising: obtaining an editing region in an original image and a target prompt word corresponding to the editing region, wherein the target prompt word is text information used to describe an expected effect of image editing; determining a target mask image according to the editing region, adding preset noise to the editing region based on the target mask image through an image editing model, and obtaining a local noise image; performing noise prediction processing and image generation processing on the editing region of the local noise image based on the target prompt word through the image editing model, and outputting a target image according to a noise prediction result and an image generation result.
2. The method of claim 1, wherein, The training method of the image editing model comprises: obtaining a sample image set; for a sample image in the sample image set, determining a target region in the sample image, performing an inflation operation on the target region to obtain a sample inflation image, and determining a sample mask image set according to the sample inflation image corresponding to the sample image set; performing image content recognition on the target region to obtain a description text corresponding to the target region, and determining a description text set according to the description text of the target region corresponding to the sample image set; training a preset editing model according to the sample image set, the sample mask image set and the description text set to obtain an image editing model.
3. The method of claim 2, wherein, The determination of the sample mask image set according to the sample inflation image corresponding to the sample image set comprises: for the sample inflation image corresponding to the sample image set, generating a reference mask image of the sample image according to the target region in the sample inflation image; performing Gaussian blur processing on the reference mask image to obtain at least two sample mask images corresponding to the sample image; determining a sample mask image set according to at least two sample mask images corresponding to the sample image in the sample image set, wherein the at least two sample mask images are associated with the description text.
4. The method of claim 3, wherein, After the Gaussian blur processing of the reference mask image to obtain at least two sample mask images corresponding to the sample image, the method further comprises: determining a blur level of the sample mask image according to the size of the Gaussian kernel used for Gaussian blur processing, and marking the corresponding sample mask image according to the blur level; The training of the preset editing model according to the sample image set, the sample mask image set and the description text set to obtain the image editing model comprises: training the preset editing model according to the sample image set, the sample mask image set, the blur level of the sample mask image corresponding to the sample image set and the description text set to obtain the image editing model.
5. The method according to any one of claims 1 to 4, wherein, The obtaining of the editing region in the original image and the target prompt word corresponding to the editing region comprises: displaying the original image in an editing interface; obtaining a region selection operation for the original image, and determining the editing region in the original image according to the region selection operation; obtaining a text input operation for the editing region, and determining the target prompt word corresponding to the editing region according to the text input operation.
6. The method according to any one of claims 1 to 5, wherein, The determination of the target mask image according to the editing region, the addition of the preset noise to the editing region based on the target mask image through the image editing model, and the obtaining of the local noise image comprise: determine a target mask according to foreground content in the region to be edited; generate a latent feature map of the original image through the image editing model, and determine a region to be edited in the latent feature map in combination with the target mask and the latent feature map; perform a preset number of noise addition operations based on the region to be edited in the latent feature map to obtain a local noise image.
7. The method of claim 6, wherein, The method for performing a preset number of noise addition operations based on the region to be edited in the latent feature map to obtain a local noise image comprises: obtain a noise feature map, wherein the noise feature map represents a preset noise; determine a noise feature region in combination with the target mask and the noise feature map; perform a preset number of noise addition operations based on the region to be edited and the noise feature region to obtain a local noise image.
8. The method according to any one of claims 1 to 7, wherein, The method for performing noise prediction processing and image generation processing on the region to be edited of the local noise image based on the target prompt word through the image editing model, and outputting a target image according to a noise prediction result and an image generation result comprises: determine a predicted noise according to an image feature in the region to be edited of the local noise image through the image editing model; generate a text feature based on the target prompt word through the image editing model, and determine a relevance between the text feature and the local noise image; generate a local editing image according to the relevance and the local noise image, and output a target image by performing a denoising operation on the local editing image based on the predicted noise.
9. The method of claim 8, further comprising: obtaining mask setting information, wherein the mask setting information comprises a blur level; The method for generating a local editing image according to the relevance and the local noise image comprises: determining a filling degree of a generated object to the region to be edited based on the blur level; generating a local editing image according to the relevance, the filling degree, and the local noise image, wherein the region to be edited in the local editing image comprises the generated object.
10. An image editing apparatus, comprising: an obtaining module configured to obtain a region to be edited in an original image and a target prompt word corresponding to the region to be edited, wherein the target prompt word is text information used to describe an expected effect of image editing; a noise addition module configured to determine a target mask according to the region to be edited, and add a preset noise to the region to be edited based on the target mask through an image editing model to obtain a local noise image; an image generation module configured to perform noise prediction processing and image generation processing on the region to be edited of the local noise image based on the target prompt word through the image editing model, and output a target image according to a noise prediction result and an image generation result.
11. An electronic device, comprising: one or more processors; a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the image editing method of any one of claims 1-9.
12. A storage medium containing computer-executable instructions, wherein, The computer executable instructions, when executed by a computer processor, are for performing the image editing method as claimed in any one of claims 1-9.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116543075A
Image generation model training method and device, equipment and storage medium
CN116958324A
Image editing method, device and equipment and readable storage medium
CN117372574A
Image processing method and device, electronic equipment and storage medium
CN117541511A
Image generation method and device, electronic equipment and storage medium
CN117670658A
Cited By
Image processing method and device based on large model, medium, electronics and product
CN121392544A
Streetscape ground object real-time semantic segmentation and geographic positioning method and system based on camera
CN121558064A
Synthetic leather defect image generation and semantic annotation method
CN122089883A
Instructional image editing method, device, storage medium and program product
CN122453977A