Image editing method, image editing device and storage medium
By replacing the image area by the bidirectional diffusion network model, the problem of poor image segmentation and generation effects in the prior art is solved, and high-quality image editing and fusion effects are achieved.
Patent Information
- Application Number
- CN202311568933.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
Smart Images

Figure CN120032017A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to an image editing method, an image editing device and a storage medium. Background Art
[0002] Image segmentation and generation are two important research directions in the field of computer vision.
[0003] In the related technology, image segmentation is the process of dividing an image into several parts or objects. The current image segmentation technology is mainly divided into the following categories: based on traditional image processing methods, such as threshold segmentation, edge detection, region growing, etc. These methods are simple and easy to understand, but they are not effective for complex image segmentation tasks. Image generation refers to the generation of new images through computer programs. The current image generation technology is mainly divided into the following categories: based on traditional image processing methods, such as interpolation, texture synthesis, image fusion, etc. These methods can generate some simple images, but they are not effective for complex image generation tasks, resulting in inaccurate segmentation results or low quality of generated images. Summary of the invention
[0004] In order to overcome the problems existing in the related art, the present disclosure provides an image editing method, an image editing device and a storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image editing method, comprising: acquiring a first image and a second image; replacing a first region image in the first image with a second region image in the second image based on a bidirectional diffusion network model to generate a target image; wherein the bidirectional diffusion network model comprises a first diffusion model and a second diffusion model, the first diffusion model is used to perform feature modeling encoding and decoding modeling on the first region image, the second diffusion model is used to fuse the second region image into the first region image, and a first encoding part of the second diffusion model and a first encoding part of the first diffusion model share the same network parameters, the second diffusion model and the first diffusion model perform network parameter fusion in the second part, the second part is the part of the first diffusion model other than the first part and the third part, and the third part is a decoding part of the first diffusion model that is symmetrical to the first part.
[0006] In one embodiment, based on the bidirectional diffusion network model, the first area image in the first image is replaced with the second area image in the second image to generate a target image, including: determining the second area image selected by the user on the second image, and upsampling the second area image to a third image of the same size as the first image; masking the first area image corresponding to the second area image in the first image, and adding noise to the masked image to generate a fourth image; inputting the third image and the fourth image into the bidirectional diffusion network model, and generating a target image based on the bidirectional diffusion network model, wherein the target image is an image obtained by fusing the third image with the masked area image in the fourth image.
[0007] In one embodiment, generating a target image based on the bidirectional diffusion network model includes: inputting the fourth image into the first diffusion model of the bidirectional diffusion network model, and inputting the third image into the second diffusion model of the bidirectional diffusion network model; encoding the fourth image based on the first encoding part of the first diffusion model to obtain a first encoded image, and encoding the third image based on the network parameters of the first encoding part shared by the second diffusion model to obtain a second encoded image; fusing the second part of the first diffusion model and the second part of the second diffusion model, fusing the first encoded image and the second encoded image to obtain a fused feature, the second part of the first diffusion model and the second part of the second diffusion model having the same network structure; decoding the fused feature based on the third part of the first diffusion model to obtain a target image.
[0008] In one embodiment, the bidirectional diffusion network model is trained in the following manner: initialize a first diffusion model and a second diffusion model; during the i-th training, update all parameters of the first diffusion model through a back propagation process, and update the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model with the same learning rate, and keep unchanged, wherein i is a positive integer; during the i+1-th training, update the parameters of the first encoding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first encoding part through a back propagation process, and update the second part of the second diffusion model and the second part of the first diffusion model by fusing the parameters of the i-th training; alternately perform the i-th training and the i+1-th training process until a bidirectional diffusion network model that meets the constraints is obtained through training.
[0009] In one embodiment, decoding the fused features based on the third part of the first diffusion model to obtain a target image includes: decoding the fused features based on the third part of the first diffusion model to obtain a sixth image; and guiding the sixth image based on the first image to obtain a target image that meets the guidance conditions.
[0010] In one embodiment, the second diffusion model and the first diffusion model perform network parameter fusion in the second part based on a cross-attention mechanism; wherein the cross-attention mechanism fuses the network parameters of the first diffusion model and the second diffusion model based on the channel dimension.
[0011] According to a second aspect of an embodiment of the present disclosure, there is provided an image editing device, comprising: an acquisition unit, for acquiring a first image and a second image; a processing unit, for replacing a first region image in the first image with a second region image in the second image based on a bidirectional diffusion network model to generate a target image; wherein the bidirectional diffusion network model comprises a first diffusion model and a second diffusion model, the first diffusion model is used to perform feature modeling encoding and decoding modeling on the first region image, the second diffusion model is used to fuse the second region image into the first region image, and the first encoding part of the second diffusion model shares the same network parameters with the first encoding part of the first diffusion model, the second diffusion model and the first diffusion model perform network parameter fusion in the second part, the second part is the part of the first diffusion model other than the first part and the third part, and the third part is the decoding part of the first diffusion model that is symmetrical to the first part.
[0012] In one embodiment, the processing unit replaces the first area image in the first image with the second area image in the second image based on a bidirectional diffusion network model to generate a target image in the following manner: determine the second area image selected by the user on the second image, and upsample the second area image to a third image of the same size as the first image; mask the first area image corresponding to the second area image in the first image, and add noise to the masked image to generate a fourth image; input the third image and the fourth image into the bidirectional diffusion network model, and generate a target image based on the bidirectional diffusion network model, wherein the target image is an image obtained by fusing the third image with the masked area image in the fourth image.
[0013] In one embodiment, the processing unit generates a target image based on the bidirectional diffusion network model in the following manner: input the fourth image into the first diffusion model of the bidirectional diffusion network model, and input the third image into the second diffusion model of the bidirectional diffusion network model; encode the fourth image based on the first encoding part of the first diffusion model to obtain a first encoded image, and encode the third image based on the network parameters of the first encoding part shared by the second diffusion model to obtain a second encoded image; fuse the second part of the first diffusion model and the second part of the second diffusion model, fuse the first encoded image and the second encoded image to obtain a fused feature, the second part of the first diffusion model and the second part of the second diffusion model have the same network structure; decode the fused feature based on the third part of the first diffusion model to obtain a target image.
[0014] In one embodiment, the bidirectional diffusion network model is trained in the following manner: initialize a first diffusion model and a second diffusion model; during the i-th training, update all parameters of the first diffusion model through a back propagation process, and update the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model with the same learning rate, and keep unchanged, wherein i is a positive integer; during the i+1-th training, update the parameters of the first encoding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first encoding part through a back propagation process, and update the second part of the second diffusion model and the second part of the first diffusion model by fusing the parameters of the i-th training; alternately perform the i-th training and the i+1-th training process until a bidirectional diffusion network model that meets the constraints is obtained through training.
[0015] In one embodiment, the processing unit decodes the fused features based on the third part of the first diffusion model to obtain a target image in the following manner: decodes the fused features based on the third part of the first diffusion model to obtain a sixth image; and guides the sixth image based on the first image to obtain a target image that meets the guidance conditions.
[0016] In one embodiment, the processing unit adopts the following method to fuse the network parameters of the second diffusion model and the first diffusion model in the second part based on the cross-attention mechanism; wherein, the cross-attention mechanism is based on the channel dimension to fuse the network parameters of the first diffusion model and the second diffusion model.
[0017] According to a third aspect of an embodiment of the present disclosure, there is provided an image editing device, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor is configured to: execute the image editing method described in the first aspect or any one of the implementations of the first aspect.
[0018] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor, the processor is enabled to execute the image editing method described in the first aspect or any one of the embodiments of the first aspect.
[0019] The technical solution provided by the embodiment of the present disclosure may include the following beneficial effects: by acquiring the first image and the second image, based on the bidirectional diffusion network model, the first region image in the first image is replaced with the second region image in the second image to generate a target image. Through the present disclosure, the edited image better fits the image quality and content style of the original image, produces a lossless image editing effect, ensures the quality of the generated target image, and improves the fusion between images.
[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0022] Figure 1 The figure is a flowchart of an image editing method according to an exemplary embodiment.
[0023] Figure 2 is a schematic diagram of an image editing method according to an exemplary embodiment.
[0024] Figure 3 The present invention is a flowchart of a method for generating a target image based on a bidirectional diffusion model according to an exemplary embodiment.
[0025] Figure 4 The figure is a flowchart of a method for generating a target image based on a model according to an exemplary embodiment.
[0026] Figure 5 The figure is a flow chart of a method for obtaining a target image according to an exemplary embodiment.
[0027] Figure 6 The figure is a schematic diagram of a virtual try-on according to an exemplary embodiment.
[0028] Figure 7The invention is a block diagram of an image editing device according to an exemplary embodiment.
[0029] Figure 8 The invention is a block diagram of an apparatus for image editing according to an exemplary embodiment. DETAILED DESCRIPTION
[0030] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure.
[0031] Image segmentation and image generation are two important research directions. Image segmentation is the process of dividing an image into several parts or objects. Image generation refers to the generation of new images through computer programs. In one example, the scene of image segmentation and image generation can be virtual try-on. Assuming that given an image depicting a person and an image of another person wearing different clothes, this virtual clothing try-on aims to visualize the appearance of clothing on a person based on the image of the person and the image of the clothing, for example: virtual try-on has the potential to enhance the online shopping experience, but most try-on methods are only effective when the body posture and body shape change is small. And either focus on detail retention without ensuring effective coordination and shape change, or forcibly retain the required shape and lack of details at the junction of the images. For complex scenes, such as occlusion, lighting changes, background noise, etc., the performance of image segmentation and generation technology may be affected, resulting in inaccurate segmentation results or low quality of generated images. This application takes the example of changing the image of a person as an example, but it is not limited to it. It can be the synthesis and replacement between other images, replacing the specified part of the reference image with an image consistent with the style of the original image.
[0032] Furthermore, the models used in related technologies are relatively complex and it is difficult to explain their internal working principles, which may affect their reliability and credibility in certain application scenarios. In terms of computing resource requirements, some image segmentation and generation technologies require a large amount of computing resources to train and infer models, which may limit their application on certain devices, such as mobile devices. At the same time, some image segmentation and generation technologies may perform better for certain specific tasks, but may perform poorly on other tasks, which may require model optimization and adjustment for different tasks.
[0033] In view of this, the present disclosure proposes an image editing method.
[0034] The disclosed embodiment describes an image editing method.
[0035] Figure 1FIG. 1 is a flowchart of an image editing method according to an exemplary embodiment. Figure 1 As shown, the method includes the following.
[0036] In step S11 , a first image and a second image are acquired.
[0037] In the disclosed embodiment, the first image can be understood as an original input image, and the second image can be understood as a reference image, wherein the original input image is an image to be processed input by a user, and the reference image is an image used to guide the original input image to generate a target image.
[0038] In step S12, based on the bidirectional diffusion network model, the first region image in the first image is replaced with the second region image in the second image to generate a target image.
[0039] In the embodiment of the present disclosure, the first region image can be understood as a region image that does not participate in the modification of the first image. The second region image can be understood as a region in the second image selected to be fused with the first image, as the second region image. The first region image in the first image is replaced with the second region image in the second image to generate a first image fused with the second region image, and the first image is used as the target image.
[0040] In the present disclosure, Figure 2 As shown, Figure 2 is a schematic diagram of an image editing method according to an exemplary embodiment. The bidirectional diffusion network model includes a first diffusion model and a second diffusion model. The first diffusion model includes a first part, a second part and a third part. The first part of the first diffusion model is Figure 2 The first five parts of the upper part, the second part of the first diffusion model is Figure 2 The sixth to twelfth parts in the middle of the upper part, the third part of the first diffusion model is Figure 2 From the last five parts of the first half, it can be understood that the first part is symmetrical with the third part. The first part is the first encoding part, the second part is the part of the first diffusion model other than the first part and the third part, and the third part is the decoding part of the first diffusion model that is symmetrical with the first part. The second diffusion model includes the first part and the second part. The first part of the second diffusion model is Figure 2 The first five parts of the second half, Figure 2 The remaining area of the lower half is the second part.
[0041] In the disclosed embodiment, the first diffusion model is used to perform feature modeling encoding and decoding modeling on the first region image, the second diffusion model is used to fuse the second region image to the first region image, and the first encoding part of the second diffusion model and the first encoding part of the first diffusion model share the same network parameters. The second diffusion model and the first diffusion model perform network parameter fusion in the second part.
[0042] In the disclosed embodiment, based on the bidirectional diffusion model, the first region image in the first image is replaced with the second region image in the second image to generate the target image. It can be used for filling missing parts and removing noise in image restoration. Based on the bidirectional diffusion network model, the first image and the second image are fused through the first diffusion model and the second diffusion model to solve the problem of image deformation of different styles, improve image editing capabilities, and achieve accurate generation of details to meet user needs.
[0043] The following is an explanation of a method for generating a target image based on a bidirectional diffusion model in an embodiment of the present disclosure.
[0044] Figure 3 FIG. 1 is a flow chart showing a method for generating a target image based on a bidirectional diffusion model according to an exemplary embodiment. Figure 3 As shown, the method includes the following.
[0045] In step S21 , a second region image selected by a user on the second image is determined, and the second region image is up-sampled into a third image of the same size as the first image.
[0046] In an embodiment of the present disclosure, in response to a user selecting a fused feature region in the second image, i.e., a second region image, the second region image is upsampled to the same dimension as the first image to ensure that the second region image has the same size as the first image, and the upsampled second region image is used as the third image.
[0047] In step S22, masking is performed on the first region image corresponding to the second region image portion of the first image, and noise is added to the masked image to generate a fourth image.
[0048] In the embodiment of the present disclosure, a masking process is performed on a first region image corresponding to a first image, wherein the first region image corresponds to a second region image. A masking operation is performed on the first region image so as to replace the second region image of the second image with the first region image, and by masking the image, it is ensured that the region not involved in the modification is not disturbed. Noise is added to the masked first region image, and the generated image is used as the fourth image.
[0049] In step S23, the third image and the fourth image are input into the bidirectional diffusion network model, and a target image is generated based on the bidirectional diffusion network model.
[0050] In the disclosed embodiment, the third image obtained by upsampling and the fourth image obtained after masking are input into the bidirectional diffusion network model together, and the target image is generated based on the bidirectional diffusion network model.
[0051] In the disclosed embodiment, based on the processing of the first region image and the second region image, the editing content and the generated content are accurately edited, which increases the fineness of the target image and improves the generation effect of the target image.
[0052] The following is an explanation of a method for generating a target image based on a model in an embodiment of the present disclosure.
[0053] Figure 4 FIG. 1 is a flow chart showing a method for generating a target image based on a model according to an exemplary embodiment. Figure 4 As shown, the method includes the following.
[0054] In step S31 , the fourth image is input into the first diffusion model of the bidirectional diffusion network model, and the third image is input into the second diffusion model of the bidirectional diffusion network model.
[0055] In the disclosed embodiment, the first region image corresponding to the second region image portion of the first image is masked, and noise is added to the masked image, and the generated fourth image is input to the first diffusion model of the bidirectional diffusion network model. The second region image is upsampled to a third image of the same size as the first image and input to the second diffusion model of the bidirectional diffusion network model.
[0056] In step S32, the fourth image is encoded based on the first encoding part of the first diffusion model to obtain a first encoded image, and the third image is encoded based on the second diffusion model sharing the network parameters of the first encoding part to obtain a second encoded image.
[0057] In the disclosed embodiment, the first part of the first diffusion model is a first coding part, and the first coding part is used to model the original features. The fourth image is encoded based on the first coding part of the first diffusion model, and the obtained image is confirmed as the first coded image. The third image is encoded based on the network parameters of the first coding of the first diffusion model shared by the second diffusion model, and the obtained image is used as the second coded image.
[0058] In step S33, the second part of the first diffusion model and the second part of the second diffusion model are fused, and the first encoded image and the second encoded image are fused to obtain a fused feature.
[0059] In the disclosed embodiment, generating a target image requires updating the relevant parameters of the image by means of parameter fusion in the bidirectional diffusion network model. This can be understood as fusing the second part of the first diffusion model and the second part of the second diffusion model, fusing the first encoded image and the second encoded image, and performing parameter updates with the same learning rate to obtain fused features.
[0060] In step S34, based on the third part of the first diffusion model, the fused features are decoded to obtain a target image.
[0061] In the disclosed embodiment, the third part of the first diffusion model is a decoding part symmetrical to the first part in the first diffusion model. The third part is used to decode and model the image. Based on the third part, the fusion features obtained by fusing the first encoded image and the second encoded image are decoded to obtain an image, and the obtained image is used as the target image.
[0062] In the disclosed embodiment, a target image is obtained based on fusion features of a first decoded image and a second decoded image. Modeling is performed based on the fusion features of the fused image to gradually establish image details. The model effect training is easy, the generation effect is stable, and there is no need for cross-modality, which reduces the guidance difficulty of the model while ensuring the generation quality of the target image.
[0063] The following is an explanation of a method for training a bidirectional diffusion network model in an embodiment of the present disclosure.
[0064] In the disclosed embodiment, when the second diffusion model maps features to the same dimension through the same first part as the first diffusion model, it is sent to the second diffusion model, and the second part of the first diffusion model and the second part of the second diffusion model have the same parameter settings and initialization at the corresponding structural positions.
[0065] In the disclosed embodiment, when the i-th training is performed, all parameters of the first diffusion model are updated through the back propagation process. It can be understood that i is a positive integer, and all parameters of the first part, the second part, and the third part of the first diffusion model are updated. And the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model are updated with the same learning rate, and remain unchanged.
[0066] In the disclosed embodiment, when the i+1th training is performed, the parameters of the first coding part shared by the second diffusion model and the first diffusion model are updated through the back propagation process, and the parameters of the third part corresponding to the first coding part in the first diffusion model are updated. Where i is a positive integer. The second part of the second diffusion model and the second part of the first diffusion model are updated by means of the parameters of the previous stage, that is, the parameters of the i-th training are integrated and updated.
[0067] In the disclosed embodiment, in the subsequent training process, the i-th training process and the (i+1)-th training process are performed alternately, and the training is performed in an alternating updating manner until a bidirectional diffusion network model that satisfies the constraint conditions is obtained through training.
[0068] In the disclosed embodiment, based on creating and training a bidirectional diffusion network model, the first encoded image and the second encoded image are modeled separately, and the high-dimensional parameters of the bidirectional diffusion network model are modeled by parameter fusion and the low-dimensional parameters are modeled by parameter sharing. This increases the fineness of the target image, reduces the amount of parameters and calculations, and generates a lossless editing effect for the target image.
[0069] The method for obtaining the target image is described below in the embodiment of the present disclosure.
[0070] Figure 5 FIG. 1 is a flow chart showing a method for obtaining a target image according to an exemplary embodiment. Figure 5 As shown, the method includes the following.
[0071] In step S41, based on the third part of the first diffusion model, the fusion feature is decoded to obtain a sixth image.
[0072] In the disclosed embodiment, the decoding part symmetrical to the first part of the first diffusion model is the third part, and the fusion feature is decoded based on the third part of the first diffusion model, wherein the fusion feature is a feature obtained by fusing the first coded image and the second coded image by fusing the second part of the first diffusion model and the second part of the second diffusion model. The image obtained by decoding the fusion feature based on the third part of the first diffusion model is used as the sixth image.
[0073] In step S42, the sixth image is guided based on the first image to obtain a target image that meets the guidance conditions.
[0074] In the disclosed embodiment, the sixth image is guided based on the first image, that is, the original input image, image details are established, and high-dimensional features are compressed to restore the image dimension of the first image, so as to obtain an image that meets the guidance conditions and is determined as the target image.
[0075] In the disclosed embodiment, guidance is provided based on the image, so that the generation effect is stable, the image quality of the generated target image and the quality of the target image are guaranteed, and the generation effect of the target image is improved.
[0076] The cross-attention mechanism is described below in the embodiments of the present disclosure.
[0077] In the disclosed embodiment, in the process of obtaining the target image by training the bidirectional diffusion network model, the second part of the first diffusion model and the second part of the second diffusion model are subjected to network parameter fusion based on the cross attention mechanism. The cross attention mechanism is based on the channel dimension, that is, an attention mechanism module based on the channel dimension is added, and the parameter training of the attention mechanism module is added to the training form of the first diffusion model and the second diffusion model, and parameter fusion is performed to fuse the network parameters of the first diffusion model and the second diffusion model, and further establish the relationship between the first image and the second image.
[0078] In the disclosed embodiment, a cross-attention mechanism of a bidirectional diffusion network model is proposed. By establishing a cross-attention method between the first image and the second image instead of a sequence of two separate tasks, the algorithm performance is further improved, and the generation effect of the target image is improved. It can also be applied to mobile video editing applications to achieve background replacement, special effects addition, editing and replacement of a part of the image, style transfer, etc. in the video, thereby improving the fusion of images and videos.
[0079] The following is an explanation of the image editing method in the embodiment of the present disclosure.
[0080] Figure 6 FIG. 1 is a schematic diagram of a virtual try-on according to an exemplary embodiment. Figure 6 As shown, the method includes the following.
[0081] In the disclosed embodiment, Unet-R is the first diffusion model, and Unet-S is the second diffusion model. In one example, assume that an image depicting a person and an image of another person wearing different clothes are given. Through the bidirectional diffusion network model, a masking area is established based on regional segmentation, and based on the generation technology of the bidirectional diffusion network model, the clothing of another person can be generated and transferred to the input person. Based on the bidirectional diffusion network model, the first image and the second image are combined, the details of the first image are retained in a single network and the two images are combined with the second image to achieve the fusion of the two images to adapt to the complete state and shape changes of the original subject. And based on the open source network, the posture style of the character is controlled to establish a delicate and harmonious generation result. The present application takes the example of changing the clothes of a character image as an example, but it is not limited to it. It can be a synthesis and replacement between other images, and the specified part of the reference image is replaced with an image consistent with the style of the original image.
[0082] In the present disclosure, Figure 2As shown, for the first image, before being sent to the bidirectional diffusion network model, a masking operation is first performed. The purpose is to expand the relevant content in the second image, that is, the second region image, to the first image later, and to ensure that the components not involved in the modification do not interfere, so it is necessary to perform a masking operation on the first region image. For the masked first image, noise is regenerated and sent to the first diffusion model. Similarly, for the second image, it is necessary to select the fused feature area, that is, the second region image, and at the same time sample the area to the same dimension as the first image by upsampling, so that the parameters of the first encoding part can be shared with the first diffusion model. The masking process and the region selection upsampling process of the second region image are both data preprocessing processes. Specifically, through the entire preprocessing process, the first region image and the second region image are "segmented" and initialized respectively. In an example, it can be initialized to a basic element image size of 128*128, and the processed first image is further used as input and sent to the bidirectional diffusion network model together with the processed second image as input.
[0083] In the disclosed embodiment, for the preprocessed first image, it is sent to the first diffusion model, and the structure of the first diffusion model is similar to the funnel-shaped structure of the original Unet, which includes modeling the original features and decoding modeling parts. At the same time, for the preprocessed second image, the features are first mapped to the same dimension as the input features, and then sent to the second diffusion model part. It should be noted that the first diffusion model and the second part of the second diffusion model have the same parameter settings and initialization at the corresponding structural positions, but the parameters are updated by parameter fusion in the later training process. Specifically, in the i-th training, the first diffusion model is fully parameterized through the back propagation process, and the parameters of the corresponding part of the second diffusion model with the first diffusion model are updated with the same learning rate. The remaining parts of the structure are not updated with parameters. Each time, the parameters of the previous training process are retained by sharing parameters with the first diffusion model. In the i+1th training, the parameters of the first encoding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first encoding part are updated through the back propagation process, and the second part of the second diffusion model and the second part of the first diffusion model are updated by fusing the parameters of the i-th training. That is, in the current step, the symmetric parts of the first diffusion model and the second diffusion model will use the same parameters as new network parameters. In the subsequent training process, the training is performed by alternating the two methods.
[0084] In the disclosed embodiment, the essence of the third part of the first diffusion model is not only to restore the modeling of the first image, but its essence is to model the fusion features after parameter fusion, and gradually establish the content details of the second image in the subsequent diffusion process. The difference from the original diffusion generation model is that it uses pictures for guidance, and fuses the guidance content into the first diffusion model through parameter fusion. The guidance method of the traditional diffusion model is text. In the field of this technology, the generation method based on text guidance requires cross-modality compared to the generation method based on picture reference, which increases the difficulty of guidance. Compared with the guidance method of the same modality, the effect model is difficult to train and the generation effect is unstable.
[0085] In the disclosed embodiment, in the specific implementation, in the two groups of diffusion models, an attention mechanism module based on the channel dimension is added, and at the same time, the parameter training of the attention mechanism module is added to the parameter fusion method in the above training form. That is, by fusing the attention mechanism parameters of the two models, the interaction between the two networks is further increased, that is, the relationship between the first image and the second image is further established. After the first diffusion model is output, the high-dimensional features need to be compressed and the dimensions of the original image need to be restored. In order to obtain a more harmonious image in the application scenario of the human role, the character posture control is performed by adding a plug-in, which is conducive to generating a more realistic and harmonious character image in the character scene, and finally generating the target image.
[0086] In the disclosed embodiments, the image editing method can be applied to scenes such as image restoration, artistic creation, and video editing. Among them, image restoration can be used for filling missing parts and removing noise in image restoration, and the content to be restored can be customized to achieve convenient and free restoration of images. Artistic creation can be used for image generation, style conversion, etc. in artistic creation, especially for style transfer, clothing replacement, color change, etc. based on a certain image. Video editing can be used in mobile video editing applications, and can be used to achieve functions such as background replacement, special effect addition, editing and replacement of a certain part of the image, and style transfer in the video. The target image generated based on the bidirectional diffusion network model improves the image editing ability and can accurately generate image details, improves the quality of the editing content and the generated content, including characters, scenes, actions, etc., so as to accurately edit the image. At the same time, the edited image or video content can be made to fit the image quality and content style of the original image or original video, and produce a lossless image and video editing effect. And based on the image guidance method and the modeling method of parameter feature fusion, the image quality and image content of the generated image content are guaranteed, and the loss of the generated target image or target video in terms of fusion, saturation, color difference, etc. is reduced.
[0087] Based on the same concept, the embodiment of the present disclosure also provides an image editing device 100 .
[0088] It is understandable that, in order to realize the above functions, the image editing device 100 provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.
[0089] Figure 7 FIG. 1 is a block diagram of an image editing device 100 according to an exemplary embodiment. Figure 7 The device includes an acquisition unit 101 and a processing unit 102.
[0090] The acquisition unit 101 is used to acquire a first image and a second image.
[0091] The processing unit 102 is used to replace the first region image in the first image with the second region image in the second image based on the bidirectional diffusion network model to generate a target image.
[0092] Among them, the bidirectional diffusion network model includes a first diffusion model and a second diffusion model. The first diffusion model is used to perform feature modeling encoding and decoding modeling on the first area image, and the second diffusion model is used to fuse the second area image into the first area image. The first encoding part of the second diffusion model and the first encoding part of the first diffusion model share the same network parameters. The second diffusion model and the first diffusion model perform network parameter fusion in the second part. The second part is the part of the first diffusion model except the first part and the third part, and the third part is the decoding part of the first diffusion model that is symmetrical to the first part.
[0093] In one embodiment, the processing unit 102 replaces the first area image in the first image with the second area image in the second image based on a bidirectional diffusion network model to generate a target image in the following manner: determine the second area image selected by the user on the second image, and upsample the second area image to a third image of the same size as the first image; mask the first area image corresponding to the second area image in the first image, and add noise to the masked image to generate a fourth image; input the third image and the fourth image into the bidirectional diffusion network model, and generate a target image based on the bidirectional diffusion network model, wherein the target image is an image obtained by fusing the third image into the masked area image in the fourth image.
[0094] In one embodiment, the processing unit 102 generates a target image based on a bidirectional diffusion network model in the following manner: input the fourth image into the first diffusion model of the bidirectional diffusion network model, and input the third image into the second diffusion model of the bidirectional diffusion network model; encode the fourth image based on the first encoding part of the first diffusion model to obtain a first encoded image, and encode the third image based on the network parameters of the first encoding part shared by the second diffusion model to obtain a second encoded image; fuse the second part of the first diffusion model and the second part of the second diffusion model, fuse the first encoded image and the second encoded image to obtain a fused feature, and the second part of the first diffusion model and the second part of the second diffusion model have the same network structure; decode the fused feature based on the third part of the first diffusion model to obtain a target image.
[0095] In one embodiment, the bidirectional diffusion network model is trained in the following manner: the first diffusion model and the second diffusion model are initialized; during the i-th training, all parameters of the first diffusion model are updated through a back propagation process, and the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model are updated with the same learning rate, and remain unchanged, i is a positive integer; during the i+1-th training, the parameters of the first coding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first coding part are updated through a back propagation process, and the second part of the second diffusion model and the second part of the first diffusion model are updated by fusing the parameters of the i-th training; the i-th training and the i+1-th training processes are performed alternately until a bidirectional diffusion network model that meets the constraint conditions is obtained through training.
[0096] In one embodiment, the processing unit 102 decodes the fused features based on the third part of the first diffusion model to obtain a target image in the following manner: decodes the fused features based on the third part of the first diffusion model to obtain a sixth image; and guides the sixth image based on the first image to obtain a target image that meets the guidance conditions.
[0097] In one embodiment, the processing unit 102 adopts the following method to fuse the network parameters of the second diffusion model and the first diffusion model in the second part based on the cross-attention mechanism; wherein, based on the cross-attention mechanism, the network parameters of the first diffusion model and the second diffusion model are fused based on the channel dimension.
[0098] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0099] Figure 81 is a block diagram of a device 200 for image editing according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0100] Reference Figure 8 , the device 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .
[0101] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.
[0102] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0103] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.
[0104] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0105] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), and when the device 200 is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 204 or sent via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.
[0106] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0107] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor assembly 214 can also detect the position change of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200 and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.
[0108] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0109] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0110] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the instructions can be executed by the processor 220 of the device 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0111] It is to be understood that in the present disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings.
[0112] It is further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or degree of importance. In fact, the expressions "first", "second", etc. can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0113] It can be further understood that, unless otherwise specified, “connection” includes a direct connection without other components between the two, and also includes an indirect connection with other components between the two.
[0114] It is further understood that, although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring the execution of all the operations shown to obtain the desired results. In certain environments, multitasking and parallel processing may be advantageous.
[0115] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present solution, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.
[0116] It should be understood that the present disclosure is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.
Claims
1. An image editing method, It is characterized in that include: Acquire a first image and a second image; Based on a bidirectional diffusion network model, replacing a first region image in the first image with a second region image in the second image to generate a target image; Among them, the bidirectional diffusion network model includes a first diffusion model and a second diffusion model, the first diffusion model is used to perform feature modeling encoding and decoding modeling on the first area image, the second diffusion model is used to fuse the second area image into the first area image, and the first encoding part of the second diffusion model and the first encoding part of the first diffusion model share the same network parameters, the second diffusion model and the first diffusion model perform network parameter fusion in the second part, the second part is the part of the first diffusion model except the first part and the third part, and the third part is the decoding part of the first diffusion model that is symmetrical to the first part.
2. The method according to claim 1, It is characterized in that The step of replacing the first region image in the first image with the second region image in the second image based on the bidirectional diffusion network model to generate a target image includes: Determine a second region image selected by a user on the second image, and upsample the second region image into a third image having the same size as the first image; performing masking processing on a first region image corresponding to a portion of the second region image in the first image, and adding noise to the masked image to generate a fourth image; The third image and the fourth image are input into the bidirectional diffusion network model, and a target image is generated based on the bidirectional diffusion network model. The target image is an image obtained by fusing the third image into the masked region image in the fourth image.
3. The method according to claim 2, It is characterized in that The generating of a target image based on the bidirectional diffusion network model comprises: Inputting the fourth image into a first diffusion model of the bidirectional diffusion network model, and inputting the third image into a second diffusion model of the bidirectional diffusion network model; encoding the fourth image based on the first encoding part of the first diffusion model to obtain a first encoded image, and encoding the third image based on the second diffusion model sharing the network parameters of the first encoding part to obtain a second encoded image; fusing the second part of the first diffusion model and the second part of the second diffusion model, performing fusion processing on the first coded image and the second coded image to obtain fusion features, wherein the second part of the first diffusion model and the second part of the second diffusion model have the same network structure; Based on the third part of the first diffusion model, the fusion feature is decoded to obtain a target image.
4. The method according to any one of claims 1 to 3, It is characterized in that The bidirectional diffusion network model is trained in the following way: Initializing a first diffusion model and a second diffusion model; During the i-th training, all parameters of the first diffusion model are updated through the back propagation process, and the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model are updated with the same learning rate, and remain unchanged, where i is a positive integer; During the i+1th training, the parameters of the first coding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first coding part are updated through the back propagation process, and the second part of the second diffusion model and the second part of the first diffusion model are updated by fusing the parameters of the i-th training; The i-th training and the i+1-th training process are performed alternately until a bidirectional diffusion network model satisfying the constraint conditions is obtained through training.
5. The method according to claim 3, It is characterized in that The third part based on the first diffusion model decodes the fusion feature to obtain a target image, including: decoding the fused features based on the third part of the first diffusion model to obtain a sixth image; The sixth image is guided based on the first image to obtain a target image that meets a guidance condition.
6. The method according to claim 1, It is characterized in that The second diffusion model and the first diffusion model are subjected to network parameter fusion based on a cross attention mechanism in the second part; The cross-attention mechanism is based on the channel dimension and integrates the network parameters of the first diffusion model and the second diffusion model.
7. An image editing device, It is characterized in that include: An acquisition unit, configured to acquire a first image and a second image; a processing unit, configured to replace a first region image in the first image with a second region image in the second image based on a bidirectional diffusion network model, so as to generate a target image; Among them, the bidirectional diffusion network model includes a first diffusion model and a second diffusion model, the first diffusion model is used to perform feature modeling encoding and decoding modeling on the first area image, the second diffusion model is used to fuse the second area image into the first area image, and the first encoding part of the second diffusion model and the first encoding part of the first diffusion model share the same network parameters, the second diffusion model and the first diffusion model perform network parameter fusion in the second part, the second part is the part of the first diffusion model except the first part and the third part, and the third part is the decoding part of the first diffusion model that is symmetrical to the first part.
8. The device according to claim 7, It is characterized in that The processing unit replaces the first region image in the first image with the second region image in the second image based on the bidirectional diffusion network model to generate a target image: Determine a second region image selected by a user on the second image, and upsample the second region image into a third image having the same size as the first image; performing masking processing on a first region image corresponding to a portion of the second region image in the first image, and adding noise to the masked image to generate a fourth image; The third image and the fourth image are input into the bidirectional diffusion network model, and a target image is generated based on the bidirectional diffusion network model. The target image is an image obtained by fusing the third image into the masked region image in the fourth image.
9. The device according to claim 8, It is characterized in that The processing unit generates a target image based on the bidirectional diffusion network model in the following manner: Inputting the fourth image into a first diffusion model of the bidirectional diffusion network model, and inputting the third image into a second diffusion model of the bidirectional diffusion network model; encoding the fourth image based on the first encoding part of the first diffusion model to obtain a first encoded image, and encoding the third image based on the second diffusion model sharing the network parameters of the first encoding part to obtain a second encoded image; fusing the second part of the first diffusion model and the second part of the second diffusion model, performing fusion processing on the first coded image and the second coded image to obtain fusion features, wherein the second part of the first diffusion model and the second part of the second diffusion model have the same network structure; Based on the third part of the first diffusion model, the fusion feature is decoded to obtain a target image.
10. The device according to any one of claims 7 to 9, It is characterized in that The bidirectional diffusion network model is trained in the following way: Initializing a first diffusion model and a second diffusion model; During the i-th training, all parameters of the first diffusion model are updated through the back propagation process, and the parameters of the second part of the second diffusion model corresponding to the fusion of the first diffusion model are updated with the same learning rate, and remain unchanged, where i is a positive integer; During the i+1th training, the parameters of the first coding part of the second diffusion model that shares parameters with the first diffusion model and the third part corresponding to the first coding part are updated through the back propagation process, and the second part of the second diffusion model and the second part of the first diffusion model are updated by fusing the parameters of the i-th training; The i-th training and the i+1-th training process are performed alternately until a bidirectional diffusion network model satisfying the constraint conditions is obtained through training.
11. The device according to claim 9, It is characterized in that The processing unit decodes the fusion feature based on the third part of the first diffusion model to obtain a target image in the following manner: decoding the fused features based on the third part of the first diffusion model to obtain a sixth image; The sixth image is guided based on the first image to obtain a target image that meets a guidance condition.
12. The device according to claim 7, It is characterized in that The processing unit fuses the network parameters of the second diffusion model and the first diffusion model in the second part based on the cross attention mechanism in the following manner; Wherein, the cross-attention mechanism is based on the channel dimension and integrates the network parameters of the first diffusion model and the second diffusion model.
13. An image editing device, It is characterized in that include: processor: a memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the image editing method according to any one of claims 1 to 6.
14. A storage medium, It is characterized in that The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor, the processor is enabled to execute the image editing method according to any one of claims 1 to 6.