Image processing method and device, storage medium and chip

Through the phased filling method, the problem of deleting objects and then filling them into the image during the image filling process is solved, achieving better filling effect and user expectations.

CN120147472APending Publication Date: 2025-06-13BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311696926.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is prone to deleting objects and being filled into the image during the image filling process, resulting in poor filling effect and unable to meet user expectations.

Method used

Using the method of staged filling, the image is initially filled in the first stage to obtain a second image, and then the second image is input into the second diffusion model for fine filling in the second stage to ensure the good integration of the fill content and the background.

Benefits of technology

By filling in stages, the situation where the deleted object is then filled into the image is avoided, and the filling effect is improved, making the obtained image more likely to meet the user's expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147472A_ABST
    Figure CN120147472A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, a storage medium and a chip, in the process of filling a first image, the first image can be divided into two stages for filling, in the first stage, only the first image is subjected to preliminary filling to obtain a second image, and then the second image is subjected to second-stage filling, so that the image processing efficiency is improved. According to the method, the filling content and the background can be better fused, and the situation that the object is easily deleted and filled into the image due to one-time filling can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, storage medium, and chip. Background Art

[0002] In related technologies, during the process of a user taking an image, there often appears an object in the captured image that the user does not want. For example, when a user takes a check-in photo of a certain landscape, at this time, the captured image may contain some tourists, and obviously these tourists are not what the user hopes to appear in the photo. Therefore, it is necessary to delete these tourists in the captured photo. Summary of the Invention

[0003] To overcome the problems existing in related technologies, the present disclosure provides an image processing method, apparatus, storage medium, and chip. During the process of filling a first image, the first image can be filled in two stages. In the first stage, only a preliminary filling of the first image is performed to obtain a second image, and then a second-stage filling is performed on the second image. In this way, not only can the filled content be better integrated with the background, but also the situation where the deleted object is easily filled back into the image due to one-time filling can be avoided.

[0004] According to a first aspect of an embodiment of the present disclosure, an image processing method is provided, including: inputting a first image into a first diffusion model to obtain a second image output by the first diffusion model, where the first image includes a region to be filled; inputting the second image into a second diffusion model to obtain a third image output by the second diffusion model, where the third image is an image obtained by filling the region to be filled of the first image.

[0005] Optionally, the filling accuracy of the first diffusion model is less than the filling accuracy of the second diffusion model.

[0006] Optionally, the first diffusion model includes a first encoding module, a first denoising module, and a first decoding module; and

[0007] The step of inputting the first image into the first diffusion model to obtain a second image output by the first diffusion model includes:

[0008] Inputting the first image into the first encoding module to obtain a first hidden vector;

[0009] Fusing the first hidden vector and first noise data to obtain first fusion data;

[0010] Inputting the first fusion data into the first denoising module to obtain first denoised data;

[0011] Input the above first denoised data into the above first decoding module to obtain the above second image.

[0012] Optionally, the above second diffusion model includes a second encoding module, a second denoising module, and a second decoding module; and

[0013] The above step of inputting the above second image into the second diffusion model to obtain the third image output by the second diffusion model includes:

[0014] Input the above second image into the second encoding module to obtain a second latent vector;

[0015] Fuse the above second latent vector and second noise data to obtain second fused data;

[0016] Input the above second fused data into the above second denoising module to obtain second denoised data;

[0017] Input the above second denoised data into the above second decoding module to obtain the above third image.

[0018] Optionally, the above step of inputting the first fused data into the above first denoising module to obtain first denoised data includes:

[0019] Using a first preset number of times as the number of loop processing times, and using a first denoising network to perform loop processing on the above first fused data;

[0020] And, the above step of inputting the above second fused data into the above second denoising module to obtain second denoised data includes:

[0021] Using a second preset number of times as the number of loop processing times, and using a second denoising network to perform loop processing on the above second fused data;

[0022] Among them, the above first number is less than the above second number.

[0023] Optionally, the above step of fusing the above first latent vector and first noise data to obtain first fused data includes:

[0024] Weight the above first image using a first image weight, and weight the above first noise data using a first noise weight, and use the first weighted sum of the above first image and the above first noise data as the above first fused data;

[0025] And, the above step of fusing the above second latent vector and second noise data to obtain second fused data includes:

[0026] Weight the second image using the second image weight, weight the second noise data using the second noise weight, and use the second weighted sum of the second image and the second noise data as the second fusion data;

[0027] Wherein, the first image weight is less than the second image weight.

[0028] Optionally, the step of obtaining the first image includes:

[0029] Display the target image;

[0030] Generate the first image based on the deletion operation of a partial area of the target image.

[0031] Optionally, the generating the first image based on the deletion operation of a partial area of the target image includes:

[0032] Detect the deletion operation of a partial area of the target image, and generate an image mask according to the partial area targeted by the deletion operation;

[0033] Generate the first image according to the target image and the image mask.

[0034] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including: a first filling module configured to input a first image into a first diffusion model to obtain a second image output by the first diffusion model, wherein the first image includes an area to be filled; a second filling module configured to input the second image into a second diffusion model to obtain a third image output by the second diffusion model, wherein the third image is an image obtained by filling the area to be filled of the first image.

[0035] According to a third aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including: a processor;

[0036] A memory for storing instructions executable by the processor;

[0037] Wherein, the processor is configured to execute the steps of implementing the image processing method provided in any optional manner of the first aspect of the present disclosure.

[0038] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of implementing the image processing method provided in any optional manner of the first aspect of the present disclosure are realized.

[0039] According to a fifth aspect of the embodiments of the present disclosure, a chip is provided, including a processor and an interface; the processor is configured to read instructions to execute the steps of the image processing method provided by any one of the optional manners in the first aspect of the present disclosure.

[0040] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Since the image is filled in stages, the first stage does not directly generate a certain filling object, but only needs to fill the background information. The second stage is carried out on the basis of the first stage, that is, on the basis of the background information, the background information filled in the second image can be made clearer. In this way, the filling of the area to be filled in the first image is completed. It can be seen that in this way, the phenomenon of filling in the object deleted by the user when filling the area to be filled in the first image can be avoided. At the same time, since the background information is extracted twice in the filling process, when filling the area to be filled, the background information can be more fully considered, so that the obtained third image is more likely to meet the user's expectations.

[0041] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0043] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment;

[0044] Figure 2 is a schematic diagram of the process of image filling in a related art shown according to an exemplary embodiment;

[0045] Figure 3 is a schematic diagram of the process of image filling shown according to an exemplary embodiment;

[0046] Figure 4 is a schematic diagram of the process of image filling shown according to an exemplary embodiment;

[0047] Figure 5 is a schematic diagram of comparison of image filling effects shown according to an exemplary embodiment;

[0048] Figure 6 is a block diagram of an image filling device shown according to an exemplary embodiment;

[0049] Figure 7 is a block diagram of a device shown according to an exemplary embodiment;

[0050] Figure 8 It is a block diagram of a device shown according to an exemplary embodiment. Detailed implementation manners

[0051] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0052] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.

[0053] As can be seen from the description of the background art, some objects in the image may not be what the user needs; that is, the user usually has the need to delete some objects in the image. However, after the objects are deleted, the image will seem incomplete. Therefore, after the user deletes some content in the image, the deleted area can also be filled to make the image complete. In this way, the effect of erasing some objects in the image can be achieved.

[0054] In the related art, in the process of filling the deleted area, a deep learning method based on a convolutional neural network can be used to fill the image hole area. However, this method often can only fill small areas due to the lack of associative ability. Then, a diffusion model emerged, and the image filling algorithm combined with the diffusion model can fill larger hole areas. However, during the filling process of the diffusion model, the stability is not high enough, and phenomena such as generating another person after eliminating a person and generating another wire after eliminating a wire often occur. This makes the filled image not meet the user's expectations.

[0055] It can be seen that the filling methods in the related art are likely to cause the same type of objects as the deleted objects to be filled as filling objects, resulting in a poor filling effect and making the finally obtained image not meet the user's expectations.

[0056] In the image processing method provided by the present disclosure, during the process of filling the first image, the first image can be filled in two stages. In the first stage, only a preliminary filling is performed on the first image to obtain a second image, and then a second-stage filling is performed on the second image. In this way, not only can the filled content be better integrated with the background, but also, by filling in stages, only a preliminary filling is performed in the first stage, and a fine filling is performed in the second stage, which can also avoid the situation where the deleted object is easily filled into the image again due to a single filling.

[0057] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an image processing method provided by the present disclosure. This image processing method can be applied to a terminal (such as a mobile phone, a tablet, etc.) or a server. As Figure 1 shown, the image processing method can include the following steps.

[0058] In S12, the first image is input into the first diffusion model to obtain a second image output by the first diffusion model.

[0059] Here, the first image includes the area to be filled.

[0060] As an example, the diffusion model can be understood as a model for filling the area to be filled.

[0061] As an example, the second image can be understood as the image after filling the area to be filled in the first image. For example, the image obtained after a rough filling of the area to be filled in the first image (for example, only a partial filling of the background of the area to be filled).

[0062] In S14, the second image is input into the second diffusion model to obtain a third image output by the second diffusion model.

[0063] Here, the third image is the image obtained by filling the area to be filled in the first image.

[0064] As an example, the second image is obtained after a preliminary filling of the area to be filled in the first image. That is, at this time, the filling of the area to be filled in the first image is not completed. Therefore, based on the second image, the filling of this area to be filled can be continued until the filling of the area to be filled is completed.

[0065] It should be noted that in the process of filling using a diffusion model, the diffusion model can combine the entire background feature information of the image to achieve filling for the area to be filled. Therefore, in the related art, when a user deletes an object A, during the filling process, the diffusion model may, by analyzing the background feature information, consider that there may be an object B here (object A and object B can be of the same type or the same object). In this way, it is very likely that the same type of object as the deleted object will be added again during the filling process for the area to be filled.

[0066] The present disclosure divides the filling process of the area to be filled into two stages. The process of inputting the first image into the first diffusion model for filling to obtain the second image can be understood as the first filling stage. Since the second image still needs to be filled using the second diffusion model, that is, the first diffusion model does not perform a complete and detailed filling of the first image, but only performs partial filling (which can also be understood as rough filling) on the first image. In this way, the second image will not be filled with too much content. At the same time, compared with the first image, the second image has filled some content. For example, compared with the first image, the second image fills some background in the area to be filled, and this part of the background is filled according to the non-filled area in the first image.

[0067] The process of inputting the second image into the second diffusion model for filling to obtain the third image can be understood as the second filling stage. Since the second-stage filling is carried out on the basis of the first-stage filling, and since the first-stage filling has completed partial filling (rough filling), the second stage can continue to fill on the basis of the first stage. In this way, when performing fine filling in the second stage, more information in the second image can be considered.

[0068] That is, the first stage performs fuzzy filling on the area to be filled, while the second stage performs fine filling on the area to be filled. In this way, during the filling process of the second stage, the image after the first-stage filling is fully considered, and more background information can be considered, making the filled image better adapt to the background. At the same time, due to the preliminary filling in the first stage (this filling process mainly fills the background information and usually does not add new filling objects), it is not easy to have the situation where the filling object deleted by the user is filled in again.

[0069] In the present disclosure, after initially filling the first image, a second image can be obtained, and then by continuing to fill the second image, a third image can be obtained. That is, in the process of filling the area to be filled of the first image, it is divided into two stages. The first stage can be understood as the process of obtaining the second image by outputting the first image to the first diffusion model. In this process, the area to be filled of the first image can be initially filled. Then, the second image is used for the second-stage fine filling to obtain the third image.

[0070] In this way, since the filling is carried out in stages, the first stage does not directly generate a certain filling object, but only needs to fill the background information. The second stage is carried out on the basis of the first stage, that is, continue to fill based on the background information, which can make the background information filled in the second image clearer. In this way, the filling of the area to be filled of the first image is completed. It can be seen that in this way, the phenomenon of filling in the object deleted by the user when filling the area to be filled of the first image can be avoided. At the same time, since the background information is extracted twice in this filling process, it can be considered more fully when filling the area to be filled, making the obtained third image more likely to meet the user's expectations.

[0071] To facilitate understanding the idea of the present disclosure, it can be combined with Figure 2-3 for illustration. Figure 2 It can be understood as a schematic diagram of the process of using a diffusion model for image filling in the related art. Figure 3 It can be understood as a schematic diagram of the process of filling an image according to the present disclosure.

[0072] In Figure 2 , the image 201 can be understood as the original schematic diagram corresponding to the first image, and the image 202 can be understood as the schematic diagram of the area to be filled. Among them, the white area in the image 202 can be understood as the area to be filled. That is, the combination of the image 201 and the image 202 can be understood as the first image. By inputting the image 201 and the image 202 into the filling model for filling, the image shown in the image 203 (that is, the filled image) can be obtained. It can be seen that, however, in the filled image, other user images are also generated.

[0073] Continuing to combine Figure 3 It can be seen that Figure 3The image 301 therein can be understood as the original schematic diagram corresponding to the first image and can be understood as the schematic diagram of the area to be filled. Among them, the white area in the image 302 can be understood as the area to be filled. That is to say, the combination of the image 301 and the image 302 can be understood as the first image. Inputting the image 301 and the image 302 into the first diffusion model, the second image, that is, the image 303, can be obtained. It can be seen that the first diffusion model only partially fills the area to be filled (background filling) of the first image. At this time, the image 303 and the image 301 (the image used to indicate the filling area) can be continuously input into the second diffusion model to continue filling the area to be filled, so that the image 304, that is, the third image, can be obtained, and there are no other user images in the third image.

[0074] In contrast, it can be seen that through the method of the present disclosure, the situation where the image deleted by the user appears in the filled image again can be better avoided, so that the erasing effect on the object is better.

[0075] In some embodiments, the filling accuracy of the first diffusion model is less than the filling accuracy of the second diffusion model.

[0076] As an example, the filling accuracy of the first diffusion model is less than that of the second diffusion model. In this way, the second image can also be a rough filling of only the area to be filled in the first image compared with the first image. For example, background filling is performed. The specific color, content, etc. filling is performed by the second diffusion model. In this way, it can better avoid the finally obtained third image from including the object deleted by the user or including the same type of object deleted by the user.

[0077] In some embodiments, the first diffusion model may include a first encoding module, a first denoising module, and a first decoding module; and, in S12, "inputting the first image into the first diffusion model to obtain the second image output by the first diffusion model" may specifically include:

[0078] Inputting the first image into the first encoding module to obtain a first latent vector;

[0079] Fusing the first latent vector and the first noise data to obtain first fusion data;

[0080] Inputting the first fusion data into the first denoising module to obtain first denoised data;

[0081] Inputting the first denoised data into the first decoding module to obtain the second image.

[0082] As an example, the first latent vector can indicate the important features of the first image and the area to be filled. Fusing the first latent vector and the first noise data can be understood as adding noise to the area to be filled, and then, through denoising and decoding, filling of the area to be filled can be achieved.

[0083] In some embodiments, the second diffusion model may include a second encoding module, a second denoising module, and a second decoding module; and, in S14, "inputting the second image into the second diffusion model to obtain a third image output by the second diffusion model" may specifically include:

[0084] Inputting the second image into the second encoding module to obtain a second latent vector;

[0085] Fusing the second latent vector and the second noise data to obtain second fusion data;

[0086] Inputting the second fusion data into the above-mentioned second denoising module to obtain second denoised data;

[0087] Inputting the second denoised data into the second decoding module to obtain a third image.

[0088] As an example, the second latent vector can indicate the important features of the second image and the area to be filled. Since the second image is obtained after the first image is preliminarily filled, the second latent vector can contain more features of the original image, so that filling can be better achieved and the appearance of deleted objects in the filled image can be better avoided.

[0089] In some embodiments, inputting the first fusion data into the first denoising module to obtain first denoised data may specifically include: using a first denoising network to perform cyclic processing on the above-mentioned first fusion data with a preset first number of times as the cyclic processing times;

[0090] And, inputting the second fusion data into the above-mentioned second denoising module to obtain second denoised data may specifically include: using a second denoising network to perform cyclic processing on the above-mentioned second fusion data with a preset second number of times as the cyclic processing times.

[0091] Here, the first number of times is less than the second number of times.

[0092] As an example, the first number of times is less than the second number of times, so that when performing the first-stage filling for the area to be filled, too much content is not filled in the area to be filled, thus affecting the second-stage filling.

[0093] It should be noted that the process of filling the area to be filled can be understood as a process of adding noise to the area to be filled and then removing the noise. In this process, by understanding the feature information of the image, the restoration of the area to be filled can be achieved. That is to say, the number of loop processes can characterize the filling degree of the area to be filled to a certain extent. And since the first number is less than the second number, it can be ensured that the filling process in the first stage only achieves a rough filling of the area to be filled, and the specific filling details are carried out in the second stage.

[0094] In some implementation manners, the first number can be 8, and the second number can be 42. That is to say, after 8 loop processes in the first stage, the background of the area to be filled can be filled, and then after 42 loop processes in the second stage, the complete filling of the filled area can be completed. In this way, although the filling is carried out in stages, the total number of loops is only 50 at this time. In this way, not only the filling accuracy of the area to be filled is ensured, but also the filling efficiency is not affected too much.

[0095] In some embodiments, the above-mentioned fusion of the first hidden vector and the first noise data to obtain the first fusion data may specifically include: weighting the first image with the first image weight, weighting the above-mentioned first noise data with the first noise weight, and taking the first weighted sum of the first image and the first noise data as the above-mentioned first fusion data; and,

[0096] The above-mentioned fusion of the second hidden vector and the second noise data to obtain the second fusion data may specifically include:

[0097] weighting the second image with the second image weight, weighting the second noise data with the second noise weight, and taking the second weighted sum of the second image and the second noise data as the second fusion data.

[0098] Here, the first image weight is less than the second image weight.

[0099] As an example, the first image weight is less than the second image weight. In this way, in the second stage, more consideration can be given to the original features of the image compared with the first stage, which helps to restore the details of the area to be filled. And since only a small amount of the original image features are considered in the first stage, this can be more conducive to filling the background of the area to be filled.

[0100] It should be understood that the more original features of the image are considered, the easier it is for the diffusion model to generate associations, making the filled image more coordinated with the background. In related technologies, weights are usually not set for images, resulting in poor coordination of the obtained image after filling the area to be filled. At the same time, since the related technologies use a one-time filling method, if weights are set for images, it is easy to misgenerate objects that the user has deleted. In the present disclosure, during the first-stage filling process, only the background needs to be blurred and filled. Therefore, the weight of the first image can be set relatively low. During the second-stage filling process, since more refined filling is required, the weight of the second image can be set relatively high.

[0101] For ease of understanding, it can be described in combination with Figure 4 as follows. Figure 4 It can be understood as a schematic diagram of two stages of filling the image. As can be seen from Figure 4 the overall process is divided into two stages (that is, Figure 4 One Stage and Two Stage in

[0102] The first stage:

[0103] The original image is multiplied by the filling area mask, and the area to be filled is masked to obtain the first image. The first image is input into the VAE Encoder (Variational Auto Encoder), and a latent space vector containing background information is obtained. After passing through the Weight-Fusion module (used to fuse the noise data with the latent space vector), the vector is fused with the Gaussian noise Zn, and the fusion method can be simply multiplying the weights and adding them.

[0104] In the present disclosure, the weight is a hyperparameter set in advance. The larger the weight, the higher the proportion of background information. In this solution, weight1 is set to 0.05 in the first stage and weight2 is set to 0.1 in the second stage. Compared with directly using the Gaussian noise Zn to restore the image in related technologies, this solution fuses Zn with the background information, and under the guidance of the background information, the image in the filled area is more coordinated with the background. At the same time, the stability of the diffusion model is also improved.

[0105] The fused noise enters the U-Net network (for cyclic denoising) and undergoes 8 cycles of denoising. 8 is a hyperparameter preset in this disclosure. This parameter setting takes two aspects into account: First, to ensure the filling efficiency, the image filling technology based on the diffusion model generally uses the Unet network to cycle 50 times for denoising; in this solution, in order not to increase too much computational load, it cycles 8 times and 42 times respectively in the first stage and the second stage for denoising. Second, to ensure the filling effect, since the output of the first stage is only used as the initial input of the second stage and does not require fine filling; at the same time, using fewer cycle times can prevent the model from restoring the complete target. Based on this, the second stage performs filling, making it easier to restore the background rather than generating new objects (for example, objects that some users want to delete are accidentally restored).

[0106] Further, the latent space vector obtained after denoising is transformed into the image space through the VAE Decoder and input into the second stage.

[0107] The second stage:

[0108] The process of the second stage is basically the same as that of the first stage. First, multiply the image output by the first stage and the filling area mask, input it into the VAE Encoder to obtain the latent space vector containing background information, and then fuse this vector with the Gaussian noise Zn. The fusion method can be simply multiplying weights and adding them. Only in the second stage, weight2 can be set to 0.1. The fused noise enters the U-Net network for 42 cycles of denoising, and then the latent space vector obtained after denoising is transformed into the image space through the VAE Decoder.

[0109] That is to say, through such a method, this disclosure can not only achieve a good filling effect but also avoid the phenomenon of filling the objects deleted by the user during the process of filling the to-be-filled area in the first image.

[0110] To better understand the idea of this disclosure, it can be combined with Figure 5 a comparison to reflect the effect of the image processing method of this disclosure, Figure 5 which can be understood as a schematic diagram of the process of filling and eliminating the filling area of the image. Image 501 can represent the original image, the white area in Image 502 can be understood as the filling area, Image 503 can be understood as the image obtained after filling by the related technology, and Image 504 can be understood as the image obtained after filling by the solution of this disclosure. It can be seen that in Image 503, the image that was originally to be deleted is filled again, while the solution of this disclosure will not fill in the objects that should have been deleted.

[0111] In some embodiments, the step of obtaining the first image may include:

[0112] Display the target image;

[0113] Generate a first image based on a deletion operation on a partial region of the target image.

[0114] In some embodiments, the above-mentioned generating a first image based on a deletion operation on a partial region of the target image may specifically include:

[0115] Detect a deletion operation on a partial region of the target image, and generate an image mask according to the partial region targeted by the deletion operation; generate the above-mentioned first image according to the target image and the image mask.

[0116] That is, in the present disclosure, the user can perform a deletion operation to determine the region to be filled.

[0117] A possible application scenario of the present disclosure may be: when the user is browsing the collected image, the user can perform an erasing operation on a partial region of the collected image; this indicates that this part of the content needs to be erased and the corresponding background is filled. Or, directly click the preset function trigger key to trigger the magic erasing function (erase the redundant lines in the image and fill the corresponding background), or the intelligent erasing function (automatically determine the object that the user wants to keep and erase and fill other objects).

[0118] That is, the user can determine the region to be filled in the image according to actual needs. After that, a first image can be generated, and then the main body can execute the image processing method of the present disclosure to fill the region to be filled in the first image.

[0119] Figure 6 It is a block diagram of an image processing device shown according to an exemplary embodiment. Refer to Figure 6 , the image processing device 600 includes a first filling module 601 and a second filling module 602.

[0120] The first filling module 601 is configured to input the first image into a first diffusion model to obtain a second image output by the first diffusion model, where the first image includes a region to be filled;

[0121] The second filling module 602 is configured to input the second image into a second diffusion model to obtain a third image output by the second diffusion model, where the third image is an image obtained by filling the region to be filled in the first image.

[0122] In some embodiments, the filling accuracy of the first diffusion model is less than the filling accuracy of the second diffusion model.

[0123] In some embodiments, the above-mentioned first diffusion model includes a first encoding module, a first denoising module, and a first decoding module; and, the first filling module 601 is configured to:

[0124] Input the above-mentioned first image into the above-mentioned first encoding module to obtain a first latent vector;

[0125] Fuse the above-mentioned first latent vector and first noise data to obtain first fusion data;

[0126] Input the above-mentioned first fusion data into the above-mentioned first denoising module to obtain first denoised data;

[0127] Input the above-mentioned first denoised data into the above-mentioned first decoding module to obtain the above-mentioned second image.

[0128] In some embodiments, the above-mentioned second diffusion model includes a second encoding module, a second denoising module, and a second decoding module; and, the second filling module 602 is configured to:

[0129] Input the above-mentioned second image into the second encoding module to obtain a second latent vector;

[0130] Fuse the above-mentioned second latent vector and second noise data to obtain second fusion data;

[0131] Input the above-mentioned second fusion data into the above-mentioned second denoising module to obtain second denoised data;

[0132] Input the above-mentioned second denoised data into the above-mentioned second decoding module to obtain the above-mentioned third image.

[0133] In some embodiments, the above-mentioned inputting the first fusion data into the above-mentioned first denoising module to obtain first denoised data includes:

[0134] Using a first denoising network to perform cyclic processing on the above-mentioned first fusion data with a preset first number as the cyclic processing times;

[0135] And, the above-mentioned inputting the second fusion data into the above-mentioned second denoising module to obtain second denoised data includes:

[0136] Using a second denoising network to perform cyclic processing on the above-mentioned second fusion data with a preset second number as the cyclic processing times;

[0137] Wherein, the above-mentioned first number is less than the above-mentioned second number.

[0138] In some embodiments, the above-mentioned fusing the above-mentioned first latent vector and first noise data to obtain first fusion data includes:

[0139] Weight the first image using a first image weight, weight the first noise data using a first noise weight, and use the first weighted sum of the first image and the first noise data as the first fused data;

[0140] Further, fusing the second latent vector and the second noise data to obtain second fused data includes:

[0141] Weight the second image using a second image weight, weight the second noise data using a second noise weight, and use the second weighted sum of the second image and the second noise data as the second fused data;

[0142] Wherein, the first image weight is less than the second image weight.

[0143] In some embodiments, the step of obtaining the first image includes:

[0144] Display a target image;

[0145] Generate the first image based on a deletion operation on a partial region of the target image.

[0146] In some embodiments, generating the first image based on a deletion operation on a partial region of the target image includes:

[0147] Detect a deletion operation on a partial region of the target image, and generate an image mask according to the partial region targeted by the deletion operation;

[0148] Generate the first image according to the target image and the image mask.

[0149] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0150] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the image processing method provided by the present disclosure are implemented.

[0151] Figure 7 It is a block diagram of an apparatus 700 for image processing shown according to an exemplary embodiment. For example, the apparatus 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0152] Refer to Figure 7, Device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output interface 712, a sensor component 714, and a communication component 716.

[0153] The processing component 702 generally controls the overall operation of the device 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-described image processing method. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.

[0154] The memory 704 is configured to store various types of data to support the operation of the device 700. Examples of these data include instructions for any application or method operating on the device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0155] The power component 706 provides power to the various components of the device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 700.

[0156] The multimedia component 708 includes a screen that provides an output interface between the above-described device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The above touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the above touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0157] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC), which is configured to receive external audio signals when the device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0158] The input / output interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0159] The sensor component 714 includes one or more sensors for providing an assessment of various aspects of the state of the device 700. For example, the sensor component 714 can detect the on / off state of the device 700, the relative positioning of components, such as the display and keypad of the device 700, the sensor component 714 can also detect a change in the position of the device 700 or a component of the device 700, the presence or absence of user contact with the device 700, the orientation or acceleration / deceleration of the device 700, and the temperature change of the device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0160] The communication component 716 is configured to facilitate communication between the device 700 and other devices in a wired or wireless manner. The device 700 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0161] In an exemplary embodiment, the apparatus 700 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described image processing method.

[0162] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the apparatus 700 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0163] In addition to being an independent electronic device, the above apparatus may also be a part of an independent electronic device. For example, in one embodiment, the apparatus may be an integrated circuit (IC) or a chip. The integrated circuit may be a single IC or a collection of multiple ICs; the chip may include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip), etc. The above integrated circuit or chip may be used to execute executable instructions (or code) to implement the above-described image processing method. The executable instructions may be stored in the integrated circuit or chip, or may be obtained from other devices or apparatuses. For example, the integrated circuit or chip includes a processor, a memory, and an interface for communicating with other devices. The executable instructions may be stored in the memory, and when the executable instructions are executed by the processor, the above-described image processing method is implemented; or, the integrated circuit or chip may receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-described image processing method.

[0164] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-described image processing method when executed by the programmable device.

[0165] Figure 8 FIG. 4 is a block diagram of an apparatus 800 for image processing shown according to an exemplary embodiment. For example, the apparatus 800 may be provided as a server. Referring to Figure 8 FIG. 4, the apparatus 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by a memory 832 for storing instructions executable by the processing component 822, such as application programs. The application programs stored in the memory 832 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 822 is configured to execute instructions to perform the above-described image processing method.

[0166] The apparatus 800 may also include a power component 826 configured to perform power management of the apparatus 800, a wired or wireless network interface 850 configured to connect the apparatus 800 to a network, and an input / output interface 858. The apparatus 800 may operate based on an operating system stored in the memory 832, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0167] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0168] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image processing method, characterized in that, it includes: Input the first image into the first diffusion model to obtain the second image output by the first diffusion model, where the first image includes an area to be filled; Input the second image into the second diffusion model to obtain the third image output by the second diffusion model, where the third image is an image obtained by filling the area to be filled of the first image.

2. The method according to claim 1, characterized in that, the filling accuracy of the first diffusion model is less than the filling accuracy of the second diffusion model.

3. The method according to claim 1, characterized in that, the first diffusion model includes a first encoding module, a first denoising module and a first decoding module; and the step of inputting the first image into the first diffusion model to obtain the second image output by the first diffusion model includes: Input the first image into the first encoding module to obtain a first latent vector; Fuse the first latent vector and first noise data to obtain first fusion data; Input the first fusion data into the first denoising module to obtain first denoised data; Input the first denoised data into the first decoding module to obtain the second image.

4. The method according to claim 3, characterized in that, the second diffusion model includes a second encoding module, a second denoising module and a second decoding module; and the step of inputting the second image into the second diffusion model to obtain the third image output by the second diffusion model includes: Input the second image into the second encoding module to obtain a second latent vector; Fuse the second latent vector and second noise data to obtain second fusion data; Input the second fusion data into the second denoising module to obtain second denoised data; Input the second denoised data into the second decoding module to obtain the third image.

5. The method according to claim 4, characterized in that, the step of inputting the first fusion data into the first denoising module to obtain first denoised data includes: Taking the preset first number as the number of loop processing times, and using the first denoising network to perform loop processing on the first fusion data; and, the step of inputting the second fusion data into the second denoising module to obtain second denoised data includes: Taking the preset second number as the number of loop processing times, and using the second denoising network to perform loop processing on the second fusion data; wherein, the first number is less than the second number.

6. The method according to claim 4, characterized in that, the step of fusing the first latent vector and first noise data to obtain first fusion data includes: Weighting the first image using the first image weight, weighting the first noise data using the first noise weight, and taking the first weighted sum of the first image and the first noise data as the first fusion data; and, the step of fusing the second latent vector and second noise data to obtain second fusion data includes: Weight the second image using a second image weight, weight the second noise data using a second noise weight, and use the second weighted sum of the second image and the second noise data as second fusion data; wherein, the first image weight is less than the second image weight.

7. The method according to claim 1, wherein, the step of obtaining the first image includes: display a target image; generate the first image based on a deletion operation on a partial area of the target image.

8. The method according to claim 7, wherein, the generating the first image based on a deletion operation on a partial area of the target image includes: detect a deletion operation on a partial area of the target image, and generate an image mask according to the partial area targeted by the deletion operation; generate the first image according to the target image and the image mask.

9. An image processing apparatus, wherein, it includes: a first filling module configured to input a first image into a first diffusion model to obtain a second image output by the first diffusion model, where the first image includes an area to be filled; a second filling module configured to input the second image into a second diffusion model to obtain a third image output by the second diffusion model, where the third image is an image obtained by filling the area to be filled of the first image.

10. An image processing apparatus, wherein, it includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, on which computer program instructions are stored, wherein, when the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A chip, wherein, it includes a processor and an interface; the processor is used to read instructions to execute the method according to any one of claims 1 to 8.