Image processing method, medium, electronic device and program product

By using a neural network model to redraw and super-resolution the foreground object and background during the portrait background replacement process, the problems of lighting disharmony and blurred edges in relighting technology are solved, and the realism and consistency of the image are improved.

CN120765472AActive Publication Date: 2025-10-10HONOR DEVICE CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202410395614.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-10-10
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

When replacing the portrait background with existing relighting technology, the lighting conditions of the foreground object and the background are not harmonious, resulting in the foreground object and background being separated in the composite image, with blurred edges or the original background remaining, which lacks realism.

Method used

By obtaining the foreground and background of different images, the pre-trained neural network model is used to redraw the outline of the foreground object and the surrounding area. Combined with super-resolution processing, the consistency of the foreground object and background lighting is ensured, and the diffusion model is used to redraw features in the latent space to improve the image fusion effect.

Benefits of technology

The lighting consistency between foreground objects and background is improved, the edges are clear, the image realism is enhanced, the foreground object distortion is avoided, and the natural and realistic image fusion is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765472A_ABST
    Figure CN120765472A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an image processing method, a medium, electronic equipment and a program product. According to the method, the edge area where the foreground object is located and other areas in the composite image with the background replaced can be redrawn respectively, the redrawing parameter of the edge area is larger than that of other areas, and it is guaranteed that image fusion is real and natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, medium, electronic device and program product. Background Art

[0002] With the development of image processing technology, users have increasingly demanded more personalized images or videos captured by electronic devices. For example, in some portrait image shooting scenarios, users may need to change the background environment of the portrait, such as changing the background from indoor to outdoor, or from daytime to nighttime.

[0003] However, when changing the background of a portrait, the lighting conditions of the portrait itself and the new background may not match. To improve this problem, relighting (RL) is usually required for the portrait after the background is changed. Relighting refers to re-determining and adding lighting information to the foreground object to make it match the lighting conditions of the new background image, such as the environment image.

[0004] Among them, the traditional relighting technology involves foreground prediction of the original portrait image, and the outline of the foreground object extracted by the foreground prediction algorithm usually has errors, which in turn leads to blurred edges of the foreground object in the composite image, or residual original background, resulting in the foreground object and background image in the composite image being relatively separated, that is, the foreground object and background in the composite image after relighting appear unrealistic. Summary of the Invention

[0005] Embodiments of the present application provide an image processing method, medium, electronic device, and program product.

[0006] In a first aspect, embodiments of the present application provide an image processing method, comprising: acquiring a first image, wherein a foreground object and a background in the first image are from different images; redrawing a first region in the first image according to a first redrawing parameter, and redrawing a second region in the first image according to a second redrawing parameter, to obtain a second image; wherein the first redrawing parameter is different from the second redrawing parameter, and the first region includes the following: an outline of the foreground object and a plurality of pixels adjacent to pixels corresponding to the outline of the foreground object; and the second region is an area of ​​the first image outside the first region. For example, the first redrawing parameter is greater than the second redrawing parameter.

[0007] Thus, in the image processing method, for the composite image after the background is changed, the edge area between the foreground object and the background in the composite image is redrawn, for example, image redrawing methods such as style transfer, image enhancement, color correction, texture creation or shape modification. Among them, the edge area (i.e., the first area) includes: the outline of the foreground object, the content adjacent to the outline inside the foreground object, and the content adjacent to the outline outside the foreground object. At this time, the edge area is the area where the foreground object and the background are fused in the composite image after relighting, so redrawing the edge area is to repair the area where the foreground object and the background are fused. For example, the present application can generate new pixels to fill the gaps in the edge area of ​​the composite image after the background is changed, and set the color of the new pixels to be the same or close to that of the adjacent pixels, thereby eliminating the problem of edge blur and improving the consistency of the lighting conditions between the foreground object and the background in the composite image. At the same time, the method also redraws other areas of the composite image except the edge area, that is, the foreground part within the edge area and the background part outside the edge area, and the redrawing parameters of the edge area are greater than the redrawing parameters of the other areas. It can be understood that the redrawing parameters of the image area refer to the magnitude of the change in the pixel value of the image area, or the degree of style migration of the image, etc. For example, the change in the pixel value of the above-mentioned edge area is greater than the change in the pixel value of the above-mentioned other areas, or the degree of style migration of the above-mentioned edge area is greater than the degree of style migration of the above-mentioned other areas. In this way, while the edge of the outline of the foreground object in the composite image is repaired, it will not cause significant changes to the foreground object itself and the background, such as the replacement of the foreground portrait will not cause significant distortion. Thus, the naturalness and reality of the image fusion can be guaranteed.

[0008] In one possible implementation of the first aspect described above, the plurality of pixels adjacent to a pixel corresponding to the outline of a foreground object include pixels within the foreground object where the shortest distance between pixels corresponding to the outline of the foreground object is less than a first distance, and pixels outside the foreground object where the shortest distance between pixels corresponding to the outline of the foreground object is less than a second distance. In this manner, the first region includes the outline of the foreground object after movement, content adjacent to the outline within the foreground object, and content adjacent to the outline outside the foreground object. For example, the first region in the first image includes the outline of the foreground object, an area where the outline of the foreground object is extended inward by 10 pixels, and an area where the outline is extended outward by 10 pixels.

[0009] In a possible implementation of the first aspect above, obtaining a first image includes: adding a foreground object in a third image whose background is to be replaced to a fourth image providing the background, and relighting the moved foreground object to obtain the first image, wherein the lighting information of the foreground object in the first image is the same as the lighting information of the fourth image.

[0010] In one possible implementation of the first aspect, the method further includes performing super-resolution processing on the second image to obtain a fifth image, wherein the resolution of the fifth image is higher than that of the second image. In this way, the image quality of the fifth image is higher than that of the redrawn second image.

[0011] In a possible implementation of the first aspect described above, a foreground object in a third image is added to a fourth image, and the foreground object is re-lit to obtain a first image, including: determining a first mask image of the foreground object from the third image, the first mask image being used to represent the outline and shape of the foreground object; performing normal prediction on the third image to obtain normal information of the third image, the normal information being used to reflect the surface direction of the foreground object in the third image; performing albedo prediction on the third image to obtain albedo information of the third image, the albedo information being used to reflect the color and material properties of the foreground object in the third image; adding the foreground object in the third image to a fourth image; and re-lighting the foreground object in the fourth image based on the first mask image, the normal information, the albedo information, and the lighting information of the fourth image to obtain the third image. It is understood that algorithms such as foreground prediction, normal prediction, and albedo prediction may have motion errors, resulting in problems such as lighting disharmony between the foreground object and the background in the re-lit first image. At this time, the present application uses a pre-trained model to perform image redrawing and super-resolution processing on the first image after relighting to ensure that the lighting between the foreground object and the background in the final fused image is harmonious, the edges are clear, and the authenticity is high.

[0012] In one possible implementation of the first aspect, adding the foreground object in the third image to the fourth image includes downsampling the third image, and adding the foreground object in the downsampled third image to the fourth image. It will be appreciated that the downsampled third image may be subjected to the above processing to reduce the computational complexity in foreground prediction, normal prediction, albedo prediction, and relighting.

[0013] In a possible implementation of the first aspect, the image redrawing of the first region in the first image according to the first redrawing parameter and the image redrawing of the second region in the first image according to the second redrawing parameter include: inputting the first image into a pre-trained first network model, redrawing the first region in the first image according to the first redrawing parameter and the second region in the first image according to the second redrawing parameter by the first network model, and outputting a second image. It can be understood that the pre-trained first network model has image redrawing capability, specifically, the capability of redrawing the edge region and other regions of the foreground object in the image according to different redrawing parameters.

[0014] In a possible implementation of the first aspect, the first network model includes the following components: a first encoder, a first decoder, and a first diffusion model. The process of the first network model processing the first image into the second image includes: inputting the first image into the first encoder, encoding the first image into a first feature in a latent space; inputting the first feature into the first diffusion model, redrawing the feature corresponding to the first region in the first feature according to the first redrawing parameter based on preset noise, and redrawing the feature corresponding to the first region in the first feature according to the second redrawing parameter based on the preset noise, to obtain a second feature; and inputting the second feature into the first decoder, decoding the second feature into the second image. The first feature and the second feature can be latent features, and the preset noise can be Gaussian noise. In this way, the foreground object and the background in the synthesized image after the background is replaced can be redrawn in the latent space, so as to improve the lighting effect and the authenticity of the fusion of the two.

[0015] In a possible implementation of the first aspect, the first diffusion model processes the first feature by superimposing the preset noise on the first feature according to a first redrawing parameter to obtain a first redrawing feature; corresponding to 1≤j≤k*(ab), during the jth iteration, inputting the jth redrawing feature into the first diffusion model for redrawing, and outputting the j+1th redrawing feature, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is the preset parameter; corresponding to j=k*(ab)+1, superimposing the preset noise on the first feature according to the second redrawing parameter to obtain The intermediate feature is replaced by the feature corresponding to the second region in the j-th redrawn feature, and the feature corresponding to the second region in the intermediate feature is replaced by the feature corresponding to the second region in the intermediate feature, to obtain the updated j-th redrawn feature; corresponding to j=k*(ab)+1, in the j-th iteration process, the updated j-th redrawn feature is input into the first diffusion model for redrawing, and the j+1-th redrawn feature is output; corresponding to j>k*(ab), in the j-th iteration process, the j-th redrawn feature is input into the first diffusion model for redrawing, and the j+1-th redrawn feature is output; corresponding to j=k*a, the j+1-th redrawn feature is used as the second feature. In this way, in the latent space, the redrawing parameters of the features of the first region where the foreground object is located in the first image obtained by replacing the background are larger and the number of redrawings is larger, while the redrawing parameters of the features of the other second regions in the first image are smaller and the number of redrawings is smaller.

[0016] In a possible implementation of the first aspect above, the first network model also includes a text encoder; the training process of the first network model includes: training the first network model according to at least one first training data to obtain a trained first network model; wherein the first training data includes a first training text and a first training image, wherein the first training text is related to the image features of the first training image, and the text encoder is used to obtain the text features of the first training text.

[0017] In a possible implementation of the first aspect above, training the first network model according to at least one first training data includes: training a first diffusion model in the first network model according to the at least one first training data to obtain a trained first diffusion model; and based on the trained first diffusion model, training a first decoder in the first network model according to the at least one first training data to obtain a trained first decoder.

[0018] In a possible implementation of the first aspect above, the processing flow of the first network model on the first training data includes: inputting the first training text in the first training data into a text encoder and outputting text features of the first training text; inputting the first training image in the first training data into the first encoder and encoding the first image into a third feature in the latent space; inputting the third feature into the first diffusion model, redrawing the features corresponding to the third region in the third feature according to a first redrawing parameter based on preset noise, and redrawing the features corresponding to the fourth region in the third feature according to a second redrawing parameter based on preset noise to obtain a fourth feature, where the third region includes the following content: the outline of the foreground object in the first training image, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object in the first training image; the fourth region is the region outside the third region in the first training image; inputting the third feature into the first decoder and decoding the fourth feature into a first training result; adjusting the network parameters of the first diffusion model based on the difference between the text features of the first training text and the third features, and / or adjusting the network parameters of the first decoder based on the difference between the first training image and the first training result.

[0019] In a possible implementation of the first aspect above, the processing flow of the third feature by the first diffusion model includes: superimposing the preset noise and the third feature according to the first redrawing parameter to obtain the first training redrawing feature; corresponding to 1≤j≤k*(ab), in the j-th iteration process, inputting the j-th training redrawing feature into the first diffusion model for redrawing, and outputting the j+1-th training redrawing feature, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is the preset parameter; corresponding to j=k*(ab)+1, superimposing the preset noise and the third feature according to the second redrawing parameter to obtain a training intermediate feature , and replace the feature corresponding to the fourth region in the jth training redraw feature with the feature corresponding to the fourth region in the training intermediate feature, obtaining an updated jth training redraw feature; corresponding to j = k*(ab)+1, during the jth iteration, the updated jth training redraw feature is input into the first diffusion model for redrawing, and the j+1th training redraw feature is output; corresponding to j>k*(ab), during the jth iteration, the jth training redraw feature is input into the first diffusion model for redrawing, and the j+1th training redraw feature is output; corresponding to j = k*a, the j+1th training redraw feature is used as the fourth feature. In this way, the first diffusion model can be enabled to redraw the image according to different redrawing parameters for the area where the foreground object is located and other areas in the image.

[0020] In one possible implementation of the first aspect, performing super-resolution processing on the second image to obtain a fifth image includes inputting the second image into a pre-trained second network model and outputting the fifth image. It is understood that the pre-trained second network model has the ability to perform super-resolution processing on images.

[0021] In a possible implementation of the first aspect above, the second network model includes the following components: a second encoder, a second diffusion model, an image encoder, a control network, a restoration network, and a second decoder; and the processing flow of the second network model on the second image includes: inputting N first blocks of the second image into the second encoder respectively, encoding the N first blocks in the latent space respectively, and obtaining N fifth features, where N is a positive integer; passing the N first blocks through the image encoder respectively to obtain N semantic features; inputting the N fifth features into the control network respectively, and outputting corresponding N sixth features, wherein the sixth feature is constrained by the corresponding semantic feature; inputting the N semantic features and the N sixth features into the second diffusion model, obtaining a corresponding seventh feature according to each sixth feature and the corresponding semantic feature, and outputting N seventh features corresponding to the N first blocks; splicing the N seventh features into an eighth feature, and upsampling the eighth feature in the latent space to obtain a ninth feature; inputting the ninth feature into the restoration network, processing data from different seventh features in the ninth feature to obtain a tenth feature; inputting the tenth feature into the second decoder, and outputting the fifth image. In this way, the present application can perform super-resolution processing on image features in latent space based on image segmentation technology and image semantic information.

[0022] In one possible implementation of the first aspect, the method further includes: training the second decoder based on at least one second training data to obtain a trained second decoder; wherein the second training data includes a first training block and a first training semantic feature, and the first training semantic feature is a semantic feature of the first training block. That is, the present application can train the second decoder based on semantic feature-image block pairs as training data.

[0023] In a possible implementation of the first aspect, the training process of the second decoder based on the second training data comprises: inputting the first training patch into the second encoder, encoding the plurality of first training patches in the latent space to obtain corresponding eleventh features; inputting the eleventh features into the control network to output corresponding twelfth features, where the twelfth features take the corresponding training semantic features as constraint conditions; inputting the first training semantic features and the twelfth features into the second diffusion model, and outputting the thirteenth features corresponding to the first training patches; inputting the thirteenth features into the second decoder to output the second training result; and adjusting the network parameters of the second decoder according to the second training result and the first training patch. In this way, the control network set can accurately restore the features of the patches in the latent space to image data in the image space with semantic features as constraint conditions.

[0024] In a possible implementation of the first aspect, the method further comprises: training the repair network according to at least one third training data to obtain a trained repair network; and wherein the third training data comprises a second training image and a third training image, and the third training image is an image obtained by reducing the resolution of the second training image.

[0025] In a possible implementation of the first aspect, the training process of the repair network based on the third training data comprises: inputting N second training patches of the third training image into the second encoder respectively, encoding the N second training patches in the latent space respectively to obtain N fourteenth features, where N is a positive integer; inputting the N fourteenth features into the second diffusion model to output N fifteenth features corresponding to the N second training patches, where the data amount of the fifteenth features is greater than that of the corresponding fourteenth features; splicing the N fifteenth features into a sixteenth feature, and performing up-sampling on the sixteenth feature in the latent space to obtain a seventeenth feature; inputting the seventeenth feature into the repair network to process the data in the seventeenth feature from different fifteenth features to obtain an eighteenth feature; inputting the second training image into the second encoder to encode the second training image in the latent space to obtain a nineteenth feature; and adjusting the network parameters of the repair network according to the difference between the eighteenth feature and the nineteenth feature.

[0026] In a possible implementation of the first aspect, the N first patches and the N second training patches are obtained by using a regular patching manner or an overlapping patching manner. For example, the overlapping patching manner can be an overlap patching manner, in which case the effect of splicing each patch is good, and the problem of distortion at the joint after splicing is less likely to occur.

[0027] In one possible implementation of the first aspect, the fourth image is an image generated based on text information input by the user, or a panoramic high dynamic range (HDR) image, or a non-HDR image. It will be appreciated that when the fourth image is an HDR image, the corresponding lighting information may be pre-collected. When the fourth image is a non-HDR image, the corresponding lighting information may be estimated using a lighting estimation algorithm.

[0028] In a possible implementation of the first aspect above, the illumination information of the fourth image is preset, pre-collected, or calculated based on the content of the fourth image (ie, estimated using an illumination estimation algorithm).

[0029] In one possible implementation of the first aspect, the fourth image is an HDR image, and the lighting information of the fourth image is pre-collected lighting information of the ambient light in the shooting scene where the fourth image is located. In this case, the image processing method provided in this application can be applied to shooting scenes in professional studios.

[0030] In a possible implementation of the first aspect, the third image includes a frame of a dynamic image, a still image, or a frame of a video; or the fourth image includes a frame of a dynamic image, a still image, or a frame of a video. Furthermore, the third and fourth images may be user-defined, and the image formats may also be user-defined.

[0031] In a possible implementation of the first aspect, the third image includes: an image in a preview interface of a photo or video mode, an image obtained by a photo operation, an image in a video obtained by video capture, an image in a video conference, an image in a video call, or an image used as a desktop wallpaper. Application scenarios of the image processing method of the present application include, but are not limited to, the above examples.

[0032] In a possible implementation of the first aspect, if there are multiple foreground objects in the third image, there may be one or more foreground objects moved to the fourth image.

[0033] In a possible implementation of the first aspect, the method further includes: displaying a first interface, the first interface including a first control, the first control being used to select at least one of the third image and the fourth image; detecting a user operation on the first control, and acquiring the third image and the fourth image based on the operation. For example, the first control may be as follows Figure 3 In the photo scene shown, the background control 13 or the first control can be Figure 6 The add image control 64 is shown as well as Figure 7 Background controls 65 are shown.

[0034] In a second aspect, an embodiment of the present application provides a readable medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the method in the above-mentioned first aspect and any possible implementation thereof.

[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the method in the above-mentioned first aspect and any possible implementation thereof.

[0036] In a fourth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device implements the method in the above-mentioned first aspect and any possible implementation thereof.

[0037] Among them, the description of the beneficial effects of the second to fourth aspects in this application can refer to the relevant description of the first aspect above, and this application will not elaborate on this. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1A According to some embodiments of the present application, a schematic diagram of an image change process during a background replacement process is shown;

[0039] Figure 1B According to some embodiments of the present application, a schematic diagram of a process for replacing the background of an image is shown;

[0040] Figure 2A According to some embodiments of the present application, a schematic diagram of an image change process during a background replacement process is shown;

[0041] Figure 2B According to some embodiments of the present application, a schematic diagram of an image change process during a background replacement process is shown;

[0042] Figure 2C According to some embodiments of the present application, a schematic diagram of a process for replacing the background of an image is shown;

[0043] Figure 3 According to some embodiments of the present application, a schematic diagram of a photographing scene is shown;

[0044] Figure 4 According to some embodiments of the present application, a schematic diagram of a text-based photo-taking scenario is shown;

[0045] Figure 5 According to some embodiments of the present application, a schematic diagram of a voice-based photo-taking scenario is shown;

[0046] Figure 6According to some embodiments of the present application, a schematic diagram of a custom wallpaper scene is shown;

[0047] Figure 7 According to some embodiments of the present application, a schematic diagram of a custom wallpaper scene is shown;

[0048] Figure 8 According to some embodiments of the present application, a flowchart of an image processing method is shown;

[0049] Figure 9 According to some embodiments of the present application, a schematic diagram of a custom wallpaper scene is shown;

[0050] Figure 10 According to some embodiments of the present application, a structural schematic diagram of a diffusion realization model is shown;

[0051] Figure 11A According to some embodiments of the present application, a schematic diagram of a process of image redrawing is shown;

[0052] Figure 11B According to some embodiments of the present application, a schematic flow chart of feature processing in image redrawing is shown;

[0053] Figure 12 According to some embodiments of the present application, a schematic diagram of a training process of a diffusion realization model is shown;

[0054] Figure 13 According to some embodiments of the present application, a flow chart of a method for training a diffusion realization model is shown;

[0055] Figure 14 According to some embodiments of the present application, a flow chart of a diffusion model training method is shown;

[0056] Figure 15 According to some embodiments of the present application, a schematic diagram of a training process of a diffusion enhancement model is shown;

[0057] Figure 16 According to some embodiments of the present application, a schematic diagram of a process of image super-resolution processing is shown;

[0058] Figure 17 According to some embodiments of the present application, a flow chart of a method for training a diffusion enhancement model is shown;

[0059] Figure 18 According to some embodiments of the present application, a schematic diagram of a training process of a decoder in a diffusion enhancement model is shown;

[0060] Figure 19According to some embodiments of the present application, a flowchart of a method for training a decoder in a diffusion enhancement model is shown;

[0061] Figure 20 According to some embodiments of the present application, a schematic diagram of a training process of a repair network in a diffusion enhancement model is shown;

[0062] Figure 21 According to some embodiments of the present application, a flow chart of a method for training a repair network in a diffusion enhancement model is shown;

[0063] Figure 22 According to some embodiments of the present application, a schematic structural diagram of a mobile phone is shown. DETAILED DESCRIPTION

[0064] Illustrative embodiments of the present application include, but are not limited to, image processing methods, media, electronic devices, and program products.

[0065] As known from the background art, a user may need to replace the background of an image, such as a portrait image, in order to change the background environment where a foreground object, such as a portrait, is located in the image.

[0066] Combine Figure 1A The image changing process during the image background replacement process is shown, and Figure 1B The image processing scenario is described below using an example of a process for replacing the background of an image.

[0067] like Figure 1A As shown, the background of the foreground image A1 to be processed is an outdoor background, and the portrait A11 is located in front of a tree. When the user wishes to change the background of the portrait A11 in image A1, for example, to the indoor background of the background image A2, the portrait A11 can be moved to the background image A2 to construct a new image A3.

[0068] like Figure 1B As shown, the image processing scenario includes a foreground processing flow and a background processing flow. Specifically, the foreground processing flow is used to process the image A1 whose background is to be replaced, specifically the portrait A11 in the image A1. The background processing flow is used to process the image A2 that provides the background.

[0069] Specifically, such as Figure 1B The image processing flow shown includes the following steps:

[0070] S1: Input foreground image A1 and background image A2.

[0071] For example, the foreground image A1 may be a portrait image captured in real time by the electronic device or a pre-generated portrait image.

[0072] S2: Down-sample the foreground image A1 to obtain a foreground image A1'.

[0073] For example, the size of the foreground image A1 is 3K×4K pixels, and the size of the downsampled foreground image A1′ is 768×1K pixels. Of course, the sizes of the foreground image A1 and the downsampled foreground image A1′ are not limited to the above examples, and can also be other sizes, which are not specifically limited in the embodiments of the present application.

[0074] S3: Perform foreground prediction on the foreground image A1' to obtain a mask image A1-1 of the foreground portrait A11.

[0075] In some embodiments, when the foreground image A1 ′ is a portrait image, the foreground prediction may be portrait segmentation, which is used to determine a portrait, such as a foreground portrait A11 , from the foreground image A1 ′.

[0076] As an example, the above-mentioned portrait segmentation usually involves semantic segmentation, that is, classifying each pixel in the image into a corresponding category (such as person, sky, building, etc.).

[0077] It can be understood that the mask image A1 - 1 is used to determine the outline and shape of the portrait A11 .

[0078] S4: Perform normal prediction on the foreground image A1' to obtain a normal image A1-2.

[0079] In some embodiments, the normal prediction of the foreground image A1 ′ may specifically be the normal prediction of the foreground portrait A11 in the foreground image A1 ′.

[0080] It can be understood that normal prediction is used to estimate the normal vector of each pixel in the image, that is, normal information. Specifically, normal information represents the degree of inclination and direction of the surface at each point, which is crucial for understanding the surface geometry of the object.

[0081] In some embodiments, normal image A1-2 represents the normal information of foreground image A1'. Specifically, normal image A1-2 can be a normal map with the same resolution as foreground image A1'. In the normal map, the color value of each pixel represents the component of the surface normal vector at the corresponding image pixel location. Typically, the red, green, and blue channels correspond to the x, y, and z components of the normal vector, respectively.

[0082] S5: Perform albedo prediction on the foreground image A1' to obtain the albedo image A1-3.

[0083] In some embodiments, the albedo prediction of the foreground image A1 ′ may specifically be the albedo prediction of the foreground portrait A11 in the foreground image A1 ′.

[0084] Among them, albedo prediction is used to predict the color and material properties of the portrait, including diffuse reflectance and specular reflectance. This information is used to simulate the material response under different lighting conditions.

[0085] In some embodiments, the albedo image A1 - 3 represents the albedo information of the foreground image A1 ′. Specifically, each pixel value in the albedo image A1 - 3 represents the albedo information of the corresponding position in the image.

[0086] In addition, in some embodiments, the foreground prediction, normal prediction, albedo prediction and other algorithms mentioned above can be implemented through a deep learning algorithm using a preset deep model.

[0087] S6: Perform image illumination estimation on the background image A2 to obtain illumination information of the background image A2.

[0088] The illumination information of the background image A2 is used to represent the characteristics of the illumination in the image, such as the direction, color, intensity, and distribution of the illumination.

[0089] S7: Move the foreground portrait A11 to the background image A2 according to the mask image A1-1, and perform physical relighting on the foreground portrait A11 in the background image A2 according to the normal image A1-2 and the albedo image A1-3 to generate a composite image A3.

[0090] Optionally, the composite image A3 may be upsampled to obtain an upsampled composite image A3.

[0091] Among them, physical relighting is a technology involving computer graphics and visual effects, which allows the lighting effects of light sources on objects to be re-simulated in a virtual environment.

[0092] As an example, the physical relighting in this application can use a physically based rendering (PBR) model, combined with normal and albedo information, to simulate the lighting effects under a new background environment. For example, based on PBR, by combining normal information A1-2, albedo image A1-3, and the lighting information of background image A2, the lighting effects of background image A2 can be simulated for portrait A11, thereby obtaining a composite image A3 after physical relighting. In this way, the lighting conditions of portrait A11 after moving are the same as those of background image A2.

[0093] However, there are errors in the above-mentioned image illumination prediction, normal prediction, and albedo prediction algorithms, which can make the image after the background is changed appear unrealistic. For example, Figure 1A The composite image A3 after relighting has problems such as blurred edges, damaged edges, or incomplete background removal, which results in a certain unrealistic feeling for the portrait A11 in the composite image A3.

[0094] In order to solve the problem of unrealistic images caused by blurred and damaged edges in composite images after background replacement, the present application provides an image processing method. In this image processing method, for the composite image after background replacement, the edge area between the foreground object and the background in the composite image is subjected to image synthesis, for example, by image redrawing methods such as style transfer, image enhancement, color correction, texture creation, or shape modification. The edge area includes: the outline of the foreground object, the content adjacent to the outline inside the foreground object, and the content adjacent to the outline outside the foreground object. In this case, the edge area is the area where the foreground object and the background are merged in the composite image after relighting. Therefore, redrawing the edge area is to repair the area where the foreground object and the background are merged. For example, in the edge area of ​​the composite image after background replacement, the present application can generate new pixels to fill the gaps in the broken or discontinuous edges of the foreground object through image redrawing, and set the color of the new pixels to be the same or close to that of the adjacent pixels, thereby eliminating the problem of edge blur and improving the consistency of lighting conditions between the foreground object and the background in the composite image.

[0095] At the same time, the method also redraws other areas of the composite image except the edge area, that is, the foreground part within the edge area and the background part outside the edge area, and the redrawing parameters of the edge area are greater than the redrawing parameters of the other areas. It can be understood that the redrawing parameters of the image area refer to the magnitude of the change in the pixel value of the image area, or the degree of style migration of the image, etc. For example, the change in the pixel value of the above-mentioned edge area is greater than the change in the pixel value of the above-mentioned other areas, or the degree of style migration of the above-mentioned edge area is greater than the degree of style migration of the above-mentioned other areas. In this way, while the edge of the outline of the foreground object in the composite image is repaired, it will not cause significant changes to the foreground object itself and the background, such as the replacement of the foreground portrait will not cause significant distortion. Thus, the naturalness and reality of the image fusion can be guaranteed.

[0096] In some embodiments, in order to solve the problem of image untruth caused by edge blur and edge damage in the above-mentioned re-lit synthetic image, the edge region where the foreground object is located in the re-lit synthetic image and other regions can be redrawn respectively, and the redrawn parameters of the edge region are greater than those of the other regions. In this way, the authenticity and naturalness of image fusion are ensured.

[0097] Of course, the synthetic image after changing the background in the present application is not limited to the above-mentioned re-lit synthetic image, but can also be an image after changing the background and without re-lighting the foreground object, and the present application embodiments do not make specific limitations thereto.

[0098] Specifically, the edge region and other regions in the synthetic image can be redrawn based on a trained neural network model with image redrawing function (hereinafter referred to as a diffusion realism model, and taken as an example of the first network model), to repair the edge region and perform color correction, thereby improving the realism of the synthetic image.

[0099] For example, in some embodiments, the training process of the diffusion realism model can be: obtaining training texts and corresponding high-definition training images, which represent the expected output results. Then, the training texts are input into a preset neural network model, so that the neural network model generates training result images corresponding to the training texts. Then, the actually generated training result images are compared with the corresponding high-definition training images, to adjust the parameters of the neural network model based on the differences between the two images. By analogy, a plurality of training texts and respective corresponding high-definition training images can be used to adjust the parameters of the neural network model, so as to obtain a trained diffusion realism model. It can be understood that the diffusion realism model is trained to have image redrawing ability through text-to-image technology, i.e., from text to image.

[0100] In some other embodiments, the training process for the diffusion realism model may include obtaining a low-quality training image and a corresponding high-quality training image, where the high-quality training image represents the desired output. For example, the low-quality training image may be a composite image generated by conventional relighting techniques, where a foreground object and background are fused together, while the high-quality training image may be an artificially restored image of the composite image, where the foreground object and background are more realistically fused. The low-quality training image is then input into a preset neural network model, causing the neural network model to generate a training result image corresponding to the low-quality training image. The generated training result image is then compared with the corresponding high-quality training image, and the parameters of the neural network model are adjusted based on the differences between the two images. Similarly, multiple low-quality training images and their corresponding high-quality training images can be used to adjust the parameters of the neural network model, thereby obtaining a trained diffusion realism model. It can be understood that the diffusion realism model is trained to acquire image redrawing capabilities by restoring high-quality images from low-quality images.

[0101] As an example, the diffusion realisation model may be a U-Net model, but is not limited thereto. Furthermore, the network structure, training and use process of the diffusion realisation model will be described in detail in this application and will not be elaborated here.

[0102] In addition, the different resolutions of the foreground image and the background image can also make the image feel unrealistic. Figure 1A In the example, since the resolution of the foreground image A1 is different from that of the background image A2, the resolution of the portrait A11 in the image A3 may be different from that of the background, further increasing the unreality of the portrait A11. Furthermore, in some embodiments, the image processing method provided by the present application can redraw the edge area between the foreground object and the background in the re-lit composite image while also redrawing the composite foreground object and background to a certain extent, such as adding detailed features to the composite image as a whole. In this way, the consistency of the lighting conditions between the foreground object and the background in the composite image is further improved.

[0103] Furthermore, in some embodiments, to further improve the quality of the relighted image, the present application may perform super-resolution processing on the redrawn image obtained by the redrawing process to obtain an enhanced image with higher resolution. Specifically, the super-resolution processing may divide the redrawn image into blocks in a latent space and fuse the latent features of each block to obtain an enhanced image.

[0104] It can be understood that latent features of an image are key information that is not directly observable but is helpful for a task, such as edges, textures, shapes, and object parts. The latent space refers to the low-dimensional continuous space to which data is mapped within a neural network. By operating on data in the latent space, various functions can be achieved, such as dimensionality reduction, denoising, and style transfer.

[0105] In some embodiments, the super-resolution processing of the image of the present application can be achieved through a network model (hereinafter referred to as a diffusion enhancement model, or a second network model).

[0106] For example, the present application can generate a diffusion enhancement model (as an instance of the second network model) based on the diffusion model generation technology to perform image enhancement on the redrawn image through the diffusion enhancement model, that is, perform super-resolution processing on the redrawn image.

[0107] As an example, the diffusion enhancement model can be a U-Net model, but is not limited thereto. In addition, the network structure, training and use process of the diffusion enhancement model will be described in detail in this application and will not be repeated here.

[0108] Next, for the convenience of explanation, Figure 2A The image changing process during the image background replacement process is shown, and Figure 2B The image processing scenario is described below using an example of a process for replacing the background of an image.

[0109] As an example, this application takes the composite image A3 after the foreground image A1 and the background image A2 are re-lit as an example to illustrate the image redrawing and super-resolution processing in the image processing method provided by this application.

[0110] exist Figure 1A On the basis of Figure 2A As shown in FIG, for the composite image A3 with the portrait A1 added, the image can be redrawn to obtain a redrawn image A4. Figure 2AAs shown, for the outline A12 of the portrait A11 in the relit composite image A3, an edge region A13 (as an example of a first region) can be determined where the outline A12 lies. Edge region A13 is formed by expanding the pixel content inward (i.e., toward the portrait A11) and outward (i.e., toward the background) along the outline A12. This means that edge region A13 includes not only the outline A12 but also portions of the portrait A11 and the background. Furthermore, an area A14 (as an example of a second region) in image A3, excluding the edge region, includes the interior of the portrait A11 and the remaining background. In this way, the present application can redraw the edge area A13 and other areas A14 in the composite image A3 according to different redrawing parameters, with the redrawing parameters for the edge area A13 being larger and the redrawing parameters for the other areas A14 being smaller, so that the area A14 will not be distorted due to the larger redrawing parameters, and will not affect the authenticity of the portrait itself and the background in the other areas A14, thereby obtaining the repaired redrawn image A4.

[0111] In some embodiments, Figure 2A On the basis of Figure 2B As shown in Figure 1, the redrawn image A4 can be divided into blocks in the latent space and super-resolution processing can be performed on the latent features of each block, such as adding detailed features to each block. Furthermore, these blocks with enhanced resolution are fused to achieve super-resolution processing of the redrawn image A4, resulting in an enhanced image A5 with a higher resolution.

[0112] exist Figure 1B On the basis of Figure 2C As shown, the image processing scenario includes a foreground processing flow and a background processing flow.

[0113] Specifically, if Figure 2B The image processing flow shown includes the following steps:

[0114] S1-S7, among which Figure 2B S1-S7 in Figure 1B S1-S7 are the same as those in the previous section and are not described here.

[0115] S8: Redraw image A3 using the diffusion realization model to obtain a redrawn image A4.

[0116] Combine Figure 2A It can be seen that the diffusion realization model redraws the edge region A13 of image A3 where the portrait A11 is located with greater parameters than the redrawing parameters of the other regions A14 outside the edge region A13 of image A3, thereby improving the image fusion effect while maintaining the realism of the portrait.

[0117] Optionally, the redrawing of the image A3 can be redrawing of the up-sampled composite image A3.

[0118] In the present application, the process of redrawing the image using the diffusion real model will be described in detail below, and will not be repeated here.

[0119] S9: using the diffusion enhancement model to perform super-resolution processing on the redrawn image A4 to obtain an enhanced image A5.

[0120] In combination Figure 2B It can be seen that the diffusion enhancement model blocks and fuses the image A4, realizes the enhancement effect of adding detailed features to the image A4, and thus improves the resolution of the image A5 and the quality of the image after changing the background.

[0121] In this way, in the scenario of changing the image background, the present application can use the diffusion model to redraw the composite image after light retouching to realize edge repair and further super-resolution processing. In this way, the image after changing the background greatly improves the realism of the fusion of the foreground object and the background on the basis of keeping the consistency of the illumination of the background and the foreground.

[0122] In the following embodiments, for the convenience of description, the foreground image to be changed in the background, the background image providing the background, the composite image after light retouching, the redrawn image after redrawing, and the enhanced image after super-resolution processing can be referred to as the third image, the fourth image, the first image, the second image and the fifth image.

[0123] It can be understood that at least one of the third image to be changed in the background and the fourth image providing the background can be selected by the user. For example, the third image and the fourth image can be selected based on the application scenario of image processing or the file to which the image belongs.

[0124] In some embodiments, the image processing method provided by the present application can be applied to different application scenarios, such as a photographing scene (such as a photographing scene of a photographing function of a system camera), a video shooting scene, a photograph preview scene, a video conference scene, a video call scene, a photo editing scene, a wallpaper changing scene, etc., but not limited thereto. For example, the above-mentioned application scenarios can also include long and short video applications, video live streaming applications, video online course applications, portrait intelligent lens moving application scenarios, video monitoring, intelligent cat eyes, etc. It can be understood that the third image is an image in these scenarios. For example, in the video call scene, the third image can be an image of the current frame in the video picture captured by the camera of the electronic device.

[0125] Optionally, the fourth image may be an image generated based on text information input by the user (i.e., the fourth image is an image generated using Vincent image technology), a panoramic high dynamic range (HDR) image, or a pre-set non-HDR image. Similarly, the third image may also be an image generated using Vincent image technology, an HDR image, or a pre-set non-HDR image.

[0126] As an example, the HDR image can be a 360-degree panoramic image, which can be acquired by a professional camera during the acquisition process. In this case, the wide-angle lighting information of the ambient light of the shooting scene, such as the intensity, direction, and color of the light, can be acquired simultaneously during the acquisition of the HDR image. Then, when the fourth image is an HDR image, the lighting information of the pre-acquired fourth image can be directly acquired while acquiring the fourth image, without the need to perform an image lighting estimation algorithm on the fourth image. For example, in some studio shooting scenes, an HDR image can be used as an image to provide a background to replace the background of a foreground image.

[0127] As an example, the aforementioned pre-set non-HDR image may be an ordinary image, such as an image obtained from a server or an image stored locally.

[0128] As an example, the fourth image generated based on the above-mentioned Wensheng graph technology can be an image automatically generated based on the text information input by the user in real time. The Wensheng graph technology can be generated by generating a neural network, such as a stable diffusion model (SD). It can be understood that the above-mentioned text information may include the foreground object, background, style, lighting information, etc. of the image. For example, the text information input by the user (also called prompt) may include "best quality, masterpiece, ultra-high resolution, original photo, movie lighting, a girl, dark style, night, city". In this case, the foreground object of the generated fourth image is a girl, and the background is a dark city night scene.

[0129] It can be understood that when the fourth image is a non-HDR image generated in real time or preset, the illumination information of the fourth image can be predicted based on an image illumination estimation algorithm.

[0130] In addition, in some embodiments, the files to which the foreground image and background image to be replaced in the present application belong may be animated images, ordinary pictures, or videos, etc. Specifically, the third image to be replaced with the background in the present application may include: a frame of image in an animated image, a frame of static image (such as an ordinary picture), or a frame of image in a video. In addition, the fourth image providing the background may include: a frame of image in an animated image, a frame of static image, or a frame of image in a video. Among them, the above-mentioned animated image refers to a file in the graphics interchange format (GIF) format, which can contain multiple frames of images and support animation and transparency.

[0131] It can be understood that the files of the third image and the fourth image can be selected by the user. For example, the user can select different files in different application scenarios. As an example, in a photo shooting scenario, the third image and the fourth image can both be static images. In a video scene or a photo editing scene, the third image and the fourth image can both be any realizable files. For example, the third image is a static portrait of a foreground object, and the fourth image is a dynamic image of a seaside that changes frame by frame from sunrise to sunset. After moving the portrait to the fourth image and relighting to obtain a composite image, the composite image can be redrawn and super-resolution processed in this application so that the lighting information of the portrait in the final image changes with the change of ambient light.

[0132] In some embodiments, the third image may not include a background. For example, in a photo editing scenario, the third image may be a foreground portrait or other foreground object that only needs to be replaced with the background.

[0133] In some embodiments, the third image may include multiple foreground objects, such as multiple foreground portraits. In this case, replacing the background of the third image, as used herein, refers to moving one or more foreground portraits in the third image to a fourth image. It will be appreciated that the foreground objects to be moved in the third image can be user-defined, and the positions of the foreground objects to be moved in the fourth image can also be user-defined.

[0134] Next, the image processing method in this application is illustrated with examples based on some application scenarios.

[0135] The following combination Figures 3 to 5 , introduce the photo-taking scenes provided by this application.

[0136] In one example, the electronic device is taken as a mobile phone for example. Figure 3 As shown in (a) of FIG, in response to the user's click operation on the camera application, the mobile phone opens the camera and displays the following Figure 3The display interface of the photo mode shown in (b) of FIG. 1 may include a shooting interface 10. The shooting interface 10 may include a viewfinder 11 and a control 12 for indicating a photo. Before detecting that a user clicks control 12, a preview image may be displayed in the viewfinder 11. In the preview image, a person is standing outdoors in front of a large tree.

[0137] When the user wants to change the background of the portrait, he can Figure 3 The shooting interface 10 in (b) is slid upwards so that Figure 3 The mobile phone display shooting interface 10 shown in (c) and (d) has additional controls 13 indicating different backgrounds and controls 14 indicating cultural images. The HDR images in each background indicated by control 13 can be marked by the symbol "HDR" in the upper right corner of the background thumbnail. Similarly, the GIF images in each background indicated by control 13 can also be marked by a symbol on the background thumbnail, for example, by the symbol "GIF".

[0138] Then, when the user desires to change the current outdoor background to an indoor background, the user can make a selection on the display interface and click on the control 13 indicating the indoor background to switch the background.

[0139] like Figure 3 As shown in (c), the user can click on the control 13 indicating the indoor background. When the mobile phone detects that the user clicks on the control 13 indicating the seaside background, it responds to the user's operation, keeps the portrait unchanged, and switches the background of the portrait from outdoor to indoor.

[0140] like Figure 3 As shown in (d) in the figure, combined with the switched indoor background, after detecting the user clicking the control 12, the mobile phone can take a photo of the portrait with the indoor background in response to the user's operation.

[0141] In some embodiments, Figure 3 (d) shows that the image A4 after the background is switched in the preview interface 11 can be an image that is re-lit based on the indoor background, and the edge area where the portrait A11 is located is redrawn based on the diffusion network, and then an enhanced image is processed by super-resolution.

[0142] In other embodiments, Figure 3 (d) in the figure shows that the image after the background is switched in preview interface 11 can be an image that has been re-lit based on the indoor background. After the user clicks control 12 to trigger the photo, the phone can redraw the edge area of ​​the image portrait A11 in preview interface 11, and then perform super-resolution processing to enhance the image, that is, take a photo to obtain the enhanced image.

[0143] In addition, in some embodiments, if the background provided in the preview interface 11 does not meet the user's needs, the user can trigger the mobile phone to generate a new background image using the text image through the control 12.

[0144] like Figure 4 As shown in , the mobile phone can obtain the text information input by the user through text input. Specifically, Figure 4 As shown in (a) of FIG. 1 , when the mobile phone detects that the user clicks the control 14, the mobile phone may display the following in response to the user's operation: Figure 4 The mobile phone detects the user's operation on the keyboard in the input control 15, and in response to the operation, obtains the text information input by the user. Then, after the user completes the input of the text information "movie light night dark style city", the mobile phone detects the user's click operation on the confirmation control 16 in the input control 15, and in response to the operation, the mobile phone automatically generates a background image of the city night scene, and displays the background image as shown in FIG. Figure 4 As shown in (c), a control 13 corresponding to the background of the city night scene is displayed in the preview interface 11, and the portrait A11 in the preview interface 11 is replaced with the city night scene.

[0145] like Figure 4 As shown in (c) in the figure, combined with the switched indoor background, after detecting the user clicking the control 12, in response to the user's operation, the mobile phone can take a photo of the portrait with the city night scene as the background.

[0146] like Figure 5 As shown in , the mobile phone can obtain the text information input by the user through voice input. Specifically, Figure 5 As shown in (a) in the figure, when the mobile phone detects that the user has long pressed the control 14, in response to the user's operation, the mobile phone can receive the user's voice input "movie lights night dark style city". Then, after the mobile phone detects that the user stops long pressing the control 14, it determines that the user has completed the voice input and can detect the text information "movie lights night dark style city" from the voice. In this way, for the text information "movie lights night dark style city", the mobile phone automatically generates a background image of the city night scene, and displays it as shown in the figure. Figure 5 As shown in (b) , a control 13 corresponding to the background of the city night scene is displayed in the preview interface 11 , and the portrait A11 in the preview interface 11 is replaced with the city night scene.

[0147] like Figure 5 As shown in (b), combined with the switched indoor background, after detecting the user's click on the control 12, in response to the user's operation, the mobile phone can take a photo of the portrait with the city night scene as the background.

[0148] akin, Figure 4 or Figure 5 The image displayed in the portrait preview interface after A11 replaces the city night scene can be a re-lit image, or an image after re-lighting, image redrawing and super-resolution processing.

[0149] The following combination Figure 6 and Figure 7 The wallpaper scene provided in the embodiment of the present application is described.

[0150] For example, take the user-defined wallpaper scene of a mobile phone as an example. Figure 6 As shown in (a) in the figure, in response to the user's click operation on the settings in the desktop interface, the mobile phone displays Figure 6 In response to the user clicking on the wallpaper option 61 in the setting interface 60, the mobile phone displays the following Figure 6 The personalized wallpaper setting interface 63 in (c) of FIG. The control 64 for adding an image in the personalized wallpaper setting interface 63 is used to select a foreground image of the wallpaper, and the control 65 for indicating different backgrounds is used to select a background image of the wallpaper.

[0151] like Figure 7 As shown, in response to the user clicking the control 64 for adding an image in the personalized wallpaper setting interface 63, the mobile phone displays the following Figure 7 In response to the user's click operation on the picture 38 in the picture selection interface 67, the mobile phone displays the following Figure 7 In response to the user's personalized wallpaper setting interface 63 shown in (b) Figure 7 Click the control 65 indicating the landscape background in the personalized wallpaper setting interface 63 shown in (b) of FIG. 1 , and the mobile phone displays the following Figure 7 The personalized wallpaper setting interface 63 shown in (c) of FIG. 6 shows an image 69 after the landscape background is changed. Figure 7 By clicking the control 66 set as wallpaper in (c), the phone can display the following Figure 7 The desktop interface is shown in (d), and the wallpaper of the desktop interface is image 69 after the landscape background is replaced.

[0152] I understand. Figure 7 The image 69 after the landscape background is replaced may be an enhanced image after relighting, image redrawing and super-resolution processing based on the lighting information of the landscape background.

[0153] In some embodiments, Figure 7In the image 69 after the replacement scenery background shown in , the scenery background is a dynamic image from sunrise to sunset. At this time, the image 69 is a dynamic image, and the lighting situation of the portrait in the image 69 changes along with the lighting situation change in the scenery background.

[0154] It can be understood that the description of other application scenarios of the image processing method of the embodiment of the present application is similar to the above-mentioned photo-taking scenario and wallpaper scenario, and the embodiment of the present application will not be repeated here.

[0155] The image processing method provided in the embodiments of the present application can be applied to various electronic devices.

[0156] In the embodiment of the present application, the electronic device 100 can be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, an in-vehicle electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, etc. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device 100.

[0157] In the embodiments of the present application, for the convenience of explanation, the electronic device is taken as an example of a mobile phone to illustrate the execution subject of the image processing method of the present application.

[0158] Further, combined Figures 8 to 21 The specific implementation of the image processing method provided in the embodiment of the present application is described in detail.

[0159] like Figure 8 FIG. 1 is a flow chart of an image processing method according to an embodiment of the present application, and the flow chart includes:

[0160] S801: Acquire a third image whose background is to be replaced and a fourth image that provides a background.

[0161] As an example, the third image may include one or more foreground objects.The fourth image may be a graphic obtained by using a Vincent graph technique, a pre-set non-HDR image, or an HDR image.

[0162] S802: Perform downsampling processing on the third image to obtain a downsampled third image.

[0163] For example, the size of the third image is 3K×4K pixels, and the size of the downsampled third image is 768×1K pixels. Of course, the sizes of the third image and the downsampled third image are not limited to the above examples and can also be other sizes, which are not specifically limited in this embodiment of the present application. As an example, the third image and the fourth image can be the foreground image A1 and background image A2 shown above, respectively.

[0164] S803: Determine a first mask image of the foreground object from the downsampled third image.

[0165] Among them, the first mask image is used to represent the outline and shape of the foreground object. It can be understood that a mask image (maskimage) generally refers to an image used to specify an operating area or hide certain areas. The mask can be a black and white binary image, in which the white part represents the region of interest (ROI) and the black part is ignored or hidden. For example, for the foreground image A1', the area where the portrait A11 is located can be the white part, and the other areas can be the black areas. In addition, the mask can also be color, in which each channel specifies a different processing method.

[0166] In some embodiments, the foreground object determined in the third image may be a foreground object to be moved, and the foreground object may be one or more foreground objects in the third image.

[0167] S804: Perform normal prediction on the downsampled third image to obtain normal information of the third image, wherein the normal information is used to reflect the direction of the surface of the foreground object in the third image.

[0168] S805: Perform albedo prediction on the downsampled third image to obtain albedo information of the third image.

[0169] Among them, albedo information is used to reflect the color and material properties of the foreground object.

[0170] S806: Add the foreground object in the downsampled third image to the fourth image.

[0171] In some embodiments, the present application can interactively add the foreground object to be moved in the third image to the fourth image.

[0172] Reference Figure 9 As shown in FIG, it is a schematic diagram of the movement of the foreground object in the wallpaper scene provided by the embodiment of the present application. Figure 9 The position of the portrait in the personalized wallpaper setting interface 63 shown in (a) is L1. In response to the user's operation of dragging the portrait at position L1 to the left, the mobile phone can move the portrait to the left and Figure 9The personalized wallpaper setting interface 63 shown in (b) displays the portrait at position L2. That is, the portrait is moved from position L1 on the landscape background to position L2. In this way, the display effect of the portrait after the background is changed can meet the user's needs.

[0173] In some embodiments, the present application may also downsample the fourth image and add the foreground object in the downsampled third image to the downsampled fourth image.

[0174] S807: Relight the foreground object in the fourth image according to the first mask image, the normal information and the albedo information, and the lighting information of the fourth image to obtain the first image.

[0175] The illumination information of the foreground object in the first image is the same as the illumination information of the fourth image.

[0176] In some embodiments, the lighting information in the fourth image can be predicted from the fourth image using an image lighting estimation algorithm. For example, when the fourth image is an image generated using Vincent image technology or a pre-set non-HDR image, its lighting information can be predicted using an image lighting estimation algorithm.

[0177] In addition, when the fourth image is an HDR image, the collected illumination information corresponding to the HDR image may be directly acquired.

[0178] It can be understood that for the above descriptions of S801 to S806, reference can be made to the above descriptions of S1 to S6, and the similarities are not repeated here.

[0179] S808: Perform upsampling processing on the first image to obtain an upsampled first image.

[0180] It is understood that the resolution of the upsampled first image is higher than the resolution of the un-upsampled image. The present application does not specifically limit the parameters for upsampling the third image, which can be determined according to actual needs.

[0181] Optionally, in some other embodiments, steps S802 and 808 may not be performed. That is, the present application may directly perform foreground prediction, normal prediction, albedo prediction, and other algorithms on the first image with a higher resolution. Accordingly, the relighting, image redrawing, and image super-resolution processing of the present application are also performed on the first image that has not been downsampled.

[0182] S809: For the upsampled first image, redraw the first region according to the first redrawing parameter, and redraw the second region according to the second redrawing parameter to obtain a second image.

[0183] The first redrawing parameter is different from the second redrawing parameter, for example, the first redrawing parameter is greater than the second redrawing parameter.

[0184] As an example, when the first redraw parameter is greater than the second redraw parameter, the number of redraws of the first region is greater than the number of redraws of the second region, and / or the proportion of new features in the first region after redrawing is greater than the proportion of new features in the second region after redrawing. For example, the present application may redraw the first region where the foreground object is located five times in the pre-trained diffusion realization model, and redraw the second region twice in the pre-trained diffusion realization model. A greater number of redraws results in more new features in the redrawn image.

[0185] Among them, the first area, that is, the edge area where the foreground object is located after moving in the third image, includes the following contents: the outline of the foreground object, the area where the partial content of the foreground object adjacent to the outline is located, and the partial content adjacent to the outline outside the foreground object; the second area is the area outside the first area in the third image.

[0186] Specifically, the first area includes the following: the outline of the foreground object, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object. For example, the multiple pixels adjacent to the pixels corresponding to the outline of the foreground object include: pixels within the foreground object whose shortest distance between pixels corresponding to the outline of the foreground object is less than a first distance, and pixels outside the foreground object whose shortest distance between pixels corresponding to the outline of the foreground object is less than a second distance. The first distance and the second distance may be the same or different, for example, both may be 10 pixels. The specific values ​​of the first distance and the second distance may also be other values, and this is not specifically limited. At this time, the first area in the first image includes: the outline of the foreground object, the area where the outline of the foreground object is extended inward by 10 pixels, and the area where the outline is extended outward by 10 pixels.

[0187] It can be understood that the shortest distance between pixel 1 corresponding to the outline of the foreground object and pixel 2 within the foreground object is: the distance between pixel 2 and pixel 1 in the direction perpendicular to the tangent line of the outline at pixel 1. In this case, the line connecting pixel 2 and pixel 1 is located on a line perpendicular to the tangent line of the outline at pixel 1.

[0188] In some embodiments, the present application may use a pre-trained diffusion realisation model to redraw an image. Specifically, the present application may input a third image into the pre-trained diffusion realisation model, redraw the third image using the diffusion realisation model, and output a redrawn fourth image.

[0189] Combine Figure 1A and Figure 2AIn the example shown, the third image may be foreground image A1, the foreground object to be moved is portrait A11, and the fourth image is background image A2. Specifically, the first image is composite image A3, and the second image is redrawn image A4. Furthermore, the edge region A13 of composite image A3, where portrait A11 is located, is the first region, while the remaining region A14 is the second region. Accordingly, the redrawn second image may be redrawn image A4.

[0190] S810: Perform super-resolution processing on the second image to obtain a fifth image.

[0191] The resolution of the fifth image is higher than that of the second image.

[0192] In some embodiments, the present application may use a pre-trained diffusion enhancement model to perform super-resolution processing on an image. Specifically, the present application may input a second image into a pre-trained diffusion enhancement model, perform resolution enhancement on the second image through the diffusion enhancement model, and output a fifth image with a higher resolution.

[0193] As an example, see Figure 2A The fifth image may be an enhanced image A5 obtained by performing super-resolution processing on the redrawn image A4. Obviously, the resolution of the enhanced image A5 is higher than that of the redrawn image A4, and the quality of the enhanced image A5 is higher.

[0194] I understand. Figure 8 The execution order of the steps is only an example, and other execution orders are also possible, for example, S804, S805, and S806 can be executed in parallel, etc. The embodiments of the present application do not make specific limitations on this.

[0195] In this way, the present application uses a pre-trained diffusion realization model to restore the edge contours of the foreground object in the relighted first image while improving the overall lighting harmony and clarity of the first image without causing significant changes to the foreground object itself or the background, such as causing significant distortion of the replaced foreground portrait. This ensures the naturalness and authenticity of the fused second image.

[0196] Application of diffusion realistic model

[0197] In some embodiments, the diffusion realization model in the present application may include multiple components, such as a first encoder, a first decoder, and a first diffusion model. The first encoder is used to obtain latent features of the image in the latent space, the first diffusion model is used to redraw the image based on the latent features, and the first decoder is used to restore the latent features in the latent space to an image in the image space.

[0198] Optionally, the first encoder and the first decoder may adopt a variational autoencoder (VAE) architecture, specifically a diffusion variational autoencoder (dVAE). In this case, the first encoder may be a dVAE encoder, and the first decoder may be a dVAE decoder.

[0199] Optionally, the first diffusion model can be implemented by a Unet model (a network model). Specifically, the first diffusion model can be a U-shaped structure in the Unet model, which is used to support the jump connection between the first encoder and the first decoder. Among them, UNet combines the shallow features and deep features of the image in a splicing manner. Among them, the shallow feature map is more inclined to express basic feature units such as points, lines, and edge contours, and contains more spatial information. The deep feature map is more inclined to express the semantic information of the image, contains less spatial information, and has more semantic features.

[0200] Reference Figure 10 , describes the structure of the diffusion reality model provided by the application embodiment. Figure 10 As shown, the diffusion realization model 100 includes the following components: a dVAE encoder 101, a diffusion model 102, and a dVAE decoder 103. For example, a relighted composite image A3 is input to the dVAE encoder 101, which extracts the original latent features of the composite image A3. The original latent features of the composite image A3 are then iteratively processed multiple times by the diffusion model 102 to obtain the redrawn latent features of the composite image A3. Furthermore, the redrawn latent features of the composite image A3 are decoded by the dVAE decoder 103 to restore the redrawn image A4.

[0201] Next, the image redrawing process based on the diffusion realization model is described in detail. Figure 11A As shown, the image redrawing process includes the following steps S1101 to S1103:

[0202] S1101: Input a first image into a first encoder, and encode the first image into a first feature in a latent space.

[0203] The first image may be data in an image space, that is, the first image may be a color digital image including three color channels: red (R), green (G) and blue (B).

[0204] It can be understood that the dVAE encoder can obtain latent feature data of the latent space (i.e., features with a low resolution of 8 bits). Specifically, the first image can be encoded into a lower-dimensional latent feature. This latent feature can be a 4-dimensional volume, where each volume slice can correspond to a specific feature or attribute in the decoded image. In this case, the latent feature has 4 channels, representing 3 color channels and image features (or parameters used to control different aspects of the image generation process).

[0205] For example, the first image is 1024×1024×3-channel image data, and the encoded first feature is 128×128×4-channel latent data. In this case, for a 128×128×4-dimensional latent feature data, this means there is a grid consisting of 128×128 feature points, each of which has four values ​​representing the feature of the point, such as data for three color channels and the value of one image feature. This data structure can be used to represent a compressed version of an image or as an intermediate step for image synthesis and editing.

[0206] S1102: Input the first feature into a first diffusion model, and process the first feature based on a preset noise to obtain a second feature, wherein a first redrawing parameter for redrawing a feature corresponding to the first region in the first feature based on the preset noise is greater than a second redrawing parameter for redrawing a feature corresponding to the second region in the first feature based on the preset noise.

[0207] Optionally, the preset noise may be Gaussian noise.

[0208] Gaussian noise follows a Gaussian distribution (normal distribution), which is characterized by randomness and continuity, and can vary across the entire image, mimicking the effects of real-world noise caused by various factors, such as sensor defects, transmission errors, and environmental interference. Gaussian noise parameters typically include its mean and standard deviation, which determine the intensity of the noise and the center of its distribution. In image remapping, different noise parameters can be selected depending on the desired noise level and image content.

[0209] In image redrawing, Gaussian noise can simulate real-world noise, making it closer to the characteristics of natural images and more realistic. Data augmentation: In tasks such as image classification and detection, adding Gaussian noise of varying intensities to images in the training set can increase data diversity and help the model learn more comprehensive knowledge.

[0210] Furthermore, Gaussian noise can prevent overfitting during the training of the diffusion realization model. It is understood that when training machine learning models (such as convolutional neural networks), adding Gaussian noise to the training data can help the model learn more robust features, thereby improving the model's ability to generalize to unseen data.

[0211] S1103: Input the second feature into the first decoder, and decode the second feature into a second image.

[0212] It can be understood that the first decoder, such as the dVAE decoder 103, can restore the second feature of the latent space to the second image in the image space, thereby completing the redrawing of the first image.

[0213] In some embodiments, the first diffusion model iteratively processes the first feature multiple times to obtain the second feature. Figure 11B As shown, the above S1102 includes S1102a to S1102g:

[0214] S1102a: Determine a first number of iterations k*a according to the first redrawing parameter a and the preset parameter k, and determine a second number of iterations k*b according to the second redrawing parameter b and the preset parameter.

[0215] For example, refer to Figure 10 , the first feature is the original latent feature of the synthetic image A3.

[0216] It can be understood that in the present application, the number of iterations of inputting the first region and the second region in the first image into the first diffusion model is different, that is, the number of times the first diffusion model redraws the images of the two regions is different.

[0217] The preset parameter k can be a preset number of iterations for the first diffusion model, for example, k = 10. The actual first number of iterations of the first diffusion model can be determined in combination with the preset number of iterations and the first redrawing parameter. For example, when the first redrawing parameter is 0.5, the actual total number of iterations of the first diffusion model, k*a, is 5, meaning that the number of iterations for the features corresponding to the first region in the first image is 5. Correspondingly, when the second redrawing parameter b is 0.2, the second number of iterations of the first diffusion model, k*b, is 2, meaning that the number of iterations for the features corresponding to the second region in the first image is 2. In this case, the second number of iterations is the number of iterations of the last several iterations of the total number of iterations, k*a.

[0218] S1102b: Superimposing the preset noise and the first feature according to the first redrawing parameter to obtain a first redrawing feature.

[0219] As an example, when the first redrawing parameter a=0.5, assuming that the first feature is i and the preset noise is n, the first redrawing feature=i*(1-a)+n*a.

[0220] S1102c: corresponding to 1≤j≤(k*ak*b), in the jth iteration process, input the jth redrawing feature into the first diffusion model for redrawing, and output the j+1th redrawing feature, where j is a positive integer.

[0221] It can be understood that when j=1, preset noise is added to the image redrawing in the first iterative process, and there is no need to add preset noise in the image redrawing in the subsequent iterative processes until the preset noise is introduced again in the image redrawing in the iterative process indicated by the second redrawing parameter.

[0222] As an example, when k=10, a=0.5, b=0.2, and k*(ab)=3, there is no need to introduce preset noise in the second and third iterations of the first diffusion model, and the input of the latter iteration is the output of the previous iteration, until the preset noise is introduced in the fourth iteration.

[0223] S1102d: Corresponding to j=(k*ak*b)+1, the preset noise is superimposed on the first feature according to the second redrawing parameter to obtain an intermediate feature, and the feature corresponding to the second area in the j-th redrawing feature is replaced with the feature corresponding to the second area in the intermediate feature to obtain an updated j-th redrawing feature.

[0224] In combination with the above example, when k=10, a=0.5, b=0.2, and k*(ab)=3, before j=4, the preset noise can be introduced again based on the second redrawing parameter. Then, the preset noise can be superimposed on the first feature according to the second redrawing parameter to obtain an intermediate feature = i*(1-b)+n*b. Furthermore, before the fourth iteration process, the feature corresponding to the second area in the fourth redrawing feature is replaced with the feature corresponding to the second area in the intermediate feature to obtain an updated fourth redrawing feature. It can be understood that the feature corresponding to the second area of ​​the first image in the updated fourth redrawing feature is obtained by introducing the preset noise according to the second redrawing parameter of 0.2 and is not input into the first diffusion model, and the feature corresponding to the first area of ​​the first image is obtained by introducing the preset noise according to the first redrawing parameter of 0.5 and iterating the first diffusion model 3 times.

[0225] Specifically, combined Figure 10, the original latent feature (i.e., the first feature) of the composite image A3 is denoted as C, the preset noise is denoted as N, the fourth redraw feature is denoted as X, the first redraw parameter is 0.5, and the second redraw parameter is 0.2. Furthermore, the feature corresponding to the first region where the portrait A11 is located in the fourth redraw feature is denoted as mask, and the feature corresponding to the second region where the portrait A11 is located in the intermediate feature is denoted as unmask. At this point, the expression: C*unMask*0.8+N*0.2->X*unMask can be used to replace the feature corresponding to the second region in the fourth redraw feature with the feature corresponding to the second region in the intermediate feature, thereby obtaining the updated fourth redraw feature.

[0226] S1102e: corresponding to j=k*(ab)+1, in the j-th iteration process, the updated j-th redrawing feature is input into the first diffusion model for redrawing, and the j+1-th redrawing feature is output.

[0227] In combination with the above example, corresponding to j=4, in the fourth iteration process, the updated fourth redrawing feature can be input into the first diffusion model for image redrawing, and the fifth redrawing feature can be output.

[0228] S1102f: corresponding to j>(k*ak*b), in the jth iteration process, the jth redrawing feature is input into the first diffusion model for redrawing, and the j+1th redrawing feature is output.

[0229] In the above example, corresponding to j=5, during the fifth iteration, the fifth redrawing feature can be input into the first diffusion model for image redrawing, and the sixth redrawing feature can be output. It can be understood that the feature corresponding to the second region of the first image in the sixth redrawing feature is obtained by introducing a preset noise according to the second redrawing parameter of 0.2 and iterating the first diffusion model twice, and the feature corresponding to the first region of the first image is obtained by introducing a preset noise according to the first redrawing parameter of 0.5 and iterating the first diffusion model five times.

[0230] As a result, the redrawing parameters for the first region of the sixth redraw feature, where the edge of the moved foreground object is located, are larger, while the redrawing parameters for the second region of the sixth redraw feature, where the interior of the moved foreground object and the background are located, are smaller. In this way, the first and second regions of the foreground object in the relit first image can be redrawn using different redrawing parameters.

[0231] S1102g: Corresponding to j=k*a, the j+1th redrawn feature is used as the second feature.

[0232] Combined with the above example, corresponding to j=5, after the fifth iteration, the sixth redrawing feature can be used as the second feature. For example, the second feature is the sixth redrawing feature, which means Figure 10 The redrawn image A4 corresponds to the features shown.

[0233] Diffusion realization model training

[0234] Furthermore, the training process of the diffusion realization model in this application is introduced.

[0235] In some embodiments, the present application uses the example of training a diffusion realization model based on language graph technology. In this case, the diffusion realization model may also include a text encoder. For example, the text encoder may be a text encoder in a contrastive language-image pre-training (CLIP) model.

[0236] In some embodiments, the present application can train a diffusion realization model based on at least one first training data to obtain a trained diffusion realization model; wherein the first training data includes a first training text and a first training image, wherein the first training text is related to the image features of the first training image, and the text encoder is used to obtain the text features of the first training text.

[0237] Reference Figure 12 The figure shows the training process of the diffusion realization model. Figure 12 The illustrated diffusion realisation model 100 also includes a text encoder 104 .

[0238] It is understood that the training dataset of the diffusion realization model 100 can be a training text-training image pair, such as a first training text and a corresponding first training image. For example, the first training text can be training text 1, and the corresponding first training image can be training image 1. Training text 1 is associated with image features of training image 1, and training image 1 can be a high-definition image, i.e., an image with a high resolution.

[0239] Specifically, if Figure 12 As shown, the training text 1 can be passed through the text encoder 104 to generate the training result 1 to train the diffusion model 102 or the dVAE decoder 103 in the diffusion realization model 100. In addition, the training image 1 can be passed through the dVAE encoder 101 to generate the result image 1 to train the dVAE decoder 103 in the diffusion realization model 100.

[0240] Optionally, the present application may first train the diffusion model 102 , and then train the dVAE decoder 103 based on the trained diffusion model 102 to complete the training of the entire diffusion realization model 100 .

[0241] Next, refer to Figure 13 , the training method process of the diffusion realization model provided in the embodiment of the present application is introduced, and the process includes the following steps:

[0242] S1301: Input a first training text in the first training data into a text encoder, and output text features of the first training text.

[0243] S1302: Input a first training image in the first training data into a first encoder, and encode the first training image into a third feature of the latent space.

[0244] S1303: Input the third feature into the first diffusion model, and process the third feature based on preset noise to obtain a fourth feature.

[0245] In some embodiments, the present application may redraw features corresponding to a third region in the third feature based on a preset noise according to a first redrawing parameter, and redraw features corresponding to a fourth region in the third feature based on a preset noise according to a second redrawing parameter. The third region includes the following: the outline of a foreground object in the first training image, and multiple pixels adjacent to pixels corresponding to the outline of the foreground object in the first training image; the fourth region is an area outside the third region in the first training image.

[0246] It can be understood that the third area and the fourth area in the first training image can refer to the relevant descriptions of the first area and the second area in the first image above, respectively, and are not repeated here.

[0247] In other embodiments, all features in the third feature may be redrawn according to a uniform redrawing parameter based on the preset noise to obtain a redrawn fourth feature.

[0248] S1304: Input the third feature into the first decoder, and decode the fourth feature into a first training result.

[0249] In addition, in some other embodiments, the first training result can also be generated based on the first training text. Specifically, the present application can also input the text features of the first training text into the first diffusion model and output the first text encoding features. Furthermore, the first text encoding features are input into the first decoder and the first training result is output.

[0250] S1305: Adjusting network parameters of the first diffusion model according to the difference between the text feature of the first training text and the third feature, and / or adjusting network parameters of the first decoder according to the difference between the first training image and the first training result.

[0251] It can be understood that the present application can set the loss function of the first diffusion model, and determine the value of the loss function based on the difference between the text features of the first training text and the third features, and adjust the network parameters of the first diffusion model through the value of the loss function.

[0252] Similarly, the present application can set the loss function of the first decoder, and determine the value of the loss function according to the difference between the first training image and the first training result, and adjust the network parameters of the first decoder according to the value of the loss function.

[0253] It can be understood that the value of the loss function is close to a stable small value, which means that the model has converged. At this time, a set of parameters that can minimize the loss function is found as the network parameters of the network model. Furthermore, this application can use different training images to input the diffusion reality model and repeat the process. Figure 13 The training process continues until convergence.

[0254] As examples, these network parameters can include weights and biases. Weights are the numerical values ​​connecting nodes in each layer of a neural network. They are acquired through training and determine how the network processes input data. Weights are randomly initialized before training begins and then continuously updated during training to minimize the loss function. Biases are another parameter used alongside the weights of each node in the network. They are added to the input of a neuron before the activation function to help adjust the probability of the neuron being activated.

[0255] Furthermore, the first diffusion model iteratively processes the third feature in the first training image multiple times to obtain the fourth feature. Figure 14 As shown, the above S1303 includes S1401 to S1406:

[0256] S1401: Determine a first number of iterations k*a according to a first redrawing parameter a and a preset parameter k, and determine a second number of iterations k*b according to a second redrawing parameter b and a preset parameter.

[0257] S1402: Superimposing the preset noise and the third feature according to the first redrawing parameter to obtain a first training redrawing feature.

[0258] S1403: corresponding to 1≤j≤k*(ab), in the j-th iteration process, the j-th training redrawing feature is input into the first diffusion model to perform image redrawing, and the j+1-th training redrawing feature is output.

[0259] S1404: Corresponding to j=k*(ab)+1, the preset noise and the third feature are superimposed according to the second redrawing parameter to obtain a training intermediate feature, and the feature corresponding to the second area in the j-th training redrawing feature is replaced with the feature corresponding to the second area in the training intermediate feature to obtain an updated j-th training redrawing feature.

[0260] S1405: corresponding to j=k*(ab)+1, in the j-th iteration process, the updated j-th training redrawing feature is input into the first diffusion model for image redrawing, and the j+1-th training redrawing feature is output.

[0261] S1406: corresponding to j>k*(ab), in the j-th iteration process, the j-th training redrawing feature is input into the first diffusion model to perform image redrawing, and the j+1-th training redrawing feature is output.

[0262] S1407: corresponding to j=k*a, the j+1th training redrawing feature is used as the fourth feature.

[0263] Similarly, the detailed description of S1401 to S1407 in this application can refer to the above description of Figure 11B The relevant descriptions of S1102a to S1102g in the above are not repeated here. As an example, the second redrawing parameter (such as 0.2) can be the redrawing parameter of the second area in the first training image, and the first redrawing parameter (such as 0.5) can be the redrawing parameter of the second area in the first training image. In addition, the present application can predict the third area in the first training image through algorithms such as foreground prediction to determine the features corresponding to the third area and the features corresponding to the fourth area in the third features of the first training image.

[0264] It can be understood that in the embodiment of the present application, multiple sets of first training texts and first training images can be used to train the diffusion realisation model, so that the diffusion realisation model has the ability to redraw in different regions.

[0265] Application of the diffusion enhancement model

[0266] Next, the diffusion enhancement model provided in the embodiments of the present application is described.

[0267] In some embodiments, the diffusion enhancement model in the present application may include multiple components, such as a second encoder, a second decoder, and a second diffusion model. For example, the second diffusion model may be a pre-trained first diffusion model or a separately trained diffusion model, which is not specifically limited in the present embodiments. In this case, the first diffusion model may be integrated into a separate component and inserted into the diffusion realization model or the diffusion enhancement model.

[0268] Optionally, the second encoder may be a dVAE encoder, and the second decoder may be a dVAE decoder.

[0269] Reference Figure 15 , describes the structure of the diffusion enhancement model provided in the embodiment of the application. Figure 15 As shown, the diffusion enhancement model 200 includes the following components: a dVAE encoder 201 , a diffusion model 202 , a dVAE decoder 203 , a control network 204 , a restoration network 205 , and an image encoder 206 .

[0270] The dVAE encoder 201 is used to input the block images of the redrawn image A4 and extract the original latent features of the block images.

[0271] In addition, the image encoder 206 is used to extract semantic features (i.e., semantic information) of the block image, such as low-level features (such as edges and textures) and high-level features (such as object categories and scene layout). For example, the image encoder 206 can be an image encoder in the CLIP model.

[0272] Control network 204 receives latent features from dVAE encoder 201 as input and adjusts them as needed, such as by introducing constraints. For example, these constraints could be semantic features extracted by image encoder 206. In this case, the semantic information introduced by control network 204 acts as a constraint to ensure that the image generated by diffusion model 202 conforms to these descriptions.

[0273] It can be understood that the control network 204 is used to support the diffusion model 202 to adjust the semantic features from the image encoder 206 and the latent features of the block images from the dVAE encoder 201, so that the adjusted latent features of the block images conform to the semantic features.

[0274] The diffusion model 202 is used to adjust the expanded latent features of the block images according to the semantic features of the block images to obtain the adjusted latent features of the block images.

[0275] The restoration network 205 is used to restore the image obtained by splicing the latent features after adjustment of the multiple block images, specifically to restore the splicing parts of the spliced ​​features, so that the content of the spliced ​​features is consistent with the content of the original redrawn image A4.

[0276] As an example, the repair network 205 may be implemented by a Unet model, but is not limited thereto.

[0277] The dVAE decoder 203 is used to decode the concatenated features to obtain an enhanced image A5.

[0278] Reference Figure 16 , the image super-resolution processing process based on the diffusion enhancement model in the implementation of this application is introduced, and the process includes the following steps:

[0279] S1601: Input the N first blocks of the second image into the second encoder respectively, encode the N first blocks in the latent space respectively, and obtain N fifth features.

[0280] Wherein, N is a positive integer.

[0281] In some implementations, the N first blocks are obtained by performing regular block division or overlap block division on the second image.

[0282] It can be understood that the regular block division method means that there is no overlapping portion between the blocks in the second image, while the overlapping block division method means that there is overlapping portion between the blocks in the second image.

[0283] Reference Figure 15 As shown, the second image is a redrawn image A4, and the redrawn image A4 is obtained by adopting a regular block method. At this time, there is no overlapping part between these blocks.

[0284] In addition, refer to Figure 17 As shown in FIG. 1 , a schematic diagram of overlapping block partitioning is shown, where, for example, block 171 overlaps with block 172. Thus, when the latent features of the blocks are subsequently spliced ​​together, the possibility of distortion caused by the fusion at the splicing point is low. For example, the seams of the clouds in the sky between the two spliced ​​blocks coincide with each other without distortion.

[0285] S1602: Pass the N first blocks through an image encoder respectively to obtain N semantic features.

[0286] It can be understood that the semantic feature of a first block is used to characterize the image content in the block. For example, the semantic feature of a first block is "night sky clouds", which means that the content of the block is the night scene with clouds in the sky.

[0287] S1603: input the N fifth features into the control network respectively, and output corresponding N sixth features, wherein the sixth feature takes the corresponding semantic feature as a constraint condition.

[0288] It can be understood that the sixth feature contains the fifth feature while taking the corresponding semantic feature as a constraint condition. Therefore, the subsequent second diffusion network supports adjusting the sixth feature to conform to the content indicated by the corresponding semantic feature through the control network.

[0289] S1604: input the N semantic features and the N sixth features into the second diffusion model, obtain corresponding seventh features according to each sixth feature and the corresponding semantic feature, and output N seventh features corresponding to the N first blocks.

[0290] It can be understood that in the embodiments of the present application, the second diffusion model can perform image enhancement on the seventh feature to increase the resolution of the image restored by the seventh feature.

[0291] In addition, the seventh feature is a feature that conforms to the semantic feature of the corresponding first block, so that the possibility of subsequent distortion of the content of the N first blocks is small.

[0292] S1605: combine the N seventh features into an eighth feature, and perform upsampling on the eighth feature in the latent space to obtain a ninth feature.

[0293] It can be understood that the N seventh features are features corresponding to the N first blocks, and at this time, the eighth feature after combination can correspond to the whole of the second image. Then, performing upsampling on the eighth feature in the latent space to obtain the ninth feature can further increase the image detail information in the ninth feature, which is conducive to reducing the possibility of subsequent distortion of the content of the N first blocks.

[0294] S1606: input the ninth feature into the repair network, process the data from different seventh features in the ninth feature, and obtain a tenth feature.

[0295] It can be understood that the repair network can repair the combination between the features from different blocks in the combined ninth feature to ensure that the content at the combination does not distort, thereby being conducive to reducing the possibility of subsequent distortion of the content of the N first blocks.

[0296] S1607: input the tenth feature into the second decoder, and output a fifth image.

[0297] It can be understood that the second decoder can decode the second feature, thereby decoding the tenth feature in the latent space into the fifth image in the image space to obtain an enhanced image after super-resolution processing (i.e. image enhancement) of the redrawn image.

[0298] Thus, the pre-trained diffusion enhancement model provided by the embodiments of the present application can enhance the redrawn image after relighting and redrawing. Specifically, each block is enhanced separately to achieve image super-resolution processing, thereby improving the overall resolution of the subsequently restored image. Furthermore, the features of each spliced ​​block can be repaired at the splicing point using a pre-trained inpainting network, ensuring that the splicing point of the spliced ​​and restored image is free of content distortion, thus ensuring the authenticity of the enhanced image.

[0299] Diffusion Enhanced Model Training

[0300] In some embodiments, the first decoder in the diffusion enhancement model and the repair network can be trained separately in the embodiments of the present application.

[0301] In some embodiments, the embodiments of the present application can train the second diffusion model based on at least one second training data to obtain a trained second diffusion model; wherein the second training data includes a first training block and a first training semantic feature, and the first training semantic feature is the semantic feature of the first training block.

[0302] It is understood that the training dataset of the diffusion realization model 100 can be a training semantic feature-training block pair, such as a first training block corresponding to a first training semantic feature. For example, the first training block can be training block 1, and the corresponding first training semantic feature is the first training semantic feature. The image content of training block 1 conforms to the semantics of training semantic feature 1, and training block 1 can be a high-definition image, that is, a block in an image with a high resolution.

[0303] It can be understood that the first training blocks mentioned above can be blocks obtained from the high-definition image by using a random blocking method, a regular blocking method or an overlapping blocking method.

[0304] Specifically, if Figure 18 As shown, the training block 1 can be passed through the text encoder 104 to generate the result block 1 to train the dVAE decoder 203 in the diffusion enhancement model 200.

[0305] Optionally, the present application may first train the diffusion model 102 , and then train the dVAE decoder 103 based on the trained diffusion model 102 to complete the training of the entire diffusion realization model 100 .

[0306] Reference Figure 19 As shown, the training process of the second decoder provided in the embodiment of the present application is described, and the process includes the following steps:

[0307] S1901: Input the first training block into the second encoder, encode multiple first training blocks in the latent space, and obtain the corresponding eleventh feature.

[0308] It can be understood that the eleventh feature is the original latent feature of the first training block in the latent space.

[0309] S1902: Input the eleventh feature into the control network and output the corresponding twelfth feature, wherein the twelfth feature uses the corresponding first training semantic feature as a constraint condition.

[0310] S1903: Input the first training semantic feature and the twelfth feature into the second diffusion model, and output the thirteenth feature corresponding to the first training block.

[0311] It is understandable that the image content indicated by the thirteenth feature is consistent with the first trained semantic feature. Specifically, in some embodiments, the second diffusion model can supplement the twelfth feature with image details or perform feature adjustments based on the first trained semantic feature, thereby improving the resolution of the segmented image while ensuring that the content of the segmented image is not distorted to a certain extent.

[0312] S1904: Input the thirteenth feature into the second decoder and output a second training result.

[0313] For example, the second training result can be Figure 18 The result block shown in FIG1 is block 1. At this time, the second training result may be an image in the image space.

[0314] S1905: Adjust network parameters of the second decoder according to the difference between the second training result and the first training block.

[0315] Similarly, the description of adjusting the network parameters in the second decoder in the present application can refer to the above description of the first decoder and the first diffusion model, which will not be repeated here.

[0316] In this way, when the difference between the second training result and the first training block is small, a trained second decoder is obtained to implement the training of the diffusion enhancement model, so that the diffusion enhancement model has the ability to improve the resolution of the block image.

[0317] Next, the training process of the repair network in the diffusion enhancement model is explained.

[0318] Reference Figure 20 , which is a schematic diagram of a repair network training process provided by an embodiment of the present application.

[0319] In some embodiments, the present application can train the restoration network based on at least one third training data to obtain a trained restoration network; wherein the third training data includes a second training image and a third training image, and the third training image is an image obtained by reducing the resolution of the second training image.

[0320] It can be understood that the second training image may be a high-definition image with a relatively high resolution, and the third training image may be an image obtained by degrading the second training image, that is, reducing its resolution.

[0321] Next, combine Figure 20 , taking the second training image as training image 2 and the third training image as training image 3 as an example, the training process of the repair network is described. Figure 20 As shown, each block in the training image 3 (such as training block 2) is passed through the encoder 201 to output the original latent features K1 of the training image 3. These are then subjected to resolution enhancement by the diffusion model 202, thereby obtaining the adjusted latent features K2 corresponding to each block in the training image 3. Furthermore, the adjusted latent features corresponding to each block in the training image 3 are concatenated and upsampled to obtain the upsampled latent features K3 of the training image 3. Furthermore, the upsampled latent features are passed through the restoration network 206 for restoration to obtain the restored latent features K4. The difference (represented by the loss) between the restored latent features K4 and the original latent features GT of the training image 2 output by the encoder 201 is used to train the restoration network 206.

[0322] In addition, in some other implementations, the encoder or diffusion model used in the training of the repair network 206 in this application can also be other networks, not limited to the above Figure 20 The network shown in FIG. 1 is not described in detail in the embodiment of the present application.

[0323] Next, refer to Figure 21 , the training process of the repair network provided in the embodiment of the present application is described in detail, and the process includes the following steps:

[0324] S2101: Inputting N second training blocks of the third training image into a second encoder respectively, encoding the N second training blocks in a latent space respectively, and obtaining N fourteenth features, where N is a positive integer.

[0325] In some embodiments, the N second training blocks are obtained by performing regular block division or overlapping block division on the third training image.

[0326] For example, a first block can be Figure 20 A block in the training image 3 is shown, and the corresponding fourteenth feature is the original latent feature of the block.

[0327] S2102: Inputting N fourteenth features into the second diffusion model, and outputting N fifteenth features corresponding to the N second training blocks, wherein the data volume of the fifteenth feature is greater than the data volume of the corresponding fourteenth feature.

[0328] It can be understood that the second diffusion model can improve the resolution of the fourteenth feature, specifically by adding image details in the latent space, thereby obtaining the corresponding fifteenth feature.

[0329] S2103: Combine N fifteenth features into a sixteenth feature, and upsample the sixteenth feature in the latent space to obtain a seventeenth feature.

[0330] S2104: Input the seventeenth feature into the repair network, process the data of the seventeenth feature that comes from different fifteenth features, and obtain the eighteenth feature.

[0331] S2105: Input the second training image into the second encoder, encode the second training image in the latent space, and obtain the nineteenth feature.

[0332] S2105: Adjust network parameters of the repair network according to the difference between the eighteenth feature and the nineteenth feature.

[0333] Similarly, the description of adjusting the network parameters in the repair network in the present application can refer to the above description of the first decoder and the first diffusion model, which will not be repeated here.

[0334] In this way, when the difference between the 18th and 19th features is small, a trained restoration network is obtained, which can be used to train the diffusion enhancement model. This allows the diffusion enhancement model to perform content-consistent restoration of the block images in the latent space, ensuring that the image after the blocks are stitched together can faithfully restore the image content. In this way, the diffusion enhancement model can fuse latent features in the latent space and eliminate the image distortion problem caused by the splicing of blocks, thereby achieving the ability to generate high-resolution images.

[0335] Next, taking a mobile phone as an example, the structure of the electronic device of the image processing method provided in the embodiment of the present application is described.

[0336] like Figure 22 As shown, the mobile phone 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, a button 101 and a display screen 102, etc.

[0337] It should be understood that the illustrated structure of the embodiment of the present invention does not constitute a specific limitation on the mobile phone 10. In other embodiments of the present application, the mobile phone 10 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0338] The processor 110 may include one or more processing units, for example, a processing module or processing circuit such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor, or a programmable logic device (FPGA). Different processing units may be independent devices or integrated into one or more processors. A storage unit may be provided in the processor 110 for storing instructions and data, such as the pre-trained diffusion realization model and diffusion enhancement model, as well as corresponding training data. In some embodiments, the storage unit in the processor 110 is a cache memory 180. For example, the processor 110 is used to redraw the edge area of ​​the foreground object of the re-lit image, and to perform super-resolution processing on the redrawn image.

[0339] The power module 140 may include a power supply, a power management component, and the like. The power supply may be a battery. The power management component manages the charging of the power supply and the supply of power to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects the power supply, the charging management module, and the processor 110. The power management module receives input from the power supply and / or the charging management module to power the processor 110, the display 102, the camera 170, and the wireless communication module 120.

[0340] The mobile communication module 130 may include but is not limited to an antenna, a power amplifier, a filter, a low noise amplifier (LNA), etc. The mobile communication module 130 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile phone 10. The mobile communication module 130 can receive electromagnetic waves through the antenna, filter, amplify and other processing on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 can be set in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 can be set in the same device as at least some modules of the processor 110. Wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wide band code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), Bluetooth (BT), global navigation satellite system (GNSS), wireless local area networks (WLAN), near field communication (NFC), frequency modulation (FM) and / or field communication (NFC), infrared technology (IR), etc.GNSS may include the global positioning system (GPS), the global navigation satellite system (GLONASS), the Beidou navigation satellite system (BDS), the quasi-zenith satellite system (QZSS) and / or the satellite based augmentation system (SBAS).

[0341] The wireless communication module 120 may include an antenna and transmit and receive electromagnetic waves via the antenna. The wireless communication module 120 may provide wireless communication solutions for the mobile phone 10, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), and the like. The mobile phone 10 may communicate with the network and other devices through wireless communication technologies.

[0342] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 may also be located in the same module.

[0343] The display screen 102 is used to display a human-computer interaction interface, images, videos, etc. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. For example, the display screen 102 provided in the present application can be a QLED screen and display each frame, such as a frame where the pixel after brightness compensation is located.

[0344] The sensor module 190 may include a proximity sensor, a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0345] The audio module 150 is used to convert digital audio information into analog audio signal output, or convert analog audio input into digital audio signal. The audio module 150 can also be used to encode and decode audio signals. In some embodiments, the audio module 150 can be provided in the processor 110, or some functional modules of the audio module 150 can be provided in the processor 110. In some embodiments, the audio module 150 can include a speaker, an earpiece, a microphone, and a headphone jack.

[0346] Camera 170 is used to capture still images or video. The lens generates an optical image of an object and projects it onto a photosensitive element. The photosensitive element converts the optical signal into an electrical signal, which is then passed to an image signal processor (ISP) for conversion into a digital image signal. Mobile phone 10 implements its camera function through the ISP, camera 170, video codec, graphics processing unit (GPU), display 102, and application processor.

[0347] The interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 via the external memory interface to implement data storage. The USB interface is used for communication between the mobile phone 10 and other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the mobile phone 1010, for example, to read the phone number stored in the SIM card or write the phone number to the SIM card.

[0348] In some embodiments, the mobile phone 10 further includes buttons 101, a motor, and an indicator. The buttons 101 may include a volume button, an on / off button, and the like. The motor is used to vibrate the mobile phone 10, for example, when a user's mobile phone 10 is called, to prompt the user to answer the call. The indicator may include a laser pointer, a radio frequency indicator, an LED indicator, and the like.

[0349] In some embodiments, the user interface may include, but is not limited to, a display (e.g., an LCD display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.

[0350] In some embodiments, the present application provides a readable medium having instructions stored thereon, which, when executed on an electronic device, causes the electronic device to execute the above method.

[0351] In some embodiments, the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the above method.

[0352] In some embodiments, the present application provides a computer program product, which includes instructions for implementing the above image processing method.

[0353] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0354] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0355] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0356] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0357] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0358] It should be noted that each unit / module mentioned in each device embodiment of the present application is a logical unit / module, and in physical, one logical unit / module can be a physical unit / module, or a part of a physical unit / module, or be realized in a combination of multiple physical unit / modules, and the physical realization of these logical units / modules is not the most important, and the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules which are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0359] It should be noted that in the examples and descriptions of the present patent, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including one" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0360] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it should be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the present application.

Claims

1. An image processing method, characterized in that: The method comprises: Acquire a first image, wherein a foreground object and a background in the first image are from different images; Redrawing a first area in the first image according to a first redrawing parameter, and redrawing a second area in the first image according to a second redrawing parameter to obtain a second image; The first redrawing parameter is different from the second redrawing parameter, and the first area includes the following contents: the outline of the foreground object, and a plurality of pixels adjacent to the pixels corresponding to the outline of the foreground object; the second area is an area outside the first area in the first image.

2. The method according to claim 1, characterized in that The multiple pixels adjacent to the pixel corresponding to the outline of the foreground object include: pixels within the foreground object whose shortest distance between pixels corresponding to the outline of the foreground object is less than a first distance, and pixels outside the foreground object whose shortest distance between pixels corresponding to the outline of the foreground object is less than a second distance.

3. The method according to claim 1, characterized in that The acquiring of the first image comprises: The foreground object in the third image is added to the fourth image, and the foreground object is re-lit to obtain the first image, wherein the lighting information of the foreground object in the first image is the same as the lighting information of the fourth image.

4. The method according to claim 1, wherein The method further comprises: Performing super-resolution processing on the second image to obtain a fifth image, wherein the resolution of the fifth image is higher than the resolution of the second image.

5. The method according to claim 3, characterized in that The step of adding the foreground object in the third image to the fourth image and relighting the foreground object to obtain the first image includes: determining a first mask image of the foreground object from the third image, wherein the first mask image is used to represent the outline and shape of the foreground object; performing normal prediction on the third image to obtain normal information of the third image, where the normal information is used to reflect the direction of the surface of the foreground object in the third image; performing albedo prediction on the third image to obtain albedo information of the third image, where the albedo information is used to reflect the color and material properties of the foreground object in the third image; adding the foreground object in the third image to the fourth image; The foreground object in the fourth image is re-lit according to the first mask image, the normal information, the albedo information, and the lighting information of the fourth image to obtain the third image.

6. The method according to claim 5, characterized in that Adding the foreground object in the third image to the fourth image comprises: Downsampling is performed on the third image, and the foreground object in the downsampled third image is added to the fourth image.

7. The method according to claim 1, characterized in that Redrawing the first area in the first image according to the first redrawing parameter, and redrawing the second area in the first image according to the second redrawing parameter, includes: The first image is input into a pre-trained first network model, and the first area in the first image is redrawn according to first redrawing parameters by the first network model, and the second area in the first image is redrawn according to second redrawing parameters, and the second image is output.

8. The method according to claim 7, characterized in that The first network model includes the following components: a first encoder, a first decoder, and a first diffusion model; The process of the first network model processing the first image into the second image includes: Inputting the first image into the first encoder, encoding the first image into a first feature in a latent space; Inputting the first feature into the first diffusion model, redrawing features in the first feature corresponding to the first region based on preset noise according to the first redrawing parameter, and redrawing features in the first feature corresponding to the second region based on the preset noise according to the second redrawing parameter, to obtain a second feature; The second feature is input into the first decoder, and the second feature is decoded into the second image.

9. The method according to claim 8, characterized in that The process of processing the first feature by the first diffusion model includes: Inputting the first feature into the first diffusion model, redrawing features corresponding to the first area in the first feature according to the first redrawing parameter based on preset noise, and redrawing features corresponding to the second area in the first feature according to the second redrawing parameter based on the preset noise, including: superimposing the preset noise and the first feature according to the first redrawing parameter to obtain a first redrawing feature; Corresponding to 1≤j≤k*(ab), in the jth iteration process, the jth redrawing feature is input into the first diffusion model for redrawing, and the j+1th redrawing feature is output, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is a preset parameter; Corresponding to j=k*(ab)+1, superimposing the preset noise with the first feature according to the second redrawing parameter to obtain an intermediate feature, and replacing the feature corresponding to the second area in the j-th redrawing feature with the feature corresponding to the second area in the intermediate feature, to obtain an updated j-th redrawing feature; Corresponding to j=k*(ab)+1, in the j-th iteration process, the updated j-th redrawing feature is input into the first diffusion model for redrawing, and the j+1-th redrawing feature is output; Corresponding to j>k*(ab), in the jth iteration process, the jth redrawing feature is input into the first diffusion model for redrawing, and the j+1th redrawing feature is output; Corresponding to j=k*a, the j+1th redrawn feature is used as the second feature.

10. The method according to claim 9, characterized in that The first network model further includes a text encoder; and the training process of the first network model includes: Training the first network model according to at least one first training data to obtain a trained first network model; The first training data includes a first training text and a first training image, wherein the first training text is related to image features of the first training image, and the text encoder is used to obtain text features of the first training text.

11. The method according to claim 10, characterized in that The training of the first network model according to at least one first training data includes: Training the first diffusion model in the first network model according to the at least one first training data to obtain a trained first diffusion model; Based on the trained first diffusion model, the first decoder in the first network model is trained according to the at least one first training data to obtain the trained first decoder.

12. The method according to claim 10 or 11, characterized in that The process of processing the first training data by the first network model includes: Inputting the first training text in the first training data into the text encoder, and outputting text features of the first training text; Inputting the first training image in the first training data into the first encoder, and encoding the first image into a third feature in the latent space; Inputting the third feature into the first diffusion model, redrawing features corresponding to a third region in the third feature based on the preset noise according to the first redrawing parameter, and redrawing features corresponding to a fourth region in the third feature based on the preset noise according to the second redrawing parameter, to obtain a fourth feature, wherein the third region includes the following: an outline of a foreground object in the first training image, and a plurality of pixels adjacent to pixels corresponding to the outline of the foreground object in the first training image; and the fourth region is an area outside the third region in the first training image; Inputting the fourth feature into the first decoder, and decoding the fourth feature into a first training result; The network parameters of the first diffusion model are adjusted according to the difference between the text feature of the first training text and the fourth feature, and / or the network parameters of the first decoder are adjusted according to the difference between the first training image and the first training result.

13. The method according to claim 12, characterized in that The process of processing the third feature by the first diffusion model includes: superimposing the preset noise and the third feature according to the first redrawing parameter to obtain a first training redrawing feature; Corresponding to 1≤j≤k*(ab), in the jth iteration process, the jth training redrawing feature is input into the first diffusion model for redrawing, and the j+1th training redrawing feature is output, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is a preset parameter; Corresponding to j=k*(ab)+1, superimposing the preset noise and the third feature according to the second redrawing parameter to obtain a training intermediate feature, and replacing the feature corresponding to the fourth area in the j-th training redrawing feature with the feature corresponding to the fourth area in the training intermediate feature, to obtain an updated j-th training redrawing feature; Corresponding to j=k*(ab)+1, in the j-th iteration process, the updated j-th training redrawing feature is input into the first diffusion model for redrawing, and the j+1-th training redrawing feature is output; Corresponding to j>k*(ab), in the jth iteration process, the jth training redrawing feature is input into the first diffusion model for redrawing, and the j+1th training redrawing feature is output; Corresponding to j=k*a, the j+1th training redraw feature is used as the fourth feature.

14. The method according to claim 4, characterized in that The performing super-resolution processing on the second image to obtain a fifth image includes: The second image is input into a pre-trained second network model, and the fifth image is output.

15. The method according to claim 14, characterized in that The second network model includes the following components: a second encoder, a second diffusion model, an image encoder, a control network, a repair network, and a second decoder; and The process of processing the second image by the second network model includes: Inputting N first blocks of the second image into the second encoder respectively, encoding the N first blocks in the latent space respectively, to obtain N fifth features, where N is a positive integer; Passing the N first blocks through the image encoder respectively to obtain N semantic features; Inputting the N fifth features into the control network respectively, and outputting corresponding N sixth features, wherein the sixth features use the corresponding semantic features as constraints; Inputting the N semantic features and the N sixth features into the second diffusion model, obtaining a corresponding seventh feature according to each of the sixth features and the corresponding semantic feature, and outputting N seventh features corresponding to the N first blocks; Combining the N seventh features into an eighth feature, and upsampling the eighth feature in the latent space to obtain a ninth feature; Inputting the ninth feature into the repair network, processing data of the ninth feature that is different from the seventh feature, to obtain a tenth feature; The tenth feature is input into the second decoder, and the fifth image is output.

16. The method according to claim 15, characterized in that The method further comprises: Training the second decoder according to at least one second training data to obtain a trained second decoder; The second training data includes a first training block and a first training semantic feature, and the first training semantic feature is a semantic feature of the first training block.

17. The method according to claim 16, characterized in that The training process of the second decoder based on the second training data includes: Inputting the first training blocks into the second encoder, encoding the multiple first training blocks in the latent space to obtain corresponding eleventh features; Inputting the eleventh feature into the control network and outputting a corresponding twelfth feature, wherein the twelfth feature uses the corresponding training semantic feature as a constraint condition; Inputting the first training semantic feature and the twelfth feature into the second diffusion model, and outputting a thirteenth feature corresponding to the first training block; Inputting the thirteenth feature into the second decoder and outputting a second training result; Adjust network parameters of the second decoder according to the second training result and the first training block.

18. The method according to claim 17, characterized in that The method further comprises: Training the repair network according to at least one third training data to obtain the trained repair network; The third training data includes a second training image and a third training image, and the third training image is an image obtained by reducing the resolution of the second training image.

19. The method according to claim 18, characterized in that The process of training the repair network based on the third training data includes: Inputting N second training blocks of the third training image into the second encoder respectively, and encoding the N second training blocks in the latent space respectively to obtain N fourteenth features, where N is a positive integer; Inputting the N fourteenth features into the second diffusion model, and outputting N fifteenth features corresponding to the N second training blocks, wherein the data volume of the fifteenth features is greater than the data volume of the corresponding fourteenth features; Combining the N fifteenth features into a sixteenth feature, and upsampling the sixteenth feature in the latent space to obtain a seventeenth feature; Inputting the seventeenth feature into the repair network, processing data of the seventeenth feature that is different from the fifteenth feature, to obtain an eighteenth feature; Inputting the second training image into the second encoder, encoding the second training image in the latent space to obtain a nineteenth feature; According to the difference between the eighteenth feature and the nineteenth feature, the network parameters of the repair network are adjusted.

20. The method according to claim 19, characterized in that The N first blocks and the N second training blocks are obtained by dividing the blocks in a regular block manner or an overlapping block manner.

21. The method according to claim 3, characterized in that The fourth image is an image generated based on text information input by the user, or a panoramic high dynamic range HDR image, or a non-HDR image.

22. The method according to claim 3, characterized in that The illumination information of the fourth image is pre-set, pre-collected, or calculated based on the content of the fourth image.

23. The method according to claim 22, characterized in that Corresponding to the fourth image being an HDR image, the lighting information of the fourth image is pre-collected lighting information of the ambient light in the shooting scene where the fourth image is located.

24. The method according to claim 21, wherein The third image includes: a frame of image in a dynamic image, a frame of static image, or a frame of image in a video, or, The fourth image includes: a frame of image in a dynamic image, a frame of static image, or a frame of image in a video.

25. The method according to claim 3, wherein The third image includes: an image in the preview interface of the photo and video modes, an image obtained by a photo operation, an image in a video obtained by video shooting, an image in a video conference, an image in a video call, or an image in a desktop wallpaper.

26. The method according to claim 25, characterized in that There are multiple foreground objects in the third image.

27. The method according to claim 3, wherein The method further comprises: displaying a first interface, wherein the first interface includes a first control, and the first control is used to select at least one of the third image and the fourth image; A user operation on the first control is detected, and the third image and the fourth image are acquired based on the operation.

28. A readable medium, characterized in that The readable medium stores instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 27.

29. An electronic device, characterized in that: include: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, which is one of the processors of the electronic device, for executing the method according to any one of claims 1 to 27.

30. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to implement the method according to any one of claims 1 to 27.

Citation Information

Patent Citations

  • Image processing method and electronic equipment

    CN115908120A

  • Image rendering method and device, equipment and storage medium

    CN116957921A

  • Image processing method and device, computer, storage medium and program product

    CN117252947A

  • Image generation method and device, equipment and medium

    CN117456033A

  • Model training method, image processing method and electronic equipment

    CN117649478A