Image processing method, medium, electronic device, and program product
By redrawing the outline of the foreground object and its neighboring and background regions using differentiated image redrawing, combined with super-resolution processing, the problem of inharmonious lighting between the foreground object and the background in the relighting technique is solved, achieving a more natural and realistic image fusion effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-03-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing relighting techniques result in inconsistent lighting between the foreground object and the background when changing the background of a portrait. This leads to a disconnect between the foreground object and the background in the composite image, with blurred edges or remnants of the original background, resulting in a lack of realism.
By acquiring the foreground and background of different images, a pre-trained neural network model is used to perform differential image redrawing of the outline of the foreground object and its neighboring and background regions. Combined with super-resolution processing, the consistency of lighting and edge sharpness of the foreground object and background are ensured.
It improves the lighting consistency and realism between foreground objects and background in the synthesized image, eliminates edge blurring, and ensures the naturalness and realism of image fusion.
Smart Images

Figure CN120765472B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method, medium, electronic device and program product. Background Technology
[0002] With the development of image processing technology, users have increasingly higher demands for personalization in images or videos captured by electronic devices. For example, in certain portrait image shooting scenarios, users need to change the background environment of the portrait, such as changing the background of the portrait from indoors to outdoors, or changing the background of the portrait from daytime to nighttime, etc.
[0003] However, when changing the background of a portrait, a problem arises where the lighting of the portrait itself is not harmonious with the lighting of the new background. To improve this, it is usually necessary to relight the portrait in the image after the background has been changed. Relighting refers to redetermining and attaching lighting information to the foreground object so that it matches the lighting of the new background image, such as the ambient image.
[0004] Traditional relighting techniques involve predicting the foreground of the original portrait image. However, the outline of the foreground object extracted by this foreground prediction algorithm usually has errors, which leads to problems such as blurred edges of the foreground object in the synthesized image or residual original background. This results in a disjointed foreground object and background image in the synthesized image, meaning that the foreground object and background in the synthesized image after relighting have an unrealistic feel. Summary of the Invention
[0005] This application provides an image processing method, a medium, an electronic device, and a program product.
[0006] In a first aspect, embodiments of this application provide an image processing method, the method comprising: acquiring a first image, wherein a foreground object and a background in the first image are from different images; redrawing a first region in the first image according to first redrawing parameters, and redrawing a second region in the first image according to second redrawing parameters, to obtain a second image; wherein the first redrawing parameters are different from the second redrawing parameters, and the first region includes the following: the outline of the foreground object, and a plurality of pixels adjacent to the pixels corresponding to the outline of the foreground object; the second region is the region outside the first region in the first image. For example, the first redrawing parameter is greater than the second redrawing parameter.
[0007] Thus, in this image processing method, for the composite image after background modification, the edge region between the foreground object and the background in the composite image is redrawn, for example, through style transfer, image enhancement, color correction, texture creation, or shape modification. This edge region (i.e., the first region) includes: the outline of the foreground object, the content inside the foreground object adjacent to the outline, and the content outside the foreground object adjacent to the outline. At this point, the edge region is the area where the foreground object and background are blended in the composite image after relighting. Therefore, redrawing this edge region is equivalent to repairing this blended area. For example, in this application, for areas where the edges of the foreground object are broken or discontinuous in the edge region of the composite image after background replacement, image redrawing can generate new pixels to fill the gaps, and the new pixels are set to have the same or similar color as adjacent pixels, thereby eliminating edge blurring and improving the consistency of lighting between the foreground object and the background in the composite image. Simultaneously, this method redraws all regions of the synthesized image except for the edge regions, including the foreground within the edge regions and the background outside the edge regions. Furthermore, the redrawing parameters for the edge regions are greater than those for the other regions. It can be understood that the redrawing parameters for an image region refer to the magnitude of the pixel value change or the degree of style transfer. For example, the pixel value change in the edge regions is greater than that in the other regions, or the style transfer in the edge regions is greater than that in the other regions. Thus, while performing edge restoration on the foreground object's contour in the synthesized image, it does not cause significant changes to the foreground object itself or the background, such as preventing significant distortion of the replaced foreground portrait. Therefore, it ensures a natural and realistic image fusion.
[0008] In one possible implementation of the first aspect described above, the plurality of pixels adjacent to the pixel corresponding to the contour of the foreground object includes: pixels within the foreground object whose shortest distance to the pixel corresponding to the contour of the foreground object is less than a first distance, and pixels outside the foreground object whose shortest distance to the pixel corresponding to the contour of the foreground object is less than a second distance. Thus, the first region includes the contour of the moved foreground object, the content inside the foreground object adjacent to the contour, and the content outside the foreground object adjacent to the contour. For example, the first region in the first image includes: the contour of the foreground object, a region extending the contour of the foreground object inward by 10 pixels, and a region extending the contour outward by 10 pixels.
[0009] In one possible implementation of the first aspect above, obtaining the first image includes: adding a foreground object from a third image whose background is to be replaced to a fourth image that provides the background, and relighting the moved foreground object to obtain the first image, wherein the lighting information of the foreground object in the first image is the same as the lighting information of the fourth image.
[0010] In one possible implementation of the first aspect described above, the method further includes: performing super-resolution processing on the second image to obtain a fifth image, wherein the resolution of the fifth image is higher than that of the second image. This results in the fifth image having a higher image quality than the redrawn second image.
[0011] In one possible implementation of the first aspect described above, adding a foreground object from the third image to the fourth image and relighting the foreground object to obtain the first image includes: determining a first mask image of the foreground object from the third image, the first mask image representing the outline and shape of the foreground object; performing normal prediction on the third image to obtain normal information of the third image, the normal information reflecting the orientation of the surface of the foreground object in the third image; performing albedo prediction on the third image to obtain albedo information of the third image, the albedo information reflecting the color and material properties of the foreground object in the third image; adding the foreground object from the third image to the fourth image; and relighting the foreground object in the fourth image based on the first mask image, normal information, albedo information, and lighting information of the fourth image to obtain the third image. It is understood that algorithms such as foreground prediction, normal prediction, and albedo prediction have shifting errors, leading to problems such as lighting disharmony between the foreground object and the background in the relighted first image. At this point, this application uses a pre-trained model to redraw and super-resolution the first image after relighting, so as to ensure that the lighting between the foreground object and the background in the final fused image is harmonious, the edges are clear, and the realism is high.
[0012] In one possible implementation of the first aspect described above, adding a foreground object from the third image to the fourth image includes: downsampling the third image and adding the foreground object from the downsampled third image to the fourth image. It is understood that this processing can be performed on the downsampled third image to reduce the computational load during foreground prediction, normal prediction, albedo prediction, and relighting.
[0013] In one possible implementation of the first aspect described above, redrawing the first region in the first image according to the first redrawing parameters and redrawing the second region in the first image according to the second redrawing parameters includes: inputting the first image into a pre-trained first network model, redrawing the first region in the first image according to the first redrawing parameters using the first network model, redrawing the second region in the first image according to the second redrawing parameters, and outputting the second image. It can be understood that the pre-trained first network model has image redrawing capabilities, specifically the ability to redraw the edge regions and other regions of the foreground object in the image according to different redrawing parameters.
[0014] In one possible implementation of the first aspect described above, the first network model includes the following components: a first encoder, a first decoder, and a first diffusion model. The process by which the first network model processes the first image into a second image includes: inputting the first image into the first encoder to encode the first image into a first feature in the latent space; inputting the first feature into the first diffusion model, redrawing the feature corresponding to the first region in the first feature according to a first redrawing parameter based on preset noise, and redrawing the feature corresponding to the first region in the first feature according to a second redrawing parameter based on preset noise to obtain a second feature; and inputting the second feature into the first decoder to decode the second feature into a second image. Wherein, the first feature and the second feature can be latent features, and the preset noise can be Gaussian noise. In this way, image redrawing of the foreground object and background in the synthesized image after background replacement can be achieved in the latent space, thereby improving the lighting effect and realism of the fusion.
[0015] In one possible implementation of the first aspect above, the processing flow of the first diffusion model for the first feature includes: superimposing a preset noise with the first feature according to the first redrawing parameters to obtain the first redrawing feature; corresponding to 1≤j≤k*(ab), in the j-th iteration, the j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is the preset parameter; corresponding to j=k*(ab)+1, after superimposing the preset noise with the first feature according to the second redrawing parameters to obtain the intermediate feature, the feature corresponding to the second region in the j-th redrawing feature is replaced with the feature corresponding to the second region in the intermediate feature to obtain the updated j-th redrawing feature; corresponding to j= k*(ab)+1, in the j-th iteration, the updated j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output; corresponding to j>k*(ab), in the j-th iteration, the j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output; corresponding to j=k*a, the (j+1)-th redrawing feature is used as the second feature. In this way, within the latent space, the redrawing parameters for the features of the first region containing the foreground object in the first image obtained by changing the background are larger and the number of redrawings is greater, while the redrawing parameters for the features of other second regions in the first image are smaller and the number of redrawings is less.
[0016] In one possible implementation of the first aspect above, the first network model further includes a text encoder; the training process of the first network model includes: training the first network model according to at least one first training data to obtain a trained first network model; wherein the first training data includes first training text and first training image, wherein the first training text is related to the image features of the first training image, and the text encoder is used to obtain the text features of the first training text.
[0017] In one possible implementation of the first aspect above, training the first network model based on at least one first training data includes: training a first diffusion model in the first network model based on at least one first training data to obtain a trained first diffusion model; and training a first decoder in the first network model based on the trained first diffusion model and at least one first training data to obtain a trained first decoder.
[0018] In one possible implementation of the first aspect described above, the processing flow of the first network model for the first training data includes: inputting the first training text in the first training data into a text encoder and outputting the text features of the first training text; inputting the first training image in the first training data into a first encoder and encoding the first image into a third feature in the latent space; inputting the third feature into a first diffusion model, redrawing the feature corresponding to the third region in the third feature according to a first redrawing parameter based on preset noise, and redrawing the feature corresponding to the fourth region in the third feature according to a second redrawing parameter based on preset noise to obtain a fourth feature, wherein the third region includes the following: the outline of the foreground object in the first training image, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object in the first training image; the fourth region is the region outside the third region in the first training image; inputting the third feature into a first decoder and decoding the fourth feature into a first training result; adjusting the network parameters of the first diffusion model according to the difference between the text features of the first training text and the third feature, and / or adjusting the network parameters of the first decoder according to the difference between the first training image and the first training result.
[0019] In one possible implementation of the first aspect above, the processing flow of the third feature by the first diffusion model includes: superimposing the preset noise with the third feature according to the first redrawing parameters to obtain the first training redrawing feature; corresponding to 1≤j≤k*(ab), in the j-th iteration, the j-th training redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th training redrawing feature is output, where j is a positive integer, a is the first redrawing parameter, b is the second redrawing parameter, and k is the preset parameter; corresponding to j=k*(ab)+1, the preset noise is superimposed with the third feature according to the second redrawing parameters to obtain the training intermediate feature. The j-th training redraw feature corresponding to the fourth region is replaced with the corresponding feature from the intermediate training features, resulting in the updated j-th training redraw feature. Corresponding to j = k*(ab) + 1, in the j-th iteration, the updated j-th training redraw feature is input into the first diffusion model for redrawing, outputting the (j+1)-th training redraw feature. Corresponding to j > k*(ab), in the j-th iteration, the j-th training redraw feature is input into the first diffusion model for redrawing, outputting the (j+1)-th training redraw feature. Corresponding to j = k*a, the (j+1)-th training redraw feature is used as the fourth feature. In this way, the first diffusion model can redraw the image according to different redraw parameters for the foreground object area and other areas in the image.
[0020] In one possible implementation of the first aspect described above, performing super-resolution processing on the second image to obtain a fifth image includes: inputting the second image into a pre-trained second network model and outputting the fifth image. It is understood that the pre-trained second network model has the capability to perform super-resolution processing on images.
[0021] In one possible implementation of the first aspect described above, the second network model includes the following components: a second encoder, a second diffusion model, an image encoder, a control network, a repair network, and a second decoder; and the processing flow of the second network model for the second image includes: inputting N first blocks of the second image into the second encoder, encoding the N first blocks in the latent space to obtain N fifth features, where N is a positive integer; passing the N first blocks through the image encoder to obtain N semantic features; inputting the N fifth features into the control network to output corresponding N sixth features, wherein the sixth features are constrained by the corresponding semantic features; inputting the N semantic features and N sixth features into the second diffusion model, obtaining corresponding seventh features based on each sixth feature and its corresponding semantic features, and outputting N seventh features corresponding to the N first blocks; concatenating the N seventh features into an eighth feature, and upsampling the eighth feature in the latent space to obtain a ninth feature; inputting the ninth feature into the repair network, processing the data from different seventh features in the ninth feature to obtain a tenth feature; and inputting the tenth feature into the second decoder to output the fifth image. Thus, this application can perform super-resolution processing on image features in the latent space based on image segmentation technology and image semantic information.
[0022] In one possible implementation of the first aspect described above, the method further includes: training the second decoder based on at least one second training data to obtain a trained second decoder; wherein the second training data includes a first training block and first training semantic features, and the first training semantic features are the semantic features of the first training block. That is, this application can train the second decoder based on semantic feature-image block pairs as training data.
[0023] In one possible implementation of the first aspect described above, the training process for the second decoder based on the second training data includes: inputting a first training block into a second encoder, encoding multiple first training blocks in the latent space to obtain a corresponding eleventh feature; inputting the eleventh feature into a control network to output a corresponding twelfth feature, wherein the twelfth feature uses the corresponding training semantic feature as a constraint; inputting the first training semantic feature and the twelfth feature into a second diffusion model and outputting a thirteenth feature corresponding to the first training block; inputting the thirteenth feature into the second decoder to output a second training result; and adjusting the network parameters of the second decoder according to the second training result and the first training block. Thus, based on the configured control network, it is possible to accurately reconstruct the features of the blocks in the latent space to image data in the image space using semantic features as constraints.
[0024] In one possible implementation of the first aspect above, the method further includes: training the repair network based on at least one third training data to obtain a trained repair network; wherein the third training data includes a second training image and a third training image, and the third training image is an image obtained by reducing the resolution of the second training image.
[0025] In one possible implementation of the first aspect described above, the process of training the repair network based on the third training data includes: inputting N second training blocks of the third training image into the second encoder, encoding the N second training blocks in the latent space to obtain N fourteenth features, where N is a positive integer; inputting the N fourteenth features into the second diffusion model and outputting N fifteenth features corresponding to the N second training blocks, wherein the data volume of the fifteenth features is greater than that of the corresponding fourteenth features; concatenating the N fifteenth features into a sixteenth feature, and upsampling the sixteenth feature in the latent space to obtain a seventeenth feature; inputting the seventeenth feature into the repair network, processing the data from different fifteenth features in the seventeenth feature to obtain an eighteenth feature; inputting the second training image into the second encoder, encoding the second training image in the latent space to obtain a nineteenth feature; and adjusting the network parameters of the repair network according to the difference between the eighteenth and nineteenth features.
[0026] In one possible implementation of the first aspect described above, the N first blocks and N second training blocks are obtained by using a regular block division method or an overlapping block division method. For example, the overlapping block division method can be an overlap block division method, in which case the splicing effect of each block is better, and it is less likely to cause distortion at the joint after splicing.
[0027] In one possible implementation of the first aspect described above, the fourth image is an image generated based on text information input by the user, a panoramic high dynamic range (HDR) image, or a non-HDR image. It can be understood that when the fourth image is an HDR image, the corresponding illumination information can be pre-acquired. When the fourth image is a non-HDR image, the corresponding illumination information can be estimated using an illumination estimation algorithm.
[0028] In one possible implementation of the first aspect described above, the illumination information of the fourth image is obtained by pre-setting, pre-collection, or calculation based on the content of the fourth image (i.e., estimation using an illumination estimation algorithm).
[0029] In one possible implementation of the first aspect described above, the fourth image is an HDR image, and the lighting information of the fourth image is the ambient light lighting information of the shooting scene where the fourth image is located, which was acquired beforehand. In this case, the image processing method provided by this application can be applied to shooting scenes in professional studios.
[0030] In one possible implementation of the first aspect described above, the third image includes: a frame of an animated image, a still image, or a frame of a video; or, the fourth image includes: a frame of an animated image, a still image, or a frame of a video. Furthermore, the third and fourth images can be user-defined, and the image format can also be user-defined.
[0031] In one possible implementation of the first aspect described above, the third image includes: an image in the preview interface of the photo and video recording modes, an image obtained by taking a photo, an image in a video captured by video recording, an image in a video conference, an image in a video call, or an image in the desktop wallpaper. Of course, the application scenarios of the image processing method of this application include, but are not limited to, the examples described above.
[0032] In one possible implementation of the first aspect described above, the number of foreground objects in the third image is multiple. Therefore, the foreground objects moved into the fourth image can be one or more.
[0033] In one possible implementation of the first aspect described above, the method further includes: displaying a first interface, the first interface including a first control, the first control being used to select at least one of a third image and a fourth image; detecting a user's operation on the first control, and acquiring the third image and the fourth image based on the operation. For example, the first control may be as described below. Figure 3 In the shown photo-taking scene, the background control 13, or the first control, can be... Figure 6 The image control 64 shown is added. Figure 7 The background control 65 is shown.
[0034] Secondly, embodiments of this application provide a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.
[0035] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for performing the methods in the first aspect and any possible implementation thereof.
[0036] Fourthly, embodiments of this application provide a computer program product that, when run on an electronic device, enables the electronic device to implement the methods described in the first aspect and any possible implementation thereof.
[0037] The description of the beneficial effects of the second to fourth aspects in this application can be referred to the relevant description of the first aspect above, and will not be repeated here. Attached Figure Description
[0038] Figure 1A According to some embodiments of this application, a schematic diagram of an image change process in an image background replacement process is shown;
[0039] Figure 1B According to some embodiments of this application, a schematic diagram of an image background replacement process is shown;
[0040] Figure 2A According to some embodiments of this application, a schematic diagram of an image change process in an image background replacement process is shown;
[0041] Figure 2B According to some embodiments of this application, a schematic diagram of an image change process in an image background replacement process is shown;
[0042] Figure 2C According to some embodiments of this application, a schematic diagram of an image background replacement process is shown;
[0043] Figure 3 According to some embodiments of this application, a schematic diagram of a photographing scene is shown;
[0044] Figure 4 According to some embodiments of this application, a schematic diagram of a text-based photo-taking scene is shown;
[0045] Figure 5 According to some embodiments of this application, a schematic diagram of a voice-based photo-taking scenario is shown;
[0046] Figure 6According to some embodiments of this application, a schematic diagram of a custom wallpaper scene is shown;
[0047] Figure 7 According to some embodiments of this application, a schematic diagram of a custom wallpaper scene is shown;
[0048] Figure 8 According to some embodiments of this application, a schematic flowchart of an image processing method is shown;
[0049] Figure 9 According to some embodiments of this application, a schematic diagram of a custom wallpaper scene is shown;
[0050] Figure 10 According to some embodiments of this application, a schematic diagram of the structure of a diffusion realization model is shown;
[0051] Figure 11A According to some embodiments of this application, a schematic diagram of an image redrawing process is shown;
[0052] Figure 11B According to some embodiments of this application, a schematic flowchart of feature processing in image redrawing is shown;
[0053] Figure 12 According to some embodiments of this application, a schematic diagram of the training process of a diffusion realization model is shown;
[0054] Figure 13 According to some embodiments of this application, a flowchart of a training method for a diffusion realization model is shown;
[0055] Figure 14 According to some embodiments of this application, a flowchart of a diffusion model training method is shown;
[0056] Figure 15 According to some embodiments of this application, a schematic diagram of the training process of a diffusion enhancement model is shown;
[0057] Figure 16 According to some embodiments of this application, a schematic flowchart of image super-resolution processing is shown;
[0058] Figure 17 According to some embodiments of this application, a flowchart of a training method for a diffusion enhancement model is shown;
[0059] Figure 18 According to some embodiments of this application, a schematic diagram of the training process of a decoder in a diffusion enhancement model is shown;
[0060] Figure 19According to some embodiments of this application, a flowchart illustrating a training method for a decoder in a diffusion enhancement model is shown.
[0061] Figure 20 According to some embodiments of this application, a schematic diagram of the training process of a repair network in a diffusion enhancement model is shown;
[0062] Figure 21 According to some embodiments of this application, a flowchart illustrating a training method for a repair network in a diffusion enhancement model is shown.
[0063] Figure 22 According to some embodiments of this application, a schematic diagram of the structure of a mobile phone is shown. Detailed Implementation
[0064] The illustrative embodiments of this application include, but are not limited to, image processing methods, media, electronic devices, and program products.
[0065] As can be seen from the background technology, users may need to change the background of images, such as portrait images, to change the background environment of foreground objects such as portraits in the image.
[0066] Combination Figure 1A The image changes during the background replacement process shown, and Figure 1B The example shown illustrates the process of changing the background of an image, thus explaining an image processing scenario.
[0067] like Figure 1A As shown, the background of the foreground image A1 to be processed is an outdoor background, and the portrait A11 is located in front of a large tree. When the user wants to change the background of the portrait A11 in image A1, for example, to the indoor background in background image A2, the portrait A11 can be moved to background image A2 to construct a new image A3.
[0068] like Figure 1B As shown, the image processing scenario includes a foreground processing flow and a background processing flow. Specifically, the foreground processing flow is used to process image A1 whose background needs to be replaced, specifically processing the portrait A11 in image A1. The background processing flow is used to process image A2, which provides the background.
[0069] Specifically, such as Figure 1B The image processing flow shown includes the following steps:
[0070] S1: Input foreground image A1 and background image A2.
[0071] For example, the foreground image A1 can be a portrait image captured in real time by an electronic device or a pre-generated portrait image.
[0072] S2: Downsample the foreground image A1 to obtain the foreground image A1'.
[0073] For example, the size of the foreground image A1 is 3K×4K pixels, and the size of the downsampled foreground image A1' is 768×1K pixels. Of course, the sizes of the foreground image A1 and the downsampled foreground image A1' are not limited to the above example, and can be other sizes. This application does not specifically limit them.
[0074] S3: Perform foreground prediction on the foreground image A1' to obtain the mask image A1-1 of the foreground portrait A11.
[0075] In some embodiments, when the foreground image A1' is a portrait image, the foreground prediction described above can be portrait segmentation, used to determine a portrait from the foreground image A1', such as the foreground portrait A11.
[0076] As an example, the above portrait segmentation usually involves semantic segmentation, which classifies each pixel in the image into a corresponding category (such as people, sky, buildings, etc.).
[0077] It is understandable that the mask image A1-1 mentioned above is used to determine the outline and shape of the portrait A11.
[0078] S4: Perform normal prediction on the foreground image A1' to obtain the normal image A1-2.
[0079] In some embodiments, the prediction of the normals of the foreground image A1' can specifically be the prediction of the normals of the foreground portrait A11 in the foreground image A1'.
[0080] As can be understood, normal prediction is used to estimate the normal vector, or normal information, for each pixel in an image. Specifically, normal information represents the degree and direction of the surface's tilt at each point, and is crucial for understanding the surface geometry of an object.
[0081] In some embodiments, the normal image A1-2 represents the normal information of the foreground image A1'. Specifically, the normal image A1-2 can be a normal map with the same resolution as the foreground image A1'. In the normal map, the color value of each pixel represents the surface normal vector component at the corresponding image pixel location. Typically, the red, green, and blue channels correspond to the x, y, and z components of the normal vector, respectively.
[0082] S5: Perform albedo prediction on the foreground image A1' to obtain the albedo image A1-3.
[0083] In some embodiments, the albedo prediction of the foreground image A1' can specifically be the albedo prediction of the foreground portrait A11 in the foreground image A1'.
[0084] Albedo prediction is used to predict the color and material properties of a portrait, including diffuse reflection and specular reflection. This information is used to simulate material responses under different lighting conditions.
[0085] In some embodiments, albedo images A1-3 represent albedo information of a foreground image A1', specifically, each pixel value in albedo images A1-3 represents the albedo information of the corresponding location in the image.
[0086] Furthermore, in some embodiments, the aforementioned algorithms such as foreground prediction, normal prediction, and albedo prediction can be implemented using deep learning algorithms based on a pre-defined deep model.
[0087] S6: Perform image illumination estimation on background image A2 to obtain the illumination information of background image A2.
[0088] The lighting information of the background image A2 is used to represent the characteristics of the lighting in the image, such as the direction, color, intensity and distribution of the lighting.
[0089] S7: Move the foreground portrait A11 into the background image A2 according to the mask image A1-1, and perform physical relighting on the foreground portrait A11 in the background image A2 according to the normal image A1-2 and the albedo image A1-3 to generate the composite image A3.
[0090] Optionally, the synthesized image A3 can be upsampled to obtain an upsampled synthesized image A3.
[0091] Physical relighting is a technique involving computer graphics and visual effects that allows for the resimulation of how light sources illuminate objects in a virtual environment.
[0092] As an example, the physically relighting in this application can use a physically based rendering (PBR) model, combining normal and albedo information, to simulate the lighting effects in a new background environment. For instance, based on PBR, combining normal information A1-2, albedo image A1-3, and lighting information from background image A2, the lighting effects in background image A2 can be simulated for portrait A11, thus obtaining the composite image A3 after physically relighting. This ensures that the lighting conditions of the moved portrait A11 are identical to those of the background image A2.
[0093] However, the aforementioned algorithms for image illumination prediction, normal prediction, and albedo prediction contain errors, which can make the altered image appear unrealistic. For example, Figure 1A The composite image A3 after relighting shows problems such as blurred edges, damaged edges, or incomplete background removal, which makes the portrait A11 in the composite image A3 look somewhat unrealistic.
[0094] To address the issue of unrealistic images caused by blurred or damaged edges in composite images after background replacement, this application provides an image processing method. In this method, for the composite image after background change, the edge region between the foreground object and the background is re-image synthesized. This re-image synthesis can be performed using methods such as style transfer, image enhancement, color correction, texture creation, or shape modification. The edge region includes: the outline of the foreground object, the content inside the foreground object adjacent to the outline, and the content outside the foreground object adjacent to the outline. This edge region is the area where the foreground object and background are blended in the re-lit composite image. Therefore, re-image synthesizing this edge region repairs the blended area. For example, for areas with broken or discontinuous edges of the foreground object in the edge region of the composite image after background change, image synthesis can generate new pixels to fill the gaps. These new pixels are set to have the same or similar color as their adjacent pixels, thereby eliminating the edge blurring problem and improving the consistency of lighting between the foreground object and the background in the composite image.
[0095] Simultaneously, this method redraws all regions of the synthesized image except for the edge regions, including the foreground within the edge regions and the background outside the edge regions. Furthermore, the redrawing parameters for the edge regions are greater than those for the other regions. It can be understood that the redrawing parameters for an image region refer to the magnitude of the pixel value change or the degree of style transfer. For example, the pixel value change in the edge regions is greater than that in the other regions, or the style transfer in the edge regions is greater than that in the other regions. Thus, while performing edge restoration on the foreground object's contour in the synthesized image, it does not cause significant changes to the foreground object itself or the background, such as preventing significant distortion of the replaced foreground portrait. Therefore, it ensures a natural and realistic image fusion.
[0096] In some embodiments, to specifically address the issue of unrealistic images caused by blurred and damaged edges in the composite image after relighting, the edge region where the foreground object is located and other regions in the composite image after relighting can be redrawn separately, with the redrawing parameters for the edge region being greater than those for the other regions. This ensures a realistic and natural image fusion.
[0097] Of course, the composite image after changing the background in this application is not limited to the composite image after relighting as described above. It can also be an image with a changed background but without relighting the foreground object. This application does not specifically limit this.
[0098] Specifically, based on a trained neural network model with image redrawing capabilities (hereinafter referred to as the diffusion realization model, and as an instance of the first network model), the edge regions and other regions in the synthesized image can be redrawn to repair and color correct the edge regions, thereby enhancing the realism of the synthesized image.
[0099] For example, in some embodiments, the training process of the diffusion realization model may involve: acquiring training text and corresponding high-resolution training images, where the training images represent the desired output. Then, the training text is input into a pre-defined neural network model, causing the model to generate a training result image corresponding to the training text. Next, the actually generated training result image is compared with the corresponding high-resolution training image to adjust the parameters of the neural network model based on the differences between the two images. Similarly, multiple training texts and their corresponding high-resolution training images can be used to adjust the parameters of the neural network model, thereby obtaining a trained diffusion realization model. It can be understood that this diffusion realization model acquires its image redrawing capability through text-to-image technology, i.e., by reconstructing images from text.
[0100] Furthermore, in some other embodiments, the training process of the aforementioned diffusion realization model can be as follows: A low-quality training image and a corresponding high-quality training image are acquired, where the high-quality training image represents the desired output result. For example, the low-quality training image can be a composite image generated by conventional relighting techniques, where the foreground object and background are fused together. The high-quality training image can be an image that has been manually repaired from the composite image, where the fusion effect between the foreground object and background is more realistic. Then, the low-quality training image is input into a preset neural network model, causing the neural network model to generate a training result image corresponding to the low-quality training image. The actually generated training result image is then compared with the corresponding high-quality training image to adjust the parameters of the neural network model based on the differences between the two images. Similarly, multiple low-quality training images and their corresponding high-quality training images can be used to adjust the parameters of the neural network model, thereby obtaining a trained diffusion realization model. It can be understood that this diffusion realization model obtains image redrawing capability by reconstructing high-quality images from low-quality images.
[0101] As an example, the diffusion realization model described above can be the U-Net model, but it is not limited to this. Furthermore, the network structure, training, and usage procedures of the diffusion realization model will be described in detail in this application, and will not be repeated here.
[0102] Furthermore, differences in resolution between the foreground and background images can also create an unrealistic appearance. For example... Figure 1A In the image A3, because the resolution of the foreground image A1 differs from that of the background image A2, the resolution of the portrait A11 may differ from that of the background, further increasing the unrealistic appearance of the portrait A11. Therefore, in some embodiments, the image processing method provided in this application, while redrawing the edge regions between the foreground object and the background in the re-lit composite image, can also perform a certain degree of image redrawing on the composite foreground object and background, for example, adding detail features to the overall composite image. This further improves the consistency of lighting conditions between the foreground object and the background in the composite image.
[0103] Furthermore, in some embodiments, to further improve the quality of the image after relighting, this application implements super-resolution processing on the redrawn image obtained from the redrawing process to obtain a higher resolution enhanced image. Specifically, the super-resolution processing can divide the redrawn image into blocks in the latent space and fuse the latent features of each block to obtain the enhanced image.
[0104] As can be understood, latent features of an image are key information in the image that is not easily observed directly but is helpful for the task, such as edges, textures, shapes, and object parts. The latent space, on the other hand, refers to the low-dimensional continuous space to which data is mapped within a neural network. By manipulating data in the latent space, various functions such as dimensionality reduction, noise reduction, and style transfer can be achieved.
[0105] In some embodiments, super-resolution processing of images in this application can be implemented by a network model (hereinafter referred to as the diffusion enhancement model, or the second network model).
[0106] For example, this application can generate a diffusion enhancement model (as an instance of a second network model) based on diffusion model generation technology, so as to perform image enhancement on the redrawn image through the diffusion enhancement model, that is, to perform super-resolution processing on the redrawn image.
[0107] As an example, the aforementioned diffusion enhancement model can be a U-Net model, but it is not limited to this. Furthermore, the network structure, training, and usage procedures of the diffusion enhancement model will be described in detail in this application, and will not be repeated here.
[0108] Next, for ease of explanation, let's combine... Figure 2A The image changes during the background replacement process shown, and Figure 2B The example shown illustrates the process of changing the background of an image, thus explaining an image processing scenario.
[0109] As an example, this application uses the composite image A3 obtained by relighting the foreground image A1 and the background image A2 as an example to illustrate the image redrawing and super-resolution processing in the image processing method provided in this application.
[0110] exist Figure 1A On the basis of, such as Figure 2A As shown, for the composite image A3 with the human figure A1 added, image redrawing can be performed to obtain the redrawn image A4. Wherein, as... Figure 2AAs shown, for the outline A12 of the portrait A11 in the re-illuminated composite image A3, the edge region A13 (as an example of the first region) where the outline A12 is located can be determined. At this time, the edge region A13 is obtained by extending some pixel content inward (i.e., towards the portrait A11 region) and outward (i.e., towards the background) along the outline A12. That is, the edge region A13 includes not only the outline A12 but also part of the portrait A11 content and part of the background content. Furthermore, the other regions A14 in image A3 besides the edge regions (as an example of the second region) include the interior region of the portrait A11 and the remaining background region. Thus, this application can redraw the edge region A13 and other regions A14 in the synthetic image A3 according to different redrawing parameters. The redrawing parameter for the edge region A13 is larger and the redrawing parameter for the other region A14 is smaller. Therefore, region A14 will not be distorted due to the larger redrawing parameter, and the authenticity of the portrait itself and the background in the other region A14 will not be affected, thus obtaining the repaired redrawn image A4.
[0111] In some embodiments, Figure 2A On the basis of, such as Figure 2B As shown, for the redrawn image A4, it can be further divided into blocks in the latent space, and super-resolution processing can be performed on the latent features of each block, such as adding detail features to each block. Then, these blocks with improved resolution are fused to achieve super-resolution processing of the redrawn image A4, resulting in a higher-resolution enhanced image A5.
[0112] exist Figure 1B On the basis of, such as Figure 2C As shown, the image processing scenario includes a foreground processing workflow and a background processing workflow. Specifically, as... Figure 2B The image processing flow shown includes the following steps:
[0113] S1-S7, among which Figure 2B S1-S7 and Figure 1B S1-S7 are the same, so they will not be elaborated here.
[0114] S8: Use the diffusion realization model to redraw image A3 to obtain redrawn image A4.
[0115] Combination Figure 2A It can be seen that the redrawing parameters of the diffusion realization model for the edge region A13 where the human figure A11 is located in image A3 are greater than the redrawing parameters for other regions A14 outside the edge region A13 in image A3. Thus, while improving the image fusion effect, it maintains the realism of the human figure.
[0116] Optionally, the above-mentioned redrawing of image A3 can be performed on the upsampled composite image A3.
[0117] The process of redrawing images using the diffusion real model in this application will be described in detail below, and will not be repeated here.
[0118] S9: The super-resolution processing of the redrawn image A4 is performed using the diffusion enhancement model to obtain the enhanced image A5.
[0119] Combination Figure 2B As can be seen, the diffusion enhancement model divides and fuses image A4, achieving enhancement effects such as adding detailed features to image A4, thereby improving the resolution of image A5 and improving the quality of the image after background replacement.
[0120] Thus, in scenarios where the image background is changed, this application can utilize a diffusion model to redraw the synthesized image after relighting, achieving edge repair and other processing, and further performing super-resolution processing. This allows the image after background replacement to maintain consistency in lighting conditions between the background and foreground, while significantly enhancing the realism of the blending between the foreground object and the background.
[0121] In the following embodiments, for ease of description, the foreground image to be replaced, the background image providing the background, the composite image after relighting, the redrawn image, and the enhanced image after super-resolution processing can be referred to as the third image, the fourth image, the first image, the second image, and the fifth image, respectively.
[0122] It is understood that at least one of the third image (the background to be replaced) and the fourth image (the background provided) can be selected by the user. For example, the third and fourth images can be selected based on the image processing application scenario or the file to which the images belong.
[0123] In some embodiments, the image processing method provided in this application can be applied to different application scenarios, such as photography scenarios (e.g., photography using the system camera's photo-taking function), video recording scenarios, shooting preview scenarios, video conferencing scenarios, video call scenarios, image editing scenarios, wallpaper changing scenarios, etc., but is not limited to these. For example, the above application scenarios may also include long and short video applications, live video applications, online video course applications, portrait intelligent camera movement applications, video surveillance, smart doorbells, etc. It can be understood that the third image is an image in these scenarios. For example, in a video call scenario, the third image can be an image of the current frame in the video footage captured by the camera of an electronic device.
[0124] Optionally, the fourth image can be an image generated based on user-input text information (i.e., the fourth image is an image generated based on text-based image processing technology), a panoramic high dynamic range (HDR) image, or a pre-set non-HDR image. Similarly, the third image can also be an image generated based on text-based image processing technology, an HDR image, or a pre-set non-HDR image, etc.
[0125] As an example, the aforementioned HDR image can be a 360-degree panoramic image, which can be captured using a professional camera. During the acquisition of this HDR image, ambient light information of the shooting scene, such as light intensity, direction, and color, can be simultaneously acquired. Therefore, when the fourth image is an HDR image, the pre-captured lighting information of the fourth image can be directly obtained simultaneously with its acquisition, without needing to perform an image lighting estimation algorithm on the fourth image. For example, in some studio shooting scenarios, HDR images can be used as background images to replace the background of foreground images.
[0126] As an example, the pre-set non-HDR image mentioned above can be a regular image, such as an image obtained from a server or an image saved locally.
[0127] As an example, the fourth image generated based on the text-based image processing technology can be an image automatically generated based on text information input by the user in real time. This text-based image processing technology can be generated using generative neural networks, such as the stable diffusion (SD) model. It can be understood that the text information can include foreground objects, background, style, lighting information, etc. For example, the text information input by the user (also called the prompt) could include "best quality, masterpiece, ultra-high resolution, original photo, cinematic lighting, a girl, dark style, night, city." In this case, the foreground object of the generated fourth image would be a girl, and the background would be a dimly lit city night scene.
[0128] It is understandable that when the fourth image is a real-time generated or pre-set non-HDR image, the illumination information of the fourth image can be predicted based on the image illumination estimation algorithm.
[0129] Furthermore, in some embodiments, the files to which the foreground and background images to be replaced in this application belong can be animated GIFs, ordinary images, or videos. Specifically, the third image to be replaced as the background in this application can include: a frame from an animated GIF, a still image (such as an ordinary image), or a frame from a video. Additionally, the fourth image providing the background can include: a frame from an animated GIF, a still image, or a frame from a video. Here, the animated GIF refers to a file in Graphics Interchange Format (GIF) format, which can contain multiple frames and supports animation and transparency.
[0130] It is understood that the files of the third and fourth images can be selected by the user; for example, the user can choose different files in different application scenarios. As an example, in a photography scenario, both the third and fourth images can be static images. In video or image editing scenarios, both the third and fourth images can be any feasible file. For example, if the third image is a static portrait of a person in the foreground, and the fourth image is a dynamic image of a beach changing frame by frame from sunrise to sunset, then after moving the portrait to the fourth image and relighting it to obtain a composite image, the composite image can be subjected to the redrawing and super-resolution processing described in this application, so that the lighting information of the portrait in the final image changes with the ambient light.
[0131] In some embodiments, the third image may not include a background. For example, in a photo editing scenario, the third image may be a foreground portrait or other foreground object for which the background needs to be replaced.
[0132] In some embodiments, the third image may include multiple foreground objects, such as multiple foreground portraits. In this case, changing the background of the third image in this application refers to moving one or more foreground portraits from the third image into a fourth image. It is understood that the foreground objects to be moved in the third image can be user-defined, and the positions of the foreground objects to be moved in the fourth image can also be user-defined.
[0133] The image processing method in this application will be illustrated with examples in the following application scenarios.
[0134] The following combination Figures 3 to 5 This application describes the photographic scenarios provided.
[0135] In one example, a mobile phone is used as an example of an electronic device, such as... Figure 3 As shown in (a), in response to a user's tap on the camera application, the phone opens the camera and displays the following: Figure 3The image (b) shows the display interface for the photo-taking mode; this display interface may include a shooting interface 10; the shooting interface 10 may include a viewfinder 11 and a control 12 for indicating that the photo should be taken; before the user clicks the control 12, a preview image may be displayed in the viewfinder 11. In the preview image, the person is standing in front of a large tree outdoors.
[0136] When a user wants to change the background of a portrait, they can... Figure 3 In (b) of the shooting interface 10, swipe upwards to make it appear as if... Figure 3 (c) and (d) show the addition of controls 13 and 14 to the mobile phone's camera display interface 10 to indicate different backgrounds. HDR images in each background indicated by control 13 can be marked with the symbol "HDR" in the upper right corner of the background thumbnail. Similarly, GIF images in each background indicated by control 13 can also be marked with symbols on the background thumbnail, such as the symbol "GIF".
[0137] So, when a user wants to change the current outdoor background to an indoor background, the user can make a selection on the display interface and click on the control 13 that indicates the indoor background to switch the background.
[0138] like Figure 3 As shown in (c), the user can click on the control 13 that indicates the indoor background. When the mobile phone detects that the user clicks on the control 13 that indicates the seaside background, it responds to the user's operation by keeping the portrait unchanged and switching the background of the portrait from outdoor to indoor.
[0139] like Figure 3 As shown in (d), combined with the switched indoor background, after detecting the user's click on control 12, the mobile phone can take a picture of the portrait with the indoor background in response to the user's operation.
[0140] In some embodiments, Figure 3 (d) shows the image A4 after the background is switched in the preview interface 11. It can be an image after relighting based on the indoor background, and the enhanced image after redrawing the edge area where the portrait A11 is located based on the diffusion network and then performing super-resolution processing.
[0141] In other embodiments, Figure 3 (d) shows that the image after changing the background in the preview interface 11 can be an image with relighting based on an indoor background. After the user clicks the control 12 to trigger the photo capture, the mobile phone can redraw the edge area of the image portrait A11 in the preview interface 11, and then perform super-resolution processing to enhance the image, that is, take the photo to obtain the enhanced image.
[0142] In addition, in some embodiments, if the background provided in the preview interface 11 does not meet the user's needs, the user can trigger the mobile phone to generate a new background image using text-based graphics through the control 12.
[0143] like Figure 4 As shown, a mobile phone can obtain text information entered by the user through text input. Specifically, as... Figure 4 As shown in (a), when the phone detects that the user has clicked control 14, in response to the user's action, the phone can display the following: Figure 4 The input control 15 is shown in (b) above. When the phone detects user interaction with the keyboard in the input control 15, it responds to this interaction by acquiring the text information entered by the user. Furthermore, after the user completes inputting the text "Cinema Lights Night Dark Style City," when the phone detects a click on the confirmation control 16 in the input control 15, it responds to this interaction by automatically generating a background image of a city night scene, and as shown... Figure 4 As shown in (c), control 13 is used to display the background of the city night scene in the preview interface 11, and the city night scene is changed for the portrait A11 in the preview interface 11.
[0144] like Figure 4 As shown in (c), in conjunction with the switched indoor background, after detecting the user's click on control 12, the mobile phone can take a picture of the portrait with the city night view as the background in response to the user's operation.
[0145] like Figure 5 As shown, mobile phones can obtain text information entered by users through voice input. Specifically, as... Figure 5 As shown in (a), when the phone detects that the user has long-pressed control 14, in response to the user's operation, the phone can receive the user's voice input "cinematic lights, nighttime, dark style city". Then, after the phone detects that the user has stopped long-pressing control 14, it determines that the user has completed voice input and can detect the text information "cinematic lights, nighttime, dark style city" from the voice. Thus, for the text information "cinematic lights, nighttime, dark style city", the phone automatically generates a background image of a city night scene, and as shown... Figure 5 As shown in (b), control 13 corresponding to the background of the city night scene is displayed in the preview interface 11, and the city night scene is changed for the portrait A11 in the preview interface 11.
[0146] like Figure 5 As shown in (b), in conjunction with the switched indoor background, after detecting the user's click on control 12, the mobile phone can take a picture of the portrait with the city night view as the background in response to the user's operation.
[0147] akin, Figure 4 or Figure 5 The image of the portrait A11 after changing to a city night scene displayed in the preview interface can be an image after relighting, or an image after relighting, image redrawing, and super-resolution processing.
[0148] The following combination Figure 6 and Figure 7 The wallpaper scenarios provided in the embodiments of this application will be described.
[0149] For example, taking the user-customized wallpaper scenario on a mobile phone as an example, such as Figure 6 As shown in (a), in response to a user's click on settings in the desktop interface, the phone displays as follows: Figure 6 The settings interface 60 in (b) of the settings interface. In response to the user's click on the wallpaper option 61 in the settings interface 60, the phone displays as follows: Figure 6 The personalized wallpaper settings interface 63 in (c) is described above. The image addition control 64 in the personalized wallpaper settings interface 63 is used to select the foreground image of the wallpaper, and the background selection control 65 is used to select the background image of the wallpaper.
[0150] like Figure 7 As shown, in response to the user's click on the image addition control 64 in the personalized wallpaper settings interface 63, the phone displays as follows: Figure 7 Image selection interface 67 is shown in (a). In response to a user's click on image 38 in image selection interface 67, the phone displays as shown... Figure 7 The personalized wallpaper settings interface 63 is shown in (b) above. It responds to user requests... Figure 7 When the control 65 indicating the landscape background is clicked in the personalized wallpaper settings interface 63 shown in (b), the phone displays the following: Figure 7 (c) shows the personalized wallpaper settings interface 63, which displays an image 69 with a changed landscape background. Furthermore, in response to the user's... Figure 7 When the control 66 in (c) is clicked, the phone can display the following: Figure 7 The desktop interface shown in (d) is a wallpaper of image 69 with a landscape background.
[0151] Understandable. Figure 7 The image 69 shown in the figure after changing the landscape background can be an enhanced image after relighting based on the lighting information of the landscape background, as well as image redrawing and super-resolution processing.
[0152] In some embodiments, Figure 7In the image 69 shown below, after the landscape background has been changed, the landscape background is a dynamic image from sunrise to sunset. At this time, image 69 is a dynamic image, and the lighting conditions of the portrait in image 69 change with the lighting conditions in the landscape background.
[0153] It is understood that the description of other application scenarios of the image processing method in this application embodiment is similar to the above-described photo taking and wallpapering scenarios, and this application embodiment will not elaborate on them.
[0154] The image processing method provided in this application can be applied to various electronic devices.
[0155] In this application embodiment, the electronic device 100 may be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, etc. This application embodiment does not impose any restrictions on the specific type of electronic device 100.
[0156] In this embodiment of the application, for ease of explanation, a mobile phone is used as an example to describe the execution subject of the image processing method of this application.
[0157] Furthermore, combined with Figures 8 to 21 The specific implementation of the image processing method provided in the embodiments of this application will be described in detail.
[0158] like Figure 8 The diagram shown is a flowchart of an image processing method provided in an embodiment of this application. The process includes:
[0159] S801: Obtain the third image of the background to be replaced and the fourth image that provides the background.
[0160] As an example, the third image may include one or more foreground objects. The fourth image may be a graphic obtained using texturing techniques, a pre-set non-HDR image, or an HDR image.
[0161] S802: Perform downsampling processing on the third image to obtain a downsampled third image.
[0162] For example, the size of the third image is 3K×4K pixels, and the size of the downsampled third image is 768×1K pixels. Of course, the size of the third image and the downsampled third image are not limited to the above example, and can be other sizes. This application does not specifically limit them. As an example, the third image and the fourth image can be the foreground image A1 and the background image A2 shown above, respectively.
[0163] S803: Determine the first mask image of the foreground object from the downsampled third image.
[0164] The first mask image is used to represent the outline and shape of the foreground object. A mask image is generally understood to be an image used to specify an area of interest (ROI) or to hide certain areas. A mask can be a binary image (black and white), where white areas represent the ROI and black areas are ignored or hidden. For example, in the foreground image A1', the area containing the person A11 can be white, and other areas can be black. Alternatively, a mask can be color, where each channel specifies a different processing method.
[0165] In some embodiments, the foreground object determined in the third image described above can be a foreground object to be moved, and the foreground object can be one or more foreground objects in the third image.
[0166] S804: Perform normal prediction on the downsampled third image to obtain the normal information of the third image. The normal information reflects the orientation of the surface of the foreground object in the third image.
[0167] S805: Perform albedo prediction on the downsampled third image to obtain the albedo information of the third image.
[0168] Among them, albedo information is used to reflect the color and material properties of the foreground object.
[0169] S806: Add the foreground object from the downsampled third image to the fourth image.
[0170] In some embodiments, this application can interactively add a foreground object to be moved in a third image to a fourth image.
[0171] Reference Figure 9 The image shown is a schematic diagram illustrating the movement of a foreground object in a wallpaper scene provided in an embodiment of this application. Figure 9 In the personalized wallpaper settings interface 63 shown in (a), the portrait is located at position L1. In response to a user's dragging operation on the portrait at position L1 to the left, the phone can move the portrait to the left, and... Figure 9The personalized wallpaper settings interface 63 shown in (b) displays the portrait at position L2. That is, the portrait is moved from position L1 on the landscape background to position L2. In this way, the display effect of the portrait after changing the background can meet the user's needs.
[0172] In some embodiments, this application may also downsample the fourth image and add the foreground object from the downsampled third image to the downsampled fourth image.
[0173] S807: Based on the first mask image, normal information, albedo information, and illumination information of the fourth image, the foreground object in the fourth image is relit to obtain the first image.
[0174] The lighting information of the foreground object in the first image is the same as that in the fourth image.
[0175] In some embodiments, the illumination information in the fourth image can be predicted from the fourth image using an image illumination estimation algorithm. For example, when the fourth image is an image generated using texturing techniques or a pre-set non-HDR image, its illumination information can be predicted using an image illumination estimation algorithm.
[0176] Furthermore, when the fourth image is an HDR image, the collected illumination information corresponding to that HDR image can be directly obtained.
[0177] It is understood that the descriptions of S801 to S806 above can be referenced to the descriptions of S1 to S6 above, and the similarities will not be repeated here.
[0178] S808: Upsample the first image to obtain an upsampled first image.
[0179] It is understood that the resolution of the first image in the previous sampling is higher than the resolution of the unsampled image. However, this application does not specifically limit the parameters for upsampling the third image; these parameters can be determined according to actual needs.
[0180] Optionally, in some other embodiments, steps S802 and 808 may not be performed. That is, this application can directly perform algorithms such as foreground prediction, normal prediction, and albedo prediction on the first image itself with high resolution. Correspondingly, the relighting processing, image redrawing, and image super-resolution processing of this application are also performed on the unsampled first image.
[0181] S809: For the upsampled first image, redraw the first region according to the first redraw parameters, and redraw the second region according to the second redraw parameters to obtain the second image.
[0182] The first redraw parameter is different from the second redraw parameter; for example, the first redraw parameter is greater than the second redraw parameter.
[0183] As an example, when the first redraw parameter is greater than the second redraw parameter, the number of redraws in the first region is greater than the number of redraws in the second region, and / or, the proportion of newly added features in the redrawn first region is greater than the proportion of newly added features in the redrawn second region. For example, this application can redraw the first region containing the foreground object 5 times in a pre-trained diffusion realization model, and redraw the second region 2 times in a pre-trained diffusion realization model. The greater number of redraws results in more newly added features in the redrawn image.
[0184] The first region, which is the edge region where the moved foreground object is located in the third image, includes the following: the outline of the foreground object, the region containing part of the outline of the foreground object, and part of the outline of the foreground object outside the foreground object; the second region is the region outside the first region in the third image.
[0185] Specifically, the first region includes the following: the outline of the foreground object, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object. For example, the multiple pixels adjacent to the pixels corresponding to the outline of the foreground object include: pixels inside the foreground object whose shortest distance to the pixels corresponding to the outline of the foreground object is less than a first distance, and pixels outside the foreground object whose shortest distance to the pixels corresponding to the outline of the foreground object is less than a second distance. The first distance and the second distance may be the same or different; for example, both may be 10 pixels. The specific values for the first distance and the second distance can also be other values, and there is no specific limitation on this. In this case, the first region in the first image includes: the outline of the foreground object, a region extending the outline of the foreground object inward by 10 pixels, and a region extending the outline outward by 10 pixels.
[0186] It can be understood that the shortest distance between pixel 1 corresponding to the outline of the foreground object and pixel 2 within the foreground object is: the distance between pixel 2 and pixel 1 in the direction perpendicular to the tangent of the outline at pixel 1. At this time, the line connecting pixel 2 and pixel 1 lies on the perpendicular line to the tangent of the outline at pixel 1.
[0187] In some embodiments, this application may employ a pre-trained diffusion realization model to redraw the image. Specifically, this application may input a third image into the pre-trained diffusion realization model, redraw the third image through the diffusion realization model, and output the redrawn fourth image.
[0188] Combination Figure 1A and Figure 2AIn the example shown, the third image can be a foreground image A1, and the foreground object to be moved is a portrait A11. The fourth image is a background image A2. Specifically, the first image is a composite image A3, and the second image is a redrawn image A4. Furthermore, the edge region A13 where the portrait A11 is located in the composite image A3 is the first region, and the other regions A14 are the second regions. Correspondingly, the redrawn second image can be the redrawn image A4.
[0189] S810: Perform super-resolution processing on the second image to obtain the fifth image.
[0190] The fifth image has a higher resolution than the second image.
[0191] In some embodiments, this application may employ a pre-trained diffusion enhancement model to perform super-resolution processing on the image. Specifically, this application may input a second image into the pre-trained diffusion enhancement model, perform resolution enhancement on the second image through the diffusion enhancement model, and output a fifth image with higher resolution.
[0192] As an example, refer to Figure 2A The fifth image can be the enhanced image A5, which is the result of super-resolution processing of the redrawn image A4. Clearly, the enhanced image A5 has a higher resolution than the redrawn image A4, and therefore, the enhanced image A5 is of higher quality.
[0193] Understandable. Figure 8 The execution order of each step is only one example, and other execution orders are also possible. For example, S804, S805, and S806 can be executed in parallel. This application does not specifically limit this.
[0194] Thus, this application, through a pre-trained diffusion realization model, can simultaneously repair the edges of the foreground object's contour in the first image after relighting, while also improving the overall lighting harmony and clarity of the first image, without causing significant changes to the foreground object itself or the background, such as preventing significant distortion of the replaced foreground portrait. This ensures the natural realism of the fused second image.
[0195] Application of diffusion realization model
[0196] In some embodiments, the diffusion realization model in this application may include multiple components, such as a first encoder, a first decoder, and a first diffusion model. The first encoder is used to acquire latent features of the image in the latent space, the first diffusion model is used to redraw the image based on the latent features, and the first decoder is used to restore the latent features in the latent space to the image in the image space.
[0197] Optionally, the first encoder and the first decoder can adopt a variational autoencoder (VAE) architecture, specifically a diffusion variational autoencoder (dVAE). In this case, the first encoder can be a dVAE encoder, and the first decoder can be a dVAE decoder.
[0198] Optionally, the first diffusion model described above can be implemented using the UNet model (a network model). Specifically, the first diffusion model can be a U-shaped structure in the UNet model, used to support skip connections between the first encoder and the first decoder. UNet combines shallow and deep features of the image in a concatenated manner. Shallow feature maps tend to express basic feature units such as points, lines, and edge contours, containing more spatial information. Deep feature maps tend to express the semantic information of the image, containing less spatial information but more semantic features.
[0199] Reference Figure 10 The structure of the diffusion realization model provided in the application embodiments is described. For example... Figure 10 As shown, the diffusion realization model 100 includes the following components: a dVAE encoder 101, a diffusion model 102, and a dVAE decoder 103. For example, the re-illuminated synthetic image A3 is input to the dVAE encoder 101, which extracts the original latent features of the synthetic image A3. Then, the original latent features of the synthetic image A3 are iteratively processed multiple times by the diffusion model 102 to obtain the redrawn latent features of the synthetic image A3. Furthermore, the redrawn latent features of the synthetic image A3 can be decoded by the dVAE decoder 103 to restore the redrawn image A4.
[0200] Next, the image redrawing process based on the diffusion realization model will be described in detail. For example... Figure 11A As shown, the image redrawing process includes the following steps S1101 to S1103:
[0201] S1101: Input the first image into the first encoder and encode the first image into a first feature in the latent space.
[0202] The first image can be data in image space, that is, the first image can be a color digital image containing three color channels: red (R), green (G) and blue (B).
[0203] It can be understood that a dVAE encoder can obtain latent feature data (i.e., features with a low resolution of 8 bits) in the latent space. Specifically, the first image can be encoded into a low-dimensional latent feature. This latent feature can be a 4-dimensional volume, where each volume slice can correspond to a specific feature or attribute in the decoded image. In this case, the latent feature has 4 channels, representing 3 color channels and image features (or parameters used to control different aspects of the image generation process).
[0204] For example, the first image is 1024×1024×3 channel image data, and the encoded first feature is 128×128×4 channel latent data. In this case, for a 128×128×4 dimensional latent feature data, this means there is a grid composed of 128×128 feature points, each feature point having 4 values to represent its feature, such as 3 color channel data and 1 image feature value. Such a data structure can be used to represent a compressed version of an image, or as an intermediate step in image compositing and editing.
[0205] S1102: Input the first feature into the first diffusion model, and process the first feature based on preset noise to obtain the second feature. The first redrawing parameter of the feature corresponding to the first region in the first feature, which is redrawn based on the preset noise, is greater than the second redrawing parameter of the feature corresponding to the second region in the first feature, which is redrawn based on the preset noise.
[0206] Optionally, the preset noise can be Gaussian noise.
[0207] Gaussian noise follows a Gaussian distribution (normal distribution), characterized by its randomness and continuity. It can vary throughout an image, mimicking noise effects in the real world caused by various factors such as sensor defects, transmission errors, and environmental interference. The parameters of Gaussian noise typically include its mean and standard deviation, which determine the noise intensity and the central location of its distribution. In image redrawing, different noise parameters can be selected depending on the desired noise level and the image content.
[0208] In image redrawing, Gaussian noise can simulate real-world noise, making it closer to the characteristics of natural images and more realistic. Data augmentation: In tasks such as image classification and detection, adding Gaussian noise of varying intensities to the images in the training set can increase data diversity, helping the model learn more comprehensive knowledge.
[0209] Furthermore, Gaussian noise can prevent overfitting during the training of diffusion-based realimagine models. It's understandable that adding Gaussian noise to the training data when training machine learning models (such as convolutional neural networks) can help the model learn more robust features, thereby improving its ability to generalize to unknown data.
[0210] S1103: Input the second feature into the first decoder and decode the second feature into the second image.
[0211] It is understandable that the first decoder, such as the dVAE decoder 103, can restore the second feature in the latent space to the second image in the image space, thereby completing the redrawing of the first image.
[0212] In some embodiments, the first diffusion model iteratively processes the first feature multiple times to obtain the second feature. For example, as... Figure 11B As shown, S1102 above includes S1102a to S1102g:
[0213] S1102a: Determine the first iteration number k*a based on the first redrawing parameter a and the preset parameter k, and determine the second iteration number k*b based on the second redrawing parameter b and the preset parameter.
[0214] For example, refer to Figure 10 The first feature is the original latent feature of the synthetic image A3.
[0215] It is understood that this application inputs the first region and the second region in the first image into the first diffusion model with different number of iterations, that is, the first diffusion model redraws the image of the two regions with different number of iterations.
[0216] The preset parameter k can be the preset number of iterations for the first diffusion model, for example, k=10. The actual number of the first iterations of the first diffusion model can be determined by combining the preset number of iterations and the first redraw parameter. For example, when the first redraw parameter is 0.5, the actual total number of iterations of the first diffusion model is k*a=5, that is, the number of iterations for the features corresponding to the first region in the first image is 5. Correspondingly, when the second redraw parameter b=0.2, the second number of iterations of the first diffusion model is k*b=2, that is, the number of iterations for the features corresponding to the second region in the first image is 2. In this case, the second iteration count is the last few iterations of the total number of iterations k*a.
[0217] S1102b: The preset noise and the first feature are superimposed according to the first redrawing parameters to obtain the first redrawing feature.
[0218] As an example, when the first redrawing parameter a=0.5, assuming the first feature is i and the preset noise is n, then the first redrawing feature = i*(1-a)+n*a.
[0219] S1102c: Corresponding to 1≤j≤(k*ak*b), in the j-th iteration, the j-th redrawn feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawn feature is output. Here, j is a positive integer.
[0220] It is understandable that when j=1, a preset noise is added to the image redrawing in the first iteration process, and no preset noise is needed in the image redrawing in the subsequent iteration processes, until the preset noise is introduced again in the image redrawing of the iteration process indicated by the second redrawing parameter.
[0221] As an example, when k=10, a=0.5, b=0.2, k*(ab)=3, then the second and third iterations of the first diffusion model do not require the introduction of preset noise, and the input of the later iteration is the output of the previous iteration, until the fourth iteration introduces preset noise again.
[0222] S1102d: Corresponding to j=(k*ak*b)+1, after superimposing the preset noise and the first feature according to the second redrawing parameters, the intermediate feature is obtained. The feature corresponding to the second region in the j-th redrawing feature is replaced with the feature corresponding to the second region in the intermediate feature to obtain the updated j-th redrawing feature.
[0223] Based on the above example, when k=10, a=0.5, b=0.2, and k*(ab)=3, before j=4, preset noise can be introduced again based on the second redrawing parameter. Then, the preset noise can be superimposed on the first feature according to the second redrawing parameter to obtain the intermediate feature = i*(1-b)+n*b. Furthermore, before the fourth iteration, the feature corresponding to the second region in the fourth redrawing feature is replaced with the feature corresponding to the second region in the intermediate feature to obtain the updated fourth redrawing feature. It can be understood that the feature corresponding to the second region of the first image in the updated fourth redrawing feature is obtained by introducing preset noise according to the second redrawing parameter 0.2 and without inputting into the first diffusion model, and the feature corresponding to the first region of the first image is obtained by introducing preset noise according to the first redrawing parameter 0.5 and iterating through the first diffusion model 3 times.
[0224] Specifically, in combination Figure 10Let C be the original latent feature (i.e., the first feature) of the synthesized image A3, N be the preset noise, and X be the fourth redraw feature. The first redraw parameter is 0.5, and the second redraw parameter is 0.2. Furthermore, let mask be the feature in the fourth redraw feature corresponding to the first region where portrait A11 is located, and let unmask be the feature in the intermediate features corresponding to the second region where portrait A11 is located. Then, the expression C *unMask*0.8 + N*0.2 -> X*unMask can be used to replace the feature in the fourth redraw feature corresponding to the second region with the feature in the intermediate features corresponding to the second region, thus obtaining the updated fourth redraw feature.
[0225] S1102e: Corresponding to j=k*(ab)+1, in the j-th iteration, the updated j-th redrawn feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawn feature is output.
[0226] Based on the above example, corresponding to j=4, in the 4th iteration, the updated 4th redraw feature can be input into the first diffusion model for image redrawing, and the 5th redraw feature can be output.
[0227] S1102f: Corresponding to j > (k*ak*b), in the j-th iteration, the j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output.
[0228] Based on the above example, corresponding to j=5, in the 5th iteration, the 5th redraw feature can be input into the first diffusion model for image redrawing, and the 6th redraw feature can be output. It can be understood that the feature corresponding to the second region of the first image in the 6th redraw feature is obtained by introducing preset noise according to the second redraw parameter 0.2 and iterating through the first diffusion model twice, and the feature corresponding to the first region of the first image is obtained by introducing preset noise according to the first redraw parameter 0.5 and iterating through the first diffusion model five times.
[0229] This results in a larger redraw parameter for the first region containing the edge of the moved foreground object in the sixth redraw feature, while the smaller redraw parameter for the second region containing the interior of the moved foreground object and the background. Thus, it is possible to redraw the first and second regions containing the foreground object in the first image after relighting using different redraw parameters.
[0230] S1102g: Corresponding to j=k*a, the (j+1)th redrawn feature is taken as the second feature.
[0231] Based on the example above, corresponding to j=5, after the 5th iteration, the 6th redraw feature can be used as the second feature. For example, the second feature mentioned above is the 6th redraw feature, which means... Figure 10 The features corresponding to the redrawn image A4 are shown.
[0232] Diffusion Realization Model Training
[0233] Furthermore, the training process of the diffusion realization model in this application is described.
[0234] In some embodiments, this application uses the training of a diffusion realization model based on text-to-image technology as an example for illustration. In this case, the diffusion realization model may also include a text encoder. For example, the text encoder may be a text encoder in a contrastive language-image pre-training (CLIP) model.
[0235] In some embodiments, this application may train a diffusion realization model based on at least one first training data to obtain a trained diffusion realization model; wherein the first training data includes a first training text and a first training image, wherein the first training text is related to the image features of the first training image, and a text encoder is used to obtain the text features of the first training text.
[0236] Reference Figure 12 The diagram shown illustrates the training process of the diffusion realization model. Figure 12 The diffusion realization model 100 shown also includes a text encoder 104.
[0237] It is understandable that the training dataset of the diffusion realization model 100 can be training text-training image pairs, such as a first training text and its corresponding first training image. For example, the first training text can be training text 1, and the corresponding first training image is training image 1. The image features of training text 1 and training image 1 are related, and training image 1 can be a high-resolution image, i.e., an image with high resolution.
[0238] Specifically, such as Figure 12 As shown, training text 1 can be processed by text encoder 104 to generate training result 1, which can then be used to train diffusion model 102 or dVAE decoder 103 in diffusion realization model 100. Furthermore, training image 1 can be processed by dVAE encoder 101 to generate result image 1, which can then be used to train dVAE decoder 103 in diffusion realization model 100.
[0239] Optionally, this application may first train the diffusion model 102, and then train the dVAE decoder 103 based on the trained diffusion model 102 to complete the overall training of the diffusion realization model 100.
[0240] Next, refer to Figure 13 The training method of the diffusion realization model provided for the embodiments of this application is described below, and the process includes the following steps:
[0241] S1301: Input the first training text from the first training data into the text encoder and output the text features of the first training text.
[0242] S1302: Input the first training image from the first training data into the first encoder and encode the first training image into the third feature of the latent space.
[0243] S1303: Input the third feature into the first diffusion model, process the third feature based on the preset noise, and obtain the fourth feature.
[0244] In some embodiments, this application may redraw the features corresponding to the third region in the third feature according to the first redrawing parameters based on preset noise, and redraw the features corresponding to the fourth region in the third feature according to the second redrawing parameters based on preset noise. The third region includes the following: the outline of a foreground object in the first training image, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object in the first training image; the fourth region is the region outside the third region in the first training image.
[0245] It is understandable that the third and fourth regions in the first training image can be referred to the relevant descriptions of the first and second regions in the first image above, which will not be repeated here.
[0246] In other embodiments, based on preset noise, all features in the third feature can be redrawn according to uniform redrawing parameters to obtain the redrawn fourth feature.
[0247] S1304: Input the third feature into the first decoder and decode the fourth feature into the first training result.
[0248] Furthermore, in some other embodiments, the first training result can also be generated based on the first training text. Specifically, this application can also input the text features of the first training text into a first diffusion model and output first text encoding features. Then, the first text encoding features are input into a first decoder and the first training result is output.
[0249] S1305: Adjust the network parameters of the first diffusion model based on the difference between the text features of the first training text and the third feature, and / or adjust the network parameters of the first decoder based on the difference between the first training image and the first training result.
[0250] It is understood that this application can set a loss function for the first diffusion model, and determine the value of the loss function based on the difference between the text features of the first training text and the third feature, and adjust the network parameters of the first diffusion model through the value of the loss function.
[0251] Similarly, this application can set a loss function for the first decoder, determine the value of the loss function based on the difference between the first training image and the first training result, and adjust the network parameters of the first decoder based on the value of the loss function.
[0252] It is understandable that when the value of the loss function approaches a stable, relatively small value, it indicates that the model has converged. At this point, a set of parameters that minimizes the loss function has been found as the network parameters for the network model. Furthermore, this application can use different training images as input to the diffusion realization model and repeatedly execute the test. Figure 13 The training process continues until convergence.
[0253] As an example, the network parameters mentioned above can include weights and biases. Weights are numerical values connecting nodes in different layers of the neural network; they are obtained through training and determine how the network processes the input data. Weights are randomly initialized before training begins and are continuously updated during training to minimize the loss function. Bias is another parameter used in conjunction with the weights of each node in the network. They are added to the input of neurons before the activation function to help adjust the probability of the neuron being activated.
[0254] Furthermore, the first diffusion model iteratively processes the third feature in the first training image multiple times to obtain the fourth feature. For example, as... Figure 14 As shown, S1303 above includes S1401 to S1406:
[0255] S1401: Determine the first iteration number k*a based on the first redrawing parameter a and the preset parameter k, and determine the second iteration number k*b based on the second redrawing parameter b and the preset parameter.
[0256] S1402: The preset noise and the third feature are superimposed according to the first redrawing parameters to obtain the first training redrawing feature.
[0257] S1403: Corresponding to 1≤j≤k*(ab), in the j-th iteration, the j-th training redraw feature is input into the first diffusion model for image redrawing, and the (j+1)-th training redraw feature is output.
[0258] S1404: Corresponding to j=k*(ab)+1, after superimposing the preset noise and the third feature according to the second redrawing parameters, the training intermediate feature is obtained. The feature corresponding to the second region in the j-th training redrawing feature is replaced with the feature corresponding to the second region in the training intermediate feature, and the updated j-th training redrawing feature is obtained.
[0259] S1405: Corresponding to j=k*(ab)+1, in the j-th iteration, the updated j-th training redraw feature is input into the first diffusion model for image redrawing, and the (j+1)-th training redraw feature is output.
[0260] S1406: Corresponding to j>k*(ab), in the j-th iteration, the j-th training redraw feature is input into the first diffusion model for image redrawing, and the (j+1)-th training redraw feature is output.
[0261] S1407: Corresponding to j=k*a, the (j+1)th training redraw feature is taken as the fourth feature.
[0262] Similarly, the detailed descriptions of S1401 to S1407 in this application can be found in the above descriptions. Figure 11B The relevant descriptions of S1102a to S1102g in the document are not repeated here. As an example, the second redrawing parameter (e.g., 0.2) can be the redrawing parameter of the second region in the first training image, while the first redrawing parameter (e.g., 0.5) can also be the redrawing parameter of the second region in the first training image. Furthermore, this application can predict the third region in the first training image using algorithms such as foreground prediction, thereby determining the features of the third region and the features of the fourth region in the third features of the first training image.
[0263] It is understood that in the embodiments of this application, multiple sets of first training text and first training images can be used to train the diffusion realization model, so that the diffusion realization model has the ability to redraw in different regions.
[0264] Application of diffusion enhancement model
[0265] Next, the diffusion enhancement model provided in the embodiments of this application will be described.
[0266] In some embodiments, the diffusion enhancement model in this application may include multiple components, such as a second encoder, a second decoder, and a second diffusion model. For example, the second diffusion model may be a pre-trained first diffusion model or a separately trained diffusion model; this embodiment does not specifically limit this. In this case, the first diffusion model may be integrated into a single component and inserted into the diffusion realization model or the diffusion enhancement model.
[0267] Optionally, the second encoder can be a dVAE encoder, and the second decoder can be a dVAE decoder.
[0268] Reference Figure 15 The structure of the diffusion enhancement model provided in the application embodiments is described. For example... Figure 15 As shown, the diffusion enhancement model 200 includes the following components: dVAE encoder 201, diffusion model 202, dVAE decoder 203, control network 204, repair network 205, and image encoder 206.
[0269] The dVAE encoder 201 is used to input the block image of the redrawn image A4 and extract the original latent features of the block image.
[0270] Furthermore, the image encoder 206 is used to extract semantic features (i.e., semantic information) from the segmented image, such as low-level features (e.g., edges and textures) and high-level features (e.g., object categories and scene layouts). For example, the image encoder 206 can be an image encoder in a CLIP model.
[0271] The control network 204 receives latent features from the dVAE encoder 201 as input and adjusts them as needed, such as by introducing constraints. For example, these constraints can be semantic features extracted by the image encoder 206. In this case, the semantic information introduced by the control network 204 serves as constraints, ensuring that the image generated by the diffusion model 202 conforms to these descriptions.
[0272] It can be understood that the control network 204 is used to support the diffusion model 202 in adjusting the latent features of the block image from the dVAE encoder 201 based on the semantic features from the image encoder 206, so that the adjusted latent features of the block image conform to the semantic features.
[0273] The diffusion model 202 is used to adjust the latent features of the expanded block image based on the semantic features of the block image, so as to obtain the adjusted latent features of the block image.
[0274] The repair network 205 is used to repair the image after the latent features of multiple segmented images are stitched together. Specifically, it repairs the stitching points in the stitched features so that the content in the stitched features is consistent with the content in the original redrawn image A4.
[0275] As an example, the above-mentioned repair network 205 can be implemented using the Unet model, but is not limited to this.
[0276] The dVAE decoder 203 is used to decode the stitched features to obtain the enhanced image A5.
[0277] Reference Figure 16 This paper describes the image super-resolution processing workflow based on the diffusion enhancement model in this application, which includes the following steps:
[0278] S1601: Input the N first blocks of the second image into the second encoder respectively, and encode the N first blocks in the latent space to obtain N fifth features.
[0279] Where N is a positive integer.
[0280] In some implementations, the N first blocks are obtained by regularly dividing or overlapping the second image.
[0281] It can be understood that the regular segmentation method means that there is no overlap between the segments in the second image, while the overlapping segmentation method means that there is overlap between the segments in the second image.
[0282] Reference Figure 15 As shown, the second image is a redrawn image A4, which was obtained by regular block division. In this case, there are no overlapping parts between these blocks.
[0283] In addition, refer to Figure 17 As shown in the schematic diagram of an overlapping segmentation method, for example, segment 171 and segment 172 overlap. Thus, when the latent features of subsequent segments are stitched together, the possibility of distortion at the stitching point is small; for example, the seam of clouds in the sky in two stitched segments matches without distortion.
[0284] S1602: The N first blocks are processed by the image encoder to obtain N semantic features.
[0285] It can be understood that the semantic features of a first block are used to characterize the image content in that block. For example, the semantic feature of a first block is "clouds in the night sky," indicating that the content of that block is a night scene with clouds in the sky.
[0286] S1603: Input N fifth features into the control network and output N corresponding sixth features, where the sixth features are constrained by the corresponding semantic features.
[0287] It can be understood that the sixth feature mentioned above includes the fifth feature while being constrained by the corresponding semantic feature. Therefore, the control network supports the subsequent second diffusion network in adjusting the sixth feature to conform to the content indicated by the corresponding semantic feature.
[0288] S1604: Input N semantic features and N sixth features into the second diffusion model, obtain the corresponding seventh features based on each sixth feature and the corresponding semantic features, and output N seventh features corresponding to the N first blocks.
[0289] It is understood that in the embodiments of this application, the second diffusion model can perform image enhancement for the seventh feature to increase the resolution of the image restored by the seventh feature.
[0290] In addition, the seventh feature is the feature that matches the semantic features of the corresponding first block, which makes it less likely that the content of the subsequent N first blocks will be distorted.
[0291] S1605: Combine the N seventh features into an eighth feature, and upsample the eighth feature in the latent space to obtain the ninth feature.
[0292] It can be understood that the N seventh features correspond to the features of the N first blocks, and the combined eighth feature can correspond to the entire second image. Therefore, upsampling the eighth feature in the latent space to obtain the ninth feature further increases the information of image details in the ninth feature, which helps reduce the possibility of distortion in the subsequent reconstruction of the N first blocks.
[0293] S1606: Input the ninth feature into the repair network, process the data from different seventh features in the ninth feature, and obtain the tenth feature.
[0294] It is understandable that the repair network can repair the junctions between features from different blocks in the assembled ninth feature to ensure that the content at the junctions is not distorted, thereby reducing the possibility that the content of the N first blocks themselves will be distorted in the subsequent restoration.
[0295] S1607: Input the tenth feature into the second decoder and output the fifth image.
[0296] It is understandable that the second decoder can decode the second feature, thereby decoding the tenth feature in the latent space into the fifth image in the image space, so as to obtain the enhanced image after super-resolution processing (i.e. image enhancement) of the redrawn image.
[0297] Thus, the pre-trained diffusion enhancement model provided in this application embodiment can enhance the redrawn image after relighting and redrawing. Specifically, it enhances each block separately to achieve image super-resolution processing, thereby improving the overall resolution of the subsequently restored image. Furthermore, the features of each stitched block can be repaired at the stitching points using a pre-trained insulation network, ensuring that there is no content distortion at the stitching points of the stitched and restored image, thus guaranteeing the realism of the enhanced image.
[0298] Diffusion-enhanced model training
[0299] In some embodiments of this application, the first decoder and the repair network in the diffusion enhancement model can be trained separately.
[0300] In some embodiments, the present application embodiments may train a second diffusion model based on at least one second training data to obtain a trained second diffusion model; wherein the second training data includes a first training block and a first training semantic feature, and the first training semantic feature is the semantic feature of the first training block.
[0301] It is understandable that the training dataset of the diffusion realization model 100 can be a training semantic feature-training block pair, such as the first training block corresponding to the first training semantic feature. For example, the first training block can be training block 1, and the corresponding first training semantic feature is the first training semantic feature. The image content of training block 1 conforms to the semantics of training semantic feature 1, and training block 1 can be a high-definition image, that is, a block in an image with high resolution.
[0302] It is understood that the first training block mentioned above can be a block obtained from a high-definition image using a random block method, a regular block method, or an overlapping block method.
[0303] Specifically, such as Figure 18 As shown, the training block 1 can be processed by the text encoder 104 to generate the result block 1, which is then used to train the dVAE decoder 203 in the diffusion enhancement model 200.
[0304] Optionally, this application may first train the diffusion model 102, and then train the dVAE decoder 103 based on the trained diffusion model 102 to complete the overall training of the diffusion realization model 100.
[0305] Reference Figure 19 As shown, the training process of the second decoder provided in this application embodiment is described, and the process includes the following steps:
[0306] S1901: Input the first training block into the second encoder, and encode multiple first training blocks in the latent space to obtain the corresponding eleventh feature.
[0307] It can be understood that the eleventh feature is the original latent feature of the first training block in the latent space.
[0308] S1902: Input the eleventh feature into the control network and output the corresponding twelfth feature, wherein the twelfth feature uses the corresponding first training semantic feature as a constraint.
[0309] S1903: Input the first training semantic feature and the twelfth feature into the second diffusion model, and output the thirteenth feature corresponding to the first training block.
[0310] It is understandable that the image content indicated by the thirteenth feature matches the first trained semantic feature. Specifically, in some embodiments, the second diffusion model can supplement the twelfth feature with image details or perform feature adjustments based on the first trained semantic feature, thereby improving the resolution of the segmented image while ensuring that the content of the segmented image is not distorted to a certain extent.
[0311] S1904: Input the thirteenth feature into the second decoder and output the second training result.
[0312] For example, the second training result mentioned above can be Figure 18 The result shown is block 1. At this point, the second training result can be an image in the image space.
[0313] S1905: Adjust the network parameters of the second decoder based on the difference between the second training result and the first training block.
[0314] Similarly, the description of adjusting the network parameters in the second decoder in this application can refer to the relevant description of the first decoder and the first diffusion model above, and will not be repeated here.
[0315] Thus, when the difference between the second training result and the first training block is small, the trained second decoder is obtained to train the diffusion enhancement model, enabling the diffusion enhancement model to enhance the resolution of the block image.
[0316] The training process of the repair network in the diffusion enhancement model will be explained next.
[0317] Reference Figure 20 The diagram shown is a schematic representation of a network repair training process provided in an embodiment of this application.
[0318] In some embodiments, this application may train the repair network based on at least one third training data to obtain a trained repair network; wherein the third training data includes a second training image and a third training image, and the third training image is an image obtained by reducing the resolution of the second training image.
[0319] It is understandable that the second training image can be a high-resolution image, and the third training image can be an image that has been downgraded from the second training image, i.e., the resolution has been reduced.
[0320] Next, combined Figure 20 Taking training image 2 as the second training image and training image 3 as the third training image as an example, the training process of the repair network will be explained. Figure 20 As shown, each block in training image 3 (such as training block 2) is processed by encoder 201 to output the original latent features K1 of training image 3, and then the resolution is upscaled by diffusion model 202 to obtain the adjusted latent features K2 corresponding to each block in training image 3. Then, the adjusted latent features corresponding to each block in training image 3 are concatenated and upsampled to obtain the upsampled latent features K3 of training image 3. Next, the upsampled latent features are processed by repair network 206 to obtain repaired latent features K4. The difference between the repaired latent features K4 and the original latent features GT output by encoder 201 of training image 2 (represented by loss) is used to train repair network 206.
[0321] Furthermore, in other embodiments, the encoder or diffusion model used during the training of the repair network 206 in this application may be other networks, not limited to those described above. Figure 20 The network shown in this application embodiment will not be described in detail.
[0322] Next, refer to Figure 21 The training process of the repair network provided in the embodiments of this application will be described in detail. The process includes the following steps:
[0323] S2101: Input the N second training blocks of the third training image into the second encoder, and encode the N second training blocks in the latent space to obtain N fourteenth features. Where N is a positive integer.
[0324] In some embodiments, the N second training blocks are obtained by regularly dividing or overlapping the third training image.
[0325] For example, a first block can be Figure 20 The fourteenth feature in the training image 3 shown is the original latent feature of that block.
[0326] S2102: Input N fourteenth features into the second diffusion model and output N fifteenth features corresponding to the N second training blocks. The data size of the fifteenth feature is greater than that of the corresponding fourteenth feature.
[0327] It is understandable that the second diffusion model can improve the resolution of the fourteenth feature by adding image details in the latent space, thereby obtaining the corresponding fifteenth feature.
[0328] S2103: Combine N fifteenth features into a sixteenth feature, and upsample the sixteenth feature in the latent space to obtain the seventeenth feature.
[0329] S2104: Input the seventeenth feature into the repair network, process the data from different fifteenth features in the seventeenth feature, and obtain the eighteenth feature.
[0330] S2105: Input the second training image into the second encoder and encode the second training image in the latent space to obtain the nineteenth feature.
[0331] S2105: Adjust the network parameters of the repair network based on the difference between the eighteenth and nineteenth features.
[0332] Similarly, the description of adjusting the network parameters in the repair network in this application can be found in the relevant description of the first decoder and the first diffusion model above, and will not be repeated here.
[0333] Thus, when the difference between the eighteenth and nineteenth features is small, the trained inpainting network is obtained, enabling the training of the diffusion enhancement model. This allows the diffusion enhancement model to perform content consistency restoration on segmented images in the latent space, ensuring that the image after stitching together the segments can faithfully reproduce the image content. In this way, the diffusion enhancement model can fuse latent features in the latent space and eliminate image distortion problems caused by stitching together segments, thereby achieving the ability to generate high-resolution images.
[0334] The structure of the electronic device for the image processing method provided in this application embodiment will be described below, taking a mobile phone as an example.
[0335] like Figure 22 As shown, the mobile phone 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, buttons 101, and a display screen 102, etc.
[0336] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone 10. In other embodiments of this application, the mobile phone 10 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0337] Processor 110 may include one or more processing units, such as processing modules or circuits of a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), microprocessor (MCU), artificial intelligence (AI) processor, or field programmable gate array (FPGA). Different processing units may be independent devices or integrated into one or more processors. Processor 110 may include storage units for storing instructions and data, such as the pre-trained diffusion realization model and diffusion enhancement model, and corresponding training data. In some embodiments, the storage unit in processor 110 is a cache memory 180. For example, processor 110 is used to redraw the edge region of the foreground object in the relit image and to perform super-resolution processing on the redrawn image.
[0338] The power module 140 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component manages the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects to the power supply and the processor 110. The power management module receives input from the power supply and / or the charging management module to supply power to the processor 110, the display 102, the camera 170, and the wireless communication module 120, etc.
[0339] The mobile communication module 130 may include, but is not limited to, antennas, power amplifiers, filters, and low-noise amplifiers (LNAs). The mobile communication module 130 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the mobile phone 10. The mobile communication module 130 can receive electromagnetic waves via the antenna, filter and amplify the received electromagnetic waves, and then transmit them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device. Wireless communication technologies can include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wide Band Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Longer Mevolution (LTE), Bluetooth (BT), Global Navigation Satellite System (GNSS), Wireless Local Area Networks (WLAN), Near Field Communication (NFC), Frequency Modulation (FM) and / or Field Communication (NFC), Infrared (IR) technology, etc.GNSS can include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0340] The wireless communication module 120 may include an antenna, which enables the transmission and reception of electromagnetic waves. The wireless communication module 120 can provide solutions for wireless communication applications on the mobile phone 10, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The mobile phone 10 can communicate with networks and other devices through wireless communication technologies.
[0341] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 may also be located in the same module.
[0342] The display screen 102 is used to display human-computer interaction interfaces, images, videos, etc. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. For example, the display screen 102 provided in this application can be a QLED screen, displaying individual frames, such as the frame containing pixels after brightness compensation.
[0343] The sensor module 190 may include proximity sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.
[0344] The audio module 150 is used to convert digital audio information into analog audio signals for output, or to convert analog audio input into digital audio signals. The audio module 150 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 150 may be located in the processor 110, or some functional modules of the audio module 150 may be located in the processor 110. In some embodiments, the audio module 150 may include a speaker, a handset, a microphone, and a headphone jack.
[0345] Camera 170 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to image signal processing (ISP) to be converted into a digital image signal. Mobile phone 10 can achieve its shooting function through ISP, camera 170, video codec, graphics processing unit (GPU), display 102, and application processor.
[0346] Interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a microSD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 through the external memory interface to perform data storage. The USB interface is used for communication between the mobile phone 10 and other electronic devices. The SIM card interface is used to communicate with the SIM card installed in the mobile phone 10, for example, to read or write phone numbers stored in the SIM card.
[0347] In some embodiments, the mobile phone 10 further includes buttons 101, a motor, and indicators. The buttons 101 may include volume buttons, a power button, etc. The motor is used to generate a vibration effect in the mobile phone 10, for example, vibrating when the user's mobile phone 10 is called to prompt the user to answer the call. The indicators may include laser indicators, radio frequency indicators, LED indicators, etc.
[0348] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.
[0349] In some embodiments, this application provides a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the methods described above.
[0350] In some embodiments, this application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for performing the methods described above.
[0351] In some embodiments, this application provides a computer program product including instructions for implementing the image processing method described above.
[0352] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0353] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0354] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0355] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other propagation signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0356] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0357] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0358] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0359] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: Acquire a first image, wherein the foreground object and background in the first image are from different images; The first image is input into a pre-trained first network model. The first network model redraws the first region in the first image according to the first redrawing parameters and redraws the second region in the first image according to the second redrawing parameters, and outputs the second image. The first redrawing parameter is different from the second redrawing parameter, and the first region includes the following: the outline of the foreground object and a plurality of pixels adjacent to the pixel corresponding to the outline of the foreground object; the second region is the region outside the first region in the first image. The first network model includes the following components: a first encoder, a first decoder, and a first diffusion model; The process by which the first network model processes the first image into the second image includes: The first image is input into the first encoder, and the first image is encoded into a first feature in the latent space; The first feature is input into the first diffusion model. Based on the preset noise, the feature corresponding to the first region in the first feature is redrawn according to the first redrawing parameters after a first number of iterations. Based on the preset noise, the feature corresponding to the second region in the first feature is redrawn according to the second redrawing parameters after a second number of iterations to obtain the second feature. The first number of iterations is equal to k*a, the second number of iterations is equal to k*b, a and b are the first redrawing parameters and the second redrawing parameters, respectively, and k is a preset parameter. The second feature is input into the first decoder, and the second feature is decoded into the second image.
2. The method according to claim 1, characterized in that, The plurality of pixels adjacent to the pixel corresponding to the outline of the foreground object include: pixels inside the foreground object whose shortest distance to the pixel corresponding to the outline of the foreground object is less than a first distance, and pixels outside the foreground object whose shortest distance to the pixel corresponding to the outline of the foreground object is less than a second distance.
3. The method according to claim 1, characterized in that, The acquisition of the first image includes: The foreground object in the third image is added to the fourth image, and the foreground object is relit to obtain the first image, wherein the lighting information of the foreground object in the first image is the same as the lighting information of the fourth image.
4. The method according to claim 1, characterized in that, The method further includes: The second image is subjected to super-resolution processing to obtain a fifth image, wherein the resolution of the fifth image is higher than that of the second image.
5. The method according to claim 3, characterized in that, The step of adding the foreground object from the third image to the fourth image and relighting the foreground object to obtain the first image includes: A first mask image of the foreground object is determined from the third image, the first mask image being used to represent the outline and shape of the foreground object; Normal prediction is performed on the third image to obtain normal information of the third image, and the normal information is used to reflect the orientation of the surface of the foreground object in the third image; Albedo prediction is performed on the third image to obtain albedo information of the third image, and the albedo information is used to reflect the color and material properties of the foreground object in the third image; Add the foreground object from the third image to the fourth image; Based on the first mask image, the normal information, the albedo information, and the illumination information of the fourth image, the foreground object in the fourth image is relit to obtain the first image.
6. The method according to claim 5, characterized in that, Adding the foreground object from the third image to the fourth image includes: The third image is downsampled, and the foreground object in the downsampled third image is added to the fourth image.
7. The method according to claim 1, characterized in that, The processing flow of the first feature in the first diffusion model includes: The first feature is input into the first diffusion model. Based on preset noise, the feature corresponding to the first region in the first feature is redrawn according to the first redrawing parameters. Then, based on the preset noise, the feature corresponding to the second region in the first feature is redrawn according to the second redrawing parameters. This includes: The preset noise is superimposed on the first feature according to the first redrawing parameters to obtain the first redrawing feature; For 1≤j≤k*(ab), in the j-th iteration, the j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output, where j is a positive integer; Corresponding to j = k*(ab) + 1, the preset noise and the first feature are superimposed according to the second redrawing parameters to obtain the intermediate feature, and the feature corresponding to the second region in the j-th redrawing feature is replaced with the feature corresponding to the second region in the intermediate feature to obtain the updated j-th redrawing feature; Corresponding to j=k*(ab)+1, in the j-th iteration, the updated j-th redrawing feature is input into the first diffusion model for redrawing, and the (j+1)-th redrawing feature is output. For j > k*(ab), in the j-th iteration, the j-th redraw feature is input into the first diffusion model for redrawing, and the (j+1)-th redraw feature is output. For j=k*a, the (j+1)th redrawn feature is taken as the second feature.
8. The method according to claim 7, characterized in that, The first network model further includes a text encoder; the training process of the first network model includes: The first network model is trained based on at least one first training data to obtain the trained first network model. The first training data includes a first training text and a first training image, wherein the first training text is related to the image features of the first training image, and the text encoder is used to obtain the text features of the first training text.
9. The method according to claim 8, characterized in that, The step of training the first network model based on at least one first training data includes: The first diffusion model in the first network model is trained based on the at least one first training data to obtain the trained first diffusion model. Based on the first diffusion model that has been trained, the first decoder in the first network model is trained according to the at least one first training data to obtain the first decoder that has been trained.
10. The method according to claim 8 or 9, characterized in that, The processing flow of the first network model for the first training data includes: The first training text in the first training data is input into the text encoder, and the text features of the first training text are output. The first training image from the first training data is input into the first encoder, and the first image is encoded as a third feature in the latent space. The third feature is input into the first diffusion model. Based on the preset noise, the feature corresponding to the third region in the third feature is redrawn according to the first redrawing parameters. Based on the preset noise, the feature corresponding to the fourth region in the third feature is redrawn according to the second redrawing parameters to obtain the fourth feature. The third region includes the following: the outline of the foreground object in the first training image, and multiple pixels adjacent to the pixels corresponding to the outline of the foreground object in the first training image. The fourth region is the region outside the third region in the first training image. The fourth feature is input into the first decoder, and the fourth feature is decoded into the first training result; Based on the difference between the text features of the first training text and the fourth feature, the network parameters of the first diffusion model are adjusted, and / or, based on the difference between the first training image and the first training result, the network parameters of the first decoder are adjusted.
11. The method according to claim 10, characterized in that, The processing flow of the third feature in the first diffusion model includes: The preset noise and the third feature are superimposed according to the first redrawing parameters to obtain the first training redrawing feature; For 1≤j≤k*(ab), in the j-th iteration, the j-th training redraw feature is input into the first diffusion model for redrawing, and the (j+1)-th training redraw feature is output, where j is a positive integer; Corresponding to j=k*(ab)+1, the preset noise and the third feature are superimposed according to the second redrawing parameters to obtain the training intermediate feature, and the feature corresponding to the fourth region in the j-th training redrawing feature is replaced with the feature corresponding to the fourth region in the training intermediate feature to obtain the updated j-th training redrawing feature; Corresponding to j=k*(ab)+1, in the j-th iteration, the updated j-th training redraw feature is input into the first diffusion model for redrawing, and the (j+1)-th training redraw feature is output. For j > k*(ab), in the j-th iteration, the j-th training redraw feature is input into the first diffusion model for redrawing, and the (j+1)-th training redraw feature is output. Corresponding to j=k*a, the (j+1)th training redraw feature is taken as the fourth feature.
12. The method according to claim 4, characterized in that, The step of performing super-resolution processing on the second image to obtain the fifth image includes: The second image is input into the pre-trained second network model, which outputs the fifth image.
13. The method according to claim 12, characterized in that, The second network model includes the following components: a second encoder, a second diffusion model, an image encoder, a control network, a repair network, and a second decoder; and, The processing flow of the second network model for the second image includes: The N first blocks of the second image are respectively input into the second encoder, and the N first blocks are respectively encoded in the latent space to obtain N fifth features, where N is a positive integer; The N first blocks are respectively processed by the image encoder to obtain N semantic features; The N fifth features are respectively input into the control network, and the corresponding N sixth features are output, wherein the sixth features are constrained by the corresponding semantic features; The N semantic features and the N sixth features are input into the second diffusion model. The corresponding seventh features are obtained according to each sixth feature and the corresponding semantic feature, and the N seventh features corresponding to the N first blocks are output. The N seventh features are combined into an eighth feature, and the eighth feature is upsampled in the latent space to obtain a ninth feature; The ninth feature is input into the repair network, and the data from different seventh features in the ninth feature are processed to obtain the tenth feature; The tenth feature is input into the second decoder, and the fifth image is output.
14. The method according to claim 13, characterized in that, The method further includes: The second decoder is trained based on at least one second training data to obtain the trained second decoder. The second training data includes a first training block and a first training semantic feature, wherein the first training semantic feature is the semantic feature of the first training block.
15. The method according to claim 14, characterized in that, The training process for the second decoder based on the second training data includes: The first training block is input into the second encoder, and the plurality of first training blocks are encoded in the latent space to obtain the corresponding eleventh feature; The eleventh feature is input into the control network, and the corresponding twelfth feature is output, wherein the twelfth feature uses the corresponding training semantic features as constraints. The first training semantic feature and the twelfth feature are input into the second diffusion model, and the thirteenth feature corresponding to the first training block is output. The thirteenth feature is input into the second decoder, and the second training result is output. Based on the second training result and the first training block, the network parameters of the second decoder are adjusted.
16. The method according to claim 15, characterized in that, The method further includes: The repair network is trained based on at least one third training data to obtain the trained repair network. The third training data includes a second training image and a third training image, wherein the third training image is an image obtained by reducing the resolution of the second training image.
17. The method according to claim 16, characterized in that, The process of training the repair network based on the third training data includes: The N second training blocks of the third training image are respectively input into the second encoder, and the N second training blocks are respectively encoded in the latent space to obtain N fourteenth features, where N is a positive integer; The N fourteenth features are input into the second diffusion model, and N fifteenth features corresponding to the N second training blocks are output, wherein the data volume of the fifteenth feature is greater than the data volume of the corresponding fourteenth feature; The N fifteenth features are combined into a sixteenth feature, and the sixteenth feature is upsampled in the latent space to obtain a seventeenth feature; The seventeenth feature is input into the repair network, and the data from different fifteenth features in the seventeenth feature are processed to obtain the eighteenth feature; The second training image is input into the second encoder, and the second training image is encoded in the latent space to obtain the nineteenth feature; Based on the difference between the eighteenth feature and the nineteenth feature, the network parameters of the repair network are adjusted.
18. The method according to claim 17, characterized in that, The N first blocks and the N second training blocks are obtained by dividing the blocks using a rule-based block division method or an overlapping block division method.
19. The method according to claim 3, characterized in that, The fourth image is an image generated based on text information input by the user, or a panoramic high dynamic range (HDR) image, or a non-HDR image.
20. The method according to claim 3, characterized in that, The illumination information of the fourth image is preset, pre-collected, or calculated based on the content of the fourth image.
21. The method according to claim 20, characterized in that, The fourth image is an HDR image, and the lighting information of the fourth image is the ambient light lighting information of the shooting scene in which the fourth image is located, which was collected in advance.
22. The method according to claim 19, characterized in that, The third image includes: a frame from a moving image, a frame from a still image, or a frame from a video. or, The fourth image includes: a frame of an animated image, a frame of a still image, or a frame of an image from a video.
23. The method according to claim 3, characterized in that, The third image includes: images in the preview interface of photo and video modes, images obtained by taking photos, images in videos obtained by video capture, images in video conferences, images in video calls, or images in desktop wallpapers.
24. The method according to claim 23, characterized in that, The number of foreground objects in the third image is multiple.
25. The method according to claim 3, characterized in that, The method further includes: Display a first interface, the first interface including a first control, the first control being used to select at least one of the third image and the fourth image; The user's operation on the first control is detected, and the third image and the fourth image are obtained based on the operation.
26. A readable medium, characterized in that, The readable medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1 to 25.
27. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, one of the processors of the electronic device, for performing the method of any one of claims 1 to 25.
28. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method of any one of claims 1 to 25.