Image processing methods, apparatus and electronic equipment
By acquiring the mask image and image restoration guide map of the LDR image, and using the diffusion model for image restoration, the problem of underexposure or overexposure in HDR images was solved, thus improving image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies suffer from underexposure or overexposure when generating high dynamic range (HDR) images, resulting in poor image quality.
By acquiring mask images and image inpainting guide maps corresponding to at least two low dynamic range (LDR) images, the area to be inpainted is determined, and a diffusion model is used for image inpainting to generate a high-quality HDR image.
It effectively avoids the quality problems of HDR images caused by underexposure or overexposure in LDR images, and improves the overall image quality, especially in cases of overexposure or underexposure.
Smart Images

Figure CN122134602A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an image processing method, apparatus, and electronic device. Background Technology
[0002] Currently, with the increasing demands for image quality in fields such as photography, surveillance, and vehicle imaging, High Dynamic Range Imaging (HDR) technology has become a hot topic due to its ability to reproduce a wide dynamic range and high clarity.
[0003] In related technologies, electronic devices can acquire multiple low dynamic range (LDR) images with different exposures, then align the pixels of the multiple LDR images with different exposures, and finally perform weighted fusion of the aligned pixel values to obtain an HDR image.
[0004] However, in the above method, when multiple LDR images with different exposures captured by the electronic device have underexposure or overexposure, the electronic device relies on the pixel values of the pixels in the LDR images with different exposures for weighted fusion. Therefore, the HDR image obtained by the electronic device through weighted fusion of the pixel values also has underexposure or overexposure. As a result, the image quality of the HDR image generated by the electronic device is poor. Summary of the Invention
[0005] The purpose of this application is to provide an image processing method, apparatus, and electronic device that can improve the image quality of HDR images when LDR images have underexposure or overexposure.
[0006] In a first aspect, embodiments of this application provide an image processing method, which includes: acquiring at least two mask images and an image restoration guide map corresponding to at least two first images, wherein the mask images are used to indicate the area to be restored in the first images, and the image restoration guide map is used to indicate the restoration target of the area to be restored, wherein the first images are LDR images; and performing image restoration processing on the at least two first images based on the at least two mask images and the image restoration guide map to obtain HDR images.
[0007] Secondly, embodiments of this application provide an image processing apparatus, comprising: an acquisition module and a processing module. The acquisition module is configured to acquire at least two mask images and an image restoration guide map corresponding to at least two first images. The mask images are used to indicate areas to be restored in the first images, and the image restoration guide map is used to indicate restoration targets for the areas to be restored. The first images are LDR images. The processing module is configured to perform image restoration processing on the at least two first images based on the at least two mask images and the image restoration guide map acquired by the acquisition module, to obtain high dynamic range (HDR) images.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, the region to be repaired in each LDR image is determined by at least two mask images corresponding to at least two LDR images, and the repair target of the region to be repaired is indicated by an image repair guide map. In the case of overexposure or underexposure in the LDR image, image repair is performed on at least two LDR images by the image repair guide map and at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the electronic device. Attached Figure Description
[0013] Figure 1 This is one of the flowcharts of an image processing method provided in the embodiments of this application;
[0014] Figure 2 This is a second flowchart of an image processing method provided in an embodiment of this application;
[0015] Figure 3 This is the third flowchart of an image processing method provided in the embodiments of this application;
[0016] Figure 4 This is the fourth flowchart of an image processing method provided in the embodiments of this application;
[0017] Figure 5 This is the fifth flowchart of an image processing method provided in the embodiments of this application;
[0018] Figure 6 This is the sixth flowchart of an image processing method provided in the embodiments of this application;
[0019] Figure 7 This is the seventh flowchart of an image processing method provided in the embodiments of this application;
[0020] Figure 8 This is the eighth flowchart of an image processing method provided in the embodiments of this application;
[0021] Figure 9 This is the ninth flowchart of an image processing method provided in the embodiments of this application;
[0022] Figure 10 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;
[0023] Figure 11 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;
[0024] Figure 12 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects. For example, a first object can be one or more, where "more" means at least two. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The terms "at least one" and "at least one of" in this application's specification refer to any one, any two, or a combination of two or more of the included objects. For example, "at least one of a, b, and c" can mean "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple, and multiple means at least two. Similarly, "at least two" means two or more, and its meaning is similar to "at least one". The identifiers in this application are text, symbols, images, etc., used to indicate information, and can use controls or other containers as carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.
[0028] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.
[0029] The image processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0030] The image processing method provided in this application can be applied to shooting scenarios, such as shooting HDR images.
[0031] Currently, with the increasing demands for image quality in fields such as photography, surveillance, and automotive imaging, HDR technology has become a hot topic due to its ability to reproduce a wide dynamic range. Traditional HDR involves acquiring multi-exposure LDR images, registering them, and then synthesizing them using methods such as weighted averaging and Laplacian pyramids. However, this approach suffers from issues such as detail loss and artifacts.
[0032] In recent years, Stable Diffusion's guided restoration technology has achieved high-fidelity restoration of damaged images thanks to its strong semantic understanding capabilities. However, most existing solutions are "pixel-level restoration first, then fusion", which does not take advantage of the latent space and lacks customized guidance strategies and dynamic mask optimization for HDR scenes, resulting in poor adaptability and high implementation costs.
[0033] Moreover, the above solution also has the following problems:
[0034] Loss of detail: Traditional pixel-level weighted blending cannot restore the texture of overexposed "white" or underexposed "black" areas, and can only perform simple interpolation smoothing, resulting in permanent loss of detail.
[0035] Fusion artifacts: Small errors in multi-exposure registration can easily cause edge blurring and ghosting; tonal differences between images exposed at different times can lead to color casts and tonal banding.
[0036] Poor adaptability: The dynamic range of the repaired pixel image does not match that of other exposed images, and there are no dedicated guiding conditions, making it difficult to guarantee semantic and tonal consistency.
[0037] Dynamic range compression: Pixel space fusion is limited by the dynamic range of LDR, resulting in compression of bright details and amplification of noise in dark areas of HDR images.
[0038] Insufficient versatility: It only adapts to specific exposure levels and static scenes, requires extensive retraining, and has weak compatibility with standard StableDiffusion models.
[0039] Based on the above-mentioned scenario applied in the embodiments of this application, the image processing method provided in the embodiments of this application determines the area to be repaired in each LDR image by using at least two mask images corresponding to at least two LDR images, and indicates the repair target of the area to be repaired by using an image repair guide map. In the case of overexposure or underexposure in the LDR image, the image repair guide map and at least two mask images are used to repair at least two LDR images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the electronic device.
[0040] The image processing method provided in this application is executed by an image processing device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an electronic device as an example to illustrate the image processing method provided in this application.
[0041] This application provides an image processing method. Figure 1 A flowchart illustrating an image processing method provided in an embodiment of this application is shown. Figure 1 As shown, the image processing method provided in this application embodiment may include the following steps 201 and 202.
[0042] Step 201: The electronic device acquires at least two mask images and image restoration guide images corresponding to at least two first images.
[0043] In this embodiment of the application, the mask image is used to indicate the area to be repaired in the first image, the image repair guide image is used to indicate the repair target of the area to be repaired, and the first image is an LDR image.
[0044] It is understood that at least two first images correspond to one image restoration guide image, and one first image corresponds to one mask image.
[0045] Optionally, in this embodiment, the at least two first images can be captured by an electronic device; or, the electronic device can download them through a third-party application. The specific method can be determined based on actual usage needs, and this embodiment does not impose any limitations.
[0046] Optionally, in this embodiment of the application, the exposure levels of the at least two first images are different.
[0047] For example, the exposure settings corresponding to each of the at least two first images can differ by 0.7EV to 2EV.
[0048] Optionally, in this embodiment, the image format of each of the at least two first images can be JPG, RAW, etc. The specific format can be determined according to actual usage requirements, and this embodiment does not impose any limitations.
[0049] Optionally, in this embodiment of the application, the electronic device can acquire at least two third images through an image sensor, and then perform image preprocessing on each of the at least two third images to obtain at least two first images.
[0050] For example, the image preprocessing described above can be pixel alignment and image filtering.
[0051] For example, taking the image preprocessing of a third image by an electronic device as an example, the electronic device can perform pixel alignment and image filtering on the third image using the following formula (1). Formula (1) can be specifically:
[0052] (1)
[0053] in, For pixels after pixel alignment and image filtering, Here, represents the coordinates of the pixel, i is the stride of the neighborhood window corresponding to the x-coordinate, and j is the stride of the neighborhood window corresponding to the y-coordinate. For a spatial Gaussian kernel, The grayscale similarity Gaussian kernel is used. To normalize the weights, I represents a 3×3 neighborhood window, and I is a third image.
[0054] It should be noted that the specific implementation process of step 201 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0055] Step 202: The electronic device performs image restoration processing on at least two first images based on at least two mask images and an image restoration guide image to obtain an HDR image.
[0056] In this embodiment of the application, the electronic device can perform image restoration processing on the areas to be restored in at least two first images using at least two mask images and an image restoration guide map to obtain an HDR image.
[0057] It should be noted that the specific implementation process of step 202 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0058] In the image processing method provided in this application embodiment, at least two mask images and an image restoration guide map corresponding to at least two first images are obtained. The mask images are used to indicate the areas to be restored in the first images, and the image restoration guide map is used to indicate the restoration targets of the areas to be restored. The first images are LDR images. Based on the at least two mask images and the image restoration guide map, image restoration processing is performed on the at least two first images to obtain a high dynamic range (HDR) image. In this scheme, the areas to be restored in each LDR image are determined by the at least two mask images corresponding to the at least two LDR images, and the restoration targets of the areas to be restored are indicated by the image restoration guide map. Thus, in the case of overexposure or underexposure in the LDR images, image restoration is performed on the at least two LDR images by the image restoration guide map and the at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR images leads to the same overexposure or underexposure in the generated HDR images, thereby improving the image quality of the HDR images generated by the electronic device.
[0059] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, the image processing method provided in this application embodiment further includes the following steps 301 and 302.
[0060] Step 301: The electronic device acquires the exposure confidence of pixels in each of at least two first images.
[0061] It should be noted that the specific implementation of step 301 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0062] Step 302: The electronic device generates at least two mask images based on each of the at least two first images and the exposure confidence of the pixels in each first image.
[0063] In this embodiment of the application, the electronic device can generate at least two mask images by using the pixel position of each pixel in at least two first images and the exposure confidence of each pixel in the first image.
[0064] It should be noted that the specific implementation process of step 302 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0065] Understandable, such as Figure 2 As shown, the execution of steps 301 and 302 can actually be performed before step 201.
[0066] In this embodiment, the electronic device can generate a mask image corresponding to each first image by using the exposure confidence of each pixel in each of at least two first images. This allows the electronic device to determine the area to be repaired in each first image based on the mask image corresponding to each first image, thereby improving the accuracy of the electronic device in determining the area to be repaired.
[0067] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, step 301 above can be implemented through step 301a below.
[0068] Step 301a: The electronic device calculates the exposure confidence of each pixel in each of the at least two first images based on the brightness value of each pixel in each of the first images and the average brightness value of each of the first images.
[0069] In this embodiment of the application, the electronic device can calculate the exposure confidence of each pixel in each first image by using the brightness value of each pixel in each of at least two first images, the average brightness value of each first image, and the neighborhood consistency coefficient.
[0070] For example, the range of the above-mentioned domain consistency coefficient can be 0.8 to 1.0.
[0071] For example, taking a pixel in a first image as an example, the electronic device can obtain the exposure confidence of a pixel in a first image through formula (2), which can be specifically:
[0072] (2)
[0073] in, Exposure confidence level for a single pixel. The brightness value of a pixel. The average brightness value of a first image. The variance of the brightness value of a single pixel. This is the domain consistency coefficient.
[0074] It should be noted that the first term in the above formula (2) is to avoid overexposed or underexposed areas by measuring how close the pixel brightness is to the global mean. The second term in the above formula (2) is to penalize areas with large saturation variance, that is, the larger the jump, the lower the reliability. The third term in the above formula (2) represents the smoothness of pixel brightness in the neighborhood. If the consistency is high, it is close to 1. The edge or noise area is low, further suppressing unreliable areas.
[0075] In this embodiment, the electronic device can calculate the exposure confidence of each pixel in each first image by using the brightness value of each pixel in each of at least two first images, the average brightness value of each first image, and the neighborhood consistency coefficient. Based on the exposure confidence, the device can determine the credibility of each pixel, thereby improving the accuracy of the electronic device in determining the exposure confidence of each pixel in each first image.
[0076] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 4 As shown, step 302 can be implemented through steps 302a and 302b below.
[0077] Step 302a: The electronic device sets the brightness value of the pixels in each of the at least two first images whose exposure confidence is greater than or equal to a preset threshold to a first value, and sets the brightness value of the pixels in each of the first images whose exposure confidence is less than the preset threshold to a second value.
[0078] In this embodiment of the application, the second value is used to indicate the pixel to be repaired.
[0079] For example, the first value can be 0 and the second value can be 1.
[0080] Optionally, in this embodiment, the aforementioned preset threshold can be preset by the electronic device or user-defined. The specific threshold can be determined based on actual usage needs, and this embodiment does not impose any limitations.
[0081] For example, the value range of the above-mentioned preset threshold can be 0.3-0.6.
[0082] Optionally, in this embodiment of the application, the electronic device can compare the exposure confidence of each pixel in each of the at least two first images with a preset threshold, thereby determining the pixels in each of the at least two first images whose exposure confidence is greater than or equal to the preset threshold and the pixels in each of the at least two first images whose exposure confidence is less than the preset threshold.
[0083] It is understandable that electronic devices can use the area composed of pixels to be repaired as the area to be repaired.
[0084] Step 302b: The electronic device generates at least two mask images based on the brightness value and position of each pixel in each of the at least two first images.
[0085] In this embodiment of the application, the electronic device can fill the pixel position of each pixel in each of at least two first images with the value of each pixel in each of at least two first images to generate at least two mask images.
[0086] For example, taking a mask image as an example, an electronic device can generate a mask image using the following formula (3), which is specifically:
[0087] (3)
[0088] in, For masked images, Exposure confidence for each pixel, This is a preset threshold.
[0089] Optionally, in this embodiment of the application, the electronic device can perform Gaussian feathering on the mask edge of the first mask image using the above formula (3) to obtain the above mask image.
[0090] In this embodiment, the electronic device can generate a mask image corresponding to each first image by using the exposure confidence of each pixel in each of at least two first images. This allows the electronic device to determine the area to be repaired in each first image based on the mask image corresponding to each first image, thereby improving the accuracy of the electronic device in determining the area to be repaired.
[0091] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 5 As shown, the "obtaining image repair guide maps corresponding to at least two first images" in step 201 above can be specifically implemented through steps 401 and 402 below.
[0092] Step 401: The electronic device acquires edge feature information of each pixel in each of at least two first images and image semantic information of each first image.
[0093] Optionally, in this embodiment of the application, the electronic device can extract edge feature information of each pixel in each of at least two first images using the Canny operator.
[0094] Optionally, in this embodiment of the application, the electronic device can obtain the image semantic information of each first image through the first model.
[0095] For example, the first model described above can be an artificial intelligence (AI) model or a neural network model. The specific model can be determined according to actual usage requirements, and this application embodiment does not impose any limitations.
[0096] For example, the neural network described above can be a ResNet neural network.
[0097] Step 402: The electronic device generates an image restoration guide map based on edge feature information and image semantic information.
[0098] In this embodiment of the application, the electronic device can generate an image restoration guide map using the edge feature information of each first image, the image semantic information of each first image, and the exposure normalization map of each first image.
[0099] For example, an electronic device can generate an image restoration guide map using the following formula (4), which is specifically:
[0100] (4)
[0101] in, Guided image for image restoration. The weights are the edge feature maps corresponding to at least two of the first images. For at least two edge feature maps corresponding to the first image, The weights are the image semantic graphs corresponding to at least two first images. For image semantic maps of at least two first images, The weights are the exposure normalization maps corresponding to at least two first images. For at least two first images, the exposure normalization plot is generated.
[0102] It can be understood that the aforementioned edge feature map includes edge feature information of each pixel in each of at least two first images; the aforementioned image semantic map includes image semantic information of each pixel in each of at least two first images.
[0103] In this embodiment, the electronic device combines an image restoration guidance map to enable the diffusion model to understand the image content. Edge feature information and image semantic information guide the diffusion model to restore reasonable details in local areas. Furthermore, the introduction of edge feature information in the image restoration guidance map forces the diffusion model to prioritize edge clarity and accuracy during image restoration, avoiding the generation of blurry or distorted boundaries. Image semantic information reflects the category of image content, helping the diffusion model understand the semantic attributes of different regions, thus maintaining semantic consistency during restoration. For example, a smooth gradient is generated in the sky area, and details are preserved in texture areas, while avoiding artifacts that do not conform to the scene content during restoration. Exposure information guides the model to focus on the brightness distribution of different exposure areas, ensuring a natural transition in brightness in the restored HDR image and avoiding unnatural restoration of overexposed or underexposed areas.
[0104] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 6 As shown, step 202 can be implemented through steps 202a and 202b below.
[0105] Step 202a: The electronic device performs image restoration processing on the area to be restored in each first image based on at least two mask images and an image restoration guide image, to obtain at least two second images.
[0106] In this embodiment of the application, the electronic device can perform image restoration processing on the region to be restored in each first image based on at least two mask images and an image restoration guide map using a diffusion model, thereby obtaining at least two second images.
[0107] For example, the above diffusion model can be the Stable Diffusion Inpainting model.
[0108] It should be noted that the specific implementation of step 202a above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0109] Step 202b: The electronic device performs pixel fusion on at least two second images to obtain an HDR image.
[0110] In this embodiment of the application, the electronic device can perform a weighted summation of each pixel in at least two second images to obtain an HDR image.
[0111] It should be noted that the specific implementation of step 202b above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0112] In this embodiment, an image restoration guide map indicates the restoration target of the area to be restored. Therefore, even when there is overexposure or underexposure in the LDR image, image restoration is performed on at least two first images using the image restoration guide map and at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the electronic device.
[0113] Optionally, in the embodiments of this application, combined with Figure 6 ,like Figure 7 As shown, step 202a can be implemented through steps 202a1 and 202a2 as described below.
[0114] Step 202a1: The electronic device performs latent space encoding on the at least two first images and the mask image corresponding to each first image.
[0115] In this embodiment of the application, the electronic device can perform latent space encoding on at least two first images and the mask image corresponding to each first image by means of a VAE latent space encoder.
[0116] For example, the electronic device can perform latent space coding on at least two first images using the following formula (5), which is specifically:
[0117] (5)
[0118] Where Z represents at least two first images encoded in the latent space, and I represents at least two first images. These are the pre-trained parameters in the VAE latent space encoder.
[0119] For example, the electronic device can perform latent space encoding on the mask image corresponding to each first image using the following formula (6), which is specifically:
[0120] (6)
[0121] in, M represents at least two first images encoded in the latent space, where M is the mask image corresponding to each first image.
[0122] Step 202a2: The electronic device performs guided repair processing on the regions to be repaired in at least two first images after latent space coding, based on the image repair guide map and the mask image after latent space coding, to obtain at least two second images.
[0123] In this embodiment of the application, the electronic device can input the latent space encoded image repair guide map, the mask image, and at least two latent space encoded first images into the above-mentioned diffusion model. The diffusion model then performs guided repair processing on the areas to be repaired in the at least two latent space encoded first images based on the latent space encoded image repair guide map and the mask image, thereby obtaining at least two second images.
[0124] For example, taking a second image as an example, the electronic device can perform reverse denoising repair using the following formula (7) to obtain a second image, specifically:
[0125] (7)
[0126] in, For the latent features of the k-th image at step t, For variance scheduling parameters, To extend the model, t is the step size.
[0127] In this embodiment, by introducing the generative capability of a diffusion model into the latent space, when encountering regions where all input images are of poor quality (i.e., regions marked by mask M), it no longer attempts to blend from existing bad pixels as in traditional schemes. Instead, it invokes prior knowledge learned by the model on large-scale datasets, such as "the sky should be gradient" and "the leaves should have texture," to regenerate these regions. This improves the image quality of HDR images generated by electronic devices.
[0128] Optionally, in the embodiments of this application, combined with Figure 6 ,like Figure 8 As shown, step 202b can be implemented through steps 202b1 to 202b3 as described below.
[0129] Step 202b1: The electronic device calculates the fusion weight corresponding to each pixel in each of the at least two second images based on the exposure confidence of each pixel in each of the at least two second images and the sum of the exposure confidence of pixels at the same position in the at least two second images.
[0130] In this embodiment of the application, the electronic device can perform a division operation on the exposure confidence of each pixel in each of at least two second images and the sum of the exposure confidence of pixels at the same position in at least two second images to obtain the fusion weight corresponding to each pixel in each second image.
[0131] For example, taking a single pixel as an example, the electronic device can obtain the fusion weight corresponding to that single pixel using the following formula (8), which can be specifically:
[0132] (8)
[0133] in, The fusion weight is the weight corresponding to a single pixel. The exposure confidence level for this single pixel. Sum the confidence scores of a single pixel at the same pixel location in at least two second images.
[0134] It can be understood that the numerator in the above formula (8) represents the exposure confidence of the k-th image at pixel (x, y). A higher value indicates better exposure quality and is more worthy of being retained in the fusion; the denominator represents the sum of the confidence of all n second images at the same pixel position, which serves as a normalization factor to ensure that all weights sum to 1. At any pixel position, the image with higher confidence will receive a larger fusion weight, thus dominating the final value of that pixel. This avoids the drawbacks of fixed weights and allows the fusion to adaptively adjust according to the local exposure quality.
[0135] Step 202b2: The electronic device performs a weighted summation of pixels in at least two second images based on the fusion weights to obtain a third image.
[0136] For example, an electronic device can obtain a third image using the following formula (9), which is specifically:
[0137] = (9)
[0138] in,( , For constraint coefficients, ), For the third image, This is a local contrast enhancement algorithm, while HSV is a color consistency constraint algorithm. This is an image obtained by pixel fusion without using fusion weights.
[0139] It should be noted that the first term in the above formula (9) performs a weighted summation of the latent features of each image according to the fusion weight to obtain the preliminary fusion result; then, the second term in the above formula (9) performs local contrast enhancement, such as Laplacian pyramid detail extraction: extracting high-frequency details (edges, textures) from the latent features and amplifying them; gradient domain operation: calculating the gradient field of the latent features, enhancing the gradient magnitude and then reconstructing; learnable filter: learning how to enhance contrast through a small convolutional network, etc., which can compensate for the detail blur that may be caused by weighted averaging, making the texture of the fusion result clearer and the three-dimensionality stronger; the third term in the above formula (9) performs color consistency constraint, constraining the hue and saturation of the fusion result to be consistent with the reference image, preventing unnatural color shifts introduced by multi-exposure fusion, such as the sky turning purple and the skin turning green.
[0140] Step 202b3: The electronic device decodes the third image to obtain an HDR image.
[0141] In this embodiment of the application, the electronic device can decode the third image using a VAE decoder and perform adaptive gamma mapping on the decoded third image to obtain the aforementioned HDR image.
[0142] For example, the electronic device can decode the third image through the VAE decoder using the following formula (10), which is as follows:
[0143] (10)
[0144] in, This is the decoded third image. These are the decoding parameters for the VAE decoder.
[0145] For example, the electronic device can perform adaptive gamma mapping on the decoded third image using the following formula (11), which is as follows:
[0146]
[0147] in, This is an HDR image.
[0148] Optionally, in this embodiment of the application, the electronic device can perform Laplacian pyramid detail enhancement on the image after adaptive gamma mapping and output a 16-bit floating-point HDR image.
[0149] In this embodiment, an image restoration guide map indicates the restoration target of the area to be restored. In the case of overexposure or underexposure in the LDR image, image restoration is performed on at least two first images using the image restoration guide map and at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the electronic device.
[0150] This application provides an image processing method. Figure 9 A flowchart illustrating an image processing method provided in an embodiment of this application is shown. Figure 9 As shown, the image processing method provided in this application embodiment may include the following steps 1 to 10.
[0151] Step 1: The electronic device acquires multi-frame, multi-exposure LDR images.
[0152] Step 2: The electronic device performs image preprocessing on the multi-frame, multi-exposure LDR images.
[0153] Step 3: The electronic device performs exposure statistics and feature extraction on the multi-frame, multi-exposure LDR images.
[0154] Step 4: The electronic device generates a guide map and an exposure difference mask map based on exposure statistics and feature extraction.
[0155] Step 5: The electronic device performs latent space coding on the exposure difference mask map and the multi-frame multi-exposure LDR image.
[0156] Step 6: The electronic device uses a diffusion model to guide the repair of the multi-frame multi-exposure LDR image after latent space coding, based on the exposure difference mask map and guide map after latent space coding, to obtain the repaired multi-frame multi-exposure LDR image.
[0157] Step 7: The electronic device performs latent space fusion on the repaired multi-frame multi-exposure LDR images to obtain the third image.
[0158] Step 8: The electronic device decodes the third image.
[0159] Step 9: The electronic device performs tone mapping and detail enhancement processing on the decoded third image.
[0160] Step 10: The electronic device outputs an HDR image.
[0161] Thus, detail restoration is enhanced: Latent space-guided restoration accurately restores semantic details in overexposed / underexposed areas, solving the detail loss problem of traditional methods. Artifact incidence is reduced: Latent space fusion, dynamic mask feathering, and cross-exposure interaction eliminate edge blur, ghosting, and tonal banding. Dynamic range is improved: Latent space fusion breaks through the dynamic range limitations of LDR, and adaptive gamma mapping balances wide dynamic range and display adaptation. It has strong versatility: Supports 3-5 stops of exposure and multiple input formats, adapts to static / slightly dynamic scenes, and is compatible with the standard Stable Diffusion model, requiring no extensive retraining. Realism is enhanced: Multi-condition guidance and tonal consistency constraints avoid "semantic drift," ensuring the restored area is highly consistent with the original image context.
[0162] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0163] It should be noted that the image processing method provided in this application embodiment can be executed by an image processing device. This application embodiment uses an image processing device executing the image processing method as an example to illustrate the image processing device provided in this application embodiment.
[0164] Figure 10 A schematic diagram of a possible structure of the image processing apparatus involved in an embodiment of this application is shown. For example... Figure 10 As shown, the image processing device 70 may include an acquisition module 71 and a processing module 72.
[0165] The acquisition module 71 is used to acquire at least two mask images and image restoration guidance maps corresponding to at least two first images. The mask images are used to indicate the areas to be restored in the first images, and the image restoration guidance maps are used to indicate the restoration targets of the areas to be restored. The first images are LDR images. The processing module 72 is used to perform image restoration processing on at least two first images based on the at least two mask images and image restoration guidance maps acquired by the acquisition module to obtain high dynamic range (HDR) images.
[0166] In one possible implementation, the acquisition module 71 is further configured to acquire the exposure confidence of each pixel in the first image. The processing module 72 is further configured to generate at least two mask images based on each first image and the exposure confidence of each pixel in the first image.
[0167] In one possible implementation, the processing module 72 is specifically used to set the brightness value of the pixel in each first image whose exposure confidence is greater than or equal to a preset threshold to a first value, and to set the brightness value of the pixel in each first image whose exposure confidence is less than the preset threshold to a second value, the second value being used to indicate the pixel to be repaired; and to generate at least two mask images based on the brightness value and position of each pixel in each first image.
[0168] In one possible implementation, the processing module 72 is specifically used to calculate the exposure confidence of each pixel in each first image based on the brightness value of each pixel in each first image and the average brightness value of each first image.
[0169] In one possible implementation, the acquisition module 71 is specifically used to acquire the edge feature information of each pixel in each first image and the image semantic information of each first image. The processing module 72 is specifically used to generate an image restoration guidance map based on the edge feature information and the image semantic information.
[0170] In one possible implementation, the processing module 72 is specifically used to perform image restoration processing on the region to be restored in each first image based on at least two mask images and an image restoration guide map to obtain at least two second images; and to perform pixel fusion on the at least two second images to obtain an HDR image.
[0171] In one possible implementation, the processing module 72 is specifically used to perform latent space encoding on at least two first images and the mask image corresponding to each first image; and based on the image repair guide map and the latent space encoded mask image, to perform guided repair processing on the regions to be repaired in the at least two latent space encoded first images to obtain at least two second images.
[0172] In one possible implementation, the processing module 72 is specifically used to calculate the fusion weight corresponding to each pixel in each second image based on the exposure confidence of each pixel in each second image and the sum of the exposure confidence of pixels at the same position in at least two second images; and to perform weighted summation of pixels in at least two second images based on the fusion weight to obtain a third image; and to perform decoding processing on the third image to obtain an HDR image.
[0173] In the image processing apparatus provided in this application embodiment, the region to be repaired in each LDR image is determined by at least two mask images corresponding to at least two LDR images, and the repair target of the region to be repaired is indicated by an image repair guide map. In the case of overexposure or underexposure in the LDR image, image repair is performed on at least two LDR images by the image repair guide map and at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the image processing apparatus.
[0174] The image processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0175] The image processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0176] The image processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0177] Optionally, such as Figure 11As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described image processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0178] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0179] Figure 12 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0180] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0181] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0182] The processor 110 is configured to acquire at least two mask images and image restoration guide maps corresponding to at least two first images, wherein the mask images are used to indicate the areas to be restored in the first images, and the image restoration guide maps are used to indicate the restoration targets of the areas to be restored, and the first images are low dynamic range (LDR) images; and based on the at least two mask images and the image restoration guide maps, perform image restoration processing on the at least two first images to obtain high dynamic range (HDR) images.
[0183] Optionally, in this embodiment of the application, the processor 110 is further configured to obtain the exposure confidence of each pixel in each first image; and generate at least two mask images based on each first image and the exposure confidence of each pixel in each first image.
[0184] Optionally, in this embodiment of the application, the processor 110 is specifically configured to set the brightness value of the pixel in each first image whose exposure confidence is greater than or equal to a preset threshold to a first value, and set the brightness value of the pixel in each first image whose exposure confidence is less than the preset threshold to a second value, the second value being used to indicate the pixel to be repaired; and generate at least two mask images based on the brightness value and position of each pixel in each first image.
[0185] Optionally, in this embodiment of the application, the processor 110 is specifically used to calculate the exposure confidence of each pixel in each first image based on the brightness value of each pixel in each first image and the average brightness value of each first image.
[0186] Optionally, in this embodiment of the application, the processor 110 is specifically used to obtain edge feature information of each pixel in each first image and image semantic information of each first image; and to generate an image restoration guide map based on the edge feature information and image semantic information.
[0187] Optionally, in this embodiment of the application, the processor 110 is specifically used to perform image restoration processing on the area to be restored in each first image based on at least two mask images and an image restoration guide map to obtain at least two second images; and to perform pixel fusion on the at least two second images to obtain an HDR image.
[0188] Optionally, in this embodiment of the application, the processor 110 is specifically used to perform latent space encoding on at least two first images and the mask image corresponding to each first image; based on the image repair guide map and the latent space encoded mask image, to perform guided repair processing on the regions to be repaired in the at least two latent space encoded first images to obtain at least two second images.
[0189] Optionally, in this embodiment of the application, the processor 110 is specifically configured to calculate the fusion weight corresponding to each pixel in each second image based on the exposure confidence of each pixel in each second image and the sum of the exposure confidence of pixels at the same position in at least two second images; perform weighted summation on the pixels in at least two second images based on the fusion weight to obtain a third image; and perform decoding processing on the third image to obtain an HDR image.
[0190] In the electronic device provided in this application embodiment, the area to be repaired in each LDR image is determined by at least two mask images corresponding to at least two LDR images, and the repair target of the area to be repaired is indicated by an image repair guide map. In the case of overexposure or underexposure in the LDR image, image repair is performed on at least two LDR images by the image repair guide map and at least two mask images to obtain a high-quality HDR image. This avoids the situation where overexposure or underexposure in the LDR image leads to the same overexposure or underexposure in the generated HDR image, thus improving the image quality of the HDR image generated by the electronic device.
[0191] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0192] The beneficial effects of the various implementation methods in this embodiment can be found in the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, they will not be repeated here.
[0193] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0194] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0195] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0196] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0197] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0198] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0199] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0200] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0201] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0203] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: At least two mask images and an image restoration guide map corresponding to at least two first images are obtained. The mask images are used to indicate the area to be restored in the first images, and the image restoration guide map is used to indicate the restoration target of the area to be restored. The first images are low dynamic range (LDR) images. Based on the at least two mask images and the image restoration guide map, image restoration processing is performed on the at least two first images to obtain a high dynamic range (HDR) image.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the exposure confidence of each pixel in the first image; The at least two mask images are generated based on each first image and the exposure confidence of the pixels in each first image.
3. The method according to claim 2, characterized in that, The process of generating the at least two mask images based on each first image and the exposure confidence of pixels in each first image includes: The brightness value of the pixels in each first image whose exposure confidence is greater than or equal to a preset threshold is set as a first value, and the brightness value of the pixels in each first image whose exposure confidence is less than the preset threshold is set as a second value. The second value is used to indicate the pixels to be repaired. Based on the brightness value and position of each pixel in each first image, the at least two mask images are generated.
4. The method according to claim 2, characterized in that, The step of obtaining the exposure confidence of each pixel in the first image includes: Based on the brightness value of each pixel in each first image and the average brightness value of each first image, the exposure confidence of each pixel in each first image is calculated.
5. The method according to claim 1, characterized in that, The step of obtaining image restoration guidance maps corresponding to at least two first images includes: Obtain the edge feature information of each pixel in each first image and the image semantic information of each first image; Based on the edge feature information and the image semantic information, the image restoration guidance map is generated.
6. The method according to claim 1, characterized in that, The step of performing image restoration processing on the at least two first images based on the at least two mask images and the image restoration guide map to obtain an HDR image includes: Based on the at least two mask images and the image restoration guide map, image restoration processing is performed on the region to be restored in each first image to obtain at least two second images; Pixel fusion is performed on the at least two second images to obtain an HDR image.
7. The method according to claim 6, characterized in that, Based on the at least two mask images and the image restoration guide map, image restoration processing is performed on the region to be restored in each first image to obtain at least two second images, including: Latent space encoding is performed on the at least two first images and the mask image corresponding to each first image; Based on the image restoration guide map and the mask image after latent space coding, the regions to be restored in at least two first images after latent space coding are subjected to guided restoration processing to obtain the at least two second images.
8. The method according to claim 6, characterized in that, The step of performing pixel fusion on the at least two second images to obtain an HDR image includes: Based on the exposure confidence of each pixel in each second image and the sum of the exposure confidence of pixels at the same position in the at least two second images, the fusion weight corresponding to each pixel in each second image is calculated; Based on the fusion weights, the pixels in the at least two second images are weighted and summed to obtain a third image; The third image is decoded to obtain the HDR image.
9. An image processing apparatus, characterized in that, The device includes: an acquisition module and a processing module; The acquisition module is used to acquire at least two mask images and an image restoration guide map corresponding to at least two first images. The mask images are used to indicate the area to be restored in the first images, and the image restoration guide map is used to indicate the restoration target of the area to be restored. The first images are low dynamic range (LDR) images. The processing module is used to perform image restoration processing on the at least two first images based on the at least two mask images and the image restoration guide map obtained by the acquisition module, so as to obtain a high dynamic range imaging HDR image.
10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the image processing method as described in any one of claims 1 to 8.