A semantic-guided extreme overexposure image inpainting method
By using a semantically guided diffusion model, the brightness imbalance is corrected pixel by pixel and the details of the brightness saturation area are restored using semantic information. This solves the problems of brightness imbalance and detail loss in extremely overexposed images and achieves visually and semantically consistent image restoration.
Patent Information
- Application Number
- CN202410939387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-07-15
AI Technical Summary
Existing technologies struggle to effectively address the overall brightness imbalance in extremely overexposed images and the recovery of details in saturated areas, especially when detail information is lost in saturated areas but semantic information remains, often leading to artifacts or the inability to recover details.
A semantically guided diffusion model is adopted, which corrects brightness imbalance pixel by pixel through an overexposure correction network and dynamic convolution. It combines semantic information to guide the information recovery of brightness saturation areas and uses the robust prior knowledge and semantic information of the diffusion model for repair.
It effectively corrects brightness imbalances, restores detailed information in saturated areas, maintains the visual appeal and semantic consistency of images, avoids blurry or unrealistic restorations, and produces visually superior results.
Smart Images

Figure CN119205573B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of extreme overexposure image restoration, and mainly to a semantically guided method for extreme overexposure image restoration. Background Technology
[0002] Overexposure is a common challenge in digital photography, often caused by overly bright environments or incorrect camera settings (such as exposure time and aperture). This problem can result in saturated areas of brightness in an image, especially in the sky, loss of detail, and the image appearing either too bright or completely white. This overexposure not only reduces the visual appeal of an image but also complicates post-processing.
[0003] Restoring extremely overexposed images presents two major challenges. First, these images are often affected by an overall brightness imbalance, requiring brightness correction to a more normal range. Second, the excessive brightness creates saturated areas in the image, where detail is lost; these areas typically appear white and lack surface information, making restoration difficult.
[0004] Recent methods attempt to address overexposure through exposure correction and high dynamic range (HDR) reconstruction. While these methods improve visual quality by adjusting contrast and brightness, they often introduce artifacts or fail to recover lost details in saturated areas of brightness under extreme overexposure. On the other hand, single-frame HDR struggles to handle extreme overexposure due to insufficient information in the original image, while multi-frame HDR requires multiple exposures, placing demands on the input image.
[0005] This paper addresses the challenge of restoring overall brightness imbalance in extremely overexposed images and recovering details in brightly saturated regions where surface texture information is lost but semantic information remains. A novel approach is proposed, employing an overexposure correction network and dynamic convolution to effectively correct image brightness while minimizing noise and artifacts. For brightly saturated regions with lost surface texture, a generative model such as a diffusion model is used to fully recover the lost information using the learned data distribution. Recognizing that these regions often retain meaningful semantic information, this semantic information is introduced to guide the diffusion model's recovery process. The proposed method avoids artifacts at the edges of brightly saturated regions and uses semantic information to guide the recovery, producing visually superior results. Summary of the Invention
[0006] Objective of the Invention: Addressing the problems existing in the aforementioned background technology, this invention provides a semantically guided diffusion model for the restoration of extremely overexposed images. The method has two main contributions: (i) This invention designs a novel framework specifically for restoring overexposed images. This framework can correct overexposed images pixel-by-pixel, resolving brightness saturation areas caused by over-illumination. Furthermore, utilizing robust prior knowledge of the diffusion model to restore brightness saturation areas improves the fidelity of the restoration process. (ii) This invention designs a unique approach that combines semantic information to provide effective guidance and constraints for image restoration. By guiding stable diffusion with semantic information, it ensures that the restored brightness saturation areas maintain the original semantic structure and consistency of the image. This method effectively prevents blurry or unrealistic restorations, producing visually appealing and semantically coherent results. Extensive experimental comparisons, including visual quality assessment, no-reference / full-reference image quality assessment, and human subjective investigation, demonstrate that the proposed method outperforms current state-of-the-art methods.
[0007] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows:
[0008] A semantically guided method for restoring extremely overexposed images includes:
[0009] Obtain a dataset containing overexposed images I. h The dataset is divided into a training set and a test set; the overexposed input image I is extracted pixel by pixel. h The brightness information is used to obtain brightness map B; an overexposed image brightness correction network model is constructed, and the overexposed images I in the training set are used as the basis for the model. h And the overexposed image I h The corresponding brightness map B is input into the overexposed image brightness correction network model; it is then determined whether the overexposed image I has been corrected. h The training process involves completing all training rounds to obtain a well-trained overexposed image brightness correction network model; the overexposed image I... h The corrected image I is obtained by feeding the trained overexposed image brightness correction network model into it. c For the corrected image I c After semantic segmentation and brightness extraction operations, a semantic map S and a mask M for the brightness saturation region are obtained. A pre-trained brightness saturation completion network model is then built to perform correction on the image I. c As input, a semantic graph S is introduced to complete the information of the mask M of the brightness saturation region using semantic information. At the same time, a brightness prompt word P is introduced to control the brightness of the information-completed part to avoid overexposure or underexposure. The brightness saturation region after information completion and the non-saturation region after brightness correction are fused according to the mask M of the brightness saturation region to obtain the final output Io.
[0010] Preferably, the overexposed images I in the dataset h Convert to grayscale image. Since each pixel value in the grayscale image represents its corresponding brightness value, the grayscale image can be directly used as the brightness image B.
[0011] Preferably, the overexposed image brightness correction network model includes seven concatenated convolutional layers, wherein the first and last layers use ordinary standard convolutional layers, and the second to sixth layers use brightness-aware variable convolutional layers. For activation functions, a ReLU activation function is concatenated after each of the first six convolutional layers, and a Tanh activation function is concatenated after the seventh convolutional layer.
[0012] Using the input image I h The luminance map B extracted and the standard convolution X are used. For the standard convolution X, the kernel generator generates two basic convolution kernels k1 and k2. The luminance map B is input into a luminance-aware variable convolution layer to predict guiding features, which are used to determine pixel weights w1 and w2. Under the guidance of weights w1 and w2, the basic kernels k1 and k2 are assigned to luminance-saturated or unsaturated regions.
[0013] Preferably, the overexposed image I h The image is input into the overexposed image brightness correction network model to obtain the corrected image I. c ;
[0014] For the corrected image I c Using exposure control loss L exp Color constancy loss L col Local smoothing term loss L tv Spatial consistency loss L spa The overexposed image brightness correction network model was trained in an end-to-end manner:
[0015]
[0016] Where N represents the number of non-overlapping local regions of size 16×16, k is a parameter, and Y k This enhances the average intensity value of local areas in the image. E refers to the gray level in the RGB color space, and is set to 0.6.
[0017]
[0018] Where R is the number of iterations and r is a parameter. and These represent the horizontal and vertical gradient operations, respectively. The graph represents the curve parameters, where ξ represents the color channel;
[0019]
[0020] J p J represents the average intensity value of p channels in the enhanced image. q This represents the average intensity value of q channels in the enhanced image, where μ represents a pair of channels (p, q).
[0021]
[0022] Where K is the number of local regions, i is the i-th local region, and Ω(i) represents the four adjacent regions centered on local region i: top, bottom, left, and right; Y i Y represents the average intensity value of the i-th local region in the enhanced image. j I represents the average intensity value of the i-th local region in the enhanced image. i I represents the average intensity value of the i-th local region in the input image. j This represents the average intensity value of the i-th local region in the input image.
[0023] Preferably, the overexposed image is input into a trained overexposed image brightness correction network model to obtain the corrected image I. c : to I c After semantic segmentation and brightness extraction operations, the corresponding semantic map S and the mask M of the brightness saturation region are obtained. The semantic map S is obtained by a general semantic segmentation model Oneformer. The extraction of the mask M of the brightness saturation region is to first calculate the brightness value of each pixel. When the brightness value is greater than the threshold H, the pixel is considered to belong to the brightness saturation region.
[0024] Preferably, the brightness saturation completion network model includes a variational autoencoder (VAE), a denoising U-Net, a semantic control network, and a CLIP text encoder: the variational autoencoder (VAE) processes the corrected image I. c To generate latent features, they are then subjected to random noise to generate noisy latent features z. t To ensure that the oversaturated portions of the recovered image are semantically consistent with the original features, the semantic map S is input into a semantic control network to extract semantic conditions c. This guides a pre-trained noise prediction network ∈ θ The noise is predicted and added to the latent features; furthermore, the CLIP text encoder extracts the text guidance condition y, which associates the cue words with the brightness of the generated image; the guidance noise prediction is represented as:
[0025]
[0026] Here, k is a factor that adjusts the guidance strength, and t refers to the number of time steps. This indicates that the guiding condition is empty;
[0027] After noise prediction, the DDIM sampler is used to accelerate the sampling process and gradually denoise to obtain the denoised features z0:
[0028]
[0029] η is usually set to 0, α and σ are predefined hyperparameters, and ∈ represents random Gaussian noise.
[0030] Preferably, to complete the restoration, the predicted z0 is passed through decoder D to obtain the image generated by the brightness saturation completion network model. Then, the final image I is synthesized using the following formula. o Guided by the mask M in the brightness saturation region, the corrected image I is processed region by region. c Images generated by the brightness saturation completion network model To merge:
[0031]
[0032] A terminal device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the semantically guided diffusion model for repairing extremely overexposed images according to any one of claims 1 to 7.
[0033] A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the semantically guided diffusion model for restoring extremely overexposed images according to any one of claims 1 to 7.
[0034] Beneficial effects:
[0035] (1) This invention designs a novel framework specifically for recovering overexposed images. This framework can correct overexposed images pixel by pixel, resolving brightness saturation areas caused by over-illumination. Furthermore, by utilizing robust prior knowledge inherent in the diffusion model to recover brightness saturation areas, the fidelity of the recovery process is improved.
[0036] (2) This invention designs a unique approach that combines semantic information to provide effective guidance and constraints for image restoration. By guiding stable diffusion with semantic information, it ensures that the restored brightness saturation regions maintain the original semantic structure and consistency of the image. This method effectively prevents blurry or unrealistic restorations, producing visually appealing and semantically coherent results.
[0037] (3) Through extensive experimental comparisons, including visual quality assessment, no-reference / full-reference image quality assessment and human subjective surveys, this invention demonstrates that the proposed method is superior to the current state-of-the-art methods. Attached Figure Description
[0038] Figure 1 This is a flowchart of the semantically guided diffusion model provided by the present invention for the restoration of extremely overexposed images;
[0039] Figure 2 This is an algorithmic framework diagram of the semantically guided diffusion model provided by the present invention for the restoration method of extremely overexposed images;
[0040] Figure 3 This is a structural diagram of the overexposed image brightness correction network model in the semantically guided diffusion model for the restoration method of extremely overexposed images provided by this invention.
[0041] Figure 4 This is a structural diagram of the brightness saturation completion network model in the semantically guided diffusion model for restoring extremely overexposed images provided by this invention. Detailed Implementation
[0042] The present invention will be further described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0043] This invention provides a semantically guided method for restoring extremely overexposed images, comprising the following steps:
[0044] A semantically guided method for restoring extremely overexposed images includes the following steps:
[0045] Step S1: Obtain the dataset, which includes overexposed image I. h The dataset is divided into a training set and a test set;
[0046] To effectively train the model, a dataset containing a large number of extremely overexposed images is needed. The training set is derived from the MIT-Adobe FiveK dataset, which includes 5000 raw RGB images and 5 professionally rendered RGB images for each raw image. Overexposed images were simulated in Photoshop using the Adobe Camera Raw SDK, with a relative EV of +1.5, and these were adjusted to extremely overexposed levels as input to the training set. It is worth noting that since the overexposure correction network is trained in an unsupervised manner, the training dataset consists only of these simulated overexposed images, without any real images. For testing, 124 overexposed images were randomly selected from the SICE dataset as the test set, including low-contrast images at different exposure levels and their corresponding high-quality counterparts.
[0047] Step S2: Extract the overexposed input image I pixel by pixel. h The brightness information is used to obtain the corresponding brightness map B;
[0048] Specifically, the input RGB image I h Convert to a grayscale image. In a grayscale image, the value of each pixel is its corresponding brightness value. Since each pixel value in a grayscale image represents its corresponding brightness value, the grayscale image is directly used as the brightness map B.
[0049] Step S3: Construct an overexposed image brightness correction network model, and use overexposed images I in the training set. h The corresponding brightness map B is input into the overexposed image brightness correction network model to obtain the corrected image I. c ;
[0050] Specifically, refer to Figure 3 The overexposed image brightness correction network model consists of seven concatenated convolutional layers. The first and last layers use standard convolutional layers, while layers 2-6 use brightness-aware variable convolutional layers. For activation functions, a ReLU activation function is concatenated after each of the first six convolutional layers, and a Tanh activation function is concatenated after the seventh convolutional layer.
[0051] Brightness-aware variable convolutional layers are designed to adaptively adjust the correction intensity based on varying overexposure conditions. This is achieved using the input image I... hThe extracted luminance map B and the standard convolution X are used to generate a luminance-aware variable convolution kernel by predicting a set of basic convolution kernels and guiding features. For the standard convolution X, the kernel generator generates two basic convolution kernels, k1 and k2. The luminance map B is fed into the convolutional layer to predict guiding features, which are used to determine pixel weights w1 and w2. Guided by these weights, the basic kernels k1 and k2 are assigned to luminance-saturated or unsaturated regions, enabling the luminance-aware variable convolution kernel to perform targeted feature extraction.
[0052] Step S4: Determine whether the overexposed image I has been properly processed. h Whether the training has been completed in all training rounds to obtain a well-trained overexposed image brightness correction network model;
[0053] For the corrected image I c Using exposure control loss L exp Color constancy loss L col Local smoothing term loss Ltv and spatial consistency loss L spa The overexposed image brightness correction network model was trained in an end-to-end manner.
[0054] For the brightness correction network model of overexposed images, the exposure control loss L is used. exp It controls the exposure of an image by measuring the deviation between the average brightness of a local area and the brightness required for correct exposure.
[0055]
[0056] Where N represents the number of non-overlapping local regions of size 16×16, Y is the average intensity value of local regions in the enhanced image, and E refers to the gray level in the RGB color space, set to 0.6.
[0057] In addition, to help the network accurately estimate illumination and parametric maps, a local smoothing loss term, denoted as L, is applied during the brightness correction process. tv :
[0058]
[0059] Where R is the number of iterations. and These represent horizontal and vertical gradient operations, respectively. A represents the curve parameter graph, and ξ represents the color channel.
[0060] To minimize the offset and ensure a balanced distribution of the RGB channels, a color constancy loss, denoted as L, is employed. col :
[0061]
[0062] J p J represents the average intensity value of p channels in the enhanced image. q This represents the average intensity value of q channels in the enhanced image, where μ represents a pair of channels (p, q).
[0063] In addition, using L spa This is used to measure changes in pixel value differences between adjacent regions after correction, aiming to maintain spatial consistency.
[0064]
[0065] Where K is the number of local regions, and Ω(i) are the four adjacent regions (top, bottom, left, right) centered on region i. Let Y and I represent the average intensity values of local regions in the enhanced version and the input image, respectively.
[0066] Step S5: Transfer the overexposed image I h The corrected image I is obtained by feeding the trained overexposed image brightness correction network model into it. c , to I c After semantic segmentation and brightness extraction operations, the corresponding semantic map S and the mask M of the brightness saturation region are obtained;
[0067] to I c After semantic segmentation and luminance extraction operations, the corresponding semantic map S and the luminance saturation region mask M are obtained. The semantic map S is obtained by a general semantic segmentation model, Oneformer. The extraction of the luminance saturation region mask M involves first calculating the luminance value of each pixel, and considering the pixel to belong to the luminance saturation region when the luminance value is greater than the threshold H.
[0068] Step S6: Construct a pre-trained brightness saturation completion network model and apply it to the corrected image I. c As input, a semantic graph S is introduced to complete the information of the brightness saturation region M using semantic information. At the same time, a brightness cue word S is introduced to control the brightness of the completed part to avoid overexposure or underexposure.
[0069] Specifically, refer to Figure 4 The brightness saturation completion network model includes a variational autoencoder (VAE), a denoising U-Net, a semantic control network, and a CLIP text encoder: the variational autoencoder (VAE) processes the corrected image I... c To generate latent features, they are then subjected to random noise to generate noisy latent features z. t To ensure that the oversaturated portions of the recovered image are semantically consistent with the original features, the semantic map S is input into a semantic control network to extract semantic conditions c. This guides a pre-trained noise prediction network ∈ θThe noise is predicted and added to the latent features; furthermore, the CLIP text encoder extracts the text guidance condition y, which associates the cue words with the brightness of the generated image; the guidance noise prediction is represented as:
[0070]
[0071] Here, k is a factor that adjusts the guidance strength, and t refers to the number of time steps. This indicates that the guiding condition is empty;
[0072] After noise prediction, the DDIM sampler is used to accelerate the sampling process and gradually denoise to obtain the denoised features z0:
[0073]
[0074] η is usually set to 0, and α and σ are predefined hyperparameters.
[0075] Step S7: Merge the saturated region after information completion with the unsaturated region after brightness correction according to the mask M of the saturated region to obtain the final output I. o To complete the restoration, the predicted z0 is used by decoder D to generate an image. Then, the final image is synthesized using the following formula, guided by the mask M of the brightness saturation region, by region, the corrected image I... c Images generated by the brightness saturation completion network model To merge:
[0076]
[0077] This method ensures that saturated areas are restored while preserving the integrity of unsaturated areas, resulting in a coherent and visually pleasing final image.
[0078] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method of semantic-guided extreme overexposure image inpainting, characterized in that, Comprising: acquiring a dataset, comprising within said dataset overexposed images I h dividing said dataset into a training set and a test set; Extracting the overexposed input image pixel by pixel h The brightness information is used to obtain brightness map B; an overexposed image brightness correction network model is constructed, and the overexposed images I in the training set are used as the basis for the model. h And the overexposed image I h The corresponding brightness map B is input into the overexposed image brightness correction network model; it is then determined whether the overexposed image I has been corrected. h The training process involves completing all training rounds to obtain a well-trained overexposed image brightness correction network model; the overexposed image I... h The corrected image I is obtained by feeding the trained overexposed image brightness correction network model into it. c For the corrected image I c After semantic segmentation and brightness extraction operations, a semantic map S and a mask M for the brightness saturation region are obtained. A pre-trained brightness saturation completion network model is then built to perform correction on the image I. c As input, a semantic graph S is introduced to complete the information of the mask M of the brightness saturation region using semantic information. At the same time, a brightness cue word P is introduced to control the brightness of the information-completed part to avoid overexposure or underexposure. The brightness saturation region after information completion and the non-saturation region after brightness correction are fused according to the mask M of the brightness saturation region to obtain the final output I. o ; The brightness saturation completion network model comprises a variational autoencoder VAE, a denoising U-Net, a semantic control network and a CLIP text encoder: the variational autoencoder VAE processes the corrected image I c to generate latent features, which are then processed with random noise to generate noisy latent features z t ; in order to ensure that the brightness oversaturated part of the recovered image is semantically consistent with the original features, the semantic graph S is input into the semantic control network to extract the semantic condition c, which guides the pre-trained noise prediction network ∈ θ to predict the noise added to the latent features; in addition, the CLIP text encoder extracts the text guide condition y, which links the prompt words to the brightness of the generated image; the guide noise prediction is expressed as: where k is a factor to adjust the guidance strength, t refers to the time step, represents that the guidance condition is empty; After noise prediction, a DDIM sampler is used to accelerate the sampling process and gradually denoise to obtain the denoised feature z0: where η is usually set to 0, α and σ are predefined hyperparameters, and ∈ represents random Gaussian noise.
2. The method of claim 1, wherein, The overexposed images I h converted into a grayscale image, since each pixel value in the grayscale image represents its corresponding luminance value, the grayscale image is directly used as the luminance map B.
3. The method of claim 1, wherein, The overexposed image brightness correction network model comprises 7 convolutional layers connected in series, wherein the first layer and the last layer use ordinary standard convolutional layers, the 2nd-6th layers use brightness-aware variable convolutional layers; for the activation function, a ReLU activation function is connected after each of the first 6 convolutional layers, and a Tanh activation function is connected after the 7th convolutional layer; Using the luminance map B extracted from the input image I h and the standard convolution X, for which the kernel generator generates two elementary convolution kernels k1 and k2; The brightness map B is input into the brightness-aware variable convolutional layer to predict the guide feature, which is used to determine the pixel weights w1 and w2; under the guidance of the weights w1 and w2, the basic kernels k1 and k2 are assigned to the brightness saturated or non-saturated regions.
4. The method of claim 1, wherein, the overexposed image I h input into the overexposed image brightness correction network model to obtain the corrected image I c ; to the corrected image I c , using an exposure control loss L exp , a color constancy loss L col , a local smoothness term loss L tv and a spatial consistency loss L spa , the overexposed image brightness correction network model is trained in an end-to-end manner: where N denotes the number of non-overlapping local regions of size 16x16, k is a parameter, Y k is the average intensity value of the local region in the enhanced image, E refers to the gray level in the RGB color space, and is set to 0.6; where R is the number of iterations, r is a parameter, and denote horizontal and vertical gradient operations, respectively, represents a parametric plot, and ξ denotes a color channel; where J p denotes the average intensity value of p channels in the enhanced image, J q denotes the average intensity value of q channels in the enhanced image, μ denotes a pair of channels (p, q); where K is the number of local regions, i is the i-th local region, Ω(i) is the four neighboring regions of top, bottom, left and right centered at the local region i; Y i represents the average intensity value of the i-th local region in the enhanced image, Y j represents the average intensity value of the j-th local region in the enhanced image, I i represents the average intensity value of the i-th local region in the input image, I j represents the average intensity value of the j-th local region in the input image.
5. The method of claim 1, wherein, input the overexposed image into the trained overexposed image brightness correction network model to obtain a corrected image I c : I c respectively undergo semantic segmentation and brightness extraction operations to obtain corresponding semantic graph S and mask M of the brightness saturated area; The semantic map S is obtained by a general semantic segmentation model Oneformer; the extraction of the mask M of the brightness saturated region is to first calculate the brightness value of each pixel, and when the brightness value is greater than the threshold H, the pixel belongs to the brightness saturated region.
6. The method of claim 1, wherein, To complete the restoration, the predicted z0 is passed through the decoder D to obtain the image generated by the brightness saturation completion network model D(z0); then the final image I is synthesized using the following formula o The corrected image I is fused regionally under the guidance of the mask M of the brightness saturation area c The image generated by the brightness saturation completion network model is fused:
7. A terminal device, characterized by comprising: The memory and the processor, the memory has a computer program stored therein, the computer program is executed by the processor, so that the processor executes the steps of the semantic guided extreme overexposed image repair method in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium has a computer program stored thereon, and the computer program is executed by the processor, so that the processor executes the steps of the semantic guided extreme overexposed image repair method in any one of claims 1 to 6.