Image defogging method based on feature and chromaticity collaborative guide diffusion model
By converting hazy images to the YCbCr space and extracting phase consistency and chromaticity correction features, combined with a conditional diffusion model, the halo artifacts and color distortion problems of existing dehazing methods are solved, achieving high-quality image restoration.
Patent Information
- Application Number
- CN202510999916.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for dehazing suffer from issues such as halo artifacts, color distortion, or excessive detail enhancement, making it difficult to meet the application requirements of complex real-world scenarios. Furthermore, deep learning methods are prone to insufficient detail recovery and color deviation.
A feature- and chromaticity-based guided diffusion model is adopted. By converting the hazy image from the RGB color space to the YCbCr color space, the phase consistency features of the luminance component and the chromaticity correction features of the chromaticity component are extracted. Combined with the conditional diffusion model, iterative calculations are performed to repair the luminance component. Finally, the result is converted back to the RGB space.
It significantly improves the dehazing effect of images, maintains overall brightness consistency while providing clearer texture and contour information, and achieves high-quality restoration of hazy images.
Smart Images

Figure CN120876301A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing and application technology, and in particular to an image dehazing method based on a feature and chromaticity co-guided diffusion model. Background Technology
[0002] With the rapid development of information technology, computer vision and image processing technologies have been widely applied in various fields such as autonomous driving, intelligent monitoring, and remote sensing image processing. However, haze is one of the key factors affecting image quality and subsequent analysis. Haze not only causes image blurring and reduced contrast, but it can also hide important details in images, seriously affecting the accuracy of image understanding and subsequent image analysis tasks. Therefore, how to effectively remove haze and restore image clarity and detail during image processing has always been a major problem that urgently needs to be solved in the field of computer vision.
[0003] In the development of image dehazing research, researchers have proposed many traditional enhancement methods based on atmospheric physical models and visual perception characteristics. Early classic algorithms include: histogram equalization based on statistical features, which improves the visualization of hazy images by stretching global and local contrast; gamma correction based on atmospheric scattering models, which uses the law of motion transformation to compensate for atmospheric transmittance attenuation; and white balance techniques to address the color cast problem in haze, which restore the intrinsic colors of the scene by adjusting color temperature parameters. In spatial domain processing, bilateral filtering combining spatial distance and color similarity is used to maintain edge smoothness in dehazing, while the Laplacian operator enhances detail sharpness by enhancing high-frequency components. In physics-driven methods, dark channel priors estimate transmittance through statistical priors, and homomorphic filtering achieves atmospheric light correction by separating the illumination reflection components. Although these traditional methods have clear physical interpretability, their performance is limited by the accuracy and reliability of manual modeling. When processing non-uniform haze and dense fog scenes, they are prone to problems such as halo artifacts, color distortion, or excessive detail enhancement, making it difficult to meet the application requirements of complex real-world scenarios.
[0004] With the development of deep learning technology, neural networks, with their superior feature learning capabilities, have gradually replaced traditional manual modeling methods. This approach not only effectively overcomes the limitations of previous methods in feature extraction and modeling, but also automatically mines complex patterns from massive amounts of data, significantly improving the adaptability and robustness of dehazing models to diverse scenes. However, these methods, whether estimating the parameters of atmospheric scattering models or obtaining clear images through end-to-end methods, are prone to insufficient detail recovery, color deviation, and artifacts, making it difficult to achieve ideal restoration results for hazy images. Summary of the Invention
[0005] In view of this, and to address the above shortcomings, it is necessary to propose an image dehazing method based on a feature and chromaticity collaborative guided diffusion model to improve the dehazing effect of hazy images and achieve high-quality restoration of hazy images.
[0006] In a first aspect, the present invention provides an image dehazing method based on a feature- and chromaticity co-guided diffusion model, comprising:
[0007] Obtain the foggy image to be dehazed;
[0008] The foggy image is converted from the RGB color space to the YCbCr color space to obtain the luminance and chrominance components.
[0009] Phase consistency is extracted from the brightness component to obtain phase consistency features;
[0010] The chromaticity components are chromaticity corrected to obtain chromaticity correction features.
[0011] The luminance component, phase consistency feature and chromaticity correction feature are fused and used as input to the conditional diffusion model for iterative calculation to repair the luminance component that is severely affected by haze degradation until the noise loss function of the conditional diffusion model meets the preset accuracy requirements.
[0012] The repaired luminance component is output and fused with the chromaticity correction feature, then converted to RGB space for restoration to obtain the restored dehazed image.
[0013] Preferably, the step of converting the hazy image from the RGB color space to the YCbCr color space includes:
[0014] According to the preset conversion method between YCbCr and RGB space, the RGB values of the foggy image are converted into YCbCr values; wherein, the YCbCr values include: the initially decoupled luminance component Y value, and the initially decoupled chrominance component Cb value and Cr value.
[0015] By introducing nonlinear functions and constructing nonlinear mapping relationships between Cb and Y values, as well as Cr and Y values, through deep learning, the final decoupled luminance component Y and chrominance components Cb′ and Cr′ are obtained.
[0016] Preferably, the constructed nonlinear mapping relationship includes:
[0017]
[0018] θ i (a,b)=Sigmoid(Conv(a))⊙Conv(F ab )⊙Sigmoid(Conv(b)),
[0019] F ab =Cat(a,b)⊙Softmax(Conv(AvgPool(Cat(a,b)))),
[0020] Where Y(x) represents the luminance component, Cb(x) and Cr(x) are the initially decoupled chrominance components, Cb′(x) and Cr′(x) are the finally decoupled chrominance components, Cat(·) represents tensor concatenation in the channel dimension, AvgPool(·) represents the average pooling operation, ⊙ represents element-wise multiplication, and a represents the network learning parameter θ. i (·,·) represents the luminance component of the input, b represents the chrominance component of the network learning parameters, and F ab For θ i The intermediate parameter of (·,·) is Sigmoid, which is the activation function.
[0021] Preferably, the step of extracting phase consistency of the luminance component includes:
[0022] Phase consistency features are extracted from the luminance component of an image using the following formula:
[0023]
[0024] Where o represents the direction indicator, ω o (x,y) is the adaptively selected weight function in the o-th direction, T is the noise compensation parameter, and Δφ no (x,y) and A no (x,y) represents the phase deviation function and magnitude of the final decoupled Y in the o-th direction and at the n-th scale. This indicates rounding down, and ε is a real number to prevent the denominator from being zero.
[0025] Preferably, the step of performing chromaticity correction on the chromaticity components includes: correcting the chromaticity components Cb′ and Cr′ respectively based on a self-attention mechanism to obtain chromaticity correction features. and
[0026] Preferably, the fusion of the luminance component, phase consistency feature, and chromaticity correction feature includes:
[0027] The luminance component Y and the phase coherence feature PC(Y) are fused using the following formula:
[0028] Y = Y⊙Sigmiod(Conv(PC(Y))).
[0029] Wherein, Y represents the result of fusing the luminance component Y and the phase consistency feature PC(Y);
[0030] Color correction features and spliced together into the fused result The input image for the conditional diffusion model is obtained above. Where H and W represent the height and width of the input image, the number of channels is 3, the first channel of x0 is Y, and the second and third channels are respectively... and
[0031] Preferably, the diffusion process of the conditional diffusion model includes:
[0032] In the forward diffusion process: given a data sample x0~q(x0) sampled from the real data distribution, a series of noise samples x0,x1,…,x following a Gaussian distribution are generated. t ,…,x T And add it to the input sample, as shown in the following formula:
[0033]
[0034] Where T represents the total diffusion steps, and β represents the variance of the noise growth level. t ∈(0,1) are hyperparameters that follow a Gaussian distribution, I is the identity matrix, and N(·) represents a normal distribution;
[0035] In the process of reverse diffusion: introducing conditional information The mean and variance of each step are predicted using a neural network, as shown in the following formula:
[0036]
[0037] Where, μ θ and Σ θ These are the mean and variance predicted by the conditional diffusion model, α t =1-β t , This refers to the noise predicted by the model.
[0038] Preferably, the noise loss function of the conditional diffusion model is:
[0039]
[0040] Among them, L cond Let ε be the noise loss function of the conditional diffusion model, and ε be the actual noise. Let be the mathematical expectation.
[0041] Preferably, the loss function used for image restoration is a composite loss function, which is expressed as follows:
[0042] Ltotal =αL Y +βL CbCr +γL rec ,
[0043]
[0044] Among them, L total For the composite loss function, L Y Let L be the brightness loss function. CbCr Let L be the chromaticity loss function. rec The reconstruction loss function uses α, β, and γ as weighting coefficients, with reference image X. GT In the YCbCr color space, each channel is Y GT ,Cb GT ,Cr GT , and For color correction features, The brightness component is the output of the conditional diffusion model. For the final restored image, The similarity calculation represents the feature, where ε is the real noise. θ This is the noise predicted by the conditional diffusion model.
[0045] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform any of the methods described in the first aspect.
[0046] As can be seen from the above technical solution, the image dehazing method based on a feature- and chromaticity co-guided diffusion model provided in this embodiment of the invention first acquires a foggy image to be dehazed, then converts the foggy image from the RGB color space to the YCbCr color space to obtain luminance and chromaticity components. Further, phase consistency features are extracted from the luminance components, and chromaticity correction features are obtained from the chromaticity components. The luminance components, phase consistency features, and chromaticity correction features are then fused and used as input to a conditional diffusion model for iterative computation to repair the luminance components severely affected by fog degradation. Finally, the repaired luminance components are output, fused with the chromaticity correction features, and converted back to the RGB space for restoration, resulting in the restored dehazed image. Therefore, this solution first converts the image from the RGB space to the YCbCr space. In the YCbCr space, fog has a smaller impact on the chromaticity channel, preserving more color information and thus helping to improve the restoration quality of the foggy image. Furthermore, since fog significantly affects the luminance channel in the YCbCr space, it can lead to the loss of luminance information. This solution improves the color fidelity of images by processing the chromaticity component through chromaticity correction. Simultaneously, it extracts phase consistency features from the luminance component to effectively capture important structural regions such as image edges, corners, and textures. Furthermore, the luminance component, phase consistency features, and chromaticity correction features are used together to guide a conditional diffusion model to repair the luminance channel, significantly improving the model's detail integrity and structural clarity in edge regions. This allows dehazed images to maintain overall luminance consistency while possessing clearer texture and contour information, thus significantly improving the dehazing effect of hazy images and achieving high-quality restoration of hazy images. Attached Figure Description
[0047] Figure 1 The flowchart illustrates an image dehazing method based on a feature- and chromaticity-coordinated guided diffusion model, as provided in this embodiment of the invention.
[0048] Figure 2 This is a model framework diagram of this solution.
[0049] Figure 3 The logical structure diagram of the color correction module provided for this solution.
[0050] Figure 4 This is a visual comparison of the present invention with different image dehazing methods on the Dense-Haze dataset.
[0051] Figure 5 This is a visual comparison of the present invention with different image dehazing methods on the NH-Haze dataset.
[0052] Figure 6 This is a visual comparison of the present invention with different image dehazing methods on the RTTS dataset.
[0053] Figure 7 This is a visual comparison of the present invention with different image dehazing methods on the SateHaze1k dataset. Detailed Implementation
[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] like Figure 1 As shown, this embodiment of the invention provides an image dehazing method based on a feature- and chromaticity co-guided diffusion model, which may include the following steps:
[0056] Step 101: Obtain the foggy image to be dehazed;
[0057] Step 102: Convert the foggy image from the RGB color space to the YCbCr color space to obtain the luminance component and chrominance component;
[0058] Step 103: Extract phase consistency from the luminance component to obtain phase consistency features;
[0059] Step 104: Perform colorimetric correction on the colorimetric components to obtain colorimetric correction features.
[0060] Step 105: The luminance component, phase consistency feature and chromaticity correction feature are fused and used as input to the conditional diffusion model for iterative calculation to repair the luminance component that is severely affected by haze degradation, until the noise loss function of the conditional diffusion model meets the preset accuracy requirements.
[0061] Step 106: Output the repaired luminance component, fuse it with the chromaticity correction feature, and convert it to RGB space for restoration to obtain the restored dehazed image.
[0062] In this embodiment, the image is first converted from RGB space to YCbCr space. In YCbCr space, haze has a smaller impact on the chroma channel, preserving more color information and thus improving the restoration quality of hazy images. Furthermore, since haze significantly affects the luminance channel in YCbCr space, it leads to the loss of luminance information. This solution processes the chroma component through chroma correction, improving the color fidelity of the image. Simultaneously, phase consistency features are extracted from the luminance component to effectively capture important structural regions such as image edges, corners, and textures. Furthermore, the luminance component, phase consistency features, and chroma correction features are used together to guide a conditional diffusion model to repair the luminance channel, significantly improving the model's detail integrity and structural clarity in edge regions. This allows the dehazed image to maintain overall luminance consistency while possessing clearer texture and contour information, thereby significantly improving the dehazing effect of hazy images and achieving high-quality restoration of hazy images. The implementation framework of this solution is as follows: Figure 2 As shown, the framework includes a color space learning module and a chromaticity correction module. The former aims to decouple the nonlinear coupling relationship between the luminance and chromaticity channels, while the latter is used to correct the chromaticity difference between hazy and clear images. The two work together to effectively restore the chromaticity information of the image.
[0063] For step 101, obtain the foggy image to be dehazed.
[0064] This step involves acquiring a hazy image that requires dehazing, specifically data in RGB space. This hazy image can be captured by a photograph or extracted from a video stream.
[0065] For step 102, the hazy image is converted from the RGB color space to the YCbCr color space to obtain the luminance and chrominance components; this step is performed by... Figure 2 The Color Space Learning (CCA) module is implemented in [the system / platform].
[0066] Although the YCbCr color space (where Y represents the luminance component, Cb represents the blue chromaticity component, and Cr represents the red chromaticity component) theoretically divides an image into a luminance axis (Y) and a chromaticity plane (CbCr), achieving a preliminary decoupling of luminance and color information, this separation is not complete in practical applications. Due to changes in illumination during imaging and the response characteristics of the photosensitive element, the luminance channel still has a certain influence on the chromaticity component, resulting in nonlinear coupling between the two. This coupling relationship is particularly evident in image degradation conditions, such as low-light environments and hazy environments, often manifesting as color distortion or color cast. Therefore, when converting the hazy image from the RGB color space to the YCbCr color space, it is further considered to utilize neural networks to deeply decouple the nonlinear coupling relationship between the YCbCr channels. For example, this can be achieved in the following way:
[0067] S11: Convert the RGB values of the foggy image to YCbCr values according to the preset YCbCr to RGB space conversion method; wherein, the YCbCr values include: the initially decoupled luminance component Y value, and the initially decoupled chrominance component Cb value and Cr value.
[0068] In this step, the following conversion formula can be considered for initial decoupling:
[0069]
[0070] x represents the pixel position. It can be seen that the luminance component Y(x) is a weighted average of the RGB three-channel values, while the chrominance components Cb(x) and Cr(x) are calculated by different linear combinations.
[0071] S12: Introduce nonlinear functions and construct nonlinear mapping relationships between Cb and Y values, and Cr and Y values through deep learning, to obtain the final decoupled luminance component Y and chrominance components Cb′ and Cr′.
[0072] In this step, we consider using the following nonlinear mapping relationship for final decoupling:
[0073]
[0074] θ i (a,b)=Sigmoid(Conv(a))⊙Conv(F ab )⊙Sigmoid(Conv(b)),
[0075] F ab =Cat(a,b)⊙Softmax(Conv(AvgPool(Cat(a,b)))),
[0076] Where Y(x) represents the luminance component, Cb(x) and Cr(x) are the initially decoupled chrominance components, Cb′(x) and Cr′(x) are the finally decoupled chrominance components, Cat(·) represents tensor concatenation in the channel dimension, AvgPool(·) represents the average pooling operation, ⊙ represents element-wise multiplication, and a represents the network learning parameter θ. i (·,·) represents the luminance component of the input, b represents the chrominance component of the network learning parameters, and F ab For θ i The intermediate parameter of (·,·) is Sigmoid, which is the activation function.
[0077] It is worth noting that this transition from linear transformation to deep learning-driven nonlinear transformation is one of the innovations of this module. It can not only capture the interference patterns of brightness on chromaticity in foggy images, but also adapt to the imaging characteristics of different scenes, achieving robust color space representation. Ultimately, the Cb'(x) and Cr'(x) output by this module no longer simply rely on the fixed weights in traditional formulas, but are dynamically adjusted according to the actual content of the image, more closely matching the true color distribution of the image and contributing to the overall improvement of image quality.
[0078] For step 103: Phase consistency extraction is performed on the luminance component to obtain phase consistency features. This step is... Figure 2 The phase congruency extraction module is implemented in the system.
[0079] Image degradation not only results in the loss of color information but also often involves the blurring of texture structures and the loss of detail. In this scheme, phase consistency is extracted not only as a structural feature representation but also as a structural prior used to guide the generation process of the diffusion model. First, using the two-dimensional PC function given by Kovesi, phase consistency is extracted from the image... Extracting the phase consistency feature map can be achieved through the following calculation formula:
[0080]
[0081] Where o represents the direction indicator, ω o (x,y) is the adaptively selected weight function in the o-th direction, T is the noise compensation parameter, and Δφ no (x,y) and A no (x,y) represents the phase deviation function and magnitude of the final decoupled Y in the o-th direction and at the n-th scale. This indicates rounding down, and ε is used to prevent real numbers with a denominator of 0, such as 0.0001.
[0082] In this embodiment, the calculated phase consistency map PC(x,y) can effectively capture important structural regions such as edges, corners, and textures without depending on image intensity, exhibiting good illumination invariance and structural sensitivity. Within this framework, this structural map not only represents detailed regions of the image but also serves as a conditional input to guide the generative network. Compared to using only a degraded brightness map, introducing structural priors significantly improves the model's detail integrity and structural clarity in edge regions, enabling the dehazed image to maintain overall brightness consistency while possessing clearer texture and contour information.
[0083] For step 104: Perform chromaticity correction on the chromaticity components to obtain chromaticity correction features. This step is performed by... Figure 2 The chroma correction module (CRM) is implemented in the system.
[0084] This step considers correcting the chromaticity components Cb′ and Cr′ separately based on a self-attention mechanism to obtain chromaticity correction features. and Specifically, after the nonlinear decoupling of the luminance and chrominance channels, the luminance and chrominance components of the image are characterized more independently and robustly. However, due to the complexity of environmental conditions during imaging, hazy images and their sharp reference images may still exhibit significant shifts in the chrominance components Cb' and Cr'. This shift often manifests as color imbalance or hue drift, leading to a decrease in the color fidelity of the restored image and further affecting the perceived quality by the human eye and the effectiveness of subsequent processing tasks.
[0085] To alleviate the aforementioned problems, a chromaticity correction module is introduced. This module takes the chromaticity features Cb' and Cr' output by the color space learning module as input, thereby achieving more refined local color correction. The CRM employs a symmetrical dual-channel design to process the two chromaticity components separately, effectively improving the model's adaptability and specificity in handling complex color degradation problems.
[0086] Specifically, this module takes two input feature maps, Cb' and Cr', as input, which are intermediate feature representations from different paths. First, the two feature maps are adjusted in dimension through 1×1 convolutions, and a LayerNorm operation is applied to stabilize the training process. Subsequently, positional encoding is introduced to enhance the features, giving the attention mechanism a certain degree of spatial awareness.
[0087] During the attention modeling phase, the two paths generate query (Q), key (K), and value (V) vectors respectively, but the attention computation for each path is performed across branches: Cb' in the left branch uses its own generated query vector Q. Cb The key and value vector K of the right branchCr V Cr Perform matrix operations to obtain the attention representation of Cr'; similarly, the attention representation of Cr' in the right branch is obtained using its own Q. Cr To perceive K on the left branch Cb V Cb This captures the complementary relationship between the two feature branches. The resulting attention features are activated by Sigmoid and then multiplied element-wise with the original input features to achieve a weighted adjustment. Next, a 3×3 convolution further integrates local contextual information and forms a residual connection with the original input, ultimately outputting the enhanced features. and The specific logic is as follows Figure 3 As shown.
[0088] In this embodiment, the CRM module guides the network to focus on regions with abnormal chromaticity distribution, achieving adaptive adjustment of pixel-level chromaticity response and avoiding color deviation and local distortion caused by traditional uniform chromaticity correction methods in complex images. Simultaneously, this module integrates contextual semantics and spatial structure information, ensuring that the correction results, while maintaining color consistency, fully preserve key details such as image edges and textures, thereby enhancing the perceptual rationality of color restoration. Through an attention-based chromaticity optimization strategy, CRM not only improves the model's color adaptability under complex lighting changes and non-uniform haze coverage conditions, but also makes the corrected chromaticity components statistically closer to a clear image while maintaining spatial coherence.
[0089] For step 105: The luminance component, phase consistency feature, and chromaticity correction feature are fused and used as input to the conditional diffusion model for iterative calculation to repair the luminance component severely affected by haze degradation, until the noise loss function of the conditional diffusion model meets the preset accuracy requirements. This step is... Figure 2 The ConditionalDiffusionModel module is implemented in [the system / platform].
[0090] In this step, phase consistency features are first combined with chroma-corrected color information as conditional input to guide the diffusion model to generate high-quality images. Specifically, structural information is derived from the phase consistency map calculated using a formula. Then, the convolutional kernel and activation function learn their weights and apply them to the luminance map Y, enabling the model to focus on feature regions related to phase information. The formula is as follows:
[0091] Y = Y⊙Sigmiod(Conv(PC(Y))).
[0092] Furthermore, color correction features and Directly concatenated into the formula The input image for the diffusion model is obtained above. Where H and W represent the height and width of the input image, and the number of channels is 3, meaning that the first channel of x0 is Y, and the second and third channels are respectively... and This allows the model to focus on both the detailed texture of the image and the overall color distribution during the sampling process, thereby achieving higher accuracy and consistency in image reconstruction.
[0093] Furthermore, conditional diffusion includes both forward diffusion and backward diffusion. Conditional diffusion models are a class of generative models that introduce external conditional information to regulate the backward diffusion generation process, thereby achieving controllable image generation. Its basic idea is to maintain the forward diffusion process q(x) without altering the forward diffusion process. 1:T Under the premise of |x0), learn a conditional backdiffusion process. in This is conditional information; q(·) represents the probability density function in the forward process, and p... θ (·) represents the probability density function that the network learns during the reverse process. Specifically:
[0094] In the forward diffusion process: given a data sample x0~q(x0) sampled from the real data distribution, a series of noise samples x0,x1,…,x following a Gaussian distribution are generated. t ,…,x T And add it to the input sample, as shown in the following formula:
[0095]
[0096] Where T represents the total diffusion steps, and β represents the variance of the noise growth level. t ∈(0,1) are hyperparameters that follow a Gaussian distribution, I is the identity matrix, and N(·) represents a normal distribution;
[0097] In the process of reverse diffusion: introducing conditional information The mean and variance of each step are predicted using a neural network, as shown in the following formula:
[0098]
[0099] Where, μ θ and Σ θ These are the mean and variance predicted by the conditional diffusion model, α t =1-β t , For the noise predicted by the model, condition It can be image features, labels, prior information, etc.
[0100] To enable the model to more accurately predict noise perturbations at each step of the diffusion process, we consider using Mean Square Error (MSE) as the training noise loss function to minimize the noise predicted by the model. The difference between the actual noise ε and the noise level. The noise loss function of the conditional diffusion model is:
[0101]
[0102] Among them, L cond Let ε be the noise loss function of the conditional diffusion model, and ε be the actual noise. Let be the mathematical expectation.
[0103] Step 106: Output the repaired luminance component, fuse it with the chromaticity correction feature, and convert it to RGB space for restoration to obtain the restored dehazed image.
[0104] In this step, the channels of the hazy image X in the YCbCr color space are defined as Y, Cb, and Cr, respectively. The reference image X... GT In the YCbCr color space, each channel is Y GT ,Cb GT ,Cr GT The chromaticity channels Cb and Cr are converted to color after passing through the chromaticity correction module. and The image generated by the conditional diffusion model after applying diffusion model to the brightness channel Y is Compared with color-corrected and After stitching, the images are converted back to RGB space to obtain the final restored image. Based on the above estimation, the composite loss function for the luminance channel can be constructed as follows:
[0105]
[0106] For color channels and The corresponding pixel-level loss is constructed as follows:
[0107]
[0108] Subject to FSIM c Inspired by the index, consider establishing a color space based on YCbCr. Indicators. Specifically, first calculate the images Y and Y'. GT Phase coherence PC(Y), PC(Y) GT ), and gradient magnitude G(Y), G(Y) GT ), thus obtaining S PC (x) and SG (x):
[0109]
[0110] S L (x)=S PC (x)⊙S G (x),
[0111] Where x represents the pixel position.
[0112] Chroma channels recovered using the model and Calculate S Cb (x) and S Cr (x):
[0113]
[0114] S C (x)=S Cb (x)⊙S Cr (x),
[0115] Where T1 = 0.85, T2 = 160, T3 = T4 = 200. Finally... The formula for calculating the indicator is:
[0116]
[0117] Where λ = 0.03, PC m (x)=max(PC(X),PC(X GT )).
[0118] To measure the overall difference between the restored image and the reference image, the constructed... To address this metric, consider introducing the following reconstruction loss function:
[0119]
[0120] Therefore, the final total loss can be composed of a weighted sum of the above three parts, as expressed below:
[0121] L total =αL Y +βL CbCr +γL rec ,
[0122] The weighting coefficients α = 10, β = 1, and γ = 10 are set empirically to balance the contributions of brightness, chromaticity, and structural reconstruction.
[0123] The following sections utilize Python software, with the model framework implemented using PyTorch on a single NVIDIA GeForce RTX 4090 GPU. Training and testing were performed on natural image datasets using NH-HAZE, Dense-Haze, and RTTS. Qualitative and quantitative comparisons were conducted with several mainstream dehazing methods, including DCP, DehazeNet, FFA-Net, PSD, Dehamer, FSDGN, RIDCP, WeatherDiff, FCDM, and DehazeDDPM. Furthermore, on the SateHaze1k remote sensing image dataset, the proposed model was also qualitatively and quantitatively compared with SkyGAN, GPD-Net, YOLY, SNSPGAN, and UME-Net models. Quantitative evaluations were performed using Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), FogAware Density Evaluator (FADE), and Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE).
[0124] Specifically, such as Figure 4 This paper presents comparative experiments on the Dense-Haze dataset, comparing the proposed method with different image dehazing models. Existing methods often suffer from significant structural loss when removing dense fog, resulting in unsatisfactory dehazing effects. In contrast, the proposed method better preserves image details and color information, producing dehazing results that visually approximate the clear reference image.
[0125] Figure 5 The paper presents comparative experiments on the NH-Haze dataset, comparing the invention with different image dehazing models. The invention demonstrates significant advantages in the accuracy of detail and color recovery, effectively restoring details in hazy areas.
[0126] Figure 6 The paper presents comparative experiments on the RTTS dataset, comparing the proposed invention with different image dehazing models. The images restored by the proposed invention do not exhibit uneven color patches or artifacts, and the generated images appear more realistic and natural.
[0127] Figure 7 The paper presents comparative experiments on the SateHaze1k dataset, comparing the invention with different image dehazing models. The invention effectively recovers the color information of images, providing better perceptual quality.
[0128] Table 1 below shows the quantitative comparison results of different methods on the Dense-Haze and NH-Haze datasets. From the various evaluation metrics, this invention demonstrates significant advantages.
[0129] Table 1
[0130]
[0131] Table 2 below shows the quantitative comparison results of different methods on the RTTS dataset. Compared with other methods, the present invention has a significant advantage in the absence of reference indicators.
[0132] Table 2
[0133]
[0134] Table 3 shows the quantitative comparison results of different methods on the SateHaze1k remote sensing image dataset. In remote sensing images where high accuracy is required, the present invention still shows good results.
[0135] Table 3
[0136]
[0137] In summary, this invention achieves high-quality restoration of hazy images by introducing phase consistency and chromaticity features and utilizing a conditional diffusion model. Specifically, it first reveals the impact of haze on various channels in different color spaces. In the RGB color space, there is a significant difference between the hazy image and the reference image. This is because the image's brightness is closely coupled with the three RGB color channels, affecting the detail and contrast of each channel. In contrast, in the HSV color space, because haze significantly affects the image's brightness, the mean square error of the hazy image and the reference image in the lightness (V) channel is large. Simultaneously, the hue (H) and saturation (S) channels are also affected by haze, resulting in a significant reduction in color vibrancy and an overall dull and distorted effect. Finally, in the YCbCr color space, haze has a significant impact on the lightness (Y) channel, leading to the loss of brightness information. However, compared to the HSV space, the YCbCr space shows certain advantages in processing chromaticity information. The blue chromaticity (Cb) and red chromaticity (Cr) channels are relatively stable, and haze has a smaller impact on these two channels, preserving more color information.
[0138] Next, due to the linear relationship between the YCbCr and RGB color spaces, a color space learning module is proposed. This module utilizes a convolutional neural network to decouple the luminance and chrominance channels, maximizing the alignment of the chrominance channels between the hazy image and the reference image. Furthermore, a chrominance correction module is designed to ensure semantic consistency between the blue and red chrominance channels through interaction. Finally, by guiding a diffusion model to repair the luminance channel using the corrected chrominance components and the phase consistency extracted from the image, high-quality restoration of the hazy image is successfully achieved.
[0139] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods in any of the embodiments of the specification.
[0140] This specification also provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method in any of the embodiments of the specification.
[0141] The device embodiments provided by the present invention are based on the same inventive concept as the method embodiments in this specification. For details, please refer to the description in the method embodiments of this specification, which will not be repeated here.
[0142] The modules or units in the device of this invention can be merged, divided, and deleted according to actual needs. The above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the invention. Those skilled in the art will understand that implementing all or part of the processes of the above embodiments and making equivalent changes according to the claims of this invention still fall within the scope of the invention.
Claims
1. An image dehazing method based on a feature- and chromaticity co-guided diffusion model, characterized in that, include: Obtain the foggy image to be dehazed; The foggy image is converted from the RGB color space to the YCbCr color space to obtain the luminance and chrominance components. Phase consistency is extracted from the brightness component to obtain phase consistency features; The chromaticity components are chromaticity corrected to obtain chromaticity correction features. The luminance component, phase consistency feature and chromaticity correction feature are fused and used as input to the conditional diffusion model for iterative calculation to repair the luminance component that is severely affected by haze degradation until the noise loss function of the conditional diffusion model meets the preset accuracy requirements. The repaired luminance component is output and fused with the chromaticity correction feature, then converted to RGB space for restoration to obtain the restored dehazed image.
2. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 1, characterized in that, The step of converting the hazy image from the RGB color space to the YCbCr color space includes: According to the preset conversion method between YCbCr and RGB space, the RGB values of the foggy image are converted into YCbCr values; wherein, the YCbCr values include: the initially decoupled luminance component Y value, and the initially decoupled chrominance component Cb value and Cr value. By introducing nonlinear functions and constructing nonlinear mapping relationships between Cb and Y values, as well as Cr and Y values, through deep learning, the final decoupled luminance component Y and chrominance components Cb′ and Cr′ are obtained.
3. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 2, characterized in that, The constructed nonlinear mapping relationship includes: θ i (a,b)=Sigmoid(Conv(a))⊙Conv(F ab )⊙Sigmoid(Conv(b)), F ab =Cat(a,b)⊙Softmax(Conv(AvgPool(Cat(a,b)))), Where Y(x) represents the luminance component, Cb(x) and Cr(x) are the initially decoupled chrominance components, Cb′(x) and Cr′(x) are the finally decoupled chrominance components, Cat(·) represents tensor concatenation in the channel dimension, AvgPool(·) represents the average pooling operation, ⊙ represents element-wise multiplication, and a represents the network learning parameter θ. i (·,·) represents the luminance component of the input, b represents the chrominance component of the network learning parameters, and F ab For θ i The intermediate parameter of (·,·) is Sigmoid, which is the activation function.
4. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 2, characterized in that, The step of extracting phase consistency of the luminance component includes: Phase consistency features are extracted from the luminance component of an image using the following formula: Where o represents the direction indicator, ω o (x,y) is the adaptively selected weight function in the o-th direction, T is the noise compensation parameter, and Δφ no (x,y) and A no (x,y) represents the phase deviation function and magnitude of the final decoupled Y in the o-th direction and at the n-th scale. This indicates rounding down, and ε is a real number to prevent the denominator from being zero.
5. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 2, characterized in that, The step of performing chromaticity correction on the chromaticity components includes: correcting the chromaticity components Cb′ and Cr′ respectively based on a self-attention mechanism to obtain chromaticity correction features. and 6. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 5, characterized in that, The fusion of the luminance component, phase consistency feature, and chromaticity correction feature includes: The luminance component Y and the phase coherence feature PC(Y) are fused using the following formula: Y = Y⊙Sigmiod(Conv(PC(Y))). Wherein, Y represents the result of fusing the luminance component Y and the phase consistency feature PC(Y); Color correction features and spliced together into the fused result The input image for the conditional diffusion model is obtained above. Where H and W represent the height and width of the input image, the number of channels is 3, the first channel of x0 is Y, and the second and third channels are respectively... and 7. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 6, characterized in that, The diffusion process in the conditional diffusion model includes: In the forward diffusion process: given a data sample x0~q(x0) sampled from the real data distribution, a series of noise samples x0,x1,…,x following a Gaussian distribution are generated. t ,…,x T And add it to the input sample, as shown in the following formula: Where T represents the total diffusion steps, and β represents the variance of the noise growth level. t ∈(0,1) are hyperparameters that follow a Gaussian distribution, I is the identity matrix, and N(·) represents a normal distribution; In the process of reverse diffusion: introducing conditional information The mean and variance of each step are predicted using a neural network, as shown in the following formula: Where, μ θ and Σ θ These are the mean and variance predicted by the conditional diffusion model, α t =1-β t , This refers to the noise predicted by the model.
8. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 7, characterized in that, The noise loss function of the conditional diffusion model is: Among them, L cond Let ε be the noise loss function of the conditional diffusion model, and ε be the actual noise. Let be the mathematical expectation.
9. The image dehazing method based on a feature- and chromaticity co-guided diffusion model according to claim 1, characterized in that, The loss function used for image restoration is a composite loss function, which is expressed as follows: L total =αL Y +βL CbCr +γL rec , Among them, L total For the composite loss function, L Y Let L be the brightness loss function. CbCr Let L be the chromaticity loss function. rec The reconstruction loss function uses α, β, and γ as weighting coefficients, with reference image X. GT In the YCbCr color space, each channel is Y GT ,Cb GT ,Cr GT , and For color correction features, The brightness component is the output of the conditional diffusion model. For the final restored image, The similarity calculation represents the feature, where ε is the real noise. θ This is the noise predicted by the conditional diffusion model.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in any one of claims 1-9.