Image defogging method based on combination of double-domain haze prior and color residual gating
By combining dual-domain haze prior and color residual gating in the DFCCNet network module, the problems of insufficient detail recovery and color distortion in single-image dehazing are solved, achieving more efficient image dehazing and color restoration effects.
Patent Information
- Application Number
- CN202510830721.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-28
AI Technical Summary
Existing image dehazing methods suffer from insufficient detail recovery and color distortion in single-image dehazing tasks, especially in high fog density or complex color scenes where it is difficult to balance detail recovery and color naturalness.
An image dehazing method based on dual-domain fog prior and color residual gating is adopted. Through the DFCCNet network module, combined with fog concentration estimation module, color correction residual gating module, multi-scale learnable fusion block and enhanced residual attention block, multi-scale frequency domain feature extraction and color correction are performed to achieve adaptive fog concentration estimation and color restoration.
It significantly improves image dehazing performance, enhances detail recovery and color fidelity, and improves PSNR, SSIM, and visual subjective quality evaluation results.
Smart Images

Figure CN120852226A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to an image dehazing method based on a combination of dual-domain haze prior and color residual gating. Background Technology
[0002] Atmospheric scattering and haze not only cause a sharp decrease in overall image contrast, blurring the edges of distant targets and causing loss of detail, but also cause significant color shifts, such as... Figure 1 As shown, issues such as blue or brownish-yellow tints are particularly critical in visual perception scenarios with extremely high real-time and security requirements, such as autonomous driving, video surveillance, and drone inspection. Since single-image dehazing cannot rely on scene depth or multi-view information, it can only depend on the brightness, color, and texture features contained in the image itself. How to strike a balance between not introducing artifacts and fully restoring details and color naturalness has always been a core challenge in this field.
[0003] Traditional physical prior methods, such as dark channel priors, color attenuation priors, and fog line priors, offer good interpretability but rely on explicit statistical assumptions, making them unsuitable for scenarios with high fog density or complex colors. Meanwhile, recent deep learning methods have achieved excellent quantitative performance through end-to-end training, but often neglect the impact of haze on frequency domain distribution and color baseline.
[0004] Meanwhile, due to the highly spatially non-uniform distribution of haze in real images, as well as the significant influence of object surface reflectivity and lighting conditions, estimating haze concentration by integrating multi-scale and multi-domain features has become an important research direction for improving defogging performance. Existing methods {he2010single, zhang2024density, wang2024dfr}, while simulating haze distribution well, still suffer from some issues related to insufficient color and contrast.
[0005] Therefore, a new solution is needed to address the above problems. Summary of the Invention
[0006] The purpose of this invention is to provide an image dehazing method based on a combination of dual-domain haze prior and color residual gating, which aims to solve the two major challenges of insufficient detail recovery and color distortion in single-image dehazing tasks.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an image dehazing method based on dual-domain haze prior and color residual gating, comprising at least the following steps:
[0008] S1: Construct the DFCCNet network module, which integrates a haze concentration estimation module, a color correction residual gating module, a multi-scale learnable fusion block, a Haar wavelet-based enhanced residual attention block, and an enhanced residual block. The haze concentration estimation module is HEBlock, the color correction residual gating module is CRG, the multi-scale learnable fusion block is MSLFB, the enhanced residual attention block is WRABlock, and the enhanced residual block is WRBlock.
[0009] S2: Design a loss function for the DFCCNet network module, use the loss function to train and optimize the DFCCNet network module, and then use the optimized DFCCNet network module for image dehazing.
[0010] S3: First, the haze concentration estimation module is used to obtain a more accurate haze density prior;
[0011] S4: Subsequently, wavelet transform is used to extract multi-scale frequency domain sub-band features and fuse them with spatial domain features. By using Haar wavelet-based enhanced residual attention blocks and enhanced residual blocks, the features are enhanced to represent haze layers and edge structures. Under the guidance of haze priors, the enhanced features effectively avoid misjudgment of background information and blurring of edge structures.
[0012] S5: The color correction residual gating module is used to solve the color shift problem that often occurs after halo removal. The color correction residual gating module effectively restores a realistic and natural color distribution through color correction and residual signal modulation.
[0013] S6: Finally, based on multi-scale learnable fusion blocks, adaptive interaction and reconstruction between different features are realized, thereby promoting visual friendliness and perception-oriented enhancement.
[0014] Furthermore, the fog concentration estimation module extracts the amplitude spectrum through two-dimensional Fourier transform and performs logarithmic compression, which can intuitively characterize the fog intensity changes at different spatial frequencies in the frequency domain, as shown in the following formula:
[0015]
[0016] Where (u,v) represents the frequency coordinates; I(x,y) represents the input image; e -j2π(ux+vy) It is a complex exponential function in the Fourier transform, representing the transform weight of the image in the frequency domain;
[0017] To reflect the overall frequency distribution characteristics, its amplitude spectrum is calculated and the dynamic range is compressed by taking its logarithm:
[0018] M(u,v)=log(1+|F(u,v)|)
[0019] log(1+|F(u,v)|) performs logarithmic compression on the amplitude, which is used to expand the energy of smaller frequency components in the frequency domain and enhance the visualization of low-frequency components.
[0020] M(u,v) emphasizes the distribution of low-frequency and mid-to-high-frequency textures in the image, and has a significant ability to distinguish areas of reduced contrast and weakened edges caused by fog.
[0021] Furthermore, in order to enhance the resolution of local spatial details, a two-dimensional Haar wavelet transform is applied to the input image, and the representation of haze layer and edge structure is enhanced by Haar wavelet-based enhanced residual attention blocks and enhanced residual blocks.
[0022] After undergoing a two-dimensional Haar wavelet transform, the input image can be separated to reveal detailed variations in different directions (horizontal, vertical, and diagonal) and scales, thus supplementing spatial information that Fourier transform cannot locate. The wavelet decomposition result is divided into four sub-bands, as shown in the following formula:
[0023] W(k)(x,y)=(I*ψ(k))(x,y),k=1,2,3,4
[0024] Where * denotes convolution operation; I denotes the spatial domain representation of the input image; ψ(k) denotes the k-th wavelet basis function, used to extract features of different scales and orientations from the image;
[0025] By using spatial downsampling and upsampling, the wavelet features are restored to the same resolution as the input image, so that they can be stitched together and fused with the Fourier amplitude spectrum features, as shown in the following formula:
[0026] Φ=Concat(M,W(1),W(2),W(3),W(4))
[0027] By combining the explicit attention map A(x,y) generated by the lightweight convolutional attention module, adaptive weighting of multi-domain features is achieved, resulting in the final fog concentration estimation map D^(x,y).
[0028] D^(x,y)=F fusion (Φ(x,y)·A(x,y))
[0029] Where F fusion This indicates a fused convolutional network.
[0030] Furthermore, the color correction residual gating module combines dynamic routing self-attention (HDR) with wavelet self-attention multi-scale color correction, and introduces explicit residual signals and gating mechanisms to more accurately adaptively adjust color deviations.
[0031] Furthermore, a hierarchical dynamic routing module is built using dynamic routing self-attention to adaptively fuse multi-scale context information;
[0032] Input feature X is processed by average pooling and convolution at three different scales to generate scaled features:
[0033]
[0034] in This indicates that average pooling is performed on X, and each F... i Upsampled to the original resolution, and then aligned by filling or truncating the channels before being stitched together with X;
[0035] Next, route weights are generated using two 1×1 convolutional layers and Softmax:
[0036] W r =Softmax(Conv 1×1 (GELU(Conv 1×1 ([X,F1,F2,F3])))).
[0037] W r After reshaping to [B,C,H,W], the features at each scale are weighted accordingly and fused by adding the residuals:
[0038]
[0039] Among them, from W r The extracted i-th routing weight determines the i-th scale feature F. i The weighting coefficients; Proj i This is a projection operation.
[0040] Furthermore, the wavelet self-attention multi-scale color correction (MWCC), given feature X, calculates the low-frequency component LL and the self-attention feature X, respectively. ` The preliminary corrected feature F is then obtained using the following formula. * :
[0041] F = GELU(Conv) 1×1 (Concat(X`,Up(LL(X)))))
[0042] F * =GELU(Conv 1×1 Conv 3×3 (F))+F
[0043] Where Up(·) represents upsampling;
[0044] The residual gating is used to avoid overcorrection and adaptively adjust the color residual;
[0045] First, global average pooling (GAP) is applied to the input feature X, followed by two 1×1 convolutions and ReLU activation. Then, the residual gate weights G are obtained through the Sigmoid function.
[0046] G=σ(Conv 1×1 (ReLU(Conv 1×1 (GAP(X)))))
[0047] Where σ represents the Sigmoid function and GAP(·) represents the pooling layer;
[0048] Finally, the gated weighted color residual is fused with the original feature X to obtain the output feature.
[0049]
[0050] Where F * This refers to the result after preliminary correction by MWCC, where ⊙ represents element-wise multiplication.
[0051] Furthermore, the multi-scale learnable fusion block integrates multi-scale spatial attention (MSA) and pyramid channel attention (MCA), and introduces a learnable fusion module (LFM) for further calibration at the channel level, allowing automatic suppression of redundant information and enhancement of important signals to ensure feature complementarity.
[0052] Furthermore, the pyramid channel attention extracts channel attention at multiple downsampling scales to perceive global statistics under different receptive fields, for a given scale set b. i Average pooling and channel attention are performed to obtain attention features at each scale. These features are then upsampled back to the original resolution and stitched together.
[0053]
[0054] Where F is the feature map, CALayer(·) represents the channel layer, and Pool(·) represents the pooling layer;
[0055] The multi-scale spatial attention fusion of spatial attention outputs under the receptive fields of multiple convolutional kernels can focus on both subtle textures and capture larger structures.
[0056] Using a set of spatial attention blocks with different kernel sizes, generate respectively And weight the original features:
[0057]
[0058] Where S k This represents the feature results generated by different kernel sizes;
[0059] S will then k The concatenation is performed along the channel dimension, and then fused with 3×3 convolution, followed by Tanh activation and residual addition.
[0060] F MSA =tanh(Conv 3×3 (Concat(S 1 ,S 3 ,S 5 ,S 7 )))+F
[0061] Here, tanh is the activation function.
[0062] Furthermore, the learnable fusion module adaptively corrects the initial fusion features and attention fusion features to reduce information loss;
[0063] First, in the channel dimension, the initial fusion features F0 and F... mix splicing:
[0064] F cat =Concat(F0,F mix )
[0065] Then, a fusion-gated mapping is constructed using two layers of 3×3 convolutions:
[0066]
[0067] Where b1 and b2 are both bias terms in the convolution operation; r is the channel compression ratio, i.e., the bottleneck ratio; σ is the Sigmoid activation function; These are the parameters of the convolutional layer; the fusion weights P∈[0,1] B×2C×H×W By weighting the spliced features element by element, the final fused representation is obtained:
[0068] F LFM =P⊙F cat
[0069] Where ⊙ represents element-wise multiplication;
[0070] Finally, to maintain consistency with subsequent channel dimensions, one more convolution is performed.
[0071]
[0072] Here, b3 is the bias term in the convolution operation.
[0073] Furthermore, the loss function employs a weighted combination of pixel-level reconstruction loss L1 and contrast preservation loss (ContrastLoss) to simultaneously ensure the brightness accuracy and local contrast of the restored result.
[0074] Pixel reconstruction loss, denoted as the network prediction result. The corresponding sharp image is denoted as J, with image size C×H×W and batch size B. The loss is defined as follows:
[0075]
[0076] Contrast Preservation Loss: To further enhance the fidelity of local structure and texture, a contrast preservation loss is introduced, denoted as . Contrast preservation loss achieves explicit constraints on detail contrast by measuring the difference in contrast features between the predicted and ground truth maps within pixel neighborhoods or multi-scale windows. Formally, it can be expressed as:
[0077]
[0078] in The comparison operator is defined as n traversing all spatial locations or sample moments, with a total of N terms.
[0079] The aforementioned pixel-level reconstruction loss and contrast preservation loss are expressed using hyperparameters λ1 and λ2. CR Weighted combination yields the total loss:
[0080]
[0081] Set λ1 = 1.0 and λ CR =0.1.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] This invention proposes the DFCCNet (Dual-domain Fusion and Color Correction Network), which extracts multi-scale frequency domain features by introducing Haar wavelet transform and fuses them with temporal convolution features, thus supplementing low-frequency fog layer and high-frequency detail information, which is beneficial for accurately distinguishing fog regions from structural regions.
[0084] Furthermore, the proposed fog concentration estimation module is the HazyEstimateBlock, or HEB module, which integrates Fourier-wavelet domain and attention mechanism. By jointly capturing global frequency features and local detail textures, it achieves more accurate spatial adaptive fog concentration estimation.
[0085] Furthermore, a Color Residual Gating (CRG) module is proposed, which introduces wavelet low-frequency components as priors, and can directly extract the overall brightness and color information from the original image;
[0086] Furthermore, a multi-scale learnable fusion block is proposed, and a learnable fusion module (LFM) is proposed within the multi-scale learnable fusion block to further adaptively correct the initial fusion features and attention fusion features, thereby reducing information loss.
[0087] Compared to mainstream dehazing and enhancement methods, our method shows significant improvements in PSNR, SSIM, and visual subjective quality assessment, fully demonstrating its superiority in single-image dehazing and color restoration tasks. Attached Figure Description
[0088] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0089] Figure 1 This is a schematic diagram illustrating the specific impact of haze concentration on images according to the present invention;
[0090] Figure 2 This is a comparative schematic diagram of the processing of the present invention;
[0091] Figure 3 This is a schematic diagram illustrating the quantitative analysis comparison of the present invention;
[0092] Figure 4 This is a flowchart of the single-image dehazing process of DFCCNet in this invention;
[0093] Figure 5 This is a schematic diagram of wavelet self-attention multi-scale color correction according to the present invention;
[0094] Figure 6 This is a schematic diagram of the multi-scale learnable fusion block of the present invention;
[0095] Figure 7 This is a schematic diagram illustrating the qualitative analysis of the present invention on the sots-indoor dataset;
[0096] Figure 8 This is a schematic diagram illustrating the qualitative analysis of the present invention on the sots-outdoor dataset;
[0097] Figure 9 This is a schematic diagram of the qualitative analysis of the present invention on the Haze4K dataset;
[0098] Figure 10 This is a schematic diagram of the qualitative analysis of this paper on the RW2AH dataset;
[0099] Figure 11This is a schematic diagram illustrating the qualitative analysis of our results after adding CRG and the basic network in the R, G, and B channels. Detailed Implementation
[0100] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0101] This invention first analyzes the specific impact of haze concentration on images, such as... Figure 1 As shown, haze not only blurs but also obscures fine structural details. Furthermore, in high-frequency images, haze suppresses edge content while shifting the overall color baseline, resulting in significant color distortion. Therefore, this invention proposes HazyEstimateBlock (HEB), which integrates Fourier wavelet domain information and an attention mechanism. Specifically, HEB extracts global frequency features through Fourier transform to perceive large-scale haze structures, while wavelet transform captures local multi-scale details, facilitating fine-grained haze estimation in textured regions. In addition, a channel-space joint attention mechanism is introduced to adaptively enhance key haze features and suppress redundant information, thereby improving haze density estimation in complex scenes, such as… Figure 2 As shown, from left to right: fog image, DCP fog estimation, DADMNet, and our proposed method.
[0102] like Figure 3 As shown, from left to right, the images represent the ground truth image, the dehazed image generated by DEANet, and the quantitative analysis of the RGB three channels. Experimental results demonstrate that this method improves the recognition capability of haze estimation and provides effective prior guidance. The estimation presented in this invention better matches the haze distribution; it preserves structural details without significant distortion or artifacts, providing a more accurate representation of the haze state for subsequent dehazing.
[0103] Inspired by DEANet, this invention further introduces Haar wavelet transform into DEConv, extracting multi-scale frequency features by explicitly separating LL (low-frequency) and LH / HL / HH (high-frequency) components and fusing them with spatial features. This helps to supplement low-frequency haze layers and high-frequency detail cues, allowing for more accurate differentiation between haze and structural regions. Subsequently, to address the common color shift problem after deinking, a Color Residual Gate (CRG) module is designed. This module constructs a residual color path and introduces a gating mechanism to adaptively adjust the enhancement and suppression of color features, thereby improving the realism and consistency of color restoration. Based on the Content Guided Attention (CGA) mechanism, multi-scale attention enhancement extensions are performed, and a learnable fusion module (LFM) is designed to handle adaptive interactions and reconstructions between different features, enhancing visual friendliness and perceptual perception capabilities. Finally, based on this, a novel end-to-end dehazing network and dual-domain fusion color correction network (DFCCNet) is proposed. This network is a wavelet enhancement network that fuses spatial and frequency domain features to extract and fuse local and global detail information. It also combines this with a haze density map to obtain richer features, while introducing a color correction mechanism to adaptively restore a realistic and natural color distribution. Through dual-domain feature fusion and color consistency constraints, the proposed method achieves complementary optimization in enhancing image details and color fidelity, significantly improving dehazing performance in complex haze scenes.
[0104] The details are as follows:
[0105] An image dehazing method based on a combination of dual-domain haze prior and color residual gating includes at least the following steps:
[0106] S1: Construct the DFCCNet network module. The DFCCNet network module integrates a haze concentration estimation module, a color correction residual gating module, a multi-scale learnable fusion block, a Haar wavelet-based enhanced residual attention block, and an enhanced residual block. The haze concentration estimation module is HEBlock, the color correction residual gating module is CRG, the multi-scale learnable fusion block is MSLFB, the enhanced residual attention block is WRABlock, and the enhanced residual block is WRBlock. The overall framework is as follows: Figure 4 As shown;
[0107] S2: Design a loss function for the DFCCNet network module, use the loss function to train and optimize the DFCCNet network module, and then use the optimized DFCCNet network module for image dehazing.
[0108] S3: First, the haze concentration estimation module is used to obtain a more accurate haze density prior;
[0109] S4: Subsequently, wavelet transform is used to extract multi-scale frequency domain sub-band features and fuse them with spatial domain features. By using Haar wavelet-based enhanced residual attention blocks and enhanced residual blocks, the features are enhanced to represent haze layers and edge structures. Under the guidance of haze priors, the enhanced features effectively avoid misjudgment of background information and blurring of edge structures.
[0110] S5: The color correction residual gating module is used to solve the color shift problem that often occurs after halo removal. The color correction residual gating module effectively restores a realistic and natural color distribution through color correction and residual signal modulation.
[0111] S6: Finally, based on multi-scale learnable fusion blocks, adaptive interaction and reconstruction between different features are realized, thereby promoting visual friendliness and perception-oriented enhancement.
[0112] The fog concentration estimation module extracts the amplitude spectrum through two-dimensional Fourier transform and performs logarithmic compression, which can intuitively characterize the fog intensity changes at different spatial frequencies in the frequency domain. The formula is shown below:
[0113]
[0114] Where (u,v) represents the frequency coordinates; I(x,y) represents the input image; e -j2π(ux+vy) It is a complex exponential function in the Fourier transform, representing the transform weight of the image in the frequency domain;
[0115] To reflect the overall frequency distribution characteristics, its amplitude spectrum is calculated and the dynamic range is compressed by taking its logarithm:
[0116] M(u,v)=log(1+|F(u,v)|)
[0117] log(1+|F(u,v)|) performs logarithmic compression on the amplitude, which is used to expand the energy of smaller frequency components in the frequency domain and enhance the visualization of low-frequency components.
[0118] M(u,v) emphasizes the distribution of low-frequency and mid-to-high-frequency textures in the image, and has a significant ability to distinguish areas of reduced contrast and weakened edges caused by fog.
[0119] To enhance the resolution of local spatial details, a two-dimensional Haar wavelet transform is applied to the input image. Enhanced residual attention blocks and enhanced residual blocks based on Haar wavelets are used to enhance the representation of haze layers and edge structures.
[0120] After undergoing a two-dimensional Haar wavelet transform, the input image can be separated to reveal detailed variations in different directions (horizontal, vertical, and diagonal) and scales, thus supplementing spatial information that Fourier transform cannot locate. The wavelet decomposition result is divided into four sub-bands, as shown in the following formula:
[0121] W(k)(x,y)=(I*ψ(k))(x,y),k=1,2,3,4
[0122] Where * denotes convolution operation; I denotes the spatial domain representation of the input image; ψ(k) denotes the k-th wavelet basis function, used to extract features of different scales and orientations from the image;
[0123] By using spatial downsampling and upsampling, the wavelet features are restored to the same resolution as the input image, so that they can be stitched together and fused with the Fourier amplitude spectrum features, as shown in the following formula:
[0124] Φ=Concat(M,W(1),W(2),W(3),W(4))
[0125] By combining the explicit attention map A(x,y) generated by the lightweight convolutional attention module, adaptive weighting of multi-domain features is achieved, resulting in the final fog concentration estimation map D^(x,y).
[0126] D^(x,y0=F fusion (Φ(x,y)·A(x,y))
[0127] Where F fusion This indicates a fused convolutional network.
[0128] The color correction residual gating module combines dynamic routing self-attention (HDR) with wavelet self-attention multi-scale color correction, and introduces explicit residual signals and gating mechanisms to more accurately adaptively adjust color deviations.
[0129] A hierarchical dynamic routing module is built using dynamic routing self-attention to adaptively fuse multi-scale context information.
[0130] Input feature X is processed by average pooling and convolution at three different scales to generate scaled features:
[0131]
[0132] in This indicates that average pooling is performed on X, and each F... i Upsampled to the original resolution, and then aligned by filling or truncating the channels before being stitched together with X;
[0133] Next, route weights are generated using two 1×1 convolutional layers and Softmax:
[0134] W r =Softmax(Conv 1×1 (GELU(Conv 1×1([X,F1,F2,F3])))).
[0135] W r After reshaping to [B,C,H,W], the features at each scale are weighted accordingly and fused by adding the residuals:
[0136]
[0137] Among them, from W r The extracted i-th routing weight determines the i-th scale feature F. i The weighting coefficients; Proj i This is a projection operation.
[0138] Wavelet self-attention multiscale color correction, i.e., MWCC, such as Figure 5 As shown, given feature X, the low-frequency component LL and the self-attention feature X' are calculated respectively, and then the preliminary corrected feature F is obtained by the following formula. * :
[0139] F = GELU(Conv) 1×1 (Concat(X`,Up(LL(X)))))
[0140] F * =GELU(Conv 1×1 Conv 3×3 (F))+F
[0141] Where Up(·) represents upsampling;
[0142] Residual gating is used to avoid overcorrection and adaptively adjust color residuals;
[0143] First, global average pooling (GAP) is applied to the input feature X, followed by two 1×1 convolutions and ReLU activation. Then, the residual gate weights G are obtained through the Sigmoid function.
[0144] G=σ(Conv 1×1 (ReLU(Conv 1×1 (GAP(X)))))
[0145] Where σ represents the Sigmoid function and GAP(·) represents the pooling layer;
[0146] Finally, the gated weighted color residual is fused with the original feature X to obtain the output feature.
[0147]
[0148] Where F *This refers to the result after preliminary correction by MWCC, where ⊙ represents element-wise multiplication.
[0149] The multi-scale learnable fusion block integrates multi-scale spatial attention (MSA) and pyramid channel attention (MCA), and introduces a learnable fusion module (LFM) for further calibration at the channel level, such as... Figure 6 It allows for the automatic suppression of redundant information and the enhancement of important signals to ensure feature complementarity.
[0150] Pyramid channel attention extracts channel attention at multiple downsampling scales to perceive global statistics under different receptive fields, for a given scale set b. i Average pooling and channel attention are performed to obtain attention features at each scale. These features are then upsampled back to the original resolution and stitched together.
[0151]
[0152] Where F is the feature map, CALayer(·) represents the channel layer, and Pool(·) represents the pooling layer;
[0153] Multi-scale spatial attention integrates spatial attention output from receptive fields of various convolutional kernels, enabling it to focus on both subtle textures and larger structures.
[0154] Using a set of spatial attention blocks with different kernel sizes, generate respectively And weight the original features:
[0155]
[0156] Where S k This represents the feature results generated by different kernel sizes;
[0157] S will then k The concatenation is performed along the channel dimension, and then fused with 3×3 convolution, followed by Tanh activation and residual addition.
[0158] F MSA =tanh(Conv 3×3 (Concat(S 1 ,S 3 ,S 5 ,S 7 )))+F
[0159] Here, tanh is the activation function.
[0160] The learnable fusion module further adaptively corrects the initial fusion features and the attention-based fusion features.
[0161] Reduce information loss;
[0162] First, in the channel dimension, the initial fusion features F0 and F... mix splicing:
[0163] F cat =Concat(F0,F mix )
[0164] Then, a fusion-gated mapping is constructed using two layers of 3×3 convolutions:
[0165]
[0166] Where b1 and b2 are both bias terms in the convolution operation; r is the channel compression ratio, i.e., the bottleneck ratio; σ is the Sigmoid activation function; These are the parameters of the convolutional layer; the fusion weights P∈[0,1] B×2C×H×W By weighting the spliced features element by element, the final fused representation is obtained:
[0167] F LFM =P⊙F cat
[0168] Where ⊙ represents element-wise multiplication;
[0169] Finally, to maintain consistency with subsequent channel dimensions, one more convolution is performed.
[0170]
[0171] Here, b3 is the bias term in the convolution operation.
[0172] The loss function uses a weighted combination of pixel-level reconstruction loss L1 and contrast preservation loss to simultaneously ensure the brightness accuracy and local contrast of the restored result.
[0173] Pixel reconstruction loss, denoted as the network prediction result. The corresponding sharp image is denoted as J, with image size C×H×W and batch size B. The loss is defined as follows:
[0174]
[0175] Contrast Preservation Loss: To further enhance the fidelity of local structure and texture, a contrast preservation loss is introduced, denoted as . Contrast preservation loss achieves explicit constraints on detail contrast by measuring the difference in contrast features between the predicted and ground truth maps within pixel neighborhoods or multi-scale windows. Formally, it can be expressed as:
[0176]
[0177] in The comparison operator is defined as n traversing all spatial locations or sample moments, with a total of N terms.
[0178] The aforementioned pixel-level reconstruction loss and contrast preservation loss are expressed using hyperparameters λ1 and λ2. CR Weighted combination yields the total loss:
[0179]
[0180] Set λ1 = 1.0 and λ CR =0.1.
[0181] To demonstrate the effectiveness of the DFCCNet network module, the following experimental scheme is proposed:
[0182] Dataset
[0183] In this experiment, we comprehensively evaluated the proposed DFCCNet using both synthetic and real images. For the synthetic data, we employed the realistic single-image DEhazing (residence) dataset. It consists of five subsets: Indoor Training Set (ITS), Outdoor Training Set (OTS), Comprehensive Objective Test Set (SOTS), Real-World Task-Driven Test Set (RTTS), and Hybrid Subjective Test Set (HSTS). During training, we used the Indoor Training Set (ITS), which contains 1399 haze-free images, each used to generate 10 synthetic haze images of different densities based on an atmospheric scattering model. We also optimized the network parameters using the Outdoor Training Set (OTS), which consists of approximately 296,000 pairs of blurred and sharp synthetic images.
[0184] For evaluation, we used the Comprehensive Objective Test Set (SOTS), which includes indoor (SOTS-Indoor) and outdoor (SOTS-Outdoor) datasets, to assess the model's generalization ability within their respective training domains. Additionally, DFCCNet was evaluated using the Haze4k dataset, which contains 3000 synthetic training images and 1000 synthetic test images. Furthermore, to further evaluate the model's generalization and robustness to real-world data, we tested it on the O-Haze and RW2AH real-world datasets, respectively. (See [link to relevant documentation]). Figure 7-11 , and Tables 1 to 6.
[0185] Table 1 shows the quantitative analysis of the sots-indoor dataset.
[0186]
[0187] Table 2 shows the quantitative analysis of the sots-outdoor dataset.
[0188]
[0189]
[0190] Table 3 shows the quantitative analysis of the Haze4K dataset.
[0191]
[0192] Table 4 shows the quantitative analysis on the real dataset 0-HAZE.
[0193]
[0194] Table 5 shows the quantitative analysis on the real dataset RW2AH.
[0195]
[0196] Table 6 shows the ablation experiments performed on SOTS-Outdoor.
[0197]
[0198] Comparison methods and evaluation indicators
[0199] EPDN, FFANet, SANet, C2PNet, DEANet, DFRNet, MCPNet, DehazeFormer, T-Net, MitNet, and ConvIRare were used to evaluate performance on labeled datasets. Two additional methods, UMENet and SGDN, were also included for comparative experiments on real-world datasets. The dehazing performance of the synthetic datasets was evaluated using two commonly used image quality metrics in computer vision: Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). Furthermore, LPIPS, FID, and CIEDE2000 scores were calculated to assess the fidelity of color restoration and the perceptual quality of the restored images.
[0200] Implementation details
[0201] We implemented the proposed DFCCNet model on the PyTorch deep learning platform using a single NVIDIA RTX 3090 GPU. We chose to use the Adam optimizer, with β1, β2, and ε set to their default values of 0.9, 0.999, and 1e-8, respectively. The batch size was set to 8, and the initial learning rate was set to 0.0001, which was then adjusted to 1e-6 using cosine annealing. During training, we set the image size to 256×256×3 and employed two data augmentation techniques: 90°, 180°, or 270° rotation and vertical or horizontal flipping. The model underwent a total of 1500K iterations during the entire training phase.
[0202] Comparison with SOTA
[0203] Results on synthetic datasets
[0204] Table 1-3 shows the quantitative evaluation results of our proposed method on SOTS and Haze4k, comparing it with more than 10 other state-of-the-art methods proposed in recent years. As shown in the table, DFCCNet achieved the highest average SSIM and Psnr (highlighted in red) on the three synthetic datasets: a Psnr of 42.3 and an SSIM of 0.996 for SOTS-indoor, and a Psnr of 37.13 and an SSIM of 0.991 for SOTS-outdoor. This indicates that our method outperforms other comparative methods in terms of brightness, contrast, structure, and distortion suppression. Furthermore, to demonstrate its strong performance in color fidelity and perceived image quality, we also compared the LPIPS, FID, and CIEDE2000 scores of each method, showing that our method also achieved excellent results. In addition, we visualized the restored images from SOTS-indoor and SOTS-outdoor using the methods described above. It can be seen that EPDN and TNet methods still exhibit some haze in the restored images. The results from FFANet, DehazeFormer, MITNet, and SANet exhibit significant color distortion and some artifacts. While DEANet achieves a fairly good restoration, some differences in color detail remain. In contrast, our method produces the most natural restoration, preserving more detail and involving less color distortion.
[0205] Results on real datasets
[0206] We also evaluated the proposed DFCCNet on real-world datasets, including o-haze and RW2AH. It's worth noting that real-world images present greater challenges due to their more complex characteristics. However, as shown in Tables 4-5, our method remains highly competitive on both datasets. We also visualized the results on... Figure 10 Although none of the compared methods produced very good results in reconstructing haze-free images, our method produced the most ideal images, successfully eliminating most of the haze.
[0207] ablation experiment
[0208] We analyzed the effectiveness of different components of the proposed DFCCNet, including the fog estimation module (HEBlock), the color correction module (CRG), and the learnable fusion module (LFM). Our base network is DEANet, and we then set five variables for the final experiments. Our ablation experiments were primarily tested on the SOTS-outdoor dataset.
[0209] 1) v2 (base+HEBlock): Adds prior fog information extracted by HEBlock to the U-Net structure. 2) v3 (base+HEBlock+CRG-1): Adds a color correction residual gating module during the first downsampling. 3) v4 (base+HEBlock+CRG-2): Adds a color correction residual gating module during the second downsampling. 4) v5 (base+HEBlock+CRG): Adds a color correction residual gating module during both downsampling operations of the network. 5) v6 ((base+HEBlock+CRG+LFM)): Based on previous experiments, adds a learnable fusion module for splicing during multi-scale attention when performing skip connections.
[0210] Effectiveness of HEBlock: HEBlock extracts fog information from the input image from a frequency domain perspective, combining Fourier and wavelet transforms. Adding it to subsequent networks allows for better learning of fog information, resulting in improved dehazing performance. Therefore, as shown in Table 6, HEBlock achieves a 0.79dB improvement over the base+HEBlock network.
[0211] CRG Effectiveness: Our goal was to determine which part of the network benefits most from integrating the CRG module. To this end, we designed three experiments, adding CRG in the first downsampling stage, the second downsampling stage, and both stages. The results show that applying CRG in the first and second downsampling stages yields the best performance. As shown in Table 6, the two-stage configuration improves PSNR by 0.68 dB compared to not using CRG or adding CRG in a single stage. We also visualized the network's comparison with the CRG module and the baseline network, as shown in Table 6. Figure 11 As shown in the figure. The results indicate that the R, G, and B color channel distributions obtained by this method are closer to the ground truth image.
[0212] Effectiveness of LFM: Building upon CGA, we introduce the concept of multi-scale features and further incorporate a learnable fusion module to adaptively refine the initial and attention-enhancing features after fusion, thereby reducing information loss. As shown in Table 6, our final network improves upon the original baseline by 1.88 dB.
[0213] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An image dehazing method based on dual-domain haze prior and color residual gating, characterized in that: At least the following steps are included: S1: Construct the DFCCNet network module, which integrates a haze concentration estimation module, a color correction residual gating module, a multi-scale learnable fusion block, a Haar wavelet-based enhanced residual attention block, and an enhanced residual block. The haze concentration estimation module is HEBlock, the color correction residual gating module is CRG, the multi-scale learnable fusion block is MSLFB, the enhanced residual attention block is WRABlock, and the enhanced residual block is WRBlock. S2: Design a loss function for the DFCCNet network module, use the loss function to train and optimize the DFCCNet network module, and then use the optimized DFCCNet network module for image dehazing. S3: First, the haze concentration estimation module is used to obtain a more accurate haze density prior; S4: Subsequently, wavelet transform is used to extract multi-scale frequency domain sub-band features and fuse them with spatial domain features. By using Haar wavelet-based enhanced residual attention blocks and enhanced residual blocks, the features are enhanced to represent haze layers and edge structures. Under the guidance of haze priors, the enhanced features effectively avoid misjudgment of background information and blurring of edge structures. S5: The color correction residual gating module is used to solve the color shift problem that often occurs after halo removal. The color correction residual gating module effectively restores a realistic and natural color distribution through color correction and residual signal modulation. S6: Finally, based on multi-scale learnable fusion blocks, adaptive interaction and reconstruction between different features are realized, thereby promoting visual friendliness and perception-oriented enhancement.
2. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 1, characterized in that: The fog concentration estimation module extracts the amplitude spectrum through two-dimensional Fourier transform and performs logarithmic compression, which can intuitively characterize the fog intensity changes at different spatial frequencies in the frequency domain, as shown in the following formula: Where (u,v) represents the frequency coordinates; I(x,y) represents the input image; e -j2π(ux+vy) It is a complex exponential function in the Fourier transform, representing the transform weight of the image in the frequency domain; To reflect the overall frequency distribution characteristics, its amplitude spectrum is calculated and the dynamic range is compressed by taking its logarithm: M(u,v)=log(1+|F(u,v)|) log(1+|F(u,v)|) performs logarithmic compression on the amplitude, which is used to expand the energy of smaller frequency components in the frequency domain and enhance the visualization of low-frequency components. M(u,v) emphasizes the distribution of low-frequency and mid-to-high-frequency textures in the image, and has a significant ability to distinguish areas of reduced contrast and weakened edges caused by fog.
3. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 2, characterized in that: To enhance the resolution of local spatial details, a two-dimensional Haar wavelet transform is applied to the input image. Enhanced residual attention blocks and enhanced residual blocks based on Haar wavelets are used to enhance the representation of haze layers and edge structures. After the input image undergoes a two-dimensional Haar wavelet transform, detailed variations in different directions and scales can be separated, thus supplementing the spatial information that Fourier transform cannot locate. The wavelet decomposition result is divided into four sub-bands, and the formula is shown below: W(k)(x,y)=(I*ψ(k))(x,y),k=1,2,3,4 Where * denotes convolution operation; I denotes the spatial domain representation of the input image; ψ(k) denotes the k-th wavelet basis function, used to extract features of different scales and orientations from the image; By using spatial downsampling and upsampling, the wavelet features are restored to the same resolution as the input image, so that they can be stitched together and fused with the Fourier amplitude spectrum features, as shown in the following formula: Φ=Concat(M,W(1),W(2),W(3),W(4)) By combining the explicit attention map A(x,y) generated by the lightweight convolutional attention module, adaptive weighting of multi-domain features is achieved, resulting in the final fog concentration estimation map D^(x,y). D^(x,y)=F fusion (Φ(x,y)·A(x,y)) Where F fusion This indicates a fused convolutional network.
4. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 1, characterized in that: The color correction residual gating module combines dynamic routing self-attention with wavelet self-attention multi-scale color correction, and introduces explicit residual signals and gating mechanisms to more accurately adaptively adjust color deviations.
5. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 4, characterized in that: A hierarchical dynamic routing module is built using dynamic routing self-attention to adaptively fuse multi-scale context information. Input feature X is processed by average pooling and convolution at three different scales to generate scaled features: F i =Conv 3×3 (AvgPool 2i (X)),i=1,2,3, in This indicates that average pooling is performed on X, and each F... i Upsampled to the original resolution, and then aligned by filling or truncating the channels before being stitched together with X; Next, route weights are generated using two 1×1 convolutional layers and Softmax: W r =Softmax(Conv 1×1 (GELU(Conv 1×1 ([X,F1,F2,F3])))). W r After reshaping to [B,C,H,W], the features at each scale are weighted accordingly and fused by adding the residuals: Among them, from W r The extracted i-th routing weight determines the i-th scale feature F. i The weighting coefficients; Proj i This is a projection operation.
6. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 5, characterized in that: The wavelet self-attention multi-scale color correction (MWCC) calculates the low-frequency component LL and the self-attention feature X' given feature X, and then obtains the preliminary corrected feature F using the following formula. * : F=GELU(Conv 1×1 (Concat(X`,Up(LL(X))))) F * =GELU(Conv 1×1 Conv 3×3 (F))+F Where Up(·) represents upsampling; The residual gating is used to avoid overcorrection and adaptively adjust the color residual; First, global average pooling is applied to the input feature X, followed by two 1×1 convolutions and ReLU activation. Then, the residual gating weights G are obtained through the Sigmoid function. G=σ(Conv 1×1 (ReLU(Conv 1×1 (GAP(X))))) Where σ represents the Sigmoid function and GAP(·) represents the pooling layer; Finally, the gated weighted color residual is fused with the original feature X to obtain the output feature. Where F * This refers to the result after preliminary correction by MWCC, where ⊙ represents element-wise multiplication.
7. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 1, characterized in that: The multi-scale learnable fusion block integrates multi-scale spatial attention and pyramid channel attention, and introduces a learnable fusion module for further calibration at the channel level, allowing automatic suppression of redundant information and enhancement of important signals to ensure feature complementarity.
8. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 7, characterized in that: The pyramid channel attention extracts channel attention at multiple downsampling scales to perceive global statistics under different receptive fields, for a given scale set b. i Average pooling and channel attention are performed to obtain attention features at each scale. These features are then upsampled back to the original resolution and stitched together. Where F is the feature map, CALayer(·) represents the channel layer, and Pool(·) represents the pooling layer; The multi-scale spatial attention fusion of spatial attention outputs under the receptive fields of multiple convolutional kernels can focus on both subtle textures and capture larger structures. Using a set of spatial attention blocks with different kernel sizes, generate respectively And weight the original features: Where S k This represents the feature results generated by different kernel sizes; S will then k The concatenation is performed along the channel dimension, and then fused with 3×3 convolution, followed by Tanh activation and residual addition. F MSA =tanh(Conv 3×3 (Concat(S 1 ,S 3 ,S 5 ,S 7 )))+F Here, tanh is the activation function.
9. The image dehazing method based on dual-domain haze prior and color residual gating as described in claim 8, characterized in that: The learnable fusion module further adaptively corrects the initial fusion features and attention fusion features to reduce information loss; First, in the channel dimension, the initial fusion features F0 and F... mix splicing: F cat =Concat(F0,F mix ) Then, a fusion-gated mapping is constructed using two layers of 3×3 convolutions: Where b1 and b2 are both bias terms in the convolution operation; r is the channel compression ratio, i.e., the bottleneck ratio; σ is the Sigmoid activation function; These are the parameters of the convolutional layer; the fusion weights P∈[0,1] B×2C×H×W By weighting the spliced features element by element, the final fused representation is obtained: F LFM =P⊙F cat Where ⊙ represents element-wise multiplication; Finally, to maintain consistency with subsequent channel dimensions, one more convolution is performed. Here, b3 is the bias term in the convolution operation.
10. The image dehazing method based on dual-domain haze prior and color residual gating according to claim 9, characterized in that: The loss function is a weighted combination of pixel-level reconstruction loss L1 and contrast preservation loss, so as to simultaneously ensure the brightness accuracy and local contrast of the restored result. Pixel reconstruction loss, denoted as the network prediction result. The corresponding sharp image is denoted as J, with image size C×H×W and batch size B. The loss is defined as follows: Contrast Preservation Loss: To further enhance the fidelity of local structure and texture, a contrast preservation loss is introduced, denoted as . Contrast preservation loss achieves explicit constraints on detail contrast by measuring the difference in contrast features between the predicted and ground truth maps within pixel neighborhoods or multi-scale windows. Formally, it is expressed as: in The comparison operator is defined as n traversing all spatial locations or sample moments, with a total of N terms. The aforementioned pixel-level reconstruction loss and contrast preservation loss are expressed using hyperparameters λ1 and λ2. CR Weighted combination yields the total loss: Set λ1 = 1.0 and λ CR =0.1.
Citation Information
Cited By
Underwater image enhancement and restoration method based on double-domain adaptive fusion module
CN122023163A