Defogging method based on frequency information differential fusion
Through a dehazing method based on differential fusion of frequency information, utilizing a U-shaped encoder-decoder architecture and a staged training strategy, the problem of existing methods failing to fully utilize frequency characteristic differences is solved, achieving efficient image dehazing effects and improving visual quality.
Patent Information
- Application Number
- CN202511163952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing dehazing methods fail to fully utilize the differences in frequency characteristics of images, resulting in limited dehazing effects. They also ignore the processing requirements of high-frequency and low-frequency information at different stages, affecting the visual quality and detail integrity of the dehazed image.
A dehazing method based on differential fusion of frequency information is adopted. Through a U-shaped encoder-decoder architecture, discrete wavelet transform and feature extraction modules are combined to process high-frequency and low-frequency components respectively. The network performance is optimized through a staged training strategy to ensure adaptive restoration of image details and global structures at different stages.
The clarity and detail fidelity of dehazed images are significantly improved, and higher dehazing quality and more stable generalization performance can be achieved in various haze levels and complex scenes.
Smart Images

Figure CN120725904A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a defogging method based on differential fusion of frequency information. Background Art
[0002] In adverse weather conditions such as haze and dust, the air contains a large number of tiny particles. These particles scatter and absorb light, causing image clarity loss, color shift, lack of contrast, and blurred details in images captured by imaging devices. This not only affects the human visual experience but also adversely impacts intelligent perception systems that rely on image information. Therefore, image dehazing technology, which targets image restoration and enhancement in these degraded environments, has become a key research area in computer vision.
[0003] Image dehazing, a typical low-level visual processing task, aims to restore a clearer image closer to the real scene from hazy images. In hazy environments, light is scattered multiple times by particles during transmission, resulting in reduced image contrast, unnatural colors, and a loss of detail. Directly applying these degraded images to high-level visual tasks (such as intelligent monitoring, autonomous driving, and drone navigation) often results in a significant decrease in recognition accuracy. Image dehazing technology not only effectively improves image visualization quality but also enhances the robustness and accuracy of subsequent computer vision algorithms.
[0004] Image dehazing techniques can be broadly categorized into two main categories: traditional methods based on handcrafted priors and deep learning-based methods. The former primarily relies on the atmospheric scattering model (ASM) and manually designed prior knowledge. Traditional methods based on handcrafted priors typically estimate transmission maps and atmospheric light intensity, then combine them with the ASM to restore haze-free images. Their advantage is strong interpretability, but their disadvantage is their heavy reliance on prior assumptions, which can easily fail in complex scenes. In recent years, with the advancement of computer hardware computing power, the development of deep learning techniques such as convolutional neural networks (CNNs), and the construction of large-scale image dehazing datasets, deep learning-based dehazing methods have gradually become a research hotspot, replacing traditional methods. Early deep learning dehazing models typically combined the ASM with a deep learning network. The network predicted transmission maps or atmospheric light values, and then used the ASM to reconstruct the image. Recent end-to-end methods, however, dispense with explicit physical model estimation and directly learn the mapping relationship from hazy to haze-free images. These methods not only maintain high dehazing performance under a wide range of weather and lighting conditions, but also excel in computational efficiency and visual quality.
[0005] Currently, most dehazing methods focus primarily on image restoration in the spatial domain, failing to fully exploit the differences in frequency characteristics between hazy and haze-free images, thus limiting further improvement in dehazing effectiveness. To address this issue, some studies have introduced tools such as Fourier transforms and wavelet transforms to decompose image features into high-frequency and low-frequency components, and then process this frequency information to reconstruct a haze-free image. However, existing dehazing networks that incorporate frequency information still have significant limitations. First, existing methods fail to design differentiated processing strategies for the characteristics of high- and low-frequency information, resulting in inefficient utilization of frequency information. For example, high-frequency information typically corresponds to edges and details, while low-frequency information focuses more on overall structure and illumination. However, current methods often treat both types of frequency information identically, failing to fully leverage their respective strengths. Second, the network's reliance on high- and low-frequency information varies at different stages. For example, in the shallow stages of feature extraction, due to the high image resolution, the network can capture more detailed information and therefore relies more on high-frequency information to enhance edge and texture extraction. In the deep stages of feature extraction, as the input image resolution gradually decreases, the network tends to extract global information (such as illumination and overall structure), and the role of low-frequency information becomes more prominent. However, existing methods generally ignore this stage-by-stage requirement, resulting in a lack of adaptability in the processing of frequency information. In addition, after the introduction of frequency domain processing, some methods rely too much on frequency domain optimization and fail to fully balance the restoration requirements in the spatial domain, thus affecting the visual quality and detail integrity of the dehazed image. Although frequency domain processing can effectively capture global characteristics and local details, ignoring the restoration of the spatial domain may lead to texture distortion or color deviation. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a defogging method based on frequency information differential fusion with good defogging effect.
[0007] The technical solution of the present invention to solve the above technical problems is: a defogging method based on differential fusion of frequency information, comprising the following steps:
[0008] S1: Create a training dataset containing original fog-free images and corresponding foggy images;
[0009] S2: Perform data augmentation on the training dataset, including cropping and flipping the input image;
[0010] S3: Construct a dehazing network based on differential fusion of frequency information;
[0011] The constructed dehazing network adopts a U-shaped encoder-decoder architecture to perform image dehazing; given any foggy image I, I∈R 3×H×W, H and W represent the height and width of the image size respectively, R is the real number domain, the dehazing network first applies 3×3 convolution to generate shallow features of size C×H×W, where C represents the number of channels; then, the shallow features pass through three U-shaped feature extraction modules to obtain multi-scale features;
[0012] In the encoding stage, the features are transformed by discrete wavelet transform (DWT) and decomposed into low-frequency components and high-frequency components. The low-frequency components and high-frequency components are fused with the main path features through the high- and low-frequency information differential fusion module;
[0013] In the decoding stage, the features are gradually restored to high-resolution features through another three U-shaped feature extraction modules and upsampling to obtain the final output;
[0014] S4: Use the training data set to train the dehazing network until the preset loss function converges. Combine the loss function and the frequency domain loss as the objective function of the network training to obtain the trained dehazing network.
[0015] S5: Input the image to be dehazed into the trained dehazing network for testing and verification to obtain the dehazing result.
[0016] In the above-mentioned defogging method based on differential fusion of frequency information, in step S3, the operation steps of the U-shaped feature extraction module are as follows:
[0017] S311: Perform a 3×3 convolution operation on the input feature Fin to generate the initial feature map Fe1, and enhance the linear representation of the feature through the activation function DyT. The expression of DyT is:
[0018] Fe1=γ·tanh(α·(Fin))+β;
[0019] Among them, γ, α, and β are all learnable parameters; tanh is the hyperbolic tangent function;
[0020] S312: Perform two more 3×3 convolution operations and activation function DyT on the feature Fe1. The specific operations are as follows:
[0021] Fe2=DyT·(Conv(Fe1));
[0022] Fe3=DyT·(Conv(Fe2));
[0023] Among them, Fe2 is the intermediate output feature, Fe3 is the encoder output feature; Conv represents the 3×3 convolution operation;
[0024] S313: In the bottleneck stage, channel attention CA and spatial attention SA mechanisms are applied to Fe3 in sequence to enhance the selective extraction of key features.
[0025] In the above-mentioned dehazing method based on differential fusion of frequency information, in step S313, the process of the channel attention CA mechanism is as follows:
[0026] Perform a global average pooling GAP operation on Fe3 to compress the spatial dimension, then sequentially undergo 1×1 convolution, GELU activation, 1×1 convolution, and Sigmoid activation to generate a channel attention weight vector WCA. The channel attention weight vector is multiplied element-by-element by Fe3 to obtain the enhanced feature Fe4, which is expressed as:
[0027] Fe4= Fe3·Sigmoid(PWConv(GELU(PWConv(GAP(Fe3)))));
[0028] Among them, GELU represents the GELU activation function, PWConv represents 1×1 convolution, and Sigmoid represents the Sigmoid activation function.
[0029] In the above-mentioned dehazing method based on differential fusion of frequency information, in step S313, the process of the spatial attention SA mechanism is: 3×3 convolution, GELU activation, 3×3 convolution and Sigmoid activation are performed on Fe4 in sequence to generate a spatial attention weight map WSA. The spatial attention weight map is multiplied element-by-element with Fe4 to obtain the enhanced feature Fe5, which is expressed as:
[0030] Fe5= Fe4·Sigmoid (Conv(GELU(Conv(Fe4)))).
[0031] The above dehazing method based on differential fusion of frequency information performs three 3×3 convolution operations on Fe5 at the decoder stage. Each convolution layer is followed by an upsampling operation to restore the spatial resolution of the feature map. The specific operations are as follows:
[0032] Fd1 = ↑(DyT(conv(Fe5)));
[0033] Fd2 = ↑(DyT(conv(Fd1)));
[0034] Fd3 = ↑(DyT(conv(Fd2)));
[0035] Among them, ↑ represents the upsampling operation, Fd1 and Fd2 are intermediate output features, and Fd3 is the decoder output feature.
[0036] In the above-mentioned dehazing method based on differential fusion of frequency information, the high- and low-frequency information differential fusion module in step S3 includes three inputs and one output. It fuses the high-frequency information HF and low-frequency information LL obtained by wavelet transform with the original feature RF extracted by the U-shaped feature extraction module. The operation process of the high- and low-frequency information differential fusion module is as follows:
[0037] The high- and low-frequency information differential fusion module first performs differential processing on the three inputs it receives. The RF is input to the Convs module for further feature extraction to obtain the enhanced feature RF'; the LL is processed by the low-frequency enhancement channel attention block to enhance the global structure and illumination information in the low-frequency features to generate the enhanced feature LL'; the HF is sent to the high-frequency information amplification block to enhance the representation ability of the local structure of the image by amplifying the edge and detail information in the high-frequency features, and obtain the final feature HF';
[0038] LL' is multiplied by the learnable parameter a, HF' is multiplied by 1-a, and the two multiplication results are finally added to RF' to generate the final output O; the specific mathematical expression is as follows:
[0039] RF' = Convs(RF);
[0040] LL' = LFECA (LL);
[0041] HF' = HFEB (HF);
[0042] O = RF'+ a·LL'+(1-a)·HF';
[0043] Among them, Convs represents the Convs module, LFECA represents the low-frequency enhanced channel attention block, and HFEB represents the high-frequency amplification block.
[0044] In the above-mentioned dehazing method based on differential fusion of frequency information, the operation steps of the low-frequency enhancement channel attention block are as follows: the low-frequency information LL first independently extracts the spatial features of each channel through a 3×3 channel-by-channel convolution layer, focusing on the global structure of the low-frequency features; then, the inter-channel information is integrated through a 1×1 point-by-point convolution layer to further enhance the semantic expression of the features; finally, the low-frequency enhancement channel attention block introduces a channel attention mechanism, extracts channel-level global information through global average pooling, and uses two layers of 1×1 convolution with ReLU activation in the middle to generate channel attention weights, which are normalized to [0,1] using a Sigmoid function.
[0045] In the above-mentioned dehazing method based on differential fusion of frequency information, the operation steps of the high-frequency information amplification block are as follows: the high-frequency information amplification block first performs a feature shift operation on the input HF, shifts the HF by one pixel to the lower right direction, and fills the empty edge part with zeros to generate a shifted feature map; then, the shifted feature map is subtracted from the HF element by element, and the absolute value of the result is taken to obtain a difference map; then, the difference map is input into a sequence including convolution, batch normalization, GELU activation function and convolution for further processing; after that, the feature is normalized by Sigmoid activation and added to the all-one matrix, and then multiplied element by element with the HF to obtain the enhanced feature; finally, the enhanced feature is residually connected with the convolved HF to generate the final feature HF'.
[0046] In the above-mentioned dehazing method based on differential fusion of frequency information, a staged training strategy is introduced in step S4. The staged training strategy optimizes the multi-scale feature learning and overall performance of the dehazing network by combining deep supervision and curriculum learning. The staged training strategy divides the training process into four stages with n rounds as the boundary. The specific steps are as follows:
[0047] In the first stage, the loss is calculated for the low-resolution output and the corresponding low-resolution real image, guiding the network to prioritize learning the basic structure and global features of the image;
[0048] In the second stage, supervision is shifted to medium-resolution output, and the loss is calculated with the corresponding medium-resolution real image, gradually enhancing the ability to restore local textures and details;
[0049] In the third stage, supervision focuses on the high-resolution output and calculates the loss with the corresponding high-resolution real image to ensure the overall consistency and fine-grained details of the dehazing results;
[0050] In the fourth stage, the losses of all resolution outputs are optimized to enable the network to fully capture multi-scale features.
[0051] The beneficial effects of the present invention are as follows: the present invention introduces a feature processing mechanism based on the differential fusion of frequency information, which can extract global information and detailed texture information from low-frequency components and high-frequency components respectively, thereby significantly improving the clarity and detail fidelity of the dehazed image. In addition, the present invention combines a staged training strategy, organically integrates deep supervision and curriculum learning, and uses a progressive training method from low resolution to high resolution, from structure to detail, so that the network can first grasp the global outline and main structure in the early stage, and gradually strengthen the ability to restore local details and textures in the later stage, and realize the comprehensive optimization of multi-scale features in the final stage. Therefore, the present invention can effectively avoid the problem of the network falling into the problem of overfitting details or ignoring global information in the early stage of training, thereby obtaining higher dehazing quality and more stable generalization performance under various degrees of fog and complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is the overall flow chart of the present invention.
[0053] Figure 2 This is a schematic diagram of the structure of the defogging network of the present invention.
[0054] Figure 3 Schematic diagram of the structure of the U-shaped feature extraction module.
[0055] Figure 4 Schematic diagram of the bottleneck design structure of the U-shaped feature extraction module.
[0056] Figure 5 Schematic diagram of the structure of the Convs module.
[0057] Figure 6 Schematic diagram of the structure of the high- and low-frequency information differential fusion module.
[0058] Figure 7 Schematic diagram of the structure of the low-frequency enhancement channel attention block.
[0059] Figure 8 Schematic diagram of the structure of the high-frequency information amplification block.
[0060] Figure 9 Schematic diagram of the staged training strategy.
[0061] Figure 10 Three examples of dehazing results are provided for the embodiments of the present invention; (a) is a foggy image, (b), (c), and (d) are the images after dehazing using the DCP, FFANet, and GridDehaze methods, respectively, (e) is the image after dehazing using the present invention, and (f) is the original fog-free image. DETAILED DESCRIPTION
[0062] The present invention will be further described below with reference to the accompanying drawings and examples.
[0063] like Figure 1 As shown, a defogging method based on differential fusion of frequency information includes the following steps:
[0064] S1: Create a training dataset containing original fog-free images and corresponding foggy images.
[0065] S2: Perform data augmentation on the training dataset, including cropping and flipping the input image. Specifically, the input image is randomly cropped into 256×256 image blocks and data augmented by randomly rotating it by 90°, 180°, or 270°.
[0066] S3: Construct a dehazing network based on differential fusion of frequency information.
[0067] The constructed dehazing network adopts a U-shaped encoder-decoder architecture to perform image dehazing, such as Figure 2 As shown; given any foggy image I, I∈R 3×H×W , H and W represent the height and width of the image size respectively, R is the real number domain, the dehazing network first applies 3×3 convolution to generate shallow features of size C×H×W, where C represents the number of channels; then, the shallow features pass through three U-shaped feature extraction modules to obtain multi-scale features;
[0068] During the encoding phase, features are transformed using discrete wavelet transforms (DWTs) and decomposed into low-frequency and high-frequency components (LH, HL, and HH). These components are then fused with the main path features using a high- and low-frequency information differential fusion module. During the decoding phase, features are gradually restored to high-resolution using three additional U-shaped feature extraction modules and upsampling to obtain the final output. Assuming an input image size of 256×256 pixels, the 128×128 and 64×64 stages obtained through downsampling are low-resolution feature maps, while the output feature maps gradually restored to 256×256 through upsampling are high-resolution outputs. During this process, skip connections are used between the encoder and decoder features to assist in image dehazing, and 1×1 convolutions are applied to reduce the number of channels. Furthermore, the present invention employs a staged training strategy (PTS), which monitors only the lower-resolution outputs in early training rounds to ensure robust learning and training of multi-scale features.
[0069] like Figure 3 As shown in Figure 2, the operation steps of the U-shaped feature extraction module are:
[0070] S311: Perform a 3×3 convolution operation on the input feature Fin to generate the initial feature map Fe1, and enhance the linear representation of the feature through the activation function DyT. The expression of DyT is:
[0071] Fe1=γ·tanh(α·(Fin))+β;
[0072] Among them, γ, α, and β are all learnable parameters, and their initial values are all set to 1; tanh is the hyperbolic tangent function;
[0073] S312: Perform two more 3×3 convolution operations and activation function DyT on the feature Fe1, gradually reducing the spatial resolution of the feature map while increasing the channel depth to capture high-level semantic information. The specific operations are as follows:
[0074] Fe2=DyT·(Conv(Fe1));
[0075] Fe3=DyT·(Conv(Fe2));
[0076] Among them, Fe2 is the intermediate output feature, Fe3 is the encoder output feature, which contains high-level semantic information; Conv represents a 3×3 convolution operation;
[0077] S313: In the bottleneck stage, channel attention CA and spatial attention SA mechanisms are applied to Fe3 in sequence to enhance the selective extraction of key features, such as Figure 4 As shown, Figure 4 The S-shaped function in is the Sigmoid activation function.
[0078] The process of the channel attention CA mechanism is:
[0079] Perform a global average pooling GAP operation on Fe3 to compress the spatial dimension, then sequentially undergo 1×1 convolution, GELU activation, 1×1 convolution, and Sigmoid activation to generate a channel attention weight vector WCA. The channel attention weight vector is multiplied element-by-element by Fe3 to obtain the enhanced feature Fe4, which is expressed as:
[0080] Fe4= Fe3·Sigmoid(PWConv(GELU(PWConv(GAP(Fe3)))));
[0081] Among them, GELU represents the GELU activation function, PWConv represents 1×1 convolution, and Sigmoid represents the Sigmoid activation function.
[0082] The process of the spatial attention SA mechanism is: perform 3×3 convolution, GELU activation, 3×3 convolution and Sigmoid activation on Fe4 in sequence to generate a spatial attention weight map WSA. The spatial attention weight map is multiplied element-by-element with Fe4 to obtain the enhanced feature Fe5, which is expressed as:
[0083] Fe5= Fe4·Sigmoid (Conv(GELU(Conv(Fe4)))).
[0084] In the decoder stage, three 3×3 convolution operations are performed on Fe5, and each convolution layer is followed by an upsampling operation to restore the spatial resolution of the feature map. The specific operations are as follows:
[0085] Fd1 = ↑(DyT(conv(Fe5)));
[0086] Fd2 = ↑(DyT(conv(Fd1)));
[0087] Fd3 = ↑(DyT(conv(Fd2)));
[0088] Among them, ↑ represents the upsampling operation, Fd1 and Fd2 are intermediate output features, and Fd3 is the decoder output feature.
[0089] like Figure 6 As shown in Figure 1, the high- and low-frequency information differential fusion module contains three inputs and one output, which fuses the high-frequency information HF and low-frequency information LL obtained by wavelet transform with the original feature RF extracted by the U-shaped feature extraction module; the operation process of the high- and low-frequency information differential fusion module is as follows:
[0090] The high- and low-frequency information differential fusion module first performs differential processing on the three inputs it receives. The RF is input to the Convs module for further feature extraction to obtain the enhanced feature RF'; the LL is processed by the low-frequency enhancement channel attention block to enhance the global structure and illumination information in the low-frequency features to generate the enhanced feature LL'; the HF is sent to the high-frequency information amplification block to enhance the representation ability of the local structure of the image by amplifying the edge and detail information in the high-frequency features, and obtain the final feature HF';
[0091] LL' is multiplied by the learnable parameter a, HF' is multiplied by 1-a, and the two multiplication results are finally added to RF' to generate the final output O; the specific mathematical expression is as follows:
[0092] RF' = Convs(RF);
[0093] LL' = LFECA (LL);
[0094] HF' = HFEB (HF);
[0095] O = RF'+ a·LL'+(1-a)·HF';
[0096] Among them, Convs represents the Convs module, and the Convs module results are as follows Figure 5 As shown; LFECA represents the low-frequency enhanced channel attention block, and HFEB represents the high-frequency amplification block.
[0097] like Figure 7As shown in the figure, the low-frequency enhancement channel attention block is a core component in the high- and low-frequency information differential fusion module designed specifically for low-frequency information enhancement. It aims to enhance the global structure and illumination information in low-frequency features through lightweight operations, while improving computational efficiency to meet the requirements of real-time dehazing tasks. The low-frequency enhancement channel attention block operates as follows: the low-frequency information LL first independently extracts spatial features from each channel through a 3×3 channel-by-channel convolution layer, focusing on the global structure of the low-frequency features. Subsequently, a 1×1 point-by-point convolution layer integrates inter-channel information to further enhance the semantic expression of the features. Finally, the low-frequency enhancement channel attention block introduces a channel attention mechanism, extracting channel-level global information through global average pooling. It then generates channel attention weights using two layers of 1×1 convolutions supplemented with ReLU activations in between. These weights are then normalized to [0, 1] using a sigmoid function, thereby adaptively highlighting channels that contribute significantly to the dehazing task, such as features related to global contrast and illumination.
[0098] like Figure 8 As shown in the figure, the high-frequency information amplification block is a core component in the high- and low-frequency information differential fusion module designed specifically for enhancing high-frequency information. Its goal is to improve the ability to represent the local structure of the image by amplifying the edge and detail information in the high-frequency features, thereby effectively restoring the details blurred by haze in the image dehazing task. The operation steps of the high-frequency information amplification block are as follows: the high-frequency information amplification block first performs a feature shift operation on the input HF, shifting the HF one pixel to the lower right, and filling the edge portion left empty after the shift with zeros to generate a shifted feature map A1; then, the shifted feature map A1 is subtracted element-by-element from the HF, and the absolute value of the result is taken to obtain the difference map A2. Figure 8 in represents absolute value subtraction; then, the difference map is input into a sequence including convolution, batch normalization, GELU activation function and convolution for further processing; after that, the feature is normalized by Sigmoid activation and added to the all-one matrix A3, and then multiplied element-wise with HF to obtain the enhanced feature; finally, the enhanced feature is residually connected with the convolved HF to generate the final feature HF'.
[0099] S4: Use the training data set to train the dehazing network until the preset loss function converges; combine the loss function and the frequency domain loss as the objective function of the network training to obtain the trained dehazing network.
[0100] In step S4, a staged training strategy is introduced. The staged training strategy optimizes the multi-scale feature learning and overall performance of the dehazing network by combining deep supervision and curriculum learning. The staged training strategy is bounded by n rounds, such as Figure 9 As shown, the training process is divided into four stages. The specific steps are as follows:
[0101] In the first stage, the loss is calculated for the low-resolution output and the corresponding low-resolution real image, guiding the network to prioritize learning the basic structure and global features of the image;
[0102] In the second stage, supervision is shifted to medium-resolution output, and the loss is calculated with the corresponding medium-resolution real image, gradually enhancing the ability to restore local textures and details;
[0103] In the third stage, supervision focuses on the high-resolution output and calculates the loss with the corresponding high-resolution real image to ensure the overall consistency and fine-grained details of the dehazing results;
[0104] In the fourth stage, the losses of all resolution outputs are optimized to enable the network to fully capture multi-scale features.
[0105] S5: Input the image to be dehazed into the trained dehazing network for testing and verification to obtain the dehazing result.
[0106] This paper conducts experiments using the NH-HAZE dataset, a publicly available dataset of non-uniform properties. Images in NH-HAZE are generated using a professional haze camera that simulates real-world haze conditions. The images have a resolution of 1600×1200 and contain 55 pairs of real hazy and fog-free images. For the NH-HAZE dataset, this paper uses the last five pairs as the test set, and the remaining images in the dataset as the training set.
[0107] Figure 10 The results of dehazing some images on the test dataset are shown, where (a) is a foggy image, (b), (c), and (d) are images after dehazing using the Dark Channel Prior (DCP), Feature Fusion Attention Network (FFANet), and Grid Dehaze Network (GridDehaze) methods, respectively; (e) is the dehazed image of the present invention; and (f) is the original fog-free image. It can be seen from the figure that the dehazed images generated by other methods contain artifacts left by incomplete dehazing, while the method proposed in the present invention can effectively restore the detailed information of the image and generate a dehazed image that is close to the real image.
Claims
1. A defogging method based on differential fusion of frequency information, characterized in that: The following steps are involved: S1: Create a training dataset containing original fog-free images and corresponding foggy images; S2: Perform data augmentation on the training dataset, including cropping and flipping the input image; S3: Construct a dehazing network based on differential fusion of frequency information; The constructed dehazing network adopts a U-shaped encoder-decoder architecture to perform image dehazing; given any foggy image I, I∈R 3×H×W , H and W represent the height and width of the image size respectively, R is the real number domain, the dehazing network first applies 3×3 convolution to generate shallow features of size C×H×W, where C represents the number of channels; then, the shallow features pass through three U-shaped feature extraction modules UFEM to obtain multi-scale features; In the encoding stage, the features are transformed by discrete wavelet transform (DWT) and decomposed into low-frequency components and high-frequency components. The low-frequency components and high-frequency components are fused with the main path features through the high-low frequency information differential fusion module (HLFDFM); In the decoding stage, the features are gradually restored to high-resolution features through another three U-shaped feature extraction modules and upsampling to obtain the final output; S4: Use the training data set to train the dehazing network until the preset loss function converges. Combine the loss function and the frequency domain loss as the objective function of the network training to obtain the trained dehazing network. S5: Input the image to be dehazed into the trained dehazing network for testing and verification to obtain the dehazing result.
2. The defogging method based on frequency information differential fusion according to claim 1 is characterized in that: In step S3, the operation steps of the U-shaped feature extraction module are: S311: Perform a 3×3 convolution operation on the input feature Fin to generate the initial feature map Fe1, and enhance the linear representation of the feature through the activation function DyT. The expression of DyT is: Fe1=γ·tanh(α·(Fin))+β; Among them, γ, α, and β are all learnable parameters; tanh is the hyperbolic tangent function; S312: Perform two more 3×3 convolution operations and activation function DyT on the feature Fe1. The specific operations are as follows: Fe2=DyT·(Conv(Fe1)); Fe3=DyT·(Conv(Fe2)); Among them, Fe2 is the intermediate output feature, Fe3 is the encoder output feature; Conv represents a 3×3 convolution operation; S313: In the bottleneck stage, channel attention CA and spatial attention SA mechanisms are applied to Fe3 in sequence to enhance the selective extraction of key features.
3. The defogging method based on frequency information differential fusion according to claim 2 is characterized in that: In step S313, the process of the channel attention CA mechanism is as follows: Perform a global average pooling GAP operation on Fe3 to compress the spatial dimension, then sequentially undergo 1×1 convolution, GELU activation, 1×1 convolution, and Sigmoid activation to generate a channel attention weight vector WCA. The channel attention weight vector is multiplied element-by-element by Fe3 to obtain the enhanced feature Fe4, which is expressed as: Fe4= Fe3·Sigmoid(PWConv(GELU(PWConv(GAP(Fe3))))); Among them, GELU represents the GELU activation function, PWConv represents 1×1 convolution, and Sigmoid represents the Sigmoid activation function.
4. The defogging method based on frequency information differential fusion according to claim 3 is characterized in that: In step S313, the process of the spatial attention SA mechanism is: 3×3 convolution, GELU activation, 3×3 convolution and Sigmoid activation are performed on Fe4 in sequence to generate a spatial attention weight map WSA. The spatial attention weight map is multiplied element-by-element with Fe4 to obtain the enhanced feature Fe5, which is expressed as: Fe5= Fe4·Sigmoid (Conv(GELU(Conv(Fe4)))).
5. The defogging method based on frequency information differential fusion according to claim 4 is characterized in that: In the decoder stage, three 3×3 convolution operations are performed on Fe5, and each convolution layer is followed by an upsampling operation to restore the spatial resolution of the feature map. The specific operations are as follows: Fd1 = ↑(DyT(conv(Fe5))); Fd2 = ↑(DyT(conv(Fd1))); Fd3 = ↑(DyT(conv(Fd2))); Among them, ↑ represents the upsampling operation, Fd1 and Fd2 are intermediate output features, and Fd3 is the decoder output feature.
6. The defogging method based on frequency information differential fusion according to claim 1 is characterized in that: The high- and low-frequency information differential fusion module in step S3 contains three inputs and one output, and fuses the high-frequency information HF and low-frequency information LL obtained by wavelet transform with the original feature RF extracted by the U-shaped feature extraction module; The operation process of the high- and low-frequency information differential fusion module is as follows: The high- and low-frequency information differential fusion module first performs differential processing on the three inputs it receives. The RF is input to the Convs module for further feature extraction to obtain the enhanced feature RF'; the LL is processed by the low-frequency enhancement channel attention block to enhance the global structure and illumination information in the low-frequency features to generate the enhanced feature LL'; the HF is sent to the high-frequency information amplification block to enhance the representation ability of the local structure of the image by amplifying the edge and detail information in the high-frequency features, and obtain the final feature HF'; LL' is multiplied by the learnable parameter a, HF' is multiplied by 1-a, and the two multiplication results are finally added to RF' to generate the final output O; the specific mathematical expression is as follows: RF' = Convs(RF); LL' = LFECA (LL); HF' = HFEB (HF); O = RF'+ a·LL'+(1-a)·HF'; Among them, Convs represents the Convs module, LFECA represents the low-frequency enhanced channel attention block, and HFEB represents the high-frequency amplification block.
7. The defogging method based on frequency information differential fusion according to claim 6 is characterized in that: The operation steps of the low-frequency enhancement channel attention block are as follows: the low-frequency information LL first extracts the spatial features of each channel independently through a 3×3 channel-by-channel convolution layer, focusing on the global structure of the low-frequency features; then, the inter-channel information is integrated through a 1×1 point-by-point convolution layer to further enhance the semantic expression of the features; finally, the low-frequency enhancement channel attention block introduces a channel attention mechanism, extracts channel-level global information through global average pooling, and uses two layers of 1×1 convolution with ReLU activation in the middle to generate channel attention weights, which are normalized to [0,1] by the Sigmoid function.
8. The defogging method based on frequency information differential fusion according to claim 7, characterized in that: The operation steps of the high-frequency information amplification block are as follows: the high-frequency information amplification block first performs a feature shift operation on the input HF, shifts the HF one pixel to the lower right, and fills the empty edge part with zeros to generate a shifted feature map; then, the shifted feature map is subtracted from the HF element by element, and the absolute value of the result is taken to obtain a difference map; then, the difference map is input into a sequence containing convolution, batch normalization, GELU activation function and convolution for further processing; after that, the feature is normalized by Sigmoid activation and added to the all-one matrix, and then multiplied with the HF element by element to obtain the enhanced feature; finally, the enhanced feature is residually connected with the convolved HF to generate the final feature HF'.
9. The defogging method based on frequency information differential fusion according to claim 1, characterized in that: In step S4, a staged training strategy is introduced. The staged training strategy optimizes the multi-scale feature learning and overall performance of the dehazing network by combining deep supervision and curriculum learning. The staged training strategy divides the training process into four stages with n rounds as the boundary. The specific steps are as follows: In the first stage, the loss is calculated for the low-resolution output and the corresponding low-resolution real image, guiding the network to prioritize learning the basic structure and global features of the image; In the second stage, supervision is shifted to medium-resolution output, and the loss is calculated with the corresponding medium-resolution real image, gradually enhancing the ability to restore local textures and details; In the third stage, supervision focuses on the high-resolution output and calculates the loss with the corresponding high-resolution real image to ensure the overall consistency and fine-grained details of the dehazing results; In the fourth stage, the losses of all resolution outputs are optimized to enable the network to fully capture multi-scale features.
Citation Information
Patent Citations
HFRC-Diff-based low-illumination superposition fog image enhancement method
CN119048381A
Double-branch defogging method based on Laplacian pyramid
CN119887582A
Cited By
Haze concentration detection method based on hyperbolic tangent function image
CN121527039A
A method for detecting image haze concentration based on hyperbolic tangent function
CN121527039B