Dual-branch Dehazing Method Based on Laplacian Pyramid
By introducing a two-branch structure based on Laplace pyramid in the image defogging method, the main features and high and low-frequency features of the image are extracted and fused, the problem of limitations in structural information extraction in the prior art is solved, and a high-quality image defogging effect is achieved.
Patent Information
- Application Number
- CN202510370165.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing deep learning-based image defogging method has limitations in extracting structural information, which leads to distortion of reconstruction results in complex areas and affects the defogging effect.
The dual-branch defogging method based on the Laplace pyramid is adopted, and the main features and high and low-frequency lighting characteristics of the image are extracted respectively through multi-feature fusion focusing branches and Laplace pyramid complementary branches, and the adaptive fusion of the main features and high and low-frequency characteristics is achieved through the learnable feature fusion module.
This method can fully capture the global structure information of the image, the generated foggy-free image will not be distorted in the local area, and the restored image detail information is close to the real image.
Smart Images

Figure CN119887582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image processing, and specifically provides a dual-branch dehazing method based on Laplacian pyramid. Background Art
[0002] Due to the presence of a large number of suspended particulate matters in the atmosphere under bad weather conditions such as haze, images are often greatly affected, resulting in problems such as decreased visual quality, blurring, color shift, decreased contrast, and buried details. This phenomenon will seriously affect the performance of computer vision systems. Therefore, image dehazing technology has become an important research direction in the field of image enhancement.
[0003] Image dehazing is an important low-level vision task, aiming to restore haze-free images from hazy images. In a hazy environment, due to the presence of suspended particulate matters in the air, the captured images will have color distortion, detail loss, etc. When such low-quality images are applied to high-level vision tasks such as object detection, semantic segmentation, or autonomous driving, the algorithm performance will be significantly reduced. Image dehazing restores hazy images to haze-free images through image processing technology, thereby effectively improving the execution accuracy of subsequent vision tasks. In recent years, accurately clear dehazed images have also attracted more and more attention.
[0004] With the wide application of deep learning in image processing tasks, many scholars have proposed many deep learning-based dehazing methods. Deep learning-based dehazing methods utilize their powerful non-linear fitting ability to automatically learn prior knowledge from a large amount of image data and directly map hazy images to haze-free images. Existing deep learning-based dehazing methods mainly use convolutional neural networks (CNNs). However, due to the continuous downsampling of convolutional neural networks, some structural information is easily lost, resulting in serious distortion of the reconstruction results in some complex regions, thereby affecting the dehazing effect.
[0005] In order to overcome the limitations of convolutional neural networks in extracting structural information, in the prior art:
[0006] 1. SDBAD-NET (Spatial Dual-Branch Attention Dehazing Network) in the prior art is a technology widely used in the fields of deep learning and computer vision, aiming to improve the performance of the model by dynamically adjusting and optimizing feature representations. By dynamically fusing shallow structural features and deep context semantic information, the structural loss of convolutional neural networks is compensated.
[0007] 2. The Dynamic Feature Enhancement Module (DFE) in the prior art is a defogging network based on the Meta-Former paradigm, aiming to solve the problem of image defogging in complex scenarios. Among them, deformable convolution is introduced to dynamically expand the receptive field and fuse more structured information.
[0008] The above methods consider the importance of structural information in image restoration and retain the structural details of the original image to a certain extent. However, limited by the receptive field of the convolution operation, these methods lack the ability of global modeling and are difficult to fully capture the global structural information of the image. Therefore, the defogged images generated by these methods may be distorted in local areas. Summary of the Invention
[0009] To solve the above problems, the present invention provides a dual-branch defogging method based on the Laplacian pyramid, which can effectively restore the detailed information of the image and generate a defogged image close to the real image.
[0010] A dual-branch defogging method based on the Laplacian pyramid provided by the present invention includes the following steps:
[0011] S1. Establish a training dataset A containing the original clear image and the corresponding foggy image;
[0012] S2. Perform data augmentation on the training dataset A; including cropping and data augmentation on the input image;
[0013] S3. Construct a defogging network LPSDNet based on the Laplacian pyramid;
[0014] The defogging network LPSDNet includes a multi-feature fusion and focusing branch and a Laplacian pyramid supplementary branch; when the multi-feature fusion and focusing branch and the Laplacian pyramid supplementary branch complete feature extraction, the features learned by the multi-feature fusion and focusing branch and the Laplacian pyramid supplementary branch are further fused through a learnable fusion module;
[0015] The multi-feature fusion and focusing branch includes an encoder, a decoder, and skip connections; among them, the encoder includes multiple 3×3 convolutional layers, a supplementary fusion module, and a hybrid focusing attention module, and the main feature information is extracted through the encoder, the size of the feature map is reduced, and the number of feature channels is increased; the decoder restores the feature map output by the encoder to the same size as the original image through an upsampling layer, an SK fusion module, a hybrid focusing attention module, and multiple 3×3 convolutional layers; a skip connection of a 1×1 convolutional layer is set between the encoder and the decoder to dynamically fuse shallow and deep features;
[0016] S4. Use the training dataset A to train the haze removal network LPSDNet until the pre-set loss function converges; combine the L1 loss function and the structural loss function as the objective function for network training to obtain the trained haze removal network LPSDNet;
[0017] S5. Input the image to be dehazed into the trained haze removal network LPSDNet for testing and verification to obtain the dehazing result.
[0018] Furthermore, the supplementary fusion module in step S3 includes two inputs and one output, and fuses the structural information obtained by decomposing the Laplacian pyramid supplementary branch with the features extracted by the current branch; the operation steps of the supplementary fusion module include:
[0019] S31a. For the feature F extracted by the multi-feature fusion focusing branch d and the image feature F obtained by decomposing the Laplacian pyramid supplementary branch p , first perform a 1×1 convolution operation on the input feature F p to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map ;
[0020] S32a. Concatenate the features of and F d in the channel dimension to obtain the concatenated feature F cat ; at the same time, perform an addition operation on and F d in the pixel dimension to obtain the feature F add ;
[0021] S33a. Perform a convolution operation on the concatenated feature F catt to achieve the fusion of the features and F d ;
[0022] S34a. Use the spatial attention module to generate a feature attention weight map from the fused features; among them, the spatial attention module consists of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence;
[0023] S35a. Multiply the obtained attention weight map by F add to obtain the feature F frame , highlighting the spatial structure information in the feature F add ;
[0024] S36a. Add the feature F frame and the feature F d to obtain the output feature F output of the supplementary fusion module.
[0025] Furthermore, the learnable fusion module in step S3 has two inputs and one output, and fuses the structural information obtained by decomposing the Laplacian pyramid supplementary branch with the features extracted by the current branch. The operation steps of the learnable fusion module are as follows:
[0026] S31b. For the features F d extracted by the multi-feature fusion focusing branch p and the image features F p obtained by decomposing the Laplacian pyramid supplementary branch, first perform a 1×1 convolution operation on the input feature F
[0027] to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map; d S32b. Concatenate the features of and F cat in the channel dimension to obtain the concatenated feature F d ; at the same time, perform an addition operation on and F add in the pixel dimension to obtain the feature F
[0028] S33b. Perform a convolution operation on the concatenated feature F cat to realize the fusion of the features and F d ;
[0029] S34b. Use the spatial attention module to generate a feature attention weight map from the fused features; among them, the spatial attention module consists of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence;
[0030] S35b. Multiply the obtained attention weight map by F add to obtain the feature F frame and highlight the spatial structure information in the feature F add ;
[0031] S36b. Introduce a learnable parameter a, multiply the parameter a by the feature F frame and then add it to the original feature F d ;
[0032] Furthermore, the network in step S3 introduces the multi-scale structural information extracted by the Laplacian pyramid supplementary branch into the multi-feature fusion focusing branch. The Laplacian pyramid supplementary branch uses the Laplacian pyramid decomposition module to perform multi-scale decomposition on the input foggy image to obtain four different-scale features. The output of the Laplacian pyramid supplementary branch is obtained by the following steps:
[0033] Among the four different-scale features, the first three multi-scale features are passed to the corresponding supplementary fusion modules of the multi-feature fusion focusing branch, and the last multi-scale feature is passed to the corresponding supplementary fusion module of the multi-feature fusion focusing branch after doubling the size of the feature map through pixel padding operation;
[0034] S32c. Enhance the high-frequency features and low-frequency features by using the supplementary fusion module and the frequency enhancement module;
[0035] S33c. Then reconstruct the enhanced high-frequency features and low-frequency features by using the Laplacian pyramid reconstruction module to obtain the output of the Laplacian pyramid supplementary branch.
[0036] Furthermore, high-frequency and low-frequency feature information is extracted through the Laplacian pyramid decomposition module and the Laplacian pyramid reconstruction module. Specifically, for the input image I, the Laplacian pyramid decomposition module decomposes the input image I into a low-frequency component and several high-frequency components, and the calculation process is as follows:
[0037] The calculation formula for the low-frequency component L is: L = G N (I), where represents the result of performing N times of Gaussian blur on the image I;
[0038] The calculation formula for the high-frequency component Hi of the i-th layer is as follows:
[0039]
[0040] where 1 ≤ i < N; represents the Gaussian pyramid image of the i-th layer, represents performing blur processing using a 5×5 two-dimensional Gaussian kernel, represents upsampling the image to the original size;
[0041] Gaussian pyramid The calculation formula is:
[0042]
[0043] where represents downsampling the image to half of the original size;
[0044] The Laplacian pyramid reconstruction is the inverse process of the above decomposition.
[0045] Furthermore, the frequency enhancement module in step S32c includes two residual blocks for extracting global structural information and focusing on key features. Specifically: The first residual block normalizes the input feature map x using a normalization layer; three convolutional layers with convolutional kernels of 3×3, 5×5, and 7×7 respectively are used to extract multi-scale features of the image in parallel, and the obtained multi-scale features are concatenated to obtain the feature map x2; the input feature x2 is added to the input feature x to obtain the output feature y of the first residual block, and convolutional operations with different convolutional kernel sizes are used to extract high-frequency and low-frequency feature information in the image; the second residual block includes a spatial attention layer for strengthening the captured features and making full use of the frequency features.
[0046] Furthermore, the hybrid focus attention module in step S3 includes two residual blocks. The operation steps of the hybrid focus attention module are as follows:
[0047] The first residual block:
[0048] S31d. For the input feature map x, the first residual block normalizes it using a normalization layer;
[0049] S32d. Three convolutional layers with convolutional kernels of 3×3, 5×5, and 7×7 respectively are used to extract multi-scale features of the image in parallel, and the obtained multi-scale features are concatenated to obtain the feature map x2;
[0050] S33d. The input feature x2 is added to the input feature x to obtain the output feature y of the first residual block;
[0051] The second residual block:
[0052] S34d. The second residual block further extracts information from the feature map y using multiple attention layers;
[0053] S35d. The feature map y is normalized using normalization to obtain the feature y1;
[0054] S36d. The feature y1 is respectively input into the tangent attention, square attention, and re-attention layers to obtain different focused features F1, F2, and F3;
[0055] S37d. The focused features F1, F2, and F3 are concatenated, and the channel attention layer is used to further extract the key features in the image to obtain the feature y3; the channel attention layer consists of global average pooling, a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence; among them, the first 1×1 convolution changes the number of channels of the feature map, and the second 1×1 convolution is used for feature extraction.
[0056] S38d, y3 passes through a multi-layer perceptron and is added to the feature y to obtain the output O of the final hybrid focus attention module.
[0057] Furthermore, in step S36d, the tangent attention layer consists of a residual connection and two consecutive 1×1 convolutional layers. The first convolutional layer is followed by a GELU activation function, and the second convolutional layer is followed by a Tanh activation function. The formula of the tangent attention layer is expressed as follows:
[0058] F1 = y1 ⊙ Tanh(PWConv(GELU(PWConv(X))))
[0059] Where Tanh represents the hyperbolic tangent function, and ⊙ represents element-wise multiplication; the hyperbolic tangent function can map the attention weights to the range between -1 and 1. Among them, the weight values less than zero indicate that the data at the corresponding input positions are unimportant or harmful to the current task.
[0060] Furthermore, in step S36d, the formula of the square attention layer is expressed as follows:
[0061] F2 = y1 ⊙ Sigmoid(PWConv(GELU(PWConv(X)))) 2
[0062] Where Sigmoid represents the Sigmoid activation function. In the square attention layer, the Sigmoid activation function can map the feature values to the attention weights from 0 to 1.
[0063] Furthermore, in step S36d, the re-attention layer uses consecutive spatial attention to gradually focus the network on important features, which specifically includes the following steps:
[0064] S36d-1: The re-attention layer applies spatial attention to obtain preliminary attention weights;
[0065] S36d-2: Multiply the preliminary attention weights by the input features to obtain intermediate features;
[0066] S36d-3: Perform a spatial attention operation on the intermediate features to obtain refined attention weights;
[0067] S36d-4: Multiply the obtained attention weights by the input to obtain the aggregated feature F3; the formula of the aggregated feature F3 is expressed as follows:
[0068] X^ = y1 ⊙ Sigmoid(PWConv(GELU(PWConv(X))))
[0069] F3 = y1⊙Sigmoid(PWConv(GELU(PWConv(X^))))
[0070] Wherein, X^ represents the intermediate feature generated by the re-attention layer.
[0071] Compared with the prior art, the present invention can achieve the following beneficial effects: The image dehazing method in the present invention can respectively extract the main features and high- and low-frequency illumination features of the image through the multi-feature fusion focusing branch and the Laplacian pyramid supplementary branch, and then through the learnable feature fusion module, realize the adaptive fusion of the main features and the high- and low-frequency features. Finally, the final dehazed image is obtained by using the reconstruction technology. It has the ability of global modeling, can fully capture the global structure information of the image, and the generated haze-free image will not be distorted in the local area. Description of the Drawings
[0072] Figure 1 is a schematic structural diagram of a Laplacian pyramid double-branch network provided by an embodiment of the present invention;
[0073] Figure 2 is a schematic structural diagram of a supplementary fusion module provided by an embodiment of the present invention;
[0074] Figure 3 is a schematic structural diagram of a hybrid focusing attention module provided by an embodiment of the present invention;
[0075] Figure 4 is a schematic structural diagram of a frequency enhancement module provided by an embodiment of the present invention;
[0076] Figure 5 is an example of the dehazing result provided by an embodiment of the present invention; wherein, (a) is a hazy image, (b), (c), and (d) respectively represent the images after dehazing by the DCP, FFANet, and GridDehaze methods, (e) is the image after dehazing by the present invention, and (f) is the original haze-free image;
[0077] Figure 6 is the overall flowchart of the double-branch dehazing method provided by an embodiment of the present invention;
[0078] Figure 7 is Figure 5 the enlarged view of the upper part of the upper figure in FIG. (a);
[0079] Figure 8 is Figure 5 the enlarged view of the upper part of the upper figure in FIG. (b);
[0080] Figure 9 is Figure 5 the enlarged view of the upper part of the upper figure in FIG. (c);
[0081] Figure 10 is Figure 5 an enlarged view of the upper figure in FIG. (d) of
[0082] Figure 11 is Figure 5 an enlarged view of the upper figure in FIG. (e) of
[0083] Figure 12 is Figure 5 an enlarged view of the upper figure in FIG. (f) of
[0084] Figure 13 is Figure 5 an enlarged view of the middle figure in FIG. (a) of
[0085] Figure 14 is Figure 5 an enlarged view of the middle figure in FIG. (b) of
[0086] Figure 15 is Figure 5 an enlarged view of the middle figure in FIG. (c) of
[0087] Figure 16 is Figure 5 an enlarged view of the middle figure in FIG. (d) of
[0088] Figure 17 is Figure 5 an enlarged view of the middle figure in FIG. (e) of
[0089] Figure 18 is Figure 5 an enlarged view of the middle figure in FIG. (f) of
[0090] Figure 19 is Figure 5 an enlarged view of the lower figure in FIG. (a) of
[0091] Figure 20 is Figure 5 an enlarged view of the lower figure in FIG. (b) of
[0092] Figure 21 is Figure 5 an enlarged view of the lower figure in FIG. (c) of
[0093] Figure 22 is Figure 5 an enlarged view of the lower figure in FIG. (d) of
[0094] Figure 23 is Figure 5 an enlarged view of the lower figure in FIG. (e) of
[0095] Figure 24 is Figure 5 an enlarged view of the lower figure in FIG. (f) of
[0096] Figure 25Yes Figure 1 Magnified view of the foggy image;
[0097] Figure 26 Yes Figure 1 Magnified view of the de - fogged image;
[0098] Figure 27 Yes Figure 1 Magnified view of the H×W×3 image;
[0099] Figure 28 Yes Figure 1 Magnified view of the H / 2×W / 2×3 image;
[0100] Figure 29 Yes Figure 1 Magnified view of the H / 4×W / 4×3 image;
[0101] Figure 30 Yes Figure 1 Magnified view of the H / 8×W / 8×3 image;
[0102] Figure 31 Yes Figure 2 Magnified view of Fp;
[0103] Figure 32 Yes Figure 2 Magnified view of Fd;
[0104] Figure 33 Yes Figure 2 Magnified view of the output;
[0105] Figure 34 Yes Figure 1 Magnified view of the Laplacian pyramid supplementary branch. Detailed implementation method
[0106] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail in conjunction with the attached Figures 1 - 34 drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0107] A double - branch de - fogging method based on Laplacian pyramid, as Figure 6 shown, includes the following steps:
[0108] S1. Establish a training data set A containing the original clear image and the corresponding foggy image.
[0109] S2. Perform data augmentation processing on the training data set A; including cropping and data augmentation processing on the input image, specifically including: for an image with a size of The input image is randomly cropped into image patches of 256×256, and data augmentation is performed by randomly rotating it by 90°, 180°, or 270°.
[0110] S3. Construct a haze removal network LPSDNet based on the Laplacian pyramid, and the specific structure is as Figure 1 shown.
[0111] The haze removal network LPSDNet includes a multi-feature fusion and focusing branch and a Laplacian pyramid supplementary branch; after the multi-feature fusion and focusing branch and the Laplacian pyramid supplementary branch complete feature extraction, the features learned by the multi-feature fusion and focusing branch and the Laplacian pyramid supplementary branch are further fused through a learnable fusion module.
[0112] The multi-feature fusion and focusing branch includes an encoder, a decoder, and skip connections. Among them, the encoder includes multiple 3×3 convolutional layers, a supplementary fusion module (SFM), and a mixed focusing attention module (MFAM). The main feature information is extracted through the encoder, the size of the feature map is reduced, and the number of feature channels is increased. The decoder restores the feature map output by the encoder to the same size as the original image through an upsampling layer, an SK fusion module, a mixed focusing attention module, and multiple 3×3 convolutional layers. In order to retain spatial detail information, a skip connection of 1×1 convolutional layer is set between the encoder and the decoder to dynamically fuse shallow and deep features.
[0113] The supplementary fusion module in step S3 is as Figure 2 shown, including two inputs and one output, and fusing the structural information decomposed by the Laplacian pyramid supplementary branch with the features extracted by the current branch; the operation steps of the supplementary fusion module include:
[0114] S31a. For the feature F d extracted by the multi-feature fusion and focusing branch and the image feature F p decomposed by the Laplacian pyramid supplementary branch, first perform a 1×1 convolution operation on the input feature F p to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map;
[0115] S32a. Perform feature concatenation on and F d in the channel dimension to obtain the concatenated feature F cat ; at the same time, perform an addition operation on and F d in the pixel dimension to obtain the feature F add ;
[0116] S33a. Perform a convolution operation on the concatenated feature F cat to achieve effective fusion of the feature and F d ;
[0117] S34a. Generate a feature attention weight map from the fused features using a spatial attention module; wherein, the spatial attention module consists of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence;
[0118] S35a. Multiply the obtained attention weight map by F add to obtain the feature F frame , highlighting the spatial structure information in the feature F add ;
[0119] S36a. Add the feature F frame and the feature F d to obtain the output feature F output of the supplementary fusion module.
[0120] To enhance the structural features, the network introduces the multi-scale structural information extracted by the Laplacian pyramid supplementary branch into the multi-feature fusion focusing branch. The Laplacian pyramid supplementary branch uses a Laplacian pyramid decomposition module to perform multi-scale decomposition on the input hazy image to obtain four features of different scales. The output of the Laplacian pyramid supplementary branch is obtained through the following steps:
[0121] S31c. Among the four features of different scales, the first three multi-scale features are directly passed to the corresponding supplementary fusion module of the multi-feature fusion focusing branch, and the last multi-scale feature is passed to the corresponding supplementary fusion module of the multi-feature fusion focusing branch after doubling the feature map size through pixel padding operation;
[0122] S32c. Use the supplementary fusion module and the frequency enhancement module (FEM) to enhance the high-frequency features and low-frequency features;
[0123] S33c. Then use the Laplacian pyramid reconstruction module to reconstruct the enhanced high-frequency features and low-frequency features to obtain the output of the Laplacian pyramid supplementary branch.
[0124] After the multi-feature fusion focusing branch and the Laplacian pyramid supplementary branch complete feature extraction, the features learned by the two branches are further fused through the learnable fusion module in step S3. The learnable fusion module has two inputs and one output, and fuses the structural information decomposed by the Laplacian pyramid supplementary branch with the features extracted by the current branch; the operation steps of the learnable fusion module include:
[0125] S31b. For the feature F d extracted by the multi-feature fusion focusing branch and the image feature F p decomposed by the Laplacian pyramid supplementary branch, first, for the input feature F pPerform a 1×1 convolution operation to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map ;
[0126] S32b. On the channel dimension, and F d Perform feature concatenation to obtain the concatenated feature F cat ; At the same time, on the pixel dimension, and Fd are added to obtain the feature F add ;
[0127] S33b. Perform a convolution operation on the concatenated feature F cat to achieve the fusion of features and F d ;
[0128] S34b. Use the spatial attention module to generate a feature attention weight map from the fused features; among them, the spatial attention module consists of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence;
[0129] S35b. Multiply the obtained attention weight map by F add to obtain the feature F frame , highlighting the spatial structure information in the feature F add ;
[0130] S36b. Introduce a learnable parameter a, multiply the parameter a by the feature F frame and then add it to the original feature F d . Introducing the learnable parameter a can enable the model to dynamically adjust the relative importance of the outputs of the two branches, achieve a more flexible and effective fusion of the main features and high- and low-frequency features, and generate a clearer and more accurate dehazed image.
[0131] Extract high-frequency and low-frequency feature information through the Laplacian pyramid decomposition module and the Laplacian pyramid reconstruction module. Specifically, for the input image I, the Laplacian pyramid decomposition module decomposes the input image I into a low-frequency component and several high-frequency components. The calculation process is as follows:
[0132] The calculation formula for the low-frequency component L is: L = G N (I), where, represents the result of performing N Gaussian blurs on the image I;
[0133] The calculation formula for the high-frequency component Hi of the i-th layer is as follows:
[0134]
[0135] where 1 ≤ i < N; Represents the Gaussian pyramid image of the i-th layer. Indicates blurring using a 5×5 two-dimensional Gaussian kernel. Indicates upsampling the image to the original size.
[0136] Gaussian pyramid The calculation formula is:
[0137]
[0138] Where Indicates downsampling the image to half of the original size;
[0139] The Laplacian pyramid is reconstructed as the inverse process of the above decomposition.
[0140] In step S3, the hybrid focus attention module is as Figure 3 shown, including two residual blocks. The operation steps of the hybrid focus attention module include:
[0141] The first residual block:
[0142] S31d. For the input feature map x, the first residual block normalizes it using a normalization layer;
[0143] S32d. Adopt three convolutional layers with convolutional kernels of 3×3, 5×5, and 7×7 respectively to extract multi-scale features of the image in parallel, and splice the obtained multi-scale features to obtain the feature map x2;
[0144] S33d. Add the input feature x2 to the input feature x to obtain the output feature y of the first residual block;
[0145] S34d. The second residual block further extracts information from the feature map y using multiple attention layers;
[0146] The second residual block:
[0147] S35d. Normalize the feature map y using normalization to obtain the feature y1;
[0148] S36d. Input the feature y1 into the tangent attention, square attention, and re-attention layers respectively to obtain different focused features F1, F2, and F3;
[0149] S37d. Splice the focused features F1, F2, and F3, and use a channel attention layer to further extract the key features in the image to obtain the feature y3; The channel attention layer consists of global average pooling, 1×1 convolution, GELU activation function, 1×1 convolution, and Sigmoid function in sequence; Among them, the first 1×1 convolution is to change the number of channels of the feature map, and the second 1×1 convolution is to extract features.
[0150] After passing through a multi-layer perceptron, S38d and y3 are added to the feature y to obtain the output O of the final hybrid focus attention module.
[0151] In step S36d, the tangent attention layer consists of a residual connection and two consecutive 1×1 convolutional layers. The first convolutional layer is followed by a GELU activation function, and the second convolutional layer is followed by a Tanh activation function. The formula of the tangent attention layer is expressed as follows:
[0152] F1 = y1⊙Tanh(PWConv(GELU(PWConv(X))))
[0153] Among them, Tanh represents the hyperbolic tangent function, and ⊙ represents element-wise multiplication; the hyperbolic tangent function can map the attention weights to the range between -1 and 1. Among them, the weight values less than zero indicate that the data at the corresponding input positions are unimportant or harmful to the current task. Therefore, the attention weight map generated by the Tanh activation function can effectively suppress irrelevant or harmful signals using negative weight values and has stronger feature discrimination ability.
[0154] In step S36d, the formula of the square attention layer is expressed as follows:
[0155] F2 = y1⊙Sigmoid(PWConv(GELU(PWConv(X)))) 2
[0156] Among them, Sigmoid represents the Sigmoid activation function. In the square attention layer, the Sigmoid activation function can map the feature values to attention weights from 0 to 1. The square attention layer performs a point-wise square operation on the generated attention weight map. Although the point-wise square operation will make all the values in the attention weight map smaller, the values close to 0 will become even smaller compared to the values close to 1. This is equivalent to increasing the larger values in the attention weight map and weakening the smaller values in the attention weight map, thereby focusing the model's attention more on the key regions.
[0157] In step S36d, the re-attention layer uses consecutive spatial attention to gradually focus the network on important features, specifically including the following steps:
[0158] S36d-1: The re-attention layer applies spatial attention to obtain preliminary attention weights;
[0159] S36d-2: Multiply the preliminary attention weights by the input features to obtain intermediate features;
[0160] S36d-3: Perform a spatial attention operation on the intermediate features to obtain fine attention weights;
[0161] S36d-4: Multiply the obtained attention weights with the input to obtain the aggregated feature F3; the formula for the aggregated feature F3 is expressed as follows:
[0162] X^ = y1 ⊙ Sigmoid(PWConv(GELU(PWConv(X))))
[0163] F3 = y1 ⊙ Sigmoid(PWConv(GELU(PWConv(X^))))
[0164] where X^ represents the intermediate feature generated by the re-attention layer. After passing X^ through an attention mechanism to generate an attention weight map and then performing an element-wise multiplication operation with the input feature y1, the network can suppress the useless information in the feature map again, focus on the key information, and effectively improve the attention degree to the key regions of the image by adopting the continuous attention method, enhance the expression of useful information, and suppress the interference of useless information.
[0165] The frequency enhancement module in step S32C is as Figure 4 shown. Specifically, the frequency enhancement module includes two residual blocks for extracting global structural information and focusing on key features. Specifically, the first residual block normalizes the input feature map x using a normalization layer; three convolutional layers with convolutional kernels of 3×3, 5×5, and 7×7 respectively are used to extract the multi-scale features of the image in parallel, and the obtained multi-scale features are concatenated to obtain the feature map x2; the input feature x2 is added to the input feature x to obtain the output feature y of the first residual block, and convolutional operations with different convolutional kernel sizes are used to extract the high-frequency and low-frequency feature information in the image; the second residual block contains a spatial attention layer for strengthening the captured features and making full use of the frequency features.
[0166] S4. Use the training dataset A to train the dehazing network LPSDNet. Specifically, put the training dataset A into the dehazing network LPSDNet for training until the pre-set loss function converges; combine the L1 loss function and the structural loss function as the objective function for network training to obtain the trained dehazing network LPSDNet.
[0167] S5. Input the image to be dehazed into the trained dehazing network LPSDNet for testing and verification to obtain the dehazing result and output the dehazed image.
[0168] The above method is used to conduct experiments on the RESIDE-IN (ITS) public dataset. ITS includes 13,990 pairs of clean images and foggy images of indoor scenes. The foggy images are generated from the clean images according to the atmospheric scattering model, where the medium extinction coefficient β is selected in the range of [0.6, 1.8], and the global atmospheric light A is selected in the range of [0.7, 1.0]. We train the proposed method for 500 rounds and test the trained network on the SOTS dataset.
[0169] Figure 5 The results of partial image dehazing on the test dataset are shown. Among them, (a) is the foggy image, and (b), (c), and (d) are the images after dehazing by the DCP, FFANet, and GridDehaze methods respectively. (e) is the image after dehazing by the present invention, and (f) is the original fog-free image. It can be seen from the figure that the dehazed images generated by other methods contain artifacts left over from incomplete dehazing, while the method proposed by the present invention can effectively restore the detailed information of the image and generate a dehazed image close to the real image.
[0170] Figure 5 The enlarged view of each figure in Figures 7 - 24 is specifically as Figure 1 shown. The enlarged view of the image in Figures 25 - 30 is specifically as Figure 2 shown. The enlarged view of the image in Figures 31 - 33 is specifically as
[0171] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A dual-branch dehazing method based on Laplacian pyramid, characterized in that: The steps include: S1. Establish a training data set A including original clear images and corresponding foggy images; S2, performing data enhancement processing on the training data set A, including cropping and data enhancement processing on the input image; S3. Construct a dehazing network LPSDNet based on Laplacian pyramid; The dehazing network LPSDNet includes a multi-feature fusion focusing branch and a Laplacian pyramid supplementing branch. After the multi-feature fusion focusing branch and the Laplacian pyramid supplementing branch complete feature extraction, the features learned by the multi-feature fusion focusing branch and the Laplacian pyramid supplementing branch are further fused through the learnable fusion module. The multi-feature fusion focusing branch includes an encoder, a decoder, and a jump connection; the encoder includes multiple 3×3 convolutional layers, a supplementary fusion module, and a hybrid focus attention module, through which the encoder extracts the main feature information, reduces the size of the feature map, and increases the number of feature channels; the decoder restores the feature map output by the encoder to the same size as the original image through an upsampling layer, an SK fusion module, a hybrid focus attention module, and multiple 3×3 convolutional layers; a jump connection of a 1×1 convolutional layer is set between the encoder and the decoder to dynamically fuse shallow and deep features; S4. Use the training data set A to train the defogging network LPSDNet until the preset loss function converges; combine the L1 loss function and the structural loss function as the objective function of the network training to obtain the trained defogging network LPSDNet; S5, input the image to be defogged into the trained defogging network LPSDNet for testing and verification to obtain the defogging result; In step S3, the network introduces the multi-scale structural information extracted by the Laplacian pyramid supplementary branch into the multi-feature fusion focusing branch. The Laplacian pyramid supplementary branch uses the Laplacian pyramid decomposition module to perform multi-scale decomposition on the input foggy image to obtain features of four different scales. The output of the Laplacian pyramid supplementary branch is obtained by the following steps: S31c, among the four features of different scales, the first three multi-scale features are passed to the supplementary fusion module corresponding to the multi-feature fusion focusing branch, and the last multi-scale feature is doubled in size after the pixel padding operation and then passed to the supplementary fusion module corresponding to the multi-feature fusion focusing branch; S32c, using a supplementary fusion module and a frequency enhancement module to enhance the high-frequency features and the low-frequency features; S33c, reconstruct the enhanced high-frequency features and low-frequency features using the Laplace pyramid reconstruction module to obtain the output of the Laplace pyramid supplementary branch.
2. The dual-branch defogging method based on Laplacian pyramid according to claim 1, characterized in that: The supplementary fusion module in step S3 includes two inputs and one output, and fuses the structural information obtained by decomposing the supplementary branch of the Laplacian pyramid with the features extracted by the current branch; the operation steps of the supplementary fusion module include: S31a, for the feature Fd extracted by the multi-feature fusion focus branch and the image feature Fp obtained by the Laplace pyramid supplementary branch decomposition, first perform a 1×1 convolution operation on the input feature Fp to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map S32a, in the channel dimension and Fd to obtain the concatenated feature Fcat; at the same time, Add it to Fd to get the feature Fadd; S33a, perform convolution operation on the concatenated feature Fcat to achieve feature Fusion with Fd; S34a, using a spatial attention module to generate a feature attention weight map from the fused features; wherein the spatial attention module is composed of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence; S35a, multiplying the obtained attention weight map by Fadd to obtain the feature Fframe, highlighting the spatial structure information in the feature Fadd; S36a, adding the feature Fframe to the feature Fd to obtain the output feature Foutput of the supplementary fusion module.
3. The double-branch defogging method based on Laplacian pyramid according to claim 1, characterized in that: The learnable fusion module in step S3 includes two inputs and one output, and fuses the structural information obtained by decomposing the supplementary branch of the Laplacian pyramid with the features extracted by the current branch; The operation steps of the learnable fusion module include: S31b, for the feature Fd extracted by the multi-feature fusion focus branch and the image feature Fp obtained by the Laplacian pyramid supplementary branch decomposition, first perform a 1×1 convolution operation on the input feature Fp to expand the number of channels of the feature map; then perform a 3×3 convolution operation to obtain the feature map S32b, in the channel dimension and Fd to obtain the concatenated feature Fcat; at the same time, Add it to Fd to get the feature Fadd; S33b, perform convolution operation on the concatenated feature Fcat to achieve feature Fusion with Fd; S34b, using a spatial attention module to generate a feature attention weight map from the fused features; wherein the spatial attention module is composed of a 1×1 convolution, a GELU activation function, a 1×1 convolution, and a Sigmoid function in sequence; S35b, multiplying the obtained attention weight map by Fadd to obtain the feature Fframe, highlighting the spatial structure information in the feature Fadd; S36b. Introduce a learnable parameter a, multiply the learned parameter a with the feature Fframe and then add it to the original feature Fd.
4. The double-branch defogging method based on Laplacian pyramid according to claim 1, characterized in that: The high-frequency and low-frequency feature information is extracted through the Laplacian pyramid decomposition module and the Laplacian pyramid reconstruction module. Specifically, for the input image I, the Laplacian pyramid decomposition module decomposes the input image I into a low-frequency component and several high-frequency components. The calculation process is as follows: The calculation formula of the low-frequency component L is: L = GN (I), where G N (·) represents the result of Gaussian blurring image I N times; The calculation formula of the high-frequency component Hi of the i-th layer is as follows: H i =G i (I)-B(G i+1 (I)↑2) Among them, 1≤i <N;G i (·) represents the Gaussian pyramid image of the i-th layer, B(·) represents the blurring process using a 2D Gaussian kernel of size 5×5, and ↑2 represents upsampling the image to the original size; Gaussian Pyramid G i The calculation formula of (I) is: Among them, ↓2 means downsampling the image to half the size of the original image; The Laplacian pyramid reconstruction is the inverse process of the above decomposition.
5. The double-branch defogging method based on Laplacian pyramid according to claim 1, characterized in that: The frequency enhancement module in step S32c includes two residual blocks for extracting global structural information and focusing on key features; Specifically: the first residual block uses a normalization layer to normalize the input feature map x; three convolution layers with convolution kernels of 3×3, 5×5, and 7×7 are used in parallel to extract multi-scale features of the image, and the obtained multi-scale features are concatenated to obtain the feature map x2; the input feature x2 is added to the input feature x to obtain the output feature y of the first residual block, and convolution operations with different convolution kernel sizes are used to extract high-frequency and low-frequency feature information in the image; the second residual block contains a spatial attention layer to enhance the captured features and make full use of frequency features.
6. The double-branch defogging method based on Laplacian pyramid according to claim 1, characterized in that: The hybrid focus attention module in step S3 includes two residual blocks, and the operation steps of the hybrid focus attention module include: The first residual block: S31d, for the input feature map x, the first residual block normalizes it using a normalization layer; S32d, using three convolution layers with convolution kernels of 3×3, 5×5, and 7×7 to extract multi-scale features of the image in parallel, and concatenating the obtained multi-scale features to obtain a feature map x2; S33d, adding the input feature x2 to the input feature x to obtain the output feature y of the first residual block; The second residual block: S34d, the second residual block uses multiple attention layers to further extract information from the feature map y; S35d, using normalization to normalize the feature map y to obtain feature y1; S36d, input feature y1 into the tangent attention layer, square attention layer, and re-attention layer respectively to obtain different focus features F1, F2, and F3; S37d, concatenate the focused features F1, F2, and F3, and use the channel attention layer to further extract the key features in the image to obtain feature y3; the channel attention layer is composed of global average pooling, 1×1 convolution, GELU activation function, 1×1 convolution, and Sigmoid function in sequence; S38d and y3 are added to the feature y after passing through the multi-layer perceptron to obtain the final output O of the hybrid focus attention module.
7. The double-branch defogging method based on Laplacian pyramid according to claim 6, characterized in that: In step S36d, the tangent attention layer consists of a residual connection and two consecutive 1×1 convolutional layers. The first convolutional layer is followed by a GELU activation function, and the second convolutional layer is followed by a Tanh activation function. The formula of the tangent attention layer is expressed as follows: F1=y1⊙Tanh(PWConv(GELU(PWConv(X)))) Among them, Tanh represents the hyperbolic tangent function, ⊙ represents pixel-by-pixel multiplication; the hyperbolic tangent function can map the attention weight to a range between -1 and 1, where a weight value less than zero indicates that the data at the corresponding input position is unimportant or harmful to the current task.
8. The double-branch defogging method based on Laplacian pyramid according to claim 6, characterized in that: In step S36d, the formula of the square attention layer is expressed as follows: F2=y1⊙Sigmoid(PWConv(GELU(PWConv(X)))) 2 Among them, Sigmoid represents the Sigmoid activation function. In the square attention layer, the Sigmoid activation function can map the feature value to an attention weight from 0 to 1.
9. The double-branch defogging method based on Laplacian pyramid according to claim 6, characterized in that: In step S36d, the attention layer uses continuous spatial attention to gradually focus the network on important features, which specifically includes the following steps: S36d-1: The attention layer applies spatial attention to get preliminary attention weights; S36d-2: Multiply the initial attention weight with the input feature to obtain the intermediate feature; S36d-3: Perform spatial attention operations on intermediate features to obtain refined attention weights; S36d-4: Multiply the obtained attention weight by the input to obtain the aggregate feature F3; the formula of the aggregate feature F3 is as follows: X^=y1⊙Sigmoid(PWConv(GELU(PWConv(X)))) F3=y1⊙Sigmoid(PWConv(GELU(PWConv(X^)))) Among them, X^ represents the intermediate features generated by the re-attention layer.
Citation Information
Patent Citations
Construction method of double-branch remote sensing image defogging network
CN115578280A
Unmanned aerial vehicle image defogging method based on global and local double-branch network
CN116542864A