Defogging method based on image block importance adaptive learning
Through the defog removal method based on the adaptive learning of image block importance, an adaptive image block learning network is constructed, which solves the problem of poor defog removal effect in complex real-life scenarios, and achieves excellent performance and robustness on the real fog image dataset.
Patent Information
- Application Number
- CN202510326343.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-27
AI Technical Summary
The existing single-image defog removal method does not perform well in complex real-life scenarios, especially when dealing with complex and diverse fog distributions, and the effect is not ideal.
Adaptive fog removal method based on image block importance is adopted, and the complex fog distribution in real images is processed by constructing an adaptive image block learning network, including an automatic haze generator and an adaptive density-aware defog network, using U-net network, density map generation module and fusion module, combined with multi-negative contrast defog loss, to process complex fog distribution in real images.
It achieves defog removal performance better than existing methods on various real fog image datasets, has the robustness of processing complex and diverse fog distributions, and the image defog removal effect is more accurate.
Smart Images

Figure CN120047359A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and specifically, to a defogging method based on adaptive learning of image block importance. Background Art
[0002] The task of single-image defogging aims to restore visual information from a single observed blurred image. Due to the complex haze distribution and the lack of real training data, the visual information restoration effect of most models in real haze scenes is not ideal enough. Therefore, in the task of single-image defogging, it is an essential step to adaptively learn and remove haze with different concentrations, and restore the original background, objects, etc. in the image. Most existing defogging methods can be divided into prior-based methods and data-driven methods.
[0003] Prior-based methods mainly use heuristics observed and summarized manually to constrain images, helping the defogging algorithm to map blurred images to clear images. However, these carefully designed prior knowledge often fails in complex real-world scenarios. Data-driven methods achieve excellent performance by training on large-scale datasets. These methods directly learn the mapping from blurred to clear images using their strong fitting ability, and incorporate various carefully designed modules to enhance the defogging performance. However, these methods highly rely on large and high-quality datasets, and often perform poorly when dealing with uniform and non-uniform real-world blurred images. Summary of the Invention
[0004] To overcome at least one deficiency in the prior art, this application provides a defogging method based on adaptive learning of image block importance.
[0005] In a first aspect, there is provided a defogging method based on adaptive learning of image block importance, including:
[0006] Constructing an adaptive image block learning network, where the adaptive image block learning network includes an automatic haze generator and an adaptive density-aware defogging network;
[0007] Training the adaptive image block learning network using a training dataset to obtain a trained adaptive density-aware defogging network, including:
[0008] The training dataset is input into the automatic haze generator to obtain an extended training dataset; the extended training dataset is input into the adaptive density-aware dehazing network to obtain the predicted clear image; the samples in the training dataset are image pairs, and the image pairs include clear image samples and blurred image samples. The automatic haze generator is used to generate an editable haze map guide based on the blurred image samples and combine it with the clear image samples to obtain a blurred composite image; the adaptive density-aware dehazing network includes a U-net network, a density map generation module, and a fusion module. An image patch enhancement module is set in the U-net network. The U-net network is used to extract features from the input image to obtain a feature map. The image patch enhancement module is used to extract patch-based embedded features. The density map generation module is used to generate a density map based on the feature map. The fusion module is used to fuse the density map and the feature map to obtain the predicted clear image;
[0009] The blurred image to be processed is input into the trained adaptive density-aware dehazing network to obtain a clear image.
[0010] In one embodiment, the automatic haze generator includes a plurality of sequentially connected residual groups, and each residual group includes a plurality of sequentially connected residual blocks; the residual block includes a convolutional layer and an attention mechanism.
[0011] In one embodiment, the U-net network includes an encoder and a decoder. The encoder includes a first encoding layer, a second encoding layer, and a third encoding layer connected in sequence. The decoder includes a first decoding layer, a second decoding layer, and a third decoding layer connected in sequence;
[0012] The image patch enhancement module includes a pre-trained image patch enhancement unit, a first image patch enhancement unit, a second image patch enhancement unit, and a third image patch enhancement unit;
[0013] The first encoding layer and the third decoding layer are connected through a pre-trained image patch enhancement unit and a first image patch enhancement unit; the second encoding layer and the second decoding layer are connected through a second image patch enhancement unit; the third encoding layer and the first decoding layer are connected through a third image patch enhancement unit.
[0014] In one embodiment, the pre-trained image patch enhancement unit and the first image patch enhancement unit have the same structure, including:
[0015] A micro encoder, a micro decoder, and a dilated convolutional layer. The micro encoder includes a first micro encoding layer, a second micro encoding layer, and a third micro encoding layer connected in sequence. The micro decoder includes a first micro decoding layer, a second micro decoding layer, and a third micro decoding layer connected in sequence;
[0016] The first micro-encoding layer and the third micro-decoding layer are connected by a first image block partitioning unit, the second micro-encoding layer and the second micro-decoding layer are connected by a second image block partitioning unit, and the third micro-encoding layer and the first micro-decoding layer are connected by a third image block partitioning unit;
[0017] The structures of the first micro-encoding layer, the second micro-encoding layer, and the third micro-encoding layer are the same, and each includes: a convolutional layer and an adaptive image block;
[0018] The structures of the first micro-decoding layer, the second micro-decoding layer, and the third micro-decoding layer are the same, and each includes: a convolutional layer and an attention block.
[0019] In one embodiment, the adaptive image block includes a frequency domain feature extraction branch, a spatial domain feature extraction branch, and an adaptive residual branch;
[0020] The frequency domain feature extraction branch is used to perform image block average pooling and convolutional operations on the input to obtain a first attention weight; and perform a convolutional operation on the input to obtain a first convolutional result; the first attention weight and the first convolutional result are multiplied term by term to obtain a frequency domain feature;
[0021] The spatial domain feature extraction branch is used to perform image block average pooling and convolutional operations on the input to obtain a second attention weight; and perform an FFT operation on the input to obtain an FFT result; the second attention weight and the FFT result are multiplied term by term, and an iFFT operation is performed to obtain a spatial domain feature;
[0022] The spatial domain feature and the frequency domain feature are added term by term to obtain an addition result;
[0023] The adaptive residual branch includes an MLP. After the input passes through the MLP and sigmoid transformation, it is multiplied by the addition result to obtain a multiplication result; the multiplication result is added to the input to obtain the output of the adaptive image block.
[0024] In one embodiment, the density map generation module includes a CNN for generating a density map based on the feature map.
[0025] In one embodiment, the fusion module includes a lightweight convolutional layer. The lightweight convolutional layer performs lightweight convolution on the density map to obtain a convolutional result; the convolutional result is fused with the feature map to obtain a predicted clear image.
[0026] In one embodiment, the loss function used during training includes: an automatic haze generator loss, an adaptive density-aware dehazing network loss, and a joint loss;
[0027] The automatic haze generator loss is:
[0028] Loss=L(pred hazy ,xhazy ) + λ 1 L(pred clear , x clear )) + λ 2 L adv (pred hazy , x hazy )
[0029] Among them, Loss is the loss of the automatic haze generator, L is the L 1 loss, pred hazy and pred clear are the density map and the clear image predicted by the adaptive density-aware haze removal network respectively, x hazy is the blurred image sample, x clear is the clear image sample, λ 1 and λ 2 are both weights, and L adv is the adversarial loss;
[0030] The loss of the adaptive density-aware haze removal network is:
[0031]
[0032] Among them, L total is the loss of the adaptive density-aware haze removal network, L con is the contrast loss, A is the clear image predicted by the adaptive density-aware haze removal network, P is the clear image sample, is the blurred synthetic image generated by the automatic haze generator, ω is the weight, L is the L 1 loss, D 预测 is the predicted density map, and D 真实 is the true density map;
[0033] The combined loss is:
[0034]
[0035] Among them, L joint is the combined loss, ω 1 , ω 2 , ω 3 are all weights, L sl1 is the smooth L 1 loss, L MS-SSIM is the multi-scale structural similarity loss, and L con is the contrast loss.
[0036] Second, a haze removal device based on adaptive learning of image block importance is provided, including:
[0037] A network construction module, which is used to construct an adaptive image patch learning network. The adaptive image patch learning network includes an automatic haze generator and an adaptive density-aware dehazing network;
[0038] A training module, which is used to train the adaptive image patch learning network with a training data set to obtain a trained adaptive density-aware dehazing network, including:
[0039] The training data set is input into the automatic haze generator to obtain an extended training data set; the extended training data set is input into the adaptive density-aware dehazing network to obtain a predicted clear image; the samples in the training data set are image pairs, and the image pairs include clear image samples and blurred image samples. The automatic haze generator is used to generate an editable haze map guide according to the blurred image samples and combine it with the clear image samples to obtain a blurred synthetic image; the adaptive density-aware dehazing network includes a U-net network, a density map generation module and a fusion module. An image patch enhancement module is set in the U-net network. The U-net network is used to extract features from the input image to obtain a feature map. The image patch enhancement module is used to extract embedded features based on image patches. The density map generation module is used to generate a density map according to the feature map. The fusion module is used to fuse the density map and the feature map to obtain a predicted clear image;
[0040] A prediction module, which is used to input the blurred image to be processed into the trained adaptive density-aware dehazing network to obtain a clear image.
[0041] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the above-mentioned dehazing method based on adaptive learning of image patch importance.
[0042] Compared with the prior art, the present application has the following beneficial effects: The dehazing method based on adaptive learning of image patch importance in the present application expands the natural haze image data set by constructing an automatic haze generator, develops an adaptive density-aware dehazing network to process complex haze distributions in real images, and introduces a multi-negative contrast dehazing loss to use multiple negative samples to ensure robust dehazing performance on a limited real haze image data set. It achieves better performance than existing methods on various real haze image data sets and has the advantage of being robust to handle complex and diverse haze distributions. Experiments show that the method of the present application has optimal dehazing performance on various uniform and non-uniform real-world blurred images. Compared with multiple baseline models, it not only shows excellent performance in quantitative evaluation indicators, but also more accurately restores information on the visual main body and details of the image. Description of the Drawings
[0043] This application can be better understood by referring to the description given below in conjunction with the accompanying drawings, which are included in and form a part of this specification together with the following detailed description. In the drawings:
[0044] Figure 1 A schematic diagram of an adaptive image patch learning network is shown;
[0045] Figure 2 A schematic diagram of an automatic haze generator is shown;
[0046] Figure 3 A schematic diagram of a pre-trained image patch enhancement unit and a first image patch enhancement unit is shown;
[0047] Figure 4 A structural block diagram of a haze removal device based on adaptive learning of image patch importance is shown. Detailed implementation manners
[0048] Exemplary embodiments of the present application will be described below in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many specific embodiment-specific decisions may be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary with different embodiments.
[0049] Here, it should also be noted that, in order to avoid obscuring the present application with unnecessary details, only the device structures closely related to the solution according to the present application are shown in the drawings, while other details less related to the present application are omitted.
[0050] It should be understood that the present application is not limited to the described embodiments due to the following description with reference to the accompanying drawings. In this document, where feasible, embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.
[0051] An embodiment of the present application provides a haze removal method based on adaptive learning of image patch importance, which mainly includes the following steps:
[0052] Step S1, constructing an adaptive image patch learning network, where the adaptive image patch learning network includes an automatic haze generator (AHG) and an adaptive density perception haze removal network (ADD Net). Figure 1 A schematic diagram of the adaptive image patch learning network is shown.
[0053] Step S2, training the adaptive image patch learning network with a training data set to obtain a trained adaptive density perception haze removal network, including:
[0054] The training dataset is input into the automatic haze generator to obtain an extended training dataset; the extended training dataset is input into the adaptive density-aware dehazing network to obtain the predicted clear image; the samples in the training dataset are image pairs, and the image pairs include clear image samples and blurred image samples. The automatic haze generator is used to generate an editable haze map guide based on the blurred image samples and combine it with the clear image samples to obtain a blurred synthetic image; the adaptive density-aware dehazing network includes a U-net network, a density map generation module, and a fusion module. An image patch enhancement module is set in the U-net network. The U-net network is used to extract features from the input image to obtain a feature map. The image patch enhancement module is used to extract patch-based embedded features. The density map generation module is used to generate a density map based on the feature map. The fusion module is used to fuse the density map and the feature map to obtain the predicted clear image.
[0055] Step S3, input the blurred image to be processed into the trained adaptive density-aware dehazing network to obtain a clear image.
[0056] In this embodiment, an automatic haze generator (AHG) is used to generate a blurred synthetic image with a corresponding haze distribution based on an editable haze map guide. The haze is high where the dark channel value is high and low where the dark channel value is low, greatly enriching the training dataset and allowing the network to observe a wide range of haze distributions. Therefore, the generalization ability of the network under different haze conditions is greatly enhanced. At the same time, an image patch enhancement module is set in the U-net network, and this network can effectively use a large amount of training data generated by the AHG to learn a robust mapping from complex blurred images to clear images.
[0057] In one embodiment, Figure 2 shows a schematic diagram of the automatic haze generator (AHG). See Figure 2 , the automatic haze generator includes a plurality of sequentially connected residual groups, and each residual group includes a plurality of sequentially connected residual blocks; the residual block includes a convolutional layer and an attention mechanism.
[0058] In this embodiment, the clear image sample is spliced with the editable haze map guide and input into the automatic haze generator to generate a blurred synthetic image with different haze levels. To generate the haze, first, the haze concentration guide is derived using the dark channel prior, and then the haze guide is scaled by a factor k or a random guide is created by selecting and scaling specific regions to generate an image with different haze levels in different regions.
[0059] In one embodiment, see Figure 1, the U-net network includes an encoder and a decoder. The encoder includes a first encoding layer (E1), a second encoding layer (E2), and a third encoding layer (E3) connected in sequence. The decoder includes a first decoding layer (D1), a second decoding layer (D2), and a third decoding layer (D3) connected in sequence. Here, the first encoding layer, the second encoding layer, the third encoding layer, the first decoding layer, the second decoding layer, and the third decoding layer have the same structure, all including a 3×3 convolutional layer and a LeakyReLU non-linear activation function layer for feature extraction.
[0060] The image patch enhancement module includes a pre-trained image patch enhancement unit (PPEM), a first image patch enhancement unit PEM, a second image patch enhancement unit PEM, and a third image patch enhancement unit PEM;
[0061] The first encoding layer and the third decoding layer are connected through the pre-trained image patch enhancement unit and the first image patch enhancement unit; the second encoding layer and the second decoding layer are connected through the second image patch enhancement unit; the third encoding layer and the first decoding layer are connected through the third image patch enhancement unit. The embedding information in the encoder is transmitted to the decoder via skip connections to facilitate image reconstruction.
[0062] Specifically, the pre-trained image patch enhancement unit and the first image patch enhancement unit have the same structure, Figure 3 The schematic diagrams of the pre-trained image patch enhancement unit and the first image patch enhancement unit are shown, see Figure 3 , including:
[0063] A micro encoder, a micro decoder, and a dilated convolutional layer (Dilated CONV). The micro encoder includes a first micro encoding layer, a second micro encoding layer, and a third micro encoding layer connected in sequence. The micro decoder includes a first micro decoding layer, a second micro decoding layer, and a third micro decoding layer connected in sequence. The input of the image patch enhancement unit is first divided into a series of image patches and then input into the micro encoder.
[0064] The first micro encoding layer and the third micro decoding layer are connected through a first image patch division unit (PR), the second micro encoding layer and the second micro decoding layer are connected through a second image patch division unit (PR), and the third micro encoding layer and the first micro decoding layer are connected through a third image patch division unit (PR);
[0065] The first micro encoding layer, the second micro encoding layer, and the third micro encoding layer have the same structure, all including: a 3×3 convolutional layer (CONV) and an adaptive image patch (APR) to extract specific information of each image patch;
[0066] The structures of the first micro-decoding layer, the second micro-decoding layer, and the third micro-decoding layer are the same, and each includes: a 3×3 convolutional layer (CONV) and an attention block (Att Block).
[0067] In this embodiment, after extracting features in each micro-encoding layer of the micro-encoder, an image patch partitioning unit (PR) is used at the bottleneck and during skip connections to restore the image patches to the downsampled feature map, and then it is sent to the decoder. The micro-decoding layer consists of a 3×3 convolution and an attention block, which allows the convolution and attention modules therein to operate on the entire image and maintain the overall consistency of the image. After the decoder, a dilated convolutional layer is added to expand the receptive field and restore the overall consistency in the reconstructed image, and the design of the image patch size varies according to the size of the feature map in the skip connection path of this module.
[0068] It should be noted that since using APR in deeper networks does not produce significant improvements, APR is only applied in the PPEM and the first image patch enhancement unit. APR is not used in the second image patch enhancement unit and the third image patch enhancement unit, and the remaining structures are the same.
[0069] Specifically, referring to Figure 3 , the adaptive image patch (APR) includes a frequency-domain feature extraction branch, a spatial-domain feature extraction branch, and an adaptive residual branch;
[0070] The frequency-domain feature extraction branch is used to perform image patch average pooling and convolution operations on the input to obtain the first attention weight; and perform convolution operations on the input to obtain the first convolution result; the first attention weight is multiplied term by term with the first convolution result to obtain the frequency-domain feature;
[0071] The spatial-domain feature extraction branch is used to perform image patch average pooling and convolution operations on the input to obtain the second attention weight; and perform FFT operations on the input to obtain the FFT result; the second attention weight is multiplied term by term with the FFT result and then perform iFFT operations to obtain the spatial-domain feature;
[0072] The spatial-domain feature and the frequency-domain feature are added term by term to obtain the addition result;
[0073] The adaptive residual branch includes an MLP. After the input passes through the MLP and sigmoid transformation, it is multiplied with the addition result to obtain the multiplication result; the multiplication result is added to the input to obtain the output of the adaptive image patch.
[0074] In this embodiment, the APR uses two branches to extract features in the spatial domain and the frequency domain, and then applies an adaptive residual connection. This operation allows the PEM to customize the corresponding feature extraction for each image patch. Residual connections are commonly used in defogging tasks in deep networks, but different image patches may require different levels of processing: severely blurred image patches require more extensive feature refinement, while image patches with little or no blur may only require basic connections. Therefore, after summing the results of processing the two branches, an adaptive residual connection is applied.
[0075] In one embodiment, the density map generation module includes a Convolutional Neural Network (CNN) for generating a density map based on the feature map.
[0076] Specifically, the fusion module includes a lightweight convolutional layer that performs lightweight convolution on the density map to obtain a convolution result; the convolution result is fused with the feature map to obtain a predicted clear image.
[0077] In the above embodiment, in order to enable the Adaptive Density-aware Dehazing Network (ADD Net) to obtain information about the haze distribution, the predicted density map is merged into the network. After the U-Net extracts the feature map, a haze perception training framework is implemented. The features extracted by the U-Net are processed through two branches composed of deep convolutions with large kernels. The first branch predicts the density map through a convolutional neural network, and the process is constrained by a specific loss function. Then, the density map is embedded into the feature space, passed through a lightweight convolutional layer and fused with the features from the other branch. This fused information provides the haze distribution to assist the second branch in predicting the clear image.
[0078] In one embodiment, the loss function used during training includes: Auto Haze Generator Loss, Adaptive Density-aware Dehazing Network Loss, and Joint Loss;
[0079] The Auto Haze Generator Loss is:
[0080] Loss = L(pred hazy , x hazy ) + λ 1 L(pred clear , x clear ) + λ 2 L adv (pred hazy , x hazy )
[0081] where Loss is the Auto Haze Generator Loss, L is the L 1 loss, pred hazy and pred clearThe density map and clear image predicted by the adaptive density-aware dehazing network, respectively, x hazy is the blurred image sample, x clear is the clear image sample, λ 1 and λ 2 are both weights, set to 0.5 and 0.0001 respectively, to balance the contributions of different loss components, L adv is the adversarial loss;
[0082] Here, calculate the loss of the automatic haze generator, using L 1 The loss and the adversarial loss compare the output of the network with the original blurred image, and at the same time calculate the L 1 loss between the network output and the original clear image, which helps the network learn to generate haze images with different haze concentrations according to different guidance values, rather than just copying the haze images in the dataset.
[0083] The loss of the adaptive density-aware dehazing network is:
[0084]
[0085] where, L total is the loss of the adaptive density-aware dehazing network, L con is the contrast loss, A is the clear image predicted by the adaptive density-aware dehazing network, P is the clear image sample, is the blurred synthetic image generated by the automatic haze generator, ω is the weight, L is the L 1 loss, D 预测 is the predicted density map, D 真实 is the true density map; the true density map D 真实 = I dark (x clean ) - I dark (x haze ), I dark is the dark channel value.
[0086] The combined loss is:
[0087]
[0088] where, L joint is the combined loss, ω 1 , ω 2 , ω 3 are all weights, L sl1 is the smooth L 1 loss, L MS-SSIM is the multi-scale structural similarity loss, L con is the contrast loss, that is, the multi-negative contrast dehazing loss (MNCD Loss). ω 1 , ω2 and ω 3 are both weights, which are set to 1, 0.3, and 0.005 respectively.
[0089] Specifically, this application adopts a dynamic data selection strategy to enhance the training of the dehazing network. Initially, the probability of selecting simulated data or real data from the dataset is equal (0.5). As the training progresses, the probability of selecting simulated data gradually decreases, which allows the network to first learn from different ranges of haze concentrations and distributions; in the later stage, the training focuses more on real data to reduce the gap between the simulated scenario and the real scenario. Although the models trained using this strategy initially converge slower due to the slight differences between simulated data and real data, they ultimately achieve better performance than the models without this enhanced data training.
[0090] It is found during the actual training process that adding the contrast loss at the beginning of training will slow down the convergence speed of the network. This is because the network cannot effectively reconstruct the image in the early stage, and adding the contrast loss will cause the network's prediction to blindly deviate from the negative samples, rather than necessarily moving towards the positive samples, resulting in chaotic predictions. Therefore, during the training process, the contrast loss is introduced after 50 training epochs and applied once the network can roughly reconstruct a clear image, which further enhances the performance of the network.
[0091] To further verify the effectiveness of the method of this application, the following experiments were conducted:
[0092] Four real-world haze image datasets with very limited data were selected to evaluate and test the proposed method, namely Dense-Haze, NH-Haze, O-Haze, and I-Haze. The three methods compared with the method of this application are DCP, FFA, and Fourmer respectively.
[0093] The peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are used to quantitatively evaluate the restored images. PSNR is a commonly used metric for measuring image quality, especially when evaluating image reconstruction or image compression algorithms. It is calculated by comparing the original image and the distorted image, and measures the ratio of the maximum possible power in the image to the noise power that affects the image quality; SSIM is a metric for measuring the visual similarity of two images. It takes into account not only the differences in brightness and contrast but also the structural information of the images. The value range of SSIM is between 0 and 1, where 1 indicates that the two images are exactly the same and 0 indicates that they are completely different.
[0094] As shown in Table 1, compared with the other three models, the model of the present application achieved the best PSNR and SSIM scores on all datasets. On the O-Haze dataset, its average performance was 1.54 dB (PSNR) and 0.044 (SSIM) higher. On the I-Haze dataset, it exceeded the best model by 1.84 dB (PSNR) and 0.09204 (SSIM). On the NH-Haze dataset, it was 0.39 dB (PSNR) and 0.01887 (SSIM) higher than the existing best model.
[0095] Table 1 Comparison of Image Dehazing on Real-World Datasets
[0096]
[0097] Adopting the same inventive concept as the dehazing method based on adaptive learning of image patch importance, this embodiment also provides a corresponding dehazing device based on adaptive learning of image patch importance. Figure 4 The structural block diagram of the dehazing device based on adaptive learning of image patch importance is shown, including:
[0098] A network construction module 41, configured to construct an adaptive image patch learning network, and the adaptive image patch learning network includes an automatic haze generator and an adaptive density-aware dehazing network;
[0099] A training module 42, configured to train the adaptive image patch learning network with a training dataset to obtain a trained adaptive density-aware dehazing network, including:
[0100] The training dataset is input into the automatic haze generator to obtain an extended training dataset; the extended training dataset is input into the adaptive density-aware dehazing network to obtain a predicted clear image; the samples in the training dataset are image pairs, and the image pairs include clear image samples and blurred image samples. The automatic haze generator is used to generate an editable haze map guide according to the blurred image samples, and combine it with the clear image samples to obtain a blurred composite image; the adaptive density-aware dehazing network includes a U-net network, a density map generation module and a fusion module. An image patch enhancement module is set in the U-net network. The U-net network is used to extract features from the input image to obtain a feature map. The image patch enhancement module is used to extract patch-based embedding features. The density map generation module is used to generate a density map according to the feature map. The fusion module is used to fuse the density map and the feature map to obtain a predicted clear image;
[0101] A prediction module 43, configured to input the blurred image to be processed into the trained adaptive density-aware dehazing network to obtain a clear image.
[0102] The haze removal device based on adaptive learning of image block importance in this embodiment has the same inventive concept as the above-mentioned haze removal method based on adaptive learning of image block importance. Therefore, the specific implementation of this device can be seen in the embodiment part of the haze removal method based on adaptive learning of image block importance in the previous text, and its technical effect corresponds to that of the above method, which will not be elaborated here.
[0103] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the above-mentioned haze removal method based on adaptive learning of image block importance.
[0104] As mentioned above, the above are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A defogging method based on adaptive learning of image block importance, characterized in that: include: Constructing an adaptive image block learning network, wherein the adaptive image block learning network includes an automatic haze generator and an adaptive density-aware dehazing network; The adaptive image block learning network is trained using a training data set to obtain a trained adaptive density-aware defogging network, including: The training data set is input into the automatic haze generator to obtain an extended training data set; the extended training data set is input into the adaptive density-aware dehazing network to obtain a predicted clear image; the samples in the training data set are image pairs, the image pairs include clear image samples and blurred image samples, the automatic haze generator is used to generate an editable haze map guide according to the blurred image samples, and obtain a blurred composite image in combination with the clear image samples; the adaptive density-aware dehazing network includes a U-net network, a density map generation module and a fusion module, the U-net network is provided with an image block enhancement module, the U-net network is used to extract features of the input image to obtain a feature map, the image block enhancement module is used to extract embedded features based on image blocks, the density map generation module is used to generate a density map according to the feature map, and the fusion module is used to fuse the density map and the feature map to obtain a predicted clear image; The blurred image to be processed is input into the trained adaptive density-aware dehazing network to obtain a clear image.
2. The method according to claim 1, characterized in that The automatic haze generator includes a plurality of sequentially connected residual groups, each of which includes a plurality of sequentially connected residual blocks; the residual block includes a convolutional layer and an attention mechanism.
3. The method according to claim 1, characterized in that The U-net network includes an encoder and a decoder, wherein the encoder includes a first encoding layer, a second encoding layer, and a third encoding layer connected in sequence, and the decoder includes a first decoding layer, a second decoding layer, and a third decoding layer connected in sequence; The image block enhancement module includes a pre-trained image block enhancement unit, a first image block enhancement unit, a second image block enhancement unit and a third image block enhancement unit; The first encoding layer and the third decoding layer are connected via the pre-trained image block enhancement unit and the first image block enhancement unit; the second encoding layer and the second decoding layer are connected via the second image block enhancement unit; and the third encoding layer and the first decoding layer are connected via the third image block enhancement unit.
4. The method according to claim 3, characterized in that The pre-trained image block enhancement unit and the first image block enhancement unit have the same structure, including: A micro encoder, a micro decoder and an expanded convolutional layer, wherein the micro encoder comprises a first micro encoding layer, a second micro encoding layer and a third micro encoding layer connected in sequence, and the micro decoder comprises a first micro decoding layer, a second micro decoding layer and a third micro decoding layer connected in sequence; The first micro coding layer and the third micro decoding layer are connected via a first image block division unit, the second micro coding layer and the second micro decoding layer are connected via a second image block division unit, and the third micro coding layer and the first micro decoding layer are connected via a third image block division unit; The first micro-coding layer, the second micro-coding layer, and the third micro-coding layer have the same structure, and all include: a convolutional layer and an adaptive image block; The first micro-decoding layer, the second micro-decoding layer, and the third micro-decoding layer have the same structure, and all include: a convolutional layer and an attention block.
5. The method according to claim 4, characterized in that The adaptive image block includes a frequency domain feature extraction branch, a spatial domain feature extraction branch, and an adaptive residual branch; The frequency domain feature extraction branch is used to perform image block average pooling and convolution operations on the input to obtain a first attention weight; And performing a convolution operation on the input to obtain a first convolution result; multiplying the first attention weight and the first convolution result item by item to obtain a frequency domain feature; The spatial domain feature extraction branch is used to perform image block average pooling and convolution operations on the input to obtain a second attention weight; And perform an FFT operation on the input to obtain an FFT result; multiply the second attention weight by the FFT result item by item, and perform an iFFT operation to obtain a spatial domain feature; The spatial domain features and the frequency domain features are added item by item to obtain an addition result; The adaptive residual branch includes an MLP, and the input is multiplied with the addition result after being transformed by the MLP and sigmoid to obtain a multiplication result; the multiplication result is added to the input to obtain an output of an adaptive image block.
6. The method according to claim 1, characterized in that The density map generation module includes a CNN, which is used to generate a density map according to the feature map.
7. The method according to claim 1, characterized in that The fusion module includes a lightweight convolution layer, which performs lightweight convolution on the density map to obtain a convolution result; the convolution result is fused with the feature map to obtain a predicted clear image.
8. The method according to claim 1, characterized in that The loss functions used in the training process include: automatic haze generator loss, adaptive density-aware dehazing network loss, and joint loss; The automatic haze generator loss is: Loss=L(before hazy ,x hazy )+λ1L(before clear ,x clear ))+λ2L adv (before hazy ,x hazy ) Among them, Loss is the automatic haze generator loss, L is the L1 loss, and pred hazy and pred clear They are the density map and clear map predicted by the adaptive density-aware dehazing network, x hazy is the blurred image sample, x clear is a clear image sample, λ1 and λ2 are weights, L adv For adversarial losses; The adaptive density-aware dehazing network loss is: Among them, L total is the adaptive density-aware dehazing network loss, L con is the contrast loss, A is the clear image predicted by the adaptive density-aware dehazing network, P is the clear image sample, is the blurred synthetic image generated by the automatic haze generator, ω is the weight, L is the L1 loss, D 预测 is the predicted density map, D 真实 is the real density map; The combined loss is: Among them, L joint is the joint loss, ω1, ω2, ω3 are weights, L sl1 is the smooth L1 loss, L MS-SSiM is the multi-scale structural similarity loss, L con is the comparison loss.
9. A defogging device based on adaptive learning of image block importance, characterized in that: include: A network construction module, used to construct an adaptive image block learning network, wherein the adaptive image block learning network includes an automatic haze generator and an adaptive density-aware dehazing network; A training module, used to train the adaptive image block learning network using a training data set to obtain a trained adaptive density-aware defogging network, comprising: The training data set is input into the automatic haze generator to obtain an extended training data set; the extended training data set is input into the adaptive density-aware dehazing network to obtain a predicted clear image; the samples in the training data set are image pairs, the image pairs include clear image samples and blurred image samples, the automatic haze generator is used to generate an editable haze map guide according to the blurred image samples, and obtain a blurred composite image in combination with the clear image samples; the adaptive density-aware dehazing network includes a U-net network, a density map generation module and a fusion module, the U-net network is provided with an image block enhancement module, the U-net network is used to extract features of the input image to obtain a feature map, the image block enhancement module is used to extract embedded features based on image blocks, the density map generation module is used to generate a density map according to the feature map, and the fusion module is used to fuse the density map and the feature map to obtain a predicted clear image; The prediction module is used to input the blurred image to be processed into the trained adaptive density-aware defogging network to obtain a clear image.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the defogging method based on image block importance adaptive learning is implemented as described in any one of claims 1 to 8.
Citation Information
Cited By
Low-light image enhancement method and system based on space-frequency domain characteristic resolution self-adjustment
CN120672637A
Low-light image enhancement method and system based on spatial-frequency domain feature resolution self-adjustment
CN120672637B
Target-aware triple-head network device for joint haze removal and target protection and image haze removal method
CN122865917A