Picture defogging algorithm based on fog concentration information guidance
By introducing a fog concentration information estimation module and a multi-scale feature alignment module, the problem of poor fog removal effect in the prior art is solved, and precise fog removal of fog areas with different concentrations is achieved and the effective recovery of image details is achieved.
Patent Information
- Application Number
- CN202510212015.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The recovery effect of the prior art when processing thick fog pictures is far less than that of processing mist pictures. The traditional physical fog removal algorithm is unstable and the restored image texture details are not clear enough. The deep learning-based methods cannot make full use of the fog concentration information, resulting in serious fog artifacts in the recovered images.
A fog concentration information-guided picture defog removal method is proposed. Through the fog concentration estimation module and the multi-scale feature alignment module, haze of different thicknesses in different areas of the picture is adaptively removed, so as to achieve effective defog removal and detail recovery of synthetic mist and natural thick fog pictures.
It significantly improves the image fog removal effect, can effectively remove dense fog areas, restore clearer image details, and can handle non-uniform foggy pictures in natural environments to avoid fog artifacts.
Smart Images

Figure CN120147181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to the field of image dehazing under natural conditions. The present invention mainly solves the problem of image quality degradation in a natural thick fog environment, and can effectively restore the details and textures of the picture and improve the image quality. Background Art
[0002] As a common natural phenomenon, fog has brought great adverse effects to people's outdoor activities such as transportation and high-altitude operations. Images taken in a foggy environment usually have problems such as color distortion, contrast reduction, and visibility reduction, which lead to serious degradation of the image and bring many challenges to advanced vision tasks in the field of computer vision (such as object detection, classification, and image segmentation). Therefore, single-image dehazing has become an important research direction in the field of computer vision, aiming to restore a clear image from a blurred image.
[0003] The current image dehazing methods are mainly divided into two categories: physical model dehazing and deep learning dehazing. The early image dehazing methods mainly relied on the physical dehazing method based on the atmospheric scattering model (ASM), which was achieved by establishing a mapping relationship between the foggy image and its clear corresponding image. The mathematical expression of the atmospheric scattering model is:
[0004] t(x) = e -βd(x)
[0005] I(x) = J(x)t(x) + A(1 - t(x))
[0006] Where, I(x) represents the foggy image, J(x) represents the clear image, t and A respectively represent the air transmittance and the global atmospheric light, and d(x) represents the imaging depth. However, the dehazing method based on the atmospheric scattering model can only handle simple light fog removal tasks, and has disadvantages such as instability, too dark restored picture color, and overexposure in the sky area.
[0007] With the rapid development of deep learning technology in various fields, researchers have begun to introduce deep learning technology into the field of image dehazing. Deep learning-based methods usually reconstruct clear images by estimating physical parameters or directly generate clear images in an end-to-end manner. For example, FFANet proposed feature attention modules in the channel and pixel dimensions to restore clear images, and DANet introduced horizontal and vertical attention mechanisms to effectively retain the details and semantic content of images. Although these methods perform well in dealing with hazy images in the synthetic domain, their dehazing effects in real environments are not satisfactory because they fail to fully explore the distribution of fog density information in images. To solve this problem, some research methods in recent years have tried to incorporate fog density information into their method systems. For example, ACERNet incorporated hazy images as negative samples into the loss function through contrastive learning, thus significantly improving the performance of the dehazing model; C2PNet introduced fog distribution constraints to gradually guide the dehazing model to continuously learn, and finally restored fog-free images with better effects. However, these methods are still insufficient in utilizing fog distribution information, resulting in problems such as heavy fog distribution artifacts and blurred edge details in the restored natural hazy images. In addition, these methods usually also have the disadvantage of a large number of model parameters.
[0008] In summary, the existing dehazing algorithms have much worse restoration effects when dealing with thick fog images than when dealing with thin fog images. Specifically, traditional physical dehazing algorithms have a single application scenario, unstable dehazing effects, and the texture details of the restored images are not clear enough; while deep learning-based dehazing algorithms perform excellently on synthetic datasets, but due to the inability to fully utilize fog concentration information, the restored images have serious fog artifacts. Therefore, it is of great significance to propose a dehazing method that can simultaneously consider fog concentration information and multi-scale feature fusion to achieve precise dehazing of different fog concentration regions and image detail restoration. Summary of the Invention
[0009] To solve the deficiencies of the existing technology, the purpose of the present invention is to provide a fog concentration information-guided image dehazing method to effectively remove the fog in non-uniform dehazed images, and at the same time use a multi-scale feature alignment module to effectively extract the texture details of the images, solving the problems of incomplete dehazing effect and color texture distortion existing in the existing technology.
[0010] To achieve the above purpose, the present invention provides the following technical solutions: A fog concentration information-guided image dehazing algorithm, including the following steps:
[0011] S1. Obtain the training dataset of the image dehazing network, and collect non-uniform hazy images and corresponding clear images in real natural scenes, thin fog images and their corresponding clear images under synthetic conditions as the network training dataset.
[0012] S2. Propose a fog concentration estimation module to obtain the global fog concentration estimation map of the foggy image using the global attention mechanism. Specifically, the extracted fog concentration estimation map is aligned with the depth information of the foggy image, and the module uses the difference between the product of the depth information of the foggy image, the fog concentration mask, and the original foggy image as the loss value to adjust the parameters of the network through backpropagation.
[0013] S3. Propose a new multi-scale feature alignment module to effectively restore the texture details of the image. Construct a shallow feature extraction network embedded with the multi-scale feature alignment module to effectively restore the texture details of the image.
[0014] S4. Construct a fog concentration information-guided thick fog removal network, which is an encoder-decoder network embedded with a fog concentration estimation module and dynamic convolution, aiming to make full use of the fog concentration information and adaptively restore the thick fog area in the image to further improve the defogging effect.
[0015] Preferably, step S2 specifically includes:
[0016] Based on the dataset collected in S1, first use the fog concentration estimation module to estimate the fog concentration information of the input image of the model. Specifically, the proposed fog concentration estimation module is mainly implemented through the multi-head self-attention mechanism, which includes an initial convolution layer, a multi-head self-attention layer, and a sigmoid activation function. First, the input foggy image is processed by the fog concentration estimation module. The specific steps are as follows: The initial convolution layer uses a 3x3 convolution kernel and several filters to perform preliminary feature extraction on the input image and introduce non-linearity through the ReLU activation function to generate the initial feature map. Then, these feature maps are processed by the multi-head self-attention layer, and the multi-head self-attention layer captures the global and local features in the input image through multiple parallel attention heads to generate a feature representation with rich context information. Subsequently, the output of the multi-head self-attention layer passes through a 1x1 convolution kernel to convert the feature map into a single-channel fog concentration mask map. Finally, the feature map is mapped to the [0,1] interval through the sigmoid activation function to generate the feature attention weight map.
[0017] Preferably, step S3 specifically includes:
[0018] Based on the dataset collected in S1, apply a shallow feature extraction network to the input image data of the model to extract the shallow texture features of the image. Specifically, the shallow feature extraction network mainly consists of the proposed multi-scale feature alignment module. Among them, the network includes three convolutional blocks: a pre-conversion module, a post-conversion module, and a multi-scale feature alignment module. First, the input image passes through the pre-convolution module. The first layer of convolution uses 32 3x3 convolutional kernels to convolve the input three-channel (RGB) image to extract initial features. After convolution, non-linear processing is performed through the ReLU activation function, and instance normalization is carried out. Instance normalization is performed on the height and width dimensions of the image, so that the normalization information comes from itself, integrating and adjusting the global information. Then, the feature map is input into the second layer of convolution. The second layer of convolution uses 64 3x3 convolutional kernels to convolve the feature map to further extract features. After convolution, non-linear processing is performed through the ReLU activation function, and instance normalization is carried out again. The processed feature map is input into the multi-scale feature fusion module. The multi-scale feature fusion module uses three different scales of convolutional kernels: 3x3, 5x5, and 7x7 to extract low, medium, and high-scale features respectively. The extracted feature maps are upsampled to the same resolution through bilinear interpolation. Then, the hybrid coefficient extraction module is used to perform weighted fusion on the three features, and the key features are enhanced through the channel attention mechanism. Finally, the fused features are added to the original input image to obtain an enhanced feature representation. Finally, the image enters the post-convolution module. The third layer of convolution uses 64 3x3 convolutional kernels to further process the fused feature map. The fourth layer of convolution uses 32 3x3 convolutional kernels to reduce the number of feature channels to 32. Finally, the fifth layer of convolution uses 3 3x3 convolutional kernels to restore the feature map to a three-channel image. In addition, non-linear transformation is performed through the ReLU activation function after each layer of convolution.
[0019] Preferably, step S4 specifically includes:
[0020] Based on the image data input by S1, the fog concentration information estimation map obtained by S2, and the shallow feature map obtained by S3, S4 constructed a thick fog removal network guided by fog concentration information. Specifically, the thick fog removal network is built based on an encoder-decoder architecture, specifically including an encoder part that enhances feature expressiveness through multiple dynamic convolutional layers and a channel attention mechanism. This part first uses two consecutive dynamic convolutional layers, extracts features from the input 7-channel features using a 3x3 convolutional kernel, and combines the ReLU activation function to generate 128-channel high-dimensional features. Subsequently, the feature information is further strengthened through the channel attention mechanism; in the decoder part, first use a dynamic convolutional layer to convert the above features from 128 channels to 64 channels, and then gradually convert the features to 32 channels and the output 3 channels through two conventional 3x3 convolutional layers. At the same time, each layer of convolution is followed by a ReLU activation function for non-linear transformation. Finally, the Sigmoid activation function is used to ensure that the pixel values of the output image fall within the range of [0,1], thereby effectively removing and restoring the thick fog area.
[0021] The advantages of the present invention are:
[0022] Compared with the prior art, the present invention first proposes a fog concentration information estimation module, introduces a fog concentration information map into the backbone fog removal network, and can adaptively remove haze of different thicknesses in different regions of the picture, significantly improving the picture dehazing effect;
[0023] Compared with the prior art, the present invention introduces a shallow feature extraction network embedded with a multi-scale feature alignment module, which can effectively extract the shallow features of the picture and clearly restore the texture details of the picture.
[0024] The present invention can not only effectively remove fog from synthetic light fog pictures, but also avoid cross-domain difficulties and remove fog from non-uniform foggy pictures in natural environments.
[0025] The present invention can make full use of the fog concentration information of the picture for adaptive dehazing, effectively remove the thick fog area in the picture, improve the picture dehazing effect, and make the processed picture more in line with the visual perception of the human eye. Brief Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0027] Figure 1 For the abstract drawing.
[0028] Figure 2Schematic diagram of the structure of a fog - concentration - information - guided image defogging model disclosed in an embodiment of the present application.
[0029] Figure 3 Schematic diagram of the structure of a multi - scale feature alignment module proposed in an embodiment of the present application.
[0030] Figure 4 Schematic diagram of the structure of a shallow - feature extraction network disclosed in an embodiment of the present application.
[0031] Figure 5 Schematic diagram of the structure of a fog - concentration information estimation module proposed in an embodiment of the present application.
[0032] Figure 6 Schematic diagram of the structure of a fog - concentration - information - guided thick - fog removal network disclosed in an embodiment of the present application.
[0033] Figure 7 Schematic diagram of the working process of a fog - concentration - information - guided image defogging model disclosed in an embodiment of the present application.
[0034] Figure 8 Effect diagram of test data of a fog - concentration - information - guided image defogging model disclosed in an embodiment of the present application.
[0035] Figure 9 Effect diagram of natural - scene defogging of a fog - concentration - information - guided image defogging model disclosed in an embodiment of the present application.
[0036] Figure 10 Comparison diagram of defogging effects between a fog - concentration - information - guided image defogging model and other traditional models disclosed in an embodiment of the present application Detailed implementation manner
[0037] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the following describes the specific implementation method of the present invention in detail with reference to the accompanying drawings of the specification.
[0038] The following description will fully and detailedly explain the embodiments with specific details, but the specific embodiments described herein are only used to explain the present invention, and the present invention can also be implemented in other ways different from this description. Therefore, the present invention is not limited by the disclosed specific embodiments.
[0039] Embodiment 1:
[0040] In this embodiment, the specific network structure of the image defogging model is provided, as Figure 2 shown:
[0041] The image dehazing method proposed by the present invention includes two parts: a shallow feature extraction network and a deep dehazing network. The shallow feature extraction network is composed of a general convolutional network and a multi-scale feature alignment module, as Figure 4 shown. The deep dehazing network consists of two parts, namely a fog concentration information estimation module and a basic encoder-decoder network composed of dynamic convolutional layers, as Figure 6 shown.
[0042] Specifically, the shallow feature extraction network is mainly composed of the proposed multi-scale feature alignment module. This network includes three convolutional blocks: a pre-conversion module, a post-conversion module, and a multi-scale feature alignment module. First, the input image passes through the pre-convolution module. The first layer of convolution uses 32 3x3 convolutional kernels to convolve the input three-channel (RGB) image and extract preliminary features. After convolution, non-linear processing is performed through the ReLU activation function, and instance normalization is carried out, such that the normalization information comes from the height and width dimensions of the image. Then, the feature map enters the second layer of convolution, which uses 64 3x3 convolutional kernels to further extract features, performs non-linear processing through the ReLU activation function, and performs instance normalization again. The processed feature map is input into the multi-scale feature fusion module, which uses three different scales of convolutional kernels (3x3, 5x5, 7x7) to extract low, medium, and high-scale features respectively. The extracted feature maps are upsampled to the same resolution through bilinear interpolation, and then a hybrid coefficient extraction module is used to weight-fuse the three features, and the key features are enhanced through the channel attention mechanism. Finally, the fused features are added to the original input image to obtain an enhanced feature representation. Finally, the image enters the post-convolution module. The third layer of convolution uses 64 3x3 convolutional kernels to further process the fused feature map. The fourth layer of convolution uses 32 3x3 convolutional kernels to reduce the number of feature channels to 32. Finally, the fifth layer of convolution uses 3 3x3 convolutional kernels to restore the feature map to a three-channel image. Non-linear transformation is performed through the ReLU activation function after each layer of convolution.
[0043] The deep dehazing network includes a fog concentration information estimation module and a basic encoder-decoder structure. Among them, the fog concentration information estimation module, as Figure 5As shown, it is implemented by means of a multi-head self-attention mechanism, including an initial convolutional layer, a multi-head self-attention layer, and a Sigmoid activation function. First, the input foggy image is processed by the fog concentration estimation module. The specific steps are as follows: The initial convolutional layer uses a 3x3 convolutional kernel and several filters to perform preliminary feature extraction on the input image, and generates an initial feature map through the ReLU activation function. Then, these feature maps are processed by the multi-head self-attention layer, and the global and local features of the input image are captured through multiple parallel attention heads to generate a feature representation with rich context information. Next, the output of the multi-head self-attention layer uses a 1x1 convolutional kernel to convert the feature map into a single-channel fog concentration mask map. Finally, the feature map is mapped to the [0,1] interval through the Sigmoid activation function to generate a feature attention weight map.
[0044] The encoder part enhances the feature expressiveness through multiple dynamic convolutional layers and a channel attention mechanism. Specifically, first, the first dynamic convolutional layer is adopted, with a convolutional kernel size of 3x3, an input channel of 7, including a fog concentration weight map, a shallow feature extraction map, and the original fog map, an output channel of 64, and 4 convolutional kernels. It performs feature extraction on the 7-channel input features and generates a 64-channel feature map in combination with the ReLU activation function. Then, the second dynamic convolutional layer is adopted, with a convolutional kernel size of 3x3, an input channel of 64, an output channel of 128, and 4 convolutional kernels. It generates high-dimensional features of 128 channels through further feature extraction and combines the ReLU activation function. Subsequently, the feature information of the 128 channels is further strengthened through the channel attention mechanism to enhance the feature expressiveness.
[0045] The decoder part gradually restores the image through dynamic convolutional layers and conventional convolutional layers, and ensures the quality of the output image. Specifically, first, a dynamic convolutional layer is used to convert the 128-channel features into 64 channels, with a convolutional kernel size of 3x3 and 4 convolutional kernels, and a non-linear transformation is performed in combination with the ReLU activation function. Then, the second conventional convolutional layer further converts the 64-channel features into 32 channels, with a convolutional kernel size of 3x3, and the features are processed through the ReLU activation function. Finally, the third conventional convolutional layer converts the 32-channel features into the final 3-channel output image, with a convolutional kernel size of 3x3, and the pixel values are mapped to the [0,1] interval using the Sigmoid activation function. Through the above design, the encoder uses multiple dynamic convolutional layers and a channel attention mechanism to enhance the feature expressiveness, and the decoder gradually restores the image through dynamic convolutional layers and conventional convolutional layers, and ensures that the pixel values of the output image fall within the [0,1] range, so as to effectively remove and restore the thick fog area.
[0046] Example 2:
[0047] In this embodiment, the datasets used in the experiment are all open-source datasets on the Internet, as follows:
[0048] NTIRE2019 Dataset: The NTIRE2019 dataset contains 55 pairs of images, including heavily foggy images generated in indoor or outdoor environments and their corresponding ground truth (fog-free) images. These images are used to evaluate and train the performance of image dehazing algorithms. The NTIRE2019 dataset provides a series of images with complex fog variations, providing rich data support for the development and verification of dehazing methods.
[0049] NTIRE20 - 21 Dataset: The NTIRE20 - 21 dataset includes two sub-datasets, NTIRE2020 and NTIRE2021. The specific introduction is as follows:
[0050] 1. NTIRE2020 Dataset: It contains 45 pairs of training images. These image pairs consist of foggy images and their corresponding fog-free images.
[0051] 2. NTIRE2021 Dataset: It contains 25 pairs of training images. These image pairs consist of foggy images and their corresponding fog-free images.
[0052] RESIDE - Indoor Training Set (ITS): It contains a large number of synthetic foggy and fog-free image pairs of indoor scenes, approximately 13,000 synthetic foggy images and their corresponding fog-free images. These foggy images are obtained by synthesizing different concentrations of fog on indoor fog-free pictures.
[0053] RESIDE - Outdoor Training Set (OTS): It contains a large number of synthetic foggy and fog-free image pairs of outdoor scenes, approximately 72,000 synthetic foggy images and their corresponding fog-free images. These foggy images are obtained by synthesizing different concentrations of fog on outdoor fog-free pictures.
[0054] By including two sub-datasets, NTIRE2020 and NTIRE2021, the NTIRE20 - 21 dataset provides more diverse image scenes and fog conditions for the training and evaluation of the model.
[0055] OHaze Dataset (NTIRE2018): The OHaze dataset is part of the NTIRE2018 challenge. This dataset contains 35 pairs of images, which consist of foggy images and their corresponding fog-free images. The OHaze dataset is generated in real indoor and outdoor environments, capturing various fog variations of different concentrations, providing valuable data for training and evaluating dehazing algorithms.
[0056] IHaze Dataset: The IHaze dataset contains 25 pairs of images, which consist of hazy images and their corresponding haze-free images. The IHaze dataset pays particular attention to hazy images generated under different lighting and environmental conditions and provides high-quality ground truth images for verifying the effectiveness and robustness of dehazing algorithms.
[0057] For all non-RESIDE datasets (i.e., real-scene datasets), the pictures were translated, cropped, and rotated to obtain 1000 clear pictures and hazy pictures for training.
[0058] The present invention divides the above datasets into three parts, namely the pre-training dataset, the training dataset, and the test dataset. Each dataset contains two types of pictures: hazy pictures and clear pictures. The first part of the dataset selects the RESIDE-Indoor dataset and the RESIDE-Outdoor dataset, which together include 750 pairs of artificial hazy pictures as the pre-training data for the network. The second part of the dataset consists of real hazy picture data, which are 800 pairs of real-scene pictures and serve as the training data for the network. The third part includes 200 pairs of real-scene hazy pictures for testing the network.
[0059] Example 3:
[0060] This embodiment provides the entire training process of the picture dehazing method, as Figure 7 shown, and specifically includes the following steps:
[0061] S1. Construct the entire picture dehazing model
[0062] S2. Send the picture into the shallow feature extraction network and the fog concentration estimation module to obtain the shallow feature map f(shallowfeature) of the hazy picture and the corresponding fog concentration weight f(hazemap).
[0063] S3. Send the shallow feature map and the fog concentration map obtained in S2 into the deep dehazing network to obtain the deep dehazed image f(deephazeremoval).
[0064] S4. Weightedly fuse the deep dehazed map, the shallow dehazed map, and the original input to obtain the final dehazed clean picture, thus realizing picture dehazing.
[0065] Specifically, in the process, the original hazy image f(origin) is input. We send it into the shallow feature extraction network and the fog concentration information estimation module respectively to obtain the shallow feature map f(shallowfeature) and the fog concentration weight map f(hazemap) corresponding to the hazy image. Among them, f(shallowfeature) helps us more comprehensively restore the texture details of the hazy image, while the fog concentration weight map helps guide the deep dehazing network to adaptively remove the fog concentration in different regions of the real scene. Then, we splice f(orgin), f(shallowfeature), and f(hazemap) and send them into the deep dehazing network to obtain the deep dehazed image f(deephazeremoval). This network can help the image better remove haze in different regions. Finally, by performing weighted fusion on the shallow feature map, the deep dehazed image, and the original input, image dehazing can be achieved, and the final dehazed image can be obtained.
[0066] Based on the above process, the specific method for training the image dehazing method proposed by the present invention is as follows:
[0067] A1. Add corresponding loss functions to the shallow feature extraction network, the fog concentration information estimation module, and the deep dehazing network in the image dehazing model proposed by the present invention.
[0068] A2. Send the pre-training dataset in Embodiment 1 into the image dehazing method proposed by the present invention, and perform 300 rounds of pre-training on the network model.
[0069] A3. Send the hazy image into the image dehazing method proposed by the present invention.
[0070] A4. After the image dehazing method completes the entire process of the training data, each loss function will backpropagate the loss value back to each network, and then optimize and adjust the corresponding network weight values to complete the training of the corresponding image dehazing method.
[0071] In the above process, first perform 300 rounds of pre-training on the entire image dehazing network using the pre-training artificial dehazing dataset, and train the shallow image extraction network and the fog concentration information estimation module to a good degree; then perform 3000 rounds of training on the entire image dehazing model using the training dataset in the real scene; after the training is completed, save the trained model, and use the test dataset in the real scene to test the model. It can be seen that the dehazing results under the test set are as Figure 8 shown. It can be seen that the network already shows good restoration ability in the case of thick fog, can accurately reconstruct the shape and color of objects in the case of thick fog, and the restoration of the sky area can also adapt to the overall effect of the image. The effect generated by the network is close to the expectation from the perspective of the human eye.
[0072] Example 4:
[0073] In this example, the effect comparison of the method of the present invention and other image dehazing methods on the dataset Figure 9 and quantitative analysis graphs are provided. Among them, Figure 9 shows the experimental effects of dehazing pictures using the method of the present invention on different natural hazy pictures; from left to right are the original hazy pictures, the dehazed pictures of DCP, the dehazed pictures of AODNet, the dehazing results of FFANet, the dehazing results of ACERNet, the dehazing results of the present invention, and the corresponding haze-free pictures. Figure 9 It can be seen that from the perspective of human visual perception, the dehazing effect of the method of the present invention is clearer compared with other algorithms, and the haze removal in different regions is more thorough.
[0074] Quantitative analysis: Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are the most commonly used indicators to measure the difference between two images. Randomly select some hazy pictures, and calculate PSNR and SSIM by comparing the dehazing results of traditional dehazing algorithms and the method used in the present invention with clear pictures respectively, and we can get Figure 10 :
[0075] Specifically, Figure 10 The horizontal axis from left to right represents I-HAZE, O-HAZE, NTIRE2019, NTIRE2020 - 2021, and the average index value in turn; the vertical axis lists the original hazy pictures, DCP method, AODNet method, FFANet method, ACERNet method, SCANet method, and the method of the present invention from top to bottom in turn. From Figure 10 it can be seen that both the PSNR index and the SSIM index obtained by the algorithm proposed in the present invention are better than the existing image dehazing algorithms. Among them, the PSNR index performs excellently among many algorithms, and the SSIM index also exceeds most image dehazing methods, which shows that the algorithm proposed in the present invention has achieved excellent effects under the objective evaluation system of these two indicators.
[0076] From the effect comparison of each dehazing algorithm Figure 9 and quantitative analysis Figure 10 it can be seen that the algorithm used in the present invention is better than most existing algorithms both in visual effects and quantitative calculation indicators.
Claims
1. A method for image defogging based on fog concentration information, characterized in that: It includes a network model training phase and a model application phase, and the network model training phase specifically includes the following steps: S1. Obtain a training dataset for the image defogging network. Collect non-uniform foggy images and corresponding clear images in real natural scenes, and misty images under synthetic conditions and their corresponding clear images as network training datasets. S2. Build and train an image defogging network guided by fog concentration information, divide the dataset into different batches, and send the dataset into the defogging network in batches. S3. Take the original foggy data as input and use the proposed fog concentration estimation module to obtain the corresponding fog concentration mask map. S4. Take the original foggy data as input and use the shallow feature extraction network embedded with a multi-scale feature fusion module to obtain the shallow feature map of the data. S5. Take the original foggy data, the fog concentration mask map obtained in S3, and the shallow feature map of the data obtained in S4 as input, and use a deep defogging network embedded with a dynamic convolutional layer and a channel attention mechanism to obtain deep defogging image data. S6. Perform weighted fusion of the deep dehazed image, the original input image and the shallow dehazed image using the Mix module to obtain the final dehazed result.
2. The fog density estimation module according to claim 1, characterized in that: The fog density estimation module is mainly implemented through a multi-head self-attention mechanism, which includes an initial convolution layer, a multi-head self-attention layer and a sigmoid activation function. First, the input foggy picture is processed by the initial convolution layer. The initial convolution layer uses a 3x3 convolution kernel and several filters to perform preliminary feature extraction on the input image, and introduces nonlinearity through the ReLU activation function to generate the initial feature map. Then, the multi-head self-attention layer captures the global and local features of the input image through multiple parallel attention heads to generate a feature representation with rich contextual information. Subsequently, the feature map is converted into a single-channel fog density mask map through a 1x1 convolution kernel. Finally, the feature map is mapped to the [0,1] interval through the sigmoid activation function to generate a feature attention weight map.
3. The shallow feature extraction network embedded with a multi-scale feature fusion module according to claim 1, characterized in that: The input foggy image first passes through a shallow feature extraction network, which includes a pre-conversion module, a post-conversion module, and a multi-scale feature fusion module. First, the input image passes through the pre-convolution module. The first layer of convolution uses 32 3x3 convolution kernels to convolve the input three-channel RGB image to extract the initial features. After convolution, it is nonlinearly processed and instance normalized by the ReLU activation function. Instance normalization is performed on the high and wide dimensions of the image, so that the normalization information comes from itself, integrating and adjusting global information. Then, the feature map is input to the second layer of convolution. The second layer of convolution uses 64 3x3 convolution kernels for convolution to further extract features. After convolution, it is nonlinearly processed and instance normalized again by the ReLU activation function. The processed feature map is input to the multi-scale feature fusion module, which uses three different scale convolution kernels (3x3, 5x5, 7x7) to extract low, medium, and high scale features respectively. The extracted feature map is upsampled to the same resolution by bilinear interpolation. Subsequently, the three features are weightedly fused using the mixing coefficient extraction module, and the key features are enhanced through the channel attention mechanism. Finally, the fused features are added to the original input image to obtain an enhanced feature representation. Finally, the image enters the post-convolution module. The third convolution layer uses 64 3x3 convolution kernels to further process the fused feature map. The fourth convolution layer uses 32 3x3 convolution kernels to reduce the number of feature channels to 32. Finally, the fifth convolution layer uses 3 3x3 convolution kernels to restore the feature map to a three-channel image. After each convolution layer, a nonlinear transformation is performed through the ReLU activation function.
4. The deep defogging network according to claim 1, characterized in that: The dense fog removal network is built based on an encoding-decoding architecture, specifically including an encoder part that enhances feature expression through multiple dynamic convolutional layers and a channel attention mechanism. The encoder first uses two consecutive dynamic convolutional layers to extract the 7-channel features of the input using a 3x3 convolution kernel, and combines the ReLU activation function to generate 128-channel high-dimensional features, and then further enhances the feature information through a channel attention mechanism. The decoder part first uses a dynamic convolutional layer to convert the above features from 128 channels to 64 channels, and then uses two conventional 3x3 convolutional layers to gradually convert the features into 32 channels and 3 channels of the output. After each convolution layer, a nonlinear transformation is performed through the ReLU activation function. Finally, the Sigmoid activation function is used to ensure that the pixel values of the output image fall within the range of [0,1], thereby achieving effective removal and restoration of dense fog areas.
Citation Information
Cited By
Image defogging system and method for low-altitude scene
CN120598820A