Image defogging method and device based on deep learning
By introducing multi-scale enhanced downsampling module and cross-attention gating module into the image defog removal model, the problem of poor image defog removal effect in complex haze conditions in the prior art is solved, and more efficient defog removal effect and lower computational complexity are achieved.
Patent Information
- Application Number
- CN202510220593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When the prior art deals with image defog under complex haze conditions, it is difficult to effectively restore detailed information, and the calculation complexity is high, requiring a large number of data sets for training.
A multi-scale enhanced downsampling module and cross-attention gating module are introduced to build an image defog removal model for improved U-Net networks. Through the multi-head self-attention mechanism and cross-attention calculation, the multi-scale feature weights are dynamically adjusted to improve the defog removal effect.
It significantly improves the effect and robustness of image defog removal, especially in complex haze conditions, which can better restore image details, reduce the computational complexity and the need for large-scale data sets.
Smart Images

Figure CN120163734A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly, to an image dehazing method and apparatus based on deep learning. Background Art
[0002] With the rapid development of computer vision technology, image dehazing, as an important method to improve image quality and visual clarity, has wide application value in many fields such as autonomous driving, remote sensing monitoring, intelligent security, and medical image analysis. Affected by haze weather, the light in the image will be distorted due to scattering and absorption, resulting in a decrease in visual quality and loss of key information. This degradation phenomenon poses a great challenge to high-precision image analysis tasks.
[0003] Traditional image dehazing methods are mainly based on physical models. For example, the classic Atmospheric Scattering Model describes the relationship between direct radiation light and ambient light by assuming the mathematical form of image degradation. Representatives of these methods include Dark Channel Prior (DCP), Color Attenuation Prior (CAP), etc., which can restore the image by inverting optical parameters. However, these methods based on physical models usually rely on fixed prior assumptions, are sensitive to changes in lighting conditions, and are difficult to effectively process in complex scenes or thick fog conditions.
[0004] With the introduction of deep learning, the field of image dehazing has changed greatly. CNN-based methods (such as methods for directly estimating the transmission map and atmospheric light) bypass the need for traditional image priors, thus greatly improving the dehazing performance. These models have been proven to be effective in enhancing image visibility and restoring details. However, despite the success, CNNs also have their drawbacks. For example, although CNNs are good at capturing local features, they may not be able to effectively handle global dependencies, which are crucial for restoring details at different haze levels.
[0005] To address the limitations of CNNs, attention mechanisms have been increasingly incorporated into dehazing networks. Attention mechanisms allow the model to prioritize relevant features, enhancing its ability to handle different haze levels and improving the quality of the restored images. As a result, Transformer-based models have been developed and show great potential in image dehazing. Transformers are particularly successful in capturing long-range dependencies within images, and their self-attention mechanism can model the relationships between distant pixels, which is especially beneficial for solving the problem of uneven light scattering caused by haze. Transformer-based dehazing models have demonstrated excellent performance, especially when dealing with complex haze patterns, but they have higher computational complexity and require large datasets for training.
[0006] In addition to CNNs and Transformers, generative adversarial networks (GANs) have also been used for image dehazing. Using GANs for image dehazing can generate high-quality dehazed images with visual appeal. However, GAN-based methods typically require large amounts of paired training data, which may not be easily obtainable, especially in remote sensing environments. Moreover, GANs are prone to problems such as mode collapse, where the generator produces a limited variety of outputs, resulting in reduced diversity in the dehazing results.
[0007] In the prior art, diffusion models are becoming a common method for tasks involving image creation and restoration. Diffusion models obtain the ability to generate image data by gradually adding random noise to an image and learning the reverse process. Due to the powerful image generation ability of diffusion models, they can be applied to image dehazing. However, due to the complexity of the distribution of hazy images, directly combining diffusion models with image dehazing remains challenging. Additionally, similar to Transformers, diffusion models generally require a large amount of computational resources and large datasets to achieve optimal performance.
[0008] These latest developments reflect the trend of increasingly complex models and stronger dehazing performance. Although these newer methods (e.g., Transformers, GANs, and diffusion models, etc.) offer significant improvements in dehazing quality, they also pose challenges, especially in terms of computational efficiency and the need for large-scale training data. Moreover, the practical integration of these advanced technologies into more diverse dehazing scenarios, such as remote sensing dehazing, remains a key issue. Summary of the Invention
[0009] 1. Technical Problems to be Solved
[0010] In view of the deficiencies in the existing technology, the present invention provides an image dehazing method and device based on deep learning. By introducing a multi-scale enhanced downsampling module and a cross-attention gating module, it can effectively extract the detailed information in the image and dynamically adjust the multi-scale feature weights, thus significantly improving the dehazing effect. Especially when processing remote sensing images and ordinary scene images under complex haze conditions, it has higher robustness and generalization ability.
[0011] 2. Technical solutions
[0012] The object of the present invention is achieved through the following technical solutions.
[0013] An image dehazing method based on deep learning, comprising the following steps:
[0014] Collect a hazy image dataset and preprocess the hazy image dataset;
[0015] Construct an image dehazing model based on the improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the skip connection structure in the original U-Net network with a cross-attention gating module;
[0016] Use the preprocessed hazy image dataset to iteratively train the image dehazing model to obtain a trained image dehazing model;
[0017] Based on the trained image dehazing model, dehaze the hazy image to generate a dehazed image.
[0018] As a further improvement of the present invention, the image dehazing model includes a 1×1 convolutional layer, an encoder, a decoder, and a 1×1 convolutional layer connected in sequence;
[0019] The encoder includes several multi-scale enhanced downsampling modules and 3×3 convolutional layers, and the encoder generates encoder feature maps;
[0020] The decoder includes an upsampling module, and the decoder generates decoder feature maps;
[0021] The image dehazing model further includes a cross-attention gating module, and the cross-attention gating module connects the multi-scale enhanced downsampling module and the upsampling module.
[0022] As a further improvement of the present invention, the multi-scale enhanced downsampling module includes a 3×3 convolutional layer, a max pooling layer, a batch normalization layer, a parallel convolutional layer, a concat module, a concat module, and a 1×1 convolutional layer connected in sequence;
[0023] The multi-scale enhanced downsampling module further includes a multi-head self-attention mechanism.
[0024] As a further improvement of the present invention, the parallel convolution layer includes three convolution layers connected in parallel after the batch normalization layer, and the three convolution layers are a 3×3 convolution layer, a 5×5 convolution layer, and a 7×7 convolution layer respectively.
[0025] As a further improvement of the present invention, the pre-processed foggy image is processed by the multi-scale enhancement downsampling module to generate a multi-scale feature map, and the steps include:
[0026] Input the pre-processed foggy image into the image dehazing model, and generate a shallow feature map through the 1×1 convolution layer downsampling operation in the image dehazing model;
[0027] Input the shallow feature map into the multi-scale enhancement downsampling module, and the 3×3 convolution layer and the max pooling layer perform downsampling operations on the input feature map to generate a feature map, and then perform batch normalization operations on the feature map;
[0028] Process the batch-normalized feature map through the parallel convolution layer to generate a first feature map, a second feature map, and a third feature map;
[0029] The first feature map, the second feature map, and the third feature map are concatenated through the concat module to generate a multi-scale feature map.
[0030] As a further improvement of the present invention, the batch-normalized feature map is enhanced by the multi-head self-attention mechanism to generate an attention-enhanced feature map.
[0031] As a further improvement of the present invention, the multi-scale feature map and the attention-enhanced feature map are concatenated through the concat module to generate an encoder feature map.
[0032] As a further improvement of the present invention, the decoder feature map and the encoder feature map are subjected to cross-attention calculation through the cross-attention gating module to obtain a gating value, and the calculation formula is:
[0033] gate = Cross-attention(F prev , F skip )
[0034] where, gate represents the gating value, Cross-attention represents cross-attention calculation, F prev represents the decoder feature map, and F skip represents the encoder feature map.
[0035] As a further improvement of the present invention, after adjusting the contribution of the decoder feature map by the gating value, it is combined with the encoder feature map to obtain an output feature map, and the calculation formula is:
[0036] F out = gate × F prev + F skip
[0037] Among them, F out represents the output feature map.
[0038] An image dehazing device based on deep learning, comprising:
[0039] A data processing module, which collects a hazy image dataset and preprocesses the hazy image dataset;
[0040] A model construction module, which constructs an image dehazing model based on an improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the skip connection structure in the original U-Net network with a cross-attention gating module;
[0041] A model training module, which iteratively trains the image dehazing model using the preprocessed hazy image dataset to obtain a trained image dehazing model;
[0042] A model detection module, which dehazes the hazy image based on the trained image dehazing model to generate a dehazed image.
[0043] 3. Beneficial effects
[0044] Compared with the prior art, the advantages of the present invention are as follows:
[0045] (1) For the image dehazing method and device based on deep learning of the present invention, by introducing a multi-scale enhanced downsampling module, it can effectively capture image detail information at different scales, enhance the image texture features, ensure a clear visual effect can still be restored under thick fog conditions, thereby improving the image dehazing performance; by calculating the similarity weights between different feature maps through the cross-attention gating module, dynamically fusing the information between the encoder and the decoder, enhancing the feature synergy, dynamically fusing the feature maps from multiple paths, and by focusing on the specific encoder features required for dehazing, reducing the redundancy risk in feature fusion, which helps to provide a clearer dehazing output and significantly improves the dehazing effect.
[0046] (2) For the image dehazing method and device based on deep learning of the present invention, by constructing an image dehazing model based on an improved U-Net network, it is not only applicable to image dehazing in ordinary scenes, but also optimized for the complex characteristics of remote sensing images, with strong practicability and wide applicability. In addition, through the image dehazing model based on the improved U-Net network, while improving the dehazing performance, the network structure is optimized through modular design, taking into account both the dehazing effect and the calculation efficiency, meeting the requirements for real-time and high efficiency in practical applications. Description of the Drawings
[0047] Figure 1 It is a flowchart of the method according to the embodiment of the present invention;
[0048] Figure 2 It is a schematic structural diagram of the image dehazing model according to the embodiment of the present invention;
[0049] Figure 3 It is a schematic structural diagram of the multi-scale enhanced downsampling module according to the embodiment of the present invention;
[0050] Figure 4 It is a schematic structural diagram of the cross-attention gating module according to the embodiment of the present invention;
[0051] Figure 5 It is the hazy image input according to the embodiment of the present invention;
[0052] Figure 6 It is the dehazed image after processing according to the embodiment of the present invention. Detailed Embodiments
[0053] The present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0054] Embodiment
[0055] As Figure 1 shown, a deep learning-based image dehazing method provided in this embodiment includes the following steps: collecting a hazy image dataset and preprocessing the hazy image dataset; constructing an image dehazing model based on an improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the skip connection structure in the original U-Net network with a cross-attention gating module; iteratively training the image dehazing model using the preprocessed hazy image dataset to obtain a trained image dehazing model; and dehazing the hazy image based on the trained image dehazing model to generate a dehazed image.
[0056] Specifically in this embodiment, paired clear images and hazy images in indoor environments, outdoor environments, or remote sensing environments are obtained, and the clear images and hazy images are paired to construct a paired hazy image and clear image dataset. The hazy images and corresponding clear images are collected by devices such as drones, satellites, or cameras. Thus, in this embodiment, if the hazy image is set as I and its corresponding clear image is J, the image dehazing model receives the hazy image I as input and outputs a prediction of the clear image J.
[0057] It should be noted that in this embodiment, an image dataset is collected through the existing Haze1K public dataset, and the images in this dataset are satellite remote sensing images. In this embodiment, the foggy image dataset is preprocessed, including normalizing the pixel values of the foggy images to the interval [0, 1] and adjusting the resolution to a fixed size. In this embodiment, 256 pixels × 256 pixels are used to adapt to the network input requirements.
[0058] Furthermore, an image dehazing model based on an improved U-Net network is constructed. In this embodiment, as Figure 2 shown, the image dehazing model includes a 1×1 convolutional layer, an encoder, a decoder, and a 1×1 convolutional layer connected in sequence.
[0059] It is worth noting that in this embodiment, the downsampling module in the original U-Net network encoder is replaced with a multi-scale enhanced downsampling module. Thus, the encoder includes a number of multi-scale enhanced downsampling modules and 3×3 convolutional layers connected in sequence, and the encoder is used to generate encoder feature maps. The decoder includes a number of upsampling modules connected in sequence, and the decoder is used to generate decoder feature maps.
[0060] In this embodiment, the image dehazing model also replaces the skip connection structure in the original U-Net network with a cross-attention gating module. The cross-attention gating module connects the multi-scale enhanced downsampling module in the encoder and the downsampling module in the decoder.
[0061] As Figure 3 shown, the multi-scale enhanced downsampling module includes a 3×3 convolutional layer, a max pooling layer, a batch normalization layer, a parallel convolutional layer, a concat module, a concat module, and a 1×1 convolutional layer connected in sequence. The parallel convolutional layer includes three convolutional layers connected in parallel after the batch normalization layer, and the three convolutional layers are a 3×3 convolutional layer, a 5×5 convolutional layer, and a 7×7 convolutional layer respectively.
[0062] In this embodiment, the multi-scale enhanced downsampling module also includes a multi-head self-attention mechanism. By integrating multi-kernel convolution and the multi-head self-attention mechanism into the downsampling module of the U-Net network, a multi-scale enhanced downsampling module is obtained, which improves the information encoding ability in the encoder stage.
[0063] In the multi-scale enhanced downsampling module, not only the traditional convolutional layer and the max pooling with a window size of 2 are used to reduce the spatial dimension of the feature map and double the number of channels, but also the multi-kernel convolution and the multi-head self-attention mechanism are combined to extract multi-scale information from the feature map to enhance the information encoding ability of the encoder.
[0064] Furthermore, the preprocessed foggy image Input into the image dehazing model, and through the downsampling operation of the 1×1 convolutional layer in the image dehazing model, a shallow feature map is generated. Among them, H represents the height of the hazy image, W represents the width of the hazy image, 3 represents the number of channels of the hazy image, H' represents the height of the shallow feature map, W' represents the width of the shallow feature map, and C represents the number of channels of the shallow feature map.
[0065] In this embodiment, a three-stage symmetric encoder-decoder structure is used for feature extraction and image reconstruction, so as to convert the shallow feature map into a deep feature map.
[0066] In the encoder, the multi-scale enhanced downsampling module is used to perform downsampling operations on the shallow feature map F to reduce the spatial dimension of the shallow feature map F. Through compression, the image dehazing model can capture a wider range of context information and reduce the computational cost. At the same time, as the spatial dimension decreases, the number of feature channels in each subsequent layer doubles, enabling the image dehazing model to capture more complex and abstract features.
[0067] Specifically, the shallow feature map F is input into the multi-scale enhanced downsampling module, and the 3×3 convolutional layer and the max pooling layer perform downsampling operations on the input feature map to generate a feature map, and then batch normalization operations are performed on the feature map. In this embodiment, a series of convolutional layers with the same convolutional kernel size are connected to form a convolutional path, and a ReLU activation function is connected behind each convolutional layer on the convolutional path. The batch-normalized feature map is processed through parallel convolutional layers to generate the first feature map, the second feature map, and the third feature map. The first feature map, the second feature map, and the third feature map are concatenated through the concat module to generate a multi-scale feature map. At the same time, the batch-normalized feature map is also input into the multi-head self-attention mechanism, and feature enhancement is performed through the multi-head self-attention mechanism to generate an attention-enhanced feature map. The multi-scale feature map and the attention-enhanced feature map are concatenated through the concat module to generate an encoder feature map.
[0068] Thus, in this embodiment, by integrating the multi-scale information extraction mechanism in the downsampling process, the detailed features of the hazy image at different scales can be fully captured, effectively improving the dehazing performance. In addition, it can not only well restore the overall appearance of the hazy image, but also take into account the fine texture and detail restoration in the hazy image.
[0069] In the decoder, the upsampling module receives feature maps as inputs from two paths, including the feature maps output by the 3×3 convolutional layer of the encoder connected to the upsampling module, and the encoder feature maps passed through the skip connection of the cross-attention gating module. After the decoder fuses the features of these two paths, it performs an upsampling operation through a transposed convolutional layer with a stride of 2, gradually restoring the dimensions and number of channels of the feature maps that were changed during the downsampling process to match the dimensions of the target output. Thus, in the decoder, a high-definition dehazed image is reconstructed through layer-by-layer upsampling operations.
[0070] It should be noted that the skip connection structure in the traditional U-Net network connects features after convolution, fuses the skip connection with the main path, and directly transmits low-level features from the encoder to the corresponding layer in the decoder through the connection, helping the U-Net network retain the spatial information lost during the downsampling stage and playing a role in preventing gradient disappearance. However, the skip connection structure in the traditional U-Net network is not sufficient to well reflect the relationship between the features in the decoder and the low-level features from the encoder. Thus, in this embodiment, a cross-attention mechanism is introduced into the skip connection of the U-Net network, and by calculating the similarity weights between different feature maps, the information between the encoder and the decoder is dynamically fused to enhance the feature synergy. By replacing the skip connection structure in the original U-Net network with a cross-attention gating module, the feature maps from multiple paths can be dynamically fused.
[0071] As Figure 4 shown, the gating value is obtained by performing cross-attention calculation on the decoder feature map and the encoder feature map through the cross-attention gating module. The calculation formula is:
[0072] gate = Cross-attention(F prev , F skip )
[0073] where gate represents the gating value, Cross-attention represents cross-attention calculation, F prev represents the decoder feature map, and F skip represents the encoder feature map.
[0074] After adjusting the contribution of the decoder feature map through the gating value and then combining it with the encoder feature map, the output feature map is obtained. The calculation formula is:
[0075] F out = gate × F prev + F skip
[0076] where F out represents the output feature map.
[0077] Thus, in this embodiment, the feature weights in the skip connection are dynamically adjusted through the cross-attention gating module, which helps the decoder selectively suppress irrelevant information and ensures that only important information can participate in the defogging process. Compared with the skip connection method of the traditional U-Net network, by focusing on specific encoder features required for defogging, this embodiment reduces the redundancy risk in feature fusion, helps to provide a clearer defogging output without modifying unaffected areas.
[0078] Finally, after the upsampling is completed, a 1×1 convolutional layer is used to reduce the feature dimension of the output feature map from C to 3, obtaining the final output of the image defogging model, that is, the generated defogged image.
[0079] In this embodiment, the image defogging model is iteratively trained using the preprocessed hazy image dataset, and the change trend of the loss function is observed until the loss converges, obtaining a trained image defogging model. Based on the trained image defogging model, the hazy image is defogged to generate a defogged image. As Figure 5 shown, it is the input hazy image, as Figure 6 shown, it is the defogged image processed by the method of this embodiment.
[0080] A deep learning-based image defogging method provided in this embodiment can effectively extract detailed information in the image and dynamically adjust the multi-scale feature weights by introducing a multi-scale enhanced downsampling module and a cross-attention gating module, thus significantly improving the defogging effect. Especially when processing remote sensing images and ordinary scene images under complex haze conditions, it has higher robustness and generalization ability.
[0081] A deep learning-based image defogging method provided in this embodiment is comprehensively improved and optimized in a modular manner. The modular design not only improves the flexibility of the network structure but also facilitates subsequent functional expansion, enabling it to quickly integrate new functions. At the same time, this design can be flexibly adjusted according to the specific requirements of different application scenarios, and the demand for computing resources is significantly reduced, enabling efficient operation with lower hardware configurations. Therefore, it can also maintain excellent performance on resource-constrained devices (such as embedded systems or edge computing devices), significantly expanding its scope of practical applications and having strong practicality and wide usability.
[0082] This embodiment also provides an image dehazing device based on deep learning, including a data processing module, a model construction module, a model training module, and a model detection module. The data processing module collects a hazy image dataset and preprocesses the hazy image dataset. The model construction module constructs an image dehazing model based on an improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the skip connection structure in the original U-Net network with a cross-attention gating module. The model training module iteratively trains the image dehazing model using the preprocessed hazy image dataset to obtain a trained image dehazing model. The model detection module dehazes the hazy image based on the trained image dehazing model to generate a dehazed image. The image dehazing device based on deep learning provided in this embodiment can implement any of the methods of the image dehazing method based on deep learning, and the specific working process of the image dehazing device based on deep learning can refer to the corresponding process in the embodiment of the image dehazing method based on deep learning. The methods and devices provided in this embodiment can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed connections or communication connections with each other can be indirect couplings or communication connections through some interfaces, devices or units, and can also be electrical, mechanical or other forms of connections.
[0083] This embodiment also provides a computer device. A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the image dehazing method based on deep learning described above.
[0084] This embodiment also provides a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the image dehazing method based on deep learning described in this embodiment. Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device; the program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0085] The above has schematically described the present invention and its implementation manners. This description is not restrictive. Without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Any reference signs in the claims should not limit the claims involved. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of this creation, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention. In addition, the term "comprising" does not exclude other elements or steps, and the term "a" before an element does not exclude including "a plurality of" such elements. The plurality of elements stated in the product claims can also be implemented by one element through software or hardware. The terms such as "first" and "second" are used to represent names and do not indicate any specific order.
Claims
1. A deep learning-based image defogging method comprising the following steps: Collect foggy image data sets and preprocess the foggy image data sets; Construct an image dehazing model based on an improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the jump connection structure in the original U-Net network with a cross-attention gating module; The image defogging model is iteratively trained using the preprocessed foggy image dataset to obtain a trained image defogging model; The foggy image is defogged based on the trained image defogging model to generate a defogged image.
2. The image defogging method based on deep learning according to claim 1, characterized in that: The image dehazing model includes a 1×1 convolutional layer, an encoder, a decoder and a 1×1 convolutional layer connected in sequence; The encoder includes a plurality of multi-scale enhanced downsampling modules and a 3×3 convolutional layer, and the encoder generates an encoder feature map; The decoder comprises an upsampling module, the decoder generating a decoder feature map; The image dehazing model also includes a cross-attention gating module, which connects the multi-scale enhanced downsampling module and the upsampling module.
3. The image defogging method based on deep learning according to claim 2, characterized in that: The multi-scale enhanced downsampling module includes a 3×3 convolution layer, a maximum pooling layer, a batch normalization layer, a parallel convolution layer, a concat module, a concat module and a 1×1 convolution layer connected in sequence; The multi-scale enhanced downsampling module also includes a multi-head self-attention mechanism.
4. The image defogging method based on deep learning according to claim 3, characterized in that: The parallel convolution layer includes three convolution layers connected in parallel after a batch normalization layer, and the three convolution layers are respectively a 3×3 convolution layer, a 5×5 convolution layer and a 7×7 convolution layer.
5. The image defogging method based on deep learning according to claims 1-4, characterized in that: The pre-processed foggy image is processed by the multi-scale enhanced downsampling module to generate a multi-scale feature map, and the steps include: The preprocessed foggy image is input into the image defogging model, and a shallow feature map is generated through the 1×1 convolutional layer downsampling operation in the image defogging model; The shallow feature map is input into the multi-scale enhanced downsampling module. The 3×3 convolution layer and the maximum pooling layer downsample the input feature map to generate a feature map, which is then batch normalized. Processing the batch normalized feature map through parallel convolutional layers to generate a first feature map, a second feature map, and a third feature map; The first feature map, the second feature map and the third feature map are concatenated through a concat module to generate a multi-scale feature map.
6. The image defogging method based on deep learning according to claim 5, characterized in that: The batch-normalized feature map is enhanced through a multi-head self-attention mechanism to generate an attention-enhanced feature map.
7. The image defogging method based on deep learning according to claim 6, characterized in that: The multi-scale feature map and the attention-enhanced feature map are concatenated through the concat module to generate the encoder feature map.
8. The image defogging method based on deep learning according to claim 7, characterized in that: The cross-attention gating module performs cross-attention calculation on the decoder feature map and the encoder feature map to obtain a gating value, and the calculation formula is: gate=Cross-attention(F prepv ,F skip ) Among them, gate represents the gate value, Cross-attention represents the cross-attention calculation, and F prev represents the decoder feature map, F skip Represents the encoder feature map.
9. The image defogging method based on deep learning according to claim 8, characterized in that: After adjusting the contribution of the decoder feature map through the gate value, it is combined with the encoder feature map to obtain the output feature map. The calculation formula is: F out =gate×F prep +F skip Among them, F out Represents the output feature map.
10. An image defogging device based on deep learning, characterized in that: include: A data processing module collects foggy image data sets and preprocesses the foggy image data sets; Model building module, building an image dehazing model based on the improved U-Net network, including replacing the downsampling module in the original U-Net network encoder with a multi-scale enhanced downsampling module and replacing the jump connection structure in the original U-Net network with a cross-attention gating module; The model training module uses the preprocessed foggy image dataset to iteratively train the image defogging model to obtain a trained image defogging model; The model detection module defogs the foggy image based on the trained image defogging model to generate a defogged image.
Citation Information
Cited By
Polarization perception adaptive defogging method and device based on autonomous prompt
CN121437318A
Two-stage image defogging method based on prior fog concentration characteristics
CN122453663A