An image dehazing system and method for low-altitude scenes

By using frequency domain fusion, spatial and channel interaction, and physical perception modules in a U-shaped network architecture, the problem of non-uniform fog distribution and result distortion in image dehazing in low-altitude scenarios is solved, achieving more efficient image clarity and consistency restoration.

CN120598820BActive Publication Date: 2025-10-31SHANDONG WEIRAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511086079.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-31
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Image dehazing techniques in low-altitude scenarios face challenges such as non-uniform fog distribution, insufficient utilization of frequency domain information, result distortion under specific conditions, and insufficient data and generalization issues. Existing methods are not effective under complex fog and haze conditions.

Method used

A U-shaped network architecture is adopted, which combines a frequency domain fusion module, a spatial and channel interaction module, and a physical perception module. Through wavelet transform processing and atmospheric scattering model, frequency domain and spatial information are processed in different layers of modules respectively. By using attention mechanism and physical parameter estimation, image dehazing and restoration are achieved.

Benefits of technology

It improves the clarity and consistency of images in low-altitude scenes, solves the problem of distortion in defogging results in non-uniform fog distribution and large-area fog and haze areas, and improves the defogging performance and practical application effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598820B_ABST
    Figure CN120598820B_ABST
Patent Text Reader

Abstract

This invention provides an image dehazing system and method for low-altitude scenes, belonging to the field of image restoration technology based on computer vision. It is constructed based on a U-shaped network, utilizing the network's varying sensitivities to frequency domain information to guide different frequency components to different layers. Frequency domain fusion, spatial and channel interaction, and physical perception modules are designed in shallow, intermediate, and deep layers, respectively. The frequency domain fusion module provides richer feature representations for subsequent modules. The spatial and channel interaction module in the intermediate layer simultaneously captures important features in both spatial and channel dimensions, improving the local clarity and global consistency of the dehazed image. The deep physical perception module, combined with prior information from the physical model, provides stronger constraints and interpretability for the dehazing task. This invention solves the problem of color distortion in dehazing results for non-uniform fog distribution and large-area fog / haze regions in low-altitude scenes, improving the network's ability to model image degradation characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image restoration technology based on computer vision, and particularly relates to an image dehazing system and method for low-altitude scenes. Background Technology

[0002] With the development of low-altitude technology, low-altitude operations have been widely applied in various fields such as logistics transportation, agricultural monitoring, and environmental protection. However, during low-altitude missions, especially in low-altitude flight and complex foggy weather conditions, image quality is severely affected by atmospheric particulate matter, leading to decreased image contrast, blurred details, and even complete loss of target object visibility. These problems directly affect the accuracy and reliability of key tasks in low-altitude operations, such as navigation, target detection, and object recognition. Therefore, how to effectively remove the interference of fog and haze on low-altitude images and restore clear, fog-free images has become an important research direction in the field of low-altitude visual perception.

[0003] Image dehazing technology, as a core method for solving this problem, has received widespread attention from scholars both domestically and internationally in recent years. In particular, dehazing techniques combined with deep learning have shown great application potential in low-altitude scenarios through modeling image degradation models and data-driven feature learning capabilities. However, due to the unique characteristics of low-altitude imaging, such as diverse viewpoints, complex environments, low resolution and noise interference, and uneven fog distribution, image dehazing in low-altitude scenes still faces many challenges.

[0004] Traditional image dehazing methods rely on atmospheric scattering models, estimating atmospheric light values ​​A and transmission maps t(x) to recover haze-free images based on physical principles. These methods have some application value in low-altitude scenes, but suffer from significant limitations in complex scenes. Early image dehazing research was spearheaded by the Dark Channel Prior (DCP) method proposed by He et al. in 2009. DCP assumes that at least one pixel value close to 0 exists in a local region of the haze-free image, using this prior to estimate the transmission map and recover the haze-free image. This method is computationally simple and effective, but performs poorly in smooth regions (such as the sky). Subsequently, Fattal proposed a color consistency-based dehazing method, enhancing the accuracy of dehazing by analyzing the independence of scene radiance and transmittance.

[0005] To improve dehazing performance, many researchers have refined parameter estimation methods for atmospheric scattering models. For example, methods based on color linearization priors and minima priors improve model robustness by introducing new assumptions. Furthermore, optimization-based methods (such as Laplacian regularization and total variational optimization) are widely used for smoothing transport maps, resulting in more natural-looking dehazed images with better detail and clarity at the edges.

[0006] While current methods for dehazing images in low-altitude scenarios have improved the quality of reconstructed images, they still face some key challenges. The first is non-uniform fog distribution: when low-altitude operations are conducted in complex terrain and weather conditions, fog and haze often exhibit a non-uniform distribution, causing the prior assumptions of traditional physics-based models to fail, and deep learning models also struggle to accurately model the problem.

[0007] Secondly, the utilization of frequency domain information is insufficient: Effective use of frequency domain information remains a challenge in the dehazing process of low-altitude images. Many existing methods rely primarily on spatial domain features when processing images, neglecting the important information contained in the frequency domain.

[0008] Then there's the issue of distortion in dehazing results under specific conditions: when processing images containing large areas of sky or sea, dehazing algorithms often suffer from distortion. The color and brightness distribution in these areas is relatively uniform, making it difficult for the algorithm to accurately estimate the effects of transmittance and fog, resulting in color deviations or loss of detail.

[0009] Finally, there are issues of insufficient data and generalization: there are currently few real dehazing datasets for low-altitude scenes, and existing models mostly rely on synthetic data for training. However, the difference in distribution between synthetic data and real data can lead to a decrease in the performance of models in practical applications. Summary of the Invention

[0010] To address the aforementioned problems, the present invention provides a first aspect of an image dehazing system for low-altitude scenes, which is constructed based on a U-shaped network. By utilizing the different sensitivities of the U-shaped network to frequency domain information, different frequency components are guided to different layers, and frequency domain fusion modules, spatial and channel interaction modules, and physical sensing modules are designed in shallow, intermediate, and deep layers, respectively.

[0011] The frequency domain fusion module takes three high-frequency components of the feature map after wavelet transform as input, and outputs a global frequency domain similarity map after being processed by the global attention mechanism of the frequency domain features. This map is used to capture the similarity between pixels in the hazy degraded image and to use the relatively clear light fog areas in the image to help restore the dense fog area image.

[0012] The input to the spatial and channel interaction module is the feature map after frequency domain fusion. First, the spatial attention and channel attention of the input feature map are calculated and connected in the channel dimension. Then, random shuffling and group convolution operations are performed to reduce the uneven distribution of haze images in the channel level.

[0013] The physical sensing module is designed based on an atmospheric scattering model. The input is the low-frequency component of the feature map after wavelet transform processing. The low-frequency component information is used as the key parameters of the model. The atmospheric diffusion model formula is used to reconstruct the dehazing and restored image in the feature space.

[0014] Preferably, the frequency domain fusion module uses an improved self-attention module for feature extraction and enhances the focus on key frequency domain features;

[0015] The spatial and channel interaction module integrates spatial and channel information from global average pooling and global max pooling in a multi-scale structure. The channel information needs to be randomly shuffled to promote cross-channel frequency domain information exchange, while the spatial level is responsible for highlighting areas with significant haze distribution.

[0016] The physical sensing module learns physical parameters, including global atmospheric light, through a network. and medium transmittance Furthermore, by combining low-frequency component information, physical parameters are estimated, providing physical constraints for the defogging task.

[0017] Preferably, the frequency domain fusion module specifically comprises:

[0018] The encoder part of the network performs frequency domain decomposition on the input features using discrete wavelet transform to obtain four frequency domain components. Three high-frequency components—the horizontal high-frequency subband LH, the vertical high-frequency subband HL, and the diagonal high-frequency subband HH—are selected as the input to this module. The frequency domain fusion module contains three branches. The first branch performs a concatenation operation on LH, HL, and HH. The second branch uses the horizontal and vertical high-frequency texture information (LH and HL) for attention calculation. It first reconstructs the feature dimensions and performs a 1×1 convolution to obtain channel-adjusted feature maps for the horizontal high-frequency subband LH, the vertical high-frequency subband HL, and the diagonal high-frequency subband HH. , and Then through Normalization yields an attention matrix of size C×C, where C is the number of channels of the input feature. The third branch multiplies the HH component with the attention matrix obtained from the second branch to obtain a weighted feature result, which is then added element-wise to the residual obtained from the first branch to form a frequency domain fusion feature representation.

[0019] Preferably, the space and channel interaction module specifically comprises:

[0020] The algorithm comprises two branches that learn spatial and channel information respectively. The two branches are merged along the channel dimension and then randomly shuffled. For the spatial information branch, global features in the spatial dimension are extracted from the feature map output by the previous layer using Global Average Pooling (GAP) and Global Max Pooling (GMP), followed by a concatenation operation. Then, spatial attention weights are learned through a 3×3 convolution operation to highlight areas with significant haze distribution. For the channel information branch, global features in the channel dimension are extracted using GAP and GMP, followed by a concatenation operation. Then, channel weights are learned through a 1×1 convolution operation to enhance the modeling of different semantic features. The spatial and channel information branches are added at the element-wise level, and finally, a channel shuffle operation is used to achieve feature interaction between spatial and channel attention.

[0021] The input feature map is The feature map has C, H, and W channels, and its height and length are respectively. The generation process of channel attention first involves processing the input feature map. Perform global average pooling to generate global features in the spatial dimension. Using 1×1 convolution pairs Perform feature compression and pass Activation function generates channel attention weights First, the input feature map Global average pooling and global max pooling are performed along the channel dimension to multiply the two features element-wise, and then spatial attention weights are generated through 3×3 convolution.

[0022] Preferably, the spatial and channel interaction module also introduces a channel shuffling operation, which enables the redistribution of information between channels. The channels are divided into multiple groups, and a shuffling operation is performed within each group of channels. Through the channel shuffling operation, the information across channels is rearranged and interacted. The weights after channel shuffling are processed by 3×3 convolution to generate the final channel-specific spatial importance map. The importance map is normalized using the sigmoid activation function to obtain the final attention map.

[0023] Preferably, the formula for the atmospheric diffusion model in the physical sensing module is:

[0024]

[0025] In the formula Images representing smog, Represents a clean image without fog. It represents the transmittance of the medium in the atmosphere. Representing atmospheric light parameters, the process of reconstructing the dehazing recovery results utilizes model learning. and These two key physical parameters ultimately lead to a more accurate haze-free image. ;

[0026] The input of the module serves as the deepest layer of the network, and its input is the feature map from the previous layer. It contains two branches that derive estimates in reverse order for estimating global atmospheric light. and medium transmittance For estimating atmospheric light The branch first uses global average pooling. The overall brightness of the image is learned, and then through dimensionality-up convolution and dimensionality-down convolution operations, a feature map with more aggregated useful features is obtained, which is then utilized... The final estimated value is obtained; finally, using the estimated physical parameters, an approximate clear image is calculated according to the atmospheric scattering model formula.

[0027] A second aspect of the present invention provides an image dehazing method for low-altitude scenes, comprising the following steps:

[0028] Capture and obtain images of low-altitude scenes;

[0029] The image is input into the low-altitude scene image dehazing system as described in the first aspect for target segmentation;

[0030] Output the target image after dehazing.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The system architecture designed in this invention is an improvement upon a U-shaped architecture comprising three layers of upsampling and downsampling. It mainly consists of three parts: an attention-based frequency domain fusion module, a spatial and channel interaction module, and a physical perception module. The shallow attention-based frequency domain fusion module performs frequency domain decomposition and fusion of input features in the encoder part of the network using discrete wavelet transform. The fused frequency domain features contain both global low-frequency background information and retain local high-frequency details, providing richer feature representations for subsequent modules. The intermediate spatial and channel interaction module captures important features in both spatial and channel dimensions, significantly improving the local clarity and global consistency of the dehazed image. The deep physical perception module, combined with prior information from the physical model, provides stronger constraints and interpretability for the dehazing task. This invention solves the problem of color distortion in dehazing results for non-uniform fog distribution and large-area haze regions in low-altitude scenes, improving the model's ability to model the degradation characteristics of haze images. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the following description is only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a diagram of the overall system structure of the present invention.

[0035] Figure 2 This is a schematic diagram of the frequency domain fusion module of the present invention.

[0036] Figure 3 This is a schematic diagram of the spatial and channel interaction module structure of the present invention.

[0037] Figure 4 This is a schematic diagram of the physical sensing module structure of the present invention.

[0038] Figure 5 This is a diagram showing the defogging effect of the model of the present invention in the implementation method. Detailed Implementation

[0039] The invention will be further described below with reference to specific embodiments.

[0040] The haze image data used in this embodiment comes from drone photography and the Reside public dataset. The drone photography dataset uses a DJI M3T drone as the image acquisition device, equipped with a full-frame camera and a high-sensitivity sensor with 45 million effective pixels. This embodiment sets up the drone to fly along a specific route consisting of 51 waypoints. A total of 4239 images were obtained at drone altitudes of 300 meters, 600 meters, and 900 meters. Each image in the dataset has a resolution of 4000×3000 pixels. Images taken by the drone at the same waypoint have slight errors due to the drone's limited GPS positioning accuracy, making it impossible to capture paired haze and fog-free images for training. To address this issue, this implementation only captures fog-free images as the baseline, and obtains haze images through depth information synthesis, specifically as follows:

[0041] The MonoDepth2 deep learning model was selected. It was first pre-trained on public deep datasets (such as KITTI and NYU DepthV2), and then fine-tuned using transfer learning for fog-free image scenes captured by drones.

[0042] Depth map generation process: The captured fog-free image is input into a trained depth estimation model. After processing through the model's feature extraction layer and depth prediction layer, combined with post-processing optimization (using a bilateral filtering algorithm to smooth the prediction results), a depth map corresponding to the input image pixels is output. Each pixel value in the depth map represents the actual distance between the corresponding object and the drone camera. For each pixel of the fog-free image J(x), the atmospheric scattering model formula is used for calculation:

[0043]

[0044] That is, based on the depth value of the pixel. Calculate transmittance Combined with the set atmospheric light value =0.6, solving for the corresponding pixel value of the haze image. By performing pixel-by-pixel calculations, we generated complete simulated haze images, achieving the simulation of haze effects with different concentrations and distribution characteristics. Based on 4239 clean, haze-free images captured, we obtained corresponding haze images using the above method, thus possessing a paired low-altitude scene haze training dataset.

[0045] To address the problem of image dehazing modeling for real-world scenes, the impact of haze in low-altitude scene images manifests in the frequency domain as the loss of high-frequency details and blurring of low frequencies. Therefore, this invention designs a multi-layered modeling mechanism based on frequency domain characteristics, introducing different frequency information into different layers of a U-shaped network to enhance the network's ability to model details and global information. This mechanism mainly includes an attention-based frequency domain fusion module, a spatial and channel interaction module, and a physical perception module. The overall structure is as follows: Figure 1 As shown, the hierarchical structure has three levels, with different blocks used at each level to extract corresponding features.

[0046] The shallow structure employs an attention-based frequency domain fusion module, such as... Figure 2As shown, the encoder part of the network performs frequency domain decomposition of the input features using discrete wavelet transform, using four kernels to decompose the input features into four sub-bands. LL: Low-frequency sub-band, containing background information and global structural features of the image. LH, HL, HH: High-frequency sub-bands, representing high-frequency texture information in the horizontal, vertical, and diagonal directions, respectively. For the high-frequency sub-bands (LH, HL, HH), this invention designs an attention-based fusion mechanism, using one-dimensional convolution and depthwise convolution to learn the local dependencies between high-frequency features. A weight map of high-frequency features is generated through the attention mechanism to guide the selection of important high-frequency information. Low-frequency feature enhancement: Low-frequency information is compressed using a 1×1 convolution to extract global features. The low-frequency features are combined with the high-frequency features to form a frequency-domain fused feature representation. First, the three frequency-domain features are multiplied element-wise to obtain a new feature representation. Then, a 1×1 convolution is applied to this new feature representation to further extract and fuse features. Next, a dot product operation is performed on the feature representations of LH and HL, and an activation function is used to generate attention weights. These weights are used to adjust the feature representation of HH. The adjusted HH feature representation is added to the feature representation obtained through 1×1 convolution to obtain the final fused feature.

[0047] The fused frequency domain features contain both global low-frequency background information and retain local high-frequency details, providing richer feature representations for subsequent modules. The specific process is represented by the following formula, where... This represents the channel dimension transformed by a 1×1 convolution. This represents performing a connection operation at the channel level. , and Feature maps representing the horizontal high-frequency subband (LH), vertical high-frequency subband (HL), and diagonal high-frequency subband (HH) after adjustment:

[0048]

[0049] The fused frequency domain features contain both global low-frequency background information and retain local high-frequency details, providing richer feature representations for subsequent modules.

[0050] The intermediate layer design incorporates multi-scale spatial channel information fusion using spatial and channel interaction modules, such as... Figure 3As shown, the module can simultaneously capture important features in both spatial and channel dimensions, significantly improving local sharpness and global consistency of dehazed images, especially in non-uniform haze distribution scenarios. For spatial information, Global Average Pooling (GAP) and Global Max Pooling (GMP) are used to extract global features in the spatial dimension. Spatial attention weights are learned through 3×3 convolution operations to highlight salient haze distribution areas. For channel information, GAP and GMP are used to extract global features in the channel dimension. Channel weights are learned through 1×1 convolutions to enhance the modeling of different semantic features. Finally, a channel shuffle operation is used to achieve feature interaction between spatial and channel attention. The input feature map is... The feature map has C, H, and W channels, and its height and length are respectively. The generation process of channel attention first involves processing the input feature map. Perform global average pooling to generate global features in the spatial dimension. Using 1×1 convolution pairs Perform feature compression and pass Activation function generates channel attention weights First, the input feature map Global average pooling and global max pooling are performed along the channel dimension to multiply the two features element-wise, and then spatial attention weights are generated through 3×3 convolution.

[0051] To further enhance information interaction between space and channels, this module introduces a channel shuffling operation, enabling the redistribution of information across channels and preventing the attention mechanism from focusing solely on local features while lacking global information integration. Channels are divided into multiple groups, and a shuffling operation is performed within each group. This channel shuffling operation achieves the rearrangement and interaction of information across channels. The weights after channel shuffling are processed through a 3×3 convolution to generate the final channel-specific spatial importance map. The importance map is then normalized using a sigmoid activation function to obtain the final attention map.

[0052] The network model contains a physical sensing module at a deeper level, such as Figure 4 As shown, combining prior information from the physical model can provide stronger constraints and interpretability for the defogging task. This module is based on the classic atmospheric scattering model:

[0053]

[0054] In the formula Images representing smog, Represents a clean image without fog. It represents the transmittance of the medium in the atmosphere. Representing atmospheric light parameters, the process of reconstructing the dehazing recovery results utilizes model learning. and These two key physical parameters ultimately lead to a more accurate haze-free image. .

[0055] The input to this module serves as the deepest layer of the network, and its input is the feature map from the previous layer. It contains two branches that derive estimates in reverse order for estimating global atmospheric light. and medium transmittance For estimating atmospheric light The branch first uses global average pooling. The overall brightness of the image is learned, and then through dimensionality-up convolution and dimensionality-down convolution operations, a feature map with more aggregated useful features is obtained, which is then utilized... The final estimated value is obtained. This is used to estimate the medium transmittance. The difference between this branch and the estimated atmospheric light branch is that it does not use global average pooling. The process involves processing the data. Finally, using the estimated physical parameters and the atmospheric scattering model formula, an approximately clear image is calculated. Through this process, the physical perception module combines physical models with deep learning, providing an effective solution for dehazing tasks.

[0056] Based on the above system, a method for image dehazing in low-altitude scenes is obtained. The dehazing system to be protected by this invention is a software system, and the data processing logic of the software corresponds to a specific method, including the following processes:

[0057] Capture and obtain images of low-altitude scenes;

[0058] The image is input into the image dehazing system for low-altitude scenes described above for target segmentation, including the design of frequency domain fusion modules, spatial and channel interaction modules, and physical perception modules in shallow, intermediate, and deep layers, respectively.

[0059] The frequency domain fusion module takes three high-frequency components of the feature map after wavelet transform as input. After being processed by the global attention mechanism of the frequency domain features, it outputs a global frequency domain similarity map, which is used to capture the similarity between pixels in the hazy degraded image and use the relatively clear light fog areas in the image to help restore the dense fog area image.

[0060] The input to the spatial and channel interaction module is the feature map after frequency domain fusion. First, spatial attention and channel attention are calculated and connected in the channel dimension of the input feature map. Then, random shuffling and group convolution operations are performed to reduce the uneven distribution of haze images in the channel level.

[0061] The physical perception module is designed based on an atmospheric scattering model. The input is the low-frequency component of the feature map after wavelet transform. The low-frequency component information is used as the key parameters of the model. The atmospheric diffusion model formula is used to reconstruct the dehazed and restored image in the feature space.

[0062] Based on the data processing flow of each module in the system, the target image after dehazing is output.

[0063] Table 1. Experimental results on two publicly available image dehazing datasets.

[0064]

[0065] Method: a method;

[0066] Pub. / Year: Publication of the conference and year;

[0067] Both SOTS-indoor and SOTS-outdoor are dataset names;

[0068] PSNR: Peak Signal-to-Noise Ratio;

[0069] SSIM: Structural Similarity Index;

[0070] Param: The size of the parameter;

[0071] As shown in Table 1, this embodiment tested the quantitative dehazing metrics of the model of the present invention on two publicly available image dehazing datasets. Compared with the previous best model, the model of the present invention achieved industry-leading results in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), especially with a significant improvement in PSNR.

[0072] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0073] While the specific embodiments of the present invention have been described above, they are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. An image dehazing system for low-altitude scenes, characterized in that: Based on the construction of a U-shaped network, and taking advantage of the different sensitivities of the U-shaped network to frequency domain information, different frequency components are guided to different layers. Frequency domain fusion module, spatial and channel interaction module and physical sensing module are designed in shallow, middle and deep layers respectively. The frequency domain fusion module takes as input three high-frequency components of the feature map processed by wavelet transform. After processing by a global attention mechanism of the frequency domain features, it outputs a global frequency domain similarity map, which is used to capture the similarity between pixels in the hazy degraded image and to use the relatively clear light fog areas in the image to assist in the recovery of the dense fog area image. The frequency domain fusion module is specifically as follows: The encoder part of the network performs frequency domain decomposition on the input features using discrete wavelet transform to obtain four frequency domain components. Three high-frequency components—the horizontal high-frequency subband LH, the vertical high-frequency subband HL, and the diagonal high-frequency subband HH—are selected as the input to this module. The frequency domain fusion module contains three branches. The first branch performs a concatenation operation on LH, HL, and HH along the channel dimension to increase dimensionality, and then performs a convolution operation to reduce dimensionality. The second branch uses the horizontal high-frequency texture information LH and the vertical high-frequency texture information HL for attention calculation. First, it reconstructs the feature dimensions of the horizontal high-frequency subband LH, the vertical high-frequency subband HL, and the diagonal high-frequency subband HH, and then performs a 1×1 convolution to obtain their corresponding channel-adjusted feature maps. , and Then adjust the feature map , After dot product operation Normalization yields an attention matrix of size C×C, where C is the number of channels in the input features; the third branch uses... Multiplying the result by the attention matrix obtained from the second branch yields a weighted feature result, which is then element-wise added to the residual obtained from the first branch to form a frequency domain fused feature representation. As shown below: The input to the spatial and channel interaction module is the feature map after frequency domain fusion. First, spatial attention and channel attention are calculated and concatenated along the channel dimension of the input feature map. Then, random shuffling and grouped convolution operations are performed to reduce the uneven distribution of haze images at the channel level. Specifically, the spatial and channel interaction module is as follows: It includes two branches that learn spatial and channel information respectively. The two parts are merged along the channel dimension and then randomly shuffled. For the spatial information branch, the feature map output from the previous layer is subjected to global average pooling. and global max pooling After extracting global features in the spatial dimension, a connection operation is performed. Then, spatial attention weights are learned through 3×3 convolution to highlight areas with significant haze distribution. For channel information branches, global average pooling is used. Then, the weights between channels are learned through 1×1 convolution, and then... Activation functions enhance the modeling of different semantic features; spatial information branches and channel information branches are added at the element level, and finally processed through the channels. shuffle The operation enables the interaction of features related to spatial and channel interests. The input to the spatial and channel information interaction module, i.e., the input feature map. The feature map has C, H, and W channels, height, and length, respectively. The first step is the generation of spatial attention: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Global average pooling and global max pooling are performed in the spatial dimension, and the two features are multiplied element-wise. Then, spatial attention weights are generated through convolution. Then comes the process of generating channel attention: first, the input feature map... Perform global average pooling to generate global features in the spatial dimension. Using convolution pairs Perform feature compression and pass Activation function generates channel attention weights Finally, the spatial information branch and the channel information branch are added at the element level, and then processed through the channel. shuffle The operation implements the interaction of spatial and channel-related features to obtain the output. ; The physical sensing module is designed based on an atmospheric diffusion model, and its input... Including the output of the space and channel interaction module The low-frequency component information extracted by wavelet transform is used for key parameter estimation of the model. The atmospheric diffusion model formula is then used to reconstruct the dehazed image in the feature space. The formula for the atmospheric diffusion model in the physical sensing module is: In the formula Images representing smog, Represents a clean image without fog. It represents the transmittance of the medium in the atmosphere. Representing atmospheric light parameters, the process of reconstructing the dehazing recovery results utilizes model learning. and These two key physical parameters ultimately lead to a more accurate haze-free image. ; The module's input, serving as the deepest layer of the network, has the following input: It contains two branches that derive estimates in reverse order for estimating global atmospheric light. and medium transmittance For estimating atmospheric light The branch first uses global average pooling. The overall brightness of the image is learned, and then through dimensionality-up convolution and dimensionality-down convolution operations, a feature map with more aggregated useful features is obtained, which is then utilized... The final estimate is obtained; for estimating the medium transmittance The branch first obtains feature representations through convolution, and then... The activation function fits the mapping relationship from feature representation to transmittance, and then through convolution and utilization... The estimated values ​​are obtained; finally, using the estimated physical parameters, an approximate clear image is calculated according to the atmospheric scattering model formula.

2. The image dehazing system for low-altitude scenes as described in claim 1, characterized in that: The frequency domain fusion module uses an improved self-attention module for feature extraction and enhances the focus on key frequency domain features; The spatial and channel interaction module integrates spatial and channel information from global average pooling and global max pooling in a multi-scale structure. The channel information needs to be randomly shuffled to promote cross-channel frequency domain information exchange, while the spatial level is responsible for highlighting areas with significant haze distribution. The physical sensing module learns physical parameters, including global atmospheric light, through a network. and medium transmittance Furthermore, by combining low-frequency component information, physical parameters are estimated, providing physical constraints for the defogging task.

3. The image dehazing system for low-altitude scenes as described in claim 1, characterized in that: The spatial and channel interaction module also introduces a channel shuffling operation, enabling the redistribution of information between channels. Channels are divided into multiple groups, and a shuffling operation is performed within each group. Through this channel shuffling operation, cross-channel information rearrangement and interaction are achieved. The weights after channel shuffling are then processed using a 3×3 convolution to generate the final channel-specific spatial importance map. The activation function normalizes the importance graph to obtain the final attention graph.

4. A method for image dehazing in low-altitude scenes, characterized in that, Includes the following processes: Capture and obtain images of low-altitude scenes; The image is input into the low-altitude scene image dehazing system as described in any one of claims 1 to 3 for target segmentation; Output the target image after dehazing.

Citation Information

Patent Citations

  • Unsupervised night image defogging method using high and low frequency decomposition

    CN113191964A

  • Method for constructing remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement

    CN120355601A