Frequency domain-space dual branch collaborative image defogging method based on high-low frequency decoupling

By constructing a frequency-space dual-branch collaborative dehazing network and utilizing high-low frequency decoupling and CSP hybrid attention mechanism, the problems of incomplete dehazing and detail loss in foggy images in existing technologies are solved, achieving high-precision and stable foggy image restoration effects, which are suitable for scenarios such as drone aerial photography and security monitoring.

CN122492504APending Publication Date: 2026-07-31XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN UNIV OF POSTS & TELECOMM
Filing Date
2026-05-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing dehazing algorithms for foggy images are incomplete in large-scale non-uniform fog scenes, easily lose fine-grained details, have weak generalization ability in complex scenes, and rely on manual prior parameters or limited receptive fields of local convolution, making it difficult to balance global fog degradation and detail fidelity.

Method used

A frequency-space dual-branch collaborative dehazing network based on high- and low-frequency decoupling is constructed. The frequency domain branch performs global fog degradation suppression and high- and low-frequency decoupling processing, while the spatial domain branch performs multi-scale feature filtering and detail enhancement. The CSP hybrid attention mechanism is introduced to combine discrete wavelet transform and fast Fourier convolution for feature fusion and reconstruction.

Benefits of technology

It achieves high-precision dehazing in complex fog scenes, preserves image detail information, and improves the stability and adaptability of the dehazing effect, making it suitable for high-quality image restoration in fields such as drone aerial photography and security monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492504A_ABST
    Figure CN122492504A_ABST
Patent Text Reader

Abstract

This invention presents a frequency-space dual-branch collaborative image dehazing method based on high- and low-frequency decoupling, addressing the problem of incomplete removal of non-uniform fog in traditional single-domain dehazing algorithms. The method includes the construction of a frequency-space dual-branch collaborative dehazing network, spatial domain dehazing enhancement, frequency domain global fog degradation suppression, and dual-domain feature fusion and global residual image reconstruction. The invention uses FFA-Net as a baseline to construct the frequency-space dual-branch collaborative dehazing network. The spatial domain branch achieves refined enhancement and filtering of fog features through a multi-scale structure embedding CSP hybrid attention; the frequency domain branch employs an encoding / decoding architecture combining discrete wavelet transform and fast Fourier convolution to achieve low-frequency fog degradation suppression and high-frequency detail preservation. After multi-scale fusion of the dual-branch features, combined with global residual reconstruction, a clear image is output. This invention offers advantages such as high dehazing accuracy and good detail fidelity, making it suitable for foggy image restoration scenarios such as UAV aerial photography and security monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology and mainly relates to dehazing and enhancement of foggy images in image restoration. Specifically, it is a frequency-space dual-branch collaborative image dehazing method based on high and low frequency decoupling. It is suitable for dehazing and image quality restoration of single images in complex foggy scenes and can provide high-quality image input for subsequent image analysis and target recognition in fields such as UAV aerial photography, security monitoring, vehicle assisted driving, and remote sensing mapping. Background Technology

[0002] Image dehazing is a core digital image processing technology that repairs image degradation caused by atmospheric scattering and restores clear images. It restores the edge texture and true color of images by suppressing problems such as decreased image contrast, blurred details, and color shift caused by fog. It is a key pre-processing technology in the field of machine vision.

[0003] Current dehazing algorithms for foggy images are mainly divided into two categories: prior-based methods based on physical models and end-to-end dehazing methods based on deep learning. Among them, traditional prior-based methods based on dark channel priors rely heavily on manually set prior assumptions. When the prior conditions are not met, color cast and artifacts are prone to occur, and their generalization ability for complex scenes such as non-uniform fog and dense fog is extremely poor. Traditional deep learning dehazing algorithms mostly perform feature processing only in the spatial domain. Constrained by the receptive field of local convolution, they do not consider the decoupling processing of global fog degradation characteristics and high-frequency detail information, making it difficult to simultaneously ensure the thoroughness of dehazing in large-scale non-uniform fog scenes and the integrity of fine-grained image details.

[0004] To address the limitations of single-domain dehazing algorithms, existing technologies have proposed a spatial-frequency domain dual-branch collaborative dehazing architecture. This architecture attempts to achieve global fog degradation modeling through the frequency domain branch and refine detail enhancement through the spatial domain branch. However, such algorithms still have significant drawbacks: First, they focus on the basic structure stacking of the frequency domain branch, achieving dual-domain fusion only through simple frequency domain transformation and feature concatenation, with insufficient design for decoupling high- and low-frequency information and collaborative fusion of dual-domain features. Second, they lack a targeted attention mechanism, resulting in poor refinement of areas with uneven fog distribution. The dehazing effect is easily affected by wavelet basis selection, image fog density, and resolution variations, leading to large fluctuations in detail fidelity.

[0005] In the existing technology, the paper "FFA-Net: Feature Fusion Attention Network for Single Image Dehazing" published by Qin et al. in 2019 uses the fusion of channel attention and pixel attention as the core of the network to improve the extraction ability of fog-related features. However, the FFA-Net network only completes the processing in the spatial domain and has insufficient global modeling ability. It cannot achieve complete dehazing in dense fog areas and is difficult to adapt to large-scale complex fog scenes such as aerial photography and long-distance monitoring.

[0006] The article "A Single Image Dehazing Method Based on Dual-Domain Feature Interaction and Local Correlation Upsampling" published by Liu Mingjie et al. in 2025 proposed a dual-domain interactive upsampling dehazing network that achieves collaborative extraction of spatial-frequency domain features and semantic optimization of the decoding layer. However, it did not design a refined hybrid attention mechanism for decoupling high and low frequencies, and the dual-domain feature hierarchical fusion strategy has the defects of single-stage static fusion and lack of cross-scale collaborative interaction, resulting in poor dehazing generalization stability in engineering scenarios.

[0007] Overall, existing dehazing methods for foggy images are either limited to optimizing convolutional structures and attention mechanisms in the spatial domain, or have significant shortcomings in the design of dual-domain collaboration. Specifically, existing methods mostly use static feature concatenation or simple linear weighting, failing to establish deep semantic connections between global features in the frequency domain and local details in the spatial domain. They also fail to achieve efficient decoupling of low-frequency fog degradation information and high-frequency detail information, making it difficult to balance global modeling and fine-grained detail preservation in large-scale complex fog scenes. This results in unstable dehazing effects and weak scene generalization ability. Summary of the Invention

[0008] The purpose of this invention is to overcome the above-mentioned defects of the prior art and provide an end-to-end frequency-space dual-branch collaborative image dehazing method based on high and low frequency decoupling. This method achieves high dehazing accuracy, good detail fidelity, and strong scene generalization ability for single foggy images, while also taking into account inference real-time performance, and has high engineering application value.

[0009] This invention is a frequency-space dual-branch collaborative image dehazing method based on high-low frequency decoupling, characterized by comprising the following steps:

[0010] Step 1: Construct a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP: Using the feature fusion attention network FFA-Net for image dehazing as the baseline architecture, and relying on the inherent global residual mechanism of FFA-Net, whose global residual input is the foggy image to be processed and whose output is the network terminal, a parallel spatial domain dehazing main branch and frequency domain feature enhancement branch architecture are constructed. The input of this architecture is connected to a shared initial feature extraction convolutional layer, and the output is connected to the feature fusion layer at the end of the network, together forming a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP; The spatial domain dehazing main branch has a cascaded residual G module with multiple sets of embedded CSP hybrid attention as the backbone structure, which sequentially includes and connects a shared initial feature extraction convolutional layer, multiple sets of cascaded residual G modules, a multi-scale cross-layer connection module, a post-CSP hybrid attention module, and a terminal dimension matching convolutional layer; The residual G modules are all based on basic residual blocks embedded with CSP hybrid attention as basic units, and the cascaded multiple sets of residual G modules and the multi-scale cross-layer connection module are sequentially connected. Cross-layer connections; the frequency domain feature enhancement branch adopts a symmetrical U-shaped encoding and decoding architecture that combines discrete wavelet transform and fast Fourier convolution. The overall structure is a three-level symmetrical encoding-decoding structure, and the backbone is constructed according to the logic of encoding-bottleneck feature optimization-decoding: the front end is set with three sets of encoding downsampling modules embedded with discrete wavelet transform, the middle bottleneck layer is embedded with fast Fourier convolution residual blocks for global frequency domain feature optimization, the front end first undergoes initial feature extraction processing through shared initial feature extraction convolutional layers, and then three sets of decoding upsampling modules corresponding one-to-one with the encoding level are set; at the same time, multi-scale cross-layer jump connections are introduced between the downsampling module and the upsampling module, and the high and low frequency features obtained by discrete wavelet transform decomposition at each level of the encoding end are fed into the corresponding level of the decoding upsampling module layer by layer to realize multi-scale frequency domain detail compensation and feature reuse; the spatial domain dehazing main branch and the frequency domain feature enhancement branch constitute a dual-branch collaborative body, and a global residual branch is introduced for image identity mapping constraints, forming an end-to-end high and low frequency decoupling and CSP combination frequency domain-spatial dual-branch collaborative dehazing network;

[0011] Step 2: Spatial Domain Main Branch Dehazing Enhancement Processing: After the foggy image is input into the spatial domain dehazing main branch, shallow initial feature extraction is first completed through the shared initial feature extraction convolutional layer at the front end of the dehazing network. Then, it is fed into multiple sets of residual G modules with embedded CSP hybrid attention to achieve hierarchical feature extraction and preliminary filtering of fog degradation features. Relying on the CSP hybrid attention module, channel dimension adaptive weight calibration, spatial fog region feature orientation enhancement, and pixel-level fog interference fine suppression are completed in sequence. While hierarchically representing the fog degradation information of the image, the focus is on modeling the fog distribution area and outputting a spatial domain fog residual feature map with strong representation capabilities.

[0012] Step 3: Global Fog Degradation Suppression Processing via Frequency Domain Feature Enhancement Branch: Global fog degradation suppression is achieved through the frequency domain feature enhancement branch. First, the foggy image is processed through a shared initial feature extraction convolutional layer at the front end of the defogging network to complete shallow feature extraction, and then fed into a three-level symmetric encoder-decoder structure. Each encoding level decomposes the features layer by layer into low-frequency subbands carrying global fog degradation information and high-frequency subbands containing edge textures through discrete wavelet transform. The first two levels of high-frequency and low-frequency subbands are directly fed into the corresponding decoding levels through multi-scale cross-layer skip connections to avoid the loss of fine-grained details during encoding and decoding. The low-frequency subbands converge to the intermediate bottleneck layer, and global modeling and degradation interference suppression of a large-scale non-uniform fog distribution are completed through fast Fourier convolution residual blocks. After decoding, reconstruction, and fusion by the frequency domain feature enhancement branch, the output is a frequency domain enhanced feature with global frequency domain constraints and high-frequency detail compensation capabilities.

[0013] Step 4: Dual-domain feature fusion and global residual image reconstruction: The fog residual features output from the spatial domain dehazing main branch and the frequency domain detail enhancement features output from the frequency domain feature enhancement branch are fused at the same scale. The fused features are then reconstructed with the original foggy image, which serves as the global residual branch, to fully leverage the synergistic advantages of multi-scale fine dehazing in the spatial domain and high- and low-frequency detail reuse in the frequency domain. This maximizes the preservation of the image's edge texture and structural information, ultimately outputting a clear dehazed image. This achieves end-to-end image dehazing based on high- and low-frequency decoupling and frequency domain-spatial dual-branch collaborative technology.

[0014] This invention constructs a dual-branch collaborative dehazing network based on FFA-Net: the spatial domain branch achieves refined enhancement and filtering of fog features at the channel, spatial, and pixel levels through a multi-scale structure embedding CSP hybrid attention; the frequency domain branch employs an encoding / decoding architecture of discrete wavelet transform and fast Fourier convolution to decouple low-frequency fog degradation information from high-frequency detail information, achieving global fog degradation suppression and preserving edge texture through skip connections. After multi-scale fusion of the dual-branch features, combined with global residual reconstruction, a clear image is output. This invention offers high dehazing accuracy, good detail fidelity, strong adaptability to all scenarios, requires no manual prior parameters, and facilitates convenient end-to-end inference. It can be widely applied to foggy image restoration in fields such as drone aerial photography and security monitoring.

[0015] This invention addresses the technical problems of incomplete dehazing of large-scale non-uniform fog, easy loss of fine-grained details, and weak generalization ability in complex scenes in traditional single-domain image dehazing algorithms. It combines frequency-domain global fog degradation modeling technology with spatial-domain refined dehazing enhancement technology. Discrete wavelet transform is used to decouple low-frequency fog degradation information from high-frequency detail information. A CSP hybrid attention mechanism is used to achieve targeted enhancement of fog-related features. A multi-scale hierarchical fusion strategy is employed to achieve collaborative optimization of dual-domain features. Finally, high-fidelity clear image restoration is achieved through global residual reconstruction. This invention improves the thoroughness and detail fidelity of dehazing in complex fog scenes, adapting to the fog image restoration needs of various scenarios such as drone aerial photography, security monitoring, and vehicle vision. It also balances dehazing accuracy and inference real-time performance, making it valuable for engineering applications.

[0016] Beneficial effects of the invention

[0017] This invention combines frequency domain global fog degradation modeling technology with spatial domain refined dehazing enhancement technology, solving the technical problems of traditional single-domain image dehazing algorithms, such as incomplete dehazing of large-scale non-uniform fog, easy loss of fine-grained details, and weak generalization ability in complex scenes. Compared with existing technologies, it has the following core advantages:

[0018] (1) Excellent defogging performance and strong adaptability to complex scenes

[0019] This invention constructs a frequency-spatial domain dual-branch collaborative dehazing architecture, breaking through the technical bottleneck of limited receptive field in traditional spatial domain convolution. It achieves global context modeling of large-scale non-uniform fog and continuous dense fog through fast Fourier convolution residual blocks, and achieves pixel-level fine-grained dehazing processing by combining CSP channel-space-pixel three-level hybrid attention. In the RESIDE standard dehazing dataset and real aerial photography scenarios, the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are significantly better than existing mainstream dehazing algorithms. It has a strong adaptability to all scenarios such as light fog, medium fog, dense fog, and non-uniformly distributed fog.

[0020] (2) High fidelity in detail, completely solving the problem of lost details.

[0021] This invention achieves decoupling of low-frequency fog degradation information and high-frequency detail information through discrete wavelet transform. High-frequency detail features, including image edges and textures, are directly transferred to the reconstruction stage through multi-scale cross-layer skip connections, completely preserving fine-grained details of the image. This completely solves the pain points of edge smoothing and texture loss in the dehazing process of traditional dehazing algorithms. The dehazed image is free of artifacts and color shifts, with strong visual realism, and can provide high-quality input for subsequent high-level vision tasks such as object detection and feature extraction.

[0022] (3) End-to-end deployment is user-friendly and highly adaptable to engineering implementation.

[0023] The FDF-Net of this invention is a fully convolutional end-to-end architecture. It does not require manual estimation of prior parameters such as transmittance and global atmospheric light, and can directly adapt to input images of any size, outputting clear images after dehazing end-to-end. At the same time, through lightweight structure optimization, the number of model parameters and computational load are controllable, and it can still meet the real-time processing requirements on embedded edge computing platforms. It can be seamlessly adapted to various engineering deployment scenarios such as UAV airborne inspection, security monitoring, and vehicle-mounted assisted driving. Attached Figure Description

[0024] Figure 1 This is a flowchart of the present invention;

[0025] Figure 2 This is a flowchart illustrating the overall architecture of the defogging network in this invention.

[0026] Figure 3 This is a schematic diagram of the overall structure of the defogging network of the present invention;

[0027] Figure 4 This is a structural diagram of module G of the present invention;

[0028] Figure 5 This is a structural diagram of the basic residual block of the present invention;

[0029] Figure 6 This is a structural diagram of the CSP hybrid attention module of the present invention;

[0030] Figure 7 This is a structural diagram of the frequency domain branch sampling module of the present invention;

[0031] Figure 8 This is a structural diagram of the FFC (Fast Fourier Convolution) unit of the present invention;

[0032] Figure 9 This is a structural diagram of the FFC residual block of the present invention;

[0033] Figure 10 This is a comparison of the dehazing performance of the FDF-Net of this invention on the RESIDE indoor dataset;

[0034] Figure 11 This is a comparison of the defogging effect of the FDF-Net of this invention on the RESIDE outdoor dataset.

[0035] Detailed Description of Embodiments: To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings.

[0036] Example 1: Image dehazing, as a core pre-processing technology in digital image processing and machine vision, is a key means to improve the accuracy of high-level visual tasks such as target detection and image recognition in foggy scenes. For many years, performance optimization of single-image dehazing algorithms has been a research hotspot in computer vision. The performance of a dehazing algorithm directly determines the effectiveness of subsequent visual tasks. Higher dehazing accuracy and better detail fidelity provide high-quality image input for machine vision applications in complex scenes. The restoration effect of traditional image dehazing algorithms is greatly affected by external factors such as scene fog distribution characteristics and lighting conditions, resulting in limitations on dehazing accuracy, scene generalization ability, and detail fidelity. Dual-domain collaborative dehazing architecture is an effective way to solve the performance limitations of traditional single-domain dehazing algorithms. While some existing dual-domain dehazing algorithms have advantages such as strong global modeling capabilities and a wide dehazing range, they still suffer from drawbacks such as unstable dehazing effects, easy loss of details, and insufficient real-time performance for engineering applications. These shortcomings of existing technologies limit the improvement of restoration accuracy during image dehazing, leading to unstable dehazing effects in complex scenes.

[0037] This invention addresses the aforementioned issues by researching and exploring a frequency-domain-spatial dual-branch collaborative image dehazing method based on high- and low-frequency decoupling. In complex foggy visual imaging scenarios, this invention constructs a frequency-domain-spatial dual-branch collaborative end-to-end dehazing network model. The frequency domain branch performs high- and low-frequency decoupling and global degradation suppression on foggy images, while the spatial domain branch performs multi-scale feature filtering and detail enhancement. Combining discrete wavelet transform, fast Fourier convolution, and a three-level hybrid attention mechanism (CSP), the dual-branch features are fused at the same scale and global residuals are reconstructed, achieving image dehazing that balances global uniformity and local fidelity.

[0038] This invention is a frequency-domain-spatial dual-branch collaborative image dehazing method based on high- and low-frequency decoupling. In complex foggy visual imaging scenarios, this invention constructs a frequency-domain-spatial dual-branch collaborative end-to-end dehazing network model. Using FFA-Net as a baseline, the spatial domain branch extracts spatial features at different levels step by step through three sets of cascaded residual G modules combined with multi-hop connections and CSP three-level hybrid attention to achieve multi-scale fog feature filtering and detail enhancement. The frequency domain branch, based on the high- and low-frequency decoupling logic of discrete wavelet transform, performs global degradation suppression and detail enhancement on the foggy image through a multi-scale downsampling-upsampling structure and FFC residual blocks. The output features of the two branches are fused by addition at the same scale and then form a global residual connection with the original foggy image to achieve dehazing image reconstruction that balances global dehazing uniformity and local texture fidelity. See [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart of the present invention, which includes the following steps:

[0039] Step 1: Construct a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP: Using the feature fusion attention network FFA-Net for image dehazing as the baseline architecture, and relying on the inherent global residual mechanism of FFA-Net, the foggy image to be processed is used as the global residual input branch. The global residual input is the foggy image to be processed, and the output is the end of the network. A parallel spatial domain dehazing main branch and frequency domain feature enhancement branch architecture are constructed. The input of this architecture is connected to a shared initial feature extraction convolutional layer, and the output is connected to the feature fusion layer at the end of the network. Together, they constitute a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP, referred to as the dehazing network. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the overall structure of the dehazing network of the present invention. The spatial domain dehazing main branch uses cascaded residual G modules with embedded CSP hybrid attention as the backbone structure. The residual G module, also called the residual group module, sequentially includes and connects a shared initial feature extraction convolutional layer, cascaded residual G modules, a multi-scale cross-layer connection module, a post-CSP hybrid attention module, and a terminal dimension matching convolutional layer. Each residual G module uses a basic residual block with embedded CSP hybrid attention as the basic unit. The cascaded residual G modules and the multi-scale cross-layer connection module are sequentially connected across layers, and the cross-layer connection realizes the cross-layer fusion and information reuse of features at different levels. The frequency domain feature enhancement branch adopts a symmetrical U-shaped encoding and decoding architecture that combines discrete wavelet transform and fast Fourier convolution. The overall structure is a three-level symmetrical encoding-decoding structure, and the backbone is constructed according to the logic of encoding-bottleneck feature optimization-decoding. The architecture consists of three layers: the front end first performs initial feature extraction through a shared initial feature extraction convolutional layer, followed by three sets of encoding downsampling modules embedded with discrete wavelet transform, and the middle bottleneck layer embeds fast Fourier transform residual blocks for global frequency domain feature optimization. The back end is configured with three sets of decoding upsampling modules that correspond one-to-one with the encoding level. Simultaneously, multi-scale cross-layer skip connections are introduced between the downsampling and upsampling modules. The high and low frequency features obtained by discrete wavelet transform decomposition at each encoding level are fed into the corresponding decoding upsampling module layer by layer to achieve multi-scale frequency domain detail compensation and feature reuse, thereby enhancing the ability to suppress global fog degradation and restore high-frequency textures. The spatial domain defogging main branch and the frequency domain feature enhancement branch constitute a dual-branch collaborative main body. A global residual branch is introduced to perform image identity mapping constraints, forming an end-to-end high and low frequency decoupling and CSP-combined frequency domain-spatial dual-branch collaborative defogging network.

[0040] Step 2: Spatial Domain Main Branch Dehazing Enhancement Processing: After the foggy image is input into the spatial domain dehazing main branch, shallow initial feature extraction is first completed through the shared initial feature extraction convolutional layer at the front end of the dehazing network. Then, it is fed into multiple sets of residual G modules with embedded CSP hybrid attention to achieve hierarchical feature extraction and preliminary filtering of fog degradation features. Relying on the CSP hybrid attention module, channel dimension adaptive weight calibration, spatial fog region feature orientation enhancement, and pixel-level fog interference fine suppression are completed in sequence. While hierarchically representing the fog degradation information of the image, the focus is on modeling the fog distribution area and outputting a spatial domain fog residual feature map with strong representation capabilities.

[0041] Step 3: Global Fog Degradation Suppression Processing via Frequency Domain Feature Enhancement Branch: Global fog degradation suppression is achieved through the frequency domain feature enhancement branch. First, the foggy image is processed through a shared initial feature extraction convolutional layer at the front end of the defogging network to complete shallow feature extraction, and then fed into a three-level symmetric encoder-decoder structure. Each encoding level decomposes the features layer by layer into low-frequency subbands carrying global fog degradation information and high-frequency subbands containing edge textures through discrete wavelet transform. The first two levels of high-frequency and low-frequency subbands are directly fed into the corresponding decoding levels through multi-scale cross-layer skip connections to avoid the loss of fine-grained details during encoding and decoding. The low-frequency subbands converge to the intermediate bottleneck layer, and global modeling and degradation interference suppression of a large-scale non-uniform fog distribution are completed through fast Fourier convolution residual blocks. After decoding, reconstruction, and fusion by the frequency domain feature enhancement branch, the output is a frequency domain enhanced feature with global frequency domain constraints and high-frequency detail compensation capabilities.

[0042] In this invention, step 2, the spatial domain main branch defogging enhancement processing, and step 3, the frequency domain feature enhancement branch global fog degradation suppression processing, are performed simultaneously.

[0043] Step 4: Dual-domain feature fusion and global residual image reconstruction: The fog residual features output from the spatial domain dehazing main branch and the frequency domain detail enhancement features output from the frequency domain feature enhancement branch are fused at the same scale. The fused features are then reconstructed with the original foggy image, which serves as the global residual branch, to fully leverage the synergistic advantages of multi-scale fine dehazing in the spatial domain and high- and low-frequency detail reuse in the frequency domain. This maximizes the preservation of the image's edge texture and structural information, ultimately outputting a clear dehazed image. This achieves end-to-end image dehazing based on high- and low-frequency decoupling and frequency domain-spatial dual-branch collaborative technology.

[0044] This invention addresses complex foggy visual imaging scenarios by constructing a frequency-domain-spatial dual-branch collaborative end-to-end dehazing network model. Using FFA-Net as a baseline, the spatial domain branch extracts spatial features at different levels through three sets of cascaded residual G modules combined with multi-hop connections and CSP three-level hybrid attention to achieve multi-scale fog feature filtering and detail enhancement. The frequency domain branch, based on the high-low frequency decoupling logic of discrete wavelet transform, performs global degradation suppression and detail enhancement on foggy images through a multi-scale downsampling-upsampling structure and FFC residual blocks. The output features of the two branches are fused by addition at the same scale and then form a global residual connection with the original foggy image to achieve dehazing image reconstruction that balances global dehazing uniformity and local texture fidelity.

[0045] This invention is a holistic technical solution, in which steps 1 to 4 are complementary, closely coordinated, and indispensable. By constructing a parallel frequency-space dual-branch architecture, this invention deeply integrates multi-scale local feature extraction in the spatial domain with high- and low-frequency decoupling and global degradation suppression in the frequency domain. This holistic design effectively overcomes the technical shortcomings of existing technologies, such as insufficient dual-domain coordination, easy loss of high-frequency details, and weak generalization ability in complex scenes, achieving high-quality image restoration that balances global dehazing uniformity and local detail fidelity.

[0046] Example 2: Frequency-Spatial Dual-Branch Collaborative Image Dehazing Method Based on High-Low Frequency Decoupling. Similar to Example 1, the cascaded multiple sets of embedded CSP hybrid attention residual G modules in the spatial domain dehazing main branch described in step 1 serve as the core unit for spatial domain feature extraction and local dehazing. Specifically designed for the characteristics of weak target features and complex background noise in foggy images, it consists of three sets of serially arranged multi-scale feature extraction residual G modules cascaded together. All residual G modules adopt a dual-path parallel structure, referencing... Figure 4 , Figure 4 This is a structural diagram of the G module of the present invention. One path is the main branch, consisting of three cascaded basic residual blocks connected in series with convolutional layers. This branch is responsible for extracting multi-scale deep semantic features and performing local dehazing. The other path is a direct-connection residual branch, constructing a gradient highway to effectively alleviate the gradient vanishing problem in deep networks. The input features of the residual G module are directly passed to the end of the main branch and fused with the output features of the main branch. This achieves complementary fusion of original detail features and deep semantic features, maximizing the preservation of small target detail information in foggy images. The basic residual block is a CSP hybrid attention module with two-level residual connections, as shown in the reference diagram. Figure 5 , Figure 5This is the basic residual block structure diagram of the present invention. Its first-level residual structure is composed of a convolutional layer and a ReLU activation function connected in series and combined with module input features. It is responsible for extracting basic features such as bottom edge and texture. At the same time, the residual connection avoids the loss of shallow features. The second-level residual structure is composed of a convolutional layer and a CSP hybrid attention module connected in series and combined with input features. Based on the CSP cross-stage local network architecture, it integrates coordinate attention, spatial attention and channel attention to achieve adaptive enhancement of the target region and accurate suppression of fog background. The two-level residual structures are spliced ​​in series to form the basic residual block. Through the progressive feature extraction of the two-level residuals, the network training stability is maintained while enhancing the feature expression ability.

[0047] Based on the above structure, combined with the appendix Figure 4 As shown, the three multi-scale feature extraction modules G-1, G-2, and G-3 are cascaded in order from shallow to deep. The G-1 module is used to further extract features from the output of the shared initial feature extraction convolutional layer at the front end of the dehazing network architecture, mainly to obtain low-level information of the image, including fine-grained features such as edge contours, local textures, and brightness changes. The G-2 module performs further feature transformation based on the output of G-1, expanding the network's receptive field to extract mid-scale structural information, thereby enhancing the ability to express fog transition regions. The G-3 module is used to extract high-level semantic features, focusing on restoring the overall structural information of the image and the contour information of distant targets to improve the global consistency of the dehazing results.

[0048] Combined with appendix Figure 5 As shown, the basic residual block in this invention is a CSP hybrid attention module with two levels of residual connections. Its internal structure consists of two levels of residual enhancement units. The first-level residual structure is composed of a convolutional layer and a ReLU activation function connected in series, which is used to perform basic nonlinear transformation on the input features and enhance local response capabilities, thereby strengthening the edge and texture information in the image. The second-level residual structure is composed of a convolutional layer and a CSP hybrid attention module connected in series. It introduces an attention mechanism during the feature extraction process. By jointly modeling the channel dimension and the spatial dimension, it achieves adaptive weighting of key feature regions, thereby highlighting regions that play an important role in dehazing and suppressing redundant information. The first-level residual structure and the second-level residual structure are connected in series and together form the B module, so that feature enhancement and attention modeling can work together.

[0049] This architecture employs a dual-path residual parallel design for G modules and a two-level residual nested architecture for basic residual blocks. It deeply embeds CSP hybrid attention into the basic feature extraction unit, overcoming the shortcomings of traditional networks where feature extraction and dehazing enhancement are isolated. This effectively strengthens multi-scale fog feature extraction capabilities, balancing global fog modeling with local detail preservation. Multi-level residual connections alleviate the gradient vanishing problem in deep networks, improving model training stability and scene generalization ability, and reducing defects such as edge blurring, detail loss, and color distortion during dehazing. In practical applications, it is necessary to ensure uniform feature dimensions and resolution across all modules, strictly adhere to the serialization logic of attention modules, and rationally configure the number of stacked modules based on the complexity of the fog scene and deployment requirements to ensure overall network stability and optimal dehazing results.

[0050] Example 3: Frequency-Spatial Dual-Branch Collaborative Image Dehazing Method Based on High-Low Frequency Decoupling. Similar to Examples 1-2, the frequency domain feature enhancement branch described in step 1 serves as the core link for global fog degradation suppression and high-frequency detail enhancement. It is specifically designed to address the characteristics of thin fog distribution globally and weak high-frequency features of small targets in foggy images. The front-end three sets of encoding downsampling modules are divided into one initial downsampling module and two activation downsampling modules, referencing... Figure 7 , Figure 7This is a structural diagram of the frequency domain branch sampling module of the present invention. The initial downsampling module adopts a dual-branch structure. The main branch consists of a convolutional layer and a batch normalization layer connected in series. The auxiliary branch embeds a discrete wavelet transform layer and a convolutional layer. Utilizing the orthogonality of the discrete wavelet transform, frequency domain decomposition without information loss is achieved. The input features are decomposed into low-frequency sub-bands and high-frequency sub-bands by the discrete wavelet transform layer. The low-frequency sub-band mainly contains the overall brightness, large-scale structure, and global fog information of the image, while the high-frequency sub-band contains the details of the image's edges, corners, textures, and small targets. After feature mapping by the convolutional layer, the low-frequency sub-band is fused with the features of the main branch. The high-frequency sub-band is then transmitted to the corresponding upsampling module at the decoding end through a cross-layer skip connection. The structures of the two sets of activation downsampling modules are basically the same as those of the initial downsampling module. The difference lies in the addition of a Leaky ReLU activation function at the end of the main branch. This differentiated activation function design is a precise optimization for feature distributions at different depths: the initial downsampling module omits the activation function, fully preserving the weak small target information in the original image; the activation downsampling module introduces Leaky ReLU. The ReLU activation function addresses the gradient vanishing problem in deep networks, enhancing the non-linear expressive power of features. The three backend decoding upsampling modules correspond one-to-one with the encoding downsampling modules, both employing a dual-input branch structure to receive low-frequency and high-frequency features transmitted from the encoding end, respectively. The low-frequency input branch consists of a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in series, while the high-frequency input branch uses only a single convolutional layer to complete feature mapping. The features from both branches are fused to output the upsampling result of the current layer. Simultaneously, a fast Fourier convolutional residual block is embedded in the intermediate bottleneck layer to complete global frequency domain feature optimization. Multi-level skip connections enable the reuse of high and low-frequency features and texture detail compensation.

[0051] Based on the above structure, combined with the appendix Figure 7As shown, the frequency domain feature enhancement branch presents an overall "encoding-decoding" hierarchical structure. Three downsampling modules sequentially perform scale compression and frequency decomposition on the input features to progressively extract low-frequency structural information and high-frequency detail information from the image. In the downsampling module, discrete wavelet transform is introduced to perform frequency domain decomposition on the input features, explicitly dividing them into low-frequency and high-frequency sub-bands before they enter the deep network, thus providing a foundation for subsequent high-low frequency decoupling processing. The low-frequency sub-band mainly contains the overall contour and brightness distribution information of the image, which is fused with the main branch features after convolution processing to enhance structural information representation. The high-frequency sub-band mainly contains edge and texture detail information, which is transmitted across layers to the corresponding upsampling module, thus avoiding the loss of high-frequency information during multiple downsampling processes. The subsequent two downsampling modules with Leaky ReLU activation functions continue the dual-branch structure of the initial module. By introducing the Leaky ReLU activation function into the main branch, the network gains stronger nonlinear expressive power during feature transformation, while mitigating the "neuron death" problem that may be caused by traditional ReLU, thereby improving the model's robustness in complex foggy scenarios. Each downsampling module progressively reduces spatial resolution while continuously strengthening the abstract representation of low-frequency information, and effectively decoupling and reusing frequency domain information by preserving the cross-layer transmission path of high-frequency information.

[0052] Corresponding to the downsampling process, three upsampling modules are used to progressively recover low-resolution features. Each module employs a dual-input branch structure: the low-frequency input branch recovers and reconstructs low-frequency features from the next level, while the high-frequency input branch receives high-frequency sub-band information from the corresponding downsampling stage. The low-frequency input branch progressively recovers and enhances features through a concatenation of convolutional layers, batch normalization layers, and ReLU activation function layers; the high-frequency input branch only performs feature mapping on the high-frequency sub-bands through convolutional layers to maintain the integrity of its detailed information. Finally, the outputs of the low-frequency and high-frequency branches are fused to achieve the collaborative reconstruction of structural and detailed information.

[0053] The core innovation and improvement of the frequency domain branch in this invention lies in the construction of a three-level differentiated downsampling-symmetric upsampling U-shaped encoding and decoding architecture based on DWT discrete wavelet transform. This architecture achieves differentiated processing of global fog information and detailed textures through layer-by-layer high-low frequency decoupling. Combined with a fusion mechanism that directly connects high-frequency features across layers to corresponding upsampling, and differentiated activation function configuration, it solves the pain points of traditional frequency domain dehazing, such as high-low frequency coupling processing, easy loss of details, and unstable gradients in deep networks. Compared with existing solutions, its advantages include stronger global fog degradation modeling capabilities, significantly improved dehazing uniformity and detail fidelity, better generalization to complex fog scenes, and efficient complementarity with the spatial domain main branch. During implementation, it is crucial to strictly ensure a one-to-one correspondence between upsampling and downsampling levels. The initial downsampling module should disable Leaky ReLU to preserve the integrity of the original features. Reasonable configuration of activation function parameters and wavelet basis selection is essential to ensure matching of feature dimensions in high-low frequency fusion and dual-branch outputs.

[0054] Example 4: Frequency-Spatial Dual-Branch Collaborative Image Dehazing Method Based on High-Low Frequency Decoupling. Similar to Examples 1-3, the CSP hybrid attention module in the spatial domain main branch dehazing enhancement process described in step 2 serves as the core execution unit for local spatial domain dehazing. It integrates channel attention, spatial attention, and pixel-level attention to construct a three-level progressive feature enhancement system from coarse to fine. This system is specifically designed to address the pain points of low signal-to-noise ratio of target features and uneven fog distribution in foggy images. (Refer to...) Figure 6 , Figure 6 This is a structural diagram of the CSP hybrid attention module of the present invention. The processing flow is divided into three parts in sequence: channel dimension adaptive weighting, spatial region feature orientation enhancement, and pixel-level fog interference fine-tuning suppression.

[0055] 2.1 Channel Dimension Adaptive Weighting: As the first-level coarse-grained feature selection mechanism, it distinguishes effective target features from invalid fog noise features from the channel dimension. It uses global average pooling to extract global spatial information of each channel from the module input feature map. It utilizes the natural noise resistance of global average pooling to aggregate the overall statistical features of each channel and avoid local noise interference. It generates channel weight coefficients through two convolutional layers with ReLU and Sigmoid activation functions. It learns the dependency relationship between channels through nonlinear transformation and automatically evaluates the contribution of each channel to the dehazing task. It multiplies the weight coefficients with the input feature map element by element to complete the channel dimension weight calibration.

[0056] 2.2 Spatial Region Feature-Oriented Enhancement: As the second-level medium-granularity feature enhancement mechanism, the target region and the fog background region are located in the spatial dimension. Global max pooling and global average pooling are performed in parallel on the feature maps after channel calibration. Global max pooling extracts the most salient features at each spatial location and captures the sharp edges of small targets; global average pooling extracts the overall statistical features at each spatial location and captures the smooth distribution of fog. The combination of the two achieves accurate differentiation between the target and the background. A spatial weight map is generated through convolutional layers to enhance the features of densely textured areas of the image and to suppress features in smooth backgrounds and fog-rich areas.

[0057] 2.3 Pixel-level fine-grained fog interference suppression: As the third-level fine-grained dehazing mechanism, it solves the problem of local non-uniform fog that the first two levels of attention cannot handle. It uses a double convolutional layer and activation function to generate a pixel-level weight matrix for the spatially enhanced feature map. Through pixel-by-pixel nonlinear transformation, it learns the fog concentration distribution at each pixel position, realizes adaptive weighting for fog of different concentrations, and performs fine-grained weighted suppression for pixels with non-uniform fog distribution. Finally, it outputs a spatial domain fog residual feature map with strong representation ability.

[0058] Based on the above processing procedure, combined with the appendix Figure 6As shown, the CSP hybrid attention module is embedded in multiple feature extraction stages within the main branch of the spatial domain, used for adaptive enhancement and filtering of features at different scales. Specifically, in step 2.1, global average pooling is used to compress the input features spatially, allowing each channel to obtain global receptive field information, thus reflecting the channel's response intensity in the overall image. Subsequently, two convolutional layers and a non-linear activation function are used to model the relationships between channels, generating channel weight coefficients. These weight coefficients adaptively adjust the importance of different channels, strengthening channels that contribute significantly to the dehazing task while suppressing redundant or interfering channels. In step 2.2, two operations, global max pooling and global average pooling, are introduced in parallel on the channel-weighted feature map to extract spatial information from different statistical perspectives. Global max pooling highlights salient response regions, while global average pooling reflects the overall distribution trend. The results of both operations are fused through a convolutional layer to generate a spatial weight map. This weight map can selectively enhance features in the spatial dimension, strengthening textured and structurally significant regions while suppressing smooth regions and areas with concentrated fog, thereby improving the network's ability to discriminate spatial information. In step 2.3, after completing the spatial dimension enhancement, a pixel-level fine-tuning modeling mechanism is introduced. A pixel-level weight matrix of the same size as the input features is generated through continuous convolutional layers and activation functions. This matrix can independently weight each pixel position, enabling the network to finely process the problem of uneven fog distribution. Through this process, the impact of fog interference in local regions on feature expression can be effectively reduced while preserving real structural information, thereby improving the detail clarity and visual realism of the dehazing result.

[0059] The CSP hybrid attention module of this invention achieves full-level dehazing optimization from global to local and from coarse to fine granular by constructing a three-level progressive enhancement link of channel-space-pixel. This overcomes the limitations of traditional attention modules, which can only perform coarse-grained feature selection and cannot adapt to fine-grained processing of non-uniform fog scenes. Its advantages include accurate filtering of multi-scale fog interference, complete preservation of image edge texture details, and effective solutions to pain points in traditional dehazing such as uneven dehazing, loss of details, and artifact distortion. It is highly compatible with the spatial domain multi-scale feature extraction link and can significantly enhance the performance of dual-branch collaborative dehazing. During implementation, it is necessary to strictly follow the cascade order of channel-space-pixel to ensure that the input and output feature dimensions of pooling and convolution operations in each step match, reasonably control the convolution kernel size to avoid feature distortion, and ensure compatibility with the feature scale of the preceding and following modules.

[0060] Example 5: Frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling. Similar to Examples 1-4, step 3 uses the Haar wavelet basis as the discrete wavelet transform function. Wavelet decomposition is performed layer by layer on the input features at each coding level. Utilizing its orthogonality, the feature map is precisely decoupled into a single low-frequency approximate sub-band carrying global fog degradation information and three high-frequency detail sub-bands containing edge texture information. The low-frequency approximate sub-band concentrates more than 90% of the image energy, mainly containing global atmospheric light, large-scale background, and uniformly distributed fog information. The three high-frequency detail sub-bands correspond to edge features in the horizontal, vertical, and diagonal directions, respectively, and contain almost all the contour, texture, and detail information of small targets, with no information overlap or redundancy among them. This wavelet decomposition process has no additional learnable parameters and does not increase the number of network parameters.

[0061] This invention employs a Haar wavelet basis without additional learnable parameters to achieve precise decoupling of high and low frequencies in the frequency domain. It can specifically decompose low-frequency subbands of global fog information and high-frequency subbands of multi-directional edge details to meet dehazing requirements, solving the pain points of traditional frequency domain transformations, such as large number of parameters, high model complexity, and insufficient targeting of high and low frequency decoupling. Its advantages lie in achieving precise decomposition of high and low frequency features without increasing the network's computational burden and number of parameters, perfectly matching the core requirements of global fog modeling and detail fidelity in dehazing tasks, while also possessing excellent lightweight characteristics and adaptability to edge deployment. During implementation, it is necessary to ensure that the wavelet decomposition subbands and subsequent upsampling and downsampling modules are precisely matched in terms of hierarchy and dimension, strictly distinguish the processing links of high and low frequency subbands, and avoid destroying the inherent characteristic attributes of the subbands.

[0062] Example 6: Frequency-Spatial Dual-Branch Collaborative Image Dehazing Method Based on High-Low Frequency Decoupling. Similar to Examples 1-5, the Fast Fourier Convolution Residual Block described in step 3 serves as the core execution unit for global frequency-domain dehazing. It specifically addresses the problem of modeling large-scale continuous fog and long-distance fog distributions, which is difficult to handle with traditional spatial domain convolution. It consists of two serially cascaded Fast Fourier Convolution units. (Refer to...) Figure 9 , Figure 9 This is a structural diagram of the FFC residual block of the present invention. Each fast Fourier convolution unit has local branches and global branches. (Refer to...) Figure 8 , Figure 8This is a structural diagram of the FFC (Fast Fourier Convolution) unit of this invention. The local branch uses two parallel traditional spatial convolutions to extract local detail textures. The global branch uses one parallel spatial convolution and one parallel frequency domain mapping. The frequency domain mapping maps spatial features to the frequency domain through a two-dimensional real-valued Fast Fourier Transform. Utilizing the conjugate symmetry of the real-valued Fast Fourier Transform, the computational load and memory usage are reduced by 50% while fully preserving information, achieving efficient spatial-frequency domain conversion. After frequency domain convolution, the features are mapped back to the spatial domain through an inverse Fourier transform. Subsequently, the output features of the local branch and the global branch are fused in parallel using a two-way weighted fusion. The fused features are then processed by batch normalization and the ReLU activation function, ultimately outputting enhanced features that simultaneously contain local details and global context information, achieving complementarity between local texture details and global fog distribution modeling. This Fast Fourier Convolution residual block completes global context modeling for large-scale continuous fog and non-uniformly distributed fog, and effectively suppresses fog degradation interference.

[0063] This invention employs a serial stacking approach to connect two Fast Fourier Convolutional (FFT) units sequentially in the residual block. This allows features to undergo two interactive modeling processes in the frequency and spatial domains within the same module, thereby progressively enhancing feature representation capabilities. By fusing the input features with the features processed by the two-stage FFT convolutional units through residual connections, not only is the original information preserved, but gradient decay in deep networks is also effectively mitigated, improving the stability and convergence speed of model training.

[0064] Combined with appendix Figure 8 As shown, each Fast Fourier Transform (FFT) convolutional unit employs a dual-path structure with local and global branches. The local branch processes the input features in the spatial domain using standard convolution operations, primarily extracting local details such as edges, textures, and local contrast variations, ensuring that details are not destroyed during dehazing. This branch focuses on high-frequency information representation, effectively compensating for potential local detail loss during frequency domain modeling. Unlike the local branch, the global branch maps the input features from the spatial domain to the frequency domain using a two-dimensional real-valued Fast Fourier Transform (FFT), performing convolution operations in the frequency domain to achieve global information exchange. Since the frequency domain representation inherently possesses a global receptive field, this branch can efficiently model the overall structural features of large-scale continuous fog and non-uniformly distributed fog in images. Subsequently, the processed frequency domain features are remapped back to the spatial domain using an inverse Fourier transform, allowing them to be fused with the output of the local branch, thus achieving a collaborative expression of global and local information.

[0065] The Fast Fourier Convolution residual block of this invention adopts a serially cascaded dual-unit spatial-frequency dual-path parallel architecture. It preserves local detail texture through traditional convolution in local branches, and achieves long-distance global context modeling of large-scale continuous fog and non-uniform fog by combining it with two-dimensional real-valued Fast Fourier Transform in global branches. This overcomes the limitations of traditional spatial convolution in terms of limited receptive field and insufficient ability to model complex fog conditions. Its advantages are that it balances the preservation of fine-grained details with accurate modeling of global fog degradation features, greatly improves the uniformity of defogging in complex fog scenes, and is highly compatible with the high-low frequency decoupling architecture of the frequency domain branches, which can enhance the performance of dual-branch collaborative defogging. In implementation, it is necessary to strictly ensure that the dimension and spatial size of the output features of the dual branches are matched, standardize the FFT and inverse transform parameters to avoid frequency domain artifacts, and ensure that the feature scale is adapted to the preceding and following modules.

[0066] Example 7: Frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling. Similar to Examples 1-6, the dual-domain feature fusion and global residual image reconstruction in step 4 use the spatial domain dehazing main branch and the frequency domain feature enhancement branch to perform same-scale feature fusion, and then perform residual compensation reconstruction with the global residual branch. Specifically, it includes the following steps:

[0067] Step 4.1 Spatial Domain Multi-Scale Feature Aggregation: The main branch of spatial domain dehazing is configured with three sets of cascaded feature processing residual G modules. Each residual G module is arranged sequentially along the feature propagation direction. The output features of each set of residual G modules are passed to the back end of the cascaded residual G modules through long skip connections to complete the multi-scale aggregation of features of different scales and different receptive fields. The aggregated features are input into the CSP hybrid attention module to complete the fine dehazing enhancement, and then processed by the convolutional layer to output the spatial domain fog residual features.

[0068] Step 4.2 Frequency Domain Multi-Scale Detail Compensation: The frequency domain feature enhancement branch adopts a three-level encoding and decoding structure that matches the scale of the spatial domain module. The high and low frequency features output by each level of the downsampling module are fed into the input of the corresponding scale upsampling module through cross-layer jump connections. They are added and fused element by element with the upsampling input features to achieve multi-scale frequency domain detail compensation and feature reuse.

[0069] Step 4.3 Dual-domain collaborative fusion and global residual reconstruction: The spatial domain fog residual features and the frequency domain detail enhancement features are fused at the same scale. The fused features are then reconstructed with the original foggy image as a global residual branch, and finally the clear image after defogging is output, realizing end-to-end image defogging based on frequency domain-spatial dual-branch collaborative fusion of high and low frequency decoupling.

[0070] The dual-domain feature fusion and global residual reconstruction architecture of this invention constructs a collaborative fusion strategy of three sets of hierarchical features in the spatial domain and multi-scale features in the frequency domain, combined with a dual-hop connection mechanism of long skip connections in the spatial domain and cross-level detail skip connections in the frequency domain. Combined with global residual reconstruction of the original image through global long skip connections, it overcomes the limitations of traditional dual-domain dehazing feature fusion, easy loss of multi-scale details, and color shift artifacts in reconstruction. Its advantages are that it can fully aggregate dehazing features of dual domains, multi-scale, and multi-receptive fields, taking into account both the preservation of local details in the spatial domain and the ability to model global fog conditions in the frequency domain, effectively avoiding the problems of detail loss, color shift distortion, and artifacts, and significantly improving the image restoration quality in complex fog scenes. In implementation, it is necessary to strictly ensure that the scales of the three sets of features in the spatial domain correspond one-to-one with the three levels of encoding and decoding in the frequency domain, ensure that the feature dimensions and resolutions of each skip connection are accurately matched, and align the feature channels and sizes before dual-domain fusion to avoid fusion misalignment affecting the reconstruction effect.

[0071] The present invention relates to a frequency-space dual-branch collaborative image dehazing method based on high and low frequency decoupling, which is in the field of image dehazing technology and aims to solve the technical problems of traditional single-domain dehazing algorithms, such as incomplete removal of large-scale non-uniform fog, easy loss of fine-grained image details, and weak generalization ability in complex scenes.

[0072] This invention uses FFA-Net as the baseline architecture and constructs a frequency-spatial domain dual-branch collaborative dehazing network architecture. The spatial domain branch achieves refined dehazing enhancement, while the frequency domain branch completes global fog degradation suppression. A multi-scale hierarchical fusion dual-domain feature collaborative optimization link is then established, combined with global residual reconstruction to achieve high-fidelity, clear image restoration. This invention decouples low-frequency fog degradation information from high-frequency detail information through discrete wavelet transform, and uses a three-level hybrid attention mechanism (CSP channel-space-pixel) to achieve targeted fog feature enhancement. This ensures thorough dehazing in non-uniform fog scenes while preserving fine-grained details such as image edges and textures. The entire inference process does not require manual estimation of prior parameters such as transmittance and atmospheric light, enabling end-to-end dehazing inference. It boasts advantages such as high dehazing accuracy, good detail fidelity, and strong scene generalization, and can be widely applied to foggy image restoration in fields such as UAV aerial photography, security monitoring, vehicle-mounted assisted driving, and remote sensing mapping.

[0073] Example 8: Frequency-Spatial Dual-Branch Collaborative Image Dehazing Method Based on High-Low Frequency Decoupling (same as Examples 1-7). This invention is an image dehazing method based on a frequency-spatial dual-branch approach. For a single input foggy image, it performs end-to-end dehazing processing through a dual-branch collaborative network, outputting a high-fidelity clear image. The spatial domain dehazing main branch and the frequency domain feature enhancement branch are parallel input structures, sharing the same original foggy image input. The features output by the two branches are fused at the same scale. The network as a whole is a fully convolutional architecture, without any manual prior parameter dependency. Specifically, it includes the following steps:

[0074] Step 1: Construct the dehazing network architecture: Build a frequency-domain / spatial dual-branch collaborative dehazing network in the PyTorch deep learning framework. See [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart of the overall architecture of the dehazing network of this invention. The spatial domain dehazing main branch is based on a cascaded set of basic residual blocks embedded with CSP hybrid attention modules. It sets up a 3×3 initial convolutional layer, three cascaded feature processing groups (G-1, G-2, G-3), CSP hybrid attention modules, and convolutional layers. Each feature processing group contains basic residual blocks embedded with CSP hybrid attention modules, completing the hierarchical extraction and enhancement of spatial domain features. The frequency domain feature enhancement branch is based on an encoding and decoding structure with discrete wavelet transform and fast Fourier convolution as its core, and adopts a three-dimensional array with a scale matching the three sets of modules in the spatial domain. The encoding / decoding structure consists of an initial downsampling module, two downsampling modules, three cascaded Fast Fourier Convolution (FFC) residual blocks, two upsampling modules, and an output convolutional layer. Each downsampling module integrates a discrete wavelet transform unit based on the Haar wavelet basis, decomposing the input feature map into one low-frequency approximate sub-band and three high-frequency detail sub-bands. The downsampling factor is 2, corresponding to feature scales G-1, G-2, and G-3 in the spatial domain. Each level has a cross-layer skip connection, and the high-frequency sub-bands are directly passed to the upsampling stage through multi-scale skip connections, ultimately constructing a complete end-to-end dehazing network.

[0075] Step 2: Training Dataset Construction and Preprocessing: The network was trained using the RESIDE standard dehazing dataset, which included 13,990 indoor synthetic foggy images and 296,695 outdoor synthetic foggy images. The test set consisted of SOTS indoor and outdoor subsets. The images in the training set were normalized to a uniform resolution of 256×256. Data augmentation strategies such as random brightness adjustment, contrast transformation, random cropping, and horizontal flipping were employed to expand the diversity of the training data and avoid network overfitting.

[0076] Step 3: Network Training Parameter Settings: The training hardware uses an NVIDIA RTX 4090 graphics card, and the training environment is built based on the PyTorch deep learning framework; the Adam adaptive optimizer is used, with a first-order moment exponential decay rate. =0.9, second-order moment exponential decay rate =0.999, initial learning rate set to 0.001, batch size set to 16, total training epochs set to 200; the corresponding joint loss function is used as the objective function for network optimization.

[0077] Step 4 Network Iterative Training and Weight Saving: Complete end-to-end iterative training of the network according to the set training parameters. Perform accuracy verification on the validation set after each training round and record the network's PSNR and SSIM metrics. When the validation set metrics no longer improve after 20 consecutive rounds, terminate the training early. Save the weight file with the best validation set metrics during the training process as the final inference weights of the FDF-Net of this invention.

[0078] Step 5: Fog Image Dehazing Inference: The trained network weights are loaded into the inference environment. An input foggy image of any size is provided. Without manual preprocessing, the network first extracts features through an initial convolutional layer, then performs refined dehazing enhancement via the spatial domain main branch. This is achieved through a CSP hybrid attention module, sequentially performing channel-dimensional adaptive weighting, spatial region feature orientation enhancement, and pixel-level fog interference filtering, removing fog degradation features from the image layer by layer while simultaneously enhancing image edges and fine-grained texture details. A frequency domain branch performs global fog degradation suppression—the input foggy image is decomposed into low-frequency and high-frequency sub-bands using discrete wavelet transform. For the low-frequency sub-band, a large-scale non-uniform fog distribution is globally modeled and interference suppressed using Fast Fourier Convolutional Residual Blocks. For the high-frequency sub-band, multi-scale skip connections are used to directly pass the data to the upsampling stage. Finally, through dual-domain feature fusion and global residual reconstruction, a clear dehazed image is output end-to-end.

[0079] This invention presents a frequency-space dual-branch collaborative image dehazing method based on high- and low-frequency decoupling. This method achieves decoupling of global fog degradation suppression and local detail enhancement through a frequency-space dual-branch collaborative architecture. Combined with a CSP hybrid attention module, it enables refined filtering and enhancement of fog-related features at the channel, spatial, and pixel levels. This solves the problems of incomplete dehazing, easy loss of details, and weak generalization ability in complex scenes found in traditional dehazing algorithms. This invention uses a fully convolutional end-to-end architecture, requiring no manual estimation of any prior parameters. After training, it can directly perform dehazing processing on foggy images, achieving high dehazing accuracy, fast inference speed, and strong scene adaptability. It can be widely applied to foggy image restoration in fields such as UAV aerial photography, security monitoring, vehicle-mounted assisted driving, and remote sensing mapping.

[0080] The technical effects of the present invention will be further explained below through comparative experimental results.

[0081] Example 9: To verify the actual dehazing performance of the frequency-spatial dual-branch dehazing network of the present invention, this example selects the authoritative RESIDE standard dataset in the field of single-image dehazing and builds indoor and outdoor dual-scene comparison experiments. The present invention is compared with existing mainstream dehazing algorithms under the same conditions. The comparison algorithms include the traditional physical prior algorithm DCP, the early lightweight deep learning dehazing algorithm AOD-Net, the classic CNN dehazing network DehazeNet, the single-domain context aggregation dehazing network GCANet, and the baseline attention dehazing network FFA-Net. This comprehensively verifies the dehazing accuracy, detail fidelity, and scene generalization of the present invention in different scenarios. The image dehazing method based on the frequency-spatial dual-branch is the same as in Examples 1-4. The specific process of the comparison experiment is as follows:

[0082] Experimental conditions: The hardware and software environment of this experiment is completely consistent with that of Example 4. The hardware uses an NVIDIA RTX4090 graphics card, and the software is built on the PyTorch 1.12 deep learning framework and Python 3.8 environment. The training hyperparameters and preprocessing procedures of all deep learning algorithms are completely consistent to ensure that the experimental results are fair and reproducible.

[0083] The experiment used the official training and testing sets of the RESIDE standard dataset to conduct defogging performance tests in both indoor and outdoor scenarios.

[0084] Evaluation metrics and experimental design: In this experiment, we selected Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), which are commonly used in the field of image dehazing, as the core quantitative evaluation metrics. The unit of PSNR is dB. The higher the value, the smaller the pixel deviation between the dehazed image and the ground truth image, and the higher the restoration accuracy. The value of SSIM ranges from 0 to 1. The closer the value is to 1, the more complete the image details, textures and structures are preserved, and the better the visual fidelity.

[0085] DCP is a traditional manual algorithm that requires no training and is directly tested through inference. AOD-Net, DehazeNet, GCANet, FFA-Net, and the FDF-Net of this invention were each trained end-to-end for 200 rounds on the RESIDE training set for their respective scenarios. After freezing the optimal weights, they were inferred frame by frame on the SOTS indoor and outdoor test sets. The average PSNR and average SSIM of all test samples were calculated, and the dehazing effect was compared and visualized. GCANet was not tested on the outdoor dataset, so its corresponding metrics are empty.

[0086] Experimental Content: This example selects the authoritative RESIDE standard dataset in the field of image dehazing to build an indoor and outdoor dual-scene comparison simulation experiment. Five classic dehazing models, FFA-Net, DCP, AOD-Net, DehazeNet, and GCANet, are selected as control groups, while the image dehazing method based on frequency domain-spatial dual branches proposed in this invention is used as the experimental group. Based on a unified hardware and software environment, training hyperparameters, and preprocessing procedures, the training convergence and optimal weight freezing of each deep learning model are completed.

[0087] Batch inference tests were conducted on the SOTS indoor and SOTS outdoor test subsets accompanying the RESIDE dataset. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) core quantitative indicators for the dehazing restoration results of each model were calculated for each sample. Visual comparison charts of the dehazing effects of each model were generated simultaneously. Figure 10 and Figure 11 .

[0088] To verify the dehazing performance of the FDF-Net of this invention, comparative experiments were conducted using indoor and outdoor test subsets of the RESIDE standard dataset. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were selected as evaluation metrics. The invention was compared with existing mainstream dehazing methods such as DCP, AOD-Net, DehazeNet, GCANet, and FFA-Net. The experimental results are shown in Table 1. Data that are "not disclosed" in the table are indicated by a hyphen "-".

[0089] Table 1. Performance comparison of different dehazing methods on the RESIDE dataset.

[0090]

[0091] Table 1 compares the performance of different dehazing methods on the RESIDE dataset. As shown in the table, the method of this invention achieves a PSNR of 37.94 dB and an SSIM of 0.9921 on the indoor dataset, and a PSNR of 35.12 dB and an SSIM of 0.9886 on the outdoor dataset. All these metrics are superior to DCP, AOD-Net, DehazeNet, GCANet, and the baseline network FFA-Net, demonstrating that the frequency-spatial dual-branch structure and CSP hybrid attention mechanism of this invention can significantly improve image dehazing accuracy and detail fidelity, and have stronger adaptability to complex fog scenes.

[0092] Experimental results and analysis: See Figure 10 and Figure 11In the figure, "Our" refers to "the model of this invention," which shows the dehazing effect of the proposed model on the RESIDE indoor and outdoor datasets, and compares it with algorithms such as DCP, AOD-Net, and DehazeNet. Experimental results show that DCP, due to its reliance on prior assumptions, is prone to significant color distortion; AOD-Net's dehazing is incomplete, and the overall brightness of the output image is low; DehazeNet restores the brightness of the image to a high level; GCANet still performs poorly in restoring high-frequency details such as texture, edges, and blue sky; FFA-Net, as the baseline architecture of this invention, although superior to the above methods in terms of dehazing effect and color restoration, still has certain shortcomings in detail preservation, such as the slight smoothing blurring of the wood texture of the chair and the outline of the chandelier in the indoor scene.

[0093] In contrast, the frequency-spatial dual-branch collaborative image dehazing method proposed in this invention, based on high-low frequency decoupling, achieves three-level refined dehazing enhancement at the channel, spatial, and pixel levels through the CSP hybrid attention module of the spatial domain dehazing main branch. Simultaneously, it completes global fog degradation suppression through the Haar wavelet basis discrete wavelet transform and fast Fourier convolution residual block of the frequency domain feature enhancement branch. Combining dual-domain feature fusion and global residual image reconstruction link, it achieves multi-scale collaborative optimization of features, resulting in superior performance in detail restoration and color fidelity. It can effectively improve the overall image clarity and visual realism, while achieving efficient removal of large-scale non-uniform fog.

[0094] In summary, the frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling of the present invention solves the technical problems of incomplete dehazing of large-scale non-uniform fog, easy loss of fine-grained details, and weak generalization ability in complex scenes in traditional single-domain image dehazing algorithms. It is realized by constructing a frequency-spatial dual-branch collaborative dehazing network architecture; completing fine dehazing enhancement through the spatial domain dehazing main branch; completing global fog degradation suppression through the frequency domain feature enhancement branch; constructing a dual-domain feature fusion and global residual image reconstruction link; and then completing end-to-end network training optimization and application.

[0095] The dual-branch collaborative dehazing architecture constructed in this invention decouples spatial domain refinement enhancement from frequency domain global modeling. It achieves feature enhancement at the channel, spatial, and pixel levels through a CSP hybrid attention module, ensuring simultaneous improvement in both dehazing thoroughness and detail fidelity. This invention is a fully convolutional end-to-end architecture, with parallel input structures in the spatial and frequency domains, sharing the same original foggy image input. It has no dependency on prior parameters, requires no manual intervention, can adapt to input images of any size, has fast inference speed, and is easy to implement in engineering. The dehazing effect achieved through dual-domain collaborative optimization is stable and has strong scene generalization ability, making it widely applicable to foggy image restoration in fields such as UAV aerial photography, security monitoring, vehicle-mounted assisted driving, and remote sensing mapping.

Claims

1. A frequency domain-space dual-branch collaborative image defogging method based on high-low frequency decoupling, characterized in that, It includes the following steps: Step 1: Construct a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP: Using the feature fusion attention network FFA-Net for image dehazing as the baseline architecture, and relying on the inherent global residual mechanism of FFA-Net, whose global residual input is the foggy image to be processed and whose output is the network terminal, a parallel spatial domain dehazing main branch and frequency domain feature enhancement branch architecture are constructed. The input of this architecture is connected to a shared initial feature extraction convolutional layer, and the output is connected to the feature fusion layer at the end of the network, together forming a frequency-domain-spatial dual-branch collaborative dehazing network combining high- and low-frequency decoupling and CSP; The spatial domain dehazing main branch has a cascaded residual G module with multiple sets of embedded CSP hybrid attention as the backbone structure, which sequentially includes and connects a shared initial feature extraction convolutional layer, multiple sets of cascaded residual G modules, a multi-scale cross-layer connection module, a post-CSP hybrid attention module, and a terminal dimension matching convolutional layer; The residual G modules are all based on basic residual blocks embedded with CSP hybrid attention as basic units, and the cascaded multiple sets of residual G modules and the multi-scale cross-layer connection module are sequentially connected. Cross-layer connections; the frequency domain feature enhancement branch adopts a symmetrical U-shaped encoding and decoding architecture that combines discrete wavelet transform and fast Fourier convolution. The overall structure is a three-level symmetrical encoding-decoding structure, and the backbone is constructed according to the logic of encoding-bottleneck feature optimization-decoding: the front end first undergoes initial feature extraction processing through a shared initial feature extraction convolutional layer, followed by three sets of encoding downsampling modules embedded with discrete wavelet transform, the middle bottleneck layer embeds fast Fourier convolution residual blocks for global frequency domain feature optimization, and the back end is configured with three sets of decoding upsampling modules that correspond one-to-one with the encoding level; at the same time, multi-scale cross-layer jump connections are introduced between the downsampling module and the upsampling module, feeding the high and low frequency features obtained by discrete wavelet transform decomposition at each level of the encoding end into the corresponding level of the decoding upsampling module layer by layer, realizing multi-scale frequency domain detail compensation and feature reuse; the spatial domain dehazing main branch and the frequency domain feature enhancement branch constitute a dual-branch collaborative body, and a global residual branch is introduced for image identity mapping constraints, forming an end-to-end high and low frequency decoupling and CSP combined frequency domain-spatial dual-branch collaborative dehazing network; Step 2 Spatial Domain Main Branch Dehazing Enhancement Processing: After the foggy image is input into the spatial domain dehazing main branch, shallow initial feature extraction is first completed through the shared initial feature extraction convolutional layer at the front end of the dehazing network; Subsequently, the data is fed into multiple sets of residual G modules with embedded CSP hybrid attention to achieve hierarchical feature extraction and preliminary filtering of fog degradation features. Relying on the CSP hybrid attention module, the channel dimension adaptive weight calibration, spatial fog region feature orientation enhancement, and pixel-level fog interference fine suppression are completed in sequence. While hierarchically representing the fog degradation information of the image, the modeling of the fog distribution area is focused on, and a spatial domain fog residual feature map with strong representation ability is output. Step 3: Global Fog Degradation Suppression Processing via Frequency Domain Feature Enhancement Branch: Global fog degradation suppression is achieved through the frequency domain feature enhancement branch. First, the foggy image is processed through a shared initial feature extraction convolutional layer at the front end of the defogging network to complete shallow feature extraction, and then fed into a three-level symmetric encoder-decoder structure. Each encoding level decomposes the features layer by layer into low-frequency subbands carrying global fog degradation information and high-frequency subbands containing edge textures through discrete wavelet transform. The first two levels of high-frequency and low-frequency subbands are directly fed into the corresponding decoding levels through multi-scale cross-layer skip connections to avoid the loss of fine-grained details during encoding and decoding. The low-frequency subbands converge to the intermediate bottleneck layer, and global modeling and degradation interference suppression of a large-scale non-uniform fog distribution are completed through fast Fourier convolution residual blocks. After decoding, reconstruction, and fusion by the frequency domain feature enhancement branch, the output is a frequency domain enhanced feature with global frequency domain constraints and high-frequency detail compensation capabilities. Step 4: Dual-domain feature fusion and global residual image reconstruction: The fog residual features output from the spatial domain dehazing main branch and the frequency domain detail enhancement features output from the frequency domain feature enhancement branch are fused at the same scale. The fused features are then reconstructed with the original foggy image, which serves as the global residual branch, to fully leverage the synergistic advantages of multi-scale fine dehazing in the spatial domain and high- and low-frequency detail reuse in the frequency domain. This maximizes the preservation of the image's edge texture and structural information, ultimately outputting a clear dehazed image. This achieves end-to-end image dehazing based on high- and low-frequency decoupling and frequency domain-spatial dual-branch collaborative technology.

2. The frequency domain-space dual branch collaborative image defogging method based on high-low frequency decoupling according to claim 1, characterized in that: The cascaded residual G modules with embedded CSP hybrid attention in the spatial domain dehazing main branch described in step 1 are composed of three sets of serially arranged multi-scale feature extraction residual G modules. Each residual G module adopts a dual-path parallel structure. One path is the main branch, which is composed of three sets of cascaded basic residual blocks and convolutional layers connected in series. The other path is a direct-connection residual branch, which directly transmits the input features of the residual G modules to the end of the main branch and fuses them with the output features of the main branch. The basic residual block is a CSP hybrid attention module with two-level residual connections. Its first-level residual structure is composed of a convolutional layer and a ReLU activation function connected in series and combined with the module input features. The second-level residual structure is composed of a convolutional layer and a CSP hybrid attention module connected in series and combined with the input features. The two levels of residual structures are serially spliced ​​together to form the basic residual block.

3. The frequency domain-space dual branch collaborative image defogging method based on high-low frequency decoupling according to claim 1, characterized in that: The frequency domain feature enhancement branch described in step 1 consists of three sets of front-end coding downsampling modules, divided into one initial downsampling module and two sets of activation downsampling modules. The initial downsampling module adopts a dual-branch structure. The main branch is composed of convolutional layers and batch normalization layers connected in series. The auxiliary branch embeds discrete wavelet transform layers and convolutional layers. The input features are decomposed into low-frequency sub-bands and high-frequency sub-bands by the discrete wavelet transform layer. The low-frequency sub-bands are fused with the main branch features after feature mapping by the convolutional layer. The high-frequency sub-bands are transmitted to the corresponding level of the upsampling module at the decoding end through cross-layer skip connections. The two sets of activation downsampling modules have a structure that is basically the same as the initial downsampling module, except that a Leaky ReLU activation function is added at the end of the main branch. The three sets of back-end decoding upsampling modules correspond one-to-one with the coding downsampling modules and all adopt a dual-input branch structure to receive the low-frequency features and high-frequency features transmitted from the coding end, respectively. The low-frequency input branch consists of a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in series, while the high-frequency input branch only uses a single convolutional layer to complete feature mapping. The features of the two branches are fused to output the upsampling result of the current layer. At the same time, the intermediate bottleneck layer embeds a fast Fourier convolutional residual block to complete global frequency domain feature optimization. High and low frequency feature reuse and texture detail compensation are achieved by relying on multi-level skip connections.

4. The frequency domain-space dual branch collaborative image defogging method based on high-low frequency decoupling according to claim 1, characterized in that, The CSP hybrid attention module in the spatial domain main branch dehazing enhancement process described in step 2 consists of three parts in sequence: channel dimension adaptive weighting, spatial region feature-oriented enhancement, and pixel-level refined suppression of fog interference. 2.1 Channel Dimension Adaptive Weighting: Global average pooling is used to extract global spatial information of each channel from the input feature map of the module. Channel weight coefficients are generated by two convolutional layers with ReLU and Sigmoid activation functions. The weight coefficients are multiplied element-wise with the input feature map to complete the channel dimension weight calibration. 2.2 Spatial Region Feature Targeting Enhancement: Global max pooling and global average pooling are performed in parallel on the feature maps after channel calibration. Spatial weight maps are generated through convolutional layers to enhance the features of dense texture regions in the image, and feature suppression is performed on smooth backgrounds and fog-rich regions. 2.3 Pixel-level refined suppression of fog interference: A pixel-level weight matrix is ​​generated by using a double convolutional layer and activation function on the spatially enhanced feature map. Refined weighted suppression is performed on pixels with non-uniform fog distribution, and the final output is a spatial domain fog residual feature map with strong representation ability.

5. The frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling according to claim 1, characterized in that: Step 3 uses the Haar wavelet basis as the discrete wavelet transform function, and performs wavelet decomposition on the input features layer by layer at each coding level. By utilizing its orthogonality, the feature map is precisely decoupled into a single low-frequency approximate subband carrying global fog degradation information and three high-frequency detail subbands containing edge texture information. This wavelet decomposition process has no additional learnable parameters and does not increase the number of network parameters.

6. The frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling according to claim 1, characterized in that: The Fast Fourier Convolution Residual Block described in Step 3 consists of two serially cascaded Fast Fourier Convolutional Units. Each Fast Fourier Convolutional Unit has a local branch and a global branch. The local branch uses two traditional spatial domain convolutions in parallel to extract local detail textures. The global branch uses one spatial domain convolution and one frequency domain mapping in parallel. The frequency domain mapping maps the spatial features to the frequency domain through a two-dimensional real-valued Fast Fourier Transform, and after frequency domain convolution, it is mapped back to the spatial domain through an inverse Fourier transform. Subsequently, the output features of the local branch and the output features of the global branch are fused in parallel. The fused features are then processed by batch normalization and ReLU activation function, and the final output is an enhanced feature that contains both local details and global context information, achieving complementarity between local texture details and global fog distribution modeling. This Fast Fourier Convolution Residual Block completes global context modeling of large-scale continuous fog and non-uniformly distributed fog, and effectively suppresses fog degradation interference.

7. The frequency-spatial dual-branch collaborative image dehazing method based on high-low frequency decoupling according to claim 1, characterized in that: Step 4, the dual-domain feature fusion and global residual image reconstruction, employs the spatial domain dehazing main branch and the frequency domain feature enhancement branch to perform same-scale feature fusion, followed by residual compensation reconstruction with the global residual branch. Specifically, it includes the following steps: Step 4.1 Spatial Domain Multi-Scale Feature Aggregation: The main branch of spatial domain dehazing is configured with three sets of cascaded feature processing residual G modules. Each residual G module is arranged sequentially along the feature propagation direction. The output features of each set of residual G modules are passed to the back end of the cascaded residual G modules through long skip connections, completing the multi-scale aggregation of features of different scales and different receptive fields. The aggregated features are input into the CSP hybrid attention module to complete the fine dehazing enhancement, and then processed by the convolutional layer to output the spatial domain fog residual features. Step 4.2 Frequency Domain Multi-Scale Detail Compensation: The frequency domain feature enhancement branch adopts a three-level encoding and decoding structure that matches the scale of the spatial domain module. The high and low frequency features output by each level of the downsampling module are fed into the input of the upsampling module of the corresponding scale through cross-layer jump connections. They are added and fused element by element with the upsampling input features to achieve multi-scale frequency domain detail compensation and feature reuse. Step 4.3 Dual-domain collaborative fusion and global residual reconstruction: The spatial domain fog residual features and the frequency domain detail enhancement features are fused at the same scale. The fused features are then reconstructed with the original foggy image as a global residual branch, and finally the clear image after defogging is output, realizing end-to-end image defogging based on frequency domain-spatial dual-branch collaborative fusion of high and low frequency decoupling.