Defogging method based on interlayer multi-scale sequence interaction and Fourier domain frequency-space enhancement

By introducing inter-layer multi-scale sequence interaction module and Fourier domain frequency-space enhancement module into the defogging network, the problems of insufficient inter-layer feature interaction and insufficient frequency-domain high-frequency information processing in the prior art are solved, and more efficient image defogging effect and detail recovery are achieved.

CN120163739AActive Publication Date: 2025-06-17CHINA THREE GORGES UNIV

Patent Information

Application Number
CN202510226121.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing defogging method is difficult to fully realize the deep interaction of interlayer features in complex haze scenarios, resulting in insufficient fusion of multi-scale features and insufficient high-frequency information processing in frequency domain feature mining, affecting image detail recovery and structural consistency.

Method used

Using the defog removal method based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, multi-scale interaction and frequency-space collaboration processing of features is realized by constructing a U-shaped defog network, inter-layer multi-scale sequence interaction module IMSIM and Fourier domain frequency-space enhancement module FDFSEM.

Benefits of technology

It significantly improves the effect of image defog removal, improves image detail expression ability and structural consistency, enhances the global and local feature expression ability of the network, and improves the defog removal performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163739A_ABST
    Figure CN120163739A_ABST
Patent Text Reader

Abstract

The invention discloses a defogging method based on interlayer multi-scale sequence interaction and Fourier domain frequency-space enhancement, and the method comprises the following steps: S1, collecting a public foggy data set containing different concentrations of haze, and carrying out the preprocessing of an image in the public foggy data set, so as to improve the data quality and a model training effect; s2, constructing a defogging network framework based on interlayer multi-scale sequence interaction and Fourier domain frequency-space enhancement, wherein the framework comprises a U-shaped defogging network, an interlayer multi-scale sequence interaction module IMSIM and a Fourier domain frequency-space enhancement module FDFSEM; s3, training the network by using the marked image data set, defining a proper loss function, and training by using an optimizer; s4, evaluating a defogging effect by using two indexes, namely a peak signal-to-noise ratio (PSNR) and a structural similarity index (SSIM), optimizing the trained model, and deploying the model to practical application for online reasoning and real-time defogging; through the steps, defogging of the input feature map can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a defogging method based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement. Background Art

[0002] With the acceleration of the urbanization and industrialization processes, hazy weather has become a common phenomenon globally, having a profound impact on social operation and economic activities. In a hazy environment, the quality of images collected by visible light imaging systems significantly deteriorates, manifested as problems such as reduced contrast, lost details, and color deviation. These image degradation phenomena not only weaken the perception ability of the human visual system but also pose major challenges to computer vision tasks that rely on high-quality visual data, such as autonomous driving, object detection, and remote sensing image analysis. Therefore, studying defogging techniques for effectively restoring image clarity and details in a hazy environment has important theoretical research value and practical application significance.

[0003] Existing image defogging methods can generally be divided into two categories: defogging methods based on parameter priors and defogging methods based on deep learning. Defogging methods based on parameter priors usually rely on classical atmospheric scattering models and process hazy images by estimating intermediate parameters such as global atmospheric light and transmittance to restore clear images. Although such methods have a certain defogging ability under specific conditions, they have inherent limitations. For example, in scenes with uneven illumination or high haze concentration, estimation errors of intermediate parameters may lead to problems such as artifacts, color distortion, and lost details. In addition, due to the parameter prior assumptions usually being targeted at specific scene conditions, the generalization ability of such methods in complex or unknown environments is poor.

[0004] In recent years, with the development of deep learning technology, end-to-end deep learning dehazing methods have gradually become the mainstream. Such methods directly learn the mapping relationship between hazy images and clear images, avoiding the dependence on intermediate parameters of the atmospheric scattering model and significantly improving the dehazing performance and generalization ability. For example, MSBDN proposed by Dong et al. enhances the dehazing effect and detail restoration ability of the network in complex hazy scenes through multi-scale feature fusion and dense connections, and performs particularly well under different weather conditions. GridDehazeNet designed by Liu et al. successfully achieves efficient dehazing of complex hazy images by introducing a grid structure, multi-scale feature extraction, and residual modules. DehazeFormer developed by Song et al. improves the image dehazing effect by introducing the Transformer architecture and self-attention mechanism, and performs particularly well in global information modeling and long-range dependence capture. KFA-Net proposed by Jiang et al. combines asymmetric size feature cascading, k-means pixel attention network, and channel attention network FCA, enhancing the feature extraction, thick fog area focusing, and frequency domain information attention for non-uniform hazy remote sensing images, and significantly improving the image dehazing performance.

[0005] Although existing dehazing methods have improved the image quality to a certain extent, it is still difficult to completely eliminate artifacts and detail distortion when dealing with complex hazy scenes. The specific problems are as follows:

[0006] (1) Most dehazing networks adopt a U-shaped structure, usually ignoring the inter-layer feature interaction between encoding layers, resulting in limited efficiency and effect of multi-scale feature fusion. In addition, the limitations of convolutional operations make the network receptive field mainly concentrated in local areas, making it difficult to effectively model long-range dependence relationships and restricting the capture and expression ability of global features. This is particularly obvious when dealing with complex hazy scenes, and it is difficult to simultaneously restore image details and maintain the consistency of the global structure.

[0007] (2) Many existing methods mainly focus on time-domain or spatial-domain features, ignoring the exploration of frequency-domain information. Although frequency-domain information such as global structure and high-frequency details is crucial for improving image clarity and contrast, the combination of time-frequency domain information is still insufficient. This restricts the restoration of image details and textures, further affecting the visual effect of the image and the accuracy of subsequent processing. Summary of the Invention

[0008] The present invention aims to solve several key problems in existing image dehazing methods, specifically: most dehazing networks based on the U-shaped structure fail to fully achieve deep interaction of inter-layer features, resulting in insufficient multi-scale feature fusion. Moreover, in complex haze scenarios, it is difficult to effectively combine global information and local details, thereby causing blurred image texture and lost details after dehazing. In addition, there are deficiencies in frequency-domain feature mining in the prior art, especially in the processing of high-frequency information. High-frequency information is crucial for detail restoration and edge enhancement, but many methods fail to fully utilize high-frequency components and global structural features in the frequency domain, leading to unclear image details and poor structural consistency. More importantly, the combination of time-frequency domain information is still insufficient, further restricting the improvement of image restoration effect.

[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] A dehazing method based on inter-layer multi-scale sequence interaction and Fourier-domain frequency-space enhancement, comprising the following steps:

[0011] Step S1: Collect a publicly available hazy dataset containing different concentrations of haze, and preprocess the images therein to improve data quality and model training effect;

[0012] Step S2: Construct a dehazing network framework based on inter-layer multi-scale sequence interaction and Fourier-domain frequency-space enhancement, which framework includes a U-shaped dehazing network, an inter-layer multi-scale sequence interaction module IMSIM, and a Fourier-domain frequency-space enhancement module FDFSEM;

[0013] Step S3: Use the labeled image dataset to train the network, define an appropriate loss function, and use an optimizer for training;

[0014] Step S4: Use two metrics, namely peak signal-to-noise ratio PSNR and structural similarity index SSIM, to evaluate the dehazing effect, and compare with other methods to verify the model performance. Optimize the trained model and deploy it to actual applications for online inference and real-time dehazing;

[0015] Through the above steps, dehazing of the input feature map can be achieved.

[0016] In step S1, collect a publicly available hazy dataset with different concentrations of haze, and then screen and sort the images; then, preprocess the images in the dataset, including image normalization and data augmentation, so as to improve the generalization ability of the model.

[0017] In step S2, when constructing the U-shaped dehazing network, it specifically includes: an encoding layer feature extraction module, a residual module, a decoding layer image restoration module, and a third decoding layer; the encoding layer feature extraction module includes a first encoding layer, a second encoding layer, and a third encoding layer; the residual module; the decoding layer image restoration module includes a third decoding layer, a second decoding layer, and a first decoding layer;

[0018] When constructing the inter-layer multi-scale sequence interaction module IMSIM, local and global context information in the features of different encoding layers is extracted through multi-scale convolution. At the same time, the sequence modeling mechanism mLSTM is used to capture the long-range dependencies between layers, thereby strengthening the interaction and expression ability of multi-scale features. The features output by the module are further optimized through the residual module to generate aggregated features that fuse edge details and semantic information, and finally transmitted to the third decoding layer to significantly improve the dehazing effect of the network;

[0019] When constructing the Fourier domain frequency-space enhancement module FDFSEM, in the decoding stage, for the corresponding encoding layer, frequency domain features are extracted through Fourier transform to strengthen high-frequency information, and spatial domain features are extracted in combination with convolution operations to maintain texture integrity. After frequency-space collaborative processing, the optimized features are transmitted to the corresponding decoding layer to remove artifacts and enhance structural consistency.

[0020] Specifically, in the encoding stage, a hazy image is input, and multi-scale features are gradually extracted through the first encoding layer, the second encoding layer, and the third encoding layer. At the same time, in combination with the inter-layer multi-scale sequence interaction module IMSIM, the features of different encoding layers are modeled and interacted to improve the fusion ability of cross-layer information. During this process, the encoding layer not only captures local details but also extracts global context information to provide high-quality features for the decoding stage. In the encoding stage, for the encoding features corresponding to each decoding layer, the Fourier domain frequency-space enhancement module FDFSEM is introduced to strengthen the expression of high-frequency details through the collaborative processing of the frequency domain and the spatial domain. The optimized features are directly transmitted to the decoding stage through skip connections to ensure the efficient transmission of multi-scale information. In the decoding stage, the spatial resolution of the image is restored through gradual upsampling. The third decoding layer, the second decoding layer, and the first decoding layer combine the encoding features and the output of the frequency-space enhancement module to achieve feature fusion through residual connections, effectively improving the image detail expression ability. Finally, the network outputs a clear dehazed image. To optimize the network performance, the network adopts a series of loss functions to constrain the pixel consistency, detail restoration ability, and global feature expression of the dehazed image, thereby learning the mapping relationship between the input hazy image and the target clear image, and significantly improving the dehazing effect and image quality.

[0021] Among them, when constructing the U-shaped network, the features of the input image are extracted layer by layer through four-fold downsampling operations. The first encoding layer completes the preliminary feature extraction, the second encoding layer further extracts features, and the third encoding layer captures deep semantic features and context information. Each encoding layer first performs feature downsampling through convolutional operations to reduce the spatial dimension and extract higher-level features, then normalizes the feature map to ensure consistency and stability, and enhances the non-linear expression ability of the network through the ReLU activation function. The downsampled feature map is further subjected to deep feature extraction through multiple residual modules to capture richer semantic information and global context.

[0022] In the decoding stage, the features extracted by the encoder gradually restore the spatial resolution of the image through upsampling operations. Among them, the third decoding layer restores the global features, the second decoding layer further fuses the feature information, and the first decoding layer gradually restores the detail information of the image and restores it to the spatial size of the input image. Through layer-by-layer upsampling and feature fusion, the final defogged image is output to complete the defogging task.

[0023] When constructing the inter-layer multi-scale sequence interaction module IMSIM, since different encoding layers of the U-shaped defogging network carry important information respectively, such as shallow features focusing on capturing detailed texture information, while deep features contain more global semantic information such as color and brightness. Therefore, in order to make full use of the information of features at different levels in the encoding stage, reduce the loss of feature information during transmission, enhance the reconstruction effect of the defogged image in the subsequent decoding stage, optimize the feature interaction and fusion between layers, and further improve the overall defogging performance of the network. For the input feature F k ∈R C×H×W (k ∈ {i, j}), first perform the Layer Normalization operation on it to standardize the feature distribution and enhance the training stability. Subsequently, the feature F k is successively passed through convolutional operations with three different receptive fields, namely 3*3 convolution, 5*5 convolution, and 7*7 convolution, to gradually extract local and global context information. The overall feature extraction process can be described as:

[0024] F ms,k =Conv 7×7 (Conv 5x5 Conv 3×3 (LN(F k ))), (k ∈ {i, j})

[0025] Among them, F ms,k ∈R C×H×W represents the feature map obtained after multi-scale convolutional operations. The sequential operations of the three convolutional kernels can capture feature information under different receptive fields, thereby strengthening the global semantic representation while retaining edge details. The extracted multi-scale feature Fms,k is transformed into a two-dimensional sequence form F seq,k ∈R (H·W)×C to adapt to the sequence modeling requirements:

[0026] F seq,k = Flatten(F ms,k ),(k ∈ {i, j})

[0027] The feature sequence is input into the mLSTM module, and the feature representation is further optimized by capturing the long-range dependencies between sequences, generating the updated sequence feature F' seq,k ∈R (H·W)×C :

[0028] F' seq,k = mLSTM(F seq,k ), k ∈ {i, j}

[0029] Through this operation, the module can effectively extract the long-range dependency information across layers, making the relationship between shallow and deep features closer and laying a foundation for subsequent feature fusion. Then, the cross-layer feature relationship calculation and fusion are carried out. Calculate the relationship matrix between the transpose of the shallow feature F' seq,i ∈R (H·W)×C of the shallow feature and the deep sequence feature F' seq,i T ∈R C×(H·W) and F' seq,j ∈R (H·W)×C :

[0030] W = Softmax(F' seq,i T · F' seq,j )

[0031] The weight matrix W ∈ R C×C is used to measure the interaction relationship between the two layers of features and captures the complementary information of the features from different encoding layers. Using the weight matrix W, the features F' seq,i of the shallow path are weighted and summed to generate the fused features:

[0032] F weighted = F' seq,i · W

[0033] The weighted features F weighted ∈R (H·W)×C are restored to the original feature map shape F res ∈R C×H×W , and the fused features are further compressed in channel information through 1*1 convolution to generate the final module output feature F final ∈R C×H×W, this operation can integrate multi-scale information across different levels, while optimizing the spatial and semantic representations of features, providing more accurate and efficient feature support for subsequent haze removal decoding.

[0034] The inter-layer multi-scale sequence interaction module IMSIM extracts feature representations under different receptive fields through multi-scale convolution, captures long-range dependencies across different levels using sequence modeling, and further optimizes the inter-layer interaction relationship through feature fusion. This module strengthens the integration of edge details and global semantic information, improves the collaborative ability of shallow and deep features, alleviates the feature dilution problem of traditional U-shaped networks, enhances context modeling, provides comprehensive feature support for the haze removal task, and effectively improves the network performance.

[0035] When constructing the Fourier domain frequency-space enhancement module FDFSEM, through the collaborative processing of frequency domain and spatial domain features, the module effectively extracts and strengthens high-frequency information while retaining rich texture details, provides high-quality multi-dimensional feature support for the decoding stage, and significantly improves the haze removal performance.

[0036] In the feature processing stage, the input encoded feature T n ∈R C×H×W is first evenly divided into two parts in the channel dimension, and each part of the feature contains C / 2 channels, which are respectively used for frequency domain processing and spatial domain processing. Among them, T p ∈R C / 2×H×W represents the feature assigned to the frequency domain, and T q ∈R C / 2×H×W represents the feature assigned to the spatial domain. In this way, the feature is divided into two independent components in the channel dimension, laying the foundation for subsequent feature enhancement in the frequency domain and spatial domain respectively. Specifically expressed as:

[0037] T p =T n [C / 2,:,:], T q =T n [C / 2,:,:]

[0038] In the Fourier domain processing, the input encoded feature T p is first transformed to the frequency domain through the two-dimensional fast Fourier transform FFT to obtain the frequency domain feature T freq ∈R C×H×W :

[0039] T freq =FFT(T p )

[0040] The frequency domain feature T freq , contains the real part Re(T freq ) ∈R C / 2×H×W and the imaginary part Im(T freq ) ∈RC / 2×H×W Perform 3*3 convolution operations on the real and imaginary parts of the frequency-domain features respectively to extract high-frequency information, and combine Batch Normalization and ReLU activation functions to optimize the feature representation:

[0041] T freq-enhanced = ReLU(BN(Conv 3×3 (Re(T freq )))) + ReLU(BN(Conv 3×3 (Im(T freq ))))

[0042] Subsequently, the enhanced frequency-domain feature T freq-enhanced ∈R C×H×W is mapped back to the spatial domain through the two-dimensional inverse Fourier transform IFFT, thereby generating an enhanced feature T reconstructed ∈R C / 2×H×W :

[0043] T reconstructed = IFFT(T freq-enhanced )

[0044] where T reconstucted is the output feature in the frequency domain. In this way, high-frequency components and global features are fully exploited to optimize the edges, textures, and global consistency in the image. The spatial domain focuses on extracting local information and edge details of the image. The input encoded feature T q passes through two 3*3 convolution operations in sequence to extract local information, and is further optimized by the BN and ReLU activation functions. The spatial domain retains the edge and texture information in the input feature map, thus providing supplementation for details not covered by the frequency-domain processing. The specific formula is as follows:

[0045] T spatial = ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (T q ))))))

[0046] where T spatial ∈R C / 2×H×W is the output feature in the spatial domain. In this way, the spatial domain provides rich local information for subsequent fusion operations. The enhanced feature T reconsturcted from the frequency domain and the detailed feature T spatial from the spatial domain are fused through the concatenation operation Concat in the channel dimension to generate a fused feature T fused ∈R C×H×W . The formula for the fusion operation is as follows:

[0047] T fused= Concat(T reconstructed , T spatial )

[0048] Among them, it represents the fused feature map. This fusion method combines the global features in the frequency domain and the local details in the spatial domain to generate a richer multi-scale feature representation. The fused feature T fused Compresses the channel dimension through a 1*1 convolution to improve the compactness and representational ability of the features. Immediately following the channel attention module CA, it further optimizes the importance distribution among channels. The channel attention mechanism enhances the contribution of key channels to the image dehazing task by learning the importance weights of different channels. The optimized features are then further compressed in terms of the number of channels through a 1*1 convolution to generate the final output feature T output ∈R C×H×W . The calculation formula is as follows:

[0049] T output = Conv 1×1 (CA(Conv 1×1 (T fused )))

[0050] Among them, T outpout represents the output feature of the Fourier domain frequency-space enhancement module FDFSEM.

[0051] In step S3, the network is trained using the labeled image dataset. First, appropriate loss functions are defined, including L1 loss, perceptual loss, multi-scale structural similarity loss, and adversarial loss, to measure the difference between the model prediction results and the true labels; then the Adam optimizer is selected to adjust the parameters of the network, and the model is continuously optimized through the backpropagation algorithm, thereby improving the dehazing performance and accuracy.

[0052] In step S4, the peak signal-to-noise ratio PSNR and the structural similarity index SSIM are used to evaluate the dehazing effect. PSNR measures the image reconstruction quality, and the higher the value, the higher the similarity between the dehazed image and the real image; SSIM evaluates the structural similarity of the image, and the closer the value is to 1, the closer the structure and visual perception of the generated image are to the real image. These metrics can quantitatively analyze the dehazing effect of the model and be compared with other methods to verify the advantages of the model in terms of detail restoration, clarity, and visual quality. The optimized model will be deployed to the actual application environment to achieve real-time dehazing processing.

[0053] The constructed dehazing network based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, the specific structure of this network is:

[0054] Such as Figure 1As shown, the foggy image with a size of 3×H×W is padded by ReflectionPad2d to preserve the image edge information. Then the image is input into the first encoding layer, using a 7*7 convolutional kernel with a stride of 1, and after normalization and ReLU activation function, an initial feature map with a size of 64×H×W is generated;

[0055] The output of the first encoding layer is connected to the input of the second encoding layer. The initial feature map enters the second encoding layer, and through a 3*3 convolutional kernel with a stride of 2, the downsampling operation is completed, and the size of the feature map is reduced from 64×H×W to 128×H / 2×W / 2. The output of the first encoding layer and the output of the second encoding layer are respectively connected to the input of the first multi-scale sequence interaction module to capture the multi-scale interaction information between texture and semantic features.

[0056] The output of the second encoding layer is connected to the input of the third encoding layer. Through a 3*3 convolutional kernel with a stride of 2, further downsampling is completed, and the size of the feature map is reduced from 128×H / 2×W / 2 to 256×H / 4×W / 4. The output of the third encoding layer and the output of the first multi-scale sequence interaction module are respectively connected to the input of the second multi-scale sequence interaction module to interact the obtained feature maps and further integrate the inter-layer multi-scale information.

[0057] The output of the second multi-scale sequence interaction module is connected to the input of multiple consecutive residual block modules. The residual block modules perform deep semantic modeling on the features to extract high-level features while maintaining the stability of information flow. The output of the residual block modules is connected to the input of the third decoding layer after being fused with the output of the third Fourier domain frequency-space enhancement module.

[0058] The third decoding layer completes the upsampling through a transposed convolution operation, restoring the size of the feature map from 256×H / 4×W / 4 to 128×H / 2×W / 2. At the same time, the output of the third decoding layer is connected to the input of the second decoding layer after being fused with the output of the second Fourier domain frequency-space enhancement module.

[0059] The features are upsampled again to restore the size of the feature map from 128×H / 2×W / 2 to 64×H×W, and the output of the second decoding layer is connected to the input of the first decoding layer after being fused with the output of the first Fourier domain frequency-space enhancement module.

[0060] The first decoding layer performs a final mapping operation. First, it uses ReflectionPad2d to preserve the edge information, then convolves the feature map through a 7*7 convolutional kernel to map the number of channels to the target number of channels, and finally outputs a clear de-fogged image with a resolution of 3×H×W through the Tanh activation function. This process not only realizes the restoration of the resolution but also ensures the preservation of edge details and the naturalness of the output image.

[0061] Among them, the structure of the constructed inter-layer multi-scale sequence interaction module IMSIM is specifically as follows:

[0062] As Figure 2 shown, first, the outputs of the first feature map and the second feature map are respectively connected to the inputs of the first normalization LayerNormalization block and the second normalization Layer Normalization block to standardize the feature distribution, enhance the training stability, and reduce the impact of feature scale differences on subsequent calculations.

[0063] Subsequently, the output feature map of the first normalization Layer Normalization block is input to the first 3×3 convolution block to obtain the first 3×3 convolution block feature map. The first 3×3 convolution block feature map is input to the first 5×5 convolution block to obtain the first 5×5 convolution block feature map. The first 5×5 convolution block feature map is input to the first 7×7 convolution block to obtain the first 7×7 convolution block feature map. The output feature map of the second normalization Layer Normalization is input to the second 3×3 convolution block to obtain the second 3×3 convolution block feature map. The second 3×3 convolution block feature map is input to the second 5×5 convolution block to obtain the second 5×5 convolution block feature map. The second 5×5 convolution block feature map is input to the second 7×7 convolution block to obtain the second 7×7 convolution block feature map. This multi-scale convolution design can capture context semantic information at different scales, making the feature representation more rich and comprehensive.

[0064] For the first 7×7 convolution block feature map and the second 7×7 convolution block feature map, they are respectively input to the first Flatten block and the second Flatten block to convert them from the spatial dimension C×H×W to the sequence form, obtaining the first feature sequence and the second feature sequence with the shape of (H·W)×C. Then, the first feature sequence and the second feature sequence are respectively input to the first mLSTM block and the second mLSTM block to obtain the enhanced first sequence feature and the enhanced second sequence feature. This process can model the context relationship in the spatial and channel dimensions, thereby strengthening the inter-layer information interaction.

[0065] To capture the interaction relationship between the enhanced first sequence feature and the enhanced second sequence feature, first, the enhanced first sequence feature is input to the Transpose block for a transpose operation to obtain the transposed sequence, changing its shape from (H·W)×C to C×(H·W) so that it can perform matrix multiplication operations with the enhanced second sequence feature with the shape of (H·W)×C, calculate the interaction relationship between the two, and input the calculation result to the Softmaax block to obtain the inter-layer attention weight matrix with the shape of C×C. Subsequently, calculate the weighted sum of the inter-layer attention weight matrix and the enhanced first sequence feature to obtain the fused sequence feature, fully capturing the interaction information between the shallow features and the deep features.

[0066] The fused sequence features are input into the Reshape block to restore the spatial dimensions of the original feature map, obtaining a fused feature map with the shape of C×H×W. To further integrate the information in the channel dimension, the fused feature map is input into a 1*1 convolutional block operation to obtain the final output feature with the shape of C×H×W. This operation further optimizes the semantic representation and channel expression ability of the features, providing high-quality feature support for subsequent modules.

[0067] Among them, the structure of the constructed Fourier domain frequency-space enhancement module FDFSEM is specifically as follows:

[0068] As Figure 3 shown, the input feature map of the encoding layer has a size of C×H×W. First, the input feature map of the encoding layer is input into the Split block to obtain a frequency domain feature map and a spatial domain feature map, and the number of channels in each part is C / 2. Among them, the frequency domain feature map is processed in the frequency domain to capture global structural information and high-frequency details; the spatial domain feature map is processed in the spatial domain to strengthen local details and texture information.

[0069] The frequency domain feature map after channel equalization is input into the two-dimensional fast Fourier transform FFT-2D block to obtain a real-imaginary combined frequency domain feature map. The real-imaginary combined frequency domain feature map is connected to the input of the 3*3 convolutional block in the frequency domain. The output of the 3*3 convolutional block in the frequency domain is connected to the input of the batch normalization block BN and the ReLU block in the frequency domain. The outputs of the batch normalization block BN and the ReLU block in the frequency domain are connected to the input of the inverse fast Fourier transform IFFT-2D block. After being processed by the inverse fast Fourier transform IFFT-2D block, an enhanced frequency domain feature map is obtained, and its size remains C / 2×H×W.

[0070] The spatial domain feature map after channel equalization is connected to the input of the first 3*3 convolutional block in the spatial domain. The output of the first 3*3 convolutional block in the spatial domain is connected to the input of the first batch normalization block BN and the first ReLU block in the spatial domain. The outputs of the first batch normalization block BN and the first ReLU block in the spatial domain are connected to the input of the second 3*3 convolutional block in the spatial domain. The output of the second 3*3 convolutional block in the spatial domain is connected to the input of the second batch normalization block BN and the second ReLU block in the spatial domain, obtaining an enhanced spatial domain feature map with the same size of C / 2×H×W.

[0071] The frequency domain feature map and the spatial domain feature map are spliced ​​in the channel dimension to form a frequency-space fusion feature map with a size of C×H×W. This fusion strategy can effectively combine the global modeling capability of the frequency domain with the local optimization capability of the spatial domain, thereby significantly improving the expression effect of high-frequency details. Subsequently, the frequency-space fusion feature map is input into the first 1*1 convolution block to obtain the first 1*1 convolution block feature map to improve the compactness and representation ability of the feature. Next, the first 1*1 convolution block feature map is input into the channel attention module CA to obtain a weighted feature map. Finally, the weighted feature map is input into the second 1*1 convolution block to obtain the final output feature map, and the size remains C×H×W.

[0072] Compared with the prior art, the present invention has the following technical effects:

[0073] 1) In order to overcome the problem of feature information dilution in traditional dehazing networks and enhance the global and local feature expression capabilities of the network, this paper proposes an inter-layer multi-scale sequence interaction module IMSIM. This module connects the features of adjacent layers in the encoder and uses multi-scale convolution and sequence modeling mechanisms to capture cross-level feature dependencies, while fusing multi-scale context information and passing features containing rich semantic and texture information to the decoding layer, thereby effectively alleviating the feature dilution problem and improving the image reconstruction quality;

[0074] 2) The present invention designs a Fourier domain frequency-space enhancement module FDFSEM based on frequency domain and space domain collaborative enhancement. This module combines the frequency domain features extracted by Fourier transform with the local features extracted by spatial convolution, and by collaboratively processing high-frequency details and global structural information, it enhances the network's ability to retain edge information and remove artifacts, effectively improving the accuracy and visual effect of image defogging;

[0075] 3) This paper proposes a defogging network framework based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement for foggy images in complex scenes. By combining cross-level feature dependency modeling and frequency-space information collaborative enhancement, this paper significantly improves the image defogging performance and generalization ability of the network in various scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:

[0077] Figure 1 It is the overall network framework structure diagram of the present invention;

[0078] Figure 2 for Figure 1 The structure diagram of the middle-level multi-scale sequence interaction module IMSIM;

[0079] Figure 3 for Figure 1The structural diagram of the Fourier domain frequency-space enhancement module FDFSEM. Detailed implementation manners

[0080] A haze removal method based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, comprising the following steps:

[0081] Step S1: Collect a publicly available hazy dataset containing different concentrations of haze, and preprocess the images therein to improve data quality and model training effect;

[0082] Step S2: Construct a haze removal network framework based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, the framework comprising a U-shaped haze removal network, an inter-layer multi-scale sequence interaction module IMSIM, and a Fourier domain frequency-space enhancement module FDFSEM;

[0083] Step S3: Use the labeled image dataset to train the network, define a suitable loss function, and use an optimizer for training;

[0084] Step S4: Use two metrics, namely peak signal-to-noise ratio PSNR and structural similarity index SSIM, to evaluate the haze removal effect, and compare with other methods to verify the model performance. Optimize the trained model and deploy it to actual applications for online inference and real-time haze removal.

[0085] The haze removal of the input feature map can be achieved through the above steps.

[0086] Step S1 is specifically as follows:

[0087] Collect a publicly available hazy dataset with different concentrations of haze, and then screen and sort the images; then, preprocess the images in the dataset, including image normalization and data augmentation, so as to improve the generalization ability of the model.

[0088] Step S2 is specifically as follows:

[0089] As Figure 1As shown in the figure, the dehazing network framework based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement includes a U-shaped dehazing network, an inter-layer multi-scale sequence interaction module IMSIM, and a Fourier domain frequency-space enhancement module FDFSEM. Specifically, in the encoding stage, a hazy image is input, and multi-scale features are gradually extracted through the first encoding layer 100, the second encoding layer 101, and the third encoding layer 102. At the same time, in combination with the inter-layer multi-scale sequence interaction module IMSIM, the features of different encoding layers are modeled and interacted to improve the fusion ability of cross-layer information. During this process, the encoding layer not only captures local details but also extracts global context information, providing high-quality features for the decoding stage. In the encoding stage, for the encoding features corresponding to each decoding layer, the Fourier domain frequency-space enhancement module FDFSEM is introduced, and the high-frequency detail expression is strengthened through the collaborative processing of the frequency domain and the spatial domain. The optimized features are directly transmitted to the decoding stage through skip connections to ensure the efficient transmission of multi-scale information. In the decoding stage, the spatial resolution of the image is restored through gradual upsampling. The third decoding layer 103, the second decoding layer 104, and the first decoding layer 105 combine the encoding features with the output of the frequency-space enhancement module, and feature fusion is achieved through residual connections, effectively improving the image detail expression ability. Finally, the network outputs a clear image after dehazing. To optimize the network performance, the network adopts a series of loss functions to constrain the pixel consistency, detail restoration ability, and global feature expression of the dehazed image, thereby learning the mapping relationship between the input hazy image and the target clear image, and significantly improving the dehazing effect and image quality.

[0090] Among them, when constructing the U-shaped network, the features of the input image are extracted layer by layer through quadruple downsampling operations. The first encoding layer 100 completes the preliminary feature extraction, the second encoding layer 101 further extracts features, and the third encoding layer 102 captures deep semantic features and context information. Each encoding layer first performs feature downsampling through convolution operations to reduce the spatial dimension and extract higher-level features, then normalizes the feature map to ensure consistency and stability, and enhances the nonlinear expression ability of the network through the ReLU activation function. The downsampled feature map is further subjected to deep feature extraction through multiple residual modules 6 to capture more abundant semantic information and global context. In the decoding stage, the features extracted by the encoder are gradually restored to the spatial resolution of the image through upsampling operations. Among them, the third decoding layer 103 restores the global features, the second decoding layer 104 further fuses the feature information, and the first decoding layer 105 gradually restores the detail information of the image and restores it to the spatial size of the input image. Through gradual upsampling and feature fusion, a dehazed image is finally output to complete the dehazing task.

[0091] Such as Figure 2As shown, when constructing the inter-layer multi-scale sequence interaction module IMSIM, since different encoding layers of the U-shaped dehazing network carry important information respectively, such as shallow features focusing on capturing detailed texture information, and deep features containing more global semantic information such as color and brightness. Therefore, in order to make full use of the information of features at different levels in the encoding stage, reduce the loss of feature information during transmission, enhance the reconstruction effect of the dehazed image in the subsequent decoding stage, optimize the feature interaction and fusion between layers, and further improve the overall dehazing performance of the network. For the input feature F in the encoding stage k ∈R C×H×W (k ∈ {i, j}), first perform Layer Normalization operation on it to standardize the feature distribution and enhance the training stability. Subsequently, the feature F k is successively passed through convolution operations with three different receptive fields, namely 3*3 convolution, 5*5 convolution, and 7*7 convolution, to gradually extract local and global context information. The overall feature extraction process can be described as:

[0092] F ms,k = Conv 7×7 (Conv 5×5 Conv 3×3 (LN(F k ))), (k ∈ {i, j})

[0093] Among them, F ms,k ∈R C×H×W represents the feature map obtained after multi-scale convolution operations. The successive operations of the three convolution kernels can capture feature information under different receptive fields, thereby strengthening the global semantic representation while retaining edge details. The extracted multi-scale feature F ms,k is transformed into a two-dimensional sequence form F seq,k ∈R (H·W)×C to adapt to the requirements of sequence modeling:

[0094] F seq,k = Flatten(F ms,k ), (k ∈ {i, j})

[0095] The feature sequence is input into the mLSTM module to further optimize the feature representation by capturing the long-range dependencies between sequences, generating the updated sequence feature F' seq,k ∈R (H·W)×C :

[0096] F' seq,k = mLSTM(F seq,k ), k ∈ {i, j}

[0097] Through this operation, the module can effectively extract long-distance dependence information across levels, making the relationship between shallow and deep features closer and laying a foundation for subsequent feature fusion. Then, cross-layer feature relationship calculation and fusion are performed. Calculate the transpose F' of the shallow features from F seq,i ∈R (H·W)×C of the shallow features seq,i T ∈R C×(H·W) and F' seq,j ∈R (H·W)×C the relationship matrix between the deep sequence features:

[0098] W = Softmax(F' seq,i T · F' seq,j )

[0099] The weight matrix W ∈ R C×C is used to measure the interaction relationship between two layers of features and captures complementary information from features of different encoding layers. Using the weight matrix W, the features F' of the shallow path seq,i are weighted and summed to generate the fused features:

[0100] F weighted = F' seq,i · W

[0101] The weighted features F weighted ∈ R (H·W)×C are restored to the original feature map shape F res ∈ R C×H×W through Reshape. The fused features further compress the channel information through 1*1 convolution to generate the final module output features F final ∈ R C×H×W . This operation can integrate multi-scale information across levels, optimize the spatial and semantic representations of features, and provide more accurate and efficient feature support for subsequent defogging decoding.

[0102] The inter-layer multi-scale sequence interaction module IMSIM extracts feature representations under different receptive fields through multi-scale convolution, captures long-distance dependencies across levels using sequence modeling, and further optimizes the inter-layer interaction relationship through feature fusion. This module strengthens the integration of edge details and global semantic information, improves the collaborative ability of shallow and deep features, alleviates the feature dilution problem of traditional U-shaped networks, enhances context modeling, provides comprehensive feature support for the defogging task, and effectively improves the network performance.

[0103] Such as Figure 3As shown, when constructing the Fourier-domain frequency-space enhancement module FDFSEM, through the collaborative processing of frequency-domain and space-domain features, the module effectively extracts and strengthens high-frequency information while retaining rich texture details, providing high-quality multi-dimensional feature support for the decoding stage and significantly improving the defogging performance.

[0104] In the feature processing stage, the input encoded feature T n ∈R C×H×W is first evenly divided into two parts in the channel dimension, with each part of the feature containing C / 2 channels, which are respectively used for frequency-domain processing and space-domain processing. Among them, T p ∈R C / 2×H×W represents the feature assigned to the frequency domain, and T q ∈R C / 2×H×W represents the feature assigned to the space domain. In this way, the feature is divided into two independent components in the channel dimension, laying the foundation for subsequent frequency-domain and space-domain feature enhancement. Specifically expressed as:

[0105] T p = T n [C / 2, :, :], T q = T n [C / 2, :, :]

[0106] In the Fourier-domain processing, the input encoded feature T p is first transformed to the frequency domain through the two-dimensional fast Fourier transform FFT to obtain the frequency-domain feature T freq ∈R C×H×W :

[0107] T freq = FFT(T p )

[0108] The frequency-domain feature T freq , contains the real part Re(T freq ) ∈R C / 2×H×W and the imaginary part Im(T freq ) ∈R C / 2×H×W . 3*3 convolution operations are respectively performed on the real part and the imaginary part of the frequency-domain feature to extract high-frequency information, and the feature expression is optimized by combining Batch Normalization and ReLU activation function:

[0109] T freq-enhanced = ReLU(BN(Conv 3×3 (Re(T freq )))) + ReLU(BN(Conv 3×3 (Im(T freq ))))

[0110] Subsequently, the enhanced frequency-domain feature Tfreq-enhanced ∈R C×H×W Map it back to the spatial domain through two-dimensional inverse Fourier transform IFFT, thereby generating an enhanced feature T containing global structure information reconstructed ∈R C / 2×H×W :

[0111] T reconstructed = IFFT(T freq-enhanced )

[0112] where T reconstucted is the output feature in the frequency domain. In this way, the high-frequency components and global features are fully exploited to optimize the edges, textures, and global consistency in the image. The spatial domain focuses on extracting the local information and edge details of the image. The input encoded feature T q , successively passes through two 3*3 convolution operations to extract local information and is further optimized by the BN and ReLU activation functions. The spatial domain retains the edge and texture information in the input feature map, thus providing supplements for the details not covered by the frequency domain processing. The specific formula is as follows:

[0113] T spatial = ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (T q ))))))

[0114] where T spatial ∈R C / 2×H×W is the output feature in the spatial domain. In this way, the spatial domain provides rich local information for the subsequent fusion operation. The enhanced feature T reconsturcted from the frequency domain and the detail feature T spatial from the spatial domain are fused through the concatenation operation Concat along the channel dimension to generate the fused feature T fused ∈R C×H×W . The formula for the fusion operation is as follows:

[0115] T fused = Concat(T reconstructed ,T spatial )

[0116] where represents the fused feature map. This fusion method combines the global features in the frequency domain and the local details in the spatial domain to generate a richer multi-scale feature representation. The fused feature T fusedCompress the channel dimension through a 1×1 convolution to enhance the compactness and representational ability of features. Immediately following the channel attention module CA, further optimize the importance distribution among channels. The channel attention mechanism enhances the contribution of key channels to the image defogging task by learning the importance weights of different channels. The optimized features are then further compressed in terms of the number of channels through a 1×1 convolution to generate the final output feature T output ∈R C×H×W . The calculation formula is as follows:

[0117] T output =Conv 1×1 (CA(Conv 1×1 (T fused )))

[0118] where T outpout represents the output feature of the Fourier domain frequency-space enhancement module FDFSEM.

[0119] Step S3 is specifically as follows:

[0120] Use the labeled image dataset to train the network. First, define appropriate loss functions, including L1 loss, perceptual loss, multi-scale structural similarity loss, and adversarial loss, to measure the difference between the model prediction result and the true label; then select the Adam optimizer, adjust the parameters of the network, and continuously optimize the model through the backpropagation algorithm, thereby improving the defogging performance and accuracy.

[0121] Step S4 is specifically as follows:

[0122] Use the peak signal-to-noise ratio PSNR and the structural similarity index SSIM to evaluate the defogging effect. PSNR measures the image reconstruction quality, and the higher the value, the higher the similarity between the defogged image and the real image; SSIM evaluates the structural similarity of the image, and the closer the value is to 1, the closer the structure and visual perception of the generated image are to the real image. These metrics can quantitatively analyze the defogging effect of the model and compare it with other methods to verify the advantages of the model in terms of detail restoration, clarity, and visual quality. The optimized model will be deployed to the actual application environment to achieve real-time defogging processing.

[0123] Regarding the network constructed in the present invention, it is specifically as follows:

[0124] In step S2, a defogging network based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement is constructed. The structure of this network is specifically:

[0125] As Figure 1As shown, the foggy image with a size of 3×H×W is padded by ReflectionPad2d to preserve the image edge information. Then the image is input into the first encoding layer 100, using a 7*7 convolutional kernel with a stride of 1, and after normalization and ReLU activation function, an initial feature map with a size of 64×H×W is generated;

[0126] The output of the first encoding layer 100 is connected to the input of the second encoding layer 101. The initial feature map enters the second encoding layer 101, and through a 3*3 convolutional kernel with a stride of 2, the downsampling operation is completed, and the size of the feature map is reduced from 64×H×W to 128×H / 2×W / 2. The outputs of both the first encoding layer 100 and the second encoding layer 101 are respectively connected to the inputs of the first multi-scale sequence interaction module 1 to capture the multi-scale interaction information between texture and semantic features.

[0127] The output of the second encoding layer 101 is connected to the input of the third encoding layer 102. Through a 3*3 convolutional kernel with a stride of 2, further downsampling is completed, and the size of the feature map is reduced from 128×H / 2×W / 2 to 256×H / 4×W / 4. The outputs of both the third encoding layer 102 and the first multi-scale sequence interaction module 1 are respectively connected to the inputs of the second multi-scale sequence interaction module 2 to interact the obtained feature maps and further integrate the inter-layer multi-scale information.

[0128] The output of the second multi-scale sequence interaction module 2 is connected to the inputs of multiple consecutive residual block modules 6. The residual block module 6 performs deep semantic modeling on the features to extract high-level features while maintaining the stability of information flow. The output of the residual block module 6 is connected to the input of the third decoding layer 103 after being fused with the output of the third Fourier domain frequency-space enhancement module 5 in a residual manner.

[0129] The third decoding layer 103 completes the upsampling through a transposed convolution operation, restoring the size of the feature map from 256×H / 4×W / 4 to 128×H / 2×W / 2. At the same time, the output of the third decoding layer 103 is connected to the input of the second decoding layer 104 after being fused with the output of the second Fourier domain frequency-space enhancement module 4 in a residual manner.

[0130] The features are upsampled again to restore the size of the feature map from 128×H / 2×W / 2 to 64×H×W, and the output of the second decoding layer 104 is connected to the input of the first decoding layer 105 after being fused with the output of the first Fourier domain frequency-space enhancement module 3 in a residual manner.

[0131] The first decoding layer 105, through a final mapping operation, first uses ReflectionPad2d for reflection padding to retain edge information, then convolves the feature map with a 7×7 convolutional kernel to map the number of channels to the target number of channels, and finally outputs a dehazed clear image with a resolution of 3×H×W through the Tanh activation function. This process not only restores the resolution but also ensures the retention of edge details and the naturalness of the output image.

[0132] Among them, the specific structure of the constructed inter-layer multi-scale sequence interaction module IMSIM is as follows:

[0133] As Figure 2 shown, first, the outputs of the first feature map 200 and the second feature map 210 are respectively connected to the inputs of the first Layer Normalization block and the second Layer Normalization block to standardize the feature distribution, enhance the training stability, and reduce the impact of feature scale differences on subsequent calculations.

[0134] Subsequently, the output feature map of the first Layer Normalization block is input to the first 3×3 convolutional block to obtain the first 3×3 convolutional block feature map 201. The first 3×3 convolutional block feature map 201 is input to the first 5×5 convolutional block to obtain the first 5×5 convolutional block feature map 202. The first 5×5 convolutional block feature map 202 is input to the first 7×7 convolutional block to obtain the first 7×7 convolutional block feature map 203. The output feature map of the second Layer Normalization block is input to the second 3×3 convolutional block to obtain the second 3×3 convolutional block feature map 211. The second 3×3 convolutional block feature map 211 is input to the second 5×5 convolutional block to obtain the second 5×5 convolutional block feature map 212. The second 5×5 convolutional block feature map 212 is input to the second 7×7 convolutional block to obtain the second 7×7 convolutional block feature map 213. This multi-scale convolutional design can capture context semantic information at different scales, making the feature representation more rich and comprehensive.

[0135] For the first 7×7 convolutional block feature map 203 and the second 7×7 convolutional block feature map 213, they are respectively input to the first Flatten block and the second Flatten block to convert them from the spatial dimension C×H×W to the sequence form, obtaining the first feature sequence 204 and the second feature sequence 114 with the shape of (H·W)×C. Then, the first feature sequence 204 and the second feature sequence 114 are respectively input to the first mLSTM block and the second mLSTM block to obtain the enhanced first sequence feature 205 and the enhanced second sequence feature 215. This process can model the context relationship in the spatial and channel dimensions, thereby strengthening the inter-layer information interaction.

[0136] To capture the interaction relationship between the enhanced first sequence feature 205 and the enhanced second sequence feature 215, first, the enhanced first sequence feature 205 is input into a Transpose block for transposition operation to obtain a transposed sequence 220, which changes from the shape of (H·W)×C to C×(H·W), so that it can perform matrix multiplication operation with the enhanced second sequence feature 215 with the shape of (H·W)×C to calculate the interaction relationship between the two. The calculation result is input into a Softmax block to obtain an inter-layer attention weight matrix 221 with the shape of C×C. Subsequently, the weighted sum of the inter-layer attention weight matrix 221 and the enhanced first sequence feature 205 is calculated to obtain a fused sequence feature 222, fully capturing the interaction information between shallow features and deep features.

[0137] The fused sequence feature 222 is input into a Reshape block to restore the spatial dimension of the original feature map to obtain a fused feature map 223 with the shape of C×H×W. To further integrate the information in the channel dimension, the fused feature map 223 is input into a 1*1 convolution block operation to obtain a final output feature 224 with the shape of C×H×W. This operation further optimizes the semantic representation and channel expression ability of the features, providing high-quality feature support for subsequent modules.

[0138] Among them, the structure of the constructed Fourier domain frequency-space enhancement module FDFSEM is specifically as follows:

[0139] As Figure 3 shown, the input feature map 300 of the encoding layer with the size of C×H×W is first input into a Split block to obtain a frequency-domain feature map 310 and a spatial-domain feature map 320, and the number of channels of each part is C / 2. Among them, the frequency-domain feature map 310 is processed in the frequency domain to capture global structure information and high-frequency details; the spatial-domain feature map 320 is processed in the spatial domain to strengthen local details and texture information.

[0140] The frequency-domain feature map 310 after channel equalization is input into a two-dimensional fast Fourier transform FFT-2D block to obtain a real-imaginary combined frequency-domain feature map 311. The real-imaginary combined frequency-domain feature map 311 is connected to the input of a frequency-domain 3*3 convolution block, the output of the frequency-domain 3*3 convolution block is connected to the input of a frequency-domain batch normalization block BN and a frequency-domain ReLU block, and the output of the frequency-domain batch normalization block BN and the frequency-domain ReLU block is connected to the input of an inverse fast Fourier transform IFFT-2D block. After being processed by the inverse fast Fourier transform IFFT-2D block, an enhanced frequency-domain feature map 312 is obtained, and its size remains C / 2×H×W.

[0141] The spatially domain feature map 320 after channel equalization is connected to the input of the first spatially domain 3*3 convolutional block. The output of the first spatially domain 3*3 convolutional block is connected to the inputs of the first spatially domain batch normalization block BN and the first spatially domain ReLU block. The outputs of the first spatially domain batch normalization block BN and the first spatially domain ReLU block are connected to the input of the second spatially domain 3*3 convolutional block. The output of the second spatially domain 3*3 convolutional block is connected to the inputs of the second spatially domain batch normalization block BN and the second spatially domain ReLU block, obtaining the enhanced spatially domain feature map 321 with the same size of C / 2×H×W.

[0142] The frequency domain feature map 312 and the spatially domain feature map 321 are concatenated in the channel dimension to form the frequency-spatial fusion feature map 330 with the size of C×H×W. This fusion strategy can effectively combine the global modeling ability of the frequency domain and the local optimization ability of the spatial domain, thus significantly improving the expression effect of high-frequency details. Subsequently, the frequency-spatial fusion feature map 330 is input into the first 1*1 convolutional block to obtain the first 1*1 convolutional block feature map 331 to enhance the compactness and representational ability of the features. Then, the first 1*1 convolutional block feature map 331 is input into the channel attention module CA to obtain the weighted feature map 332. Finally, the weighted feature map 332 is input into the second 1*1 convolutional block to obtain the final output feature map 333 with the size remaining C×H×W.

[0143] For the convenience of those of ordinary skill in the art to better understand the present invention, the further description is as follows:

[0144] 1) Parameter setting

[0145] The experiment used an NVIDIA RTX 4090 graphics card with 24GB of video memory. The code was implemented based on the Pytorch framework, and the CUDA version was 11.8. The Adam optimizer was used as the optimization strategy, and the momentum decay exponents were set to β1 = 0.9 and β2 = 0.999. The initial learning rate was 0.001, and the MultiStepLR strategy was used for dynamic adjustment. The batch size was 4. The experiment used two public datasets, SateHaze1k and HRSD, to evaluate the performance of the model. The SateHaze1k dataset consists of three subsets: SateHaze1k-Thin, SateHaze1k-Moderate, and SateHaze1k-Thick. Each subset contains 320 pairs of training images, 35 pairs of validation images, and 45 pairs of test images, covering different scenarios of light, moderate, and heavy haze, aiming to evaluate the effect of the dehazing model under various haze conditions. The HRSD dataset includes two subsets, LHID and DHID. The LHID subset consists of 30,517 images, which are from Google Earth and simulate light fog weather by adding light haze. The DHID subset contains 14,990 images, which are generated from the DLR 3k Munich vehicle aerial photography dataset and are added with dense haze to simulate a thick fog environment, suitable for evaluating the performance of the dehazing algorithm under strong haze conditions. To evaluate the model performance, the present invention adopted two metrics, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), which are common standards for measuring the quality of dehazed images.

[0146] 2) Experimental results

[0147] To verify the effectiveness of the algorithm of the present invention, the present invention was quantitatively compared with current excellent dehazing algorithms, including DCP, DCRD, FCTF, SGID, MAXIM, DCI-Net, and Trinity-Net. The datasets included SateHaze1k and HRSD.

[0148] Table 1 shows the PSNR and SSIM metric results of each dehazing algorithm on the above datasets, and the specific analysis is as follows:

[0149] On the SateHaze1k-Thin dataset with light fog, the PSNR of the algorithm of the present invention reaches 26.74 dB and the SSIM is 0.927, significantly superior to other algorithms. Compared with the second-place DCI-Net, it is improved by 2.27 dB and 0.045 in PSNR and SSIM respectively, demonstrating excellent ability to remove light fog. On the SateHaze1k-Moderate dataset with moderate fog, the method of the present invention reaches 27.65 dB and 0.950 in PSNR and SSIM respectively, also superior to other algorithms. Compared with the closest-performing MAXIM, the method of the present invention shows obvious advantages in detail restoration and global consistency. On the SateHaze1k-Thick dataset with thick fog, the PSNR of the method of the present invention is 24.68 dB and the SSIM is 0.882, and the overall performance is still superior to DCI-Ne and MAXIM, proving its robustness in complex fog scenarios. On the HRSD-LHID dataset, the algorithm of the present invention reaches 28.68 dB and 0.887 in PSNR and SSIM respectively, significantly superior to other comparison algorithms. This performance improvement is mainly due to the inter-layer multi-scale sequence interaction module IMSIM's modeling of cross-layer long-distance dependencies, which can effectively capture special hierarchical features in complex scenarios. On the HRSD-DHID dataset, the algorithm of the present invention reaches 28.23 dB and 0.893 in PSNR and SSIM respectively, exceeding all comparison methods. This indicates that the fog removal method proposed by the present invention has better detail restoration ability in high dynamic range and complex fog scenarios. The algorithm of the present invention performs excellently in various test scenarios, especially in complex thick fog and high dynamic range scenarios. With the multi-scale feature interaction ability of the inter-layer multi-scale sequence interaction module IMSIM and the high-frequency information capture ability of the Fourier domain frequency-space enhancement module FDFSEM, the PSNR and SSIM are significantly improved. Compared with the SOTA method, the algorithm of the present invention shows excellent performance in detail restoration and context information modeling, reaching the SOTA level.

[0150]

[0151] Table 1 Comparison with SOTA methods on SateHaze1k and HRSD datasets

[0152] 3) Ablation experiments

[0153] To evaluate the effectiveness of each module of the invention, we designed ablation experiments according to the framework and the proposed innovations, which include 4 experiments: (1) Base represents the U-shaped basic framework, mainly including an encoding layer, six residual blocks, and a decoding layer, where the encoding and decoding layers are directly connected by skip connections. (2) Base + IMSIM module. (3) Base + FDFSEM module. (4) Base + IMSIM module + FDFSEM module.

[0154]

[0155] Table 2 PSNR and SSIM Results on the SateHaze1k-Moderate Dataset

[0156] As can be seen from Table 2, the Base framework has PSNR and SSIM values of 23.97 dB and 0.844 respectively, demonstrating its basic performance as a benchmark model. By introducing the inter-layer multi-scale sequence interaction module IMSIM, the PSNR of the model is increased to 26.16 dB and the SSIM is improved to 0.891, representing an increase of 2.19 dB and 0.047 respectively compared to the benchmark model. This indicates that the IMSIM module can effectively model the long-range dependencies of cross-level features and enhance the network's comprehensive expression ability for global semantic information and detailed textures. Further, by adding the Fourier domain frequency-space enhancement module FDFSEM to the Base framework, the PSNR and SSIM of the model are respectively increased to 25.39 dB and 0.882, an improvement of 1.42 dB and 0.038 respectively compared to the benchmark model. This result verifies the effectiveness of the FDFSEM module in capturing high-frequency information and enhancing the expression of spatial features, indicating that it can significantly improve the model's adaptability to complex scenarios. When both the IMSIM and FDFSEM modules are introduced simultaneously, the model performance reaches the best level, with the PSNR and SSIM respectively increased to 27.65 dB and 0.950, an improvement of 3.68 dB and 0.106 respectively compared to the benchmark model. Compared with the case of introducing IMSIM or FDFSEM alone, the combined use of the two further enhances the ability of feature interaction and multi-scale feature capture, verifying the synergistic advantages and effectiveness of the modules proposed in the present invention in improving the overall performance of the dehazing network.

Claims

1. A dehazing method based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, characterized in that: The following steps are involved: Step S1: Collect public foggy datasets containing different concentrations of haze and preprocess the images to improve data quality and model training effect; Step S2: construct a defogging network framework based on inter-layer multi-scale sequence interaction and Fourier domain frequency-space enhancement, which includes a U-shaped defogging network, an inter-layer multi-scale sequence interaction module IMSIM and a Fourier domain frequency-space enhancement module FDFSEM; Step S3: Use the labeled image dataset to train the network, define a suitable loss function, and use the optimizer for training; Step S4: Use the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) to evaluate the defogging effect, optimize the trained model, and deploy it in practical applications for online reasoning and real-time defogging; The above steps can achieve dehazing of the input feature map.

2. The method according to claim 1, characterized in that In step S1, a foggy public dataset with different concentrations of haze is collected, and the images are then screened and sorted; then, the images in the dataset are preprocessed, including image normalization and data enhancement, to improve the generalization ability of the model.

3. The method according to claim 2, characterized in that In step S2, when constructing a U-shaped defogging network, it specifically includes: a coding layer feature extraction module, a residual module (6), a decoding layer image restoration module and a third decoding layer; the coding layer feature extraction module includes a first coding layer (100), a second coding layer (101), a third coding layer (102), a residual module (6), and a decoding layer image restoration module includes a third decoding layer (103), a second decoding layer (104), and a first decoding layer (105); When constructing the inter-layer multi-scale sequence interaction module IMSIM, local and global context information in the features of different coding layers is extracted through multi-scale convolution; at the same time, the sequence modeling mechanism mLSTM is used to capture the long-distance dependencies between layers, thereby enhancing the interaction and expression capabilities of multi-scale features; the features output by the module are further optimized by the residual module (6) to generate aggregated features that integrate edge details and semantic information, and finally passed to the third decoding layer (103) to significantly improve the dehazing effect of the network; When constructing the Fourier domain frequency-space enhancement module FDFSEM, in the decoding stage, for the corresponding coding layer, the frequency domain features are extracted by Fourier transform to enhance the high-frequency information, and the spatial domain features are extracted by combining the convolution operation to maintain the texture integrity; after the frequency-space coordinated processing, the optimized features are passed to the corresponding decoding layer to remove artifacts and enhance the structural consistency.

4. The method according to claim 3, characterized in that Specifically, in the encoding stage, a foggy image is input, and multi-scale features are gradually extracted through the first encoding layer (100), the second encoding layer (101) and the third encoding layer (102); at the same time, the inter-layer multi-scale sequence interaction module IMSIM is combined to model and interact the features of different encoding layers, thereby improving the fusion capability of cross-layer information; in this process, the encoding layer not only captures local details, but also extracts global context information, thereby providing high-quality features for the decoding stage; In the encoding stage, for the encoding features corresponding to each decoding layer, a Fourier domain frequency-space enhancement module FDFSEM is introduced to enhance the expression of high-frequency details through the coordinated processing of the frequency domain and the spatial domain; the optimized features are directly transmitted to the decoding stage through jump connections to ensure the efficient transmission of multi-scale information; in the decoding stage, the spatial resolution of the image is restored through step-by-step upsampling; the third decoding layer (103), the second decoding layer (104), and the first decoding layer (105) combine the encoding features with the output of the frequency-space enhancement module, realize feature fusion through residual connections, and effectively improve the image detail expression capability; finally, the network outputs a clear image after defogging; in order to optimize the network performance, the network adopts a series of loss functions to constrain the pixel consistency, detail restoration capability and global feature expression of the defogging image, thereby learning the mapping relationship between the input foggy image and the target clear image, and significantly improving the defogging effect and image quality.

5. The method according to claim 4, characterized in that in, When constructing a U-shaped network, the features of the input image are extracted layer by layer through a four-fold downsampling operation. The first coding layer (100) completes preliminary feature extraction, the second coding layer (101) further extracts features, and the third coding layer (102) captures deep semantic features and context information. Each coding layer first performs feature downsampling through a convolution operation to reduce spatial dimensions and extract higher-level features, and then normalizes the feature map to ensure consistency and stability, and enhances the nonlinear expression ability of the network through a ReLU activation function. The downsampled feature map is further subjected to deep feature extraction through multiple residual modules (6) to capture richer semantic information and global context.

6. The method according to claim 4, characterized in that In the decoding stage, the features extracted by the encoder are gradually restored to the spatial resolution of the image through upsampling operations, wherein the third decoding layer (103) restores the global features, the second decoding layer (104) further fuses the feature information, and the first decoding layer (105) gradually restores the detail information of the image and restores it to the spatial size of the input image. Through layer-by-layer upsampling and feature fusion, a dehazed image is finally output to complete the dehazing task.

7. The method according to any one of claims 4 to 6, characterized in that When constructing the inter-layer multi-scale sequence interaction module IMSIM, since different coding layers of the U-shaped defogging network carry important information, such as shallow features focus on capturing detailed texture information, and deep features contain more global semantic information such as color and brightness; therefore, in order to make full use of the information of different layers of features in the encoding stage and reduce the loss of feature information in the transmission process, while enhancing the reconstruction effect of the defogging image in the subsequent decoding stage, optimizing the feature interaction and fusion between layers, and further improving the overall defogging performance of the network; for the input feature F in the encoding stage k ∈R C×H×W (k∈{i,j}), firstly perform Layer Normalization on it to standardize the feature distribution and enhance the training stability; then, feature F k It is sequentially subjected to three convolution operations with different receptive fields, namely 3*3 convolution, 5*5 convolution, and 7*7 convolution, to gradually extract local and global context information; the overall feature extraction process can be described as: F ms,k =Conv 7×7 (Conv 5×5 Conv 3×3 (LN(F k ))),(k∈{i,j}) Among them, F ms,k ∈R C×H×W represents the feature map obtained after the multi-scale convolution operation; the step-by-step operation of the three convolution kernels can capture the feature information under different receptive fields, thereby strengthening the global semantic representation while retaining the edge details; the extracted multi-scale feature F ms,k is converted into a two-dimensional sequence form F seq,k ∈R (H·W)×C To adapt to sequence modeling needs: F seq,k =Flatten(F ms,k ),(k∈{i,j}) The feature sequence is input into the mLSTM module, which further optimizes the feature representation by capturing the long-distance dependencies between sequences and generates the updated sequence feature F' seq,k ∈R (H·W)×C : F' seq,k =mLSTM(F seq,k ),k∈{i,j} Through this operation, the module can effectively extract cross-layer long-distance dependency information, making the relationship between shallow and deep features closer, laying the foundation for subsequent feature fusion; then, cross-layer feature relationship calculation and fusion are performed; calculation from F' seq,i ∈R (H·W)×C Transpose F' of shallow features seq,i T ∈R C×(H·W) and F' seq,j ∈R (H·W)×C The relationship matrix between deep sequence features: W=Softmax(F' seq,i T ·F' seq,j ) Weight matrix W∈R C×C It is used to measure the interaction between the two layers of features and capture the complementary information from the features of different coding layers. The weight matrix W is used to measure the interaction between the two layers of features. seq,i Perform weighted summation to generate fused features: F weighted =F' seq,i ·W Weighted feature F weighted ∈R (H·W)×C After Reshape, it is restored to the original feature map shape F res ∈R C×H×W The fused features are further compressed through 1*1 convolution to generate the final module output feature F final ∈R C×H×W ,This operation can integrate multi-scale information across layers and optimize the spatial and semantic representation of features, providing more accurate and efficient feature support for subsequent dehazing decoding; The inter-layer multi-scale sequence interaction module IMSIM extracts feature representations under different receptive fields through multi-scale convolution, captures long-distance dependencies across layers using sequence modeling, and further optimizes the inter-layer interaction relationship through feature fusion; this module strengthens the integration of edge details and global semantic information, improves the coordination ability of shallow and deep features, alleviates the feature dilution problem of traditional U-shaped networks, enhances context modeling, provides comprehensive feature support for dehazing tasks, and effectively improves network performance.

8. The method according to any one of claims 1 to 4, characterized in that: When constructing the Fourier domain frequency-space enhancement module FDFSEM, the module effectively extracts and enhances high-frequency information through the coordinated processing of frequency domain and space domain features, while retaining rich texture details, providing high-quality multi-dimensional feature support for the decoding stage and significantly improving the dehazing performance; In the feature processing stage, the input encoding feature T n ∈R C×H×W First, it is divided into two parts in the channel dimension, and each part of the features contains C / 2 channels, which are used for frequency domain processing and spatial domain processing respectively; among them, T p ∈R C / 2×H×W represents the features assigned to the frequency domain, T q ∈R C / 2×H×W Represents the features assigned to the spatial domain; in this way, the features are divided into two independent components in the channel dimension, laying the foundation for subsequent feature enhancement in the frequency domain and spatial domain respectively; specifically expressed as: T p =T n [C / 2,:,:],T q =T n [C / 2,:,:] In Fourier domain processing, the input encoding feature T p First, the two-dimensional fast Fourier transform FFT is used to convert it to the frequency domain to obtain the frequency domain feature T freq ∈R C×H×W : T freq =FFT(T p ) Frequency domain characteristics T freq , including the real part Re(T freq )∈R C / 2×H×W and the imaginary part Im(T freq )∈R C / 2×H×W ; Perform 3*3 convolution operations on the real and imaginary parts of the frequency domain features to extract high-frequency information, and combine Batch Normalization and ReLU activation function to optimize feature expression: T freq-enhanced =ReLU(BN(Conv 3×3 (Re(T freq ))))+ReLU(BN(Conv 3×3 (Im(T freq )))) Then, the enhanced frequency domain features T freq-enhanced ∈R C×H×W Through the two-dimensional inverse Fourier transform IFFT mapping back to the spatial domain, the enhanced feature T containing global structure information is generated. reconstructed ∈R C / 2×H×W : T reconstructed =IFFT(T freq-enhanced ) Among them, T reconstucted is the output feature in the frequency domain; in this way, the high-frequency components and global features are fully exploited to optimize the edges, textures, and global consistency of the image; the spatial domain focuses on extracting local information and edge details of the image; the input encoding feature T q , two 3*3 convolution operations are performed in sequence to extract local information, and further optimized by BN and ReLU activation functions; the spatial domain retains the edge and texture information in the input feature map, thereby providing supplements for the details not covered by the frequency domain processing; the specific formula is as follows: T spatial =ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (T q )))))) Among them, T spatial ∈R C / 2×H×W is the output feature of the spatial domain; in this way, the spatial domain provides rich local information for subsequent fusion operations; the enhanced features T from the frequency domain reconsturcted and the detail features T in the spatial domain spatial Through the concatenation operation Concat fusion of the channel dimension, the fusion feature T is generated. fused ∈R C×H×W ; The formula for the fusion operation is as follows: T fused =Concat(T reconstructed ,T spatial ) Among them, represents the fused feature map; this fusion method combines the global features in the frequency domain and the local details in the spatial domain to generate a richer multi-scale feature representation; the fused feature T fused The channel dimension is compressed through a 1*1 convolution to improve the compactness and representation ability of the feature; followed by the channel attention module CA, which further optimizes the importance distribution between channels; the channel attention mechanism improves the contribution of key channels to the image defogging task by learning the importance weights of different channels; the optimized features are then further compressed through 1*1 convolution to generate the final output feature T output ∈R C×H×W ; The calculation formula is as follows: T output =Conv 1×1 (CA(Conv 1×1 (T fused ))) Among them, T outpout Represents the output characteristics of the Fourier domain frequency-space enhancement module FDFSEM.

9. The method according to claim 1, characterized in that: In step S3, the network is trained using the labeled image dataset. First, appropriate loss functions are defined, including L1 loss, perceptual loss, multi-scale structural similarity loss, and adversarial loss, to measure the difference between the model prediction results and the true labels. Then, the Adam optimizer is selected to adjust the parameters of the network and the model is continuously optimized through the back-propagation algorithm to improve the dehazing performance and accuracy.

10. The method according to claim 1, characterized in that In step S4, the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are used to evaluate the dehazing effect; PSNR measures the quality of image reconstruction, and a higher value indicates a higher similarity between the dehazed image and the real image; SSIM evaluates the structural similarity of images. The closer the value is to 1, the closer the structure and visual perception of the generated image are to the real image. These indicators can quantitatively analyze the dehazing effect of the model and compare it with other methods to verify the advantages of the model in detail recovery, clarity and visual quality. The optimized model will be deployed in actual application environments to achieve real-time dehazing processing.

Citation Information

Patent Citations

  • Image defogging method based on global and local feature fusion

    CN110544213A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on U-Net combined with Transform

    CN117994274A

  • Diffusion model, multi-scale and attention module medical ultrasonic image segmentation method

    CN119180826A

  • Systems and methods for image denoising using deep convolutional networks

    US20230043310A1

Cited By

  • Unmanned aerial vehicle image enhancement method and system

    CN120953059A

  • Image deblurring method and device

    CN121582101A

  • Lightweight image defogging device and method based on noise robustness evaluation and optimization

    CN122473032A