A lightweight underwater image enhancement method and system based on low-frequency suppression and space-frequency domain interaction

CN122675666APending Publication Date: 2026-09-01贵州电子科技职业学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610802947.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

具体表现为:基于像素调整的方法容易导致色彩偏移或过度增强;基于物理成像模型的方法依赖难以准确估计的参数;现有空频域结合的Transformer方法无法有效抑制低频退化对高频细节表达的干扰

Benefits of technology

[0026] (1) Effectively suppresses low-frequency degradation interference and improves detail recovery capability. This invention attenuates the low-frequency amplitude components in the central region of the spectrum in the Fourier domain through a low-frequency suppression module, which can specifically weaken the low-frequency redundant response caused by backscattering and uneven illumination, and reduce the fogging effect and gray screen interference commonly found in underwater images. Since this process maintains the phase information unchanged, it can better preserve the structural contour and detailed texture information of the scene while suppressing degradation components, providing more effective discrimination information for subsequent feature extraction and reconstruction processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122675666A_ABST
    Figure CN122675666A_ABST
Patent Text Reader

Abstract

The application discloses a kind of light weight underwater image enhancement method and system based on low frequency suppression and space-frequency domain interaction, belong to image processing technical field.The method includes: shallow feature is extracted by encoder and multi-stage feature extraction is carried out using stacked low-frequency suppression based on transformer block;In the low-frequency suppression transformer block, the input feature is processed in frequency domain by low-frequency suppression module, and the low-frequency degradation response in the center region of frequency spectrum is suppressed;Feature interaction is carried out by low-frequency suppression driven self-attention module, and low-frequency suppression operation is introduced into query matrix and key matrix;Feature reconstruction is carried out by decoder;In the attention module of low-frequency suppression optimization in skip connection, the skip feature is screened and recalibrated;fusion output enhances the underwater image.The application effectively suppresses the low-frequency degradation interference in underwater image, improves the texture detail recovery ability and color fidelity, the parameter quantity is 1.75M, the calculation complexity is 11.47G, the PSNR reaches 25.126dB, the SSIM reaches 0.954, with light weight and efficient characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically relating to a lightweight underwater image enhancement method and system based on low-frequency suppression and spatial-frequency domain interaction. Background Technology

[0002] Light underwater imaging technology is increasingly used in marine biodiversity surveys, infrastructure inspections, and search and rescue. However, the underwater environment presents complex optical degradation problems: as water depth increases, light absorption causes the red, orange, yellow, and green light components to disappear in sequence, while scattering causes background light interference and energy diffusion, resulting in visual degradation phenomena such as reduced contrast, haze, and noise, which seriously affect the performance of underwater applications.

[0003] Traditional underwater image enhancement methods are mainly divided into two categories: pixel-based methods and physical imaging model-based methods. Pixel-based methods indiscriminately correct pixel values, often leading to color shifts or over-enhancement. Physical imaging model-based methods attempt to separate the direct transmission component from the scattering component, offering greater physical interpretability, but their performance typically relies on several idealized assumptions and parameters that are difficult to accurately estimate in real-world scenarios.

[0004] In recent years, deep learning-based image enhancement methods have learned complex degradation patterns from large-scale data, achieving more robust enhancement results across different scenes and degradation types. Traditional U-Net and CNN architectures effectively improve local detail recovery capabilities through multi-scale features and skip connections. Since Transformer can model long-range dependencies and is more helpful in recovering global color, some studies have incorporated Transformer and CNN into the U-Net framework to achieve co-modeling of local and global features.

[0005] Furthermore, since frequency domain analysis helps reveal fine details (high-frequency components) and global background and contours (low-frequency components), methods such as Spectroformer and Phaseformer fuse spatial and frequency domain information to improve image visibility, color accuracy, and contrast. However, existing Transformer methods that combine spatial and frequency domains ignore the inherent imbalance between the degradation of low-frequency and high-frequency components during underwater imaging, where low-frequency components are often more severely degraded. In severely degraded scenarios, this imbalance is further exacerbated, thus hindering the full exploitation of frequency domain representation. Specifically, low-frequency fogging and veil effects in underwater images can obscure high-frequency structural information, leading to insufficient detail recovery and inadequate color fidelity.

[0006] Therefore, effectively suppressing low-frequency degradation interference in underwater images while preserving high-frequency texture details has become a key technical issue in improving the performance of underwater image enhancement. Summary of the Invention

[0007] The technical problem this invention aims to solve is that existing underwater image enhancement methods fail to fully exploit frequency domain information and ignore the inherent imbalance between low-frequency and high-frequency component degradation during underwater imaging. Low-frequency fogging and veil effects in underwater images severely obscure high-frequency structural information, leading to shortcomings in detail recovery, color fidelity, and overall visual quality in existing methods. Specifically, pixel-based methods are prone to color shifts or over-enhancement; methods based on physical imaging models rely on parameters that are difficult to estimate accurately; and existing spatial-frequency domain combined Transformer methods cannot effectively suppress the interference of low-frequency degradation on high-frequency detail representation.

[0008] To address the aforementioned technical problems, this invention proposes a lightweight underwater image enhancement method and system based on low-frequency suppression and spatial-frequency domain interaction. The core idea of ​​this invention is to reshape the feature energy distribution in the Fourier domain, suppress low-frequency masking effects through controllable frequency domain filtering, and integrate the low-frequency suppression mechanism into the self-attention module and skip-connection attention module of the Transformer architecture, achieving synergistic enhancement of global modeling and local compensation.

[0009] This invention provides an underwater image enhancement method based on low-frequency suppression and spatial-frequency domain interaction, comprising the following steps:

[0010] S1: Encoding Stage. The water depth image to be enhanced is acquired, and shallow features are extracted using an encoder. The encoder employs a stacked low-frequency suppression-based Transformer block for multi-stage feature extraction, with downsampling used between stages to reduce spatial dimensionality. The encoder first uses a 3×3 convolution to extract shallow features, then performs local and global representation extraction using a four-stage stacked low-frequency suppression-based Transformer block. Pixel inverse rearrangement downsampling is used between stages to reduce the spatial dimensionality of the image. The low-frequency suppression-based Transformer block adopts an architecture based on pre-normalization and a dual residual structure, consisting of a low-frequency suppression-driven self-attention module and a feedforward network. Information fidelity is achieved through residual connections outside each sub-layer.

[0011] S2: Low-frequency suppression module processing. In the low-frequency suppression-based Transformer block, the input features are processed in the frequency domain by the low-frequency suppression module. The low-frequency suppression module performs the following operations: performs a two-dimensional Fourier transform on each channel to map the spatial domain features to the frequency domain; moves the low-frequency components to the center of the spectrum through a spectrum centering operation; constructs a radially symmetric suppression mask to attenuate the amplitude components of the central low-frequency region while maintaining the phase information; and performs inverse centering and inverse Fourier transform on the suppressed spectrum to return the features from the frequency domain to the spatial domain.

[0012] The formula for the suppression mask is:

[0013]

[0014] in, Indicates frequency domain position Normalized radial distance to the center of the spectrum. The low-frequency retention intensity of the control center The suppression range is controlled. This mask is multiplied point-by-point with the amplitude spectrum to obtain the suppressed amplitude components. The suppressed spectrum can be expressed as:

[0015]

[0016] in Indicates the spectral amplitude. This represents the spectral phase. Through the above design, the low-frequency suppression module can specifically weaken the low-frequency redundant response caused by backscattering and uneven illumination in the frequency domain, thereby mitigating the fogging effect and gray screen interference commonly found in underwater images, while better preserving the structural contours and detailed texture information of the scene.

[0017] S3: Low-frequency suppression-driven self-attention module processing. In the low-frequency suppression-based Transformer block, feature interaction is performed through a low-frequency suppression-driven self-attention module. This module introduces the low-frequency suppression operation into the query matrix and key matrix respectively, while keeping the value matrix unchanged. Attention computation is used to achieve a synergistic enhancement of global modeling and local compensation.

[0018] The low-frequency suppression-driven self-attention module employs a multi-head attention approach. Specifically, it performs a 1×1 convolutional linear mapping on the normalized features, then models the local spatial context using a 3×3 depthwise separable convolution to generate a query matrix. Key matrix Sum matrix Perform low-frequency suppression operations on the query matrix and key matrix respectively to obtain... and The value matrix remains unchanged; for and After L2 normalization, the attention matrix is ​​calculated, and a learnable temperature parameter is introduced to adjust the response intensity of each attention head. The value matrix is ​​then weighted and aggregated using the attention matrix. This design enables the attention weight generation process to actively reduce degradation-related low-frequency interference, allowing the attention mechanism to focus more on discriminative structural contours and texture details, while ensuring that the overall brightness distribution and edge contours are fully transmitted during the weighted aggregation process.

[0019] S4: Decoding Stage. The decoder reconstructs the features output by the encoder. The decoder employs stacked low-frequency suppression-based Transformer blocks for multi-stage feature processing, with upsampling between each stage to increase the spatial dimension. The decoder uses a symmetrical design; the final output of the encoder passes through three stacked low-frequency suppression-based Transformer blocks to capture rich frequency domain information, and pixel rearrangement upsampling is performed before each stage to increase the spatial dimension of the pixels.

[0020] S5: Low-frequency suppression optimized attention module processing. In the skip connections between the encoder and decoder, the skip features transmitted by the encoder are filtered and recalibrated by a low-frequency suppression optimized attention module. This module first weakens the low-frequency degradation response in the central region of the spectrum using the low-frequency suppression module, then extracts channel-level statistical descriptions through global average pooling, and generates channel weights by modeling cross-channel dependencies using one-dimensional convolution to recalibrate the skip features. Since the weight estimation is driven by the low-frequency suppressed features, the low-frequency suppression optimized attention module can more accurately distinguish the importance of effective and degraded information in different channels, thereby suppressing invalid responses while preserving shallow details and providing more favorable fusion features for the decoder.

[0021] S6: Fusion Output. The recalibrated skip features are fused with the decoder features, and the enhanced underwater image is output through convolution. At the s-th scale, the decoder features are upsampled with pixel rearrangement to obtain features with the same resolution as the encoder skip features. At the same time, the encoder skip features are effectively extracted through the attention module optimized by low-frequency suppression. Then, these two features are concatenated along the channel dimension, and a 1×1 convolution is used to maintain the consistency of the number of channels. Finally, a 3×3 convolution is used to obtain the enhanced underwater image.

[0022] This invention also provides an underwater image enhancement system based on low-frequency suppression and spatial-frequency domain interaction, comprising: an encoding module for acquiring the underwater downgraded image to be enhanced, extracting shallow features, and performing multi-stage feature extraction using stacked low-frequency suppression-based Transformer blocks; a low-frequency suppression module for performing frequency domain processing on the input features; a low-frequency suppression-driven self-attention module for performing feature interaction within the low-frequency suppression-based Transformer blocks; a decoding module for reconstructing the features output by the encoding module; a low-frequency suppression-optimized attention module for filtering and recalibrating skip features in skip connections between the encoding and decoding modules; and a fusion output module for fusing the recalibrated skip features with the features from the decoding module and outputting the enhanced underwater image.

[0023] Furthermore, this invention also includes a phased loss optimization strategy: a joint loss consisting of Charbonnier loss, gradient loss, MS-SSIM loss, and perceptual loss is used for supervision. In the early stages of training, adaptive weights are used to learn the importance of each loss term. In the later stages of training, the weights are fixed based on the optimal peak signal-to-noise ratio (PSNR) result on the validation set. The adaptive weights are learned parameters mapped to normalized weights using a softmax function, enabling the model to automatically learn a more reasonable loss combination in the early stages of training and maintain weight stability in the later stages to avoid continuous fluctuations negatively impacting convergence stability.

[0024] Beneficial effects

[0025] The present invention has the following advantages over the prior art:

[0026] (1) Effectively suppresses low-frequency degradation interference and improves detail recovery capability. This invention attenuates the low-frequency amplitude components in the central region of the spectrum in the Fourier domain through a low-frequency suppression module, which can specifically weaken the low-frequency redundant response caused by backscattering and uneven illumination, and reduce the fogging effect and gray screen interference commonly found in underwater images. Since this process maintains the phase information unchanged, it can better preserve the structural contour and detailed texture information of the scene while suppressing degradation components, providing more effective discrimination information for subsequent feature extraction and reconstruction processes.

[0027] (2) Enhance image contrast and restore color fidelity. This invention integrates a low-frequency suppression mechanism into the self-attention module, enabling the attention weight generation process to actively weaken low-frequency interference dominated by degradation, allowing the attention mechanism to focus more on discriminative structural contours and texture details. Simultaneously, the value branch retains the original feature response, ensuring that the overall brightness distribution and edge contours are fully transmitted during the weighted aggregation process. This design achieves enhanced image contrast and more accurate color correction and fidelity, while also improving the modeling ability for high-frequency discriminative details.

[0028] (3) Promote multi-scale feature fusion and improve the efficiency of structural information transmission. This invention embeds a low-frequency suppression optimized attention module in the skip connection to adaptively filter and recalibrate the features passed from the encoder to the decoder. Since the weight estimation is driven by the low-frequency suppressed features, this module can more accurately distinguish the importance of effective information and degenerate information in different channels, thereby suppressing invalid responses while preserving shallow details and promoting the effective transmission of structural information and detailed textures during multi-scale feature fusion.

[0029] (4) Lightweight design reduces computational complexity. This invention adopts a lightweight Transformer architecture with 1.75M parameters and a computational complexity of 11.47G FLOPs, ranking best among all compared methods. While ensuring enhanced performance, it reduces the number of model parameters and computational complexity, demonstrating excellent lightweight characteristics, which is beneficial for practical deployment and application.

[0030] (5) Improved overall enhancement performance. Experimental results show that the peak signal-to-noise ratio (PSNR) of the present invention reaches 25.126 dB and the structural similarity index (SSIM) reaches 0.954 on the UIEB dataset, which are approximately 0.695 dB and 0.021 dB higher than the existing best methods, respectively; the learned perceptual image patch similarity (LPIPS) is 0.073 and the depth image structure and texture similarity (DISTS) is 0.080, both achieving the best results. On the U45 and UCCS real datasets, the present invention outperforms the comparison methods on most unreferenced evaluation metrics, demonstrating strong competitiveness and cross-scene generalization ability. Attached Figure Description

[0031] Figure 1 This is a schematic diagram illustrating the principle of low-frequency suppression for underwater image reconstruction. It shows that the local texture and detail response in the reconstructed image are significantly enhanced after the degraded underwater image is processed by low-frequency suppression, and the principle of achieving a balance between detail enhancement and overall visual naturalness through complementary fusion of the original image and the low-frequency suppressed image.

[0032] Figure 2 This is the overall network architecture diagram, demonstrating the complete structure of the proposed space-frequency domain interactive lightweight Transformer architecture. The architecture includes core components such as an encoder, decoder, Low-Frequency Suppression Transformer Block (LFSTB), Low-Frequency Suppression Module (LFS), Low-Frequency Suppression Driven Self-Attention Module (LFS-SA), and Low-Frequency Suppression Optimized Attention Module (OLCA). The encoder uses a four-stage stacked LFSTB for feature extraction, and the decoder uses a three-stage stacked LFSTB for feature reconstruction. The encoder and decoder are connected via the OLCA module.

[0033] Figure 3This image compares the enhancement effects of different methods on the UIEB dataset, showcasing the subjective visual effects of the proposed method compared to 10 other methods, including WWPE, WFAC, UNTV, ICSP, Water-Net, PUGAN, DGD-cGAN, Spectroformer, Phaseformer, and PyUIE, on the UIEB synthetic dataset. The results show that the proposed method can better enhance the hierarchical relationship between the cave entrance area and the surrounding rocks, restore clearer local textures, and make the color representation and overall tone of the coral area closer to the reference image.

[0034] Figure 4 This is a comparison chart of the enhancement effects of different methods on the U45 and UCCS datasets, demonstrating the subjective visual effect comparison between the method of this invention and the 10 comparative methods mentioned above on real underwater datasets. The results show that the method of this invention can effectively remove green or blue color casts, restore a more natural color distribution, and significantly enhance the texture details of areas such as shells and divers, improving image contrast and clarity while maintaining overall brightness and structural consistency. Detailed Implementation

[0035] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings.

[0036] Example 1: Overall Network Architecture

[0037] This embodiment details the overall design of the lightweight Transformer architecture with spatial-frequency domain interaction proposed in this invention. For example... Figure 2 As shown, given a color image of a waterway descent with height H and width W. The goal of this invention is to generate enhanced images with improved visibility, corrected color distortion, and restored structural details. The entire enhancement process is divided into two stages: encoding and decoding.

[0038] Encoding stage: The encoder first uses a 3×3 convolution to extract shallow features. , where C represents the initial number of channels. After four stages of stacked Low Frequency Suppression Transformer Block (LFSTB)-based local and global representation extraction, pixel unshuffle downsampling is used between each stage to reduce the spatial dimension of the image. Pixel unshuffle splits and reassembles the high-resolution image into the channel dimension at a fixed ratio, achieving parameter-free spatial downsampling. This preserves the integrity of spatial information while avoiding the information loss caused by traditional pooling or convolution downsampling.

[0039] Decoding Stage: The decoder employs a symmetrical design. The encoder's final output undergoes a three-stage stacked LFSTB to capture rich frequency domain information, with pixel shuffle upsampling performed before each stage to increase the spatial dimension of the pixels. Pixel shuffle is the inverse operation of pixel inverse shuffle, rearranging channel-dimensional information into the spatial dimension to achieve efficient spatial upsampling.

[0040] Skip Connections: To efficiently transfer information features from the encoder to the corresponding decoder, this invention introduces a Low-Frequency Suppression Optimized Attention Module (OLCA) into the encoder-decoder connection. At the s-th scale, the encoder skip features are represented as follows: The decoder features are represented as First of all, Perform pixel rearrangement upsampling to obtain the same as Features with the same resolution At the same time, through OLCA... Effective feature extraction was performed to obtain Then, these two features are concatenated along the channel dimension, and a 1×1 convolution kernel is used to maintain the consistency of the number of channels, which facilitates subsequent decoding. Finally, the transmitted features of the codec are obtained. :

[0041]

[0042] At the end of the decoder, a 3×3 convolution is used to obtain the enhanced underwater image. .

[0043] Example 2: Low-Frequency Suppression Transformer Block (LFSTB) Structure

[0044] This embodiment details the internal structure of LFSTB. LFSTB employs a Transformer-based architecture with pre-normalization and dual residual structures, consisting of a low-frequency suppression-driven self-attention module (LFS-SA) and a feedforward network (FFN). Given the input tensor of the i-th LFSTB module... The LFSTB processing flow is as follows:

[0045] First, LFS-SA is used to suppress low-frequency degradation and preserve structural details in the frequency domain, while performing feature interactions in the spatial dimension:

[0046]

[0047] in The representation layer is normalized. Then, local textures and details are further recovered using FFN:

[0048]

[0049] By introducing residual connections outside each sub-layer, a stable optimization process and high-fidelity information transfer are achieved. This dual-residual structure effectively alleviates the gradient vanishing problem while ensuring the layer-by-layer accumulation and transfer of feature information, providing rich multi-scale representations for subsequent reconstruction processes.

[0050] Example 3: Low Frequency Suppression Module (LFS)

[0051] This embodiment details the design principle and implementation of the low-frequency suppression module. To reduce low-frequency fogging and gray-screen effects in underwater images, and to mitigate the influence of low frequencies on the structural and detail information characterized by high-frequency components, this invention designs a low-frequency suppression module based on the Fourier domain.

[0052] Let the input features be First, a two-dimensional Fourier transform is performed on each channel to map the spatial domain features to the frequency domain. In the default arrangement of the discrete spectrum, low-frequency components are located at the four corners. To facilitate explicit modeling of the low-frequency region, this invention further moves the low-frequency components to the center of the spectrum through a spectrum centering operation. Its frequency domain features can be expressed as:

[0053]

[0054] in Represents two-dimensional frequency coordinates. This represents a two-dimensional Fast Fourier Transform. This represents the spectrum centering operation. For ease of description, the centered spectrum can be represented in polar coordinates for amplitude and phase:

[0055]

[0056] in Indicates the spectral amplitude. This represents the spectral phase. This invention does not simply correlate amplitude with low frequencies and phase with high frequencies, but rather suppresses the central low-frequency region at the frequency position, and manifests this as modulation of the amplitude component in the complex representation.

[0057] To address the low-frequency response in the central region, this invention constructs a radially symmetric suppression mask. :

[0058]

[0059] in Indicates frequency domain position The normalized radial distance to the center of the spectrum is used to characterize the positional relationship of the current frequency component relative to the low-frequency center; The intensity of low-frequency retention in the control center; Controlling the suppression range. In this embodiment, and The values ​​are 0.5 and 0.1 respectively. This mask is multiplied point-by-point with the amplitude spectrum to obtain the suppressed amplitude components, where... This indicates element-wise multiplication. Therefore, the suppressed spectrum can be expressed as:

[0060]

[0061] This method attenuates the low-frequency amplitude response while preserving the phase information. Finally, inverse centering and inverse Fourier transform are performed on the suppressed spectrum to return the features from the frequency domain to the spatial domain, yielding the spatial domain features. :

[0062]

[0063] in This represents the two-dimensional inverse fast Fourier transform. This indicates a spectrum inverse centralization operation. This indicates taking the real part.

[0064] Through the above design, LFS can specifically attenuate the low-frequency redundant response caused by backscattering and uneven illumination in the frequency domain, thereby mitigating the fogging effect and gray screen interference commonly found in underwater images. Since this process mainly attenuates the amplitude component in the central region of the spectrum while maintaining phase information, it can better preserve the structural contours and detailed texture information of the scene while suppressing degenerative components. The resulting features have clearer edge representation and higher contrast, providing more effective discriminative information for subsequent feature extraction and reconstruction processes.

[0065] Example 4: Low-frequency suppression driven self-attention module (LFS-SA)

[0066] This embodiment details the design principle and implementation of the low-frequency suppression-driven self-attention module. In the proposed low-frequency suppression-based Transformer block, the features modulated by LFS are further used for multi-head self-attention modeling, enabling each attention head to focus more on discriminative structural contours and texture details, thereby forming an LFS-SA module with selective suppression capability for degenerate low-frequency components.

[0067] In LFS-SA, given a normalized tensor First, a 1×1 convolution is used to linearly map the query, and then a 3×3 depthwise separable convolution is used to further model the local spatial context, thereby generating the query. ,key Sum Features. Unlike conventional attention, to mitigate the low-frequency redundant response caused by backscattering and uneven illumination during underwater degradation, this invention further... and A low-frequency suppression (LFS) operation is introduced into the branch, while The branches remain unchanged, and then the branches are rearranged in a multi-head manner. This process can be represented as:

[0068]

[0069] in Indicates the number of heads of attention. This represents the channel dimension corresponding to each attention head. To stabilize attention computation and improve the robustness of the similarity metric, [the following is used]: and L2 normalization was performed on the last dimension.

[0070] Based on this, make the query matrix Bond projection matrix The dot product results in a matrix of size 1. The attention matrix is ​​used, and a learnable temperature parameter is introduced. Adaptive adjustment of the response intensity of each attention head:

[0071]

[0072] Finally, the value features are weighted and aggregated using the attention matrix, and the multi-head results are reassembled and mapped through a 1×1 convolution to obtain the output of the LFS-SA module. .

[0073] Since low-frequency suppression only applies to the query and key branches, the attention weight generation process can proactively weaken degradation-related low-frequency interference, allowing the attention mechanism to focus more on discriminative structural contours and texture details. Meanwhile, the value branch retains the original feature response, ensuring that the overall brightness distribution and edge contours are fully transmitted during weighted aggregation. This design is equivalent to using discriminative features after low-frequency suppression to guide the attention distribution, thereby enhancing details while preserving the integrity of the original information. Thus, this module can globally model cross-channel relationships while reducing the adverse effects of low-frequency degradation components on attention learning, thereby enhancing image contrast and restoring color fidelity.

[0074] In this embodiment, the number of attention heads for each LFS-SA module is set to 8, and the initial number of channels is set to 16.

[0075] Example 5: Low-Frequency Suppression Optimized Attention Module (OLCA)

[0076] This embodiment details the design principles and implementation of the low-frequency suppression optimized attention module. Between the encoder and decoder, skip connections are responsible for transmitting shallow features to supplement the edge, texture, and local details required in the decoding stage. However, in underwater scenes, these features often contain not only effective details but also low-frequency redundant responses caused by backscattering and lighting perturbations. If these features are directly fused with decoder features without filtering, these degenerate components will also be introduced into the subsequent reconstruction process, affecting feature utilization efficiency.

[0077] Based on this, the present invention embeds an attention module optimized by low-frequency suppression (OLCA) into the skip connection. For the skip feature of the s-th scale transmitted from the encoder... OLCA first weakens the low-frequency degradation response in the central region of the spectrum using a low-frequency suppression module (LFS), thereby highlighting effective components related to structural contours and texture details. Subsequently, channel-level statistical descriptions are extracted using global average pooling (GAP), and local cross-channel dependencies are modeled using one-dimensional convolution to generate channel weights to recalibrate the original skip features. The formula for OLCA is:

[0078]

[0079] Since the weight estimation is driven by the features after low-frequency suppression, the OLCA module can more accurately distinguish the importance of effective information and degraded information in different channels, thereby suppressing invalid responses while preserving shallow details and providing more favorable fusion features for the decoder.

[0080] Example 6: Loss Function and Staged Optimization Strategy

[0081] This embodiment details the design of the loss function and the phased optimization strategy. To balance pixel-level reconstruction, structure preservation, and perceptual quality, this invention employs Charbonnier loss, gradient loss, MS-SSIM loss, and perceptual loss jointly to supervise the network during the training phase. The total loss function is defined as:

[0082]

[0083] in These are the adaptive weights for each loss term. The definitions and functions of each loss function are as follows:

[0084] Charbonnier loss is used to enhance images with pixel-level constraints. Compared with reference image The difference lies in the fact that, compared to traditional L1 loss, Charbonnier loss exhibits better robustness, reducing reconstruction error while preventing abnormal pixels from excessively impacting the training process. It is defined as follows:

[0085]

[0086] in Indicates pixel position, Represents the total number of pixels in the image. This loss is used to maintain numerical stability. It primarily improves the overall reconstruction accuracy of the image and enhances the recovery of basic texture and brightness.

[0087] Gradient loss is used to constrain the consistency of the enhanced image and the reference image in terms of edge and local texture structure. Since gradient information effectively reflects high-frequency details and contour changes in an image, introducing gradient loss helps enhance edge fidelity and reduce detail blurring. Its calculation formula is as follows:

[0088]

[0089] in and These represent the gradient operators of the image in the horizontal and vertical directions, respectively.

[0090] MS-SSIM loss is used to measure the consistency between the enhanced image and the reference image from the perspective of multi-scale structural similarity. Unlike simple pixel-level loss, MS-SSIM focuses more on the overall similarity of the image at the brightness, contrast, and structural levels, thus better constraining the structural fidelity of the enhancement result. It is defined as follows:

[0091]

[0092] This loss helps improve the model's ability to recover global structural information and enhances the visual consistency of reconstruction results at different scales.

[0093] Perceptual loss imposes higher-level visual perceptual constraints on the network by comparing the differences in high-level semantic representations between the enhanced image and the reference image in the pre-trained feature space. Compared to relying solely on pixel-level supervision, perceptual loss better preserves the semantic structure and natural visual effect of the image. Its expression is:

[0094]

[0095] in This indicates that the pre-trained feature extraction network is in the first... Feature maps extracted from layers, , , These represent the number of channels, height, and width of the feature map at this layer, respectively. This loss, by constraining the consistency of the high-level feature distribution, makes the enhancement result more natural in overall visual perception and helps alleviate the oversmoothing problem.

[0096] Adaptive Weight Allocation Strategy: Due to differences in the numerical scale and optimization objective of different loss terms, directly using fixed weights for linear combination often fails to yield stable and effective optimization results. Therefore, this invention proposes an adaptive loss weight allocation strategy. To ensure the weights are non-negative and satisfy normalization constraints, this invention introduces learnable parameters. The softmax function is then used to map these weights to normalized weights.

[0097]

[0098] In the early stages of training, the loss weights and network parameters are updated together, enabling the model to automatically learn more reasonable loss combinations. To avoid instability caused by continuous weight fluctuations in the later stages of training, this invention further introduces a weight freezing strategy based on the validation set PSNR. After each epoch, the PSNR is calculated on the validation set, and the corresponding loss weights are recorded, retaining only the 5 weights with the best PSNR. When the validation set PSNR no longer improves for several consecutive epochs, the average of these 5 best weights is taken as the final fixed weights, which remain unchanged in subsequent training, with only the network parameters being updated. This strategy combines the adaptive learning capability in the early stages with the stable optimization process in the later stages, helping to improve the robustness of model training and the final recovery performance.

[0099] Example 7: Training Parameter Configuration

[0100] This embodiment details the training parameter configuration of the model. The proposed model is implemented using Python 3.7 and PyTorch 1.13, and all experiments were conducted on a PC equipped with an NVIDIA RTX A6000 GPU.

[0101] During training, the batch size was set to 16, the patch size to 256, and the Adam optimizer was used for 700 epochs. The initial learning rate was set to... The learning rate strategy employs a hold-and-decay approach, meaning it remains constant for the first 400 epochs and then gradually decreases to near zero over the next 300 epochs. Random horizontal and vertical flipping is used as a data augmentation strategy, and all images are normalized to the range [-1, 1] before being input into the model.

[0102] This invention uses the UIEB dataset, which contains degraded underwater images and corresponding reference images, for training. The UIEB dataset contains 890 underwater images, covering a wide range of scenes including corals, barriers, and marine life. The corresponding reference images are generated by 12 image augmentation methods. Following the partitioning method of Underwater Ranker, 800 pairs of training samples and 90 pairs of test samples were obtained.

[0103] With the above training configuration, the model of this invention has 1.75M parameters and a computational complexity of 11.47G FLOPs, which is the best among all the comparison methods, demonstrating good lightweight characteristics.

[0104] The above description is merely a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. All equivalent technical solutions made based on the concept of the present invention should be covered within the scope of protection of the present invention.

Claims

1. An underwater image enhancement method based on low-frequency suppression and spatial-frequency domain interaction, characterized in that, Includes the following steps: S1: Obtain the water downgrading image to be enhanced, extract shallow features through an encoder. The encoder uses stacked low-frequency suppression Transformer blocks for multi-stage feature extraction, and downsampling is used between each stage to reduce the spatial dimension. S2: In the low-frequency suppression Transformer block, the input features are processed in the frequency domain by the low-frequency suppression module. The low-frequency suppression module performs the following: performing a two-dimensional Fourier transform on each channel to map the spatial domain features to the frequency domain; moving the low-frequency components to the center of the spectrum through a spectrum centering operation; constructing a radially symmetric suppression mask to attenuate the amplitude components of the central low-frequency region while keeping the phase information unchanged; and performing inverse centering and inverse Fourier transform on the suppressed spectrum to return the features from the frequency domain to the spatial domain. S3: In the low-frequency suppression-based Transformer block, feature interaction is performed through a low-frequency suppression-driven self-attention module. The low-frequency suppression-driven self-attention module introduces the low-frequency suppression operation into the query matrix and the key matrix respectively, while keeping the value matrix unchanged. The attention calculation achieves the synergistic enhancement of global modeling and local compensation. S4: The features output by the encoder are reconstructed by the decoder. The decoder uses stacked low-frequency suppression Transformer blocks for multi-stage feature processing, and upsampling is used between each stage to increase the spatial dimension. S5: In the skip connection between the encoder and the decoder, the skip features transmitted by the encoder are filtered and recalibrated by the low-frequency suppression optimized attention module. The low-frequency suppression optimized attention module first weakens the low-frequency degradation response in the center region of the spectrum through the low-frequency suppression module, then extracts the channel-level statistical description through global average pooling, and generates channel weights by modeling cross-channel dependencies through one-dimensional convolution to recalibrate the skip features. S6: The recalibrated jump features are fused with the decoder features, and the enhanced underwater image is output through convolution.

2. The method according to claim 1, characterized in that, The low-frequency suppression-based Transformer block adopts an architecture based on pre-normalization and dual residual structure, including a low-frequency suppression-driven self-attention module and a feedforward network, and achieves information fidelity transmission through residual connections outside each sub-layer.

3. The method according to claim 1, characterized in that, The formula for the suppression mask is: in, Indicates frequency domain position Normalized radial distance to the center of the spectrum. The low-frequency retention intensity of the control center Control the range of inhibition.

4. The method according to claim 1, characterized in that, The low-frequency suppression-driven self-attention module adopts a multi-head attention approach. It performs linear mapping and depthwise separable convolution on the normalized features to generate a query matrix, a key matrix, and a value matrix. After L2 normalization of the query matrix and the key matrix, it calculates the attention matrix. It introduces a learnable temperature parameter to adjust the response intensity of each attention head, and uses the attention matrix to perform weighted aggregation on the value matrix.

5. The method according to claim 1, characterized in that, The encoder comprises a four-stage stacked low-frequency suppression-based Transformer block, and the decoder comprises a three-stage stacked low-frequency suppression-based Transformer block. The method is characterized by employing a pixel-inverse rearrangement method for downsampling and a pixel rearrangement method for upsampling. It also includes a staged loss optimization strategy, using a joint loss of Charbonnier loss, gradient loss, MS-SSIM loss, and perceptual loss for supervision. In the early stages of training, the importance of each loss term is learned through adaptive weights, and in the later stages of training, the weights are fixed based on the optimal peak signal-to-noise ratio result of the validation set. The adaptive weights are characterized by using learnable parameters mapped to normalized weights via a softmax function.

6. An underwater image enhancement system based on low-frequency suppression and spatial-frequency domain interaction, characterized in that, include: The encoding module is used to acquire the water downgrading image to be enhanced, extract shallow features, and use stacked low-frequency suppression-based Transformer blocks for multi-stage feature extraction. Downsampling is used between each stage to reduce the spatial dimension. The low-frequency suppression module is used to process the input features in the frequency domain. It performs the following steps: performing a two-dimensional Fourier transform on each channel to map the spatial domain features to the frequency domain; moving the low-frequency components to the center of the spectrum through a spectrum centering operation; constructing a radially symmetric suppression mask to attenuate the amplitude components of the central low-frequency region while keeping the phase information unchanged; and performing inverse centering and inverse Fourier transform on the suppressed spectrum to bring the features back from the frequency domain to the spatial domain. A low-frequency suppression-driven self-attention module is used to perform feature interaction in the low-frequency suppression-based Transformer block. The low-frequency suppression module's operation is introduced into the query matrix and key matrix respectively, while the value matrix remains unchanged. The attention calculation achieves a synergistic enhancement of global modeling and local compensation. The decoding module is used to reconstruct the features output by the encoding module. It adopts a stacked low-frequency suppression-based Transformer block for multi-stage feature processing, and upsampling is used between each stage to increase the spatial dimension. The low-frequency suppression optimized attention module is used to filter and recalibrate the skip features passed by the encoding module in the skip connection between the encoding module and the decoding module. First, the low-frequency suppression module weakens the low-frequency degradation response in the center region of the spectrum. Then, channel-level statistical descriptions are extracted by global average pooling. Channel weights are generated by modeling cross-channel dependencies through one-dimensional convolution to recalibrate the skip features. The fusion output module is used to fuse the recalibrated jump features with the features from the decoding module, and output the enhanced underwater image through convolution.

7. The system according to claim 6, characterized in that, The low-frequency suppression-based Transformer block adopts an architecture based on pre-normalization and dual residual structure, including a low-frequency suppression-driven self-attention module and a feedforward network, and achieves information fidelity transmission through residual connections outside each sub-layer.

8. The system according to claim 6, characterized in that, The formula for the suppression mask is: in, Indicates frequency domain position Normalized radial distance to the center of the spectrum. The low-frequency retention intensity of the control center Control the range of inhibition.

9. The system according to claim 6, characterized in that, The low-frequency suppression-driven self-attention module adopts a multi-head attention approach. It performs linear mapping and depthwise separable convolution on the normalized features to generate a query matrix, a key matrix, and a value matrix. After L2 normalization of the query matrix and the key matrix, it calculates the attention matrix. It introduces a learnable temperature parameter to adjust the response intensity of each attention head, and uses the attention matrix to perform weighted aggregation on the value matrix.

10. The system according to claim 6, characterized in that, The encoding module includes four stacked low-frequency suppression-based Transformer blocks, and the decoding module includes three stacked low-frequency suppression-based Transformer blocks. The downsampling uses a pixel inverse rearrangement method, and the upsampling uses a pixel rearrangement method. The system also includes a staged loss optimization module, which uses a joint loss of Charbonnier loss, gradient loss, MS-SSIM loss, and perceptual loss for supervision. In the early stage of training, the importance of each loss term is learned through adaptive weights. In the later stage of training, the weights are fixed based on the optimal peak signal-to-noise ratio result of the validation set. The adaptive weights are normalized weights mapped to learnable parameters through the softmax function.