An underwater image enhancement method based on a multi-scale cross-domain fusion network

CN122510115BActive Publication Date: 2026-09-22HOHAI UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202611001744.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-22
Estimated Expiration
2046-07-07

AI Technical Summary

Technical Problem

中国专利申请提出一种带有小波变换和融合注意力机制的水下图像增强方法(公开号CN119784616A,公开日2025.04.08),通过生成器和判别器构建网络,结合复合损失函数训练,虽能提升图像质量和细节还原度,但针对低照度偏色场景,其色彩校正仅依赖全局特征匹配,难以精准应对局部色偏问题

Benefits of technology

1、本发明提出了一种新颖的多尺度跨域融合网络(MSCDF),采用U-Net风格的编码器-解码器架构,在不同尺度级别实现空间-频率交互和融合,精准捕捉水下图像的全局照明偏差、局部散射、色彩失真等多尺度退化特征,显著提升了方法对复杂水下场景的适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510115B_ABST
    Figure CN122510115B_ABST
Patent Text Reader

Abstract

The application discloses an underwater image enhancement method based on a multi-scale cross-domain fusion network, and aims at the problems of color distortion, insufficient contrast, detail loss, poor multi-scale degradation adaptability and the like of existing underwater image enhancement methods. The application proposes a multi-scale cross-domain fusion network, adopts an encoder-decoder architecture in the style of U-Net, realizes bidirectional depth interaction between a spatial domain and a frequency domain through a space-frequency fusion block, introduces a multi-band decomposition unit to perform band-level decomposition and adaptive enhancement on an amplitude spectrum, combines a color correction module to perform global color calibration, effectively retains image structural details and natural color fidelity while eliminating color deviation and improving contrast, realizes a good balance between processing efficiency and enhancement effect, and provides a more effective solution for image restoration in a complex underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an underwater image enhancement method based on a multi-scale cross-domain fusion network, belonging to the field of underwater image enhancement technology. Background Technology

[0002] With the rapid development of underwater imaging technology, underwater images have become an indispensable information carrier in fields such as marine resource exploration, ecological monitoring, and engineering operations. However, due to the selective absorption of light by water and the scattering of suspended particles, underwater images generally suffer from severe degradation phenomena such as color distortion, low contrast, and blurred details. Traditional enhancement methods face many challenges in dealing with these complex degradations, while frequency domain analysis techniques provide a new research direction for simultaneously solving color correction and detail enhancement.

[0003] In traditional image enhancement methods, physics-based approaches estimate scene irradiance and transmission maps by establishing underwater optical imaging models, but these often rely on assumptions of an ideal environment. Non-physics-based methods, such as histogram equalization and Retinex-based methods, directly adjust the image pixel distribution, which can improve overall contrast but is prone to over-enhancement and color distortion. In recent years, with the development and application of deep learning theory, learning-based underwater image enhancement algorithms have also made some progress. However, supervised learning algorithms rely on underwater image / reconstructed image pairs, but high-quality reconstructed images are difficult to obtain. To overcome the limitations of image pairing, related technologies have proposed self-supervised enhancement algorithms, which complete self-supervised training using traditional enhancement algorithms. However, these methods still only focus on adjusting image brightness and do not judge or process color shifts, thus failing to eliminate color casts.

[0004] In recent years, frequency domain analysis methods have attracted widespread attention due to their ability to effectively separate global illumination and local structural features of images. A Chinese patent application (CN119784616A, April 8, 2025) proposes an underwater image enhancement method with wavelet transform and fusion attention mechanisms. While this method improves image quality and detail reproduction by constructing a network using a generator and discriminator and training with a composite loss function, it relies solely on global feature matching for color correction in low-light color-shifting scenarios, making it difficult to accurately address local color shifts. Another Chinese patent application (CN120707394A, September 26, 2025) designs a multi-branch fusion structure based on the Retinex algorithm, including color correction and contrast adjustment branches, which can dynamically adjust fusion weights. However, under low-light conditions, its feature confidence calculation is sensitive to brightness noise, easily leading to over- or under-color correction. Chinese patent application (publication number CN120807323A, publication date 2025.10.17) adopts a dual-domain attention U-Net network, which integrates RGB and NIR modal features to achieve frequency-space feature weighted fusion. Although it can correct color deviation, it relies on multimodal data input and its applicability is limited in low-light color-biased image scenes containing only a single RGB channel.

[0005] However, existing frequency domain methods still have significant limitations: most methods rarely involve frequency band-level operations to recover degradation across multiple frequency bands; they lack targeted design for processing amplitude and phase components; more importantly, the interaction between frequency and spatial domain features often remains at the level of simple stitching or weighted summation, failing to achieve true bidirectional deep fusion. Therefore, frequency domain analysis methods provide a new technical path for underwater image enhancement, but existing methods still have significant shortcomings in terms of scale adaptability, frequency band-level operations, and cross-domain fusion depth. Summary of the Invention

[0006] The technical problem to be solved by this invention is to provide an underwater image enhancement method based on a multi-scale cross-domain fusion network, which provides a more effective solution for image restoration in complex underwater environments by integrating multi-scale cross-domain fusion, bidirectional space-frequency interaction and multi-band decomposition strategies.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: An underwater image enhancement method based on a multi-scale cross-domain fusion network is described below: A multi-scale cross-domain fusion network is constructed to achieve underwater image enhancement tasks; the multi-scale cross-domain fusion network includes an initial feature extraction module, an encoder, a bottleneck layer, a decoder, and a color correction module; The initial feature extraction module performs initial feature extraction on the input raw underwater image; The encoder employs a three-level hierarchical structure, establishing feature encoding results at different levels by progressively downsampling the initial features. The initial features are used as input to the first level of the encoder, and the encoder's second level... The input features of each level are processed sequentially by residual blocks and spatial-frequency fusion blocks to obtain the first level. Hierarchical coding features, the first The encoded features of the hierarchical level are downsampled and used as the first level. Hierarchical input, ; The bottleneck layer aggregates the coding features of the three levels across layers to obtain aggregated features. At the same time, the coding features of the third level are processed by residual blocks and spatial-frequency fusion blocks to obtain deep coding features. The deep coding features and aggregated features are fused to obtain global guiding features. The decoder and encoder have a symmetrical structure. The decoder upsamples and fuses features step by step from the third level based on global guided features to obtain decoded features with the same resolution as the original underwater image. The color correction module performs color correction on the decoded features with the same resolution as the original underwater image to obtain the final enhanced result; The parameters of the multi-scale cross-domain fusion network are optimized by training with a multi-constraint loss function to achieve underwater image enhancement.

[0008] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects: 1. This invention proposes a novel multi-scale cross-domain fusion network (MSCDF), which adopts a U-Net-style encoder-decoder architecture to achieve spatial-frequency interaction and fusion at different scale levels. It accurately captures multi-scale degradation features of underwater images, such as global illumination deviation, local scattering, and color distortion, and significantly improves the adaptability of the method to complex underwater scenes.

[0009] 2. The present invention designs a spatial-frequency fusion block (SFF), which includes a parallel dual-domain feature decoupler and a cross-domain fusion unit. Each dual-domain branch integrates a corresponding attention module to enhance the selectivity of domain-specific features, thereby achieving bidirectional deep fusion of spatial and frequency domain features.

[0010] 3. This invention decomposes the amplitude spectrum into multiple frequency bands through a multi-band decomposition unit (MBD) and performs adaptive weighting on each frequency band. It performs differentiated enhancement for degradation modes in different frequency ranges, captures global illumination patterns in the low-frequency band, and encodes fine textures and structural details in the high-frequency band. This overcomes the limitation of existing frequency domain methods that ignore video band-level operations and enhances the adaptability of information in different frequency bands.

[0011] 4. Based on the multi-scale spatial-frequency bidirectional fusion network, this invention further introduces a color correction module (CC) to carry out global color calibration. It extracts global color statistics through global average pooling and learns channel-level correction weights to reduce underwater color cast, achieving synergistic optimization of color fidelity and contrast enhancement. This effectively eliminates the common blue-green cast problem in underwater images, and the resulting enhancement has a natural visual appearance. Attached Figure Description

[0012] Figure 1 This is an overall structural diagram of the underwater image enhancement method based on multi-scale cross-domain fusion network proposed in this invention; Figure 2 This is a structural diagram of the space-frequency fusion block (SFF) proposed in this invention; Figure 3 This is a structural diagram of the multi-band decomposition unit (MBD) proposed in this invention; Figure 4 This is a visual comparison between the original underwater image and the enhanced result obtained by applying the method of this invention; Figure 5 This is a comparison diagram of the enhancement effects of the method of the present invention and the current state-of-the-art technology, wherein (a) is the input image, (b) is Fusion, (c) is PUIE-Net, (d) is Semi-UIR, (e) is SFGNet, (f) is WF-Diff, (g) is CDF-UIE, (h) is the method of the present invention, and (i) is the reference image. Detailed Implementation

[0013] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0014] like Figure 1 As shown, this invention proposes an underwater image enhancement method based on a multi-scale cross-domain fusion network. This method employs a U-Net-style encoder-decoder architecture, achieving bidirectional depth interaction between the spatial and frequency domains through a spatial-frequency fusion block. Furthermore, a multi-band decomposition unit is embedded in the frequency branch to perform band-level decomposition and adaptive weighting of the amplitude spectrum. Combined with a color correction module, global color calibration is performed, and the enhanced image is output under the guidance of a multi-constraint loss function. The method specifically includes the following steps: Initial feature extraction: Input underwater image of precipitation ,in Corresponding to RGB three channels, and These represent the height and width of the image, respectively. First, a shallow feature extraction is performed on the input image through an initial convolutional layer to obtain an initial feature representation: , in, This indicates the initial convolution operation. , This represents the initial number of feature channels. This step maps the RGB channels of the input image to a high-dimensional feature space, serving as the first layer input for feature encoding.

[0015] Feature Encoding: The encoder employs a hierarchical structure with three scale levels, establishing feature representations at different scales through progressive downsampling. Specifically, in the first... hierarchical ( ), its input features The data is processed sequentially through a residual block (RES) and a spatial-frequency fusion block (SFF). Specifically, the processing procedure for the residual block (RES) is as follows: , in, This indicates a convolution operation with a kernel size of 3×3. These are residual refinement features obtained from residual blocks, containing richer local texture and structural information. Following this, After spatial-frequency fusion (SFF) block processing, the fused coding features of this level are obtained. : , Correspondingly, the feature transfer relationship between the different levels of the encoder can be expressed as: , in, Indicates the encoder at the 1st Hierarchical input, Indicates the first Hierarchical encoding features, functions This indicates a 2x downsampling operation.

[0016] like Figure 2 As shown, the specific processing procedure of the Space-Frequency Fusion Block (SFF) is as follows: The Space-Frequency Fusion Block (SFF) mainly consists of a Space-Frequency Dual-Domain Feature Decoupler (SFDU) and a Cross-Domain Fusion Unit (CDFU). The Space-Frequency Dual-Domain Feature Decoupler (SFDU) includes two branches: one in the spatial domain and one in the frequency domain. (1) In the spatial domain branch, for the received residual refinement features A spatial attention mechanism is employed to capture local texture and spatial relationships, highlighting salient regions and suppressing irrelevant background information, thus obtaining the first... Spatial domain features at different levels : , in, This indicates the average pooling operation. This indicates a maximum pooling operation. Indicates feature concatenation, operators Represents matrix multiplication. Activation function. The sigmoid function is used to calculate spatial attention weights, enabling the enhancement or suppression of different spatial features.

[0017] (2) In the frequency domain branch, the refined features of the received residual. ,like Figure 3 As shown, a Fast Fourier Transform is performed using a Multi-Band Decomposition Unit (MBD). Achieving frequency domain transformation and amplitude-phase separation: , in, and They represent The amplitude spectrum and phase spectrum. Then, the amplitude spectrum... Decomposed into Different frequency bands: , Among them, operators This represents element-wise matrix multiplication. (Symbol) For a bandpass filter, it represents the first... Binarization masks for each frequency band. Specifically, Covering low-frequency areas, Covers high-frequency areas.

[0018] After frequency band decomposition, for each frequency band Convolution and weighted summation are performed to obtain the reconstructed amplitude spectrum. : , in, These are learnable scalar weight parameters. The reconstructed amplitude spectrum... With the original phase spectrum Reassemble and perform inverse Fourier transform. Enhanced features : , Furthermore, regarding features Using a channel attention mechanism, appropriate weights are assigned to different channels to obtain the... Frequency domain features at each scale level : , in, This represents the sigmoid activation function, which outputs channel attention weights. These weights are multiplied by the original feature map to enhance attention along the channel frequency dimension.

[0019] The corresponding features are obtained by calculating through two branches: the spatial domain and the frequency domain. and After that. Regarding the first... Hierarchical structure, using a cross-domain fusion engine (CDFU) to implement spatial domain features. and frequency domain characteristics The two-way interaction and fusion first utilizes weighted combination to reconstruct features: , , in, and These are learnable scalar parameters used to adaptively adjust the contribution strength of cross-domain features. Then, the reconstructed features... and By splicing, we obtain the first... Hierarchical encoding features after fusion : , in, This indicates feature splicing, which complements the features in the spatial and frequency domains.

[0020] Bottleneck layer: After the encoder completes feature extraction at three scale levels, it enters the bottleneck layer to perform multi-scale feature fusion. Specifically, the encoded output features at each scale level of the encoder are fused together. Perform cross-layer aggregation to obtain aggregated features. : , in, This indicates a 4x downsampling. Furthermore, the aggregated features are fused with the deep coding features: , income As a global guiding feature, it provides cross-scale image features for the decoder, making up for the shortcomings of single-scale features in global perception.

[0021] Decoder: The decoder has a symmetrical structure with the encoder, also containing three scale levels, and progressively reconstructs the decoded features starting from the third level, which is the lowest level of the decoder. Specifically, for the lowest level of the decoder, the third level, it will take features from the bottleneck layer. For the sake of consistency in expression, as information at the previous level, the following convention is adopted: ; The decoder in its first In terms of hierarchy, the features are first decoded and output from the previous level. Upsampling is performed to adjust the spatial resolution to the size corresponding to the encoder level. , in, This indicates the transpose convolution operation. This is the upsampling result. Subsequently, the upsampling features... With the encoder Hierarchical coding features The pieces are stitched together to integrate the fine-grained spatial details from the encoding phase: , Among them, the spliced ​​features The data is then processed sequentially through a residual block (RES) for local detail refinement, followed by a spatial-frequency fusion block (SFF) to achieve bidirectional interaction between the spatial and frequency domains, resulting in the decoded output features at this scale. : , Thus, after progressive upsampling and feature fusion at three scale levels, the top-level output of the decoder, i.e., the final decoded features, is obtained. It has the same resolution as the original input image.

[0022] Color Correction Module (CC): Reconstructs a feature image with the original resolution from the decoder. Subsequently, considering the common blue-green bias in underwater images, further color correction was carried out.

[0023] First, extract using global average pooling. Global color information in each channel is used to optimize features: , in, This represents the global average pooling operation, used to compress the feature values ​​of each channel into a 1×1 statistic, thus serving as... Attention weights for each channel.

[0024] Furthermore, the features before and after optimization are weighted to obtain the final enhanced result: , in, These are learnable scalar weight parameters. This is the final output image.

[0025] Multi-constraint loss function optimization training: The network parameters are optimized using a multi-constraint loss function. The total loss function is defined as follows: , in, All are weights of the comprehensive loss function. This indicates a loss of reconstruction accuracy. Indicates frequency consistency loss, This represents the perceived quality VGG loss. Indicates a loss of color fidelity. The specific formulas for representing image contrast loss are as follows: , in, The total number of samples in a batch. This represents the enhanced image. This refers to a labeled image, i.e., a clear underwater image; , in, Indicates the weighting coefficient; , in, This represents the high-level features extracted by the VGG pre-trained network; , in, Indicates the CIEDE2000 standard color distance; , in, It is a contrast measure based on local mean / variance.

[0026] For underwater image enhancement tasks, this invention effectively overcomes degradation problems such as color cast, low contrast, blurred details, and hazy water in the original image. Figure 4 As can be seen, on the one hand, the blue-green bias was significantly corrected, and the natural tones of corals, fish and other objects were realistically restored; on the other hand, the image contrast was greatly improved, and details such as coral textures and biological patterns were clearly presented. At the same time, water scattering and noise were suppressed, and the transparency of the image was enhanced. Overall, the observability and readability of underwater scenes were significantly improved.

[0027] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings: Data preparation: The publicly available UIEB dataset was used, containing 800 pairs of training images, 90 pairs of test images (Test-U90), and 60 unpaired degraded images (Test-C60). The LSUI dataset, containing 4279 pairs of underwater images, was divided into 3864 training pairs and 415 test pairs. Data augmentation strategies such as horizontal / vertical flipping, random cropping, and color dithering were used on the training set to expand sample diversity and improve the model's generalization ability.

[0028] Network parameter configuration: The input image size is adjusted to 256×256, RGB color space input, and both the MSCDF encoder and decoder adopt a three-level scale structure (original scale, 1 / 2 scale, 1 / 4 scale). Each scale level contains a residual block (RES) and a spatial-frequency fusion block (SFF). Downsampling uses a convolution operation with a stride of 2, and upsampling uses a transposed convolution operation.

[0029] In the Spatial-Frequency Fusion Block (SFF) module, the spatial domain branch employs a spatial attention mechanism, generating a spatial weight map through convolutional layers. The frequency domain branch performs multi-band decomposition using a Multi-Band Decomposition Unit (MBD) and employs a channel attention mechanism for adaptive weighting. In the cross-domain fusion unit, scalar parameters can be learned. and Adaptive adjustment of cross-domain contribution.

[0030] The multi-band decomposition unit (MBD) uses 2D Fast Fourier Transform (FFT) for frequency domain transformation, decomposing the amplitude spectrum into... Each frequency band ( (Set to 3), each frequency band is equipped with a dedicated 1×1 convolutional layer for independent processing. Additionally, learnable weight parameters are used. Adaptive weighted fusion of each frequency band is performed, and the frequency bands are converted back to the spatial domain through inverse fast Fourier transform (IFFT).

[0031] The color correction module uses global average pooling to extract global color statistics, learns channel-level color correction weights through convolutional layers, and further utilizes learnable scalar parameters. Features before and after adaptive fusion correction.

[0032] Training process: The optimizer used was AdamW with an initial learning rate of 2e-4, weight decay of 0.005, and gradient clipping maximum norm of 1.0; the training cycle was 200 epochs, with the first 5 epochs being linear warm-up and the last 195 epochs using cosine annealing; the weights of the comprehensive loss function were set to... {5.0, 2.0, 0.5, 1.0, 2.0}.

[0033] Performance Validation: Ninety pairs of real images were extracted from UIEB as the test set, denoted as Test-R90, and quantitative analysis was performed using evaluation metrics. For the Test-R90 dataset, Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Patch Similarity (LPIPS) were used as full-reference evaluation metrics. Lower MSE and higher PSNR indicate that the results are closer to the reference images in terms of image content. Higher SSIM scores indicate that the results are more similar to the reference images in terms of image structure and texture. LPIPS, similar to perceptual loss, measures perceptual similarity by measuring high-level semantic features between images, making it more consistent with human perception when judging image similarity.

[0034] Table 1 shows the comparison results of the method of this invention with other methods on the full reference evaluation indicators (PSNR, SSIM, LPIPS) (bold values ​​are the best, and underlined values ​​are the second best). It can be seen that all indicators of this invention exceed or approach the highest level.

[0035] Table 1 Comparison of Full Reference Image Quality Evaluation Indicators

[0036] Table 2 shows the comparison results of no-reference image quality evaluation indicators such as UIQM (Underwater Image Quality Measurement), UCIQE (Underwater Color Image Quality Evaluation), and URanker. All indicators of the present invention exceed or are close to the optimal values, indicating that the method in this paper has good color calibration capabilities.

[0037] Table 2 Comparison of Quality Evaluation Indicators for No-Reference Images

[0038] Experimental results on benchmark datasets such as UIEB and LSUI show that MSCDF outperforms existing state-of-the-art methods in metrics such as PSNR, SSIM, LPIPS, UIQM, UCIQE, and URanker, especially in color restoration and detail preservation.

[0039] Figure 5In the diagram, (a)-(i) represent comparisons of the enhancement results of the method of this invention with other state-of-the-art technologies. Fusion is a method that utilizes image fusion technology to improve image quality, dynamic range, or information richness. PUIE-Net (Probabilistic U-Net for Image Enhancement) is a neural network model for image enhancement. Semi_UIR (Semi-supervised Underwater Image Restoration) is an underwater image restoration framework based on contrastive semi-supervised learning. SFGNet (Spectral Feature Grouping Network) is a deep learning model for image enhancement tasks. WF-Diff is an underwater image enhancement (UIE) framework based on wavelet transform and diffusion models, short for Wavelet-based Fourier information interaction with frequencyDiffusion adjustment. CDF-UIE (Cross-Domain Fusion for Underwater Image Enhancement) is a deep learning network for underwater image enhancement based on a cross-domain fusion mechanism, aiming to solve the problems of color distortion, low contrast, and blurred details in underwater images. It can be seen that the results of this invention ( Figure 5 The (h) in the image can better restore the color and detail of the image, and obtain enhanced results that are more in line with visual perception.

[0040] Based on the same inventive concept, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned underwater image enhancement method based on a multi-scale cross-domain fusion network.

[0041] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned underwater image enhancement method based on a multi-scale cross-domain fusion network.

[0042] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0043] This invention is described with reference to flowchart illustrations of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each step in the flowchart, and combinations of steps in the flowchart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the steps in the flowchart. Figure 1 A device for a function specified in one or more processes.

[0044] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 The function specified in one or more processes.

[0045] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 Steps of a specified function in one or more processes.

[0046] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.

Claims

1. An underwater image enhancement method based on a multi-scale cross-domain fusion network, characterized in that, The method is as follows: A multi-scale cross-domain fusion network is constructed to achieve underwater image enhancement tasks; the multi-scale cross-domain fusion network includes an initial feature extraction module, an encoder, a bottleneck layer, a decoder, and a color correction module; The initial feature extraction module performs initial feature extraction on the input raw underwater image; The encoder employs a three-level hierarchical structure, establishing feature encoding results at different levels by progressively downsampling the initial features. The initial features are used as input to the first level of the encoder, and the encoder's second level... The input features of each level are processed sequentially by residual blocks and spatial-frequency fusion blocks to obtain the first level. Hierarchical coding features, the first The encoded features of each level are downsampled and used as the first level. Hierarchical input, ; The bottleneck layer aggregates the coding features of the three levels across layers to obtain aggregated features. At the same time, the coding features of the third level are processed by residual blocks and spatial-frequency fusion blocks to obtain deep coding features. The deep coding features and aggregated features are fused to obtain global guiding features. The decoder and encoder have a symmetrical structure. The decoder upsamples and fuses features step by step from the third level based on global guided features to obtain decoded features with the same resolution as the original underwater image. The color correction module performs color correction on the decoded features with the same resolution as the original underwater image to obtain the final enhanced result; The parameters of the multi-scale cross-domain fusion network are optimized by training with a multi-constraint loss function to achieve underwater image enhancement.

2. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 1, characterized in that, The initial feature extraction module includes a convolutional layer, batch normalization, and a ReLU activation function connected in sequence. The initial features extracted by the initial feature extraction module are... Represented as: , in, Represents the ReLU activation function. Indicates batch normalization, This indicates the initial convolution operation. Represents the original underwater image. , , This represents the RGB three channels of the original underwater image. and These are the height and width of the original underwater image, respectively. The initial number of feature channels, Each element of the three-dimensional tensor of the input image is represented.

3. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 1, characterized in that, In the encoder The process for handling hierarchical residual blocks is as follows: , in, This refers to the refined residual features obtained after residual block processing. For encoder number Hierarchical input features Represents the residual block. This indicates a convolution operation with a kernel size of 3×3. Represents the ReLU activation function; After spatial-frequency fusion block processing, the fused coding features of this layer are obtained. : , in, Represents a space-frequency fusion block. Indicates the first Hierarchical coding features; No. The encoded features of the level are downsampled by 2 times and used as the first level. Hierarchical input: , in, For encoder number Hierarchical input features This indicates a 2x downsampling.

4. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 3, characterized in that, The spatial-frequency fusion block includes a spatial-frequency dual-domain feature decoupler and a cross-domain fusion unit, wherein the spatial-frequency dual-domain feature decoupler includes a spatial domain branch and a frequency domain branch; In the encoder Hierarchical, spatial domain branching refines the received residual features. A spatial attention mechanism is used to capture local textures and spatial relationships to obtain the first... Hierarchical spatial domain features : , in, This indicates the average pooling operation. This indicates a maximum pooling operation. Indicates feature splicing, This represents a convolution operation with a kernel size of 7×7. This represents the sigmoid activation function. Represents matrix multiplication; Spatial domain branch refines the received residual features Frequency domain transformation and amplitude-phase separation are achieved by using a multi-band decomposition unit to perform a fast Fourier transform: , in, Represents the Fast Fourier Transform. and They represent The amplitude spectrum and phase spectrum, The imaginary unit; Amplitude spectrum Decomposed into Different frequency bands: , in, express The Each frequency band This represents element-wise matrix multiplication. For bandpass filters, The number of frequency bands obtained from the decomposition; right Convolution of each frequency band and weighted summation are performed to obtain the reconstructed amplitude spectrum. : , in, These are learnable scalar weight parameters. This indicates a convolution operation with a kernel size of 1×1; The reconstructed amplitude spectrum Phase spectrum Recombined and subjected to inverse Fourier transform, the enhanced features are obtained. : , in, Indicates the inverse Fourier transform; Enhanced features Using a channel attention mechanism, assign appropriate weights to different channels to obtain the... Hierarchical frequency domain features : ; In the encoder Hierarchical, cross-domain fusion processors utilize weighted combination to reconstruct features in the feature space domain. and frequency domain characteristics : , , in, and All of them are learnable scalar parameters. and These are the reconstructed spatial domain features and frequency domain features, respectively; for and After concatenation, the result is sequentially processed by 3×3 convolution, batch normalization, and ReLU activation function to obtain the [number of] [units]. Hierarchical coding features : , in, This indicates batch normalization.

5. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 1, characterized in that, The bottleneck layer aggregates the encoded features from the three levels across layers to obtain aggregated features, represented as: , in, As an aggregation feature, Indicates feature splicing, This indicates a 4x downsampling. This indicates a 2x downsampling. This represents a convolution operation with a kernel size of 1×1. and These represent the coding features of levels 1, 2, and 3, respectively. After processing with residual blocks and spatial-frequency fusion blocks, deep coding features are obtained, which are then fused with aggregated features to obtain global guiding features, represented as follows: , in, As a global guiding feature, Represents the residual block. This represents a space-frequency fusion block.

6. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 1, characterized in that, In the decoder Hierarchy, decoding features of the previous level Perform upsampling to adjust the spatial resolution to the size corresponding to the encoder level: , in, For upsampling features, This represents the transpose convolution operation; it transforms the global guided features output by the bottleneck layer. As the decoding feature preceding the third level of the decoder, i.e. ; Upsampled features With encoder number Hierarchical coding features The pieces are stitched together to integrate the fine-grained spatial details from the encoding phase: , in, Features after splicing Indicates feature splicing; Features after splicing After local detail refinement via residual blocks, and bidirectional interaction between the spatial and frequency domains via spatial-frequency fusion blocks, the first... Hierarchical decoding features : , in, Represents the residual block. Represents a space-frequency fusion block. This indicates a convolution operation with a kernel size of 1×1.

7. The underwater image enhancement method based on a multi-scale cross-domain fusion network according to claim 6, characterized in that, Level 1 decoding features That is, to obtain decoding features with the same resolution as the original underwater image, the color correction module extracts them through global average pooling. Feature optimization is performed on the global color information of each channel: , in, For the optimized features, This represents the sigmoid activation function. This represents a convolution operation with a kernel size of 1×1. This indicates a global average pooling operation; Optimized features and By weighting, we obtain the final enhanced result: , in, These are learnable scalar weight parameters. To ultimately enhance the results.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the underwater image enhancement method based on a multi-scale cross-domain fusion network as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater image enhancement method based on a multi-scale cross-domain fusion network as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Underwater image enhancement method with wavelet transform and fusion attention mechanism

    CN119784616A

  • Underwater image enhancement method based on Retinex algorithm

    CN120707394A

  • Underwater image enhancement method based on double-domain attention U-Net

    CN120807323A

  • Underwater image enhancement method based on graph structure and space-frequency cooperation

    CN122199345A

  • KR20210096926A