Dual-guided underwater image multi-scale enhancement method

By introducing a depth estimation subnetwork and a dynamic multi-scale attention model, combined with a brightness inversion module, the problem of unbalanced global and local information processing in underwater image enhancement is solved, and efficient enhancement of underwater images is achieved.

CN120707798AActive Publication Date: 2025-09-26QINGDAO UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510669142.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-26
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods have difficulty in simultaneously adapting to dynamic degradation, local feature extraction, and global brightness balance when processing complex underwater scenes, resulting in poor underwater vision system performance.

Method used

A dual-guided multi-scale enhancement method for underwater images is adopted. The scene depth information is obtained through the depth estimation sub-network DBEDNet. Combined with the dynamic multi-scale attention model DMAM and the brightness inversion module RBM, attention weights are dynamically allocated, multi-scale feature fusion and brightness compensation are performed, and multi-level residual units and loss functions are designed to improve image quality.

Benefits of technology

It effectively solves the problems of blurred details and unclear edges in underwater images, realizes dynamic adjustment of global structure and local details, improves the overall brightness balance and visual effect of the image, and is suitable for different underwater scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707798A_ABST
    Figure CN120707798A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and digital image processing, and particularly relates to a dual-guidance underwater image multi-scale enhancement method, in particular to an underwater image color cast correction, contrast improvement and detail recovery method combining a dynamic multi-scale attention mechanism, depth perception and a brightness reversal fusion technology. The method is suitable for image preprocessing and quality optimization of scenes such as underwater robot navigation, marine resource exploration and underwater monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of computer vision and digital image processing, and specifically relates to a dual-guided multi-scale enhancement method for underwater images. Background Art

[0002] Underwater images are affected by factors such as water absorption, scattering effects and uneven illumination, and generally suffer from problems such as color cast, blur, low contrast and loss of local details, which seriously restrict the application effect of underwater vision systems.

[0003] Existing deep learning-based underwater image enhancement methods (such as UWCNN, WaterNet, and UIE-DAL) improve image quality to a certain extent by designing specific network structures or loss functions. UWCNN enhances visual effects through color correction and contrast stretching, significantly improving the overall visual quality of underwater images, especially in correcting color casts. However, it struggles to adapt to the dynamic degradation of complex underwater scenes. WaterNet uses physical models to guide network training, enhancing its ability to model complex degradation processes and improve the algorithm's generalization. However, it makes insufficient use of depth information, resulting in limited structural recovery capabilities. UIE-DAL introduces an attention mechanism to enhance local features, strengthening local feature extraction capabilities and enabling enhanced optimization of specific areas. However, it lacks a multi-scale dynamic weight allocation mechanism, resulting in poor results in the coordinated optimization of global brightness balance and detail recovery. Summary of the Invention

[0004] Based on the above problems, this application introduces a dynamic multi-scale module to adaptively allocate attention weights to better compensate for the shortcomings of the static model. Its technical solution is: A dual-guided underwater image multi-scale enhancement method, the steps are as follows: S1. Acquire underwater images and perform preprocessing; S2. Construct a depth estimation subnetwork DBEDNet based on the Encoder-Decoder framework, which introduces spatial structure information and can estimate scene depth information from underwater images; S3. Use the backbone network to further extract high-level image features and combine them with depth information for enhanced processing; S4. Using the brightness reversal map R generated earlier, the detail recovery capability of the low-brightness area of ​​the image is enhanced; S5. Using the dynamic multi-scale attention model DMAM, after fusing multi-scale features with depth guidance information, it dynamically assigns importance weights of different scales to different regions; S6. Reconstruction of fused feature map; S7. Design a loss function and output the predicted underwater image.

[0005] Preferably, the pre-processing in step S1 includes: Normalization processing: normalize the original pixel values ​​of the input underwater image I to the interval [0,1]; Brightness inversion generation: Design a brightness inversion mapping function for the dark areas in the image : ; By inverting the original image to obtain the brightness reversal map R, the network can focus on the recovery of details in low-brightness areas in the subsequent process; Convolutional preprocessing layer: After the input image is normalized and brightness inverted, the initial feature extraction layer is set, which can be expressed as: ; in, is the convolution kernel parameter, is the bias term, then Represents the initially extracted low-level features, and its size is consistent with the input.

[0006] Preferably, the depth estimation subnetwork DBEDNet in step S2 includes an encoder and a decoder; The encoder adopts a multi-level convolution stacking structure. Assuming that the encoder has multiple layers, the processing process of each layer is described as follows: ; ; PReLU is the activation function, represents the convolution kernel weight matrix of the encoder layer i, represents the bias term of the encoder layer i; A dual attention module (DAM) is introduced in the encoder to capture inter-channel and spatial dependencies. Its structure includes a self-attention mechanism and a position attention mechanism: In specific implementation, channel attention is first obtained through global average pooling, and then spatial attention is obtained through local convolution; The output of DAM is fused with the standard convolutional features, and the formula is as follows: ; in, and represent channel attention and spatial attention functions respectively, is the fusion weight; The decoder uses upsampling operation to gradually restore the feature map to its original size. Let each layer of the decoder be: ; in represents the upsampling convolution operation, represents the convolution kernel weight matrix of the jth layer of the decoder, Represents the bias item of the jth layer of the decoder; in each decoding stage, in order to maintain the encoder information, a skip connection is used to directly fuse the corresponding encoding layer features with the decoding layer features.

[0007] Preferably, the last layer of the decoder outputs a single-channel depth map D through convolution mapping. The depth map can approximately describe the spatial distance information of each area in the image. The formula is as follows: ; in, Represents the Sigmoid activation function, which constrains the value to be between [0,1], Represents the convolution kernel weight matrix for depth estimation, represents the output feature map of the last layer of the decoder, represents the bias term for depth estimation; the final depth map D will serve as the auxiliary input of the subsequent dynamic multi-scale attention model (DMAM) to guide the multi-scale attention mechanism.

[0008] Preferably, the backbone network in step S3 includes primary feature extraction, multi-scale feature fusion, and feature and depth information fusion; The primary feature extraction formula is as follows: ; in, is the original feature map, and These are the first and second convolution layers, respectively. The feature map size of the primary feature extraction output is consistent with the input; Multi-scale feature fusion: convolution kernels of different scales (such as 3×3, 5×5, etc.) are applied to the same input, and features of different receptive fields are extracted and fused through multi-branch parallel convolution; there are multiple scale branches , and its calculation formula is: ; Indicates the The bias term of the scale branch is a vector used to offset the output of each neuron after the convolution operation, increasing the flexibility and fitting ability of the model. Indicates the The convolution kernel weight matrix of the scale branch; it is used to Convolution is performed to extract features at that scale. The size and number of convolution kernels depend on the network design. The fusion operation uses channel concatenation followed by 1×1 convolution for dimensionality reduction, followed by activation processing to generate a fused multi-scale feature map. ; Feature and depth information fusion: Assume that the processed depth feature is , then the feature fusion operation is: ; in, It means that the feature maps are summed or concatenated channel by channel, and the channels are adjusted through fused convolution.

[0009] Preferably, step S4 is specifically implemented as follows: S41. The input brightness inversion map R is processed by an independent convolution branch at the same time. The branch is designed as follows: ; in, represents the convolution kernel weight matrix in the independent convolution branch that processes the brightness inversion map R, Represents the bias term in the convolution branch. After multiple layers of convolution, the output feature map of the branch is fused layer by layer with the backbone features. Brightness inversion features after convolution extraction With backbone features Hierarchical fusion: Specifically, at key locations such as the residual block, the end of the coding layer, and the entrance of the DMAM module, and Perform element-by-element summation, the formula is: ; in, is the fusion coefficient.

[0010] Preferably, the dynamic multi-scale attention DMAM model includes multi-branch attention calculation, dynamic learning of attention weights, and feature reweighting and fusion; Multi-branch attention calculation: Input features After being processed by multiple convolution branches of different scales, each branch uses a convolution kernel of different sizes to capture contextual information in different ranges; the output of each scale branch is recorded as , the calculation formula is: ; Represents the convolution kernel weight matrix of the i-th scale branch in the Dynamic Multi-Scale Attention Module (DMAM). It is used to Perform convolution operations to capture features at different scales), (representing the bias term in the scale branch), and obtaining their respective response maps by calculating the features of different scales; Dynamic learning of attention weights: After splicing the multi-scale response maps, the attention weight is learned through global average pooling and full connection layer. The spliced ​​feature map is recorded as , whose attention vector The calculation steps are as follows: ; in, is the Sigmoid activation function, and the output weight vector is in the interval [0,1]; (represents the weight matrix in attention weight learning. It is used to linearly transform the features after global average pooling (GAP) to generate the attention vector), (represents the feature map after splicing the multi-scale response map. It contains feature information from branches of different scales and is used to calculate the attention weight), (represents the bias term in attention weight learning).

[0011] Feature reweighting and fusion: Each branch feature is weighted by attention After weighting, channel splicing and 1×1 convolution fusion are used to generate a dynamically adjusted feature map. , the fusion formula is expressed as: ; in, is the corresponding branch attention weight, (convolution kernel weight matrix for fusion operation), (feature map of the i-th scale branch), (Bias term for the fusion operation).

[0012] Preferably, the image output by the dynamic multi-scale attention DMAM model is subjected to smoothing processing, including reshape, reverse and region smoothing; Reshape processing: The feature map output by the dynamic multi-scale attention DMAM model is re-blocked, and the feature map is divided into regions of fixed size to ensure that each region has a certain local consistency; the sub-block form after division is recorded as ; Reverse processing: Perform reverse processing on the features within each sub-block, that is, use the brightness inversion features generated in the early stage Compensate the area and apply the inverse mapping to each pixel within the subblock: ; in, To adjust the parameters, is the compensation term reconstructed based on the regional brightness inversion relationship; Region Smoothing Processing: Apply regional smoothing (such as mean filtering or bilateral filtering) to each processed sub-block, eliminate inter-block mutations through multiple iterations, and finally reorganize the sub-blocks into an overall feature map .

[0013] Preferably, the reconstruction of the fused feature map in step S6 is as follows: A multi-level residual unit is designed to perform nonlinear mapping and detail recovery on the fusion features, using the formula: ; (represents the input feature map of the kth layer. It contains the feature information passed from the previous layer), (Represents the weight matrix of the first convolution operation of the kth layer. Used to input feature map Perform convolution operation), (Represents the bias term of the first convolution operation in the kth layer. It is used to adjust the output after the convolution operation).

[0014] Then we pass the second convolution layer to get: ; (represents the feature map of the kth layer after the first convolution operation and activation function. It is obtained by converting the input feature map obtained by adding the residual connection), (Represents the weight matrix of the second convolution operation of the kth layer. Used to Perform further convolution operations), Represents the bias term of the second convolution operation of the kth layer, which is used to adjust the output after the convolution operation; The multi-level residual design can retain the initial image feature information and refine the enhancement effect layer by layer; The reconstruction module finally uses a 3×3 convolutional mapping layer to map the feature map output by the residual unit into a three-channel image. The formula is as follows: ; The Sigmoid activation function constrains the output to be within [0,1], and the final output is the final fusion feature map, (represents the weight matrix of the final convolutional mapping layer. It is used to map the feature map output by the residual unit into a three-channel image), (represents the final feature map after all processing), (denoting the bias term of the final convolutional map layer).

[0015] Preferably, step S7 designs a loss function and outputs the predicted underwater image. The specific method is as follows: Using mean square error loss To measure the output image With the original clear image The gap between Use the pre-trained VGG network to extract high-level features of the image and construct perceptual loss in the feature space: ; Among them, N (represents the number of feature layers in the VGG network used to calculate the perceptual loss), Represents the feature map of the jth layer of the VGG network. This loss helps to improve the high-level semantic consistency and visual realism of the image; Design structural similarity index SSIM loss: ;

[0016] This loss encourages the output image to maintain high consistency with the original image in local structure and texture; Attention Constrained Loss: Let the weight vector To meet certain sparsity and smoothness constraints, it can be designed Loss or KL divergence: ; in (represents the preset target attention vector. It is an ideal attention weight distribution used to guide the model to learn a more reasonable attention allocation) can be set according to the preset target distribution; Total loss function: The total loss function of each module is defined as: ; or, ; in, is the weight of each loss.

[0017] Compared with the prior art, this application has the following beneficial effects: 1. Through the depth estimation subnetwork and multi-scale attention mechanism, the system can not only capture global structural information but also dynamically adjust and enhance local details, effectively solving the problems of blurred details and unclear edges in underwater images.

[0018] 2. Utilizing inverted brightness image information, we perform specialized brightness inversion fusion to address typical issues of uneven underwater lighting and insufficient detail in dark areas. This ensures a more balanced overall brightness distribution in the image without introducing overexposure or darkening.

[0019] 3. Each module uses residual learning and skip connections to ensure gradient propagation and prevent network degradation. Multiple loss functions are designed to ensure stable model performance across various test datasets. Ablation experiments validate the performance improvements achieved by the DMAM, RBM, and R³S modules, demonstrating the applicability of this invention to diverse underwater scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the overall architecture of DBED-Net; Figure 2 This is the DMAM structure diagram; Figure 3 This is a comparison chart of the effects. DETAILED DESCRIPTION

[0021] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description of a frequency domain detection method and system for colored noise environment signals based on spectrum envelope extraction proposed in accordance with the present invention, in conjunction with the accompanying drawings and specific embodiments.

[0022] The aforementioned and other technical contents, features, and effects of the present invention are clearly presented in the following detailed description of the specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a deeper and more specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are provided for reference and illustration purposes only and are not intended to limit the technical solutions of the present invention.

[0023] like Figure 1 As shown in Figure 1, the overall framework of this system can be divided into an input preprocessing module, a depth estimation subnetwork (DBEDNet), a backbone feature extraction module, an RBM (Reverse Brightness Mask) enhancement module, a DMAM (Dynamic Multi-scale Attention Module), an R³S (Reshape-Reverse-Region Smoothing) module, and a final reconstruction module. The entire network uses an end-to-end training approach to achieve a direct mapping of underwater imagery from input to output. The system has the following features: 1. Depth perception guidance: Use the depth estimation sub-network to obtain scene structure information and guide the main network to perform feature enhancement.

[0024] 2. Multi-scale attention: Dynamically capture regional features at different scales through DMAM, so that local details and global information can be processed simultaneously.

[0025] 3. Brightness inversion fusion: Use the RBM module to enhance and restore dark areas, making up for the shortcomings of traditional methods in terms of uneven brightness.

[0026] 4. Residual and multi-layer fusion: A multi-level residual module and R³S module are used to ensure feature transfer and detail smoothing, achieving accurate reconstruction.

[0027] A dual-guided underwater image multi-scale enhancement method, the steps are as follows: S1. Acquire underwater images and perform preprocessing; S2. Construct a depth estimation subnetwork DBEDNet based on the Encoder-Decoder framework, which introduces spatial structure information and can estimate scene depth information from underwater images; S3. Use the backbone network to further extract high-level image features and combine them with depth information for enhanced processing; S4. Using the brightness reversal map R generated earlier, the detail recovery capability of the low-brightness area of ​​the image is enhanced; S5. Using the dynamic multi-scale attention model DMAM, after fusing multi-scale features with depth guidance information, it dynamically assigns importance weights of different scales to different regions; S6. Reconstruction of fused feature map; S7. Design a loss function and output the predicted underwater image.

[0028] The preprocessing method in step S1 is as follows: Underwater images often suffer from color cast and low contrast due to their unique light absorption and scattering characteristics. To fully preserve the original information and facilitate feature extraction, this implementation performs the following preprocessing at the input stage: 1. Image normalization: Normalize the original pixel values ​​of the input underwater image I to the range [0,1] to reduce the impact of value fluctuations on subsequent training.

[0029] 2. Brightness inversion generation: Design a brightness inversion mapping function for the dark areas in the image : ; By inverting the original image to obtain the brightness inversion map R, the network can focus on recovering details in low-brightness areas in the subsequent process. In actual implementation, this operation is performed on each color channel, and the original image and its inversion map are combined as joint input.

[0030] 3. Convolutional preprocessing layer: After the input image is normalized and brightness inverted, the initial feature extraction layer is set up, using a 3×3 convolution kernel, a stride of 1, and a padding of 1. It is then followed by a PReLU activation function and batch normalization (BN). This layer can be expressed as: ; in, is the convolution kernel parameter, is the bias term. Represents the initially extracted low-level features, and its size is consistent with the input.

[0031] Depth estimation subnetwork (DBEDNet) in step S2: To introduce spatial structure information into the image enhancement process, this implementation designs a subnetwork specifically for depth estimation, namely DBEDNet. This network is based on the Encoder-Decoder framework and can estimate scene depth information from underwater images to provide effective guidance for subsequent modules.

[0032] 1. Encoder part: The encoder uses a multi-level convolution stack structure. Each layer consists of convolution, batch normalization, and pre-reduction unit (PReLU), gradually reducing the size of the feature map and increasing the receptive field. Assuming the encoder has N layers, the processing process of each layer is described as follows: ; ; The convolution kernel size of each layer is set to 3×3 or 5×5, with a stride of 2, to achieve image downsampling. In addition to standard convolution, the encoder also incorporates a Dual-Attention Module (DAM) to capture inter-channel and spatial dependencies. Its structure essentially includes a self-attention mechanism and a positional attention mechanism. Specifically, global average pooling is used to obtain channel attention, followed by local convolution to obtain spatial attention. The output of the DAM is fused with the standard convolution features using the following formula: ; in, and represent channel attention and spatial attention functions respectively, is the fusion weight, determined through experiments.

[0033] 2. Decoder part: The multi-scale features extracted by the encoder are restored in the decoder after passing through several encoding layers. The decoder mainly uses upsampling operations to gradually restore the feature map to its original size. Let each layer of the decoder be: ; in Represents an upsampling convolution operation. At each decoding stage, to preserve encoder information, skip connections are used to directly fuse the corresponding encoding layer features with the decoding layer features. This fusion is achieved through simple addition or concatenation, followed by 1×1 convolution to a fixed number of channels.

[0034] 3. Depth map output: The last layer of the decoder outputs a single-channel depth map D through convolution mapping. This depth map can approximately describe the spatial distance information of each area in the image. The formula is as follows: ; in, represents the Sigmoid activation function, constraining the value between [0, 1]. The final depth map D will serve as the auxiliary input of the subsequent DMAM module to guide the multi-scale attention mechanism.

[0035] Backbone feature extraction module in step S3: After preprocessing and depth estimation, the backbone network's main task is to further extract high-level image features and combine them with depth information for enhancement. This module mainly includes the following parts: 1. Primary feature extraction: In addition to the aforementioned input convolutional layers, the backbone network's primary feature extraction again utilizes multiple layers of convolution (using either 3×3 or 5×5 kernels) and a residual block (ResBlock) structure. ResBlock employs identity mapping to prevent vanishing gradients. Its structure is formulated as follows: ; in, and These are the first and second convolution layers, respectively. Residual connections ensure sufficient information transfer. The feature map output by the primary feature extraction module has the same size as the input, and the number of channels is set to 64 or 128 based on actual needs.

[0036] 2. Multi-scale feature fusion: In order to fully capture the information of each scale in the image, the present invention introduces a pyramid structure into the backbone network, applies convolution kernels of different scales and different receptive fields to the same input, obtains features of different scales through multi-branch parallel convolution, and then fuses them. , and its calculation formula is: ; The fusion operation uses channel concatenation and 1×1 convolution to reduce the dimension, and then performs activation processing to generate a fused multi-scale feature map. .

[0037] 3. Feature and depth information fusion: In order to introduce depth information guidance, this embodiment upsamples the depth map D obtained from DBEDNet to the same size as the backbone network feature map through bilinear interpolation, and then processes it through a single 3×3 convolution and PReLU. Let the processed depth feature be , then the feature fusion operation is: ; in, The feature maps are summed or concatenated channel by channel and then adjusted through fused convolution. This design allows the network to focus on local texture while also taking into account overall structural information, helping to suppress noise and blur in underwater images.

[0038] RBM module in step S4 (brightness inversion auxiliary enhancement): The main purpose of the RBM module is to use the brightness reversal map R generated earlier to help enhance the ability to restore details in low-brightness areas of the image. The specific implementation steps are as follows: 1. Branch build: The input brightness inversion map R is processed by an independent convolution branch at the same time. Similarly, this branch is designed as: ; The convolution kernel size is also set to 3×3, and the number of channels is kept similar to that of the main branch. After multiple layers of convolution, the output feature map of this branch is fused layer by layer with the main features.

[0039] 2. Feature complementation and fusion: Brightness inversion features after convolution extraction With backbone features Specifically, at each key position (such as the edge of the residual block, the end of the coding layer, and the entrance of the DMAM module), and To do element-wise summation, the formula is: ; in, is the fusion coefficient, and the optimal value can be determined through preliminary experiments. This design ensures that even areas with lost dark details in underwater images can be effectively supplemented by inverting the brightness information.

[0040] DMAM module (Dynamic Multi-Scale Attention Module) in step S5: The DMAM module is the core of this invention. After fusing multi-scale features with depth-guided information, it dynamically assigns importance weights to different regions and scales, thereby improving the ability to recover key structural regions and texture details. The module's implementation involves the following steps: 1. Multi-branch attention calculation: Input features After being processed by multiple convolution branches of different scales, each branch uses a convolution kernel of different sizes (such as 3×3, 5×5, 7×7, etc.) to capture contextual information in different ranges. The output of each scale branch is recorded as , the calculation formula is: ; By calculating the features at different scales, the respective response maps are obtained.

[0041] 2. Dynamic learning of attention weights: After splicing the multi-scale response maps, the attention weight is learned through global average pooling (GAP) and fully connected layer (FC). The spliced ​​feature map is denoted as , whose attention vector The calculation steps are as follows: ; in, The sigmoid activation function outputs a weight vector in the range [0,1]. This vector is then used to weight features at each scale.

[0042] 3. Feature reweighting and fusion: After the features of each branch are weighted by the attention vector, they are fused through channel splicing and 1×1 convolution to generate a dynamically adjusted feature map. The fusion formula is expressed as: ; in, This design enables the system to adaptively adjust the proportion of information at each scale based on the importance of different regions, thereby more accurately restoring image details.

[0043] R³S module (Reshape-Reverse-Region Smoothing module): To further improve the smoothness and local consistency of the reconstructed image, this implementation designs an R³S module, which mainly includes three operations: Reshape, Reverse, and Region Smoothing.

[0044] 1.Reshape operation: The feature map output by the DMAM module is re-blocked and the feature map is divided into regions according to a fixed size (such as 8×8 or 16×16) to ensure that each region has a certain local consistency. The sub-block form after division is recorded as .

[0045] 2. Reverse operation: Perform reverse processing on the features within each sub-block, that is, use the brightness inversion features generated in the early stage Compensate the area and apply the inverse mapping to each pixel within the subblock: ; in, To adjust the parameters, This is a compensation term reconstructed based on the inversion relationship of regional brightness. This operation can enhance the details of low-brightness areas while avoiding local distortion caused by over-enhancement.

[0046] 3.Region Smoothing: Apply regional smoothing to each sub-block after processing, using mean filtering or bilateral filtering to ensure smooth transition between different sub-blocks and avoid artifacts caused by transition mutations. After multiple iterations of smoothing, each sub-block is finally reassembled into an overall feature map. .

[0047] Reconstruction module in step S6: After processing by the above modules, a fusion feature map is obtained after dynamic multi-scale attention and depth guidance, brightness inversion compensation and local smoothing. The task of the reconstruction module is to transform the feature map into the final output enhanced image , its structural design is as follows: 1. Residual reconstruction unit: A multi-level residual unit is designed to perform nonlinear mapping and detail recovery on the fused features. Each residual unit contains two layers of convolution operations, using the formula: ; Then we pass the second convolution layer to get: ; The multi-level residual design can preserve the initial image information and refine the enhancement effect layer by layer.

[0048] 2. Final convolution mapping: The reconstruction module finally uses a 3×3 convolutional mapping layer to map the feature map output by the residual unit into a three-channel image. The formula is as follows: ; The Sigmoid activation function constrains the output to be within [0,1]. For the enhanced underwater image after comprehensive processing by all modules, color correction, contrast improvement and detail restoration are achieved.

[0049] Step S7: Training strategy and loss function design: This system adopts an end-to-end joint training strategy. To ensure that each module can work together, the following multiple loss functions are designed during the training process.

[0050] 1. Reconstruction loss: The mean squared error (MSE) is used to measure the output image With the original clear image The gap between: ; This loss is mainly used to guide the global reconstruction accuracy.

[0051] 2. Perceptual loss: Use the pre-trained VGG network to extract high-level features of the image and construct perceptual loss in the feature space: ; in, Represents the feature map of the jth layer of the VGG network. This loss helps to improve the high-level semantic consistency and visual realism of the image.

[0052] 3. Structural retention loss: To address the problem of preserving structural information in underwater images, the structural similarity index (SSIM) loss is designed: ; This loss encourages the output image to maintain high consistency with the original image in local structure and texture.

[0053] 4. Attention Constraint Loss: For the dynamic learning process of attention weights in the DMAM module, a regularization term is added to make the distribution of attention weights at each scale more reasonable. Let the weight vector To meet certain sparsity and smoothness constraints, L1 loss or KL divergence can be designed: ; in Can be set according to preset target distribution.

[0054] 5. Total loss function: Combining the above losses, the total loss function of each module is defined as: ; ; in, is the weight of each loss, and the best combination is determined through cross-validation and experiments.

[0055] 6. Training details: Dataset: Public underwater image datasets such as UIEB and EUVP are used during training. At the same time, data can be amplified by the user to perform artificial synthesis of noise, blur, color cast, etc.

[0056] Optimizer: The Adam optimizer is used, and the initial learning rate is set to 1e-4. In combination with the learning rate decay strategy, the learning rate is gradually reduced when the validation set loss no longer decreases.

[0057] Batch size: Depending on the GPU memory, the batch size is generally set to 8 to 16.

[0058] Iterations: Usually, we train for about 100,000 iterations, and record the model's performance on the validation set at regular intervals. System advantages and implementation results.

[0059] Through the above design, this embodiment has the following obvious advantages: 1. Consider both global and local information: Through the depth estimation subnetwork and multi-scale attention mechanism, the system can capture global structural information while also dynamically adjusting and enhancing local details, effectively solving the problems of blurred details and unclear edges in underwater images.

[0060] 2. Adaptive brightness compensation: By utilizing the inverted brightness image information, we perform special brightness inversion fusion to address the typical problems of uneven underwater lighting and insufficient details in dark areas, ensuring a more balanced overall brightness distribution of the image without introducing overexposure or darkening.

[0061] 3. Robustness and generalization ability: Each module uses residual learning and skip connections to ensure gradient propagation and prevent network degradation. Multiple loss functions are designed to ensure stable model performance across various test datasets. Ablation experiments validate the performance improvements achieved by the DMAM, RBM, and R³S modules, demonstrating the suitability of this invention for diverse underwater scenarios.

[0062] 4. End-to-end training advantages: The entire network adopts an end-to-end joint training method, without the need for complex post-correction, and can output enhanced results in real time. It has extremely high practical value in applications such as underwater robots, remote monitoring, and marine engineering exploration.

[0063] Software and hardware implementation environment: To verify the feasibility of this implementation plan, the following training and deployment environments are recommended: 1. Software environment: The system is based on the PyTorch framework and uses Python 3.8.18 and other environments.

[0064] Dependent libraries include numpy, opencv-python, scikit-image, etc.

[0065] The training code is modularized and supports multi-GPU distributed training.

[0066] 2. Hardware configuration: We recommend using a GPU (such as NVIDIA RTX 3080 or higher) for model training to improve training efficiency.

[0067] The system runs on the Windows operating system. Taking into account the real-time reasoning requirements of industrial applications, it also supports embedded system deployment and achieves fast reasoning after conversion to the TensorRT model format.

[0068] Summary of specific implementation process: Based on the above modules, the specific implementation process of this implementation plan is as follows: 1. Normalize and invert the brightness of the input underwater image to generate the initial feature F0 and the inversion map R.

[0069] 2. Obtain low-level features through the primary feature extraction layer, and input the image into the depth estimation subnetwork DBEDNet at the same time to obtain the depth map D.

[0070] 3. Use a multi-level convolution residual module and a multi-scale pyramid structure to process the primary features to obtain the multi-scale fused features F_ms, which are then fused with the upsampled deep features F_D to obtain F_fused.

[0071] 4. Using the RBM module, the feature F_R extracted by the brightness inversion branch is added to F_fused to enhance local details and generate F_enhanced.

[0072] 5. Input F_enhanced into the DMAM module, dynamically reweight the features of each scale through multi-scale convolution branches and global attention weight learning, and output F_{DMAM}.

[0073] 6. For F_{DMAM}, the R³S module is used to perform intra-region reshaping, reverse compensation, and region smoothing to obtain the smoothed comprehensive feature F_{R^3S}.

[0074] 7. Finally, F_{R^3S} is fed into the residual reconstruction module and the final convolutional mapping layer to output the restored underwater enhanced image I_out.

[0075] 8. During the training process, MSE, perceptual loss, SSIM loss and attention constraint loss are used to jointly optimize the network parameters to ensure the best balance between global restoration and detail recovery.

[0076] Experiment and results display: In the experimental part, the proposed method is compared with existing mainstream methods (such as UWCNN, WaterNet, UIE-DAL, etc.). The experiments use public datasets such as UIEB and EUVP, and conduct the following tests.

[0077] 1. Quantitative indicators: Indicators such as PSNR and SSIM are used to measure the quality of enhanced images, and numerical results show that the proposed method is superior to traditional methods in detail restoration and color correction.

[0078] 2. Ablation experiment: Remove the DMAM, RBM, and R³S modules respectively, observe the changes in image details, and verify the improvement of each module on the final performance.

[0079] 3. Qualitative observation: Comparing the performance of images under different lighting and underwater environments proves that the present invention can work stably in various underwater scenes.

[0080] Experimental results show that this method is superior to existing methods in detail restoration, color balance, edge clarity, etc., and has strong robustness and broad application prospects.

[0081] This implementation proposes a novel underwater image enhancement method based on the fusion of depth estimation and dynamic multi-scale attention. By preprocessing the input image and introducing a DBEDNet subnetwork for depth guidance, the method combines the brightness inversion assistance of the RBM module with the multi-scale dynamic attention adjustment of the DMAM module, and finally utilizes the R³S module for local region smoothing, effectively improving the detail, contrast, and color of underwater images. Furthermore, an end-to-end joint training strategy is employed to ensure the gradual convergence of the entire model during training, demonstrating excellent enhancement results in a variety of underwater scenes.

[0082] The present invention has high industrial application value, especially in the fields of underwater robot navigation, marine resource exploration, environmental monitoring, etc., and can play a significant image enhancement role, providing more reliable input information for underwater vision systems, thereby improving the working efficiency and safety of the overall system.

Claims

1. A dual-guided underwater image multi-scale enhancement method, characterized in that: Here are the steps: S1. Acquire underwater images and perform preprocessing; S2. Construct a depth estimation subnetwork DBEDNet based on the Encoder-Decoder framework, which introduces spatial structure information and can estimate scene depth information from underwater images; S3. Use the backbone network to further extract high-level image features and combine them with depth information for enhanced processing; S4. Using the brightness reversal map R generated earlier, the detail recovery capability of the low-brightness area of ​​the image is enhanced; S5. Using the dynamic multi-scale attention model DMAM, after fusing multi-scale features with depth guidance information, it dynamically assigns importance weights of different scales to different regions; S6. Reconstruction of fused feature map; S7. Design a loss function and output the predicted underwater image.

2. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: The pre-processing in step S1 includes: Normalization processing: normalize the original pixel values ​​of the input underwater image I to the interval [0,1]; Brightness inversion generation: Design a brightness inversion mapping function for the dark areas in the image : ; By inverting the original image to obtain the brightness reversal map R, the network can focus on the recovery of details in low-brightness areas in the subsequent process; Convolutional preprocessing layer: After the input image is normalized and brightness inverted, the initial feature extraction layer is set, which can be expressed as: ; in, is the convolution kernel parameter, is the bias term, then Represents the initially extracted low-level features, and its size is consistent with the input.

3. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: In step S2, the depth estimation subnetwork DBEDNet includes an encoder and a decoder; The encoder adopts a multi-level convolution stacking structure. Assuming that the encoder has multiple layers, the processing process of each layer is described as follows: ; ; PReLU is the activation function, represents the convolution kernel weight matrix of the encoder layer i, represents the bias term of the encoder layer i; A dual attention module is introduced in the encoder to capture inter-channel and spatial dependencies. Its structure includes a self-attention mechanism and a position attention mechanism: In specific implementation, channel attention is first obtained through global average pooling, and then spatial attention is obtained through local convolution; The output of DAM is fused with the standard convolutional features, and the formula is as follows: ; in, and represent channel attention and spatial attention functions respectively, is the fusion weight; The decoder uses upsampling operation to gradually restore the feature map to its original size. Let each layer of the decoder be: ; in represents the upsampling convolution operation, represents the convolution kernel weight matrix of the jth layer of the decoder, Represents the bias item of the jth layer of the decoder; in each decoding stage, in order to maintain the encoder information, a skip connection is used to directly fuse the corresponding encoding layer features with the decoding layer features.

4. The dual-guided underwater image multi-scale enhancement method according to claim 3, characterized in that: The last layer of the decoder outputs a single-channel depth map D through convolution mapping. This depth map can approximately describe the spatial distance information of each area in the image. The formula is as follows: ; in, Represents the Sigmoid activation function, which constrains the value to be between [0,1], Represents the convolution kernel weight matrix for depth estimation, represents the output feature map of the last layer of the decoder, represents the bias term of depth estimation; the final depth map D will be used as the auxiliary input of the subsequent dynamic multi-scale attention model to guide the multi-scale attention mechanism.

5. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: In step S3, the backbone network includes primary feature extraction, multi-scale feature fusion, and feature and depth information fusion; The primary feature extraction formula is as follows: ; in, is the original feature map, and These are the first and second convolution layers, respectively. The feature map size of the primary feature extraction output is consistent with the input; Multi-scale feature fusion: Apply convolution kernels of different scales to the same input, extract features of different receptive fields through multi-branch parallel convolution, and fuse them; there are multiple scale branches , and its calculation formula is: ; Indicates the The bias term of the scale branch, Indicates the The convolution kernel weight matrix of the scale branch; it is used to Perform convolution operation to extract features at that scale; the fusion operation uses channel splicing and 1×1 convolution to reduce the dimension, and then performs activation processing to generate a fused multi-scale feature map ; Feature and depth information fusion: Assume that the processed depth feature is , then the feature fusion operation is: ; in, It means that the feature maps are summed or concatenated channel by channel, and the channels are adjusted through fused convolution.

6. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: The specific implementation steps of step S4 are as follows: S41. The input brightness inversion map R is processed by an independent convolution branch at the same time. The branch is designed as follows: ; in, represents the convolution kernel weight matrix in the independent convolution branch that processes the brightness inversion map R, Represents the bias term in the convolution branch. After multiple layers of convolution, the output feature map of the branch is fused layer by layer with the backbone features. Brightness inversion features after convolution extraction With backbone features Layered fusion, and Perform element-by-element summation, the formula is: ; in, is the fusion coefficient.

7. The dual-guided underwater image multi-scale enhancement method according to claim 6, characterized in that: The dynamic multi-scale attention DMAM model includes multi-branch attention calculation, dynamic learning of attention weights, and feature reweighting and fusion; Multi-branch attention calculation: The input features are processed by multiple convolution branches of different scales. Each branch uses a convolution kernel of different sizes to capture contextual information in different ranges. The output of each scale branch is recorded as , the calculation formula is: ; represents the convolution kernel weight matrix of the i-th scale branch in the dynamic multi-scale attention module, Represents the bias term in the scale branch, and obtains the respective response maps by calculating the features of different scales; Dynamic learning of attention weights: After splicing the multi-scale response maps, the attention weight is learned through global average pooling and full connection layer. The spliced ​​feature map is recorded as , whose attention vector The calculation steps are as follows: ; in, is the Sigmoid activation function, and the output weight vector is in the interval [0,1]; represents the weight matrix in attention weight learning, Represents the feature map after splicing the multi-scale response map, Represents the bias term in attention weight learning; Feature reweighting and fusion: Each branch feature is weighted by attention After weighting, channel splicing and 1×1 convolution fusion are performed to generate a dynamically adjusted feature map. , the fusion formula is expressed as: ; in, is the corresponding branch attention weight, The convolution kernel weight matrix used for the fusion operation, represents the feature map of the i-th scale branch, Represents the bias term of the fusion operation.

8. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: Perform smoothing processing on the image output by the dynamic multi-scale attention DMAM model, including Reshape, Reverse and RegionSmoothing; Reshape processing: The feature map output by the dynamic multi-scale attention DMAM model is re-blocked, and the feature map is divided into regions according to a fixed size to ensure local consistency within each region; the sub-block form after division is recorded as ; Reverse processing: Perform reverse processing on the features within each sub-block, that is, use the brightness inversion features generated in the early stage Compensate the area and apply the inverse mapping to each pixel within the subblock: ; in, To adjust the parameters, is the compensation term reconstructed based on the regional brightness inversion relationship; Region Smoothing Processing: Apply regional smoothing to each processed sub-block, eliminate the mutation between blocks through multiple iterations, and finally reorganize the sub-blocks into the overall feature map .

9. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: The reconstruction of the fused feature map in step S6 is as follows: A multi-level residual unit is designed to perform nonlinear mapping and detail recovery on the fusion features, using the formula: ; Represents the input feature map of the kth layer, which contains the feature information passed from the previous layer; Represents the weight matrix of the first convolution operation of the kth layer, which is used to convolution the input feature map Perform convolution operation; represents the bias term of the first convolution operation of the kth layer; Then we pass the second convolution layer to get: ; Represents the feature map of the kth layer after the first convolution operation and activation function, which is obtained by converting the input feature map Added with residual connection; Represents the weight matrix of the second convolution operation of the kth layer, which is used to convolution the feature map Perform further convolution operations; Represents the bias term of the second convolution operation of the kth layer; The multi-level residual design can retain the initial image feature information and refine the enhancement effect layer by layer; The reconstruction module finally uses a 3×3 convolutional mapping layer to map the feature map output by the residual unit into a three-channel image. The formula is as follows: ; The Sigmoid activation function constrains the output to be within [0,1], and the final output is the final fusion feature map, Represents the weight matrix of the final convolutional mapping layer, which is used to map the feature map output by the residual unit into a three-channel image; Represents the final feature map after all processing, Represents the bias term of the final convolutional map layer.

10. The dual-guided underwater image multi-scale enhancement method according to claim 1, characterized in that: Step S7 designs a loss function and outputs the predicted underwater image. The specific method is as follows: Using mean square error loss To measure the output image With the original clear image The gap between Use the pre-trained VGG network to extract high-level features of the image and construct perceptual loss in the feature space: ; Where N represents the number of feature layers in the VGG network used to calculate the perceptual loss. represents the feature map of the jth layer of the VGG network; Design structural similarity index SSIM loss: ; This loss encourages the output image to maintain high consistency with the original image in local structure and texture; Attention Constrained Loss: Let the weight vector Satisfying the set sparsity and smoothness constraints, it can be designed Loss or KL divergence: ; in Represents the preset target attention vector, which is an ideal attention weight distribution used to guide the model to learn a more reasonable attention allocation and can be set according to the preset target distribution; The total loss function of each module is defined as: ; or ; in, is the weight of each loss.

Citation Information

Patent Citations

  • Underwater image enhancement method of multi-attention mechanism guided by brightness mask

    CN116402715A

  • Underwater image enhancement method for efficiently guiding information flow

    CN117392032A

  • Monocular three-dimensional target detection method based on convolution attention and feature decoupling

    CN117557980A

  • Convolutional neural network-based image processing method and device, and unmanned aerial vehicle

    WO2020062284A1