Underwater Image Enhancement Method Based on Adaptive Multi-Scale Fusion and Attention Mechanism
Through the underwater image enhancement method of adaptive multi-scale fusion and attention mechanism, the problem of underutilization of feature information at different scales is solved, efficient enhancement of underwater images is achieved, and the clarity and contrast of images are improved.
Patent Information
- Application Number
- CN202311510657.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-11-14
AI Technical Summary
The existing deep learning-based underwater image enhancement methods fail to fully consider the importance of feature information at different scales, resulting in poor image enhancement quality.
The underwater image enhancement method is adopted with an adaptive multi-scale fusion and attention mechanism. Different scale features are extracted and fused through the adaptive multi-scale fusion module, and the improved channel space attention module focuses on key features to build an underwater image enhancement model.
提高了水下图像的增强效果,有效校正色偏,增强对比度,保留细节信息,提升了图像清晰度和质量。
Smart Images

Figure CN117314787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly relates to an underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism. Background Art
[0002] In the process of marine resource exploration and development, the clarity of underwater video images is one of the key technologies, which is crucial for improving the clarity of underwater video images and the performance of underwater detection equipment. However, underwater optical images usually suffer from serious quality degradation problems due to the absorption and scattering of light during underwater propagation, such as color distortion, blurring and fogging, low contrast, etc. And severely degraded underwater images will affect the subsequent tasks of underwater equipment. Therefore, it is necessary to enhance underwater images to improve the quality of underwater images.
[0003] Currently, the enhancement methods of underwater optical images are generally divided into: traditional image enhancement methods based on non-physical models, traditional image enhancement methods based on physical models, and image enhancement methods based on deep learning. Traditional image enhancement methods based on non-physical models do not rely on the underwater imaging physical model and are directly applied to underwater degraded images. According to the requirements, corresponding image processing methods are used to adjust pixel values to improve the visual effect and obtain enhanced underwater images. The basic idea of traditional image enhancement methods based on physical models is to construct a mathematical model according to the physical characteristics of the underwater environment and reconstruct the clear image before degradation by estimating model parameters. However, this method usually needs to consider the optical characteristics of underwater images, depends on environmental assumptions and prior knowledge, is limited by the accuracy of parameter estimation, and the mathematical model has great limitations. It is difficult to design a unified physical model for different environments, resulting in generally weak environmental adaptability of this method.
[0004] In recent years, underwater image enhancement methods based on deep learning have been widely applied. Since deep learning algorithms have powerful feature learning and representation capabilities, they can autonomously learn key features in images and end-to-end fit the complex mapping relationship between underwater degraded images and underwater enhanced images in a data-driven manner, and can better describe the content and structure of underwater images. Existing underwater image enhancement methods based on deep learning generally use deep convolutional networks (CNNs) to improve the quality of underwater images. CNNs can learn the features and patterns of images through a series of convolutional layers, pooling layers and activation functions, thereby improving the contrast and clarity of images and reducing noise. Deep convolutional networks (CNNs) are more stable and efficient in underwater enhancement tasks. The features of underwater images are usually distributed at different scales, and the importance of feature information at different scales is different. However, existing underwater image enhancement technologies based on deep convolutional networks rarely consider the importance of feature information at different scales, resulting in insufficient extraction of image feature information and poor image enhancement quality. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide an underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] An underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism, characterized in that the method includes the following contents:
[0008] Construct an underwater image enhancement model, which includes an encoder, a decoder and a bottleneck layer located between the encoder and the decoder; the encoder includes a depth convolution layer, an adaptive multi-scale fusion module and downsampling, and the adaptive multi-scale fusion module is used to extract features of different scales and fuse the feature maps of different scales. The decoder includes an improved channel spatial attention module, upsampling and a depth convolution layer;
[0009] The underwater degraded image passes through the depth convolution layer, normalization processing and activation function in sequence, and then is input into the first adaptive multi-scale fusion module. The output feature map of the first adaptive multi-scale fusion module is input into the second adaptive multi-scale fusion module after downsampling. The output feature map of the second adaptive multi-scale fusion module is input into the third adaptive multi-scale fusion module after downsampling. The output feature map of the third adaptive multi-scale fusion module is input into the bottleneck layer after downsampling; the output feature map of the bottleneck layer is concatenated with the output feature map of the third adaptive multi-scale fusion module after upsampling, and then passes through the first improved channel spatial attention module and the first depth convolution layer in sequence. The output feature map of the first depth convolution layer is concatenated with the output feature map of the second adaptive multi-scale fusion module after upsampling, and then passes through the second improved channel spatial attention module and the second depth convolution layer in sequence. The output feature map of the second depth convolution layer is concatenated with the output feature map of the first adaptive multi-scale fusion module after upsampling, and then passes through the third improved channel spatial attention module and the third depth convolution layer in sequence to obtain the output feature map of the decoder. The output feature map of the decoder is fused with the input image of the model through skip connection to obtain the enhanced underwater image;
[0010] Train the underwater image enhancement model, and use the trained underwater image enhancement model for the enhancement of underwater degraded images.
[0011] Further, the adaptive multi-scale fusion module includes an atrous convolution layer, a convolution layer, and a fully connected layer. The input feature map passes through two atrous convolution layers with different atrous rates to obtain two feature maps with different scales. After the two feature maps are concatenated along the channel dimension, they pass through a convolution layer, and then through a fully connected layer and a Relu activation function, and then through another fully connected layer and a Sigmoid activation function to obtain weights of two different scales. The feature maps extracted by the two atrous convolution layers are respectively multiplied by the weights of the corresponding scales and then concatenated along the channel dimension to obtain the output feature map of the adaptive multi-scale fusion module.
[0012] Further, the convolution kernel size of one atrous convolution layer of the adaptive multi-scale fusion module is 3×3, the atrous rate is 2, and the convolution kernel size of the other atrous convolution layer is 3×3, and the atrous rate is 1.
[0013] The bottleneck layer includes an improved channel-spatial attention module. The improved channel-spatial attention modules of the decoder and the bottleneck layer have the same structure, both including a channel attention module and a spatial attention module. In the channel attention module, the input feature map passes through max pooling and average pooling operations respectively to obtain two feature maps. These two feature maps pass through a fully connected layer and a PReLU activation function in sequence, and then through another fully connected layer and a Sigmoid activation function to obtain the weights of max pooling and average pooling. After the weights of max pooling and average pooling are subjected to element-wise summation, and then through a Sigmoid activation function to obtain the channel attention weight. The channel attention weight is multiplied by the input feature map of the channel attention module to obtain the channel attention feature map. The channel attention feature map serves as the input feature map of the spatial attention module. In the spatial attention module, the channel attention feature map passes through max pooling and average pooling operations respectively to obtain two feature maps. These two feature maps are concatenated along the channel dimension and then pass through convolution and a Sigmoid activation function to obtain the spatial attention weight. The spatial attention weight is multiplied by the input feature map of the spatial attention module to obtain the output feature map of the improved channel-spatial attention module.
[0014] Further, the convolution kernel size of the depth convolution layers of the encoder and the decoder is 3×3, and the stride is 1.
[0015] Compared with the prior art, the present invention has the following advantages:
[0016] The present invention designs an adaptive multi-scale fusion module, which extracts feature information of different scales through dilated convolutions with different dilation rates, and pays attention to the importance of different-scale features in an adaptive manner, enabling the module to dynamically adjust the degree of attention to each scale of information, improving the generalization ability of the module. By fusing feature maps of different scales, it helps to reduce the problem of information loss. To further improve the feature extraction ability of the model, an improved channel-spatial attention mechanism is added to the high-dimensional feature space. Guided by the attention, it is beneficial for the model to focus more on the features that are more critical for the underwater image enhancement task, enabling the model to better correct color cast, enhance contrast, and retain detail information, solving the problems of color distortion, low contrast, and blurred details existing in underwater degraded images, and improving the enhancement effect of underwater degraded images. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a structural diagram of the underwater image enhancement model of the present invention;
[0018] Figure 2 is a structural diagram of the adaptive multi-scale fusion module of the present invention;
[0019] Figure 3 is a structural diagram of the improved channel-spatial attention module of the present invention;
[0020] Figure 4 is a comparison chart of the results of different models on the UIEB test set. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Specific embodiments are given below in conjunction with the accompanying drawings. The specific embodiments are only used to introduce the technical solutions of the present invention in detail, and do not limit the protection scope of this application.
[0022] The present invention provides an underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism (hereinafter referred to as the method, see Figures 1 to 4 ), including the following contents:
[0023] Construct an underwater image enhancement model, such as Figure 1As shown in the figure, the underwater image enhancement model includes an encoder, a decoder, and a bottleneck layer. The bottleneck layer is located between the encoder and the decoder and plays a role in information transmission. The encoder includes a depth convolution layer, an adaptive multi-scale fusion module (AMM), and downsampling. The underwater degraded image is input into the encoder. First, it passes through a depth convolution layer with a convolution kernel size of 3×3 and a stride of 1 to adjust the number of channels. After the output feature map of the depth convolution layer is normalized and passed through an activation function, it enters the first adaptive multi-scale fusion module for feature extraction. The output feature map of the first adaptive multi-scale fusion module is input into the second adaptive multi-scale fusion module after downsampling. The output feature map of the second adaptive multi-scale fusion module is input into the third adaptive multi-scale fusion module after downsampling. The output feature map of the third adaptive multi-scale fusion module is input into the bottleneck layer after downsampling. The adaptive multi-scale fusion module is used to extract features of different scales and fuse the feature maps of different scales to achieve feature integration. Downsampling is implemented through a max pooling layer with a convolution kernel size of 2×2 and a stride of 2. The purpose of downsampling is to reduce the size of the feature map and the computational amount.
[0024] The decoder includes an improved channel spatial attention module (PCBAM), upsampling, and a depth convolution layer. After the output feature map of the bottleneck layer is upsampled and concatenated with the output feature map of the third adaptive multi-scale fusion module through channels, it then passes through the first improved channel spatial attention module and the first depth convolution layer in sequence. The output feature map of the first depth convolution layer is upsampled and concatenated with the output feature map of the second adaptive multi-scale fusion module through channels, and then passes through the second improved channel spatial attention module and the second depth convolution layer in sequence. The output feature map of the second depth convolution layer is upsampled and concatenated with the output feature map of the first adaptive multi-scale fusion module through channels, and then passes through the third improved channel spatial attention module and the third depth convolution layer in sequence to obtain the output feature map of the decoder. The output feature map of the decoder and the input image of the model are fused through a skip connection to obtain the output of the underwater image enhancement model, that is, the enhanced underwater image. The purpose of the skip connection is to better retain the original feature information and prevent the loss of the original feature information. The purpose of channel concatenation is to reduce the loss of feature information caused by the model during downsampling. Upsampling is implemented through a transposed convolution layer with a convolution kernel size of 2×2 and a stride of 2. The convolution kernel size of each depth convolution layer in the decoder is 3×3 and the stride is 1.
[0025] The adaptive multi-scale fusion module can integrate information from different-scale features in a way that pays attention to the importance of feature maps of different scales. Its structure is as Figure 2As shown in the figure; the adaptive multi-scale fusion module includes an atrous convolution layer, a convolution layer, and a fully connected layer. The input feature map of the adaptive multi-scale fusion module first passes through two atrous convolution layers with different atrous rates to extract information of different scales, obtaining feature maps of different scales; after the channel splicing of the two feature maps of different scales, a convolution layer with a kernel size of 1×1 is used to adjust the number of channels, and then through a fully connected layer and a Relu activation function, and then through a fully connected layer and a Sigmoid activation function to obtain two weights K1 and K2 of different scales; after multiplying the feature maps extracted by the two atrous convolution layers by the corresponding scale weights respectively, and then performing channel splicing, the output feature map of the adaptive multi-scale fusion module is obtained. The role of the two atrous convolution layers is to obtain information of different scales without reducing the resolution. The convolution kernel size of one atrous convolution layer is 3×3 and the atrous rate is 2, and the convolution kernel size of the other atrous convolution layer is 3×3 and the atrous rate is 1.
[0026] The bottleneck layer consists of an improved channel-spatial attention module. The structure of the improved channel-spatial attention module in the decoder and the bottleneck layer is the same. As Figure 3 shown in the figure, the improved channel-spatial attention module includes a channel attention module and a spatial attention module. The channel attention module performs max-pooling and average-pooling operations on the input feature map respectively to obtain two feature maps. These two feature maps pass through a fully connected layer and a PReLU activation function in sequence, and then through a fully connected layer and a Sigmoid activation operation to obtain the weights of max-pooling and the weights of average-pooling; after performing element-wise summation on the weights of max-pooling and the weights of average-pooling, and then through a Sigmoid activation function to obtain the channel attention weights, multiplying the channel attention weights by the input feature map of the channel attention module to obtain the channel attention feature map; the channel attention feature map is used as the input feature map of the spatial attention module. In the spatial attention module, the channel attention feature map performs max-pooling and average-pooling operations respectively to obtain two feature maps. After the channel splicing of these two feature maps, a convolution operation is performed to reduce the dimension of the feature map. The feature map after dimension reduction passes through a Sigmoid activation operation to obtain the spatial attention weights, and multiplying the spatial attention weights by the input feature map of the spatial attention module to obtain the channel-spatial attention feature map, that is, the output feature map of the improved channel-spatial attention module. In the channel attention module, the PReLU activation function is used to replace the original ReLU activation function. This is because the PReLU activation function introduces a learnable parameter, allowing the slope of the activation function to be learned during training instead of using a fixed slope, which can make the model more flexible to adapt to different data distributions and improve the feature extraction ability of the model.
[0027] The underwater image enhancement model is trained and tested using the underwater image enhancement benchmark dataset UIEB. The training set consists of 890 pairs of images composed of underwater degraded images and labeled images, and the test set consists of 60 underwater degraded images. During the training process, the Adam optimizer is used, the Batch size is 4, the learning rate is set to 0.0002, and a total of 500 epochs are trained. To better enhance underwater images, the present invention jointly optimizes the model using smooth L1 loss and structural similarity error loss, and the joint loss function is defined as follows:
[0028] L loss =σL MSSSIM +(1 - σ)SmoothL1
[0029] where L loss is the model training loss, L MSSSIM is the structural similarity error loss, SmoothL1 is the smooth L1 loss, σ is the weight coefficient, and in this embodiment, σ is set to 0.8;
[0030] In the underwater enhancement task, the smooth L1 loss helps to ensure the similarity between the image generated by the model and the real image at the pixel level, prompting the model to learn to generate smoother and more natural images. The smooth L1 loss is expressed as:
[0031]
[0032] where x i is the pixel value at the i-th pixel point in the labeled image, and y i is the pixel value at the i-th pixel point in the enhanced underwater image;
[0033] In the underwater enhancement task, the structural similarity loss helps to ensure that the image generated by the model is more similar to the real image in structure. The structural similarity loss is expressed as:
[0034]
[0035]
[0036] where n represents the total number of pixel points in the image, μ x represents the pixel mean of the labeled image, μ y represents the pixel mean of the enhanced underwater image, σ xy represents the covariance represents the variance of the pixel values of the labeled image, represents the variance of the pixel values of the enhanced underwater image, and C1 and C2 are constants. In this embodiment, C1 takes the value of 0.012 and C2 takes the value of 0.032;
[0037] The underwater image enhancement model is tested using a test set to verify the effectiveness of the model. To more objectively analyze and evaluate the model performance, the image quality evaluation index information entropy IE and the underwater image quality evaluation index UCIQE are selected as objective evaluation indexes. IE is a metric for measuring the amount of information, representing the richness of information. The larger the information entropy, the richer the information and the better the image quality. Its calculation formula is as follows:
[0038]
[0039] where j represents the pixel value of the image, p(j) represents the probability when the pixel value is j, and m represents the maximum pixel value;
[0040] UCIQE is a commonly used underwater image quality evaluation index, which is a linear combination of color concentration, saturation, and contrast, used to evaluate the impact of factors such as color distortion and contrast on image quality and give corresponding scores. The higher the UCIQE score, the better the underwater image quality. Its calculation formula is as follows:
[0041] UCIQE = c1×δ + c2×conl + c3×μ
[0042] where δ is the standard deviation of color concentration, conl is the contrast of brightness, μ is the average value of saturation, and c1, c2, and c3 are the weight values of the linear combination, and the values taken in this embodiment are 0.4680, 0.2745, and 0.2576 respectively.
[0043] Table 1 shows the IE evaluation index values of the images enhanced by different models, and Table 2 shows the UCIQE evaluation index values of the images enhanced by different models. It can be seen from the table that the vast majority of the images enhanced by the model of the present invention have obtained the best results. Compared with the other 5 models, the model proposed by the present invention has significantly improved in both IE and UCIQE evaluation indexes, indicating that the underwater image quality enhanced by the present invention is better.
[0044] Table 1 Comparison of IE evaluation indexes of different images under different models
[0045]
[0046]
[0047] Table 2 Comparison of UCIQE evaluation indexes of different images under different models
[0048]
[0049] Figure 4For the comparison of the results enhanced by different models, from the comparison results, the UDCP model exacerbates the blue color cast. The enhanced image introduces a red color cast, and the image contrast decreases, resulting in the overall image being too dark. The CLAHE model improves the contrast of underwater images and reduces the fogging problem. However, this method does not well correct the blue color cast and there is a problem of local over-enhancement. The UWCNN model has an insignificant effect on image color correction and introduces a red color cast. The UGAN model can effectively correct the image color and increase the image contrast, but there is an over-enhancement situation, resulting in the colors of some regions after enhancement being distorted compared to the real colors, and there are local over-bright or local over-dark situations. The U-Net model enhances underwater images to a certain extent, but this method fails to well correct the blue-green color cast and improve the contrast, and does not restore the real color of the image. In contrast, the model proposed in the present invention can effectively remove the blue-green color cast in the underwater environment, restore the real color of the image, enhance the contrast and improve the image clarity. To sum up, the method of the present invention can well correct the color cast of underwater images, enhance the image contrast and improve the image clarity.
[0050] To further verify the contribution of the improvement points of the present invention to the model performance, ablation experiments were carried out. The effectiveness of the adaptive multi-scale fusion module and the improved channel spatial attention module was verified through control experiments with or without the adaptive multi-scale fusion module and the improved channel spatial attention module. The IE and UCIQE evaluation indexes of different ablation experiments are shown in Table 3.
[0051] Table 3 Comparison of IE and UCIQE evaluation indexes under different model structures
[0052]
[0053] Among them, in Model A, the adaptive multi-scale fusion module was replaced with ordinary depth convolution, and the original channel spatial attention module was replaced with the improved channel spatial attention module. On this basis, Models B - D were obtained. As can be seen from Table 3, using both the adaptive multi-scale fusion module and the improved channel attention module shows better results in both IE and UIQE evaluation indexes than only using the adaptive multi-scale module or the improved channel spatial attention module. Therefore, the underwater image enhancement model of the present invention has good performance and can significantly improve the image quality.
[0054] Matters not described in the present invention are applicable to the prior art.
Claims
1. An underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism, characterized in that, The method includes the following steps: Construct an underwater image enhancement model, which includes an encoder, a decoder, and a bottleneck layer located between the encoder and the decoder; the encoder includes a depth convolution layer, an adaptive multi-scale fusion module, and downsampling. The adaptive multi-scale fusion module is used to extract features of different scales and fuse the feature maps of different scales. The decoder includes an improved channel spatial attention module, upsampling, and a depth convolution layer; The underwater degraded image passes through a depth convolution layer, normalization processing, and an activation function in sequence, and then is input into the first adaptive multi-scale fusion module. The output feature map of the first adaptive multi-scale fusion module is input into the second adaptive multi-scale fusion module after downsampling. The output feature map of the second adaptive multi-scale fusion module is input into the third adaptive multi-scale fusion module after downsampling. The output feature map of the third adaptive multi-scale fusion module is input into the bottleneck layer after downsampling. The output feature map of the bottleneck layer is concatenated with the output feature map of the third adaptive multi-scale fusion module through channels after upsampling, and then passes through the first improved channel spatial attention module and the first depth convolution layer in sequence. The output feature map of the first depth convolution layer is concatenated with the output feature map of the second adaptive multi-scale fusion module through channels after upsampling, and then passes through the second improved channel spatial attention module and the second depth convolution layer in sequence. The output feature map of the second depth convolution layer is concatenated with the output feature map of the first adaptive multi-scale fusion module through channels after upsampling, and then passes through the third improved channel spatial attention module and the third depth convolution layer in sequence to obtain the output feature map of the decoder. The output feature map of the decoder and the input image of the model are fused through skip connection to obtain the enhanced underwater image; The adaptive multi-scale fusion module includes a dilated convolution layer, a convolution layer, and a fully connected layer. The input feature map passes through two dilated convolution layers with different dilation rates to obtain two feature maps of different scales. The two feature maps are concatenated through channels and then pass through a convolution layer, and then through a fully connected layer and a Relu activation function, and then through a fully connected layer and a Sigmoid activation function to obtain weights of two different scales; the feature maps extracted by the two dilated convolution layers are multiplied by the weights of the corresponding scales respectively and then concatenated through channels to obtain the output feature map of the adaptive multi-scale fusion module; Train the underwater image enhancement model, and use the trained underwater image enhancement model for enhancing underwater degraded images.
2. The underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism according to claim 1, wherein The convolution kernel size of one dilated convolution layer of the adaptive multi-scale fusion module is 3×3, and the dilation rate is 2. The convolution kernel size of the other dilated convolution layer is 3×3, and the dilation rate is 1.
3. The underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism according to claim 1, characterized in that The bottleneck layer includes an improved channel spatial attention module. The improved channel spatial attention modules of the decoder and the bottleneck layer have the same structure, both including a channel attention module and a spatial attention module. In the channel attention module, the input feature map undergoes max pooling and average pooling operations respectively to obtain two feature maps. These two feature maps sequentially pass through a fully connected layer and a PReLU activation function, and then pass through another fully connected layer and a Sigmoid activation function to obtain the weights of max pooling and the weights of average pooling. After performing element-wise summation on the weights of max pooling and the weights of average pooling, and then passing through a Sigmoid activation function to obtain the channel attention weights, the channel attention weights are multiplied by the input feature map of the channel attention module to obtain the channel attention feature map. The channel attention feature map serves as the input feature map of the spatial attention module. In the spatial attention module, the channel attention feature map undergoes max pooling and average pooling operations respectively to obtain two feature maps. After concatenating these two feature maps along the channel dimension and then passing through a convolution and a Sigmoid activation function to obtain the spatial attention weights, the spatial attention weights are multiplied by the input feature map of the spatial attention module to obtain the output feature map of the improved channel spatial attention module.
4. The underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism according to any one of claims 1 to 3, characterized in that, The convolutional kernels of the deep convolutional layers of the encoder and the decoder both have a size of 3×3 and a stride of 1.
Citation Information
Patent Citations
Crowd counting method based on multi-scale context enhancement network
CN112132023A
Photovoltaic panel crack detection method based on dual-channel multi-scale attention mechanism
CN116402761A