Underwater image enhancement method based on double-domain attention mechanism
By introducing a two-domain attention mechanism enhancement method in underwater image processing, problems such as color distortion, low contrast, and blurred details of underwater images are solved, and a significant improvement in image quality is achieved.
Patent Information
- Application Number
- CN202510011790.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-05
- Publication Date
- 2025-05-06
AI Technical Summary
Due to the absorption and scattering effects of light in water, underwater images often face problems such as color distortion, low contrast, and blurred details. The existing technology is difficult to effectively solve these problems.
The underwater image enhancement method based on the dual-domain attention mechanism is adopted, and the downsampled feature map is obtained through the convolution module, and a network composed of encoder and decoder is used to combine the dual-domain attention module to fully explore the frequency domain and spatial domain information to achieve efficient image enhancement.
The quality of underwater images was significantly improved, and through testing on EUVP and LSUI public data sets, significant effects in color correction, detail texture recovery and deblurring were achieved, and the peak signal-to-noise ratio and structural similarity were higher than that of the existing technology.
Smart Images

Figure CN119941537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an underwater image enhancement method. Technical Background
[0002] With the rapid development of industry and the increasing demand for resources, the development and utilization of marine resources has become an important direction to promote the sustainable development strategy. As the core medium for carrying underwater information, underwater images are widely used in many fields such as ocean exploration, ecological monitoring, underwater robot navigation and control, underwater equipment inspection, and underwater archaeology. However, the physical and optical properties of the underwater environment are significantly different from those of the terrestrial environment, resulting in underwater images often facing problems such as color distortion, low contrast, and blurred details. These problems seriously restrict the application of underwater images in engineering practice.
[0003] The degradation of underwater images is mainly reflected in the following two aspects: First, due to the different degrees of absorption of light of different wavelengths in water, red light decays the fastest and blue light decays the slowest, causing underwater images to appear blue-green and the true color of objects is difficult to distinguish; second, suspended particles and dissolved substances in the water will cause light scattering, and the light entering the camera is a superposition of the reflected light and scattered light of the target object, resulting in a decrease in image clarity, loss of detailed features, and a fogging effect.
[0004] Traditional underwater image enhancement technology is usually achieved by improving hardware equipment or physical model-based methods, but the hardware equipment is expensive and difficult to adapt to complex underwater environments, and the physical model method has limited effect in complex degradation scenarios. With the rapid development of deep learning, methods based on convolutional neural networks (CNNs) have been widely used in underwater image enhancement tasks and have achieved remarkable results. CNNs are good at extracting local spatial features and can effectively restore blurred or lost detail information, but due to the limitations of their local receptive field, they are insufficient in capturing long-range dependencies and global features.
[0005] On the other hand, the Transformer model has demonstrated excellent capabilities in long-distance dependency modeling through the attention mechanism. It can capture the global relationship between distant pixels in the image and improve the image enhancement effect from a global semantic level. However, existing deep learning methods, whether CNN or Transformer, ignore the significant difference in frequency distribution between degraded images and clear images. The detail information of degraded images is usually concentrated in the high-frequency part, while the global structural information is mostly concentrated in the low-frequency part. Therefore, it is difficult to fully restore the details and structure of degraded images by relying solely on modeling the relationship between pixels. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides an underwater image enhancement method network based on a dual-domain attention mechanism, which relates to the field of image processing technology. First, the underwater image is preprocessed, and then the downsampled feature map is obtained through a convolution module, and then the enhanced image is obtained through an encoder and a decoder, and finally the network composed of the encoder and the decoder is trained. The present invention utilizes a convolution module and a dual-domain attention mechanism to fully exploit the frequency domain and spatial domain information, thereby achieving efficient enhancement of degraded images and significantly improving the quality of underwater images; the test results on the EUVP and LSUI public data sets show that significant effects have been achieved in color correction, detail texture restoration, and deblurring of underwater images. The present invention can fully exploit the frequency domain and spatial domain information, thereby achieving efficient enhancement of degraded images.
[0007] The technical solution steps adopted by the present invention are specifically as follows:
[0008] Step 1, preprocessing the underwater image;
[0009] Step 2, use the convolution module to obtain the downsampled feature map;
[0010] Step 3, inputting the underwater image into an encoder for encoding to obtain a coding feature map;
[0011] Step 4, input the encoded feature map into the decoder for decoding to obtain an enhanced image;
[0012] Step 5: Use the loss function to train the network consisting of the encoder and decoder.
[0013] Further, in step 1, the underwater image preprocessing process is as follows:
[0014] First, the underwater image size is uniformly transformed into H*W, where H represents the height of the image and W represents the width of the image;
[0015] Then, the underwater images are flipped and rotated to achieve the purpose of enhancing the underwater image dataset.
[0016] Further, in step 2, the method for obtaining the downsampled feature map is as follows:
[0017] Firstly, the preprocessed underwater images are downsampled by 2 times and 4 times respectively to obtain 2 times downsampled images and 4 times downsampled images;
[0018] Then, the 2x downsampled image and the 4x downsampled image are input into the convolution module respectively to obtain a 2x downsampled feature map and a 4x downsampled feature map;
[0019] Finally, the underwater image used as a reference is also downsampled by 2 times and 4 times to obtain a reference downsampled image.
[0020] Furthermore, the dimension of the 2x downsampled feature map is 64*128*128; the dimension of the 4x downsampled feature map is 128*64*64.
[0021] Further, in step 2, the convolution module sequentially includes a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer and a RELU activation function.
[0022] Further, in step 3, the encoder sequentially includes a resnet module, a dual-domain attention module, a feature fusion module, a resnet module, a dual-domain attention module, a feature fusion module, a resnet module and a dual-domain attention module.
[0023] Further, in step 3, the resnet module includes M identical resnet units, and the value range of M is [3,10]; each resnet unit includes a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a deep convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function and a residual connection.
[0024] Furthermore, in step 3, the dual-domain attention module first uses convolution to implicitly decompose different frequency features; then uses the channel attention mechanism to weightedly fuse different frequency domain features to obtain frequency domain attention features; then uses Transformer to extract self-attention features; finally, uses the channel attention mechanism to weightedly fuse frequency domain attention features and self-attention features to obtain a dual-domain attention feature map.
[0025] Further, in step 4, the decoder includes a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module and a resnet module in sequence.
[0026] Further, in step 4, the feature fusion module first splices the input downsampled feature map according to channels; the feature map after channel splicing passes through a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 1*1, a normalization layer and a RELU activation function in sequence; the feature fusion module finally outputs a fused feature map;
[0027] Finally, the fused feature map is input into the convolution module, and the output size is The enhanced image and size are Enhanced image.
[0028] Further, in step 6, the loss function L is as follows:
[0029]
[0030] Among them, I i and J i The predicted image and reference image are of size i respectively; FFT stands for Fast Fourier Transform; N i represents the normalization parameter under the image of size i, and the value of the normalization parameter is the product of the image length and width; i represents different image sizes, and the value range of i is [1,3]; I1 represents an underwater image of size 256*256, I2 represents an underwater image of size 128*128 downsampled by 2 times; I3 represents an underwater image of size 64*64 downsampled by 4 times; J1 represents a clear underwater image of size 256*256; J2 represents a clear underwater image of size 128*128 downsampled by 2 times; J3 represents a clear underwater image of size 64*64 downsampled by 4 times.
[0031] The technical effects of the present invention are as follows:
[0032] An underwater image enhancement method based on a dual-domain attention mechanism proposed in the present invention significantly improves the quality of underwater images; test results on the EUVP and LSUI public datasets show that remarkable effects are achieved in color correction, detail texture restoration, and deblurring of underwater images; in the LSUI public dataset, the peak signal-to-noise ratio and structural similarity of the method proposed in the present invention are higher than those of the prior art; in the EUVP public dataset, the peak signal-to-noise ratio and structural similarity of the method proposed in the present invention are also higher than those of the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a diagram of the underwater image enhancement network structure based on a dual-domain attention mechanism;
[0034] Figure 2 This is the structural diagram of the dual-domain attention module. DETAILED DESCRIPTION
[0035] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0036] In view of the shortcomings of the prior art, the present invention provides an underwater image enhancement method based on a dual-domain attention mechanism.
[0037] S1. Preprocess the underwater image. First, the underwater image size is uniformly transformed into H*W size. In this embodiment, the pixel H=W=256 is selected. Then, the image is randomly flipped and rotated to achieve the purpose of enhancing the underwater image data set.
[0038] S2. Use the convolution module to obtain the downsampled feature map: Figure 1 As shown, the preprocessed underwater image is downsampled by 2 times and downsampled by 4 times, respectively, to obtain a 2 times downsampled image and a 4 times downsampled image, respectively; the 2 times downsampled image and the 4 times downsampled image are input into the convolution module, respectively, to obtain a 2 times downsampled feature map and a 4 times downsampled feature map, respectively.
[0039] At the same time, the reference underwater images are also downsampled by 2 times and 4 times.
[0040] The feature map size obtained by downsampling the image 2 times is 64*128*128, and the feature map size obtained by downsampling the image 4 times is 128*64*64.
[0041] The convolution module includes a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer and a RELU activation function.
[0042] S3. Input the original underwater image to the encoder for encoding:
[0043] The encoder includes a resnet module, a dual-domain attention module, a feature fusion module, a resnet module, a dual-domain attention module, a feature fusion module, a resnet module, and a dual-domain attention module. The resnet module includes M identical resnet units, where the value of M is in the range of [3,10]; each resnet unit includes a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a deep convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, and a residual connection.
[0044] The dual-domain attention module in the encoder first uses convolution to implicitly decompose different frequency features; then uses the channel attention mechanism to weightedly fuse different frequency domain features to obtain frequency domain attention features; then uses Transformer to extract self-attention features; finally, uses the channel attention mechanism to weightedly fuse frequency domain attention features and self-attention features to obtain a dual-domain attention feature map.
[0045] The dual-domain attention module structure is as follows Figure 2As shown in the figure, given the feature map X output by the resnet module, a series of convolutional layers are first used to extract the multi-frequency components in the input features. The feature map X passes through two convolutional networks to obtain feature maps X1 and X2 in different frequency domains respectively;
[0046] X1=Conv1(X)
[0047] X2=Conv2(X1)
[0048] Among them, Conv1 and Conv2 represent 3×3 convolution layers corresponding to different frequency domain feature maps; X1 and X2 are added according to the corresponding elements for preliminary feature fusion, and then pass through the global average pooling layer to extract the global statistical information of each channel as the input for generating attention weights; after two convolution layers, the attention weight W generated by each frequency feature is obtained:
[0049] W = Conv(Conv(GAP(X1+X2)))
[0050] Among them, Conv represents the convolutional layer;
[0051] Next, the weights of different frequency domains are concatenated in the channel dimension, and the weights are normalized using Softmax to obtain the normalized weight W′.
[0052]
[0053] Where C represents the number of channels of X1; j represents the channel.
[0054] Split W′ into W1 and W2 according to the channel dimension, multiply the weights of different frequency components and channel attention, and obtain the frequency domain attention feature X c , frequency domain attention feature X c for:
[0055] X c =Conv(X1W1+X2W2)
[0056] At the same time, Transformer is used to extract the self-attention feature X a .
[0057] After obtaining Transformer attention and frequency domain attention, the channel attention mechanism is used to fuse the self-attention features and frequency domain attention. a and X c After splicing, it passes through the convolution layer and the GELU activation function. The convolution layer generates attention weights, and the weights are normalized using Softmax to obtain the normalized weight W. k . Use the normalized attention weights for X c Weighted, using X aUse the attention weights to reverse weight, and finally add the results to get the dual-domain attention module output feature map X out , expressed as:
[0058] W k =Sigmoid(Conv(GELU(Conv(Cat(X c ,X a )))))
[0059] X out =Conv(X c W k +X a (1-W k ))
[0060] S4. Input the encoded feature map into the decoder for decoding, gradually restore the original size of the image, and obtain an enhanced image:
[0061] The decoder consists of a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module and a resnet module in sequence.
[0062] The feature fusion module in the decoder first splices the input downsampled feature map according to the channel; the feature map after channel splicing passes through the convolution layer with a convolution kernel of 3*3, the normalization layer, the RELU activation function, the convolution layer with a convolution kernel of 1*1, the normalization layer and the RELU activation function in sequence; the feature fusion module finally outputs the fused feature map;
[0063] Finally, the fused feature map is input into the convolution module, and the output size is The enhanced image and size are Enhanced image.
[0064] S5. Use the following loss function to train the network consisting of the encoder and decoder.
[0065]
[0066] Where, I and J represent the predicted image and the reference image respectively, I1 represents an underwater image with a size of 256*256, I2 represents an underwater image with a size of 128*128 downsampled by 2 times, I3 represents an underwater image with a size of 64*64 downsampled by 4 times, J1 represents a clear underwater image with a size of 256*256, J2 represents a clear underwater image with a size of 128*128 downsampled by 2 times, J3 represents a clear underwater image with a size of 64*64 downsampled by 4 times, FFT represents fast Fourier transform, N iRepresents the normalization parameter. For an image of size i, the value of the normalization parameter is the product of the image length and width. i represents the image size, and the value range of i is [1,3].
[0067] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.
[0068] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. An underwater image enhancement method based on a dual-domain attention mechanism, characterized in that: The underwater image enhancement method comprises the following steps: Step 1, preprocessing the underwater image; Step 2, use the convolution module to obtain the downsampled feature map; Step 3, inputting the original underwater image into the encoder for encoding to obtain a coding feature map; Step 4, input the encoded feature map into the decoder for decoding to obtain an enhanced image; Step 5: Use the loss function to train the network consisting of the encoder and decoder.
2. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 1, the underwater image preprocessing process is as follows: First, the underwater image size is uniformly transformed into H*W, where H represents the height of the image and W represents the width of the image; Then, the underwater images are flipped and rotated.
3. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 2, the method for obtaining the downsampled feature map is as follows: Firstly, the preprocessed underwater images are downsampled by 2 times and 4 times respectively to obtain 2 times downsampled images and 4 times downsampled images; Then, the underwater reference image is downsampled by 2 times and 4 times to obtain a downsampled reference image; Finally, the 2x downsampled image and the 4x downsampled image are respectively input into the convolution module to obtain a 2x downsampled feature map and a 4x downsampled feature map; the dimension of the 2x downsampled feature map is 64*128*128; the dimension of the 4x downsampled feature map is 128*64*64.
4. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 2, the convolution module sequentially includes a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 1*1, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer and a RELU activation function.
5. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 3, the encoder includes a resnet module, a dual-domain attention module, a feature fusion module, a resnet module, a dual-domain attention module, a feature fusion module, a resnet module and a dual-domain attention module in sequence.
6. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 5, characterized in that: The resnet module includes M identical resnet units, and the value range of M is [3,10]; each resnet unit includes a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a deep convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function, a convolution layer with a convolution kernel of 3*3, a normalization layer, a RELU activation function and a residual connection.
7. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 5, characterized in that: The dual-domain attention module first uses convolution to implicitly decompose different frequency features; Then, the channel attention mechanism is used to weightedly fuse different frequency domain features to obtain frequency domain attention features; then Transformer is used to extract self-attention features; finally, the channel attention mechanism is used to weightedly fuse frequency domain attention features and self-attention features to obtain a dual-domain attention feature map.
8. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 4, the decoder includes a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module, a feature fusion module, a convolution module, a resnet module, a dual-domain attention module and a resnet module in sequence.
9. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 8, characterized in that: The feature fusion module first splices the input downsampled feature map according to channels; the feature map after channel splicing passes through the convolution layer with a convolution kernel of 3*3, the normalization layer, the RELU activation function, the convolution layer with a convolution kernel of 1*1, the normalization layer and the RELU activation function in sequence; the feature fusion module finally outputs the fused feature map; finally, the fused feature map is input into the convolution module, and the output size is The enhanced image and size are Enhanced image.
10. The underwater image enhancement method based on the dual-domain attention mechanism according to claim 1, characterized in that: In step 6, the loss function L is as follows: Among them, I i and J i The predicted image and reference image are of size i respectively; FFT stands for Fast Fourier Transform; N i represents the normalization parameter under the image of size i, and the value of the normalization parameter is the product of the image length and width; i represents different image sizes, and the value range of i is [1,3]; I1 represents an underwater image of size 256*256, I2 represents an underwater image of size 128*128 downsampled by 2 times; I3 represents an underwater image of size 64*64 downsampled by 4 times; J1 represents a clear underwater image of size 256*256; J2 represents a clear underwater image of size 128*128 downsampled by 2 times; J3 represents a clear underwater image of size 64*64 downsampled by 4 times.