Fast Underwater Image Enhancement Method Based on Two-Stage Attention Mechanism

The two-stage attention mechanism in the simple CNN framework addresses the high computational load and poor generalization of existing methods by capturing image features and adapting to underwater conditions, resulting in efficient and accurate image enhancement.

CN116416157BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310277944.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-07-15
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

The existing underwater image enhancement methods have large calculations and poor generalization performance, which cannot effectively deal with the quality degradation of underwater images.

Method used

A fast underwater image enhancement method based on a two-stage attention mechanism is adopted, and image features are extracted using the self-attention mechanism, and combined with the cross-attention mechanism of the transmission graph, a convolutional block, self-attention module and cross-attention module are constructed through a simple convolutional neural network design, and a variety of loss functions are used for training to improve the enhancement effect.

Benefits of technology

A lightweight underwater image enhancement network is realized, which can adapt to the diversity of different underwater scenes, generate more accurate enhanced images, and improve computing efficiency and generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416157B_ABST
    Figure CN116416157B_ABST
Patent Text Reader

Abstract

A fast underwater image enhancement method based on a two-stage attention mechanism belongs to the field of computer vision technology. The specific steps are as follows: Divide the existing open-source dataset into a training set and a test set, and adjust the image size to h*t as the input a. Use the general dark channel prior method to generate the transmission map of the dataset, and adjust the size of the transmission map to h*t as another input T. Construct an underwater image enhancement model based on two-stage attention. The underwater image enhancement model includes a convolutional block CONV BLOCK, a self-attention module SA, and a cross-attention module TA. Construct a loss function, use the training set to train the underwater image enhancement model based on two-stage attention, and obtain a trained underwater image enhancement network. Input the test set, and use the trained underwater image enhancement network to perform image enhancement and output the results. The present invention solves the technical problems of large computational complexity and poor generalization performance of existing underwater image enhancement methods while enhancing underwater images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a fast underwater image enhancement method based on a two-stage attention mechanism. Background Art

[0002] As an indispensable part of an autonomous underwater vehicle, a machine vision system has been widely used in ocean observation, exploration, and operation in extreme underwater environments. However, underwater images usually suffer from severe quality degradation caused by wavelength-dependent attenuation, such as color deviation, low contrast, and detail blurring. These degraded underwater images have a negative impact on the development of high-level vision tasks (such as recognition and tracking). Therefore, underwater image enhancement, which restores clear images from degraded images, is of great significance for vision-guided autonomous underwater vehicles.

[0003] Existing underwater image enhancement methods can be divided into non-physical model-based methods, physical model-based methods, and deep learning-based methods. Non-physical model-based methods tend to adopt general image enhancement methods, namely histogram equalization, retina algorithms, and image fusion. These methods adjust image pixel values without relying on an underwater imaging model to obtain satisfactory results. Physical model-based methods usually first establish a physical degradation model of underwater images, and then use various prior assumptions to estimate unknown model parameters. Finally, by inverting this degradation process, high-quality underwater images can be obtained. Deep learning-based methods obtain clear underwater images by constructing a deep convolutional neural network and training it with a large number of underwater images and high-quality reference images.

[0004] However, existing underwater image enhancement methods have certain drawbacks: Non-physical model-based methods are prone to introducing color deviation and artifacts and may also exacerbate noise because they do not consider the optical characteristics of underwater imaging; Physical model-based methods have poor generalization because it is difficult to accurately estimate the parameters of the imaging model due to the complex underwater environment; Deep learning-based underwater image enhancement methods can basically produce good enhancement results through complex network structures, which inevitably requires a large amount of computer memory. Therefore, it is necessary to further study suitable underwater image enhancement methods. Summary of the Invention

[0005] Technical Problems to be Solved

[0006] To avoid the deficiencies of the prior art, the present invention provides a fast underwater image enhancement method based on a two-stage attention mechanism. It is designed based on a simple convolutional neural network, extracts image features through a self-attention mechanism, and considering the special imaging characteristics underwater, introduces a transmission map combined with a cross-attention mechanism to solve the technical problems of large computational complexity and poor generalization performance of existing underwater image enhancement methods while enhancing underwater images.

[0007] The technical solution of the present invention is: a fast underwater image enhancement method based on a two-stage attention mechanism, and the specific steps are as follows:

[0008] Step 1: Divide the existing open-source dataset into a training set and a test set, and adjust the image size to h*t as the input a;

[0009] Step 2: Then use the general dark channel prior method to generate the transmission map of the dataset from the existing open-source dataset, and adjust the size of the transmission map to h*t as another input T;

[0010] Step 3: Construct an underwater image enhancement model based on two-stage attention; the underwater image enhancement model includes a convolutional block CONV BLOCK, a self-attention module SA, and a cross-attention module TA;

[0011] The self-attention module SA is used to capture more global context information and enhance feature representation; the cross-attention module TA is used to capture the dependence relationship between image features and transmission map features, improving the adaptability of the network to different turbid regions of underwater images; an instance normalization layer is added to the convolutional block to ensure the feature independence between each image, and the input image is spliced with the output of each convolutional block in a skip connection manner to avoid overfitting and obtain a more accurate enhanced image;

[0012] Step 4: Construct a loss function, and use the training set to train the underwater image enhancement model based on two-stage attention to obtain a trained underwater image enhancement network;

[0013] Step 5: Input the test set, and use the trained underwater image enhancement network to perform image enhancement and output the result.

[0014] A further technical solution of the present invention is: in the step 1, a real underwater image is acquired by an image acquisition module and stored in a readable storage medium to construct an image library as the existing open-source dataset.

[0015] A further technical solution of the present invention is that the open-source dataset includes UIEB and LISU, and UIEB and LISU are respectively randomly divided into a training set and a test set; UIEB is divided into 800 training sets and 90 test sets, and LISU is divided into 4500 training sets and 504 test sets.

[0016] A further technical solution of the present invention is that the image size h*t is 256*256.

[0017] A further technical solution of the present invention is that the convolutional block in step 3 includes three parts, and each convolutional block makes a skip connection with the input image; the first part and the second part are the same, both including a convolutional layer, a dropout layer, a RELU layer, and an instance normalization layer, and the third part includes a convolutional layer, a dropout layer, and an instance normalization layer.

[0018] A further technical solution of the present invention is that the convolutional block and the skip connection can prevent the network from overfitting to the training data, and the instance normalization layer is used to ensure the feature independence between each image, thereby supporting the generalization of the network.

[0019] A further technical solution of the present invention is that the specific steps of step 3 are as follows: the input a first undergoes an operation of a 3×3 convolutional kernel-RELU activation function to obtain a feature map F1, and enters the self-attention module SA to obtain a feature map F2; then, F2 and the input T are used as the inputs of the cross-attention module TA to generate a feature map F3. F3 passes through three parts of convolutional blocks, and the output of each convolutional block part is connected to the input image a in a skip connection manner, and finally an enhanced image is generated through a 3×3 convolutional kernel.

[0020] A further technical solution of the present invention is that the specific operation of the self-attention module SA in step 3 is to convert the feature map F1 into query, key, and value features through a 1×1 convolution: Query = Q(F1), Key = K(F1), and Value = V(F1). After passing through the softmax layer, an attention map is obtained. Finally, the attention map is multiplied by the value of the original feature F1, and the sum is taken corresponding to each element in the original feature F1 to obtain an enhanced feature. The specific formula is as follows:

[0021] F2 = F1 + softmax(Q(F1) T K(F1))V(F1).

[0022] A further technical solution of the present invention is that the specific operation of the cross-attention module TA based on the transmission map in step 3 is to obtain the value feature by using 1×1 convolution on the input feature map F2: Value = V(F2), and convert the input transmission map T into query and key features through 1×1 convolution: Query = Q(T) and Key = K(T), and obtain the spatial cross-attention Figure 1 -A, and finally obtain the fused feature. The specific formula is as follows:

[0023] F3 = F2 + softmax(1 - Q(T) T K(T))V(F2).

[0024] A further technical solution of the present invention is that in step 4, the loss function sequentially adopts the mean square loss function L2, Charbonnier_Loss L cha , the structural similarity loss function L ssim , the perceptual loss function L per , and the LAB color space loss function L lab for calculation;

[0025] S4.1: Calculate the mean square loss function, which is used to reduce the pixel difference between the enhanced image and the reference image. The formula is as follows:

[0026]

[0027] Among them, is the image enhanced by step 3, J(i, j) is the reference image, and H and W are the height and width of the image respectively;

[0028] S4.2: Calculate the Charbonnier_Loss function. The formula is as follows:

[0029]

[0030] Among them, ε = 1e -6 ;

[0031] S4.3: Calculate the structural similarity loss function, which is used to calculate the structural similarity between two images to reduce the difference. The formula is as follows:

[0032]

[0033]

[0034] Among them, μ J respectively represent the means of the images and J, σ Jrespectively represent the images and the variance of J, represents the image and the covariance of the image J, C1 and C2 are constants;

[0035] S4.4: Calculate the perceptual loss function, and the formula is as follows:

[0036]

[0037] where φ is the operation of extracting features using the VGG-19 network;

[0038] S4.5: Calculate the LAB color space loss function to remove the color scattering phenomenon of underwater images; since the color distribution in the LAB space is better than that in the RGB space, the loss function is constructed in the LAB space, and the formula is as follows:

[0039]

[0040] where L lab in L represents the luminance channel of the image, A and B are two color channels, and Q represents the quantization operation;

[0041] S4.6: Calculate the total loss function L Loss , and the formula is as follows:

[0042] L Loss =λ1L2 + λ2L cha +λ3L ssim +λ4L per +λ5L lab

[0043] where λ1 = λ2 = λ3 = λ4 = 1, λ5 = 1e -6 .

[0044] Beneficial effects

[0045] The beneficial effects of the present invention are as follows: The present invention uses self-attention to capture more global context information, thereby enhancing feature representation; designs a cross-attention module based on the transmission map, combines the uneven distribution characteristics of the degradation of underwater image quality with the feature map obtained by self-attention, facilitates capturing the dependence relationship between the image features and the transmission map features, improves the adaptability of the network to different turbid regions of underwater images, and thus enables the enhanced image to better adapt to the diversity of underwater scenes; adds an instance normalization layer to three convolutional blocks to ensure the feature independence between each image, and the input image is spliced with the output of each convolutional block in a skip connection manner to avoid overfitting, and a more accurate enhanced image is obtained.

[0046] (1) Based on a simple convolutional neural network, the present invention designs an underwater image enhancement network with a simple model and fewer parameters. Then, considering the pixel features, texture features, and color features of underwater images, five loss functions are combined to train the network, achieving lightweight and real-time performance of the underwater enhancement network while producing good enhancement results.

[0047] (2) The present invention uses a two-stage attention mechanism to extract underwater image features. First, the self-attention mechanism is used to capture more important information. Then, considering the uneven distribution of underwater image quality degradation, a cross-attention module based on the transmission map is introduced to capture the dependence relationship between image features and transmission map features, enabling the model to be widely applied in different underwater scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow module diagram of the fast underwater image enhancement method based on the two-stage attention mechanism of the present invention;

[0049] Figure 2 is a schematic flow diagram of the cross-attention module based on the transmission map of the present invention;

[0050] Figure 3 is a schematic flow diagram of 3 convolutional blocks of the present invention;

[0051] Figure 4 is a visual comparison diagram showing the enhancement results of the UIEB dataset;

[0052] Figure 5 is a visual comparison diagram showing the enhancement results of the LISU dataset;

[0053] Figure 6 is a visual comparison diagram of the test results on the UFO-120 dataset during training on Train-L4500. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0055] An embodiment of the present invention is a fast underwater image enhancement method based on a two-stage attention mechanism, which is designed based on a simple CNN network. The features of the image are extracted through the self-attention mechanism. Considering the special imaging characteristics of underwater, the combination of the transmission map and cross-attention is introduced to solve the technical problems of large computational complexity and poor generalization performance of existing underwater image enhancement methods while enhancing underwater images.

[0056] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0057] The fast underwater image enhancement method based on two-stage attention in this embodiment is as follows Figure 1 , including the following steps:

[0058] Step 1: Randomly divide the existing open-source datasets UIEB (Underwater Image Enhancement Benchmark Dataset) and LISU (Large Scale Underwater Image Dataset) into a training set and a test set respectively. UIEB is divided into 800 training images and 90 test images, and LISU is divided into 4500 training images and 504 test images. Resize the image size to 256*256 as the input a.

[0059] Step 2: Then, use the general dark channel prior method to generate the transmission map of the dataset from the existing open-source datasets, and also resize it to 256*256 as another input T.

[0060] Step 3: Construct an underwater image enhancement framework based on two-stage attention; the underwater image enhancement model includes a convolutional block CONV BLOCK, a self-attention module SA, and a cross-attention module TA; the self-attention module SA captures more global context information to enhance the feature representation; the cross-attention module TA captures the dependence between the image features and the transmission map features to improve the adaptability of the network to different turbid regions of underwater images; an instance normalization layer is added to the convolutional block to ensure the feature independence between each image. The input image is concatenated with the output of each convolutional block in a skip connection manner to avoid overfitting and obtain a more accurate enhanced image.

[0061] The convolutional block includes three parts, and each convolutional block is connected to the input image in a skip connection manner; the first part and the second part are the same, both including a convolutional layer, a dropout layer, a RELU layer, and an instance normalization layer, and the third part includes a convolutional layer, a dropout layer, and an instance normalization layer. The convolutional block and the skip connection can prevent the network from overfitting to the training data, and the instance normalization layer is used to ensure the feature independence between each image, thereby supporting the generalization of the network.

[0062] As Figure 1 shown, the input a first undergoes an operation of a 3*3 convolutional kernel-RELU activation function to obtain a feature map F1, and then obtains a feature map F2 through the self-attention module. Then, F2 and the input T are used as the inputs of the cross-attention module to generate a feature map F3. F3 passes through three convolutional blocks, and the output of each convolutional block is connected to the input image a in a skip connection manner. Each convolutional block includes a convolutional layer, a dropout layer, an instance normalization layer, and a RELU layer for feature enhancement, and finally generates an enhanced image through a convolutional operation.

[0063] Step 4: Design a loss function and use the divided training set to train the underwater image enhancement framework based on two-stage attention.

[0064] Step 5: Input the test set (90 pairs), and use the trained underwater image enhancement network to enhance the real underwater images and output the results.

[0065] The specific operation in the self-attention module described in Step 3 is to convert the feature map F1 into query, key, and value features through a 1×1 convolution: Query = Q(F1), Key = K(F1), and Value = V(F1). After passing through the softmax layer, an attention map is obtained. Finally, the attention map is multiplied by the value of the original feature F1, and the sum is calculated corresponding to each element in the original feature F1 to obtain the enhanced feature. The specific formula is as follows:

[0066] F2 = F1 + softmax(Q(F1) T K(F1))V(F1)

[0067] For the cross-attention module based on the transmission map, as Figure 2 shown, use a 1×1 convolution on the input feature map F2 to obtain the value feature: Value = V(F2). Convert the input transmission map T into query and key features through a 1×1 convolution: Query = Q(T) and Key = K(T). Obtain the spatial cross-attention Figure 1 -A, and finally obtain the fused feature. The specific formula is as follows:

[0068] F3 = F2 + softmax(1 - Q(T) T K(T))V(F2)

[0069] The convolution block in Step 3 is as Figure 3 shown. First, pass through two groups of the same convolution units, including the convolution kernel, dropout, RELU, and instance normalization. Then, change the dimension through the convolution kernel, dropout, and instance normalization layer. Adding the instance normalization layer is to ensure the feature independence between each image.

[0070] In Step 4, the loss function sequentially uses the mean square loss function (L2), Charbonnier_Loss (L cha ), structural similarity loss function (L ssim ), perceptual loss function (L per ), and LAB color space loss function (L lab ) for calculation, as follows:

[0071] Mean Square Loss Function: To reduce the pixel difference between the enhanced image and the reference image, the following formula is used for optimization:

[0072]

[0073] where is the enhanced image, J(i,j) is the reference image, and H and W are the height and width of the image respectively.

[0074] Charbonnier_Loss: Due to the addition of a regularization term, this loss function is more robust and more stable during training. Empirically, ε = 1e -6 :

[0075]

[0076] Structural Similarity Loss Function: Considering the structural similarity between the enhanced image and the reference image, the structural similarity between the two images is calculated to reduce the difference. Among them μ J represent the means of images and J respectively, σ J represent the variances of images and J respectively, represents the covariance between image and image J, and C1, C2 are constants:

[0077]

[0078]

[0079] Perceptual Loss Function: Considering the high-level semantic features of the enhanced image and the reference image, φ is an operation to extract features using the VGG-19 network:

[0080]

[0081] LAB Color Space Loss Function: To remove the color scattering phenomenon of underwater images, since the color distribution in the LAB space is better than that in the RGB space, the loss function is constructed in the LAB space:

[0082]

[0083] L lab In this case, L represents the luminance channel of the image, A and B are two color channels, and Q represents the quantization operation.

[0084] The total loss function L is as follows. According to multiple experiments, the weights λ1 = λ2 = λ3 = λ4 = 1, λ5 = 1e -6 .

[0085] L = λ1L2 + λ2L cha + λ3L ssim + λ4L per + λ5L lab 。

[0086] Example 1

[0087] In this example, UIEB and LISU are used as the main data, serving as the basic dataset for underwater image enhancement and being widely applied to underwater image enhancement tasks. The UIEB dataset contains 890 pairs of data (real underwater images and reference images), as well as 60 real underwater images (without reference images). This invention only uses the 890 pairs of the dataset and divides them into: 800 pairs of images as the training set and 90 pairs of images as the test set. The LISU dataset contains 5004 pairs of data (real underwater images and reference images) and is divided into: 4500 pairs of images as the training set and 504 pairs of images as the test set.

[0088] This invention takes real underwater images and their corresponding transmission maps as input, and calculates the loss function using reference images. All images are resized to 256*256. This invention uses an NVIDIA GeForce RTX 3080 graphics card to train the entire network. During the training process, the Adam optimizer is used to train for 400 iterations, the learning rate is set to 0.0002, and the weight of the loss function except for L lab has a weight coefficient of 10 -6 , and the weight coefficients of the remaining loss functions are all 1.

[0089] Based on the above description, Table 1 shows the test results of the method of this invention and some deep learning-based methods on the UIEB and LISU datasets, and uses the Peak Signal-to-Noise Ratio (PSNR) and the Structural Similarity Index Measurement (SSIM) to describe the accuracy of the test results. It can be found that the method of this invention has obtained the best results in these two metrics.

[0090] The calculation formulas for the Peak Signal-to-Noise Ratio (PSNR) and the Structural Similarity Index Measurement (SSIM) are as follows:

[0091]

[0092] MSE is the mean square error between the original image and the enhanced image, and MAX I : represents the maximum value of the image color;

[0093]

[0094] μ x 、 μ y represent the means of images x and y respectively, and σ x 、 σ y represent the variances of images x and y respectively, and σ xy represents the covariance of images x and y. C1 and C2 are constants, and the value range of SSIM is [0, 1]. The larger the value, the smaller the image distortion.

[0095] PSNR is used to measure the difference between two images. SSIM evaluates from the illumination, contrast, and structure of the images and quantifies the similarity degree of the two images. The higher these two metrics are, the better the enhancement effect.

[0096] Table 1: Comparison of the method of the present invention with other deep learning methods on UIEB and LISU datasets

[0097]

[0098] In addition, Figure 4 a visual comparison graph of the enhancement results of the UIEB dataset is shown, Figure 5 a visual comparison graph of the enhancement results of the LISU dataset is shown. It can be seen that the results of the present invention are closest to the reference image in both human visual perception and texture details. Shallow-Uwnet shows the most serious color artifacts and haze. The effects of UGAN and FUnIE-GAN in restoring color and structural texture details are not ideal. Ucolor produces a better color appearance but cannot enhance details well. Although U-shape has better visual quality for the human eye, there are color artifacts in some areas.

[0099] Example 2

[0100] In this example, LISU and UFO are used as the main data. 4500 pairs of images of LISU are used as the training set, and 120 pairs of data of UFO are used as the test set to conduct a cross-validation experiment. The experimental configuration is the same as that of Example 1 to prove the generalization of the present invention in the field of image enhancement.

[0101] Based on the above description, Table 2 shows the test results of the method of the present invention and some deep learning-based methods on the UFO dataset. The method of the present invention obtains the best result in terms of the PSNR metric and the second-best result in terms of SSIM.

[0102] Table 2: Comparison of the method of the present invention with other deep learning methods on the UFO dataset

[0103]

[0104] Figure 6 The visualization comparison graph of the test results on the UFO-120 dataset when training on Train-L4500 is given. It can be seen that our method achieves results closer to the reference image than other methods.

[0105] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A fast underwater image enhancement method based on a two-stage attention mechanism, characterized in that The specific steps are as follows: Step 1: Divide the existing open-source dataset into a training set and a test set, and adjust the image size to h*t as input a; Step 2: Then, use the general dark channel prior method to generate the transmission map of the dataset from the existing open-source dataset, and adjust the size of the transmission map to h*t as another input T; Step 3: Construct an underwater image enhancement model based on two-stage attention; the underwater image enhancement model includes a convolutional block CONV BLOCK, a self-attention module SA, and a cross-attention module TA; The self-attention module SA is used to capture more global context information and enhance feature representation; the cross-attention module TA is used to capture the dependence between image features and transmission map features, improving the adaptability of the network to different turbid regions of underwater images; an instance normalization layer is added to the convolutional block to ensure the feature independence between each image, and the input image is concatenated with the output of each convolutional block in a skip connection manner to avoid overfitting, obtaining a more accurate enhanced image; Step 4: Construct a loss function, and use the training set to train the underwater image enhancement model based on two-stage attention to obtain a trained underwater image enhancement network; Step 5: Input the test set, and use the trained underwater image enhancement network to perform image enhancement and output the results.

2. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 1, wherein: In Step 1, a real underwater image is obtained by an image acquisition module and stored in a readable storage medium to construct an image library as the existing open-source dataset.

3. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 1, characterized in that: The open-source dataset includes UIEB and LISU, and UIEB and LISU are randomly divided into a training set and a test set respectively; UIEB is divided into 800 training sets and 90 test sets, and LISU is divided into 4500 training sets and 504 test sets.

4. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 1, wherein: The image size h*t is 256*256.

5. A fast underwater image enhancement method based on a two-stage attention mechanism according to any one of claims 1-4, characterized in that: The convolutional block in Step 3 includes three parts, and each convolutional block is connected to the input image in a skip connection; the first part and the second part are the same, both including a convolutional layer, a dropout layer, a RELU layer, and an instance normalization layer, and the third part includes a convolutional layer, a dropout layer, and an instance normalization layer.

6. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 5, wherein: The convolutional block and the skip connection can prevent the network from overfitting to the training data, and the instance normalization layer is used to ensure the feature independence between each image, thus supporting the generalization of the network.

7. A fast underwater image enhancement method based on a two-stage attention mechanism according to claim 6, characterized in that: The specific steps of Step 3 are as follows: Input a first undergoes a 3×3 convolutional kernel-RELU activation function operation to obtain a feature map F1, and enters the self-attention module SA to obtain a feature map F2; then, F2 and input T are used as the inputs of the cross-attention module TA to generate a feature map F3, and F3 passes through three parts of the convolutional block. The output of each convolutional block part is connected to the input image a in a skip connection manner, and finally, an enhanced image is generated through a 3×3 convolutional kernel.

8. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 7, wherein: The specific operation of the self-attention module SA in step 3 is to convert the feature map F1 into query, key, and value features through a 1×1 convolution: Query = Q(F1), Key = K(F1), and Value = V(F1). After passing through the softmax layer, an attention map is obtained. Finally, the attention map is multiplied by the value of the original feature F1, and the results are summed with each element in the original feature F1 to obtain the enhanced feature. The specific formula is as follows: F2 = F1 + softmax(Q(F1) T K(F1))V(F1).

9. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 8, wherein: The specific operation of the cross-attention module TA based on the transmission map in step 3 is to use a 1×1 convolution on the input feature map F2 to obtain the value feature: Value = V(F2). The input transmission map T is converted into query and key features through a 1×1 convolution: Query = Q(T) and Key = K(T). The spatial cross-attention map 1-A is obtained through matrix multiplication, and finally the fused feature is obtained. The specific formula is as follows: F3 = F2 + softmax(1 - Q(T) T K(T))V(F2).

10. The fast underwater image enhancement method based on a two-stage attention mechanism according to claim 9, wherein: In the said step 4, the loss functions are successively calculated using the mean square loss function L2, the Charbonnier_Loss L cha , the structural similarity loss function L ssim , the perceptual loss function L per , and the LAB color space loss function L lab ; S4.1: Calculate the mean square loss function to reduce the pixel difference between the enhanced image and the reference image. The formula is as follows: Among them, is the image enhanced by step 3, J(i, j) is the reference image, and H and W are the height and width of the image respectively; S4.2: Calculate the Charbonnier_Loss function. The formula is as follows: where ε = 1e -6 ; S4.3: Calculate the structural similarity loss function to calculate the structural similarity between two images and reduce the difference. The formula is as follows: Among them, μ J respectively represent the mean values of images and J, σ J respectively represent the variances of images and J, represents the covariance of image and image J, and C1, C2 are constants; S4.4: Calculate the perceptual loss function. The formula is as follows: where φ is the operation of extracting features using the VGG-19 network; S4.5: Calculate the LAB color space loss function to remove the color scattering phenomenon in underwater images; since the color distribution in the LAB space is better than that in the RGB space, the loss function is constructed in the LAB space. The formula is as follows: Among them, L lab where L in lab represents the luminance channel of the image, A and B are two color channels, and Q represents the quantization operation; S4.6: Calculate the total loss function L Loss , and the formula is as follows: L Loss = λ1L2 + λ2L cha + λ3L ssim + λ4L per + λ5L lab Among them, λ1 = λ2 = λ3 = λ4 = 1, λ5 = 1e -6 .

Citation Information

Patent Citations

  • Multi-stage progressive underwater image enhancement method

    CN114445292A

  • RGB-T target tracking network based on global attention convolution-Transform

    CN115375948A