Underwater image enhancement method

By performing color space transformation and feature extraction on underwater images, combining the channel transmission attention module and the cross-color space Transformer module, the problem of insufficient visual effects and global consistency of the existing underwater image enhancement methods is solved, and high-quality underwater image enhancement is achieved.

CN120013839APending Publication Date: 2025-05-16DALIAN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510048638.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing underwater image enhancement methods have problems with poor visual effects and poor global consistency, and it is difficult to effectively improve the contrast and detail clarity of underwater images.

Method used

By transforming the image to be enhanced in color space, the target features and cascading feature data of RGB, LAB and HSV data are extracted, and the attention module and the cross-color space Transformer module are used to generate the final enhanced image.

Benefits of technology

It achieves the enhancement of underwater images with good visual effects, improves the contrast and detail clarity of the image, while ensuring global consistency and reducing ghosting and artifact phenomena.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013839A_ABST
    Figure CN120013839A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater image enhancement method. The method comprises the following steps: obtaining target feature data and cascade feature data; enhanced data is obtained; fusion features are obtained; and generating a final enhanced image. The to-be-enhanced image is converted into the RGB data, the LAB data and the HSV data, and the RGB data, the LAB data and the HSV data are subjected to feature extraction, so that the number of features extracted in the feature extraction process is larger, the image enhancement effect is improved, namely, a clearer final enhanced image is obtained, and the image enhancement efficiency is improved. The channel attention module of the channel transmission attention module focuses on channels with strong features, meanwhile, a reverse medium transmission image is used as an attention image, degeneration with higher quality is extracted, image enhancement is more accurate and targeted, global consistency is enhanced through a self-attention mechanism of a cross-color space Transform module, and the image enhancement accuracy is improved. Therefore, the problems of ghosting, artifacts and the like of underwater images are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and in particular to an underwater image enhancement method. Background Art

[0002] Under the condition of increasing pressure on land space and resources, the development of underwater space has become urgent. As an important carrier and presentation form of underwater information, underwater images play an irreplaceable and important role in underwater environment detection and perception. However, plankton and small non-algae particles in the water will cause light scattering and other problems, which will lead to underwater image quality degradation such as low contrast, blurred details, color distortion, poor clarity, non-uniform illumination, and limited visual distance. These problems have brought serious challenges to subsequent work. In recent years, researchers have proposed many underwater image enhancement methods, hoping to achieve excellent human eye visual effects and reduce the difficulties in subsequent work.

[0003] Initially, researchers tended to use more traditional non-physical model methods and physical model-based methods to enhance underwater images. However, non-physical model methods have poor generalization ability, sometimes introducing too much red components, and sometimes ignoring texture information. Physical model-based methods can achieve relatively excellent results when the depth estimation is accurate due to the scientific nature of their enhancement principles. However, because the depth estimation of underwater images is easily affected by factors such as the underwater environment, too much red components are often introduced when the depth estimation is wrong.

[0004] With the development of neural networks, there are more and more underwater image enhancement methods based on learning, but the effects they achieve vary. Most underwater image enhancement methods based on learning use pure CNN network structures. Among them: although the models proposed by the WaterNet method and the WaterGAN method can reduce the haze effect in underwater images and improve their contrast, the brightness of the enhanced images is too low, showing an effect of insufficient illumination, and the overall visual effect is poor; the image enhanced by the UWCNN method is not only insufficient in brightness, but also has large-area artifacts in the originally high-brightness areas; the Ucolor method can achieve good dehazing effects and restore the color of the image to a certain extent, but the global consistency of this method is poor, and artifacts will also appear in high-brightness areas such as white. Summary of the invention

[0005] In view of the technical problems that the existing underwater image enhancement has poor visual effect, poor global consistency and other poor effects, the present invention proposes an underwater image enhancement method with good visual effect and good global consistency.

[0006] A method for underwater image enhancement comprises the following steps:

[0007] A. Obtain target feature data and cascade feature data

[0008] Performing color space transformation on the image to be enhanced to obtain RGB data, LAB data and HSV data, and performing feature extraction on the RGB data, LAB data and HSV data to obtain target feature data and cascade feature data;

[0009] B. Get enhanced data

[0010] Inputting the cascade feature data into a preset channel transmission attention module to obtain enhanced data;

[0011] C. Get fusion features

[0012] Inputting the target feature data into a preset cross-color space Transformer module to obtain a fusion feature;

[0013] D. Generate the final enhanced image

[0014] A final enhanced image of the image to be enhanced is generated according to the enhanced data and the fusion features.

[0015] Furthermore, the steps of obtaining the target feature data and the cascade feature data are as follows:

[0016] A1. Input the RGB data, LAB data and HSV data into the first layer network of a preset feature extraction module to obtain a first feature for the RGB data, a second feature for the LAB data and a third feature for the HSV data;

[0017] Performing cascade processing on the first feature, the second feature and the third feature to obtain first cascade feature data of the cascade feature data;

[0018] A2. Obtain secondary RGB data, secondary LAB data, and secondary HSV data according to the first feature, the second feature, and the third feature by a preset downsampling method, and input the secondary RGB data, the secondary LAB data, and the secondary HSV data into a second layer network in a preset feature extraction module to obtain a fourth feature for the secondary RGB data, a fifth feature for the secondary LAB data, and a sixth feature for the secondary HSV data;

[0019] Performing cascade processing on the fourth feature, the fifth feature and the sixth feature to obtain second cascade feature data of the cascade feature data;

[0020] A3. Obtain primary RGB data, primary LAB data and primary HSV data according to the fourth feature, the fifth feature and the sixth feature by the preset downsampling method, and input the primary RGB data, the primary LAB data and the primary HSV data into the third layer network in the preset feature extraction module to obtain a seventh feature for the primary RGB data, an eighth feature for the primary LAB data and a ninth feature for the primary HSV data;

[0021] Performing cascade processing on the seventh feature, the eighth feature and the ninth feature to obtain third cascade feature data of the cascade feature data;

[0022] A4. Through the preset downsampling method, target RGB data, target LAB data and target HSV data are obtained according to the seventh feature, the eighth feature and the ninth feature, and the target RGB data, target LAB data and target HSV data are input into the fourth layer network in the preset feature extraction module to obtain the target feature data.

[0023] Furthermore, the preset feature extraction module is composed of eight convolution layers with kernel_size=3, stride=1, and padding=1 to form a residual module. After the data is input into the first convolution layer, the Relu function is added to sort the data of this layer to obtain the first convolution data. After the first convolution data is input into the second convolution layer, the Relu function is added to sort the data of this layer to obtain the second convolution data. After the second convolution data is input into the third convolution layer, the Relu function is added to sort the data of this layer to obtain the third convolution data. After the third convolution data is input into the fourth convolution layer, the fourth convolution data is obtained. The first convolution data is input into the fourth convolution layer. The first layer of residual features are obtained by adding the first layer of residual features to the fourth layer of convolution data at the pixel level. The first layer of residual features are input into the fifth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the fifth layer of convolution data. The fifth layer of convolution data is input into the sixth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the sixth layer of convolution data. The sixth layer of convolution data is input into the seventh layer of convolution and then the Relu function is added to sort the data of this layer to obtain the seventh layer of convolution data. The seventh layer of convolution data is input into the eighth layer of convolution to obtain the eighth layer of convolution data. The fifth layer of convolution data and the eighth layer of convolution data are added at the pixel level to obtain the target data of the feature extraction module.

[0024] Furthermore, the cascade processing method is as follows: using the Cat function to perform a connection operation on the input tensor sequence in the first dimension.

[0025] The preset downsampling method is maximum pooling with kernel_size=2, stride=2, padding=0, which is mathematically represented as:

[0026] F Downsampling =MaxPool(F input ) (1)

[0027] Among them, F Downsampling is the data obtained after downsampling, F Input is the input data, MaxPool(*) is the maximum pooling.

[0028] Furthermore, before obtaining the enhanced data in step B, a reverse medium transmission map is first obtained, and the steps are as follows:

[0029] Depth information corresponding to the image to be enhanced is obtained, and a reverse medium transmission map corresponding to the image to be enhanced is generated through an underwater imaging model according to the depth information, wherein the mathematical representation of the underwater imaging model is:

[0030] T=e -dβ (2)

[0031] Wherein, T is the medium transmission map, d is the depth map corresponding to the depth information, and β is the attenuation coefficient of light in water.

[0032] Furthermore, the preset channel transmission attention module in step B is to connect the channel attention module in series with the attention module guided by the reverse medium transmission map to enhance the feature data and obtain enhanced data. The method for obtaining the enhanced data is as follows:

[0033] B1. Perform global average pooling on the cascade feature data to obtain pooled data;

[0034] B2. Inputting the pooled data into the first convolutional network in the channel attention module to obtain convolutional data;

[0035] B3, sorting the convolution data through the Relu function in the preset channel transmission attention module to obtain initial data;

[0036] B4, inputting the initial data into the second convolutional network in the preset channel transmission attention module to obtain intermediate data;

[0037] B5, mapping the intermediate data through the S-type function in the preset channel transmission attention module to obtain weight data of each channel;

[0038] B6, performing pixel-level multiplication of the channel weight data and the cascade feature data to obtain channel weighted feature data;

[0039] B7, adding the channel weighted feature data and the cascade feature data at pixel level to obtain sum feature data;

[0040] B8. Use the reverse medium transmission map as the attention map, assign higher weight values ​​to more severely degraded areas, perform pixel-level multiplication on the sum feature data and the reverse medium transmission map to obtain transmission map weighted data, and perform pixel-level addition on the transmission map weighted data and the sum feature data to obtain enhanced data.

[0041] Furthermore, the steps of obtaining the fusion features are as follows:

[0042] C1. Input the target feature data into the first Transformer submodule in the preset cross-color space Transformer module, obtain a first self-attention weight map using the LAB color space in the target feature data as the main color space and combining the RGB color space, and obtain a first initial fusion feature corresponding to the first self-attention weight map through the GELU activation function in the first Transformer submodule; the specific method is as follows:

[0043] The Transformer submodule in the preset cross-color space Transformer module inputs the main color space feature data and the auxiliary color space feature data into the main color space channel and the auxiliary color space channel respectively, adds position embedding in the main color space, records the relative position relationship of the data blocks, obtains the main color space position embedding feature data, performs LayerNorm layer normalization on the main color space position embedding feature data and the auxiliary color space feature data, obtains the normalized main color space feature data and the normalized auxiliary color space feature data, and normalizes the normalized main color space feature data and the normalized auxiliary color space feature data. The feature data are respectively input into the series-connected 1*1 convolution layer and 3*3 depthwise convolution layer to obtain the convolution main color space feature data and the convolution auxiliary color space feature data, and the convolution main color space feature data and the convolution auxiliary color space feature data are converted into the main color space sequence and the auxiliary color space sequence through the convolution layer, and the convolution auxiliary color space sequence is array-split to obtain the first sub-vector and the second sub-vector, and the first sub-vector and the second sub-vector are matrix-multiplied to obtain the self-attention product map, and the self-attention product map is softmax-operated to obtain the self-attention weight map, and the self-attention weight map is combined with The main color space sequence is matrix multiplied to obtain a self-attention weighted feature sequence, which is converted into self-attention weighted feature data through a convolution layer, and the self-attention weighted feature data is input into a 1*1 convolution layer to obtain self-attention weighted convolution data; the main color space position embedding feature data and the self-attention weighted convolution data are added at the pixel level to obtain self-attention weighted data, the self-attention weighted data is normalized by the LayerNorm layer and then array split to obtain the first self-attention weighted data and the second self-attention weighted data, and the first self-attention weighted data and the second self-attention weighted data are added. The weighted data are respectively input into the series-connected 1*1 convolution layer and 3*3 depthwise convolution layer to obtain the first self-attention weighted convolution data and the second self-attention weighted convolution data. The second self-attention weighted convolution data is subjected to GELU operation to obtain the robust second self-attention weighted convolution data. The first self-attention weighted convolution data and the robust second self-attention weighted convolution data are multiplied at the pixel level to obtain the multiplied fusion feature data, which is input into the 1*1 convolution layer to adjust the number of channels of the multiplied fusion feature data. The adjusted multiplied fusion feature data is added to the self-attention weighted data at the pixel level to obtain the initial fusion feature.

[0044] C2, inputting the target feature data into the second Transformer submodule in the preset cross-color space Transformer module, obtaining a second self-attention weight map using the LAB color space in the target feature data as the main color space and combining the HSV color space, and obtaining a second initial fusion feature corresponding to the second self-attention weight map through the GELU activation function in the second Transformer submodule;

[0045] C3, inputting the first initial fusion feature and the second initial fusion feature into the selective kernel feature fusion module in the preset cross-color space Transformer module to obtain the fusion feature. The specific method is as follows:

[0046] The selective kernel feature fusion module first performs pixel-level addition and fusion on the first initial fusion feature and the second initial fusion feature to obtain the added fusion feature, and then uses the global average pooling operation to encode the global information on the added fusion feature to generate channel coding statistical information. Next, a compact channel coding feature is obtained through a fully connected layer for selection and dimensionality reduction. The compact channel coding feature applies a channel-level softmax operation to obtain feature information of different sizes. The feature information of different sizes is respectively multiplied and fused with the first initial fusion feature and the second initial fusion feature at the pixel level, and then added at the pixel level to obtain the fusion feature.

[0047] Furthermore, the steps of generating the final enhanced image are as follows:

[0048] D1, inputting the fused features into the fourth residual enhancement module corresponding to the fourth layer network to obtain enhanced features;

[0049] The fourth residual enhancement module has the same network structure as the preset feature extraction module, specifically:

[0050] Each layer of residual enhancement module is composed of eight layers of convolutional layers with kernel_size=3, stride=1, and padding=1. After the data is input into the first layer of convolution, the Relu function is added to sort the data of this layer to obtain the first layer of convolution data. After the first layer of convolution data is input into the second layer of convolution, the Relu function is added to sort the data of this layer to obtain the second layer of convolution data. After the second layer of convolution data is input into the third layer of convolution, the Relu function is added to sort the data of this layer to obtain the third layer of convolution data. After the third layer of convolution data is input into the fourth layer of convolution, the fourth layer of convolution data is obtained. The first layer of convolution data is combined with The fourth layer of convolution data is added at the pixel level to obtain the first layer of residual features. The first layer of residual features is input into the fifth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the fifth layer of convolution data. The fifth layer of convolution data is input into the sixth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the sixth layer of convolution data. The sixth layer of convolution data is input into the seventh layer of convolution and then the Relu function is added to sort the data of this layer to obtain the seventh layer of convolution data. The seventh layer of convolution data is input into the eighth layer of convolution to obtain the eighth layer of convolution data. The fifth layer of convolution data and the eighth layer of convolution data are added at the pixel level to obtain the target data of the residual enhancement module.

[0051] D2. Extract the enhanced features by a preset upsampling method to obtain first sampled data, and input the first sampled data and the third cascade enhanced data into a third residual enhancement module corresponding to the third layer network to obtain secondary enhanced features. The specific method is as follows:

[0052] The preset upsampling method is bilinear interpolation implemented by the interpolat function, and the formula is as follows:

[0053] F Upsampling =BilinearInterpolat(F Input ) (3)

[0054] Among them, F Upsampling is the data obtained after upsampling, F Input is the input data, BilinearInterpolat(*) is bilinear interpolation, where scale_factor=2.

[0055] D3. Extract the secondary enhanced features by a preset upsampling method to obtain second sampled data, and input the second sampled data and the second cascade enhanced data into a second residual enhancement module corresponding to the second layer network to obtain primary enhanced features;

[0056] D4. Extracting the secondary enhanced features by a preset upsampling method to obtain third sampled data, and inputting the third sampled data and the first cascade enhanced data into a first residual enhancement module corresponding to the first layer network to obtain a target enhanced feature;

[0057] D5. Generate a final enhanced image of the image to be enhanced based on the target enhancement feature.

[0058] Furthermore, the method for generating the final enhanced image of the image to be enhanced based on the target enhancement feature is as follows:

[0059] The target enhancement feature requires that the enhanced image and the image to be enhanced are consistent in pixel information and content information, and retains rich edge detail information. The Charbonnier penalty loss function L is used. Char , content-aware loss function L per , gradient loss function L G A joint loss function is constructed to jointly constrain the generation of a final enhanced image of the image to be enhanced.

[0060] Joint loss function L total Described as:

[0061] L total =λ1L Char +λ2L per +λ3L G (4)

[0062] Where: λ1, λ2, λ3 are the coefficients of Charbonnier penalty loss function, content-aware loss function, and gradient loss function, respectively, which are set according to experience.

[0063]

[0064] Where: I g and I label They represent the generated water map and reference map respectively; ε Represents the regularization term.

[0065] Content-aware loss function L per It is used to ensure the consistency of the content information of the final enhanced image and the image to be enhanced. It is calculated based on the pre-trained model of the VGG-19 network φ on the ImageNet dataset. The VGG-19 network is used as the feature extraction network. Let j represent the jth convolutional layer, φ j (*) represents the feature map in the jth convolutional layer, and the content-aware loss function L per Described as the gap between the generated water map and the target map, the formula is as follows:

[0066]

[0067] Gradient loss function L G It is used to preserve edge sharpness, generate a gradient map through the sobel operator to extract edge features, use the gradient loss as a second-order constraint to refine the edge of the water map, and force the deep network to generate a water map with clearer details. The formula is as follows:

[0068] L G =||Grad(I g )-Grad(I label )|| (7)

[0069] Where: Grad(*) represents the gradient map generated by the Sobel operator.

[0070] Furthermore, the Charbonnier penalty loss function coefficient λ1, the content-aware loss function coefficient λ2, and the gradient loss function coefficient λ3 are set to 1, 0.2, and 3, respectively.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] The present invention transforms the image to be enhanced into color space to obtain RGB data, LAB data and HSV data, extracts features from the RGB data, LAB data and HSV data, obtains target feature data and cascade feature data; inputs the cascade feature data into a preset channel transmission attention module to obtain enhanced data; inputs the target feature data into a preset cross-color space Transformer module to obtain fusion features; and generates a final enhanced image of the image to be enhanced according to the enhanced data and the fusion features. By converting the image to be enhanced into RGB data, LAB data and HSV data, and extracting features from the RGB data, the LAB data and the HSV data, more features are extracted during the feature extraction process, which is beneficial to improving the effect of image enhancement, that is, obtaining a clearer final enhanced image, focusing on channels with strong features through the channel attention module of the channel transmission attention module, and using the reverse medium transmission map as the attention map to extract higher quality degradation, so that the image enhancement is more accurate and targeted, and the self-attention mechanism of the cross-color space Transformer module is used to enhance global consistency, thereby solving the problems of underwater image ghosting and artifacts. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a flow chart of the method of the present invention.

[0074] Figure 2 It is the overall framework of the method of the present invention.

[0075] Figure 3 The method of the present invention presets a channel transmission attention module.

[0076] Figure 4 The method of the present invention presets a cross-color space Transformer module. DETAILED DESCRIPTION

[0077] The present invention will be further described below in conjunction with the accompanying drawings. Figure 1 As shown, an underwater image enhancement method comprises the following steps:

[0078] A. Obtain target feature data and cascade feature data

[0079] Performing color space transformation on the image to be enhanced to obtain RGB data, LAB data and HSV data, and performing feature extraction on the RGB data, LAB data and HSV data to obtain target feature data and cascade feature data;

[0080] C. Get enhanced data

[0081] Inputting the cascade feature data into a preset channel transmission attention module to obtain enhanced data;

[0082] C. Get fusion features

[0083] Inputting the target feature data into a preset cross-color space Transformer module to obtain a fusion feature;

[0084] D. Generate the final enhanced image

[0085] A final enhanced image of the image to be enhanced is generated according to the enhanced data and the fusion features.

[0086] Figure 2 The overall architecture of the present invention is shown as a schematic diagram, and the overall architecture consists of four parts: encoder, skip connection, neck and decoder. The encoder is used to obtain target feature data and cascade feature data. The specific steps are detailed in steps A1-A4 in the invention content.

[0087] Figure 3 The figure shows adding a preset channel transmission attention module at the jump connection to obtain the final enhanced data. For specific steps, please refer to steps B1-B8 in the invention content.

[0088] Figure 4 The figure shows that the neck uses a cross-color space Transformer module to obtain fusion features. For specific details, please refer to steps C1-C3 in the invention content.

[0089] The decoder is used to generate the final enhanced image. For specific steps, please refer to steps D1-D5 in the content of the invention.

[0090] The present invention is not limited to this embodiment, and any equivalent concepts or changes within the technical scope disclosed by the present invention are included in the protection scope of the present invention.

Claims

1. An underwater image enhancement method, characterized in that: The following steps are involved: A. Obtain target feature data and cascade feature data Performing color space transformation on the image to be enhanced to obtain RGB data, LAB data and HSV data, and performing feature extraction on the RGB data, LAB data and HSV data to obtain target feature data and cascade feature data; B. Get enhanced data Inputting the cascade feature data into a preset channel transmission attention module to obtain enhanced data; C. Get fusion features Inputting the target feature data into a preset cross-color space Transformer module to obtain a fusion feature; D. Generate the final enhanced image A final enhanced image of the image to be enhanced is generated according to the enhanced data and the fusion features.

2. The underwater image enhancement method according to claim 1, characterized in that: The steps of obtaining the target feature data and the cascade feature data are as follows: A1. Input the RGB data, LAB data and HSV data into the first layer network of a preset feature extraction module to obtain a first feature for the RGB data, a second feature for the LAB data and a third feature for the HSV data; Performing cascade processing on the first feature, the second feature and the third feature to obtain first cascade feature data of the cascade feature data; A2. Obtain secondary RGB data, secondary LAB data, and secondary HSV data according to the first feature, the second feature, and the third feature by a preset downsampling method, and input the secondary RGB data, the secondary LAB data, and the secondary HSV data into a second layer network in a preset feature extraction module to obtain a fourth feature for the secondary RGB data, a fifth feature for the secondary LAB data, and a sixth feature for the secondary HSV data; Performing cascade processing on the fourth feature, the fifth feature and the sixth feature to obtain second cascade feature data of the cascade feature data; A3. Obtain primary RGB data, primary LAB data and primary HSV data according to the fourth feature, the fifth feature and the sixth feature by the preset downsampling method, and input the primary RGB data, the primary LAB data and the primary HSV data into the third layer network in the preset feature extraction module to obtain a seventh feature for the primary RGB data, an eighth feature for the primary LAB data and a ninth feature for the primary HSV data; Performing cascade processing on the seventh feature, the eighth feature and the ninth feature to obtain third cascade feature data of the cascade feature data; A4. Through the preset downsampling method, target RGB data, target LAB data and target HSV data are obtained according to the seventh feature, the eighth feature and the ninth feature, and the target RGB data, target LAB data and target HSV data are input into the fourth layer network in the preset feature extraction module to obtain the target feature data.

3. The underwater image enhancement method according to claim 2, characterized in that: The preset feature extraction module is composed of eight convolution layers with kernel_size=3, stride=1, and padding=1 to form a residual module. After the data is input into the first convolution layer, the Relu function is added to sort the data of the current layer to obtain the first convolution data. After the first convolution data is input into the second convolution layer, the Relu function is added to sort the data of the current layer to obtain the second convolution data. After the second convolution data is input into the third convolution layer, the Relu function is added to sort the data of the current layer to obtain the third convolution data. After the third convolution data is input into the fourth convolution layer, the fourth convolution data is obtained. The first convolution data and the fourth convolution data are combined. The four-layer convolution data are added at the pixel level to obtain the first-layer residual features. The first-layer residual features are input into the fifth-layer convolution and then the Relu function is added to sort the data in this layer to obtain the fifth-layer convolution data. The fifth-layer convolution data is input into the sixth-layer convolution and then the Relu function is added to sort the data in this layer to obtain the sixth-layer convolution data. The sixth-layer convolution data is input into the seventh-layer convolution and then the Relu function is added to sort the data in this layer to obtain the seventh-layer convolution data. The seventh-layer convolution data is input into the eighth-layer convolution to obtain the eighth-layer convolution data. The fifth-layer convolution data and the eighth-layer convolution data are added at the pixel level to obtain the target data of the feature extraction module.

4. The underwater image enhancement method according to claim 2, characterized in that: The cascade processing method is as follows: using the Cat function to perform a connection operation on the input tensor sequence in the first dimension; The preset downsampling method is maximum pooling with kernel_size=2, stride=2, padding=0, which is mathematically represented as: F Downsampling =MaxPool(F input ) (1) Among them, F Downsampling is the data obtained after downsampling, F Input is the input data, MaxPool(*) is the maximum pooling.

5. The underwater image enhancement method according to claim 1, characterized in that: Before obtaining the enhanced data in step B, a reverse medium transmission diagram is obtained first, and the steps are as follows: Depth information corresponding to the image to be enhanced is obtained, and a reverse medium transmission map corresponding to the image to be enhanced is generated through an underwater imaging model according to the depth information, wherein the mathematical representation of the underwater imaging model is: T=e -dβ (2) Wherein, T is the medium transmission map, d is the depth map corresponding to the depth information, and β is the attenuation coefficient of light in water.

6. The underwater image enhancement method according to claim 1, characterized in that: The preset channel transmission attention module in step B is to connect the channel attention module and the attention module guided by the reverse medium transmission map in series to enhance the feature data and obtain enhanced data. The method for obtaining the enhanced data is as follows: B1. Perform global average pooling on the cascade feature data to obtain pooled data; B2. Inputting the pooled data into the first convolutional network in the channel attention module to obtain convolutional data; B3, sorting the convolution data through the Relu function in the preset channel transmission attention module to obtain initial data; B4, inputting the initial data into the second convolutional network in the preset channel transmission attention module to obtain intermediate data; B5, mapping the intermediate data through the S-type function in the preset channel transmission attention module to obtain weight data of each channel; B6, performing pixel-level multiplication of the channel weight data and the cascade feature data to obtain channel weighted feature data; B7, adding the channel weighted feature data and the cascade feature data at pixel level to obtain sum feature data; B8. Use the reverse medium transmission map as the attention map, assign higher weight values ​​to more severely degraded areas, perform pixel-level multiplication on the sum feature data and the reverse medium transmission map to obtain transmission map weighted data, and perform pixel-level addition on the transmission map weighted data and the sum feature data to obtain enhanced data.

7. The underwater image enhancement method according to claim 1, characterized in that: The steps of obtaining the fusion features are as follows: C1. Input the target feature data into the first Transformer submodule in the preset cross-color space Transformer module, obtain a first self-attention weight map using the LAB color space in the target feature data as the main color space and combining the RGB color space, and obtain a first initial fusion feature corresponding to the first self-attention weight map through the GELU activation function in the first Transformer submodule; The specific method is as follows: The Transformer submodule in the preset cross-color space Transformer module inputs the main color space feature data and the auxiliary color space feature data into the main color space channel and the auxiliary color space channel respectively, adds position embedding in the main color space, records the relative position relationship of the data blocks, obtains the main color space position embedding feature data, performs LayerNorm layer normalization on the main color space position embedding feature data and the auxiliary color space feature data, obtains the normalized main color space feature data and the normalized auxiliary color space feature data, and normalizes the normalized main color space feature data and the normalized auxiliary color space feature data. The feature data are respectively input into the series-connected 1*1 convolution layer and 3*3 depthwise convolution layer to obtain the convolution main color space feature data and the convolution auxiliary color space feature data, and the convolution main color space feature data and the convolution auxiliary color space feature data are converted into the main color space sequence and the auxiliary color space sequence through the convolution layer, and the convolution auxiliary color space sequence is array-split to obtain the first sub-vector and the second sub-vector, and the first sub-vector and the second sub-vector are matrix-multiplied to obtain the self-attention product map, and the self-attention product map is softmax-operated to obtain the self-attention weight map, and the self-attention weight map is combined with The main color space sequence is matrix multiplied to obtain a self-attention weighted feature sequence, which is converted into self-attention weighted feature data through a convolution layer, and the self-attention weighted feature data is input into a 1*1 convolution layer to obtain self-attention weighted convolution data; the main color space position embedding feature data and the self-attention weighted convolution data are added at the pixel level to obtain self-attention weighted data, the self-attention weighted data is normalized by the LayerNorm layer and then array split to obtain the first self-attention weighted data and the second self-attention weighted data, and the first self-attention weighted data and the second self-attention weighted data are added. The weighted data are respectively input into the series-connected 1*1 convolution layer and 3*3 depthwise convolution layer to obtain the first self-attention weighted convolution data and the second self-attention weighted convolution data. The second self-attention weighted convolution data is subjected to GELU operation to obtain the robust second self-attention weighted convolution data. The first self-attention weighted convolution data and the robust second self-attention weighted convolution data are pixel-wise multiplied to obtain the multiplication fusion feature data, which is input into the 1*1 convolution layer to adjust the number of channels of the multiplication fusion feature data. The adjusted multiplication fusion feature data is pixel-wise added to the self-attention weighted data to obtain the initial fusion feature. C2, inputting the target feature data into the second Transformer submodule in the preset cross-color space Transformer module, obtaining a second self-attention weight map using the LAB color space in the target feature data as the main color space and combining the HSV color space, and obtaining a second initial fusion feature corresponding to the second self-attention weight map through the GELU activation function in the second Transformer submodule; C3. Inputting the first initial fusion feature and the second initial fusion feature into the selective kernel feature fusion module in the preset cross-color space Transformer module to obtain the fusion feature; The specific method is as follows: The selective kernel feature fusion module first performs pixel-level addition and fusion on the first initial fusion feature and the second initial fusion feature to obtain the added fusion feature, and then uses the global average pooling operation to encode the global information on the added fusion feature to generate channel coding statistical information. Next, a compact channel coding feature is obtained through a fully connected layer for selection and dimensionality reduction. The compact channel coding feature applies a channel-level softmax operation to obtain feature information of different sizes. The feature information of different sizes is respectively multiplied and fused with the first initial fusion feature and the second initial fusion feature at the pixel level, and then added at the pixel level to obtain the fusion feature.

8. The underwater image enhancement method according to claim 1, characterized in that: The steps of generating the final enhanced image are as follows: D1, inputting the fused features into the fourth residual enhancement module corresponding to the fourth layer network to obtain enhanced features; The fourth residual enhancement module has the same network structure as the preset feature extraction module, specifically: Each layer of residual enhancement module is composed of eight layers of convolutional layers with kernel_size=3, stride=1, and padding=1. After the data is input into the first layer of convolution, the Relu function is added to sort the data of this layer to obtain the first layer of convolution data. After the first layer of convolution data is input into the second layer of convolution, the Relu function is added to sort the data of this layer to obtain the second layer of convolution data. After the second layer of convolution data is input into the third layer of convolution, the Relu function is added to sort the data of this layer to obtain the third layer of convolution data. After the third layer of convolution data is input into the fourth layer of convolution, the fourth layer of convolution data is obtained. The first layer of convolution data is combined with The fourth layer of convolution data is added at the pixel level to obtain the first layer of residual features. The first layer of residual features is input into the fifth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the fifth layer of convolution data. The fifth layer of convolution data is input into the sixth layer of convolution and then the Relu function is added to sort the data of this layer to obtain the sixth layer of convolution data. The sixth layer of convolution data is input into the seventh layer of convolution and then the Relu function is added to sort the data of this layer to obtain the seventh layer of convolution data. The seventh layer of convolution data is input into the eighth layer of convolution to obtain the eighth layer of convolution data. The fifth layer of convolution data and the eighth layer of convolution data are added at the pixel level to obtain the target data of the residual enhancement module. D2. Extract the enhanced features by a preset upsampling method to obtain first sampled data, and input the first sampled data and the third cascade enhanced data into a third residual enhancement module corresponding to the third layer network to obtain secondary enhanced features. The specific method is as follows: The preset upsampling method is bilinear interpolation implemented by the interpolat function, and the formula is as follows: F Upsampling =BilinearInterpolat(F Input ) (3) Among them, F Upsampling is the data obtained after upsampling, F Input is the input data, BilinearInterpolat(*) is bilinear interpolation, where scale_factor = 2; D3. Extract the secondary enhanced features by a preset upsampling method to obtain second sampled data, and input the second sampled data and the second cascade enhanced data into a second residual enhancement module corresponding to the second layer network to obtain primary enhanced features; D4. Extracting the secondary enhanced features by a preset upsampling method to obtain third sampled data, and inputting the third sampled data and the first cascade enhanced data into a first residual enhancement module corresponding to the first layer network to obtain a target enhanced feature; D5. Generate a final enhanced image of the image to be enhanced based on the target enhancement feature.

9. The underwater image enhancement method according to claim 8, characterized in that: The method for generating the final enhanced image of the image to be enhanced based on the target enhancement feature is as follows: The target enhancement feature requires that the enhanced image and the image to be enhanced are consistent in pixel information and content information, and retains rich edge detail information. The Charbonnier penalty loss function L is used. Char , content-aware loss function L per , gradient loss function L G A joint loss function is formed to jointly constrain the generation of a final enhanced image of the image to be enhanced; Joint loss function L total Described as: L total =λ1L Char +λ2L per +λ3L G (4) Where: λ1, λ2, λ3 are the coefficients of Charbonnier penalty loss function, content-aware loss function, and gradient loss function respectively set according to experience; Where: I g and I label They represent the generated water map and the reference map respectively; ε represents the regularization term; Content-aware loss function L per It is used to ensure the consistency of the content information of the final enhanced image and the image to be enhanced. It is calculated based on the pre-trained model of the VGG-19 network φ on the ImageNet dataset. The VGG-19 network is used as the feature extraction network. Let j represent the jth convolutional layer, φ j (*) represents the feature map in the jth convolutional layer, and the content-aware loss function L per Described as the gap between the generated water map and the target map, the formula is as follows: Gradient loss function L G It is used to preserve edge sharpness, generate a gradient map through the sobel operator to extract edge features, use the gradient loss as a second-order constraint to refine the edge of the water map, and force the deep network to generate a water map with clearer details. The formula is as follows: L G =||Grad(I g )-Grad(I label )|| (7) Where: Grad(*) represents the gradient map generated by the Sobel operator.

10. The underwater image enhancement method according to claim 9, characterized in that: The Charbonnier penalty loss function coefficient λ1, the content-aware loss function coefficient λ2, and the gradient loss function coefficient λ3 are set to 1, 0.2, and 3, respectively.

Citation Information

Cited By

  • Dense multiplexing and jump connection underwater image enhancement method for multilayer color features

    CN120495149A

  • Underwater Image Enhancement Method Based on Dense Reuse and Skip Connections for Multi-Layer Color Features

    CN120495149B

  • Underwater image enhancement method based on self-attention guidance and associated feature compensation

    CN122391007A

  • Underwater image enhancement method based on self-attention guidance and correlation feature compensation

    CN122391007B