Multi-scale Unpaired Underwater Image Enhancement Method

Through the multi-scale non-paired underwater image enhancement method, the underwater image is processed using the encoder decoder structure and attention enhancer network, which solves the problems of information loss and color unevenness in the existing methods, and achieves the generation of high-quality images.

CN115880176BActive Publication Date: 2025-07-01FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211600609.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-01
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

The existing underwater image enhancement methods have problems of information loss and color unevenness when processing non-paired images, making it difficult to generate high-quality underwater images.

Method used

Using a multi-scale non-paired underwater image enhancement method, a network of encoder decoder structure is designed, a network of image features is processed using detail-keeping modules and attention enhancer networks, and a circular generation adversarial network structure is built to improve image quality.

Benefits of technology

Effectively learn and restore information from underwater images, enhance image details, improve color contrast, reduce image blur, and generate high-quality images that conform to human visual perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880176B_ABST
    Figure CN115880176B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-scale unpaired underwater image enhancement method, which includes the following steps: Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained; Step S2: Design a multi-scale underwater image enhancement network; Step S3: Build a cyclic generative adversarial network structure and combine it with the multi-scale underwater image quality enhancement network to obtain a multi-scale unpaired underwater image enhancement network; Step S4: Design an objective loss function for training the unpaired underwater image enhancement network; Step S5: Use unpaired images to train the multi-scale unpaired underwater image enhancement network to converge to the Nash equilibrium; Step S6: Normalize the underwater image to be enhanced, then input it into the trained underwater image enhancement model, and output the enhanced image. The present invention can enhance underwater images, use unpaired underwater images for model training, and solve the problem of underwater image distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of image processing and computer vision, and particularly relates to a multi-scale unpaired underwater image enhancement method. Background Art

[0002] Underwater images are often used in underwater operations such as underwater biological discovery, seabed terrain observation, and underwater archaeology. Often, these underwater operations have high requirements for the quality of underwater images. However, due to the complex underwater environment, the quality of underwater images is low. Since the types of underwater image distortion are complex, underwater image enhancement has become an extremely challenging problem. Scattering and energy attenuation of light during underwater propagation are the main reasons for the reduction of underwater image quality. When propagating underwater, light waves have different attenuation coefficients. Generally, the attenuation of red light is faster than that of blue and green light, so underwater images mostly appear blue-green. In addition, underwater images are also affected by forward scattering and backward scattering of light. Among them, forward scattering refers to the scattering phenomenon in which the light reflected by an object in water undergoes a small-angle shift when transmitted to the camera. Forward scattering usually causes image details to be blurred. Backward scattering refers to the situation where when irradiating an object in water, impurities in the water are scattered and directly received by the camera. Backward scattering usually results in low image contrast. Moreover, many plankton, particulate matter, etc. in water introduce noise to underwater images. These adverse effects reduce the visibility, color contrast of underwater images, and even introduce color deviation to the images, resulting in a serious decline in the quality of underwater images. The above-mentioned degradation problems make underwater image enhancement a challenging task.

[0003] Existing underwater image enhancement methods are mainly divided into two categories: one is the method based on physical models. This type of method requires prior mathematical modeling of the underwater image degradation process, and a high-quality underwater image is obtained by inverting the underwater image degradation process. This type of method requires accurate estimation of model parameters. However, the underwater environment is complex and variable, resulting in difficult parameter estimation and low parameter accuracy, leading to low quality of the enhanced image. At the same time, different degradation factors are considered in different underwater environments, and the required mathematical models are also different, making these methods based on physical models have great limitations. The other is the method based on deep learning. This type of method regards the transformation from an underwater image to a high-quality image as a mapping relationship, and uses a deep learning neural network to learn this mapping change, so as to realize the migration from an underwater image to a high-quality image. However, this type of method requires a large number of paired underwater images; and many networks will cause the loss of details of underwater images during the learning process, resulting in the inability to generate high-quality underwater images.

[0004] Most of the existing underwater image quality enhancement methods based on deep learning require a large number of paired images. However, in reality, it is difficult to obtain paired underwater image datasets. Moreover, during the image enhancement process, due to insufficient network learning, information loss usually occurs, resulting in phenomena such as uneven colors, low color contrast, and blurred details in the enhanced images. Summary of the Invention

[0005] Aiming at the defects and deficiencies of the existing technology, the purpose of the present invention is to provide a multi-scale unpaired underwater image enhancement method. This method fully learns image features through a multi-scale network, which is beneficial to improving the quality of underwater images. The solution includes the following steps: Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained; Step S2: Design a multi-scale underwater image enhancement network. The network is designed as an encoder-decoder structure. In the encoder part, a detail preservation module is used to preserve details of the features, and an attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder; Step S3: Build a cyclic generative adversarial network structure and combine it with the multi-scale underwater image quality enhancement network to obtain a multi-scale unpaired underwater image enhancement network; Step S4: Design an objective loss function for training the unpaired underwater image enhancement network; Step S5: Use unpaired images to train the multi-scale unpaired underwater image enhancement network to converge to the Nash equilibrium; Step S6: Normalize the underwater image to be enhanced, then input it into the trained underwater image enhancement model, and output the enhanced image. The present invention can enhance underwater images, use unpaired underwater images for model training, and solve the problem of underwater image distortion.

[0006] The technical solution adopted by the present invention to solve its technical problems is:

[0007] A multi-scale unpaired underwater image enhancement method, characterized in that:

[0008] Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained;

[0009] Step S2: Design a multi-scale underwater image enhancement network, adopting an encoder-decoder structure. In the encoder part, a detail preservation module is used to preserve details of the features, and an attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder;

[0010] Step S3: Build a cyclic generative adversarial network structure and combine it with the multi-scale underwater image quality enhancement network to obtain a multi-scale unpaired underwater image enhancement network;

[0011] Step S4: Design an objective loss function for training the unpaired underwater image enhancement network;

[0012] Step S5: Use unpaired images to train the multi-scale unpaired underwater image enhancement network to converge to the Nash equilibrium;

[0013] Step S6: Normalize the underwater image to be enhanced, and then input it into the trained underwater image enhancement model to output the enhanced image.

[0014] Furthermore, step S1 specifically includes the following steps:

[0015] Step S11: Divide all the unpaired images to be trained into underwater images and high-quality images;

[0016] Step S12: Enlarge all the unpaired images to be trained, with the length and width enlarged to 1.12 times that of the original image. Perform data augmentation on the enlarged images through random cutting operations and random flipping operations;

[0017] Step S13: Normalize all the images to be trained. Given the image I(i, j) to be processed, its normalized image is At the pixel position (i, j), calculate the normalized value The formula is as follows:

[0018]

[0019] where (i, j) represents the position of the pixel;

[0020] Step S14: Use the normalized underwater images as the underwater images for the subsequent steps, and use the normalized high-quality images as the high-quality images for the subsequent steps.

[0021] Furthermore, step S2 specifically includes the following steps:

[0022] Step S21: Design a multi-scale underwater image enhancement network: The multi-scale underwater image quality enhancement network consists of an encoder network and a decoder network. Use an attention enhancement sub-network to enhance the output features of multiple different scales of the encoder and then connect them to the decoder, and directly connect the output features of the encoder that do not use the attention enhancement sub-network to the decoder;

[0023] The encoder consists of 7 downsampling layers and 2 detail preservation modules. The detail preservation modules are located after the 3rd and 6th downsampling layers respectively. The output features of the 2nd, 4th, and 6th downsampling layers of the encoder are enhanced using an attention enhancement subnetwork and then connected to the corresponding layers in the decoder. The output features of the 1st, 3rd, 5th, and 7th downsampling layers of the encoder are directly connected to the corresponding layers in the decoder. Each downsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 convolution. The calculation formula is as follows:

[0024] X eo = ReLU(Conv(Norm(X ei )))

[0025] Among them, X eo represents the output feature of the downsampling layer, X ei represents the input feature of the downsampling layer, ReLU() represents the activation function, Conv() represents the convolution, and Norm() represents the normalization layer;

[0026] The decoder consists of 7 upsampling layers. Each upsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 transposed convolution. The calculation formula is as follows:

[0027] X do = ReLU(TranConv(Norm(X di )))

[0028] Among them, X do represents the output feature of the upsampling layer, X di represents the input feature of the upsampling layer, ReLU() represents the activation function, TranConv() represents the transposed convolution, and Norm() represents the normalization layer;

[0029] The calculation formula of the multi-scale underwater image enhancement network is as follows:

[0030] X down1 = Down(X ni )

[0031] X down2 = Down(X down1 )

[0032] X down3 = Down(X down2 )

[0033] X down4 = Down(DRM(X down3 ))

[0034] X down5 = Down(X down4)

[0035] X down6 = Down(X down5 )

[0036] X down7 = Down(DRM(X down6 ))

[0037] X down8 = Down(X down7 )

[0038] X up1 = ADD(Up(X down8 ), X down7 )

[0039] X up2 = ADD(Up(X up1 ), AEN(X down6 ))

[0040] X up3 = ADD(Up(X up2 ), X down5 )

[0041] X up4 = ADD(Up(X up3 ), AEN(X down4 ))

[0042] X up5 = ADD(Up(X up4 ), X down3 )

[0043] X up6 = ADD(Up(X up5 ), AEN(X down2 ))

[0044] X no = ADD(Up(X up6 ), X down1 )

[0045] Among them, X ni represents the input image of the multi-scale underwater image enhancement network, X no represents the output image of the multi-scale underwater image enhancement network, X down1 , X down2 , X down3 , X down4 , X down5 , X down6 , X down7 , X down8 respectively represent the output features of the corresponding layers of the downsampling layers of the 1st to 8th layers, Xup1 , X up2 , X up3 , X up4 , X up5 , X up6 respectively represent the output features of the corresponding layers of the upsampling layers of the first to the sixth layers, Down() represents the downsampling layer, Up() represents the upsampling layer, DRM() represents the detail retention module, AEN() represents the attention enhancement subnetwork, and ADD() represents the matrix addition operation;

[0046] Step S22: Design the detail retention module in the multi-scale underwater image enhancement network: The input feature X of the detail retention module Ri passes through a channel attention module to obtain the feature X SRi , X SRi is added to X Ri to obtain X R1 ; X R1 passes through a normalization layer to obtain the feature X NR1 , X R1 passes through a 1×1 convolution to obtain the feature X C1R1 , X NR1 is added to X C1R1 to obtain X R2 ; X R2 passes through a 3×3 convolution to obtain the feature X CR2 , X R1 passes through a 1×1 convolution to obtain the feature X C2R1 , X CR2 is added to X C2R1 to obtain the output feature X of the detail retention module Ro , and the calculation formula is as follows:

[0047] X R1 = ADD(SE_Layer(X Ri ), X Ri )

[0048] X R2 = ADD(Conv 1×1 (X R1 ), Norm(X R1 ))

[0049] X Ro = ADD(Conv 3×3 (X R2 ), Conv 1×1 (X R1 ))

[0050] Among them, X Ri represents the input feature of the detail retention module, X Ro represents the output feature of the detail retention module, XR1 , X R2 represents the output features of each stage of the detail preservation module, SE_Layer() represents the channel attention module, Conv 3×3 () represents a 3×3 convolution kernel, Conv 1×1 () represents a 1×1 convolution kernel, Norm() represents the normalization layer, ADD() represents the matrix addition operation;

[0051] Step S23: Design the attention enhancement sub-network in the multi-scale underwater image enhancement network: The attention enhancement sub-network consists of an extended learning module connected in series with a 3×3 convolution, a channel attention module, a 3×3 convolution, and an attention fusion module; The extended learning module consists of three parallel branches; The first branch is a 1×1 convolution, the second branch consists of a 1×1 convolution and a 3×3 convolution connected in series, and the third branch consists of a 1×1 convolution and a max pooling layer connected in series; The attention fusion module consists of a spatial attention module residually connected to another branch, which consists of a channel attention module, a 3×3 convolution, and another channel attention module; The formula for the attention enhancement sub-network is as follows:

[0052] X A1 = Conv 1×1 (X A i)

[0053] X A2 = Conv 3×3 (Conv 1×1 (X A i))

[0054] X A1 = Conv 1×1 (X A i)

[0055] X A2 = Conv 3×3 (Conv 1×1 (X A i))

[0056] X A3 = MaxPool(Conv 1×1 (X A i))

[0057] X A4 = Conv 3×3 (SE_Layer(Conv 3×3 (Cat[X A1 , X A2 , X A3 )))

[0058] X A5 = SE_Layer(Conv 3×3 (SE_Layer(X A4 )))

[0059] X Ao = ADD(SA_Layrt(X A5 ), X A5 )

[0060] Among them, X Ai represents the input feature of the attention enhancement sub-network, X Ao represents the output feature of the attention enhancement sub-network, X A1 , X A2 , X A3 , X A4 , X A5 respectively represent the output features of each stage of the attention enhancement sub-network, SE_Layer() represents the channel attention module, SA_Layer() represents the spatial attention module, Conv 3×3 () represents a 3×3 convolution kernel, Conv 1×1 () represents a 1×1 convolution kernel, Cat[] represents the operation of concatenating features along the channel dimension, and ADD() represents the matrix addition operation.

[0061] Furthermore, step S3 specifically includes the following steps:

[0062] Step S31: Build a cyclic generative adversarial network structure, including a generator G WtoC for enhancing the quality of underwater images, a generator G CtoW for converting high-quality images into underwater images, a discriminator D C for judging the high-quality images generated by the generator, and a discriminator D W for discriminating the underwater images generated by the generator; among them, the network of generator GW toC , G CtoW uses the network structure designed in step S2;

[0063] Step S32: Input unpaired underwater images and high-quality images into the cyclic generative adversarial network structure. Generator G WtoC enhances the quality of underwater image I W to generate an enhanced image E C , and discriminator D C compares the style of enhanced image E C with high-quality image I C ; Generator G CtoW converts high-quality image I CPerform style conversion to generate an underwater image E W , discriminator D W Then, use E W to perform style comparison with the underwater image I W and output a binary classification result of 0 or 1.

[0064] Furthermore, step S4 specifically includes the following steps:

[0065] Step S41: Design the target loss function of the network. The loss includes the generator network loss and the discriminator network loss:

[0066] The total target loss function of the generator network is as follows:

[0067] l G = λ1·l GAN + λ2·l cycle + λ3·l identity

[0068] where l GAN is the generator loss, l cycle is the cycle loss, l identity is the style loss, λ1, λ2, and λ3 are the balance coefficients for balancing each loss, and · is the real number dot product operation;

[0069] The specific calculation formula for the generator loss is as follows:

[0070] l GAN = L MSE (1, D C (E C )) + L MSE (1, D W (E W ))

[0071] where L MSE () is the MSE loss; D C () is the discriminator for the enhanced image E C generated by the generator, and D W () is the discriminator for the underwater image E W generated by the generator; E C is the enhanced image generated by the generator G WtoC , and E W is the underwater image generated by the generator G CtoW ;

[0072] The specific calculation formula for the cycle loss is as follows:

[0073] l cycle = L1(R C , I C ) + L1(R W , IW )

[0074] Among them, L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, R W is the generator G CtoW for the enhanced image E C to reconstruct the obtained image, R C is the generator G WtoC for the generated underwater image E w to reconstruct the obtained image;

[0075] The specific calculation formula of the style loss is as follows:

[0076] l identity = L1(F C , I C ) + L1(F W , I W )

[0077] Among them, L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, F W is to input the underwater image I W into the generator G CtoW to generate the style image, F C is to input the high-quality image I C into the generator G WtoC to obtain the style image;

[0078] The discriminator network loss is as follows:

[0079]

[0080] Among them, is the discriminator loss of the discriminator D C , is the discriminator loss of the discriminator D W , · is the real number dot product operation;

[0081] The discriminator loss of the discriminator D C The specific calculation formula is as follows:

[0082]

[0083] Among them, L MSE () is the MSE loss, D C () is the discriminator for the enhanced image E C , E C is the enhanced image, I Cis the input high-quality image, and · is the real number dot product operation;

[0084] Discriminator D W The specific calculation formula of the discriminator loss of is as follows:

[0085]

[0086] where L MSE () is the MSE loss, and D W () is the discriminator of the underwater image E generated by the discriminator; E W is the generated underwater image; I W is the input underwater image, and · is the real number dot product operation. W is the input underwater image, and · is the real number dot product operation.

[0087] Furthermore, step S5 specifically includes the following steps:

[0088] Step S51: Randomly divide the underwater images and high-quality images into multiple batches, where each batch contains N underwater images and N high-quality images, and randomly pair the underwater images and high-quality images in each batch to obtain N pairs of images;

[0089] Step S52: Input each pair of underwater image I W and high-quality image I C in the same batch into the multi-scale unpaired underwater image enhancement network constructed in step S3 to obtain images E C 、E w 、R W 、R C 、F W 、F C ;

[0090] Step S53: According to the total objective loss function of the image enhancement network, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network;

[0091] Step S54: Repeat the image enhancement network training steps of step S51 to step S53 in batches until the value of the objective loss function of the image enhancement network converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained multi-scale unpaired underwater image enhancement model.

[0092] Furthermore, step S6 specifically includes the following steps:

[0093] Step S61: Normalize the underwater image to be enhanced;

[0094] Step S61: Input the image processed in step S61 into the generator G in the trained multi-scale unpaired underwater image enhancement model WtoC , and output the enhanced image.

[0095] Compared with the prior art, the present invention and its preferred solutions can be applied to unpaired underwater images, and this method can more effectively learn the information of underwater images, thereby effectively restoring the image distortion information, processing its color distortion, enhancing image details, removing image blurring, improving the brightness of the image tone, and the enhanced image conforms to human subjective visual perception. For the existing underwater image enhancement methods, the enhanced images are prone to problems such as blurred details, uneven color distribution, and low brightness. The present invention proposes a multi-scale unpaired underwater image enhancement method, which can fully learn the underwater image features, increase the image color contrast, brighten the image and repair the details, and generate high-quality images. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0097] Figure 1 FIG. is a flowchart for implementing the method of the embodiment of the present invention.

[0098] Figure 2 FIG. is a structural diagram of the network model in the embodiment of the present invention.

[0099] Figure 3 FIG. is a structural diagram of the detail preservation module in the embodiment of the present invention.

[0100] Figure 4 FIG. is a structural diagram of the attention enhancement sub-network in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0101] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0102] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0103] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0104] The following further specifically introduces the solution of this embodiment in conjunction with the accompanying drawings:

[0105] The present invention provides a multi-scale unpaired underwater image enhancement method, as Figures 1 - 4 shown, including the following steps:

[0106] Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained;

[0107] Step S2: Design a multi-scale underwater image enhancement network. The network is designed as an encoder-decoder structure. In the encoder part, a detail-preserving module is used to preserve the details of the features, and an attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder;

[0108] Step S3: Build a cyclic generative adversarial network structure and combine it with the multi-scale underwater image quality enhancement network to obtain a multi-scale unpaired underwater image enhancement network;

[0109] Step S4: Design an objective loss function for training the unpaired underwater image enhancement network;

[0110] Step S5: Use unpaired images to train the multi-scale unpaired underwater image enhancement network to converge to the Nash equilibrium;

[0111] Step S6: Normalize the underwater image to be enhanced, then input it into the trained underwater image enhancement model, and output the enhanced image.

[0112] Furthermore, step S1 includes the following steps:

[0113] Step S11: Divide all the unpaired images to be trained into underwater images and high-quality images;

[0114] Step S12: Enlarge all the unpaired images to be trained. The length and width are enlarged to 1.12 times the original image. Perform data augmentation on the enlarged images through random cutting operations and random flipping operations;

[0115] Step S13: Normalize all the images to be trained. Given the image I(i,j) to be processed, its normalized image is At the pixel position (i,j), calculate the normalized value The formula is as follows:

[0116]

[0117] where (i,j) represents the position of the pixel.

[0118] Step S14: Use the normalized underwater image as the underwater image for subsequent steps, and use the normalized high-quality image as the high-quality image for subsequent steps.

[0119] Further, step S2 includes the following steps:

[0120] Step S21: Design a multi-scale underwater image enhancement network. The multi-scale underwater image quality enhancement network consists of an encoder network and a decoder network. An attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder. At the same time, the output features of the encoder that do not use the attention enhancement sub-network are directly connected to the decoder.

[0121] The encoder consists of 7 downsampling layers and 2 detail preservation modules. The detail preservation modules are located after the 3rd and 6th downsampling layers respectively, and the output features of the 2nd, 4th, and 6th downsampling layers of the encoder are enhanced using an attention enhancement sub-network and then connected to the corresponding layers in the decoder. At the same time, the output features of the 1st, 3rd, 5th, and 7th downsampling layers of the encoder are directly connected to the corresponding layers in the decoder. Each downsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 convolution. Its calculation formula is as follows:

[0122] X eo = ReLU(Conv(Norm(X ei )))

[0123] where X eo represents the output feature of the downsampling layer, X ei represents the input feature of the downsampling layer, ReLU() represents the activation function, Conv() represents the convolution, and Norm() represents the normalization layer.

[0124] The decoder consists of 7 upsampling layers. Each upsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 transposed convolution. Its calculation formula is as follows:

[0125] X do = ReLU(TranConv(Norm(X di )))

[0126] where X do represents the output feature of the upsampling layer, X di represents the input feature of the upsampling layer, ReLU() represents the activation function, TranConv() represents the transposed convolution, and Norm() represents the normalization layer.

[0127] The calculation formula of the multi-scale underwater image enhancement network is as follows:

[0128] Xdown1 = Down(X ni )

[0129] X down2 = Down(X down1 )

[0130] X down3 = Down(X down2 )

[0131] X down4 = Down(DRM(X down3 ))

[0132] X down5 = Down(X down4 )

[0133] X down6 = Down(X down5 )

[0134] X down7 = Down(DRM(X down6 ))

[0135] X down8 = Down(X down7 )

[0136] X up1 = ADD(Up(X down8 ), X down7 )

[0137] X up2 = ADD(Up(X up1 ), AEN(X down6 ))

[0138] X up3 = ADD(Up(X up2 ), X down5 )

[0139] X up4 = ADD(Up(X up3 ), AEN(X down4 ))

[0140] X up5 = ADD(Up(X up4 ), X down3 )

[0141] X up6 = ADD(Up(X up5 ), AEN(X down2 ))

[0142] X no= ADD(Up(X up6 ), X down1 ))

[0143] where X ni represents the input image of the multi-scale underwater image enhancement network, X no represents the output image of the multi-scale underwater image enhancement network, X down1 , X down2 , X down3 , X down4 , X down5 , X down6 , X down7 , X down8 respectively represent the output features of the corresponding layers of the downsampling layers from the 1st to the 8th layer, X up1 , X up2 , X up3 , X up4 , X up5 , X up6 respectively represent the output features of the corresponding layers of the upsampling layers from the 1st to the 6th layer, Down() represents the downsampling layer, Up() represents the upsampling layer, DRM() represents the detail retention module, AEN() represents the attention enhancement sub-network, and ADD() represents matrix addition operation.

[0144] Step S22: Design the detail retention module in the multi-scale underwater image enhancement network. The input feature X Ri of the detail retention module passes through a channel attention module to obtain the feature X SRi , and a SRi is added to X Ri to obtain X R1 ; X R1 passes through a normalization layer to obtain the feature X NR1 , X R1 passes through a 1×1 convolution to obtain the feature X C1R1 , X NR1 is added to X C1R1 to obtain X R2 ; X R2 passes through a 3×3 convolution to obtain the feature X CR2 , X R1 passes through a 1×1 convolution to obtain the feature X C2R1 , X CR2 is added to X C2R1 to obtain the output feature X Ro of the detail retention module. The calculation formula is as follows:

[0145] X R1 = ADD(SE_Layer(X Ri ), X Ri )

[0146] X R2 = ADD(Conv 1×1 (X R1 ), Norm(X R1 ))

[0147] X Ro = ADD(Conv 3×3 (X R2 ), Conv 1×1 (X R1 ))

[0148] Among them, X Ri represents the input feature of the detail preservation module, X Ro represents the output feature of the detail preservation module, X R1 , X R2 represent the output features of each stage of the detail preservation module, SE_Layer() represents the channel attention module, Conv 3×3 () represents a 3×3 convolution kernel, Conv 1×1 () represents a 1×1 convolution kernel, Norm() represents the normalization layer, and ADD() represents the matrix addition operation.

[0149] Step S23: Design the attention enhancement sub-network in the multi-scale underwater image enhancement network. The attention enhancement sub-network consists of an extended learning module connected in series with a 3×3 convolution, a channel attention module, a 3×3 convolution, and an attention fusion module. The extended learning module consists of three branches in parallel; the first branch is a 1×1 convolution, the second branch consists of a 1×1 convolution and a 3×3 convolution connected in series, and the third branch consists of a 1×1 convolution and a max pooling layer connected in series. The attention fusion module consists of a spatial attention module connected in residual connection with another branch, and this branch consists of a channel attention module, a 3×3 convolution, and another channel attention module. The formula of the attention enhancement sub-network is as follows:

[0150] X A1 = Conv 1×1 (X A i)

[0151] X A2 = Conv 3×3 (Conv 1×1 (X A i))

[0152] X A3 = MaxPool(Conv 1×1 (X A i))

[0153] XA4 = Conv 3×3 (SE_Layer(Conv 3×3 (Cat[X A1 , X A2 , X A3 )))

[0154] X A5 = SE_Layer(Conv 3×3 (SE_Layer(X A4 )))

[0155] X Ao = ADD(SA_Layer(X A5 ), X A5 ))

[0156] Among them, X Ai represents the input feature of the attention enhancement sub-network, X Ao represents the output feature of the attention enhancement sub-network, X A1 , X A2 , X A3 , X A4 , X A5 respectively represent the output features of each stage of the attention enhancement sub-network, SE_Layer() represents the channel attention module, SA_Layer() represents the spatial attention module, Conv 3×3 () represents a 3×3 convolution kernel, Conv 1×1 () represents a 1×1 convolution kernel, Cat[] represents the feature splicing operation along the channel dimension, and ADD() represents the matrix addition operation.

[0157] Furthermore, step S3 includes the following steps:

[0158] Step S31: Build a cyclic generative adversarial network structure, which includes a generator G WtoC that enhances the quality of underwater images, a generator G CtoW that converts high-quality images into underwater images, a discriminator D C that judges the high-quality images generated by the generator, and a discriminator D W that discriminates the underwater images generated by the generator. Among them, the network of generator G WtoC , G CtoW uses the network structure designed in step S2.

[0159] Step S32: Input unpaired underwater images and high-quality images into the cyclic generative adversarial network structure, and generator G WtoC enhances the quality of underwater image I W to generate an enhanced image E C, discriminator D C Enhanced image E C is compared with high-quality image I C in terms of style. Generator G CtoW transforms the style of high-quality image I C to generate an underwater image E W , discriminator D W then compares E W with underwater image I W in terms of style and outputs a binary classification result of 0 or 1.

[0160] Furthermore, step S4 includes the following steps:

[0161] Step S41: Design the objective loss function of the network, where the loss includes the generator network loss and the discriminator network loss.

[0162] The total objective loss function of the generator network is as follows:

[0163] l G = λ1·l GAN + λ2·l cycle + λ3·l identity

[0164] where l GAN is the generator loss, l cycle is the cycle loss, l identity is the style loss, λ1, λ2, and λ3 are balance coefficients for balancing each loss, and · is the real dot product operation;

[0165] The specific calculation formula for the generator loss is as follows:

[0166] l GAN = L MSE (1, D C (E C )) + L MSE (1, D W (E W ))

[0167] where L MSE () is the MSE loss; D C () is the discriminator for the enhanced image E C generated, and D W () is the discriminator for the underwater image E W generated; E C is the enhanced image generated by generator G WtoC , and E W is the underwater image generated by generator G CtoW ;

[0168] The specific calculation formula for the cycle loss is as follows:

[0169] l cycle = L1(R C , I C ) + L1(R W , I W )

[0170] where L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, R W is the image obtained by reconstructing the enhanced image E CtoW by the generator G C , and R C is the image obtained by reconstructing the generated underwater image E WtoC by the generator G w ;

[0171] The specific calculation formula for the style loss is as follows:

[0172] l identity = L1(F C , I C ) + L1(F W , I W )

[0173] where L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, F W is the style image obtained by inputting the underwater image I W into the generator G CtoW , and F C is the style image obtained by inputting the high-quality image I C into the generator G WtoC ;

[0174] The discriminator network loss is as follows:

[0175]

[0176] where is the discriminator loss of the discriminator D C , is the discriminator loss of the discriminator D W , and · is the real number dot product operation;

[0177] The specific calculation formula for the discriminator loss of the discriminator D C is as follows:

[0178]

[0179] where L MSE() is the MSE loss, D C () is the discriminator for the discriminator-enhanced image E C ; E C is the enhanced image, I C is the input high-quality image, · is the real number dot product operation;

[0180] The discriminator D W The specific calculation formula for the discriminator loss of is as follows:

[0181]

[0182] where L MSE () is the MSE loss, D W () is the discriminator for the underwater image E generated by the discriminator W ; E W is the generated underwater image; I W is the input underwater image, · is the real number dot product operation.

[0183] Furthermore, step S5 includes the following steps:

[0184] Step S51: Randomly divide the underwater images and high-quality images into multiple batches, where each batch contains N underwater images and N high-quality images, and randomly pair the underwater images and high-quality images in each batch to obtain N pairs of images;

[0185] Step S52: Input each pair of underwater image I W and high-quality image I C in the same batch into the multi-scale unpaired underwater image enhancement network described in step S3 to obtain images E C , E w , R W , R C , F W , F C ;

[0186] Step S53: According to the total objective loss function of the image enhancement network, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network;

[0187] Step S54: Repeat the above steps S51 to S53 of the image enhancement network training steps in batches until the value of the objective loss function of the image enhancement network converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained multi-scale unpaired underwater image enhancement model.

[0188] Furthermore, step S6 includes the following steps:

[0189] Step S61: Normalize the underwater image to be enhanced;

[0190] Step S61: Input the image processed in Step S61 into the generator G in the trained multi-scale unpaired underwater image enhancement model WtoC to output an enhanced image.

[0191] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0192] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0193] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0194] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0195] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

[0196] This patent is not limited to the above best implementation mode. Anyone can obtain various other forms of multi-scale unpaired underwater image enhancement methods under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.

Claims

1. A multi-scale unpaired underwater image enhancement method, characterized in that: Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained; Step S2: Design a multi-scale underwater image enhancement network, adopting an encoder-decoder structure. In the encoder part, a detail preservation module is used to preserve the details of the features, and an attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder; Step S3: Build a cyclic generative adversarial network structure and combine it with the multi-scale underwater image quality enhancement network to obtain a multi-scale unpaired underwater image enhancement network; Step S4: Design an objective loss function for training the unpaired underwater image enhancement network; Step S5: Use unpaired images to train the multi-scale unpaired underwater image enhancement network to converge to the Nash equilibrium; Step S6: Normalize the underwater image to be enhanced, then input it into the trained underwater image enhancement model, and output the enhanced image; Step S2 specifically includes the following steps: Step S21: Design a multi-scale underwater image enhancement network: The multi-scale underwater image quality enhancement network consists of an encoder network and a decoder network. An attention enhancement sub-network is used to enhance the output features of multiple different scales of the encoder and then connect them to the decoder, and the output features of the encoder that do not use the attention enhancement sub-network are directly connected to the decoder; The encoder consists of 7 downsampling layers and 2 detail preservation modules, which are respectively located after the 3rd and 6th downsampling layers. The attention enhancement sub-network is used to enhance the output features of the 2nd, 4th, and 6th downsampling layers of the encoder, and then connect them to the corresponding layers in the decoder. The output features of the 1st, 3rd, 5th, and 7th downsampling layers of the encoder are directly connected to the corresponding layers in the decoder; Each downsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 convolution; The calculation formula is as follows: X eo = ReLU(Conv(Norm(X ei ))) Among them, X eo represents the output feature of the downsampling layer, X ei represents the input feature of the downsampling layer, ReLU() represents the activation function, Conv() represents convolution, and Norm() represents the normalization layer; The decoder consists of 7 upsampling layers; Each upsampling layer consists of a concatenation of an activation function, a normalization layer, and a 2×2 transposed convolution; The calculation formula is as follows: X do = ReLU(TranConv(Norm(X di ))) Among them, X do represents the output feature of the upsampling layer, and X di represents the input feature of the upsampling layer. ReLU() represents the activation function, TranConv() represents the transposed convolution, and Norm() represents the normalization layer; The calculation formula of the multi-scale underwater image enhancement network is as follows: X down1 = Down(X ni ) X down2 = Down(X down1 ) X down3 = Down(X down2 ) X down4 = Down(DRM(X down3 )) X down5 = Down(X down4 ) X down6 = Down(X down5 ) X down7 = Down(DRM(X down6 )) X down8 = Down(X down7 ) X up1 = ADD(Up(X down8 ), X down7 ) X up2 = ADD(Up(X up1 ), AEN(X down6 )) X up3 = ADD(Up(X up2 ), X down5 ) X up4 = ADD(Up(X up3 ), AEN(X down4 )) X up5 = ADD(Up(X up4 ), X down3 ) X up6 = ADD(Up(X up5 ), AEN(X down2 )) X no = ADD(Up(X up6 ), X down1 ) Among them, X ni represents the input image of the multi-scale underwater image enhancement network, and X no represents the output image of the multi-scale underwater image enhancement network, and X down1 , X down2 , X down3 , X down4 , X down5 , X down6 , X down7 , X down8 respectively represent the output features of the corresponding layers of the downsampling layers of the 1st to 8th layers. X up1 , X up2 , X up3 , X up4 , X up5 , X up6 respectively represent the output features of the corresponding layers of the upsampling layers of the 1st to 6th layers. Down() represents the downsampling layer, Up() represents the upsampling layer, DRM() represents the detail retention module, AEN() represents the attention enhancement sub-network, and ADD() represents the matrix addition operation; Step S22: Design the detail preservation module in the multi-scale underwater image enhancement network: The input feature X of the detail preservation module Ri passes through a channel attention module to obtain the feature X SRi , a SRi is added to X Ri to obtain X R1 ; X R1 passes through a normalization layer to obtain the feature X NR1 , X R1 passes through a 1×1 convolution to obtain the feature X C1R1 , X NR1 is added to X C1R1 to obtain X R2 ; X R2 passes through a 3×3 convolution to obtain the feature X CR2 , X R1 passes through a 1×1 convolution to obtain the feature X C2R1 , X CR2 is added to X C2R1 to obtain the output feature X of the detail preservation module Ro , and the calculation formula is as follows: X R1 = ADD(SE_Layer(X Ri ), X Ri ) X R2 = ADD(Conv 1×1 (X R1 ), Norm(X R1 )) X Ro = ADD(Conv 3×3 (X R2 ), Conv 1×1 (X R1 )) Among them, X Ri represents the input feature of the detail preservation module, and X Ro represents the output feature of the detail preservation module. X R1 , X R2 represent the output features of each stage of the detail preservation module. SE_Layer() represents the channel attention module, and Conv 3×3 () represents a convolution with a kernel of 3×3, and Conv 1×1 () represents a convolution with a kernel of 1×1. Norm() represents the normalization layer, and ADD() represents the matrix addition operation; Step S23: Design the attention enhancement sub-network in the multi-scale underwater image enhancement network: The attention enhancement sub-network consists of an extended learning module concatenated with a 3×3 convolution, a channel attention module, a 3×3 convolution, and an attention fusion module; The extended learning module consists of three parallel branches; The first branch is a 1×1 convolution, the second branch consists of a concatenation of a 1×1 convolution and a 3×3 convolution, and the third branch consists of a concatenation of a 1×1 convolution and a max pooling layer; The attention fusion module consists of a spatial attention module residually connected to another branch, which consists of a channel attention module, a 3×3 convolution, and another channel attention module; The formula of the attention enhancement sub-network is as follows: X A1 = Conv 1×1 (X Ai ) X A2 = Conv 3×3 (Conv 1×1 (X Ai )) X A3 = MaxPool(Conv 1×1 (X Ai )) X A4 = Conv 3×3 (SE_Layer(Conv 3×3 (Cat[X A1 , X A2 , X A3 ))) X A5 = SE_Layer(Conv 3×3 (SE_Layer(X A4 ))) X Ao = ADD(SA_Layer(X A5 ), X A5 ) Among them, X Ai represents the input feature of the attention enhancement sub-network, X Ao represents the output feature of the attention enhancement sub-network, X A1 , X A2 , X A3 , X A4 , X A5 respectively represent the output features of each stage of the attention enhancement sub-network, SE_Layer() represents the channel attention module, SA_Layer() represents the spatial attention module, Conv 3×3 () represents a convolution with a kernel of 3×3, Conv 1×1 () represents a convolution with a kernel of 1×1, Cat[] represents an operation of feature concatenation along the channel dimension, and ADD() represents a matrix addition operation.

2. The multi-scale unpaired underwater image enhancement method according to claim 1, wherein Step S1 specifically includes the following steps: Step S11: Divide all the unpaired images to be trained into underwater images and high-quality images; Step S12: Enlarge all the unpaired images to be trained, with the length and width enlarged to 1.12 times that of the original images, and perform data augmentation on the enlarged images through random cutting operations and random flipping operations; Step S13: Normalize all the images to be trained. Given an image I(i, j) to be processed, its normalized image is At the pixel position (i, j), calculate the normalization value using the following formula: Among them, (i, j) represents the position of the pixel; Step S14: Use the normalized underwater images as the underwater images for the subsequent steps, and the normalized high-quality images as the high-quality images for the subsequent steps.

3. The multi-scale unpaired underwater image enhancement method according to claim 1, characterized in that Step S3 specifically includes the following steps: Step S31: Build a cyclic generative adversarial network structure, including a generator G for enhancing the quality of underwater images WtoC , a generator G for converting high-quality images into underwater images CtoW , a discriminator D for judging the high-quality images generated by the generator C , a discriminator D for discriminating the underwater images generated by the generator W ; among them, the networks of the generators G WtoC and G CtoW use the network structure designed in Step S2 Step S32: Input the unpaired underwater images and high-quality images into the cycle generative adversarial network structure, where the generator G WtoC enhances the quality of the underwater image I W to generate an enhanced image E C , and the discriminator D C compares the style of the enhanced image E C with the high-quality image I C ; the generator G CtoW transforms the style of the high-quality image I C to generate the underwater image E W , and the discriminator D W then compares the style of E W with the underwater image I W and outputs a binary classification result of 0, 1.

4. The multi-scale unpaired underwater image enhancement method according to claim 3, wherein Step S4 specifically includes the following steps: Step S41: Design the objective loss function of the network, and the loss includes the loss of the generator network and the loss of the discriminator network: The total objective loss function of the generator network is as follows: l G = λ1·l GAN + λ2·l cycle + λ3·l identity where, l GAN is the generator loss, l cycle is the cycle loss, l identity is the style loss, λ1, λ2, and λ3 are the balance coefficients for balancing each loss, and · is the real dot product operation; The specific calculation formula of the generator loss is as follows: l GAN = L MSE (1, D C (E C )) + L MSE (1, D W (E W )) Among them, L MSE () is the MSE loss; D C () is the discriminator for the enhanced image E C generated; D W () is the discriminator for the underwater image E W generated; E C is the enhanced image generated by the generator G WtoC ; E W is the underwater image generated by the generator G CtoW ; The specific calculation formula of the cycle loss is as follows: l cycle = L1(R C , I C ) + L1(R W , I W ) Among them, L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, R W is the generator G CtoW for the enhanced image E C the image obtained by reconstructing, R C is the generator G WtoC for the generated underwater image E w the image obtained by reconstructing; The specific calculation formula of the style loss is as follows: l identity = L1(F C , I C ) + L1(F W , I W ) Among them, L1() is the L1 loss, I W is the input underwater image, I C is the input high-quality image, F W is the underwater image I W input into the generator G CtoW to obtain the style image, F C is the high-quality image I C input into the generator G WtoC to obtain the style image; The loss of the discriminator network is as follows: Among them, is the discriminator loss of discriminator D C , is the discriminator loss of discriminator D W , · is the real number dot product operation; Discriminator D C The specific calculation formula of the discriminator loss of C is as follows: Among them, L MSE () is the MSE loss, D C () is the discriminator for enhancing the discriminative image E C ; E C is the enhanced image, I C is the input high-quality image, · is the real number dot product operation; Discriminator D W The specific calculation formula for the discriminator loss of Among them, L MSE () is the MSE loss, D W () is the discriminator for discriminating the generated underwater image E W ; E W is the generated underwater image; I W is the input underwater image, · is the real number dot product operation.

5. The multi-scale unpaired underwater image enhancement method according to claim 4, wherein, Step S5 specifically includes the following steps: Step S51: Randomly divide the underwater images and high-quality images into multiple batches, where each batch contains N underwater images and N high-quality images, and randomly pair the underwater images and high-quality images in each batch to obtain N pairs of images; Step S52: Input each pair of underwater images I W and high-quality images I C in the same batch into the multi-scale unpaired underwater image enhancement network constructed in Step S3 to obtain images E C 、E w 、R W 、R C 、F W 、F C ; Step S53: According to the total objective loss function of the image enhancement network, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network; Step S54: Repeat the image enhancement network training steps of Step S51 to Step S53 in batches until the value of the objective loss function of the image enhancement network converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained multi-scale unpaired underwater image enhancement model.

6. The multi-scale unpaired underwater image enhancement method according to claim 5, wherein Step S6 specifically includes the following steps: Step S61: Normalize the underwater image to be enhanced; Step S61: Input the image processed in step S61 into the generator G in the trained multi-scale unpaired underwater image enhancement model WtoC , and output the enhanced image.