A non-paired underwater image enhancement method based on wavelet transform

Through the non-paired underwater image enhancement method based on wavelet transform, the problems of information loss and blurred details during non-paired image enhancement in the prior art are solved, and higher quality underwater image enhancement is achieved, which is suitable for complex environments.

CN115731199BActive Publication Date: 2025-07-01FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211486776.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-01
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

The existing underwater image enhancement methods have information loss when processing non-paired images, resulting in local blurred details of the enhanced images and low adaptability to complex environments.

Method used

The non-paired underwater image enhancement method based on wavelet transformation is adopted. The image is divided into low-frequency and high-frequency parts through wavelet transformation, and is processed and corrected separately. The adversarial network is generated by combining attention mechanisms and loops to reduce information loss and improve image quality.

Benefits of technology

It effectively reduces information loss during image transmission, retains image detail information, improves image contrast and brightness, and the enhanced image conforms to human subjective visual perception and is suitable for most complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731199B_ABST
    Figure CN115731199B_ABST
Patent Text Reader

Abstract

The present invention relates to a non-paired underwater image enhancement method based on wavelet transform. It includes: preprocessing the non-paired data to be trained; designing an underwater image enhancement network based on wavelet transform to separately process the low-frequency and high-frequency parts of the input image, and then combining the high-frequency part and the low-frequency part; building a cyclic generative adversarial network framework and combining it with the underwater image quality enhancement network based on wavelet transform to obtain a non-paired underwater image enhancement network based on wavelet transform; designing an objective loss function for training the non-paired underwater image enhancement network; using non-paired images to train the non-paired underwater image enhancement network based on wavelet transform to converge to the Nash equilibrium; normalizing the underwater image to be enhanced, and then inputting it into the trained underwater image enhancement model to output the enhanced image. The present invention can enhance underwater images, use non-paired underwater images for model training, and solve the problem of underwater image distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing and computer vision, and particularly to a non-paired underwater image enhancement method based on wavelet transform. Background Art

[0002] Underwater images have an important impact on underwater environmental applications such as seabed resource exploration, underwater target detection, and underwater salvage, all of which require a high-quality vision system to perceive the surrounding environment. However, the complex underwater environment causes a serious decline in the quality of underwater images. Due to the complexity of the underwater environment and lighting conditions, underwater image enhancement is a challenging problem. The main reasons for the decline in underwater image quality are the energy attenuation and scattering of light when it propagates underwater. When light propagates underwater, light waves have different attenuation coefficients. Generally speaking, the attenuation of red light is faster than that of blue and green light, so underwater images usually appear blue-green. In addition, underwater images are also affected by wavelength-dependent absorption and scattering, including forward scattering and backscattering. Forward scattering refers to the scattering phenomenon in which the light reflected by an object in water undergoes a small-angle shift when transmitted to the camera, resulting in blurred image details. Backscattering refers to the situation where when irradiating an object in water, impurities in the water are scattered and directly received by the camera, resulting in low image contrast. Moreover, in the ocean, many particles such as plankton, sediment, and dust settle and float up and down, introducing a lot of noise. These adverse effects reduce the visibility and contrast of underwater images and even introduce color deviations, resulting in a serious decline in the quality of underwater images. The above degradation problems make underwater image enhancement a challenging task.

[0003] Existing underwater image enhancement methods are mainly divided into two categories: one is the method based on deep learning. This kind of method regards the conversion from an underwater image to a normal image as a mapping relationship, and uses the learning and fitting ability of the network to learn this mapping change, so as to realize the conversion from an underwater image to a normal image. However, this kind of method has relatively high requirements for the data set and requires a large number of paired images. And during the learning process, there are usually encoding and decoding operations on the images, and some image information may be lost during this process, resulting in local detail blurring in the enhanced image. The other is the method based on physical models. This kind of method needs to first establish a mathematical model of the underwater image degradation process, and obtain a clear underwater image through the inversion of the degradation process. This kind of method needs to estimate the model parameters. However, the underwater environment is complex and lighting and other conditions are changeable, making it difficult to estimate the parameters and with low accuracy, resulting in generally low quality of the enhanced images. At the same time, different environments consider different factors and different models need to be established, making these methods have great limitations.

[0004] Existing methods mainly target paired images. However, it is difficult to collect a dataset of paired images in life, and information loss usually occurs during the image enhancement process, resulting in blurred details in the enhanced images. We propose an unpaired underwater image quality enhancement method based on wavelet transform. This method divides image features into low-frequency and high-frequency parts and processes them separately. Summary of the Invention

[0005] The object of the present invention is to provide an unpaired underwater image enhancement method based on wavelet transform, which is beneficial to improving the quality of underwater images.

[0006] To achieve the above object, the technical solution of the present invention is: an unpaired underwater image enhancement method based on wavelet transform, comprising the following steps:

[0007] Step S1, perform data preprocessing, data enhancement, and normalization on the unpaired data to be trained;

[0008] Step S2, design an underwater image enhancement network based on wavelet transform, and separately process the low-frequency and high-frequency parts of the input image through wavelet transform. The low-frequency part uses a low-frequency processing module to correct the low-frequency part of the image, and the high-frequency part uses a high-frequency processing module for correction. Finally, the high-frequency part and the low-frequency part are combined;

[0009] Step S3, build a cyclic generative adversarial network and combine it with the underwater image quality enhancement network based on wavelet transform to obtain an unpaired underwater image enhancement network based on wavelet transform;

[0010] Step S4, design the objective loss function of the unpaired underwater image enhancement network based on wavelet transform;

[0011] Step S5, use unpaired images to train the unpaired underwater image enhancement network based on wavelet transform to converge to the Nash equilibrium;

[0012] Step S6, perform normalization on the underwater image to be enhanced, and then input it into the trained unpaired underwater image enhancement network based on wavelet transform to output the enhanced image.

[0013] In an embodiment of the present invention, the specific implementation of step S1 is as follows:

[0014] Step S11, divide all the unpaired images to be trained into original underwater images and high-quality images;

[0015] Step S12, uniformly enlarge all the unpaired images to be trained to 1.12 times the original image, and perform data enhancement on the enlarged images through random cutting and random flipping operations;

[0016] Step S13: Normalize all the images to be trained. Given an image I(i, j), its normalized image is At the pixel position (i, j), calculate the normalized value The formula is as follows:

[0017]

[0018] where (i, j) represents the position of the pixel. The normalized underwater image is used as the underwater image for the subsequent steps, and the normalized high-quality image is used as the high-quality image for the subsequent steps.

[0019] In an embodiment of the present invention, the specific implementation of step S2 is as follows:

[0020] Step S21: Process the low-frequency and high-frequency parts of the input image through wavelet transform; the input of the underwater image enhancement network based on wavelet transform is the normalized underwater image After the image is input into the underwater image enhancement network based on wavelet transform, the image goes through a convolution with a convolution kernel of 1×1 and a stride of 1, and two convolutions with convolution kernels of 3×3 and a stride of 1 to extract the initial underwater image feature F1 out ; the underwater image feature F1 out is decomposed into a low-frequency feature ll and high-frequency features lh, hl, hh through a wavelet pooling layer;

[0021] Step S22: Design the low-frequency processing module in the underwater image enhancement network based on wavelet transform; the low-frequency feature ll enters the low-frequency processing module of the U-Net network structure combined with the attention mechanism to obtain the processed low-frequency feature ll′; the calculation formula is as follows:

[0022] ll' = LUNet(ll)

[0023] where ll represents the input low-frequency feature, ll′ represents the processed low-frequency feature, and LUNet(*) represents the low-frequency processing module;

[0024] Step S23: Design the high-frequency processing module in the underwater image enhancement network based on wavelet transform; the high-frequency processing module uses three parallel high-frequency attention modules to process the three high-frequency features lh, hl, hh respectively, and the output processed high-frequency features are lh′, hl′, hh′; the calculation formula is as follows:

[0025] lh' = HAM(lh)

[0026] hl′ = HAM(hl)

[0027] hh′ = HAM(hh)

[0028] where HAM(*) represents the high-frequency attention module;

[0029] Step S24: Fuse the low-frequency features ll′ and high-frequency features lh′, hl′, hh′ and output the enhanced image; the low-frequency features and high-frequency features first pass through a wavelet inverse pooling layer, and then through two convolutions with a convolution kernel of 4×4 and a stride of 1 to output the enhanced image E; the wavelet inverse pooling layer sets its convolution as deconvolution on the basis of the wavelet pooling layer; the calculation formula is as follows:

[0030] E = Conv(Conv(WaveUnpool(ll′, lh′, hl′, hh′)))

[0031] where E represents the enhanced image, Conv(*) represents convolution, and WaveUnPool(*) represents the wavelet inverse pooling layer.

[0032] In an embodiment of the present invention, the specific implementation of step S22 is as follows:

[0033] Step S221: Design the low-frequency processing module of the U-Net network structure with a combined attention mechanism in step S22; the low-frequency processing module consists of a U-Net network structure stacked with seven modules, and the skip connection parts of the 2nd, 4th, and 6th layers of the low-frequency processing module are each combined with a low-frequency attention module; the downsampling module in the U-Net network structure consists of an activation function, a normalization layer, and a 2×2 convolution; the calculation formula is as follows:

[0034] X out = ReLU(Conv(Norm(X in )))

[0035] where X in represents the input feature of the downsampling module, X out represents the output feature of the downsampling module, Conv(*) represents convolution, ReLU(*) represents the activation function, and Norm(*) represents the normalization layer;

[0036] The upsampling module in the U-Net network structure consists of an activation function, a normalization layer, and a 2×2 deconvolution;

[0037] The specific calculation formula of the low-frequency processing module is as follows:

[0038] X down1 = Down(X in )

[0039] X down2 = Down(X down1 )

[0040] Xdown3 = Down(X down2 )

[0041] ……X down8 = Down(X down7 )

[0042] X up1 = ADD[Up(X down8 ), X down7 )

[0043] X up2 = ADD[Up(X up1 ), LAM(X down6 )

[0044] X up3 = ADD[Up(X up2 ), X down5 )

[0045] X up4 = ADD[Up(X up3 ), LAM(X down4 )

[0046] X up5 = ADD[Up(X up4 ), X down3 )

[0047] X up6 = ADD[Up(X up5 ), LAM(X down2 )

[0048] X out = ADD[Up(X up6 ), X down1 )

[0049] Among them, X in represents the input feature of the low-frequency processing module, X out represents the output feature of the low-frequency processing module, Down(*) represents the downsampling module, UP(*) represents the upsampling module, LAM(*) represents the low-frequency attention module, and ADD[*] represents the matrix addition operation;

[0050] Step S222: Design the low-frequency attention module in step S221; the low-frequency attention module consists of an expansion network connected in series with a 3×3 convolution and a channel attention module; the expansion network is composed of three layers of convolution blocks with different numbers in parallel and a feature concatenation operation connected in series, and each convolution block is composed of a convolution and an activation function connected in series; the first layer of the expansion network is a 1×1 convolution block, the second layer is a 1×1 convolution block connected in series with a 3×3 convolution block, and the third layer is a 1×1 convolution block connected in series with two 3×3 convolution blocks; the features output by the expansion network pass through a 3×3 convolution and a channel attention module; the specific calculation formula of the low-frequency attention module is as follows:

[0051] X1 = ReLU(Conv(X in ))

[0052] X2 = ReLU(Conv(ReLU(Conv(X in ))))

[0053] X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in ))))))

[0054] X out = SE_Layer(Conv(Cat[X1,X2,X3]))

[0055] where X in represents the input features of the low-frequency attention module, X out represents the output features of the low-frequency attention module, Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation in the channel dimension, and SE_Layer(*) represents the channel attention module.

[0056] In an embodiment of the present invention, in step S23, the high-frequency attention module consists of an expansion network connected in series with a 3×3 convolution, a spatial attention module, and a channel attention module; the expansion network is composed of three layers of convolution blocks with different numbers in parallel and a feature concatenation operation connected in series, and each convolution block is composed of a convolution and an activation function connected in series; the first layer is a 3×3 convolution block, the second layer is two 3×3 convolution blocks, and the third layer is three 3×3 convolution blocks; the features output by the expansion network pass through a 3×3 convolution, a spatial attention module, and a channel attention module; the calculation formula of the high-frequency attention module is as follows:

[0057] X1 = ReLU(Conv(X in ))

[0058] X2 = ReLU(Conv(ReLU(Conv(X in ))))

[0059] X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in ))))))

[0060] X out = SE_Layer(SA_Layer(Conv(Cat[X1,X2,X3])))

[0061] Wherein, X in represents the input feature of the high-frequency attention module, and X out represents the output feature of the high-frequency attention module. Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation in the channel dimension, SA_Layer(*) represents the spatial attention module, and SE_Layer(*) represents the channel attention module.

[0062] In an embodiment of the present invention, the specific implementation of step S3 is as follows:

[0063] Step S31: Build a cyclic generative adversarial network, including a generator G for enhancing the quality of underwater images W2C , a generator G for converting high-quality images into underwater images C2W , a discriminator D for judging the generated high-quality images C , and a discriminator D for discriminating the generated underwater images W ; The network structures of generators G W2C , G C2W use the network structure of the underwater image enhancement network based on wavelet transform designed in step S2.

[0064] Step S32: Input unpaired underwater images and high-quality images into the cyclic generative adversarial network. The generator G W2C enhances the quality of the underwater image G W to generate an enhanced image E C , and the discriminator D C compares its style with the high-quality image; The generator G C2W converts the high-quality image G C into an underwater image E W , and the discriminator D W then compares its style with the underwater image.

[0065] In an embodiment of the present invention, the specific implementation of step S4 is as follows:

[0066] Design the objective loss function of the unpaired underwater image enhancement network based on wavelet transform. The loss includes the loss of the generator network and the loss of the discriminator network;

[0067] The total objective loss function of the generator network is as follows:

[0068] l G = λ1·l GAN + λ2·l cycle + λ3·l identity

[0069] where l GAN , l cycle and l identity are the generator loss, cycle loss, and style loss respectively. λ1, λ2, and λ3 are the coefficients for balancing each loss, and · is the dot product operation of real numbers. The specific calculation formulas for each loss are as follows:

[0070] l GAN = L MSE (1, D C (E C )) + L MSE (1, D W (E W ))

[0071] where D C (*) is the discriminator for the enhanced image E C , and D W (*) is the discriminator for the generated underwater image E W . E C and E W are the enhanced image generated by the generator G W2C and the underwater image generated by the generator G C2W respectively. L MSE (*) is the MSE loss;

[0072] l cycle = L1(R C , G C ) + L1(R W , G W )

[0073] where R W is the image reconstructed by the generator G C2W from the enhanced image E C , and R C is the image reconstructed by the generator G W2C from the generated underwater image E w . G W is the input underwater image, and G C is the input high-quality image. L1(*) is the L1 loss;

[0074] l identity = L1(F C , G C ) + L1(F W , G W )

[0075] wherein, F W is the image obtained by inputting the underwater image G W into the generator G C2W ; F C is the image obtained by inputting the high-quality image G C into the generator G W2C ; G W is the input underwater image, G C is the input high-quality image, and L1(*) is the L1 loss.

[0076] The discriminator network loss is as follows:

[0077]

[0078] wherein, and are respectively the discriminator loss of the discriminator D C and the discriminator loss of the discriminator D W ; · is the dot product operation of real numbers;

[0079]

[0080] wherein, D C (*) is the discriminator for the discriminant enhanced image E C ; E C is the enhanced image, G C is the input high-quality image, L MSE (*) is the MSE loss, and · is the dot product operation of real numbers;

[0081]

[0082] wherein, D W (*) is the discriminator for the discriminant generated underwater image E W ; E W is the generated underwater image; G W is the input underwater image, L MSE (*) is the MSE loss, and · is the dot product operation of real numbers.

[0083] In an embodiment of the present invention, the specific implementation of the step S5 is as follows:

[0084] Step S51: Randomly divide the underwater images and high-quality images into multiple batches, where each batch contains N underwater images and N high-quality images, and randomly pair the underwater images and high-quality images to obtain N pairs of images;

[0085] Step S52: Input each pair of underwater image G W and high-quality image G C in the same batch into the non-paired underwater image enhancement network based on wavelet transform in Step S3 to obtain images E C 、E w 、R W 、R c 、F W 、F c ;

[0086] Step S53: According to the objective loss function of the non-paired underwater image enhancement network based on wavelet transform, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network;

[0087] Step S54: Repeat the above image enhancement network training steps of Step S51 to Step S53 in batches until the value of the objective loss function of the non-paired underwater image enhancement network based on wavelet transform converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained non-paired underwater image enhancement network based on wavelet transform.

[0088] In an embodiment of the present invention, the specific implementation of Step S6 is as follows:

[0089] Step S61: Normalize the underwater image to be enhanced;

[0090] Step S62: Input the image processed in Step S61 into the generator G W2C in the trained non-paired underwater image enhancement network based on wavelet transform, and output the enhanced image.

[0091] Compared with the prior art, the present invention has the following beneficial effects: The present invention can be applied to non-paired underwater images, and this method can more effectively utilize the information of underwater images, thereby effectively restoring the distorted colors of images, removing image blurs, enhancing image contrast and brightness, and the enhanced images conform to human subjective visual perception. In the existing underwater image enhancement methods, the enhanced images often show the phenomenon of blurred details. The present invention proposes a non-paired underwater image enhancement method based on wavelet transform, which can effectively reduce the information loss in the image transmission process, retain the detailed information of the image, avoid blurred details, and can be applied to most complex scenarios at the same time. Description of the Drawings

[0092] Figure 1 It is the implementation flowchart of the method of the present invention.

[0093] Figure 2 It is the network model structure diagram in the embodiment of the present invention.

[0094] Figure 3 It is the low-frequency attention network structure diagram in the embodiment of the present invention.

[0095] Figure 4 It is the high-frequency attention network structure diagram in the embodiment of the present invention. Specific implementation manners

[0096] Next, in conjunction with the accompanying drawings, the technical solution of the present invention will be specifically described.

[0097] The present invention provides a non-paired underwater image enhancement method based on wavelet transform, as Figures 1-4 shown, including the following steps:

[0098] Step S1: Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained;

[0099] Step S2: Design an underwater image enhancement network based on wavelet transform, and separately process the low-frequency and high-frequency parts of the input image through wavelet transform. The low-frequency part uses a low-frequency processing module to correct the low-frequency part of the image, and the high-frequency part uses a high-frequency processing module for correction. Finally, the high-frequency part and the low-frequency part are combined;

[0100] Step S3: Build a cyclic generative adversarial network and combine it with the underwater image quality enhancement network based on wavelet transform to obtain a non-paired underwater image enhancement network based on wavelet transform;

[0101] Step S4: Design the objective loss function of the non-paired underwater image enhancement network based on wavelet transform;

[0102] Step S5: Use unpaired images to train the non-paired underwater image enhancement network based on wavelet transform to converge to the Nash equilibrium;

[0103] Step S6: Normalize the underwater image to be enhanced, and then input it into the trained non-paired underwater image enhancement network based on wavelet transform to output the enhanced image.

[0104] In this embodiment, the specific implementation of step S1 is as follows:

[0105] Step S11: Divide all unpaired images to be trained into original underwater images and high-quality images;

[0106] Step S12: Uniformly magnify all the unpaired images to be trained to 1.12 times the original image, and perform data augmentation on the magnified images through random cropping and random flipping operations;

[0107] Step S13: Normalize all the images to be trained. Given an image I(i,j), its normalized image is At the pixel position (i,j), calculate the normalized value The formula is as follows:

[0108]

[0109] where (i,j) represents the position of the pixel. The normalized underwater image is used as the underwater image in the subsequent steps, and the normalized high-quality image is used as the high-quality image in the subsequent steps.

[0110] In this embodiment, the specific implementation of step S2 is as follows:

[0111] Step S21: Process the low-frequency and high-frequency parts of the input image separately through wavelet transform; the input of the underwater image enhancement network based on wavelet transform is the normalized underwater image After the image is input into the underwater image enhancement network based on wavelet transform, the image passes through a convolution with a kernel size of 1×1 and a stride of 1, and two convolutions with kernel sizes of 3×3 and a stride of 1 to extract the initial underwater image feature F1 out ; the underwater image feature F1 out is decomposed into a low-frequency feature ll and high-frequency features lh, hl, hh through a wavelet pooling layer;

[0112] Step S22: Design the low-frequency processing module in the underwater image enhancement network based on wavelet transform; the low-frequency feature ll enters the low-frequency processing module of the U-Net network structure combined with the attention mechanism to obtain the processed low-frequency feature ll′; the calculation formula is as follows:

[0113] ll′ = LUNet(ll)

[0114] where ll represents the input low-frequency feature, ll′ represents the processed low-frequency feature, and LUNet(*) represents the low-frequency processing module;

[0115] The specific implementation of step S22 is as follows:

[0116] Step S221: Design the low-frequency processing module of the U-Net network structure with an attention mechanism in Step S22. The low-frequency processing module consists of a U-Net network structure stacked with seven modules. The skip connection parts of the 2nd, 4th, and 6th layers of the low-frequency processing module are each combined with a low-frequency attention module. The downsampling module in the U-Net network structure consists of an activation function, a normalization layer, and a 2×2 convolution. Its calculation formula is as follows:

[0117] X out = ReLU(Conv(Norm(X in )))

[0118] where X in represents the input feature of the downsampling module, and X out represents the output feature of the downsampling module. Conv(*) represents convolution, ReLU(*) represents the activation function, and Norm(*) represents the normalization layer;

[0119] The upsampling module in the U-Net network structure consists of an activation function, a normalization layer, and a 2×2 transposed convolution;

[0120] The specific calculation formula of the low-frequency processing module is as follows:

[0121] X down1 = Down(X in )

[0122] X down2 = Down(X down1 )

[0123] X down3 = Down(X down2 )

[0124] ……X down8 = Down(X down7 )

[0125] X up1 = ADD[Up(X down8 ),X down7

[0126] X up2 = ADD[Up(X up1 ),LAM(X down6 )]

[0127] X up3 = ADD[Up(X up2 ),X down5

[0128] X up4 = ADD[Up(X​​up3 ), LAM(X down4 )]

[0129] X up5 = ADD[Up(X up4 ), X down3 )]

[0130] X up6 = ADD[Up(X up5 ), LAM(X down2 )]

[0131] X out = ADD[Up(X up6 ), X down1 )]

[0132] where X in represents the input feature of the low-frequency processing module, X out represents the output feature of the low-frequency processing module, Down(*) represents the downsampling module, UP(*) represents the upsampling module, LAM(*) represents the low-frequency attention module, and ADD[*] represents the matrix addition operation;

[0133] Step S222, design the low-frequency attention module in step S221; the low-frequency attention module consists of an expansion network in series with a 3×3 convolution and a channel attention module; the expansion network is composed of three layers of parallel convolution blocks with different numbers and a feature concatenation operation in series, and each convolution block is composed of a convolution and an activation function in series; the first layer of the expansion network is a 1×1 convolution block, the second layer is a 1×1 convolution block in series with a 3×3 convolution block, and the third layer is a 1×1 convolution block in series with two 3×3 convolution blocks; the feature output by the expansion network passes through a 3×3 convolution and a channel attention module; the specific calculation formula of the low-frequency attention module is as follows:

[0134] X1 = ReLU(Conv(X in ))

[0135] X2 = ReLU(Conv(ReLU(Conv(X in ))))

[0136] X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in ))))))

[0137] X out = SE_Layer(Conv(Cat[X1, X2, X3]))

[0138] where X inDenote the input features of the low-frequency attention module as X out Denote the output features of the low-frequency attention module as Y. Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation along the channel dimension, and SE_Layer(*) represents the channel attention module;

[0139] Step S23: Design the high-frequency processing module in the underwater image enhancement network based on wavelet transform; The high-frequency processing module uses three parallel high-frequency attention modules to process three high-frequency features lh, hl, and hh respectively, and the output processed high-frequency features are lh', hl', and hh'; The calculation formula is as follows:

[0140] lh′ = HAM(lh)

[0141] hl' = HAM(hl)

[0142] hh' = HAM(hh)

[0143] Where HAM(*) represents the high-frequency attention module; The high-frequency attention module consists of an expansion network in series with a 3×3 convolution, a spatial attention module, and a channel attention module; The expansion network is composed of three layers of parallel convolution blocks with different numbers and a feature concatenation operation in series. Each convolution block is composed of a convolution and an activation function in series; The first layer is a 3×3 convolution block, the second layer is two 3×3 convolution blocks, and the third layer is three 3×3 convolution blocks; The features output by the expansion network pass through a 3×3 convolution, a spatial attention module, and a channel attention module; The calculation formula of the high-frequency attention module is as follows:

[0144] X1 = ReLU(Conv(X in ))

[0145] X2 = ReLU(Conv(ReLU(Conv(X in ))))

[0146] X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in ))))))

[0147] X out = SE_Layer(SA_Layer(Conv(Cat[X1,X2,X3])))

[0148] Where, X in Denote the input features of the high-frequency attention module as X outDenote the output features of the high-frequency attention module, Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation along the channel dimension, SA_Layer(*) represents the spatial attention module, and SE_Layer(*) represents the channel attention module;

[0149] Step S24: Fuse the low-frequency features ll' and high-frequency features lh', hl', hh' and output the enhanced image; the low-frequency features and high-frequency features first pass through a wavelet inverse pooling layer, and then through two convolutions with a convolution kernel of 4×4 and a stride of 1 to output the enhanced image E; the wavelet inverse pooling layer sets its convolution as deconvolution on the basis of the wavelet pooling layer; the calculation formula is as follows:

[0150] E = Conv(Conv(WaveUnpool(ll', lh', hl', hh')))

[0151] where E represents the enhanced image, Conv(*) represents convolution, and WaveUnpool(*) represents the wavelet inverse pooling layer.

[0152] In this embodiment, the specific implementation of step S3 is as follows:

[0153] Step S31: Build a cycle generative adversarial network, including a generator G for enhancing the quality of underwater images W2C , a generator G for converting high-quality images into underwater images C2W , a discriminator D for judging the generated high-quality images C , a discriminator D for discriminating the generated underwater images W ; the network structures of generators G W2C , G C2W use the network structure of the underwater image enhancement network based on wavelet transform designed in step S2.

[0154] Step S32: Input the unpaired underwater images and high-quality images into the cycle generative adversarial network. Generator G W2C enhances the quality of underwater image G W to generate an enhanced image E C , and discriminator D C compares its style with the high-quality image; generator G C2W converts the high-quality image G C into an underwater image E W , and discriminator D W then compares its style with the underwater image.

[0155] In this embodiment, the specific implementation of step S4 is as follows:

[0156] Design the objective loss function of the unpaired underwater image enhancement network based on wavelet transform. The loss includes the loss of the generator network and the loss of the discriminator network;

[0157] The total objective loss function of the generator network is as follows:

[0158] l G = λ1·l GAN + λ2·l cycle + λ3·l identity

[0159] where l GAN , l cycle and l identity are the generator loss, cycle loss, and style loss respectively. λ1, λ2, and λ3 are the coefficients for balancing each loss, and · is the dot product operation of real numbers. The specific calculation formulas for each loss are as follows:

[0160] l GAN = L MSE (1, D C (E C )) + L MSE (1, D W (E W ))

[0161] where D C (*) is the discriminator for the enhanced image E C , and D W (*) is the discriminator for the generated underwater image E W . E C and E W are the enhanced image generated by the generator G W2C and the underwater image generated by the generator G C2W respectively. L MSE (*) is the MSE loss;

[0162] l cycle = L1(R C , G C ) + L1(R W , G W )

[0163] where R W is the image reconstructed by the generator G C2W from the enhanced image E C , and R C is the image reconstructed by the generator G W2C from the generated underwater image E w . G W is the input underwater image, and G C is the input high-quality image. L1(*) is the L1 loss;

[0164] l identity = L1(F C , G C ) + L1(F W , G W )

[0165] where F W is the image obtained by inputting the underwater image G W into the generator G C2W ; F C is the image obtained by inputting the high-quality image G C into the generator G W2C ; G W is the input underwater image, and G C is the input high-quality image. L1(*) is the L1 loss.

[0166] The discriminator network loss is as follows:

[0167]

[0168] where and are the discriminator losses of the discriminator D C and the discriminator D W respectively; · is the dot product operation of real numbers;

[0169]

[0170] where D C (*) is the discriminator for the discriminator-enhanced image E C ; E C is the enhanced image, G C is the input high-quality image, L MSE (*) is the MSE loss, and · is the dot product operation of real numbers;

[0171]

[0172] where D W (*) is the discriminator for the discriminator-generated underwater image E W ; E W is the generated underwater image; G W is the input underwater image, L MSE (*) is the MSE loss, and · is the dot product operation of real numbers.

[0173] In this embodiment, the specific implementation of step S5 is as follows:

[0174] Step S51: Randomly divide the underwater images and high-quality images into multiple batches, with each batch containing N underwater images and N high-quality images. Randomly pair the underwater images and high-quality images to obtain N pairs of images;

[0175] Step S52: Input each pair of underwater image G W and high-quality image G C in the same batch into the wavelet transform-based unpaired underwater image enhancement network in Step S3 to obtain images E C 、E w 、R W 、R c 、F W 、F c ;

[0176] Step S53: According to the objective loss function of the wavelet transform-based unpaired underwater image enhancement network, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network;

[0177] Step S54: Repeat the above image enhancement network training steps of Step S51 to Step S53 in batches until the value of the objective loss function of the wavelet transform-based unpaired underwater image enhancement network converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained wavelet transform-based unpaired underwater image enhancement network.

[0178] In this embodiment, the specific implementation of Step S6 is as follows:

[0179] Step S61: Normalize the underwater image to be enhanced;

[0180] Step S62: Input the image processed in Step S61 into the generator G W2C in the trained wavelet transform-based unpaired underwater image enhancement network to output the enhanced image.

[0181] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention that do not exceed the scope of the technical solutions of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.

Claims

1. An unpaired underwater image enhancement method based on wavelet transform, characterized in that It includes the following steps: S1. Perform data preprocessing, data augmentation, and normalization on the unpaired data to be trained; S2. Design an underwater image enhancement network based on wavelet transform. Process the low-frequency and high-frequency parts of the input image separately through wavelet transform. The low-frequency part is corrected using a low-frequency processing module, and the high-frequency part is corrected using a high-frequency processing module. Finally, combine the high-frequency part and the low-frequency part. The specific implementation is as follows: S21. The input of the underwater image enhancement network based on wavelet transform is the normalized underwater image After the image is input into the underwater image enhancement network based on wavelet transform, the image undergoes a convolution with a convolution kernel of 1×1 and a stride of 1, and two convolutions with convolution kernels of 3×3 and a stride of 1 to extract the initial feature F1 of the underwater image out ; The underwater image feature F1 out is decomposed into low-frequency feature ll and high-frequency features lh, hl, hh through a wavelet pooling layer; S22. Design the low-frequency processing module in the underwater image enhancement network based on wavelet transform; The low-frequency feature ll enters the low-frequency processing module of the U-Net network structure with a combined attention mechanism to obtain the processed low-frequency feature ll'; The specific implementation is as follows: S221. Design the low-frequency processing module of the U-Net network structure with a combined attention mechanism in S22; The low-frequency processing module consists of a U-Net network structure stacked with seven modules. The skip connection parts of the 2nd, 4th, and 6th layers of the low-frequency processing module are each combined with a low-frequency attention module; The downsampling module in the network structure of U-Net consists of an activation function, a normalization layer, and a 2×2 convolution; S222. Design the low-frequency attention module in S221; The low-frequency attention module consists of an extended network in series with a 3×3 convolution and a channel attention module; The extended network is composed of three parallel convolutional blocks with different numbers and a feature splicing operation in series. Each convolutional block is composed of a convolution and an activation function in series; The first layer of the extended network is a 1×1 convolutional block, the second layer is a 1×1 convolutional block in series with a 3×3 convolutional block, and the third layer is a 1×1 convolutional block in series with two 3×3 convolutional blocks; The features output by the extended network pass through a 3×3 convolution and a channel attention module; S23. Design the high-frequency processing module in the underwater image enhancement network based on wavelet transform; The high-frequency processing module uses three parallel high-frequency attention modules to process three high-frequency features lh, hl, hh respectively, and the output processed high-frequency features are lh', hl', hh'; The high-frequency attention module consists of an extended network in series with a 3×3 convolution, a spatial attention module, and a channel attention module; The extended network is composed of three parallel convolutional blocks with different numbers and a feature splicing operation in series. Each convolutional block is composed of a convolution and an activation function in series; The first layer is a 3×3 convolutional block, the second layer is two 3×3 convolutional blocks, and the third layer is three 3×3 convolutional blocks; The features output by the extended network pass through a 3×3 convolution, a spatial attention module, and a channel attention module; S24. Fuse the low-frequency feature ll' and the high-frequency features lh', hl', hh' and output the enhanced image; The low-frequency feature and the high-frequency feature first pass through a wavelet inverse pooling layer, and then through two convolutions with a convolution kernel of 4×4 and a stride of 1 to output the enhanced image E; The wavelet inverse pooling layer sets its convolution as a transposed convolution based on the wavelet pooling layer; S3. Build a cycle generative adversarial network and combine it with the underwater image quality enhancement network based on wavelet transform to obtain an unpaired underwater image enhancement network based on wavelet transform; S4. Design the objective loss function of the unpaired underwater image enhancement network based on wavelet transform; S5. Use unpaired images to train the unpaired underwater image enhancement network based on wavelet transform to converge to the Nash equilibrium; S6. Normalize the underwater image to be enhanced, and then input it into the trained unpaired underwater image enhancement network based on wavelet transform to output the enhanced image.

2. The unpaired underwater image enhancement method based on wavelet transform according to claim 1, characterized in that The specific implementation of S1 is as follows: S11. Divide all unpaired images to be trained into original underwater images and high-quality images; S12. Uniformly magnify all unpaired images to be trained to 1.12 times the original image, and perform data augmentation on the magnified images through random cutting and random flipping operations; S13. Normalize all the images to be trained. Given an image I(i, j), its normalized image is At the pixel position (i, j), calculate the normalization value using the following formula: Among them, (i, j) represents the position of the pixel. The normalized underwater image is used as the underwater image in the subsequent steps, and the normalized high-quality image is used as the high-quality image in the subsequent steps.

3. A non-paired underwater image enhancement method based on wavelet transform according to claim 1, characterized in that In S22, the low-frequency feature ll enters the low-frequency processing module of the U-Net network structure combined with the attention mechanism to obtain the processed low-frequency feature ll'; the calculation formula is as follows: ll' = LUNet(ll) Among them, ll represents the input low-frequency feature, ll' represents the processed low-frequency feature, and LUNet(*) represents the low-frequency processing module; In S23, the high-frequency processing module uses three parallel high-frequency attention modules to process the three high-frequency features lh, hl, hh respectively, and the output processed high-frequency features are lh', hl', hh'; the calculation formula is as follows: lh' = HAM(lh) hl' = HAM(hl) hh' = HAM(hh) Among them, HAM(*) represents the high-frequency attention module; In S24, the low-frequency feature and the high-frequency feature first pass through a wavelet inverse pooling layer, and then pass through two convolutions with a convolution kernel of 4×4 and a stride of 1 to output the enhanced image E; the calculation formula is as follows: E = Conv(Conv(WaveUnpool(ll', lh', hl', hh'))) Among them, E represents the enhanced image, Conv(*) represents convolution, and WaveUnPool(*) represents the wavelet inverse pooling layer.

4. A non-paired underwater image enhancement method based on wavelet transform according to claim 3, characterized in that In S221, the downsampling module in the network structure of U-Net, the calculation formula is as follows: X out = ReLU(Conv(Norm(Xi n ))) Among them, X in represents the input feature of the downsampling module, and X out represents the output feature of the downsampling module. Conv(*) represents convolution, ReLU(*) represents the activation function, and Norm(*) represents the normalization layer; The upsampling module in the network structure of U-Net consists of an activation function, a normalization layer, and a 2×2 transposed convolution; The specific calculation formula of the low-frequency processing module is as follows: X down1 = Down(X in ) X down2 = Down(X down1 ) X down3 = Down(X down2 ) …… X down8 = Down(X down7 ) X up1 = ADD[Up(X down8 ), X down7 ​ X up2 = ADD[Up(X up1 ), LAM(X down6 )] X up3 = ADD[Up(X up2 ), X down5 ​ X up4 = ADD[Up(X up3 ),LAM(X down4 )] X up5 = ADD[Up(X up4 ), X down3 ​ X up6 = ADD[Up(X up5 ), LAM(X down2 )] X out = ADD[Up(X up6 ), X down1 ​ Among them, X in represents the input feature of the low-frequency processing module, and X out represents the output feature of the low-frequency processing module. Down(*) represents the downsampling module, Up(*) represents the upsampling module, LAM(*) represents the low-frequency attention module, and ADD[*] represents the matrix addition operation; In S222, the specific calculation formula of the low-frequency attention module is as follows: X1 = ReLU(Conv(X in )) X2 = ReLU(Conv(ReLU(Conv(X in )))) X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in )))))) X out = SE_Layer(Conv(Cat[X1, X2, X3])) Among them, X in represents the input feature of the low-frequency attention module, and X out represents the output feature of the low-frequency attention module. Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation along the channel dimension, and SE_Layer(*) represents the channel attention module.

5. The unpaired underwater image enhancement method based on wavelet transform according to claim 3, characterized in that, In S23, the calculation formula of the high-frequency attention module is as follows: X1 = ReLU(Conv(X in )) X2 = ReLU(Conv(ReLU(Conv(X in )))) X3 = ReLU(Conv(ReLU(Conv(ReLU(Conv(X in )))))) X out = SE_Layer(SA_Layer(Conv(Cat[X1,X2,X3]))) Among them, X in represents the input feature of the high-frequency attention module, and X out represents the output feature of the high-frequency attention module. Conv(*) represents convolution, ReLU(*) represents the activation function, Cat[*] represents the feature concatenation operation along the channel dimension, SA_Layer(*) represents the spatial attention module, and SE_Layer(*) represents the channel attention module.

6. The unpaired underwater image enhancement method based on wavelet transform according to claim 1, wherein The specific implementation of S3 is as follows: S31. Build a cyclic generative adversarial network, including a generator G for enhancing the quality of underwater images W2C , a generator G for converting high-quality images into underwater images C2W , a discriminator D for judging the generated high-quality images C , a discriminator D for discriminating the generated underwater images W ; The network of the generator G W2C , G C2W uses the network structure of the underwater image enhancement network based on wavelet transform designed in S2; S32. Input the unpaired underwater images and high-quality images into the cyclic generative adversarial network. The generator G W2C Enhance the quality of the underwater image G W to generate an enhanced image E C , and the discriminator D C Compare its style with the high-quality image; the generator G C2W Convert the high-quality image G C into an underwater image E W , and the discriminator D W Then compare its style with the underwater image again.

7. A non-paired underwater image enhancement method based on wavelet transform according to claim 1, characterized in that, The specific implementation of S4 is as follows: Design the objective loss function of the unpaired underwater image enhancement network based on wavelet transform, and the loss includes the loss of the generation network and the loss of the discriminator network; The total objective loss function of the generation network is as follows: Among them, and are the generator loss, cycle loss, and style loss respectively. λ1, λ2, and λ3 are coefficients for balancing each loss, and · is the dot product operation of real numbers. The specific calculation formulas for each loss are as follows: Among them, D C (*) is the discriminator for the discriminative enhanced image E C , and D W (*) is the discriminator for the discriminatively generated underwater image E W ; E C and E W are respectively the enhanced image generated by the generator G W2C and the underwater image generated by the generator G C2W , and L MSE (*) is the MSE loss; Among them, R W is the image reconstructed by the generator G C2W from the enhanced image E C ; R C is the image reconstructed by the generator G W2C from the generated underwater image E w ; G W is the input underwater image; G C is the input high-quality image, and L1(*) is the L1 loss; Among them, F W is the underwater image G W input into the generator G C2W to obtain the image, F C is the high-quality image G C input into the generator G W2C to obtain the image, G W is the input underwater image, G C is the input high-quality image, and L1(*) is the L1 loss; The loss of the discriminator network is as follows: Among them, and are the discriminator losses of discriminator D C and the discriminator losses of discriminator D W respectively; · is the dot product operation of real numbers. Among them, where D C (*) is the discriminator for discriminating the enhanced image E C ; E C is the enhanced image, G C is the input high-quality image, L MSE (*) is the MSE loss, · is the dot product operation of real numbers; Among them, D W (*) is the discriminator for discriminating the generated underwater image E W ; E W is the generated underwater image; G W is the input underwater image, L MSE (*) is the MSE loss, and · is the dot product operation of real numbers.

8. A method for enhancing unpaired underwater images based on wavelet transform according to claim 1, characterized in that, The specific implementation of S5 is as follows: S51. Randomly divide the underwater images and high-quality images into multiple batches, where each batch contains N underwater images and N high-quality images, and randomly pair the underwater images and high-quality images to obtain N pairs of images; S52. Input each pair of underwater images G W and high-quality images G C in the same batch into the unpaired underwater image enhancement network based on wavelet transform in S3 to obtain images E C 、E w 、R W 、R c 、F W 、F c ; S53. According to the objective loss function of the non-paired underwater image enhancement network based on wavelet transform, use the backpropagation method to calculate the gradients of the parameters in the image enhancement network, and use the stochastic gradient descent method to update the parameters of the image enhancement network; S54. Repeat the image enhancement network training steps of S51 to S53 in batches until the value of the objective loss function of the non-paired underwater image enhancement network based on wavelet transform converges to the Nash equilibrium, save the network parameters, complete the training process of the image enhancement network, and obtain the trained non-paired underwater image enhancement network based on wavelet transform.

9. A method for unpaired underwater image enhancement based on wavelet transform according to claim 1, characterized in that, The specific implementation of S6 is as follows: S61. Normalize the underwater image to be enhanced; S62. Input the image processed in S61 into the generator G in the trained unpaired underwater image enhancement network based on wavelet transform W2C , and output the enhanced image.