Deep Learning-Based Super-Resolution Underwater Image Enhancement Method and System
By improving the combination of the generative adversarial network model and the depth residual multiplier, the color shift and low contrast problems of underwater images are solved, efficient enhancement of underwater images is achieved, and the resolution and clarity of the image are improved.
Patent Information
- Application Number
- CN202210583321.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-05-25
AI Technical Summary
Existing underwater image enhancement technologies are difficult to effectively solve the color shift, details and low contrast problems of underwater images, especially in the real-time pre-processing of visually guided underwater robots and single-image super-resolution.
The improved generative adversarial network model based on the generative adversarial network and the depth residual multiplier is adopted. The underwater image is processed by the enhanced algorithm loss function and the depth residual multiplier, and combined with the image content and the quality index loss function, the image details recovery and color correction are achieved, and the resolution is improved.
It significantly improves the resolution and visual quality of underwater images, enhances the sharpness and contrast of the image, expands the resolution of underwater images to twice the original, and improves the overall visual effect of the image.
Smart Images

Figure CN115034965B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater image processing, and particularly to a super-resolution underwater image enhancement method and system based on deep learning. Background Art
[0002] When exploring and developing marine resources, obtaining underwater information through underwater robots is the current mainstream method. However, the underwater environment is relatively complex. For example, the turbulence of the underwater environment, the scattering of light caused by suspended matter in seawater, and the loss of color channels caused by the attenuation of light make the underwater images present color deviation, lack of details, and low contrast. Therefore, it is of great significance to improve the quality of underwater images and solve the problems of lack of details and blurriness in underwater images. Existing technologies include a real-time underwater image enhancement model based on conditional generative adversarial networks, which evaluates the perceptual image quality according to global content, color, local texture, and style information, providing an improved standard model performance for underwater object detection, human pose estimation, and saliency prediction. However, it is only applicable to real-time preprocessing of vision-guided underwater robots, and a generative model based on deep residual networks evaluates the perceptual quality of images according to the global content, color, and local style information of the images, but only improves the single-image super-resolution of underwater images. Therefore, the technology for correcting color deviation and improving the color, clarity, and contrast of underwater images needs to be further improved. Summary of the Invention
[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a super-resolution underwater image enhancement method and system based on deep learning, which can improve the resolution of underwater images and at the same time improve the visual quality of underwater images.
[0004] The first technical solution adopted by the present invention is: a super-resolution underwater image enhancement method based on deep learning, including the following steps:
[0005] Construct an improved generative adversarial network model based on a generative adversarial network and a deep residual multiplier;
[0006] Input the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image.
[0007] Further, the improved generative adversarial network model further includes:
[0008] An upsampling module, including five convolutional layers, with the activation function being the Leaky-ReLU function, used to extract the features of the underwater image;
[0009] A downsampling module, including four transposed convolutional layers, with the activation function being the Drop-out function, used to restore the details of the underwater image.
[0010] Further, the generative adversarial network is specifically as follows:
[0011]
[0012] In the above formula, G * represents the generative adversarial network, G represents the generator of the generative adversarial network, D represents the discriminator of the generative adversarial network, α, λ1, λ c and λ2 represent hyperparameters, L1 represents the first loss function, L2 represents the second loss function, L con (G) represents the image content metric loss function, L IQM represents the image quality metric loss function.
[0013] Further, the step of inputting the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image specifically includes:
[0014] Obtain the underwater image through an underwater robot;
[0015] Based on the generative adversarial network of the improved generative adversarial network model, enhance the underwater image through the enhancement algorithm loss function to obtain an enhanced underwater image;
[0016] Based on the deep residual multiplier of the improved generative adversarial network model, perform interpolation processing on the enhanced underwater image through the super-resolution algorithm to obtain an enhanced super-resolution underwater image.
[0017] Further, the step of enhancing the underwater image through the enhancement algorithm loss function based on the generative adversarial network of the improved generative adversarial network model to obtain an enhanced underwater image specifically includes:
[0018] The enhancement algorithm loss function includes the first loss function, the second loss function, the image content metric loss function, the Markov discriminator, and the image quality metric loss function;
[0019] Solve the global similarity of the underwater image through the first loss function and the second loss function to obtain a clear underwater image;
[0020] Correct the local texture and style information of the underwater images with similar content through the Markov discriminator to obtain a colorless-offset underwater image;
[0021] Map the image quality of the underwater image through the image content metric loss function and the image quality metric loss function to obtain an underwater image with similar content;
[0022] Combine the clear underwater image, the colorless-offset underwater image, and the underwater image with similar content to obtain an enhanced underwater image.
[0023] Furthermore, the step of performing mapping processing on the image quality of the underwater image through the image content metric loss function and the image quality metric loss function to obtain an underwater image with similar content specifically includes:
[0024] Input the underwater image and the desired image;
[0025] Based on the image content metric loss function, calculate the contrast gain value of the desired image in the underwater image;
[0026] Calculate the sharpness gain value of the underwater image through the Sobel operator;
[0027] Based on the image quality metric loss function, minimize the contrast gain value and the sharpness gain value, and output an underwater image with similar content.
[0028] Furthermore, the step of performing interpolation processing on the enhanced underwater image through the super-resolution algorithm by the deep residual multiplier based on the improved generative adversarial network model to obtain an enhanced super-resolution underwater image specifically includes:
[0029] The deep residual multiplier includes a convolutional layer, a residual layer, an additional convolutional layer, and a deconvolutional layer;
[0030] Through the convolutional layer of the deep residual multiplier, perform feature extraction processing on the enhanced underwater image to obtain a feature image;
[0031] Through the residual layer of the deep residual multiplier, perform non-linear mapping processing on the feature image to obtain a mapped image;
[0032] Through the additional convolutional layer and the deconvolutional layer of the deep residual multiplier, perform interpolation and magnification processing on the mapped image to obtain an enhanced super-resolution underwater image.
[0033] The second technical solution adopted by the present invention is: a super-resolution underwater image enhancement system based on deep learning, including:
[0034] A construction module, based on the generative adversarial network and the deep residual multiplier, constructs an improved generative adversarial network model;
[0035] A training module, used to input the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image.
[0036] The beneficial effects of the method and system of the present invention are as follows: Through the generative adversarial network, the present invention details the preliminary underwater images. Specifically targeting problems such as low contrast, low clarity, and color deviation in underwater images, it conducts guidance and repair through the enhanced algorithm loss function of the generative adversarial network. Based on the deep residual multiplier, interpolation processing is performed on the enhanced underwater images, expanding the resolution of the underwater enhanced images to twice the original resolution, and obtaining enhanced super-resolution underwater images. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the flowchart of the steps of the super-resolution underwater image enhancement method based on deep learning of the present invention;
[0038] Figure 2 is the structural block diagram of the super-resolution underwater image enhancement system based on deep learning of the present invention;
[0039] Figure 3 is the schematic design diagram of the generator of the improved generative adversarial network model of the present invention;
[0040] Figure 4 is the schematic diagram of the results of the underwater image without using the present invention and the underwater image using the improved generative adversarial network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following further elaborates the present invention in detail in conjunction with the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0042] Referring to Figure 1 , the present invention provides a super-resolution underwater image enhancement method based on deep learning, and the method includes the following steps:
[0043] S1. Based on the generative adversarial network and the deep residual multiplier, construct an improved generative adversarial network model;
[0044] Specifically, combine the generative adversarial network and the deep residual multiplier to construct an improved generative adversarial network model. The generative adversarial network is specifically as follows:
[0045]
[0046] In the above formula, G * represents the generative adversarial network, G represents the generator of the generative adversarial network, D represents the discriminator of the generative adversarial network, α, λ1, λ c and λ2 represent hyperparameters, L1 represents the first loss function, L2 represents the second loss function, L con(G) represents the loss function of the image content metric, L IQM represents the loss function of the image quality metric;
[0047] Referring to Figure 3 , the improved generative adversarial network model is divided into two parts: downsampling and upsampling. The downsampling layer can perform feature extraction processing on underwater images, and the upsampling layer can restore the details of underwater images. The total number of elements contained in the image matrix at its input and output is 256×256×3. The light gray square on the left represents the convolutional part of downsampling, which progressively extracts the features of the image step by step. The right side is the transposed convolutional part of upsampling, which restores the details of the image step by step. The generator of the algorithm is an encoder-decoder network, as shown by the gray horizontal line in the figure. There are connections between the mirror layers. Skip connections are applied between the convolutional layer and the transposed convolutional layer of this network model. It can accelerate the network training process while avoiding the loss of low-level image features. The main role of the encoder and decoder structures is to remove the noise of the image. The input image will obtain a string of features smaller than the original image after downsampling encoding, which is equivalent to compression; after decoding, ideally, it can be restored to the original real image. In the encoding-decoding framework, the output of each encoder is connected to the corresponding mirror decoder. In the network model of the generator, convolutional layers are used as encoding to reduce the noise in underwater images, and transposed convolutional layers are used as decoding to restore the lost details of underwater images and refine the underwater images pixel by pixel. Through the downsampling of the encoder and the upsampling of the decoder, the finally restored feature map integrates more low-level features. Among them, multiple upsamplings make the obtained picture information more refined. During the downsampling and upsampling processes, the generator adopts a step-by-step progressive method, that is, the picture size is gradually reduced during downsampling, and the picture size is gradually enlarged during upsampling. Such a progressive symmetric structure can ensure the extraction of image detail information, and the model enhances underwater images in a pixel-to-pixel manner. Among them, the convolutional part is used for denoising and retaining key detail features, while the transposed convolutional part is used to refine the details of each feature map corresponding to the convolutional layer. In the enhanced algorithm loss function, the generator model realizes fast inference. Specifically, the encoder learns 512 feature maps with a size of 16×16, and the decoder learns according to the feature maps and the input from the skip connections to generate an enhanced image as the output. Each layer of downsampling uses a convolution with a 4×4 filter, and a leaky rectified linear unit activation function and batch normalization are also applied.
[0048] S2. Input the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image.
[0049] S21. Obtain the underwater image through an underwater robot;
[0050] S22. The generative adversarial network based on the improved generative adversarial network model enhances the underwater image through the enhanced algorithm loss function to obtain an enhanced underwater image;
[0051] S221. The enhanced algorithm loss function includes a first loss function, a second loss function, an image content metric loss function, a Markov discriminator, and an image quality metric loss function;
[0052] Specifically, the first loss function is as follows:
[0053] L1(G) = E X,Y,Z [‖Y - G(X, Z)‖1]
[0054] In the above formula, L1(G) represents the first loss function, Y represents the true value, and G(·) represents the predicted value;
[0055] The second loss function is as follows:
[0056] L2(G) = E X,Y,Z [‖Y - G(X, Z)‖2]
[0057] In the above formula, L2(G) represents the second loss function;
[0058] The image content metric loss function is as follows:
[0059] L con (G) = E X,Y,Z [‖φ(Y) - φ(G(X, Z))‖2]
[0060] In the above formula, L con (G) represents the image content metric loss function, and φ(·) represents the high-level features extracted by the encoder layer of the pre-trained VGG-19 network;
[0061] The image quality metric loss function is as follows:
[0062]
[0063] In the above formula, L IQM represents the image quality metric loss function, q I represents the image quality metric, β I represents the metric weight, and i represents the set of image quality metrics
[0064] S222. Solve the global similarity of the underwater image through the first loss function and the second loss function to obtain a clear underwater image;
[0065] Specifically, in the global similarity part, a first loss function is adopted and a second loss function is introduced. Adding the first loss function and the second loss function to the objective function enables the generator to learn sampling. Among them, the first loss function helps to generate clearer images, and the second loss function is sensitive to outliers and can find more stable and closer solutions (by setting the derivative to 0), and its solution is easier than the first loss function, which is beneficial to the rapid iteration of the network and will be more stable and accurate during the optimization process. The first loss function calculates the expectation of the L1 norm of the difference between the true value Y and the predicted value G, and the second loss function calculates the expectation of the L2 norm of the difference between the true value Y and the predicted value G. Among them, the L1 norm refers to the sum of the absolute values of each element in the vector, and the L2 norm is the square root of the sum of the squares of each element of the vector.
[0066] S223. Use a Markov discriminator to correct the local texture and style information of underwater images with similar content to obtain colorless-offset underwater images;
[0067] Specifically, the image content function is defined as the high-level features extracted by the (block5_conv2) layer of the pre-trained model VGG-19 network. In the local texture and style information part, the Markov discriminator architecture is adopted to effectively capture the high-frequency information related to the local texture and style of underwater images. Therefore, through the discriminator, the consistency between the enhanced underwater degraded image and the target image in terms of local texture and style can be strengthened in an adversarial manner.
[0068] S224. Map the image quality of the colorless-offset underwater image through an image quality metric loss function to obtain an enhanced underwater image;
[0069] S225. Combine the clear underwater image, the colorless-offset underwater image, and the underwater image with similar content to obtain an enhanced underwater image.
[0070] Specifically, in the image content part, a content loss term is added to the objective function. This loss function helps to provide finer texture details of underwater images and promotes the generator to generate enhanced images with content similar to real images. In the image quality metric loss function of the enhanced algorithm loss function, a custom image quality metric loss function is added according to the image quality metric to guide the learning of the mapping between the input distorted image and the expected output enhanced image. By minimizing the image quality metric loss function, it can effectively help the generator generate high-quality underwater images.
[0071] S2241. Input the underwater image and the expected image;
[0072] S2242. Based on the image content metric loss function, calculate the contrast gain value of the expected image in the underwater image;
[0073] S2243. Calculate the sharpness gain value of the underwater image through the Sobel operator;
[0074] S2244. Minimize the contrast gain value and the sharpness gain value based on the image quality metric loss function, and output an underwater image with similar content.
[0075] Specifically, the image quality metrics include contrast and sharpness, which are obtained through Python programming and measure the loss between the predicted value and the true value during the training process. In the added image quality metric loss function, two metrics, contrast and sharpness, which are closely related to human visual perception, are selected to form the image quality metric set. Calculate the contrast gain of the enhanced image on the degraded image to make the contrast of the enhanced image as close as possible to the reference image. In the sharpness recovery gain, the Sobel operator is used to calculate the gradient intensity. After calculating the gain of each quality metric, minimize the image quality metric loss. In the loss function of the enhancement algorithm based on the generative adversarial network, use the image quality metric loss function in conjunction with the above other loss functions to guide the optimization process. Through the image quality metric error, the weights of the network can be updated.
[0076] S23. Based on the depth residual multiplier of the improved generative adversarial network model, perform interpolation processing on the enhanced underwater image through the super-resolution algorithm to obtain an enhanced super-resolution underwater image.
[0077] Specifically, the core of the depth residual multiplier is a fully convolutional depth residual block, which is used to learn the 2× interpolation of the image. It learns to magnify the spatial dimension of the input feature by two times. The depth residual multiplier includes a part of convolutional layers, followed by 8 repeated residual layers, then another convolutional layer, and finally a transposed convolutional layer for magnification. Filters, rectified linear unit functions, and batch normalization are used in each layer. As a whole, the depth residual multiplier is a 10-layer residual network. The depth residual multiplier uses a series of 2D convolutions of size 3×3 in the repeated residual blocks and a series of 2D convolutions of size 4×4 in the rest of the network to learn spatial interpolation from paired training data. For example, an underwater image with an input size of 320×240×3 can output an underwater image with a size of 640×480×3;
[0078] Refer to Figure 4 , the schematic diagram of the results of the underwater image without using the present invention and the underwater image using the improved generative adversarial network model of the present invention, and the quantitative comparison of the proposed method based on the UFO-120 dataset. The specific experimental data is shown in the following table:
[0079] Model Underwater Image Quality Measurement Structural Similarity Peak Signal-to-Noise Ratio Relative Global Histogram Stretching 2.36±0.33 0.75±0.06 20.05±3.10 Unsupervised Color Correction 2.81±0.47 0.73±0.07 20.79±4.48 Multi-Scale Fusion 2.76±0.45 0.79±0.09 21.32±3.30 MS-Retinex 2.69±0.59 0.75±0.10 21.69±3.60 Water-Net 2.83±0.48 0.79±0.05 22.46±1.90 UWIQE-GAN 2.89±0.47 0.70±0.07 23.55±2.65 The method proposed by the present invention 3.07±0.42 0.73±0.07 23.65±2.72
[0080] Refer to Figure 2, a super-resolution underwater image enhancement system based on deep learning, comprising:
[0081] A construction module, based on a generative adversarial network and a deep residual multiplier, constructs an improved generative adversarial network model;
[0082] A training module, configured to input underwater images into the improved generative adversarial network model for training to generate enhanced super-resolution underwater images.
[0083] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0084] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A super-resolution underwater image enhancement method based on deep learning, characterized in that, It includes the following steps: Based on a generative adversarial network and a deep residual multiplier, an improved generative adversarial network model is constructed; Input the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image; The step of inputting the underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image specifically includes: Obtain the underwater image through an underwater robot; Based on the generative adversarial network of the improved generative adversarial network model, enhance the underwater image through an enhancement algorithm loss function to obtain an enhanced underwater image; Based on the deep residual multiplier of the improved generative adversarial network model, perform interpolation processing on the enhanced underwater image through a super-resolution algorithm to obtain an enhanced super-resolution underwater image; The step of enhancing the underwater image through the enhancement algorithm loss function based on the generative adversarial network of the improved generative adversarial network model to obtain an enhanced underwater image specifically includes: The enhancement algorithm loss function includes a first loss function, a second loss function, an image content metric loss function, a Markov discriminator, and an image quality metric loss function; Solve the global similarity of the underwater image through the first loss function and the second loss function to obtain a clear underwater image; Correct the local texture and style information of underwater images with similar content through the Markov discriminator to obtain a colorless-offset underwater image; Map the image quality of the underwater image through the image content metric loss function and the image quality metric loss function to obtain an underwater image with similar content; Combine the clear underwater image, the colorless-offset underwater image, and the underwater image with similar content to obtain an enhanced underwater image; The step of mapping the image quality of the underwater image through the image content metric loss function and the image quality metric loss function to obtain an underwater image with similar content specifically includes: Input the underwater image and the desired image; Based on the image content metric loss function, calculate the contrast gain value of the desired image in the underwater image; Calculate the sharpness gain value of the underwater image through the Sobel operator; Minimize the contrast gain value and the sharpness gain value based on the image quality metric loss function and output an underwater image with similar content.
2. The super-resolution underwater image enhancement method based on deep learning according to claim 1, wherein The improved generative adversarial network model further includes: An upsampling module, which includes five convolutional layers, and the activation function is the Leaky-ReLU function, and is used to extract the features of the underwater image; A downsampling module, which includes four deconvolutional layers, and the activation function is the Drop-out function, and is used to restore the details of the underwater image.
3. The super-resolution underwater image enhancement method based on deep learning according to claim 2, characterized in that, The specific form of the generative adversarial network is as follows: In the above formula, G * represents a generative adversarial network, G represents the generator of the generative adversarial network, D represents the discriminator of the generative adversarial network, α, λ1, λ c and λ2 represent hyperparameters, L1 represents the first loss function, L2 represents the second loss function, L con (G) represents the loss function of the image content metric, L IQM represents the loss function of the image quality metric.
4. The super-resolution underwater image enhancement method based on deep learning according to claim 3, wherein The step of performing interpolation processing on the enhanced underwater image through the super-resolution algorithm based on the deep residual multiplier of the improved generative adversarial network model to obtain an enhanced super-resolution underwater image specifically includes: The deep residual multiplier includes a convolutional layer, a residual layer, an additional convolutional layer, and a deconvolutional layer; Through the convolutional layer of the deep residual multiplier, perform feature extraction processing on the enhanced underwater image to obtain a feature image; Through the residual layer of the deep residual multiplier, perform non-linear mapping processing on the feature image to obtain a mapped image; Through the additional convolutional layer and deconvolutional layer of the deep residual multiplier, perform interpolation magnification processing on the mapped image to obtain an enhanced super-resolution underwater image.
5. A super-resolution underwater image enhancement system based on deep learning, characterized in that, A device for executing the deep learning-based super-resolution underwater image enhancement method according to claim 1, comprising the following modules: A construction module, based on the generative adversarial network and the deep residual multiplier, constructs an improved generative adversarial network model; A training module, configured to input an underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image.
Citation Information
Patent Citations
Single image super-resolution reconstruction method and system
CN113298718A