A comprehensive underwater image sharpening method based on exposure correction

Through an exposure correction method and combined with deep learning technology, low-illumination enhancement network and improved cyclic adversarial network CycleGAN solves the problems of uneven light, color cast and blurring of underwater images, achieving high-quality enhancement of underwater images and meeting the needs of underwater operations and robot applications.

CN120125453BActive Publication Date: 2025-08-08SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510621597.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-08
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

When processing underwater images, the prior art faces problems of uneven light, color cast and blur, especially in extreme environments, and the effect is limited, making it difficult to meet the needs of high-quality underwater images.

Method used

Using an exposure correction method and combined with deep learning technology, image enhancement is performed through low-illumination enhancement network and cyclic adversarial network (CycleGAN), and using the depth curve estimation network (DCE-Net) to adjust exposure correction and pyramid image fusion, combined with the improved cyclic adversarial network CycleGAN for underwater image enhancement, solving the problems of uneven light, color cast and blur.

Benefits of technology

It significantly improves the quality and usability of underwater images, provides clearer and more accurate image information, and meets the needs of underwater operations and underwater robot applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125453B_ABST
    Figure CN120125453B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of underwater image processing, and discloses a comprehensive underwater image clarity method based on exposure correction. The present invention first inputs an underwater non-uniform illumination image and its inverted image into a low-light enhancement network, and simultaneously obtains a low-light enhancement image and a reduced-exposure image, and obtains an exposure-corrected image through pyramid image fusion. Then, based on the exposure correction, a cyclic adversarial network is introduced to perform underwater image enhancement processing. During the training process, the interaction between the forward generator and the backward generator is utilized to learn to convert the underwater image into an effect consistent with the style of a real air image, while ensuring that the generated image is consistent with the real underwater image and maintains cyclic consistency. The method of the present invention can effectively solve various common quality problems in underwater images, thereby providing clearer and more accurate image information to meet the needs of underwater operations and underwater robot applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater image processing and relates to a comprehensive underwater image clarity method based on exposure correction. Background Art

[0002] The generation process of underwater images is significantly more complex than that of terrestrial images and is affected by a variety of environmental factors. First, due to the wavelength differences of underwater light, different wavelengths of light attenuate at different rates in water, with red light attenuating most significantly. As a result, underwater images often exhibit a blue-green color cast. Second, physical factors such as suspended particles and dissolved substances in the water interfere with light propagation, causing noticeable image blur. This can severely impact image quality, particularly in deep water and low-visibility environments. To improve the brightness and clarity of underwater images, artificial light sources are often used for supplemental illumination. However, such artificial light sources often result in areas near the light point appearing too bright, while areas further away appear too dark, resulting in uneven illumination. These issues, including uneven illumination, color cast, and blurring, result in generally poor underwater image quality, severely impacting applications such as underwater operations and underwater robotics. Therefore, effectively addressing these issues and improving underwater image quality has become a pressing and critical task for researchers in this field.

[0003] Several methods have attempted to address these issues with underwater images. Traditional underwater image enhancement methods primarily focus on exposure correction, dehazing, and color correction. For example, exposure correction and dehazing algorithms based on traditional image processing use histogram equalization, color correction, and dehazing algorithms to compensate for uneven illumination, color cast, and blur in underwater images. Exposure correction typically analyzes the image's illumination distribution and adjusts the brightness of overly bright or dark areas to achieve a balanced image brightness. However, this method is limited in effectiveness in extreme underwater environments, particularly those with extremely low light levels or poor water quality, and cannot meet the requirements for high-quality underwater images. Exposure correction techniques improve image clarity by adjusting image brightness and contrast, but due to the unevenness of underwater illumination and the complexity of light wave propagation, a single exposure correction method often cannot address all issues. Underwater image enhancement algorithms based on color restoration and reflectance estimation reduce color cast by restoring the color information of underwater images. However, these methods often have shortcomings in addressing image blur and restoring detail.

[0004] In recent years, deep learning-based image enhancement methods have gained popularity, particularly the application of Cycle Generative Adversarial Networks (GANs) to underwater image enhancement, which has achieved promising results. For example, deep learning-based underwater image enhancement methods train deep neural networks to learn the characteristics of underwater images and leverage the model's inference capabilities to achieve goals such as image dehazing, decolorization, and detail enhancement. However, existing deep learning methods still face challenges when processing images in diverse water conditions, including large data requirements, complex training processes, and insufficient algorithmic robustness. Underwater image enhancement techniques based on multi-scale fusion attempt to enhance images by fusing information at different scales through multi-scale image processing techniques. While these methods can address image blurring to a certain extent, they still face limitations when navigating complex underwater environments due to the significant differences in illumination across different regions of underwater images. While these methods can mitigate image blurring and color cast, they still face challenges. For example, image enhancement effectiveness remains limited under extreme lighting conditions, and some methods struggle to process fine details. Summary of the Invention

[0005] The purpose of this invention is to propose a comprehensive underwater image clarity method based on exposure correction, aiming to solve the problems of uneven illumination, color cast and blur in the underwater image processing process in the prior art, and to combine deep learning technology for image enhancement, thereby significantly improving the quality and usability of underwater images, and thus providing clearer and more accurate image information.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:

[0007] A method for comprehensive clarity of underwater images based on exposure correction includes the following steps:

[0008] Step 1. First, the underwater non-uniform illumination image and its inverted image are input into the low-light enhancement network to obtain a low-light enhancement image and a reduced-exposure image. Then, the exposure-corrected image is obtained through pyramid image fusion.

[0009] The low-light enhancement network is a low-light enhancement algorithm based on the deep curve estimation network. A deep curve estimation network is designed to estimate a set of best-fit light enhancement curves for a given input image. The deep curve estimation network then iteratively applies the curves to map all pixels of the RGB channels of the input image to obtain the final enhanced image.

[0010] Step 2. Based on exposure correction, a cyclic adversarial network (CycleGAN) is introduced to achieve underwater image enhancement. During CycleGAN training, the interaction between the forward generator and the backward generator is utilized to learn to transform underwater images into effects consistent with the style of real air images, while ensuring that the generated images are consistent with the real underwater images to maintain cycle consistency.

[0011] The present invention has the following advantages:

[0012] As described above, the present invention proposes a comprehensive underwater image clarity method based on exposure correction, which proposes an underwater image exposure correction algorithm based on dynamic curve adjustment and an underwater image enhancement algorithm based on an improved recurrent adversarial network. The method aims to solve the problems of uneven illumination, color cast and blurring existing in the prior art, and combines deep learning technology for image enhancement, which significantly improves the quality and usability of underwater images, and can effectively solve various common quality problems in underwater images, thereby providing clearer and more accurate image information to meet the needs of underwater operations and underwater robot applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Flowchart of a method for comprehensive sharpening of underwater images based on exposure correction in an embodiment of the present invention;

[0014] Figure 2 1 is a flowchart of exposure correction based on dynamic curve adjustment in an embodiment of the present invention;

[0015] Figure 3 This is a network structure diagram of the depth curve estimation network DCE-Net in an embodiment of the present invention;

[0016] Figure 4 Schematic diagram of the multi-exposure rate fusion process in an embodiment of the present invention; wherein, Figure 4 (a) in the figure is the original image. Figure 4 (b) in the figure is the de-exposure image. Figure 4 (c) in the figure is the low-light enhanced image; Figure 4 (d) in the figure is the fused image;

[0017] Figure 5 This is a flowchart of underwater image enhancement based on an improved cyclic adversarial network in an embodiment of the present invention; wherein, Figure 5 (a) is the structure diagram of the forward network. Figure 5 (b) is the structural diagram of the backward network;

[0018] Figure 6 This is a schematic diagram of the generator network structure in an embodiment of the present invention;

[0019] Figure 7Schematic diagram of the discriminator discrimination method in an embodiment of the present invention; wherein, Figure 7 (a) is a global discrimination diagram. Figure 7 (b) is a schematic diagram of local discrimination;

[0020] Figure 8 Schematic diagram of the discriminator structure in an embodiment of the present invention;

[0021] Figure 9 Schematic diagram of subjective ablation experiment results in a specific example of the present invention; wherein, Figure 9 (a) is the original image. Figure 9 (b) is the result before the introduction of gradient structural similarity loss and conditional information. Figure 9 (c) in the figure shows the result after adding gradient structural similarity loss. Figure 9 (d) in the figure shows the result of simultaneously introducing gradient structural similarity loss and conditional information. DETAILED DESCRIPTION

[0022] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0023] Example 1

[0024] Traditional methods have limited effects in extreme underwater environments (such as extremely low light or poor water quality) and are difficult to fully restore the color, brightness, contrast and details of images. Although deep learning methods have strong image restoration capabilities, they face the problems of large data volume, complex training and insufficient robustness. This paper proposes a comprehensive underwater image clarity method based on exposure correction. Figure 1 As shown in the figure, this method first inputs the underwater non-uniform illumination image and its inverted image into the low-light enhancement network, and obtains the low-light enhanced image and the exposure-reduced image at the same time. The exposure-corrected image is obtained by pyramid image fusion. Then, the cyclic adversarial network CycleGAN is introduced on the basis of exposure correction to perform underwater image enhancement processing. During the training process, the interaction between the forward generator and the backward generator is utilized to learn to convert the underwater image into an effect consistent with the style of the real air image, while ensuring that the generated image is consistent with the real underwater image, thereby maintaining cycle consistency.

[0025] like Figure 1 As shown, the method for comprehensive underwater image sharpening based on exposure correction in this embodiment includes the following steps:

[0026] Step 1. First, the underwater non-uniform illumination image and its inverted image are input into the low-light enhancement network to obtain the low-light enhancement image and the reduced-exposure image. Then, the exposure-corrected image is obtained through pyramid image fusion.

[0027] The present invention proposes an underwater image exposure correction algorithm based on dynamic curve adjustment, which aims to solve the exposure problem of underwater images caused by uneven illumination. The exposure correction flow chart based on dynamic curve adjustment is as follows: Figure 2 shown.

[0028] Depend on Figure 2 It can be seen that the correction algorithm includes two parts: low-light image enhancement based on dynamic curve adjustment and image fusion.

[0029] Step 1.1. First, the underwater non-uniform illumination image is inverted to convert the overexposed area into underexposed area.

[0030] Step 1.2. Simultaneously input the underwater non-uniform illumination image and its inverted image into the depth curve estimation network. The network estimates the best-fit light enhancement curve for the given input image through training. The curve is iteratively applied to map all pixels of the input RGB channels to enhance the underexposed regions of both images, ultimately obtaining a low-light enhanced image and a reduced-exposure image.

[0031] The low-light enhancement algorithm based on dynamic curve adjustment regards low-light enhancement as an image-specific curve estimation task based on a deep network. A deep curve estimation network (DCE-Net) is designed to estimate a set of best-fit light enhancement curves (LE curves) for a given input image. The framework then maps all pixels of the input RGB channels by iteratively applying the curves to obtain the final enhanced image. For the overexposed areas in the original image, first The underwater non-uniform illumination image is inverted, and the overexposed areas in the image are converted to underexposed areas in the inverted image. The corresponding underexposed areas are enhanced through dynamic curve adjustment to correct the overexposed areas of the original image. The underexposed image is obtained by secondary inversion.

[0032] Inspired by the curve adjustment used in photo editing software, we design a curve that can automatically map low-light images to their enhanced versions, namely the light enhancement curve, where the adaptive curve parameters are completely dependent on the input image.

[0033] When designing such a curve, there are three goals:

[0034] 1) Each pixel value of the enhanced image should be within the normalized range of [0, 1] to avoid information loss caused by overflow truncation;

[0035] 2) This curve should be monotonic to preserve the differences (contrast) between adjacent pixels;

[0036] 3) The form of the curve should be as simple as possible and differentiable during gradient backpropagation.

[0037] In order to achieve these three goals, a quadratic curve is designed, which can be expressed as:

[0038] .

[0039] in, represents pixel coordinates, is the given input That is, an enhanced version of the underwater non-uniform illumination image, The parameters [-1, 1] are trainable curve parameters used to adjust the size of the LE curve and control the exposure level. Each pixel is normalized to [0, 1], and all operations are pixel-by-pixel. The LE curve is applied to all three RGB channels, rather than just the illumination channel. This three-channel adjustment better preserves inherent colors and reduces the risk of oversaturation.

[0040] The above light enhancement curve can adjust the dynamic range of the entire image, but it is still a global adjustment because Applied to all pixels. Global mapping tends to over- or under-enhance local areas.

[0041] To solve this problem, Formulated as pixel-level parameters, i.e., for each pixel of a given input image there is a corresponding curve that best fits To adjust its dynamic range, the quadratic curve formula is reformulated as:

[0042] ;

[0043] in, is the number of iterations, used to control the shape of the curve, is a parameter map of the same size as the given image, Represents the brightness of pixel x after iteration n-1 (usually normalized to a value between [0,1]). Represents the brightness of pixel x after the nth iteration. Represents the parameter corresponding to pixel position x, controlling the strength of this curve.

[0044] Here, it is assumed that pixels within a local region have the same intensity, that is, the same adjustment curve, so the adjacency relationship of adjacent pixels in the output image remains unchanged. To learn the mapping between the input image and its best-fit curve parameter mapping, this paper proposes a Deep Curve Estimation Network (DCE-Net).

[0045] The input of the depth curve estimation network DCE-Net is a low-light image, and the output is a set of pixel-by-pixel curve parameter maps corresponding to high-order curves, such as Figure 3 As shown in Figure 2, DCE-Net uses a common CNN with 7 convolutional layers; the first six layers are composed of 32 convolution kernels with a size of 3×3 and a stride of 1, followed by a ReLU activation function.

[0046] Discarding the downsampling layer and batch normalization layer that destroy the relationship between adjacent pixels, the last convolution layer is followed by the Tanh activation function, which produces 3n parameter maps for n iterations, and each iteration requires three curve parameter maps for three channels.

[0047] Here, for example, when n=8, 24 parameter maps are generated after 8 iterations.

[0048] In this embodiment, DCE-Net is based on the total loss Perform model training, where the total loss Loss of spatial consistency , exposure control loss , loss of color constancy and lighting smoothness loss composition.

[0049] Spatial consistency loss The spatial consistency loss encourages spatial consistency of the enhanced image by preserving the differences in adjacent regions between the input image and its enhanced version. The calculation formula is as follows:

[0050] .

[0051] in, is the number of local regions, are four adjacent regions (upper, lower, left, and right) centered on region i, and Expressed as the average intensity value of the local area in the enhanced version and the input image; Indicates the brightness of area i in the enhancement image. Indicates the brightness of region j in the enhanced image. Indicates the brightness of area i in the original image. Indicates the brightness of area j in the original image.

[0052] In this embodiment, the size of the local area is set to 4×4 based on experience.

[0053] To suppress underexposed / overexposed areas, an exposure control loss is designed To control the exposure level. Exposure control loss is measured by the average intensity value of the local area to the well-exposed level. The distance between them.

[0054] Loss of exposure control The calculation formula is as follows:

[0055] .

[0056] in, Indicates the number of non-overlapping local regions of size 16×16, To enhance the average intensity value of the local area in the image, represents the average intensity value of the kth local area in the enhanced image, is the grayscale level in the RGB color space, which is set to 0.6 in this embodiment.

[0057] Following the grayscale world color constancy assumption, i.e., the color in each channel is gray on average over the whole image, a color constancy loss is designed to correct the potential color bias in the enhanced image and establish the relationship between the three adjustment channels.

[0058] Loss of color constancy The calculation formula is as follows:

[0059] ;

[0060] in, and Represents the enhanced image Channel and The average intensity value of the channel, Represents a pair of channels.

[0061] In addition, in order to maintain the unity relationship between adjacent pixels, Add a lighting smoothness loss to the The calculation formula is as follows:

[0062] .

[0063] in, represents the number of iterations, and represent the horizontal and vertical gradients respectively, Represents the curve parameter mapping of the cth channel after the nth iteration.

[0064] Total loss Expressed as: ;

[0065] in, and Represent the weights of color constancy loss and lighting smoothness loss respectively.

[0066] This paper presents a low-light enhancement algorithm based on dynamic curve adjustment. By treating low-light enhancement as an image-specific curve estimation task based on a deep network, a deep curve estimation network (DCE-Net) is designed to estimate a set of best-fit light enhancement curves (LE curves) for a given input image. The DCE-Net network then iteratively applies the curves to map all pixels of the input RGB channels to obtain the final enhanced image.

[0067] Step 1.3. Fuse the low-light enhanced image, the de-exposed image, and the original underwater non-uniform illumination image through pyramid weighted fusion. Select the locally optimally exposed portion from each image and fuse them into a globally balanced optimized image.

[0068] Low-light image enhancement improves the exposure of underexposed pixels in degraded images, while exposure reduction suppresses overexposed pixels. To obtain uniformly illuminated underwater images, high weights are assigned to well-exposed areas in images with varying exposures, while low weights are assigned to underexposed and overexposed areas. This method performs a weighted fusion of the original image (i.e., the underwater image with non-uniform illumination), the low-light enhanced image, and the exposure reduction image to produce the final exposure-corrected image.

[0069] When calculating the weight maps for these three images, contrast, exposure, and saturation are taken into account:

[0070] (1) Contrast weight.

[0071] Contrast is related to the texture and edge information of the image and is very beneficial to maintaining the visual quality and details of the image. Apply a Laplacian filter to the grayscale image and take the absolute value of the filter response to obtain the contrast weight:

[0072] ;

[0073] in, represents a grayscale image, =1, 2 or 3, Represents three images with different exposure rates, Represents the gradient operator, which means taking the first-order derivative of the image. Represents the Laplace operator, which represents the sum of the second-order derivatives of the image; contrast weight after Gaussian smoothing Expressed as .

[0074] (2) Exposure weight.

[0075] The exposure weights balance images with different exposures during the fusion process, allowing details to be preserved in overexposed and underexposed areas. The exposure weight map is calculated by taking into account how close the pixel value is to the median value:

[0076] ;

[0077] in, represents the standard deviation, Represents the normalized pixel value of the cth channel (c∈{R,G,B}) at the pixel position (x, y) of the kth image.

[0078] (3) Saturation weight.

[0079] When an image is close to being overexposed or too dark, it is usually close to being colorless, tending towards pure black or pure white. The standard deviation of the R, G, and B channels represents the saturation of the image:

[0080] .

[0081] in, Respectively represent the pixel values of the red, green, and blue channels at the same pixel (x, y) in the k-th image; Represents the saturation weight at the position (x,y); represents the saturation weight at the position (x,y), Represents the standard deviation of the R, G, and B channels. The standard deviation of the three channel values is used to measure the color saturation.

[0082] Considering the three weight indicators of contrast, exposure and saturation, the fusion weights of different image pixels are is calculated as follows:

[0083] ;

[0084] in, 、 and Set to 1; then the weight Normalization is performed to obtain :

[0085] .

[0086] On this basis, images with different exposure rates are fused, and the calculation formula is as follows:

[0087] ;in, It represents the fusion enhanced image, that is, the exposure corrected image is obtained.

[0088] The entire fusion process, such as Figure 4 As shown. Among them, Figure 4 (a) in the figure is the original image. Figure 4 (b) in the figure is the de-exposure image. Figure 4 (c) in the figure is the low-light enhanced image, which is obtained by Figure 4 It can be seen that the fused image takes into account both bright area protection and dark area enhancement, with richer overall details and better contrast, and the visual quality is significantly better than the original image. The fused image takes into account both bright area protection and dark area enhancement, with richer overall details and better contrast, and the visual quality is significantly better than the original image.

[0089] Step 2. Based on exposure correction, a cyclic adversarial network (CycleGAN) is introduced to achieve underwater image enhancement. During CycleGAN training, the interaction between the forward generator and the backward generator is utilized to learn to transform underwater images into effects consistent with the style of real air images, while ensuring that the generated images are consistent with the real underwater images to maintain cycle consistency.

[0090] The CycleGAN in this embodiment is based on the original CycleGAN and is improved as follows: 1) Real air images are added to the forward generator as conditional information to guide the training of the generator. , so that the generated samples are as consistent as possible with the real clear images without color bias. 2) In the original cyclic adversarial network, a global discriminator was introduced. 3) In terms of loss function: the SSIM loss function was improved, and on the basis of the SSIM loss function, the gradient loss was added, and the color consistency loss function was transferred. Among them, by adding real air images as conditional information to the forward generator, the generator can be guided to better generate clear underwater images without color bias. By introducing a global discriminator to evaluate the authenticity of the entire image, and by weightedly combining the two discriminators, the global structure and local details of the image can be optimized at the same time, thereby generating higher quality and more realistic images. In addition, in the loss calculation of the cyclic adversarial network CycleGAN, the present invention achieves contrast enhancement by improving the SSIM loss function, and corrects color deviation by transferring the color consistency loss function.

[0091] like Figure 5 An underwater image enhancement algorithm based on an improved recurrent adversarial network is presented. This method achieves the color style conversion from underwater images to dehydrated images in air by learning from unpaired training sets.

[0092] The non-paired here means that the input image x and the input image y are not a complete set of paired images.

[0093] Specifically, the improved cycle adversarial network CycleGAN contains two generators 、 And two discriminators 、 , which adds real air images as conditional information to the forward generator to guide the training generator , so that the generated samples are as consistent as possible with the real clear image without color cast. The processing flow of the improved recurrent adversarial network is as follows:

[0094] The encoder E converts the input clear air image _Z into a feature vector Z0 as conditional information, and splices it with the real color-blurred image _X, which is the feature extracted from the exposure-corrected image obtained in step 1, and inputs it into the generator. In the generator Generate a clear image without color cast_Y, through the discriminator To determine whether the generated clear image _Y without color cast is consistent with the real clear image _Y without color cast; the generated clear image _Y without color cast is also input into the generator In the generator Generate a cyclically generated image_X, which is cyclically consistent with the true color-cast blurred image_X.

[0095] At the same time, the real clear image _Y without color cast is input to the generator In the generator Generate a color-biased blurred image_X, through the discriminator To determine whether the generated color cast blurred image _X is consistent with the real color cast blurred image _X; the generated color cast blurred image _X is spliced with the feature vector Z0 and input into the generator In the generator Generate a cyclically generated image_Y, which maintains cyclic consistency with the true clear image_Y without color deviation.

[0096] In order to better realize the to the air dehydrated image domain The present invention adds real air images as conditional information to the forward generator to guide the training generator. , that is, learning a mapping , so that the generated samples As far as possible, the distribution of the real air dehydration image domain Y is consistent; the backward generator remains unchanged, that is, a mapping is learned , so that the generated samples are as close as possible to the real underwater image set The distribution of All images are mapped to An image in .

[0097] At the same time, the present invention also introduces the discriminator To determine whether the generated image is real or fake, the discriminator consists of two parts: the global discriminator and the patch discriminator. The discrimination result is obtained by weighting the two parts. and ,until The output result is close to 0.5, indicating that there is no obvious difference between the generated image and the real image in the eyes of the discriminator, and Nash equilibrium is achieved.

[0098] In this equilibrium state, the generator has learned how to generate indistinguishable images, and the discriminator can no longer effectively distinguish between real and generated images. Domain to The same applies to domain style transfer.

[0099] In addition, the present invention optimizes the global and local details of the image by weighting the outputs of the global discriminator and the PatchGAN discriminator, thereby achieving the goals of color style conversion, contrast enhancement, and detail restoration.

[0100] like Figure 6 The figure shows the structure of the generator network in the improved recurrent adversarial network. In this network, the input consists of two images: an underwater image with blurred color cast and an air image with clear color cast, both of which are 256×256×3 in size.

[0101] First, the two images are each passed through a 7×7 convolutional layer with instance normalization (InstanceNorm) and Leaky ReLU activation, generating a feature map of size 256×256×64. These two feature maps are then concatenated to produce a 256×256×128 feature map. This concatenated feature map then passes through a 3×3 convolutional layer with instance normalization and Leaky ReLU activation for further feature extraction, outputting a feature map of size 128×128×128. The subsequent structure includes multiple 3×3 transposed convolutional layers, also using instance normalization and Leaky ReLU activation, to gradually upsample the feature map size from 128×128×128 to 64×64×256. In the deep layers of the network, nine residual blocks are designed, each containing multiple 3×3 convolutional layers. This aims to optimize feature extraction through residual learning and effectively alleviate the vanishing gradient problem in deep network training. Finally, after a layer of 7×7 convolution and activation function processing, the generator outputs an image of size 256×256×3 as the result of network generation.

[0102] like Figure 7As shown in the figure, the original recurrent adversarial network uses a local discriminator (PatchGAN). By introducing a Markov process, the generative adversarial network's ability to model local image structure is enhanced, enabling the discriminator to better capture image details, especially in the generation of local structures such as textures and edges. A parallel network path is added to the discriminator, accepting the entire image as input. This global discriminator judges the entire input image and ensures the authenticity of the overall structure and content of the generated image. Figure 7 Figures (a) and (b) illustrate two discriminative approaches. The global discriminator assesses the authenticity of the entire image, while the PatchGAN discriminator focuses on local details. The local and global discriminator losses are calculated separately and then weighted together using preset weights to produce the overall discriminator loss function. This weighted combination of the two discriminators simultaneously optimizes both the global structure and local details of the image, resulting in higher-quality, more realistic images.

[0103] like Figure 8 As shown in the figure, the global discriminator judges the entire input image, and the local discriminator PatchGAN divides the input image into multiple local images. For each local image, it outputs a value to judge the consistency of the local image's feature distribution with the real image domain, and outputs a matrix containing the discrimination results of each local image. The process of weighting the global and local discrimination results to obtain the final discrimination result is as follows:

[0104] First, the generated image and the real image are input and passed through four CIL modules; each CIL module includes a convolution ConV, an InstanceNorm normalization layer and a LeakyReLU non-activation function, and the convolution stride is 2; in each CIL module, image dimensionality reduction is first achieved through convolution with a stride of 2, and then the LeakyReLU nonlinear activation function is used to smoothly transmit information less than 0, and finally the InstanceNorm normalization layer is added to achieve standardization; after 4 convolutions with a stride of 2, the global discriminator outputs a single scalar indicating whether the image is real, and the PatchGAN discriminant outputs a 16×16 matrix, representing the authenticity of the corresponding 16×16 area in the image. Finally, the global and local discrimination results are weighted to obtain the final discrimination result.

[0105] In the loss calculation process of the improved recurrent adversarial network, the enhancement effect of underwater images is improved by optimizing multiple loss functions (such as adversarial loss, cycle consistency loss, gradient structure similarity loss and color consistency loss).

[0106] (1) Loss function

[0107] The adversarial loss is mainly used to more accurately extract and fuse the features of underwater distorted images. Representative generated air dehydration images Compared with the real air dehydration image domain The difference in distribution between .

[0108] The process is the adversarial loss (which means continuously reducing the difference between the distribution of the generated image in the pixel space and the distribution of the real image in this space) The definition of is expressed as:

[0109] .

[0110] in, 、 represent the underwater image domain and the air dehydrated image domain respectively, , , Representing a dataset The distribution of Indicates obey Find the mean value in the case of represents the distribution of the air dehydration image y; Indicates y obeys Find the mean value under the condition of ; Represents the output result of the discriminator DY for the real sample y, Representation Generator After inputting the underwater image x and the latent variable z, the dehydrated image is generated; Representation Discriminator Judgment results of the generated dehydrated image.

[0111] The process of generating underwater images Compared with the real underwater image domain The confrontation loss between them is:

[0112] .

[0113] in Representation Discriminator The judgment output of the real underwater image x, Representation Generator The underwater image generated after inputting the dehydrated image y, Representation Discriminator right Generate judgment output of the image.

[0114] (2) Cycle consistency loss.

[0115] The cycle consistency loss is mainly to maintain the consistency of the converted image with the original image content. and 、 and The loss between them plays the role of preserving the image content information.

[0116] Definition of cycle consistency loss Expressed as:

[0117] .

[0118] in, Represents the input underwater image x through the generator Converted into a dehydrated image and passed through the generator Restore to underwater image; represents the process of converting the dehydrated image y into an underwater image; Indicates that the input dehydrated image y first passes through the generator Converted to underwater image and then passed through the generator Restore to dehydrated image.

[0119] (3) Gradient structural similarity loss.

[0120] The CycleGAN network is used to realize the conversion from underwater color-cast blurred images to their corresponding underwater clear images without color cast, which is mainly reflected in the changes in image color and contrast, namely image color correction and contrast enhancement.

[0121] The structural similarity (SSIM) loss function has been proposed. Structural similarity is achieved by comparing the brightness of two images. , contrast and structure These three differences are used to quantify the similarity between images.

[0122] By combining the differences of these three features, we can evaluate the similarity of images more comprehensively. and , and They represent image blocks belonging to two images respectively, and their structural similarity is defined as:

[0123] .

[0124] in, and Respectively and The mean of and Respectively and The variance of express and The covariance of and represents a constant, Indicates structural similarity.

[0125] In order to ensure that the original image content and structural information can be preserved in the process of converting the underwater color-cast blurred image into the underwater clear image without color cast, and only achieve color correction without changing the image structure, the objective function, namely the structural similarity loss, is:

[0126] .

[0127] in, represents the number of pixels in the image, Represents the center pixel of the image block.

[0128] When performing color correction and contrast enhancement on underwater images, it is necessary to maintain the texture structure of the image. However, it is not reasonable to limit the consistency of image brightness and contrast at the same time.

[0129] Therefore, we add gradient difference loss to the structural similarity loss and propose gradient structural similarity loss. By comparing the gradient difference of each pixel and its adjacent pixels, the generator can not only optimize the structural similarity of the image, but also improve the clarity and detail restoration of the image. Defined as:

[0130] ;

[0131] in, 、 represents pixel coordinates, Represents a hyperparameter used to adjust the strength of the gradient difference loss, usually set to 1. The gradient structural similarity loss is defined as:

[0132] .

[0133] in, and Represents the weight parameter of the corresponding loss function, usually set to 1, represents the structural similarity loss, represents the gradient difference loss.

[0134] (4) Following the grayscale world color constancy assumption, i.e., the color in each channel is gray on average over the entire image, a color constancy loss is designed to correct the potential color bias in the generated image and establish the relationship between the three adjustment channels.

[0135] Loss of color consistency Expressed as:

[0136] .

[0137] in, and Represent the generated images Channel and The average intensity value of the channel, represents a pair of channels, , .

[0138] Total loss function It is a weighted combination of adversarial loss, cycle consistency loss, gradient structure similarity loss, and color consistency loss, namely:

[0139] .

[0140] in, and Represent the weight parameters of cycle consistency loss and color consistency loss respectively.

[0141] In addition, in order to verify the effectiveness of the method proposed in this invention, the following experiments are given.

[0142] 1. Subjective ablation experiment.

[0143] In order to verify the effectiveness of conditional information and gradient structure similarity loss in underwater image enhancement algorithm, a subjective ablation experiment was designed. The results of the subjective ablation experiment are shown in Figure 2. Figure 9 As shown, Figure 9 (a) in the figure is the original image. It can be seen that before the gradient structure similarity loss and conditional information are introduced, the enhanced result still has color deviation and blur problems, such as Figure 9 As shown in (b) in .

[0144] When the gradient structure similarity loss is added, the algorithm can improve the contrast of the image, because the gradient structure similarity loss can enable the algorithm to maintain the texture structure and improve the contrast, which helps to better clarify the image. However, there is still the problem of color deviation, such as Figure 9 On this basis, after adding conditional information, the problem of image color deviation is improved while maintaining clarity, and the visual effect of the image is improved, as shown in (c). Figure 9 This result verifies the effectiveness of conditional information and gradient structure similarity loss in improving the performance of underwater enhancement algorithms.

[0145] 2. Objective ablation experiment.

[0146] In order to objectively evaluate the effectiveness of conditional information and gradient structure similarity loss in underwater image enhancement algorithms, objective experiments were conducted using three statistical indicators: image average gradient, standard deviation, and information entropy.

[0147] As shown in Table 1, the effects of image enhancement are compared and evaluated using the above performance indicators.

[0148] Table 1 Objective evaluation of image enhancement effect

[0149]

[0150] Table 1 uses three images as an example. The objective evaluation metrics listed are evaluated based on three different underwater images. The three rows of values for each metric correspond to the calculation results for the first, second, and third images, respectively. As can be seen from Table 1, before the introduction of the gradient structural similarity loss and conditional information, the enhanced results still suffer from color deviation and blurring, resulting in minimal improvement in image color information and contrast, and therefore minimal improvement in the evaluation metrics. However, with the addition of the gradient structural similarity loss, the algorithm improves image contrast, resulting in a significant increase in the average gradient and information entropy. Furthermore, the addition of conditional information improves image color deviation while maintaining clarity, resulting in slight fluctuations in the average gradient and information entropy, and a significant improvement in the standard deviation. Therefore, the addition of the gradient structural similarity loss and conditional information can improve the performance of underwater image enhancement. Experimental results demonstrate that the proposed method outperforms traditional methods and existing deep learning methods in both subjective and objective tests, demonstrating superior performance in detail restoration and quality improvement in underwater images. This effectively addresses the challenges of underwater image enhancement and validates the effectiveness of the proposed method.

[0151] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above-mentioned embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with this field under the guidance of this specification fall within the substantive scope of this specification and should be protected by the present invention.

Claims

1. A method for comprehensive underwater image sharpening based on exposure correction, characterized in that: The steps include: Step 1. First, the underwater non-uniform illumination image and its inverted image are input into the low-light enhancement network to obtain the low-light enhancement image and the reduced-exposure image. Then, the exposure-corrected image is obtained through pyramid image fusion. The low-light enhancement network is a low-light enhancement algorithm based on the deep curve estimation network. A deep curve estimation network is designed to estimate a set of best-fit light enhancement curves for a given input image. The deep curve estimation network then iteratively applies the curve to map all pixels of the RGB channels of the input image to obtain the final enhanced image; Step 2. Based on exposure correction, a cyclic adversarial network (CycleGAN) is introduced to achieve underwater image enhancement. During CycleGAN training, the interaction between the forward generator and the backward generator is utilized to learn to transform underwater images into effects consistent with the style of real air images. At the same time, the generated images are consistent with the real underwater images to maintain cycle consistency. The recurrent adversarial network consists of two generators G X2Y , G Y2X And two discriminators D X 、D Y , which adds real air images as conditional information to the forward generator to guide the training generator G X2Y , so that the generated samples are as consistent as possible with the real clear images without color cast. The processing flow of the improved recurrent adversarial network is as follows: The encoder E converts the input clear image of air without color cast into a feature vector Z0 as conditional information, and splices it with the real color cast blurred image _X, which is the feature extracted from the exposure correction image obtained in step 1, and inputs it into the generator G. X2Y In the generator G X2Y Generate a clear image without color cast_Y, through the discriminator D Y To determine whether the generated clear image _Y without color cast is consistent with the real clear image _Y without color cast; the generated clear image _Y without color cast is also input to the generator G Y2X In the generator G Y2X Generate a cyclic generated image_X, which is cyclically consistent with the true color-biased blurred image_X; At the same time, the true color-free clear image _Y is input to the generator G Y2X In the generator G Y2X Generate a color-blurred image_X, and pass it through the discriminator D X To determine whether the generated color cast blurred image _X is consistent with the real color cast blurred image _X; the generated color cast blurred image _X is spliced with the feature vector Z0 and input into the generator G X2Y In the generator G X2Y Generate a cyclically generated image_Y, wherein the cyclically generated image_Y maintains cyclic consistency with the true clear image_Y without color deviation; The recurrent adversarial network adopts a local discriminator and introduces a Markov process to enhance the generative adversarial network's ability to model the local structure of the image. In the discriminator, a parallel network path is added to accept the entire image as input, namely the global discriminator, which judges the entire input image to ensure the authenticity of the overall structure and content of the generated image. The local discriminator loss and the global discriminator loss are calculated separately, and the two losses are weightedly fused according to the preset weight parameters to obtain the loss function of the total discriminator. By weighted combination of the above two discriminators, the global structure and local details of the image are optimized at the same time.

2. The method for comprehensive underwater image sharpening based on exposure correction according to claim 1, characterized in that: The step 1 is specifically as follows: Step 1.

1. First, the underwater non-uniform illumination image is inverted to convert the overexposed area into underexposed area; Step 1.

2. The underwater non-uniform illumination image and its inverted image are simultaneously fed into the depth curve estimation network. The network estimates the best-fit light enhancement curve for the given input image through training. The curve is then iteratively applied to map all pixels in the input RGB channels to enhance the underexposed regions of both images, ultimately yielding a low-light enhanced image and a reduced-exposure image. Step 1.

3. Fuse the low-light enhanced image, the de-exposure image, and the original underwater non-uniform illumination image through pyramid weight fusion. Select the locally best-exposed portion from each image and fuse them into an optimized image with globally balanced exposure.

3. The method for comprehensive underwater image sharpening based on exposure correction according to claim 2, characterized in that: In step 1.2, the formula for the best fitting light enhancement curve is expressed as follows: THE n (x)=THE n-1 (x)+A n ·THE n-1 (x)·(1-LE n-1 (x)); Where x represents the pixel coordinate, n is the number of iterations used to control the shape of the curve, A is a parameter map of the same size as the given non-uniformly illuminated underwater image, and E n-1 (x) represents the brightness of pixel x after the n-1th iteration, LE n (x) represents the brightness of pixel x after the nth iteration, A n Represents the parameter corresponding to pixel position x, which is used to control the strength of the curve.

4. The method for comprehensive underwater image sharpening based on exposure correction according to claim 2, characterized in that: In step 1.2, the input of the deep curve estimation network DCE-Net is a low-light image, and the output is a set of pixel-by-pixel curve parameter maps corresponding to high-order curves; DCE-Net uses a common CNN with 7 convolutional layers; DCE-Net discards the downsampling layer and batch normalization layer that destroy the relationship between adjacent pixels. The last convolutional layer is followed by a Tanh activation function, which generates 3n parameter maps for n iterations, and each iteration requires three curve parameter maps for three channels.

5. The method for comprehensive underwater image sharpening based on exposure correction according to claim 4, characterized in that: In step 1.2, DCE-Net is based on the total loss L total Perform model training, where the total loss L total By spatial consistency loss L spa , exposure control loss L exp , color constancy loss L col and the lighting smoothness loss L tvA composition; Spatial consistency loss L spa The calculation formula is as follows: Where K is the number of local regions, Ω(i) is the four adjacent regions centered on region i, and Y and I represent the average intensity values of the local regions in the enhanced version and the input image; Y i Indicates the brightness of area i in the enhanced image; Y j Indicates the brightness of area j in the enhanced image; I i Indicates the brightness of area i in the original image; I j Represents the brightness of area j in the original image; Exposure control loss L exp The calculation formula is as follows: Where M represents the number of non-overlapping local regions of size 16×16, Y is the average intensity value of the local region in the enhanced image, and E is the grayscale level in the RGB color space; Y k Represents the average intensity value of the kth local area in the enhanced image; Color constancy loss L col The calculation formula is as follows: Among them, J P and J Q Respectively represent the average intensity values of the P channel and Q channel in the enhanced image, (P, Q) represents a pair of channels; Lighting smoothness loss L tvA The calculation formula is as follows: Where N represents the number of iterations, and represent the horizontal and vertical gradients respectively, Represents the curve parameter mapping of the cth channel after the nth iteration; Total loss L total Expressed as: L total =L spa +L exp +ω col L col +ω tvA L tvA ; Among them, ω col and ω tvA Represent the weights of color constancy loss and lighting smoothness loss respectively.

6. The method for comprehensive underwater image sharpening based on exposure correction according to claim 2, characterized in that: The step 1.3 is specifically as follows: Apply a Laplacian filter to the grayscale image and take the absolute value of the filtered response to obtain the contrast weights: Among them, gray(I k (x,y)) represents a grayscale image, when k = 1, 2 or 3, I k (x,y) represents three images with different exposure rates. Represents the gradient operator, which means taking the first-order derivative of the image; Represents the Laplace operator, which represents the sum of the second-order derivatives of the image; the contrast weight C after Gaussian smoothing x,y,k Represented as C' x,y,k ; The exposure weight map is calculated as follows: Where σ represents the standard deviation, Represents the normalized pixel value of the cth channel (c∈{R,G,B}) at the pixel position (x, y) of the kth image; calculates the saturation weight, and the standard deviation of the R, G, and B channels represents the saturation of the image: S x.y,k =std(R x,y,k ,G x,y,k ,B x,y,k ); Among them, R x,y,k ,G x,y,k ,B x,y,k They represent the pixel values of the red, green, and blue channels at the same pixel (x, y) in the k-th image; S x.y,k Represents the saturation weight at the position (x, y), and std() represents the standard deviation of the R, G, and B channels; Considering the three weight indicators of contrast, exposure and saturation, the fusion weight W of different image pixels is x,y,k is calculated as follows: Among them, w c 、w e and w s Set to 1; then the weight W x,y,k Normalization is performed to obtain On this basis, images with different exposure rates are fused, and the calculation formula is as follows: Among them, I f (x,y) represents the fused enhanced image, that is, the exposure corrected image.

7. The method for comprehensive underwater image sharpening based on exposure correction according to claim 1, characterized in that: The global discriminator judges the entire input image, and the local discriminator PatchGAN divides the input image into multiple local images. For each local image, it outputs a value to judge the consistency of the local image's feature distribution with the real image domain, and outputs a matrix containing the discrimination results of each local image. The process of weighting the global and local discrimination results to obtain the final discrimination result is as follows: First, the generated image and the real image are input and passed through four CIL modules; each CIL module includes a convolution ConV, an InstanceNorm normalization layer and a LeakyReLU non-activation function, and the convolution stride is 2; in each CIL module, image dimensionality reduction is first achieved through convolution with a stride of 2, and then the LeakyReLU nonlinear activation function is used to smoothly transmit information less than 0, and finally the InstanceNorm normalization layer is added to achieve standardization; after 4 convolutions with a stride of 2, the global discriminator outputs a single scalar indicating whether the image is real, and the PatchGAN discriminant outputs a 16×16 matrix, representing the authenticity of the corresponding 16×16 area in the image. Finally, the global and local discrimination results are weighted to obtain the final discrimination result.

8. The method for comprehensive underwater image sharpening based on exposure correction according to claim 1, characterized in that: The cyclic adversarial network adopts the total loss function L(G X2Y ,G Y2X ,D X ,D Y ), train the model; Define the adversarial loss L GAN (G X2Y ,D Y ,X,Y) are as follows: Among them, X and Y represent the underwater image domain and the air dehydrated image domain respectively, x∈X, y∈Y, P data (x) represents the distribution of the data set X, Indicates that x obeys P data (x) to find the mean, P data (y) represents the distribution of air dehydration image y; Indicates that y obeys P data (y) to find the mean; D Y (y) represents the discriminator D Y The output result of the real sample y, G X2Y (x|z) represents the generator G X2Y (x|z) is the dehydrated image generated after inputting the underwater image x and the latent variable z; D Y (G X2Y (x|z)) represents the discriminator D Y judgment results of the generated dehydrated image; Define the adversarial loss L GAN (G Y2X ,D X ,Y,X) is; Among them D X (x) represents the discriminator D X The judgment output of the real underwater image x, G Y2X (y) represents the generator G Y2X The underwater image generated after inputting the dehydrated image y, D X (G Y2X (y)) represents the discriminator D X To G Y2X (y) generating judgment output of the image; Definition of cycle consistency loss L cyc (G X2Y ,G Y2X ) is expressed as: Among them, G Y2X (G X2Y (x|z)) represents the input underwater image x through the generator G X2Y Converted into a dehydrated image and then passed through the generator G Y2X Restore to underwater image; G Y2X (y)|z represents the process of converting the dehydrated image y into an underwater image; G X2Y (G Y2X (y)|z) indicates that the input dehydrated image y first passes through the generator G Y2X Converted into underwater image and then passed through generator G X2Y Restore to dehydrated image; The gradient structural similarity loss is defined as: L GSSIM =λ ssim L SSIM (x,G(x|z))+λ gdl L GDL (x,G(x|z)); Among them, λ ssim and λ gdl Represents the weight parameter of the corresponding loss function; L SSIM (x,G(x|z)) represents the structural similarity loss between the real graph and the generated graph, L GDL (x,G(x|z)) represents the gradient difference loss; G(x|z) represents the generated image output by the generator of the input underwater image x; Color consistency loss L col Expressed as: Among them, G(x|z) P and G(x|z) Q Represent the average intensity values of the P channel and Q channel in the generated image respectively, (P,Q) represents a pair of channels, (P,Q)∈ε,ε={(R,G),(R,B),(G,B)}; The total loss function L(G X2Y ,G Y2X ,D X ,D Y ) is a weighted combination of adversarial loss, cycle consistency loss, gradient structure similarity loss, and color consistency loss, namely: Among them, λ cyc and λ col Represent the weight parameters of cycle consistency loss and color consistency loss respectively.

Citation Information

Patent Citations

  • Low-illumination image judgment and image enhancement method for drilling operation site

    CN115984535A

  • Low-illumination image enhancement method and system based on improved generative adversarial network

    CN117593238A