Underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning
Through Laplace pyramid decomposition and depth separation convolution combined with frequency domain enhancement technology, the problems of underwater image enhancement technology's robustness and computational complexity in complex underwater environments are solved, and efficient and real-time image quality improvement is achieved.
Patent Information
- Application Number
- CN202510429396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing underwater image enhancement technology is not robust enough in complex underwater environments, has high computational complexity, relies on large-scale data and is difficult to real-time, and cannot be effectively applied in scenarios where resource constraints or high efficiency requirements.
The Laplace pyramid decomposition image is used to form a low-frequency residual layer and a high-frequency detail layer, combined with the depth-separable convolution and frequency domain enhancement feature module, and the network is trained through comprehensive optimization of the loss function, and progressive reconstruction and band attention weighting are carried out to achieve multi-scale enhancement of the image.
Significantly reduce the computational complexity, improve image visual quality and robustness, adapt to complex underwater scenes, meet real-time processing needs, and have efficient image detail retention and visual consistency.
Smart Images

Figure CN120339148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to underwater image processing technology in the field of underwater autonomous perception, and particularly to an underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning. Background Art
[0002] In the fields of ocean exploration, underwater archaeology, and underwater ecological monitoring, in order to obtain high-quality underwater images for subsequent target recognition, feature extraction, and environmental assessment, researchers have long been committed to developing efficient and robust underwater image enhancement and restoration methods. However, the special optical properties of the underwater environment, complex light propagation paths, and variable water medium conditions, including suspended particles, plankton, and diverse light absorption and scattering, result in the original underwater images obtained by traditional imaging systems often presenting problems such as low contrast, poor clarity, severe color deviation, and lack of detailed texture. To address this challenge, academia and industry have proposed various technical routes, including image restoration methods based on physical models, image enhancement models based on deep learning, and image optimization schemes based on multi-image fusion strategies.
[0003] Traditional methods based on physical models are based on classical underwater imaging models, which describe the degradation process of underwater images as a functional relationship affected by factors such as absorption and scattering when light propagates in the underwater medium. The advantage of physical model methods lies in their clear theoretical basis and strong interpretability, and they can achieve image restoration under non-reference conditions to a certain extent. However, these methods often require accurate modeling of imaging parameters, and the parameter measurement in the underwater environment is complex and unstable, resulting in insufficient robustness and generalization of these methods in real and complex underwater scenes.
[0004] In recent years, with the development of deep learning, researchers have gradually applied deep learning structures such as convolutional neural networks (CNNs), generative adversarial networks (GANs), and Transformers to underwater image enhancement. Deep learning methods no longer simply rely on accurate physical model assumptions, but learn the mapping relationship from degraded images to clear images in a data-driven manner.
[0005] CycleGAN is an unsupervised generative adversarial network structure that realizes domain conversion of images by learning the cyclic mapping relationship between two domains. In underwater image enhancement, CycleGAN-based methods utilize non-reference data and enhance the contrast and color of images by introducing cyclic consistency loss. However, CycleGAN has a strong dependence on the distribution of training data in practical applications, and at the same time, due to the complex network structure, the model inference speed and resource consumption are relatively high, which limits its application in scenarios with high real-time requirements.
[0006] UIE - UnFold and other depth models attempt to combine the underwater imaging mechanism with the depth network structure to achieve more refined enhancement through the adaptive learning of network weights and parameters. Although these methods have achieved better objective evaluation metrics and visual effects under experimental conditions, their training processes often require a large amount of training data with high - quality annotations. At the same time, the parameter scale of depth models is large, and the computational cost cannot be ignored, which is not conducive to deployment in resource - limited embedded systems or real - time application scenarios.
[0007] The strategy based on multi - image fusion improves the quality of underwater images by fusing multi - scale, multi - type, or multi - perspective image information. For example, WaterNet introduces multi - input image fusion, combines different color correction or enhancement results, and realizes the final high - quality output through a weight optimization strategy. Such methods can retain more details and improve colors to a certain extent, but their architectures are usually complex and the computational processes are cumbersome, which is not conducive to application in resource - constrained or high - timeliness - required occasions. In addition, the fusion strategy depends heavily on the quality of input images and the pre - processing link. Once the input image distributions vary greatly or there is strong noise interference, the fusion strategy may not be able to effectively improve the overall image quality.
[0008] In summary, the existing underwater image enhancement and restoration technologies still have the following problems to be solved urgently in practical applications:
[0009] (1) Insufficient trade - off between interpretability and robustness: Methods based on physical models have strong interpretability, but are not stable and robust enough when dealing with complex underwater scenes;
[0010] (2) Limitations in the complexity and applicability of fusion strategies: Although multi - image fusion methods have achieved improved effects in some aspects, their complex architectures and long processing chains are not conducive to practical engineering deployment and applications requiring fast response;
[0011] (3) Dependence on large - scale data and high computational cost: Although deep learning methods have outstanding enhancement effects, they have high requirements for the quantity and quality of training data, and the model inference cost is large, making it difficult to be lightweight and real - time. Summary of the Invention
[0012] The purpose of the present invention is to overcome the above - mentioned defects existing in the prior art, and propose an underwater low - quality image enhancement method based on Laplacian pyramid and contrast learning, which can be used for clarifying optical image data in the operation tasks of underwater autonomous vehicles and improve the image quality.
[0013] The technical solution of the present invention is: An underwater low - quality image enhancement method based on Laplacian pyramid and contrast learning, which includes the following steps:
[0014] S1. Image input and multi-scale decomposition: The input image is decomposed into a low-frequency residual layer and several high-frequency detail layers through Laplacian pyramid decomposition;
[0015] S2. The low-frequency residual layer is input into the global illumination and color correction sub-module to obtain an enhanced low-frequency residual layer;
[0016] S3. The enhanced low-frequency residual layer and high-frequency detail layers of multiple scales are input into the frequency domain enhanced feature module to obtain the output of high-frequency detail layers with multi-scale enhancement;
[0017] S4. The output of the enhanced low-frequency layer and the outputs of the high-frequency enhanced layers of each scale are progressively reconstructed according to the inverse process of the Laplacian pyramid;
[0018] S5. Output result and application.
[0019] In the present invention, the specific implementation process of step S1 is as follows:
[0020] Underwater optical image sequences are collected in real time, and the latest frame of image I(w, h, 3) is read in real time;
[0021] The image I(w, h, 3) is input into the Laplacian pyramid enhancement network model and decomposed into frequency components of multiple scales, including a low-frequency residual layer I1 and several high-frequency detail layers I hi , where i = 1, 2,..., n.
[0022] The Laplacian pyramid enhancement network model is trained by means of supervised learning, and a comprehensive optimization loss function is used to guide the parameter optimization of the Laplacian pyramid enhancement network model. The equation of the comprehensive optimization loss function is:
[0023] L = λ1·L 结构一致性 + λ2·L 颜色 + λ3·L 感知 + λ4·L 对比 ,
[0024] where λ1 represents the structure consistency loss weight; λ2 represents the color consistency loss weight; λ3 represents the perceptual loss weight; λ4 represents the contrast loss weight;
[0025] The function of the structure consistency loss L 结构一致性 is:
[0026] L 结构一致性 = 1 - pytorch_ssim.ssim(I 生成 , I 参考 ),
[0027] where I 生成 represents the output generated image; I 参考Indicates a reference image;
[0028] The perceptual loss L 感知 The function of is:
[0029]
[0030] Where, φ i Indicates the feature representation extracted by the pre-trained neural network for the input image on the i-th layer;
[0031] The color consistency loss L 颜色 The function of is:
[0032]
[0033] Where, S represents the set composed of a group of subsets formed after the image is partitioned by regions; s represents the index of a specific subset in the set S; N s Indicates the number of traversals required for the s-th subset; p represents the index used to index a single sampling unit in the s-th subset; x is In the refers to the horizontal coordinate of the image, used to calculate the horizontal gradient; y is In the refers to the vertical coordinate of the image, used to calculate the vertical gradient;
[0034] The contrastive loss L 对比 The function of is:
[0035]
[0036] The network structure of the global illumination and color correction sub-module successively includes a feature extraction layer, a residual block group, and an output layer;
[0037] The feature extraction layer successively includes a first DSConv layer, an InstanceNorm2d layer, a first LeakyReLU activation function layer, a second DSConv layer, and a second LeakyReLU activation function layer;
[0038] The first DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution, expanding the number of channels of the input low-frequency residual layer I1 from 3 to 16 for preliminary feature extraction; the InstanceNorm2d layer is used to perform instance normalization on the feature map; the first LeakyReLU activation function layer is used to introduce non-linearity; the second DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to expand the number of channels from 16 to 64 for further feature extraction;
[0039] The residual block group includes five residual modules, each residual module contains two DSConv layers and a skip connection; the DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution; the skip connection directly adds the input of the residual module to the outputs of the two DSConv layers;
[0040] The output layer sequentially includes a third DSConv layer, a third LeakyReLU activation function layer, and a fourth DSConv layer;
[0041] The third DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to compress the number of channels from 64 back to 16; the fourth DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to compress the number of channels from 16 back to 3; after the output layer, a tanh activation function is connected to map the output value to the range of [-1, 1].
[0042] The data processing flow of the global illumination and color correction sub-module is as follows:
[0043] S2.1. Through the feature extraction layer, primary features are extracted and transformed from the original three-channel RGB image through depthwise separable convolution, instance normalization, and LeakyReLU activation function;
[0044] S2.2. The primary features obtained in step S2.1 are passed through five residual modules to further extract and refine features;
[0045] S2.3. The output layer remaps the features back to the same channel dimension as the input to achieve the reduction process from the low-level feature space back to the original image space, and maps the output value to the range of [-1, 1] through the tanh activation function to obtain the enhanced low-frequency residual layer I 1_enhanced 。
[0046] The specific implementation process of step S3 is as follows:
[0047] S3.1. Concatenate the multi-scale high-frequency detail layer I hi , i = 1, 2, …, n, the low-frequency residual layer, and the enhanced low-frequency residual layer I 1_enhanced to obtain the input X;
[0048] S3.2. Generate the initial mask feature map Mask0 to obtain the adaptive mask Mask for weighted fusion;
[0049] S3.3. Determine the number of high-frequency levels to be processed according to the design parameter num_high;
[0050] S3.4. Perform the following operations on each high-frequency layer:
[0051] S3.4.1. Spatial Alignment and Interpolation: Upsample the generated Mask to make the mask size consistent with the high-frequency feature layer to be processed;
[0052] S3.4.2. Frequency-Domain Enhanced Feature Extraction: The frequency-domain enhanced feature module performs frequency-domain analysis and enhancement on the current high-frequency layer image features to obtain the enhanced high-frequency detail features enhanced;
[0053] S3.4.3. Mask Fusion Strategy: Combine the enhanced features enhanced with the weighting strategy guided by the Mask to obtain the high-frequency result result_highfreq;
[0054] S3.4.4. Post-Processing and Feature Fine-Tuning: Input the fused high-frequency result result_highfreq into the subsequent processing unit,
[0055] S3.5. Output Hierarchical Aggregation and Reconstruction: After completing the enhancement processing of all high-frequency levels, store the processing results of each high-frequency level in the list pyr_result in the corresponding order, and at the same time add the low-frequency enhancement result I 1_enhanced to this result list.
[0056] The specific implementation process of step S3.2 is as follows:
[0057] S3.2.1. Perform a depthwise separable convolution DSConv operation on the input X to expand the number of channels of the input X from 9 to 64. This DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution;
[0058] S3.2.2. Introduce non-linearity through the LeakyReLU activation function layer;
[0059] S3.2.3. Integrate the multi-scale feature extraction module and several residual blocks to obtain the initial mask feature map Mask0;
[0060] The multi-scale feature extraction module contains four parallel branches, which perform feature extraction using convolutional kernels of 1×1, 3×3, 5×5, and 7×7 respectively, and splice the outputs of the four branches in the channel dimension, and finally perform feature fusion through a 1×1 convolution;
[0061] S3.2.4. Apply the Sigmoid function to Mask0 to map it to a continuous range between 0 and 1 to obtain the adaptive mask Mask that can be used for weighted fusion.
[0062] In step S3.4.2, the frequency-domain enhanced feature module includes a frequency band attention module, a global context encoding module, and a frequency mask generation function;
[0063] The specific processing flow for frequency domain analysis and enhancement using the frequency domain enhancement feature module is as follows:
[0064] S3.4.2.1. Convert the current high-frequency layer image feature x from the spatial domain to the frequency domain through the dct_2d(x) function implemented based on the torch_dct library to obtain the DCT coefficients;
[0065] S3.4.2.2. Generate frequency masks using the self.frequency_band_masks(h,w) function: This function calculates the frequency position matrix based on pixel coordinates and divides the frequency domain into several frequency bands according to the distance, with each frequency band corresponding to a mask;
[0066] S3.4.2.3. Multiply the mask of each frequency band by the DCT coefficients to obtain the DCT coefficients and their amplitudes for each frequency band;
[0067] S3.4.2.4. Use the frequency band attention module to weight the information of different frequency bands:
[0068] The frequency band attention module includes two 1×1 convolutional layers, a ReLU activation function, and a Sigmoid activation function. This module concatenates the amplitude spectra of each frequency band along the channel dimension as the input and outputs the weights of each frequency band;
[0069] S3.4.2.5. Multiply the DCT coefficients of each frequency band by the corresponding frequency band weights to obtain the enhanced frequency band DCT coefficients;
[0070] S3.4.2.6. Convert the enhanced frequency band DCT coefficients, which is the sum of all frequency band coefficients, back to the spatial domain through the idct_2d() function implemented based on the torch_dct library to obtain the enhanced high-frequency detail feature enhanced.
[0071] The specific implementation process of step S4 is as follows:
[0072] S4.1. Starting from the lowest resolution layer, upsample it and align it to the size of the higher resolution feature layer through interpolation;
[0073] S4.2. Add the alignment result of step S4.1 to the high-frequency feature of this layer after frequency domain enhancement and masking processing, so as to stack more details and local textures;
[0074] Repeat steps S4.1 and S4.2, gradually fusing from the global structure at the bottom of the pyramid to the high-frequency features at the top layer, and finally obtaining a high-resolution reconstructed image with complete details.
[0075] In step S4.1, an upsampling template is constructed through torch.cat and view operations, and convolution operations are performed using a Gaussian convolution function to achieve upsampling.
[0076] The beneficial effects of the present invention are as follows:
[0077] (1) By introducing depthwise separable convolution and a progressive processing strategy, the computational complexity of the model is significantly reduced, meeting the requirements of real-time processing;
[0078] (2) The frequency domain enhanced feature module combines a frequency band attention mechanism, effectively optimizing color consistency and detail restoration, and improving visual quality.
[0079] (3) The robustness of the present invention in complex underwater scenes is enhanced, and it can adapt to the domain differences between synthetic and real data, with stronger generalization ability.
[0080] The method proposed in this application is superior to the prior art in terms of computational efficiency, image detail retention, and visual consistency, and has high practical value and promotion potential. Description of the Drawings
[0081] Figure 1 is a flowchart of the method described in the present invention;
[0082] Figure 2 is an image before being processed by the method described in this application;
[0083] Figure 3 is an image after being processed by the method described in this application. Detailed Embodiments
[0084] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be made in conjunction with the accompanying drawings.
[0085] Specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0086] This application discloses an underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning, and its method flow is as Figure 1 shown. The method specifically includes the following steps.
[0087] Step 1: Image input and multi-scale decomposition.
[0088] The input image is decomposed into frequency components of multiple scales through Laplacian pyramid decomposition, including a low-frequency residual layer and several high-frequency detail layers.
[0089] The low-frequency residual layer is used to characterize the overall brightness and color distribution information of the image, and the high-frequency detail layer is used to characterize the texture and edge detail information of the image. Through this multi-scale decomposition, it helps the present application to enhance the global features and local features respectively in a targeted manner.
[0090] In this embodiment, taking the input of a single image as an example, the specific implementation process of the present solution is introduced in detail. The specific implementation process of image input and multi-scale decomposition is as described below.
[0091] First, collect the underwater optical image sequence in real time.
[0092] In this embodiment, dual-thread real-time data acquisition is used, where:
[0093] (1) Thread 0: Construct a global empty queue Q for temporarily storing the collected image data, start the camera to obtain real-time data, and temporarily store it in the queue. If the queue reaches the maximum length, delete elements from the head; maintain the latest image data stored in real time;
[0094] (2) Thread 1: The main thread, read the latest frame of image I(w, h, 3) from the queue in real time;
[0095] Second, multi-scale Laplacian pyramid decomposition:
[0096] Input the image I(w, h, 3) into the Laplacian pyramid enhancement network model. Through Laplacian pyramid decomposition, the image I(w, h, 3) is decomposed into frequency components of multiple scales, including a low-frequency residual layer I1 and several high-frequency detail layers I hi , i = 1, 2,..., n.
[0097] In this embodiment, the Laplacian pyramid enhancement network model is trained in a supervised learning manner. The training data includes a series of underwater degraded images and their corresponding clear reference images, and a comprehensive optimization loss function is used to guide the parameter optimization of the generation network.
[0098] Specifically, this embodiment proposes a comprehensive optimization loss function for the underwater image enhancement task. This loss function combines loss terms in multiple dimensions and aims to improve the quality of the generated image and its consistency with the reference image. Its definition is as follows:
[0099] L = λ1·L 结构一致性 + λ2·L 颜色 + λ3·L 感知 + λ4·L 对比 ,
[0100] where λ1 represents the structural consistency loss weight, which takes the value of 1 in this embodiment; L结构一致性 denotes the structural consistency loss; λ2 denotes the weight of the color consistency loss, which is set to 1 in this embodiment; L 颜色 denotes the color consistency loss; λ3 denotes the weight of the perceptual loss, which is set to 10 in this embodiment; L 感知 denotes the perceptual loss; λ4 denotes the weight of the contrast loss, which is set to 10 in this embodiment; L 对比 denotes the contrast loss.
[0101] The structural consistency loss L 结构一致性 measures the structural similarity between the generated image and the reference image based on the Structural Similarity Index (SSIM). SSIM takes into account the brightness, contrast, and structural information of the image and is a commonly used image quality evaluation metric. In this embodiment, the ssim function in the pytorch_ssim library is used to calculate the SSIM value and convert it into a loss:
[0102] L 结构一致性 = 1 - pytorch_ssim.ssim(I 生成 , I 参考 ).
[0103] where I 生成 denotes the output generated image; I 参考 denotes the reference image.
[0104] The perceptual loss L 感知 is calculated based on the differences between the feature maps extracted by a pre-trained VGG19 network. In this embodiment, a pre-trained VGG19 model is loaded, and the slice1, slice2, and slice3 layers are retained, corresponding to the shallow, middle, and deep features of the network respectively. The perceptual loss uses the L1 loss to calculate the differences between these feature maps and performs a weighted sum according to the weight coefficients of different levels to obtain the final loss.
[0105]
[0106] where φ i denotes the feature representation of the input image extracted by a pre-trained neural network at the i-th layer. In the present invention, a pre-trained VGG19 model is selected for feature extraction.
[0107] The color consistency loss L 颜色 aims to constrain the similarity in color distribution between the generated image and the reference image, thereby enhancing the color naturalness and realism of the enhanced image. The implementation idea of this loss combines multi-scale local color statistical features and color gradient information.
[0108] Specifically, this loss first divides the image into local patches at multiple scales, calculates the mean, variance, skewness, and kurtosis of the three RGB channels within each patch, and then calculates the L1 loss of the generated image and the reference image in these statistics to constrain the consistency of the multi-scale local color distribution. At the same time, the horizontal and vertical gradients of the generated image and the reference image are calculated, and the L1 loss between the two gradients is calculated to constrain the consistency of the color gradient change. Finally, the above two parts of the loss are added together to obtain the color consistency loss.
[0109]
[0110] Among them, S represents the set composed of a group of subsets formed after the image is divided into regions; s represents the index of a specific subset (a certain region) in the set S; N s represents the number of traversals required for the s-th subset; p represents the index of a single sampling unit in the s-th subset; x in refers to the horizontal coordinate of the image, used to calculate the horizontal gradient; y in refers to the vertical coordinate of the image, used to calculate the vertical gradient.
[0111] The contrast loss L 对比 The core idea of the function is to shorten the distance between the "generated sample (anchor point)-Positive (positive sample)" in the feature space, while lengthening the distance between the "generated sample (anchor point)-Negative (negative sample)", and a "margin" is used to control the intensity of the contrast loss during the shortening and lengthening processes.
[0112]
[0113] Among them: a i , represents the generated output image of the i-th layer; p i , represents the reference image of the i-th layer; n i represents the unprocessed input image of the i-th layer; ||·||1 represents the L1 norm; w i is the weighting coefficient of the i-th layer; margin is the margin constant in the contrast loss.
[0114] Step 2: Enhancement of the low-frequency residual layer.
[0115] For the low-frequency residual layer, its input is fed into the global illumination and color correction sub-module to obtain the adjusted output of the low-frequency layer.
[0116] The global illumination and color correction sub-module uses depthwise separable convolutions to achieve adaptive adjustment of the overall illumination and color while maintaining the efficiency of the model.
[0117] Input the low-frequency residual layer I1 into the global illumination and color correction sub-module for enhancing brightness and color to obtain the enhanced low-frequency residual layer I. 1_enhanced 。
[0118] Specifically, the network structure of the global illumination and color correction sub-module sequentially includes a feature extraction layer, a residual block group, and an output layer.
[0119] The feature extraction layer sequentially includes a first DSConv layer, an InstanceNorm2d layer, a first LeakyReLU activation function layer, a second DSConv layer, and a second LeakyReLU activation function layer.
[0120] First, through the first DSConv layer, i.e., the depthwise separable convolution layer, expand the number of channels of the input low-frequency residual layer I1 from 3 to 16. This DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution, both with a stride of 1 and a padding of 1, for initially extracting features. The InstanceNorm2d layer is used to perform instance normalization on the feature map to accelerate model convergence and improve generalization ability. The first LeakyReLU activation function layer, with a negative slope set to 0.2, is used to introduce non-linearity and enhance the model's expressive ability. The second DSConv layer expands the number of channels from 16 to 64, also using a 3×3 depth convolution and a 1×1 pointwise convolution, with a stride of 1 and a padding of 1, to further extract features. The second LeakyReLU activation function layer has a negative slope set to 0.2.
[0121] The residual block group located between the feature extraction layer and the output layer includes five residual modules for further extracting and refining features. Each residual module contains two DSConv layers and a skip connection. The DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution, with a stride of 1 and a padding of 1. The skip connection directly adds the input of the residual module to the output of the two DSConv layers, which helps to solve the problem of gradient disappearance during the training of deep networks, enabling the network to learn the residual mapping, thus better retaining the detailed information of the image and further enhancing the expressive ability of the features.
[0122] The output layer sequentially includes a third DSConv layer, a third LeakyReLU activation function layer, and a fourth DSConv layer.
[0123] The third DSConv layer compresses the number of channels from 64 back to 16, using 3×3 depthwise convolution and 1×1 pointwise convolution, with a stride of 1 and padding of 1. The third LeakyReLU activation function layer has a negative slope set to 0.2. The fourth DSConv layer compresses the number of channels from 16 back to 3, using 3×3 depthwise convolution and 1×1 pointwise convolution, with a stride of 1 and padding of 1, so that the number of output channels is the same as that of the original input image. A tanh activation function is connected after the output layer to map the output values to the range [-1, 1].
[0124] The data processing flow of the global illumination and color correction sub-module is as follows:
[0125] First, through the feature extraction layer, primary features are extracted and transformed from the original three-channel RGB image through depthwise separable convolution, instance normalization, and the LeakyReLU activation function.
[0126] Second, the primary features obtained in the first step are passed through five residual modules to further extract and refine features. Each residual module contains two DSConv layers and a skip connection.
[0127] The skip connection directly adds the input of the residual module to the output of the two DSConv layers. This structure can not only refine and enhance features, but more importantly, solve the problem of gradient disappearance during the training of deep networks, enabling the network to learn the residual mapping, thus better retaining the detailed information of the image and further enhancing the feature expression ability. The use of residual modules enables the network to more effectively extract and refine features from the low-frequency residual layers, thereby improving the effect of global illumination and color correction.
[0128] Third, the output layer remaps the features back to the same channel dimension as the input to achieve the reduction process from the low-level feature space back to the original image space, and maps the output values to the range [-1, 1] through the tanh activation function to obtain the enhanced low-frequency residual layer.
[0129] Step 3: Enhancement of the high-frequency detail layer.
[0130] The low-frequency residual layer and high-frequency detail layers of multiple scales are input into the frequency-domain enhanced feature module. This module uses wavelet transform to map the high-frequency details in the time domain to the frequency domain space. In the frequency domain space, a frequency band attention mechanism and a dynamic mask strategy are introduced to weight and screen the information of different frequency bands, thereby highlighting the effective texture information and suppressing noise and useless details. After being processed by the frequency-domain enhanced feature module, the output of the high-frequency detail layer with multi-scale enhancement is obtained, and its texture sharpness and detail restoration are significantly improved. The specific implementation process of this step is described as follows.
[0131] First, concatenate the multi-scale high-frequency detail layers I hi of the image to be enhanced, where i = 1, 2, …, n, the low-frequency residual layer I1, and the enhanced low-frequency residual layer I 1_enhanced to obtain the input X.
[0132] Second, generate the initial mask feature map.
[0133] First, perform a depthwise separable convolution (DSConv) operation on the input X to expand the number of channels of the input X from 9 to 64. The DSConv layer consists of a 3×3 depthwise convolution and a 1×1 pointwise convolution, both with a stride of 1 and a padding of 1.
[0134] Then, introduce non-linearity through the LeakyReLU activation function layer, and the negative slope of the LeakyReLU activation function layer is set to 0.2.
[0135] Next, fuse the multi-scale feature extraction module and several residual blocks to obtain the initial mask feature map Mask0.
[0136] The multi-scale feature extraction module contains four parallel branches, which perform feature extraction using 1×1, 3×3, 5×5, and 7×7 convolutional kernels respectively, and concatenate the outputs of the four branches in the channel dimension, and finally perform feature fusion through a 1×1 convolution.
[0137] Finally, apply the Sigmoid function to Mask0 to map it to a continuous range between 0 and 1 to obtain the adaptive mask Mask that can be used for weighted fusion.
[0138] Third, determine the number of high-frequency levels n to be processed according to the design parameter num_high. For example, if num_high = 3, then the details of three high-frequency resolution levels of the image will be enhanced. Pre-define and load the corresponding enhanced frequency domain feature extraction module and subsequent processing unit for each high-frequency level.
[0139] Fourth, perform the following operations on each high-frequency layer, where the selection order of the high-frequency layers is from the lower resolution layer to the higher resolution layer in sequence, or set according to the order of actual requirements.
[0140] (1) Spatial alignment and interpolation.
[0141] Upsample the generated Mask, such as by bilinear interpolation, to the resolution of the current high-frequency layer I h[-2-i] to make the mask size consistent with the high-frequency feature layer to be processed.
[0142] (2) Frequency domain enhanced feature extraction.
[0143] Use the pre - defined frequency - domain enhanced feature module to perform frequency - domain analysis and enhancement on the current high - frequency layer image feature I h[-2-i] Perform frequency - domain analysis and enhancement.
[0144] The frequency - domain enhanced feature module includes: a frequency - band attention module, a global context encoding module, and a frequency mask generation function. This module mines high - frequency detail features from the frequency - domain level and adaptively strengthens them, so as to provide higher - fidelity and richer detail information in subsequent fusion. The obtained enhanced output is denoted as enhanced.
[0145] The specific processing flow of using the frequency - domain enhanced feature module for frequency - domain analysis and enhancement is as follows.
[0146] a. Discrete cosine transform.
[0147] Convert the current high - frequency layer image feature x from the spatial domain to the frequency domain through the dct_2d(x) function implemented based on the torch_dct library to obtain the DCT coefficients.
[0148] b. Frequency mask generation.
[0149] Use the self.frequency_band_masks(h,w) function to generate frequency masks. This function calculates the frequency position matrix according to the pixel coordinates and divides the frequency domain into multiple frequency bands according to the distance, and each frequency band corresponds to a mask.
[0150] c. Frequency - band selection.
[0151] Multiply the mask of each frequency band by the DCT coefficients to obtain the DCT coefficients and their amplitudes of each frequency band.
[0152] d. Frequency - band attention.
[0153] Use the frequency - band attention module to weight the information of different frequency bands.
[0154] The frequency - band attention module is a lightweight network, including two 1×1 convolutional layers, a ReLU activation function, and a Sigmoid activation function. It concatenates the amplitude spectra of each frequency band in the channel dimension as the input and outputs the weights of each frequency band. This module can assign different weights according to the importance of different frequency bands and enhance the information of important frequency bands.
[0155] e. Frequency - band weighting.
[0156] Multiply the DCT coefficients band of each frequency band by the corresponding frequency - band weights to obtain the enhanced frequency - band DCT coefficients.
[0157] f. Inverse discrete cosine transform (IDCT).
[0158] Through the idct_2d() function implemented based on the torch_dct library, the enhanced DCT coefficient, that is, the sum of all frequency band coefficients, is converted back to the spatial domain to obtain the enhanced high-frequency detail features enhanced.
[0159] (3) Mask fusion strategy.
[0160] Combine the enhanced feature enhanced with the weighted strategy guided by Mask to obtain the high-frequency result result_highfreq. This process is equivalent to flexibly adjusting the enhancement amount through the mask weight on the basis of maintaining the original enhancement information to avoid over-sharpening or information loss.
[0161] (4) Post-processing and feature fine-tuning.
[0162] The high-frequency result result_highfreq after the above fusion is input into the predefined post-processing unit, which is a lightweight post-processing unit including two DSConv and one LeakyReLU activation function. This unit can further refine, balance and stabilize the enhanced features, making its output more consistent and natural with the original feature structure in terms of texture, brightness, etc. The final result after the enhanced processing of this layer is obtained and retained for subsequent splicing.
[0163] The subsequent processing unit does not process each high-frequency level in isolation, but is based on frequency domain enhancement and mask fusion to further fine-tune the enhanced and fused high-frequency details.
[0164] Fifth, output level aggregation and reconstruction.
[0165] After completing the enhancement processing of all high-frequency levels, the processing results of each high-frequency level are stored in the list pyr_result from top to bottom or in the corresponding order. 1_enhanced Add to the result list so that the final output contains complete reconstruction information from high frequency to low frequency.
[0166] Step 4: Progressive reconstruction of multi-scale information.
[0167] The enhanced low-frequency layer output and the high-frequency enhancement layer output of each scale are progressively reconstructed according to the inverse process of the Laplacian pyramid. The reconstruction process adopts a bottom-up layer-by-layer synthesis strategy:
[0168] Taking the enhancement result of the low scale as a benchmark, it is upsampled and processed by depthwise separable convolution so that it can be precisely aligned and fused with the high-frequency detail layer of higher resolution. As the scale gradually increases, the high-frequency detail layer corresponding to the scale is fused into the reconstruction process to obtain the final image that takes into account both global color correction and local texture detail enhancement. Through the above progressive reconstruction, the enhanced underwater image is finally output.
[0169] Taking the bottom image (pyr[-1]) in the pyramid list as the initial basis, by successively upsampling, aligning, and fusing the higher resolution layers from bottom to top, an image close to the original resolution can be finally reconstructed. Specifically:
[0170] First, starting from the lowest resolution layer, it is upsampled and aligned to the size of the higher resolution feature layer by interpolation if necessary; then, the aligned result is added to the high-frequency features of this layer after frequency domain enhancement and masking in step three, so as to superimpose more details and local textures.
[0171] Repeat the above process, gradually fusing the global structure at the bottom of the pyramid to the high-frequency features at the top layer, and finally obtaining a high-resolution reconstructed image with complete details.
[0172] It should be noted that the upsampling in this application is not simply bilinear interpolation, but a upsampling template is constructed through torch.cat and view operations, and a Gaussian convolution function is used for convolution operation. This method can further improve the reconstruction quality.
[0173] The progressive reconstruction strategy adopted in this step can be regarded as an improvement and optimization of the existing Laplacian pyramid reconstruction method. Although the Laplacian pyramid decomposition and reconstruction itself is a classic method in the field of image processing, the usual reconstruction method is to simply stack each level directly or perform fusion through a fixed weight fusion strategy. The innovative point of the progressive reconstruction method proposed in this application is that the enhanced low-frequency component is used as a guide to gradually fuse with the enhanced high-frequency details, and the order and method of fusion from low frequency to high frequency correspond to the pyramid decomposition process. This strategy makes more full use of the guiding role of low-frequency information on high-frequency details, can better retain the detail information of the image, reduce the reconstruction error, and obtain better reconstruction quality.
[0174] Step Five: Output Results and Applications.
[0175] The enhanced underwater image is output to a display terminal or storage medium for subsequent application scenarios such as ocean exploration, target recognition, underwater archaeology, and underwater ecological monitoring. If necessary, it can be deployed and applied on an embedded platform, such as Xavier.
[0176] Figure 2is the image to be processed, Figure 3 is the image processed by the method described in this application. After Figure 2 and Figure 3 comparison, it can be seen that the image processed by the method described in this application is clearer, the color distortion is effectively corrected, and the naturalness restoration and detail enhancement of the image are realized.
[0177] The above has introduced in detail the underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning, characterized in that It includes the following steps: S1. Image input and multi-scale decomposition: decompose the input image into a low-frequency residual layer and several high-frequency detail layers through Laplacian pyramid decomposition; S2. Input the low-frequency residual layer into the global illumination and color correction sub-module to obtain an enhanced low-frequency residual layer; S3. Input the enhanced low-frequency residual layer and high-frequency detail layers of multiple scales into the frequency domain enhanced feature module to obtain the output of the multi-scale enhanced high-frequency detail layers; S4. Perform progressive reconstruction on the output of the enhanced low-frequency layer and the output of the high-frequency enhanced layers of each scale according to the inverse process of the Laplacian pyramid; S5. Output results and applications.
2. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 1, wherein The specific implementation process of step S1 is as follows: Real-time collect the underwater optical image sequence and read the latest frame of image I(w, h, 3) in real time; Input the image I(w, h, 3) into the Laplacian pyramid enhancement network model, and decompose it into frequency components of multiple scales, including a low-frequency residual layer I1 and several high-frequency detail layers I hi , where i = 1, 2, …, n.
3. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 2, characterized in that Adopt the supervised learning method to train the Laplacian pyramid enhancement network model, and use the comprehensive optimization loss function to guide the parameter optimization of the Laplacian pyramid enhancement network model. The equation of the comprehensive optimization loss function is: L = λ1·L 结构一致性 + λ2·L 颜色 + λ3·L 感知 + λ4·L 对比 , where, λ1 represents the structure consistency loss weight; λ2 represents the color consistency loss weight; λ3 represents the perceptual loss weight; λ4 represents the contrast loss weight; Structural consistency loss L 结构一致性 The function of which is: L 结构一致性 = 1 - pytorch_ssim.ssim(I 生成 , I 参考 ), Among them, I 生成 represents the output generated image; I 参考 represents the reference image; Perceptual loss L 感知 The function is as follows: Among them, φ i represents the feature representation extracted by the pre-trained neural network for the input image at the i-th layer; Color consistency loss L 颜色 The function of Among them, S represents the set composed of a group of subsets formed after the image is divided by regions; s represents the index of a specific subset in the set S; N s represents the number to be traversed for the s-th subset; p represents the index used to index a single sampling unit in the s-th subset; x refers to the horizontal coordinate of the image, which is used to calculate the horizontal gradient; y refers to the vertical coordinate of the image, which is used to calculate the vertical gradient; Contrastive loss L 对比 The function of which is:
4. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 1, characterized in that The network structure of the global illumination and color correction sub-module successively includes a feature extraction layer, a residual block group, and an output layer; The feature extraction layer successively includes a first DSConv layer, an InstanceNorm2d layer, a first LeakyReLU activation function layer, a second DSConv layer, and a second LeakyReLU activation function layer; The first DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution, which expands the number of channels of the input low-frequency residual layer I1 from 3 to 16 for preliminary feature extraction; the InstanceNorm2d layer is used to perform instance normalization on the feature map; the first LeakyReLU activation function layer is used to introduce non-linearity; the second DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to expand the number of channels from 16 to 64 for further feature extraction; The residual block group includes five residual modules, each residual module contains two DSConv layers and a skip connection; the DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution; the skip connection directly adds the input of the residual module to the output of the two DSConv layers; The output layer successively includes a third DSConv layer, a third LeakyReLU activation function layer, and a fourth DSConv layer; The third DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to compress the number of channels from 64 back to 16; the fourth DSConv layer uses a 3×3 depth convolution and a 1×1 pointwise convolution to compress the number of channels from 16 back to 3; after the output layer, a tanh activation function is connected to map the output value to the range of [-1, 1].
5. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 4, characterized in that The data processing flow of the global illumination and color correction sub-module is as follows: S2.
1. Extract and transform the primary features from the original three-channel RGB image through the feature extraction layer, using depthwise separable convolution, instance normalization, and the LeakyReLU activation function. S2.
2. Further extract and refine the features of the primary features obtained in step S2.1 through five residual modules. S2.
3. The output layer remaps the features back to the same channel dimension as the input to achieve the reduction process from the low-level feature space back to the original image space, and maps the output values to the range of [-1, 1] through the tanh activation function to obtain the enhanced low-frequency residual layer I 1_emjamced .
6. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 1, characterized in that, The specific implementation process of step S3 is as follows: S3.
1. Concatenate the multi-scale high-frequency detail layer I of the image to be enhanced, hi , where i = 1, 2, …, n, the low-frequency residual layer, and the enhanced low-frequency residual layer I 1_enhanced to obtain the input X; S3.
2. Generate the initial mask feature map Mask0 to obtain the adaptive mask Mask for weighted fusion. S3.
3. Determine the number of high-frequency levels to be processed according to the design parameter num_high. S3.
4. Perform the following operations on each high-frequency layer: S3.4.
1. Spatial alignment and interpolation: Upsample the generated Mask to make the mask size consistent with the high-frequency feature layer to be processed. S3.4.
2. Frequency-domain enhanced feature extraction: The frequency-domain enhanced feature module performs frequency-domain analysis and enhancement on the current high-frequency layer image features to obtain the enhanced high-frequency detail features enhanced. S3.4.
3. Mask fusion strategy: Combine the enhanced features enhanced with the weighted strategy guided by Mask to obtain the high-frequency result result_highfreq. S3.4.
4. Post-processing and feature fine-tuning: Input the fused high-frequency result result_highfreq into the subsequent processing unit. S3.
5. Output level aggregation and reconstruction: After completing the enhancement processing of all high-frequency levels, store the processing results of each high-frequency level in the list pyr_result in the corresponding order. At the same time, add the low-frequency enhancement result I 1_enhanced to this result list.
7. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 6, wherein The specific implementation process of step S3.2 is as follows: S3.2.
1. Perform a depthwise separable convolution DSConv operation on the input X to expand the number of channels of the input X from 9 to 64. This DSConv layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution. S3.2.
2. Introduce non-linearity through the LeakyReLU activation function layer. S3.2.
3. Fuse the multi-scale feature extraction module and several residual blocks to obtain the initial mask feature map Mask0. The multi-scale feature extraction module contains four parallel branches, which use convolution kernels of 1×1, 3×3, 5×5, and 7×7 respectively for feature extraction, and splice the outputs of the four branches in the channel dimension, and finally perform feature fusion through a 1×1 convolution. S3.2.
4. Apply the Sigmoid function to Mask0 to map it to a continuous range between 0 and 1 to obtain the adaptive mask Mask for weighted fusion.
8. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 6, characterized in that, In step S3.4.2, the frequency-domain enhanced feature module includes a frequency band attention module, a global context encoding module, and a frequency mask generation function. The specific processing flow for performing frequency-domain analysis and enhancement using the frequency-domain enhanced feature module is as follows: S3.4.2.
1. Convert the current high-frequency layer image feature x from the spatial domain to the frequency domain through the dct_2d(x) function implemented based on the torch_dct library to obtain the DCT coefficients. S3.4.2.
2. Generate frequency masks using the self.frequency_band_masks(h,w) function: This function calculates the frequency position matrix according to the pixel coordinates and divides the frequency domain into several frequency bands according to the distance, and each frequency band corresponds to a mask. S3.4.2.
3. Multiply the mask of each frequency band by the DCT coefficients to obtain the DCT coefficients and their amplitudes for each frequency band; S3.4.2.
4. Use the frequency band attention module to weight the information of different frequency bands: The frequency band attention module includes two 1×1 convolutional layers, a ReLU activation function, and a Sigmoid activation function. This module takes the amplitude spectra of each frequency band concatenated in the channel dimension as input and outputs the weights of each frequency band; S3.4.2.
5. Multiply the DCT coefficients of each frequency band by the corresponding frequency band weights to obtain the enhanced frequency band DCT coefficients; S3.4.2.
6. Through the idct_2d() function implemented based on the torch_dct library, convert the enhanced frequency band DCT coefficients, that is, the sum of all frequency band coefficients, back to the spatial domain to obtain the enhanced high-frequency detail feature enhanced.
9. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 1, characterized in that, The specific implementation process of step S4 is as follows: S4.
1. Starting from the lowest resolution layer, upsample it and align it to the size of the higher resolution feature layer through interpolation; S4.
2. Add the alignment result of step S4.1 to the high-frequency feature after frequency domain enhancement and masking processing of this layer, so as to superimpose more details and local textures; Repeat steps S4.1 and S4.2, gradually fusing the global structure at the bottom of the pyramid to the high-frequency features at the top layer, and finally obtaining a high-resolution reconstructed image with complete details.
10. The underwater low-quality image enhancement method based on Laplacian pyramid and contrast learning according to claim 9, characterized in that In step S4.1, an upsampling template is constructed through torch.cat and view operations, and convolution operations are performed using the Gaussian convolution function to achieve upsampling.
Citation Information
Patent Citations
Low-illumination image enhancement method based on deep learning and Laplacian pyramid
CN117196968A
Laplace three-layer cyclic high-definition image enhancement method
CN117611465A
Multi-channel illumination adaptive network model with Laplacian enhanced pyramid and low-light image enhancement method
CN119579448A
Underwater image dynamic enhancement method based on pyramid network and application
CN119599924A
Image processing method and apparatus, storage medium, and electronic device
WO2022017025A1
Cited By
Image diffusion enhancement method and device based on multi-scale feature extraction and fusion
CN120976039A
Image diffusion enhancement method and device based on multi-scale feature extraction and fusion
CN120976039B
Anaphora image segmentation method based on space-frequency dual tuning
CN120976550A
Reference image segmentation method based on space-frequency duality tuning
CN120976550B
Image enhancement method based on double-domain illumination prior and electronic equipment
CN121366092A