An underwater image dynamic enhancement method and application based on a pyramid network

Through the decomposition and combination method of the pyramid network, the shortcomings of underwater image enhancement technology in adaptability, color distortion and noise suppression are solved, and high-quality image enhancement in different underwater environments is achieved, which improves the visibility and recognition of the image.

CN119599924BActive Publication Date: 2025-07-29SHENZHEN MSU-BIT UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411740672.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-07-29
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The existing underwater image enhancement technology has shortcomings in adaptability, color distortion and noise suppression, making it difficult to achieve consistent high-quality image enhancement in different underwater environments.

Method used

Using a pyramid network-based method, color correction, refinement and noise suppression are performed through Laplace pyramid decomposition, deterministic color mapping network, dual convolution attention module and group expansion feedforward network, and dynamic enhancement is achieved through adaptive color correction and high-frequency component refinement.

Benefits of technology

Maintain consistent image enhancement effects in different underwater environments, effectively correct color distortion, suppress noise, preserve image details, and improve visibility and recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119599924B_ABST
    Figure CN119599924B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of underwater image enhancement processing, and in particular to an underwater image dynamic enhancement method and application based on a pyramid network, including: decomposing an input underwater image into a hierarchical image group; using a deterministic color mapping network to convert the input image of the low-frequency layer into an output image of the low-frequency layer; using a double convolutional attention module to obtain enhanced input features; using a grouped dilated feed-forward network to obtain a pixel-by-pixel mask; performing an operation on the enhanced input features and the step-by-step pixel mask to obtain a high-frequency component; using a Laplacian pyramid network to reconstruct the output image of the low-frequency layer and the high-frequency component to obtain an enhanced output underwater image. In summary, the underwater image can be dynamically enhanced, with strong adaptability, applicable to various water bodies, capable of effectively correcting color distortion, suppressing noise introduction, and retaining more image details, making the underwater image have better visibility and higher recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater image enhancement processing, and in particular to an underwater image dynamic enhancement method and application based on a pyramid network. Background Art

[0002] In the underwater environment, due to the influence of light scattering, absorption, and suspended particles in the water body, underwater images usually have problems such as color distortion, low contrast, and poor clarity. These factors seriously affect the quality of underwater images, thus bringing difficulties to underwater detection, search, monitoring and other tasks. With the wide application of underwater detection and monitoring technologies, such as in the fields of marine archaeology, marine resource development, and marine biological research, higher requirements are put forward for the processing and enhancement of underwater images. Improving the visibility and detail restoration of underwater images through image enhancement technology is crucial for enhancing the success rate of these tasks.

[0003] Existing underwater image enhancement technologies mainly focus on optical distortion and color distortion problems, and a series of methods have been developed to improve image quality. However, the existing technologies still cannot fully overcome the complex optical conditions in the underwater environment, and in the process of image restoration, it is easy to introduce noise or cause detail loss. Although traditional enhancement methods have improved the visual effect of images to a certain extent, they have poor adaptability to different water depths, lighting conditions, and different water body components, and cannot achieve real-time enhancement in dynamic scenarios.

[0004] Currently, the existing underwater image enhancement technology solutions are mainly divided into two categories: image enhancement methods based on physical models and enhancement methods based on image processing algorithms.

[0005] (1) Methods based on physical models: This type of method utilizes the physical characteristics of underwater imaging to establish a light propagation model to inversely deduce and correct the degradation of underwater images. Typical physical models include the atmospheric scattering model (DCP) and the color compensation model (CCM). By separating the scattered light and background light, the true color and contrast in the image are restored. The advantage of this type of method lies in its strong physical interpretability. However, since the model establishment depends on the optical parameters of the water body, when the water quality changes or the environment is complex, the model shows poor adaptability.

[0006] For example, the Chinese patent document with the publication number CN110851965A discloses a light source optimization method and optimization system based on a physical model, including constructing an underwater camera model, constructing an underwater light source model, constructing a water effect model, constructing a complete underwater optical imaging model, setting simulation method parameters, image quality evaluation criteria, optimization algorithms, and experimental verification.

[0007] (2) Methods based on image processing algorithms: These methods do not rely on the physical model of the underwater environment. They mainly use image enhancement algorithms such as histogram equalization and contrast enhancement to improve the visual quality of images. In recent years, deep learning technology has also been widely applied in the field of underwater image enhancement. Based on deep learning models such as convolutional neural networks (CNNs), the network is trained to automatically enhance images. However, deep learning methods rely on a large amount of labeled data and have limited generalization ability in complex underwater environments, and are prone to performance degradation in complex scenarios.

[0008] For example, the Chinese patent document with the publication number CN115034965A discloses a super-resolution underwater image enhancement method and system based on deep learning. The method includes: constructing an improved generative adversarial network model based on a generative adversarial network and a deep residual multiplier; inputting an underwater image into the improved generative adversarial network model for training to generate an enhanced super-resolution underwater image. The system includes: a construction module and a training module. By using the present invention, it is possible to improve the visual quality of underwater images while increasing the resolution of underwater images.

[0009] Thus, it can be seen that in the existing technical solutions, there are the following technical drawbacks:

[0010] Poor adaptability: Methods based on physical models usually rely on the optical properties of water bodies, such as parameters like turbidity, depth, and light direction. Therefore, when the water body environment changes, the adaptability of the model is poor. Different water depths and lighting conditions will result in unstable enhancement effects and cannot achieve consistent improvement in image quality in various complex environments.

[0011] Color distortion problem: In the underwater environment, light of specific wavelengths will be absorbed to varying degrees, resulting in color distortion in the image. Existing methods often have difficulty effectively correcting color deviations and restoring true natural colors. Even if some methods can perform color compensation, it is easy to introduce excessive color enhancement under certain conditions, resulting in an unnatural effect of the image.

[0012] Noise introduction: Due to the poor quality of underwater images themselves, especially in the case of insufficient light in deep water, image enhancement methods may amplify the noise in the image. Some techniques based on contrast enhancement are prone to amplifying the noise simultaneously when enhancing image details, resulting in a decrease in the visibility of the image.

[0013] Detail loss: Some enhancement methods are prone to ignoring the restoration of details when improving the overall contrast and color saturation of the image, especially for underwater images affected by strong scattered light. Existing technologies often cannot retain sufficient detail information while ensuring the clarity of the image, resulting in the restored image lacking fine structure.

[0014] Therefore, it is necessary to improve the existing technology to overcome the above technical drawbacks. Summary of the Invention

[0015] The present invention aims to provide a technical solution to solve the above problems in order to overcome the above deficiencies.

[0016] The present invention provides a method for dynamically enhancing underwater images based on a pyramid network, which includes the following steps:

[0017] Step 1: First, decompose the initial input underwater image into a set of hierarchical images with a total of L layers using a Laplacian pyramid network. The L-th layer is the low-frequency layer input image, and the first layer to the (L - 1)-th layer is a set of high-frequency layer input images.

[0018] Step 2: Use a deterministic color mapping network to perform color correction and illumination compensation on the low-frequency layer input image to convert it into a low-frequency layer output image.

[0019] Step 3: Use a double convolutional attention module to extract and enhance the input features of the input image of the n-th layer to obtain the enhanced input feature D of the input image n ;

[0020] Step 4: Use a grouped dilated feed-forward network to refine the output features of the output image of the (n + 1)-th layer to obtain the per-pixel mask M after the refinement operation of the output image n+1 ;

[0021] Step 5: Perform an arithmetic operation on the enhanced input feature of the input image of the n-th layer in Step 3 and the per-pixel mask after the refinement operation of the output image of the (n + 1)-th layer in Step 4 to obtain the high-frequency component h of the high-frequency layer output image of the n-th layer n ;

[0022] Step 6: Repeat the operations in Step 3, Step 4, and Step 5. After repeating (L - 1) times, obtain a set of high-frequency component sets HF, HF = [h L-1 , …, h n+1 , h n , …, h2, h1];

[0023] Step 7: Use a Laplacian pyramid network to reconstruct the low-frequency layer output image and the high-frequency component set HF to obtain the enhanced output underwater image.

[0024] As a further solution of the present invention: In Step 2, the deterministic color mapping network is a convolutional neural network, and its conversion process includes the following steps:

[0025] Step 2.1: Expand the low-frequency layer input image into a two-dimensional matrix.

[0026] Step 2.2, perform matrix multiplication using the projection matrix P(k, 3) and the two-dimensional matrix, so as to embed each pixel in the low-frequency layer input image into a k-dimensional vector space;

[0027] Step 2.3, perform matrix multiplication on the embedded k-dimensional vector space and the color mapping matrix T(k, k), so as to adjust the color;

[0028] Step 2.4, perform matrix multiplication using the projection matrix Q(3, k) and the adjusted vector space, convert the adjusted vector space back to the RGB color space, and recombine each pixel into an image of the original size, so as to output the low-frequency layer output image with enhanced color.

[0029] As a further solution of the present invention: in step 3, preferably, the double convolutional attention module specifically includes the following steps: extract the input features of the input image through two convolutional layers, and perform non-linear activation using the LeakyReLU activation function.

[0030] As a further solution of the present invention: in step 4, the refinement operation further includes using at least two grouped dilated feed-forward networks to form multiple grouped dilated feed-forward networks, and respectively perform refinement adjustment on the output features of the output image, specifically including the following steps:

[0031] Step 4.1, use multiple grouped dilated feed-forward networks for cascading in sequence, and gradually adjust the output features of the output image, so as to gradually correct the output features;

[0032] Step 4.2, the corrected features fed into each grouped dilated feed-forward network are also densely connected to the output end of the last grouped dilated feed-forward network and perform © connection to form the total output features of the multiple grouped dilated feed-forward networks;

[0033] Step 4.3, then use the channel attention module to perform adaptive weight adjustment on the channel dimension of the total output features, so as to obtain a per-pixel mask after the refinement operation on the output features.

[0034] As a further solution of the present invention: in step 4, the grouped dilated feed-forward network specifically includes the following method:

[0035] Use two cascaded 3x3 convolutional layers and combine with the GELU activation function to capture the local spatial features of the output features of the output image;

[0036] At the same time, use two cascaded 1x1 convolutional layers and combine with the GELU activation function to capture the local spatial features of the output features of the output image;

[0037] Perform a bitwise addition operation on two sets of local spatial features, and then perform a residual connection on the result with the output features of the output image;

[0038] Use an adaptive hybrid attention module for capture, form a 3x3 convolutional layer after capture, then perform a residual connection with the output features of the output image, and finally obtain a per-pixel mask after refining the output features.

[0039] As a further solution of the present invention: In step 5, set the input features after enhancing the input image of the nth layer to D n ; Set the per-pixel mask after refining the output image of the (n + 1)th layer to M n+1 ; After calculation, obtain the high-frequency component h of the high-frequency layer output image of the nth layer n ; Its specific calculation formula is:

[0040] ,

[0041] where, is a matrix multiplication operation, is a bitwise addition operation.

[0042] As a further solution of the present invention: In step 4, the per-pixel mask of the low-frequency layer image of the Lth layer is obtained by the following steps:

[0043] Use bilinear operations to upsample the low-frequency layer input image and the low-frequency layer output image respectively to improve their resolutions, and © connect the upsampled low-frequency layer input image and the low-frequency layer output image to form [low-frequency layer input image, low-frequency layer output image];

[0044] Feed the formed [low-frequency layer input image, low-frequency layer output image] into a double convolutional attention module to obtain the enhanced input feature D L ;

[0045] Use a grouped dilated feed-forward network to refine the enhanced input feature D L to obtain the per-pixel mask LF of the low-frequency layer image of the Lth layer L .

[0046] The present invention also provides an underwater image dynamic enhancement system based on a pyramid network, and the system includes:

[0047] A decomposition module for decomposing the initial input underwater image into a hierarchical image group of a total of L layers using a Laplacian pyramid network, where the Lth layer is the low-frequency layer input image, and the first layer to the (L - 1)th layer is a group of high-frequency layer input image groups;

[0048] A conversion module, which is used to perform color correction and illumination compensation processing on the input image of the low-frequency layer by using a deterministic color mapping network, so as to convert it into an output image of the low-frequency layer;

[0049] An enhancement module, which is used to extract and enhance the input features of the input image of the nth layer by using a double convolutional attention module, so as to obtain the enhanced input features of the input image;

[0050] A refinement module, which is used to refine the output features of the output image of the (n + 1)th layer by using a grouped dilated feed-forward network, so as to obtain a pixel-by-pixel mask after the refinement operation of the output image;

[0051] An operation module, which is used to perform an operation on the enhanced input features and the pixel-by-pixel mask after the refinement operation, so as to obtain the high-frequency component of the high-frequency layer output image of the nth layer;

[0052] A repetition module, which is used to repeat the enhancement module, the refinement module and the operation module. After repeating L - 1 times, a set of high-frequency component sets HF is obtained, HF = [h L-1 ,…,h n+1 ,h n ,…,h2,h1];

[0053] A reconstruction module, which is used to reconstruct the output image of the low-frequency layer and the high-frequency component set HF by using a Laplacian pyramid network, so as to obtain an enhanced output underwater image.

[0054] The present invention also provides a device, including:

[0055] A memory, which is used to store a computer program;

[0056] A processor, which is used to execute the computer program and implement the above-mentioned underwater image dynamic enhancement method based on a pyramid network.

[0057] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer processor, the above-mentioned underwater image dynamic enhancement method based on a pyramid network is implemented.

[0058] Compared with the prior art, the beneficial effects of the present invention are:

[0059] Aiming at the defects in the prior art, the present invention proposes a more robust and efficient underwater image enhancement method, aiming to achieve the following improvements:

[0060] Enhanced adaptability: By combining physical models with deep learning algorithms and introducing an adaptive color mapping matrix T and channel attention layers, the present invention can adaptively adjust the enhancement strategy in various underwater environments. Whether it is shallow water or deep water, clear or turbid water bodies, it can maintain a consistent image enhancement effect.

[0061] Effective correction of color distortion: Through a deterministic color mapping network, the present invention uses an adaptive color correction algorithm to dynamically adjust color compensation according to different water depths and lighting conditions, ensuring that the generated images have natural color representations and avoiding unnatural phenomena caused by over-enhancement.

[0062] Suppression of noise introduction: Through a grouped dilated feed-forward network and the introduction of an adaptive hybrid attention module, especially in low-contrast or poorly lit scenarios, the present invention can effectively filter out noise while enhancing image details, ensuring the clarity and visual quality of the images.

[0063] Retention of image details: By combining the input features and output features of the low-frequency layer images and performing refinement operations on the high-frequency components, the present invention realizes a detail enhancement algorithm. While enhancing the contrast, it pays attention to the restoration of image details, ensuring that the texture and structure information in the images are fully retained and improving the recognition ability of underwater targets.

[0064] Therefore, through the above improvements, the present invention can provide a dynamic underwater image enhancement method and application based on a pyramid network, thereby dynamically enhancing underwater images, with strong adaptability, applicable to a variety of different water bodies, capable of effectively correcting color distortion, suppressing noise introduction, and retaining more image details, making underwater images have better visibility and higher recognition.

[0065] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0067] Figure 1 is a flowchart of the present invention;

[0068] Figure 2 is a flowchart of the per-pixel mask of the L-th low-frequency layer image of the present invention;

[0069] Figure 3 is a schematic diagram of the technical principle framework of the present invention;

[0070] Figure 4 is a schematic diagram of the framework of the grouped expansion feedforward network of the present invention;

[0071] Figure 5 is a schematic diagram of the Laplacian pyramid decomposition of the present invention. Detailed implementation manners

[0072] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0073] Please refer to Figures 1 to 5 , in the embodiments of the present invention, an underwater image dynamic enhancement method based on a pyramid network includes the following steps:

[0074] Step 1, first decompose the initial input underwater image into a hierarchical image group with a total of L layers using a Laplacian pyramid network, and the resolutions of the hierarchical images from the first layer to the Lth layer decrease gradually, where the Lth layer is the low-frequency layer input image, and the first layer to the L-1th layer is a group of high-frequency layer input image groups;

[0075] Step 2, use a deterministic color mapping network to perform color correction and illumination compensation processing on the low-frequency layer input image to convert it into a low-frequency layer output image;

[0076] Step 3, use a double convolutional attention module to extract and enhance the input features of the input image of the nth layer to obtain the enhanced input feature D of the input image n ;

[0077] Step 4, use a grouped expansion feedforward network to refine the output features of the output image of the (n + 1)th layer to obtain the per-pixel mask M after the refinement operation of the output image n+1 ; Specifically, the output features can be either the high-frequency component features of the high-frequency layer output image or the low-frequency component features of the low-frequency layer output image.

[0078] Step 5, perform matrix multiplication and bitwise addition operations on the enhanced input feature of the input image of the nth layer in Step 3 and the per-pixel mask after the refinement operation of the output image of the (n + 1)th layer in Step 4 to obtain the high-frequency component h of the high-frequency layer output image of the nth layer n ;

[0079] Step 6, repeat the operations in Step 3, Step 4, and Step 5. After repeating L - 1 times, a set of high - frequency component sets HF is obtained, HF = [h L-1 ,…,h n+1 ,h n ,…,h2,h1];

[0080] Step 7, use the Laplacian pyramid network to reconstruct the low - frequency layer output image and the high - frequency component set HF to obtain the enhanced output underwater image.

[0081] Specifically, the present invention proposes an end - to - end framework based on the Laplacian pyramid network to reduce the computational complexity. Given an input underwater image, it is first decomposed into a Laplacian pyramid to obtain a set of high - frequency components h and a low - frequency layer image. The components in h have gradually decreasing resolutions, starting from h×w and reaching the coarsest level layer , while the low - frequency layer image is represented by pixels. This decomposition is completely reversible, allowing the original image to be reconstructed through successive operations.

[0082] Secondly, the low - frequency layer image can reflect the global attributes of the image, including color and brightness. At the same time, other high - frequency components contain the edge detail processing of the image. Therefore, the color is mainly represented by the low - frequency components. Then, a deterministic color mapping network is introduced to process the low - frequency layer image and convert it into a low - frequency layer output image, thereby adjusting the color and brightness of the underwater image.

[0083] Thirdly, the high - frequency layer image mainly contains the small structures and details of the image. The high - frequency components contain the edge detail processing of the image. Precise enhancement of the high - frequency components is crucial for restoring the small structures and fine details in the image. Since there is a strong spatial correlation between different frequency components, in order to facilitate the reconstruction of the high - frequency components, it is necessary to explore the inter - layer correlation in the Laplacian pyramid.

[0084] Therefore, an efficient and progressive upsampling strategy is adopted. The multi - block grouped dilated feed - forward network is applied to each high - frequency layer, which uses the features from the lower layer to support the refinement in the current processing layer. By learning the masks of the L - 1 layer high - frequency components and gradually expanding the masks according to the intrinsic characteristics of the Laplacian pyramid to refine the remaining high - frequency components.

[0085] After the refinement operation of the multi - block grouped dilated feed - forward network, corresponding pixel - by - pixel masks can be generated. And the pixel - by - pixel masks of each layer are updated through the multi - block grouped dilated feed - forward network. At each level from n = L - 1 to n = 1, the high - frequency components of each layer are refined and adjusted through the pixel - by - pixel masks, and then a high - frequency component set HF is obtained, HF = [h L-1 ,…,h n+1 ,h n,…,h2,h1].

[0086] Finally, by applying the Laplacian pyramid to reconstruct the output image of the low-frequency layer and the high-frequency component set HF, the enhanced underwater image J can be obtained. The enhanced underwater image can effectively correct color distortion, suppress the introduction of noise, and retain more image details, making the underwater image have better visibility and higher recognition.

[0087] In another embodiment of the present invention, preferably, in step 2, the deterministic color mapping network is a convolutional neural network, and its conversion process includes the following steps:

[0088] Step 2.1, expand the input image of the low-frequency layer into a two-dimensional matrix of size (3, h×w);

[0089] Step 2.2, perform matrix multiplication on the two-dimensional matrix using the projection matrix P(k, 3), so as to embed each pixel in the input image of the low-frequency layer into a k-dimensional vector space;

[0090] Step 2.3, perform matrix multiplication on the embedded k-dimensional vector space and the color mapping matrix T(k, k), so as to adjust the color;

[0091] Step 2.4, perform matrix multiplication on the adjusted vector space using the projection matrix Q(3, k), convert the adjusted vector space back to the RGB color space, and recombine each pixel into an image of the original size, so as to output the output image of the low-frequency layer with enhanced color.

[0092] Specifically, given the input image of the low-frequency layer, color correction and illumination compensation are first performed through a deterministic color mapping network. The convolutional neural network of deterministic color mapping is a technology for image processing and image conversion. Its core idea is to use a trained convolutional neural network to realize the mapping from the color of the input image to the target color. The combination of filters or LUTs with CNN for color mapping has obvious limitations. The deterministic color mapping network proposed in the present invention can efficiently extract high-level features in the image without relying on traditional filters or LUTs.

[0093] Among them, both P and Q are learnable matrices shared by all images in the prior art. P maps the input data from the RGB space (3 channels) to a high-dimensional embedding space through linear projection. Q maps the data from the high-dimensional embedding space back to the original RGB space through linear inverse projection. Moreover, the three matrices P, T, and Q can be applied multiple times to further refine the color adjustment. For example, taking the output image of the low-frequency layer in the above steps as the input image again, and repeating the operation of step 2, so as to perform more refined multiple adjustments on the low-frequency layer image.

[0094] Secondly, before and after the conversion by the deterministic color mapping network, the data can be subjected to the ® operation. And the ® in the technical solution framework diagram, that is, the reshape operation, refers to the operation of reshaping and dimension transformation of tensors in the prior art. The reshape operation is an important function in PyTorch for changing the shape of tensors, which allows users to redefine the dimensions of tensors without changing the tensor data. PyTorch is an open-source deep learning framework for machine learning and deep learning.

[0095] Again, given the low-frequency layer input image of the input image and the input underwater image, the input underwater image is used to generate an adaptive color mapping matrix T for deterministic color mapping. Preferably, in step 2.3, the color mapping matrix T(k, k) is obtained by the following method: The initial input underwater image is input into the encoder E, and the encoder E extracts features of the input underwater image through the MobileNetV3 model, and then outputs the extracted image features as a color mapping matrix T of size (k×k), and then through the ® operation, it is reshaped into the form of T(k, k) to obtain the color mapping matrix T(k, k). Thus, T(k, k) can be used as a color transformation matrix to adjust the color of the low-frequency layer input image of the underwater image. Among them, MobileNetV3 is a lightweight network model proposed by the Google team in 2019.

[0096] In another embodiment of the present invention, preferably, in step 3, preferably, the double convolutional attention module specifically includes the following steps: Extract the input features of the input image through two convolutional layers and perform non-linear activation using the LeakyReLU activation function.

[0097] Specifically, preferably, the kernel size of the convolutional layer of the double convolutional attention module is set to 3. The two convolutional layers in this embodiment are preferably 3x3 convolutional layers, and the LeakyReLU activation function is used for non-linear activation for each convolutional layer, thereby enhancing the feature representation ability.

[0098] Secondly, preferably, a channel attention layer is embedded between the two convolutional layers of the double convolutional attention module for adaptive weight adjustment of the channel dimension. The channel attention layer is set as the SE attention module in the prior art for adaptive weight adjustment. The double convolutional attention module extracts local features through convolutional layers and combines the global information enhancement of the channel attention layer, effectively improving the expression ability of the model for input data.

[0099] In another embodiment of the present invention, preferably, in step 4, the refinement operation further includes using at least two grouped dilated feed-forward networks to form multiple grouped dilated feed-forward networks, and respectively performing refinement adjustment on the output features of the output image, which specifically includes the following steps:

[0100] Step 4.1: Use x grouped dilated feed-forward networks to perform cascading in sequence, and gradually adjust the output features of the output image to gradually correct the output features;

[0101] Step 4.2: The corrected features fed into each grouped dilated feed-forward network are also densely connected to the output end of the last grouped dilated feed-forward network and are ©-connected to form the total output features of the multiple grouped dilated feed-forward networks;

[0102] Step 4.3: Then use a channel attention module to adaptively adjust the weights of the channel dimension of the total output features, so as to obtain a per-pixel mask after performing the refinement operation on the output features. Specifically, the feature extraction is optimized through the combination of small modules of multiple grouped dilated feed-forward networks, and the grouped dilated feed-forward network ensures the balance between computational efficiency and performance.

[0103] Specifically, the ©-connection herein refers to a simple Concatenation. It means that multiple tensors (usually the outputs of different network branches) are concatenated (cascaded) in a certain dimension to form a new tensor.

[0104] In another embodiment of the present invention, preferably, in step 4, the grouped dilated feed-forward network specifically includes the following steps: Use two cascaded 3x3 convolutional layers, combined with the GELU activation function, to capture the local spatial features of the output features of the output image; at the same time, use two cascaded 1x1 convolutional layers, combined with the GELU activation function, to capture the local spatial features of the output features of the output image; finally, perform a bitwise addition operation on the two sets of local spatial features, and the result is then subjected to a residual connection operation with the output features of the output image; subsequently, use an adaptive hybrid attention module for capture and form a captured 3x3 convolutional layer, which is then subjected to a residual connection operation with the output features of the output image, and finally obtain a per-pixel mask after performing the refinement operation on the output features.

[0105] Specifically, using two cascaded 3x3 convolutional layers, combined with the GELU activation function, can capture rich local spatial features while keeping the computational amount relatively low. In addition, the introduction of the 1x1 convolutional layer helps cross-channel feature interaction, effectively reducing the number of parameters and making the model more lightweight. Then, through the residual link, the original features are retained without losing spatial information. The adaptive hybrid attention module is a feature module in the prior art, which can exhibit lightweight features and can more effectively capture the dependencies between channels.

[0106] Among them, a residual connection refers to an operation in a neural network that directly adds the input of a certain layer to the output of that layer. This operation can help the neural network better learn the identity mapping, that is, directly pass the input to the output, thereby avoiding the problems of gradient vanishing and gradient explosion, accelerating the training of the model, and improving the accuracy of the model.

[0107] In another embodiment of the present invention, preferably, in step 4, the per-pixel mask of the low-frequency layer image of the L-th layer is obtained by the following steps: bilinear operations are used to upsample the low-frequency layer input image and the low-frequency layer output image respectively to improve their resolution to match the resolution of the high-frequency layer image of the (L - 1)-th layer, and the upsampled low-frequency layer input image and the low-frequency layer output image are ©-connected to form the form of [low-frequency layer input image, low-frequency layer output image]; the [low-frequency layer input image, low-frequency layer output image] formed by the connection is fed into a double convolutional attention module to obtain the enhanced input feature D L ; subsequently, a grouped dilated feed-forward network is used to refine the enhanced input feature D L to obtain the per-pixel mask LF of the low-frequency layer image of the L-th layer L .

[0108] Specifically, since in step 5, when refining and updating the output image of the (L - 1)-th layer, the per-pixel mask of the L-th layer is required, and the L-th layer is a low-frequency layer image, therefore, in step 4, the low-frequency layer input image and the low-frequency layer output image can also be subjected to the enhancement operation of the double convolutional attention module and the refinement operation of the grouped dilated feed-forward network. In addition, since the per-pixel mask of the high-frequency component of the high-frequency layer output image of the first layer is not used, the high-frequency component of the first layer can be not refined, thus saving computational effort.

[0109] In another embodiment of the present invention, preferably, in step 5, the enhanced input feature of the n-th layer input image is set as D n ; the per-pixel mask after the refinement operation of the (n + 1)-th layer output image is set as M n+1 ; through calculation, the high-frequency component h of the high-frequency layer output image of the n-th layer is obtained n ; its specific calculation formula is:

[0110] ,

[0111] wherein, is the matrix multiplication operation, is the bitwise addition operation.

[0112] Specifically, is the per-pixel mask after the refinement operation of the output image of the (n + 1)-th layer, D n is the input feature after enhancing the input image of the n-th layer using the dual convolutional attention module. Finally, h n can also be fine-tuned through an optional 1×1 lightweight convolutional block and the LeakyReLU activation function to obtain the high-frequency component of the output image of the high-frequency layer of the n-th layer after fine-tuning . It is equivalent to using a 1×1 lightweight convolutional layer to perform operations of aggregating and compressing connections on the adjusted total output features, and optimizing feature extraction through a combination of multiple small modules. The grouped dilated feed-forward network ensures the balance between computational efficiency and performance.

[0113] Secondly, h n can be upsampled and used as the new input of the multi-block grouped dilated feed-forward network of the (n - 1)-th layer to update the enhanced input feature D of the (n - 1)-th layer n-1 . Thus, through this step-by-step upsampling method, a set of high-frequency component sets HF, HF = [h L-1 , …, h n+1 , h n , …, h2, h1] can be obtained; to match the high-frequency components of the output images of the high-frequency layers from the 1st layer to the (L - 1)-th layer.

[0114] Again, n in the above calculation formula is the specific layer number of the hierarchical image in step 1. The following is an example to illustrate the calculation formula

[0115] h L-1 = D L-1 LF L + D L-1 , perform matrix multiplication and bitwise addition operations on the input feature of the hierarchical image of the enhanced (L - 1)-th layer and the per-pixel mask of the output image of the low-frequency layer of the L-th layer to obtain the high-frequency component of the (L - 1)-th layer;

[0116] …,

[0117] h n = D n M n+1 + D n , perform matrix multiplication and bitwise addition operations on the input feature of the hierarchical image of the enhanced n-th layer and the per-pixel mask of the (n + 1)-th layer to obtain the high-frequency component of the n-th layer;

[0118] …,

[0119] h2 = D2 M3 + D2 uses the input features of the enhanced second-layer hierarchical image and performs matrix multiplication and bitwise addition operations with the per-pixel mask of the third layer to obtain the high-frequency component of the second layer;

[0120] h1 = D1 M2 + D1 uses the input features of the enhanced first-layer hierarchical image and performs matrix multiplication and bitwise addition operations with the per-pixel mask of the second layer to obtain the high-frequency component of the first layer.

[0121] In addition, where LF L is the per-pixel mask of the low-frequency component features after performing the enhancement operation of the double convolutional attention module on the low-frequency layer output image and the low-frequency layer input image of the L-th layer, and then performing the refinement operation using multiple grouped dilated feed-forward networks. By mainly introducing the transformed features of the low-frequency layer image, the high-frequency component is further refined, so as to achieve faithful reconstruction when manipulating the attributes of a specific domain.

[0122] Generally speaking, the output of the multiple grouped dilated feed-forward network is considered to be the per-pixel mask of the high-frequency component of the L - 1 layer. As Figure 5 shown, for the image pairs in the two domains, the high-frequency components of the same level only differ slightly in global brightness. Therefore, the per-pixel mask can be interpreted as a global adjustment and is relatively easier to optimize than the mixed-frequency image.

[0123] Since the addition of unprocessed high-frequency components may cause the re-degradation of the image. To alleviate this degradation effect, we first enhance the high-frequency component of the L - 1 layer through the double convolutional attention module to obtain D L-1 . Then, by introducing the per-pixel mask LF L of the low-frequency component features of the low-frequency layer image, and using the above calculation formula for calculation, after refinement operation, the high-frequency component h L-1 of the L - 1 layer is obtained.

[0124] Finally, it is fine-tuned through an optional lightweight convolutional block to obtain the refined high-frequency component. It is upsampled and used as the new input of the multiple grouped dilated feed-forward network to update the mask. Through this step-by-step upsampling method, we obtain a set of masks [M L−2 , …, M n , …, M2, M1] to match the remaining high-frequency components.

[0125] Use the same operation to refine all high-frequency components, and together with the low-frequency layer output image, obtain the result set to reconstruct the image J.

[0126] Finally, the enhanced underwater image J is obtained by applying Laplacian pyramid reconstruction to the refined component. It should be noted that the double convolutional attention module and the multi-block grouped dilated feed-forward network module in the L-1 layer do not share parameters with other layers.

[0127] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any respect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention.

Claims

1. An underwater image dynamic enhancement method based on a pyramid network, characterized in that, It includes the following steps: Step 1: First, decompose the initial input underwater image into a set of hierarchical images with a total of L layers using a Laplacian pyramid network. Among them, the L-th layer is the low-frequency layer input image, and the first layer to the (L - 1)-th layer is a set of high-frequency layer input images; Step 2: Use a deterministic color mapping network to perform color correction and illumination compensation on the low-frequency layer input image to convert it into a low-frequency layer output image; Step 3: Use the dual convolutional attention module to perform input feature extraction and enhancement operations on the input image of the nth layer, and obtain the input feature D after enhancing the input image n ; Step 4: Use a grouped dilated feedforward network to refine the output features of the output image of the (n + 1)-th layer, obtaining a per-pixel mask M after the output image refinement operation n+1 ; Step 5: Perform an arithmetic operation on the enhanced input features of the n-th layer input image in Step 3 and the per-pixel mask after the refinement operation of the (n + 1)-th layer output image in Step 4 to obtain the high-frequency component h of the high-frequency layer output image of the n-th layer n ; Step 6, repeat the operations of Step 3, Step 4, and Step 5. After repeating L - 1 times, a set of high-frequency component sets HF is obtained, HF = [h L-1 , …, h n+1 , h n , …, h2, h1]; Step 7: Use a Laplacian pyramid network to reconstruct the low-frequency layer output image and the high-frequency component set HF to obtain an enhanced output underwater image; In Step 2, the deterministic color mapping network embeds the low-frequency layer input image into a k-dimensional vector space through a projection matrix P(k, 3), performs color adjustment using a color mapping matrix T(k, k), and converts it back to the RGB color space through a projection matrix Q(3, k). The color mapping matrix T(k, k) is adaptively adjusted according to the characteristics of the input underwater image; In Step 3, the double convolutional attention module extracts the input features of the input image through two convolutional layers and performs non-linear activation using the LeakyReLU activation function; In Step 4, the grouped dilated feed-forward network includes multiple cascaded grouped dilated feed-forward networks, and performs feature refinement through residual connections and channel attention modules.

2. The underwater image dynamic enhancement method based on a pyramid network according to claim 1, wherein, In Step 2, the deterministic color mapping network is a convolutional neural network, and its conversion process includes the following steps: Step 2.1: Unfold the low-frequency layer input image into a two-dimensional matrix; Step 2.2: Perform matrix multiplication on the two-dimensional matrix using the projection matrix P(k, 3) to embed each pixel in the low-frequency layer input image into a k-dimensional vector space; Step 2.3: Perform matrix multiplication on the embedded k-dimensional vector space and the color mapping matrix T(k, k) to adjust the color; Step 2.4: Perform matrix multiplication on the adjusted vector space using the projection matrix Q(3, k), convert the adjusted vector space back to the RGB color space, and recombine each pixel into an image of the original size to output the color-enhanced low-frequency layer output image.

3. A dynamic underwater image enhancement method based on a pyramid network according to claim 1, characterized in that In Step 4, the refinement operation further includes using at least two grouped dilated feed-forward networks to form a multi-block grouped dilated feed-forward network, and respectively performing refinement adjustment on the output features of the output image, specifically including the following steps: Step 4.1: Use a multi-block grouped dilated feed-forward network for cascading in sequence and gradually adjust the output features of the output image to gradually correct the output features; Step 4.2: The corrected features fed into each grouped dilated feed-forward network are also densely connected to the output end of the last grouped dilated feed-forward network and perform © connection to form the total output features of the multi-block grouped dilated feed-forward network; Step 4.3: Then use a channel attention module to perform adaptive weight adjustment on the channel dimension of the total output features to obtain a per-pixel mask after the refinement operation on the output features.

4. A dynamic enhancement method for underwater images based on a pyramid network according to claim 1, characterized in that, In Step 4, the grouped dilated feed-forward network specifically includes the following method: Use two cascaded 3x3 convolutional layers and combine with the GELU activation function to capture the local spatial features of the output features of the output image; Meanwhile, use two cascaded 1x1 convolutional layers and combine with the GELU activation function to capture the local spatial features of the output features of the output image; Perform a bitwise addition operation on the two sets of local spatial features, and then perform a residual connection with the output features of the output image; Use an adaptive hybrid attention module for capture and form a captured 3x3 convolutional layer, and then perform a residual connection with the output features of the output image. Finally, obtain a per-pixel mask after refining the output features.

5. A dynamic enhancement method for underwater images based on a pyramid network according to claim 1, characterized in that In step 5, set the enhanced input feature of the n-th layer input image as D n ; set the per-pixel mask after the refinement operation of the (n + 1)-th layer output image as M n+1 ; through calculation, obtain the high-frequency component h of the high-frequency layer output image of the n-th layer n ; its specific calculation formula is: , Among them, is matrix multiplication operation, is bitwise addition operation.

6. The underwater image dynamic enhancement method based on a pyramid network according to claim 1, characterized in that, In step 4, the per-pixel mask of the low-frequency layer image of the L-th layer is obtained by the following steps: Use bilinear operation to upsample the low-frequency layer input image and the low-frequency layer output image respectively to improve their resolution, and ©-connect the upsampled low-frequency layer input image and the low-frequency layer output image to form [low-frequency layer input image, low-frequency layer output image]; Feed the [low-frequency layer input image, low-frequency layer output image] formed by connection into the double convolutional attention module to obtain the enhanced input feature D L ; Refine the enhanced input feature D using a grouped dilated feedforward network L to obtain the per-pixel mask LF of the low-frequency layer image of the L-th layer L .

7. An underwater image dynamic enhancement system based on a pyramid network, characterized in that, The system includes: A decomposition module for decomposing the initial input underwater image into a hierarchical image group of a total of L layers using a Laplacian pyramid network, where the L-th layer is the low-frequency layer input image, and the first layer to the L-1-th layer is a group of high-frequency layer input image groups; A conversion module for using a deterministic color mapping network to perform color correction and illumination compensation processing on the low-frequency layer input image to convert it into a low-frequency layer output image; An enhancement module for using a double convolutional attention module to perform extraction and enhancement operations on the input features of the input image of the n-th layer to obtain the input features after enhancing the input image; A refinement module for using a grouped dilated feed-forward network to perform refinement operations on the output features of the output image of the n+1-th layer to obtain a per-pixel mask after refining the output image; An operation module for performing an operation on the enhanced input features and the per-pixel mask after refinement operations to obtain the high-frequency component of the high-frequency layer output image of the n-th layer; A repetition module, which is used to perform repetitive operations on the enhancement module, the refinement module, and the arithmetic module. After repeating L - 1 times, a set of high - frequency component sets HF is obtained, HF = [h L-1 ,…,h n+1 ,h n ,…,h2,h1]; A reconstruction module for using a Laplacian pyramid network to reconstruct the low-frequency layer output image and the high-frequency component set HF to obtain an enhanced output underwater image; In the conversion module, the deterministic color mapping network embeds the low-frequency layer input image into a k-dimensional vector space through a projection matrix P(k, 3), performs color adjustment using a color mapping matrix T(k, k), and converts back to the RGB color space through a projection matrix Q(3, k). The color mapping matrix T(k, k) is adaptively adjusted according to the characteristics of the input underwater image; In the enhancement module, the double convolutional attention module extracts the input features of the input image through two convolutional layers and performs non-linear activation using the LeakyReLU activation function; In the refinement module, the grouped dilated feed-forward network includes multiple cascaded grouped dilated feed-forward networks, and performs feature refinement through residual connection and channel attention module.

8. A device, characterized in that, Includes: A memory for storing computer programs; A processor for executing the computer program and implementing an underwater image dynamic enhancement method based on a pyramid network as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a computer processor, an underwater image dynamic enhancement method based on a pyramid network as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Light source optimization method and optimization system based on physical model

    CN110851965A

  • Super-resolution underwater image enhancement method and system based on deep learning

    CN115034965A

  • Underwater image enhancement method of multi-attention mechanism guided by brightness mask

    CN116402715A

  • Low-illumination image enhancement method based on deep learning and Laplacian pyramid

    CN117196968A