Underwater image enhancement method and device, electronic equipment and medium
Underwater images are processed through encoder and decoder with U-Net network structure, and debuffering and adaptive convolution are used to solve the problem of underwater image blurring and improve image clarity.
Patent Information
- Application Number
- CN202510369091.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-26
AI Technical Summary
The existing underwater image enhancement scheme cannot effectively improve the blurring of underwater images, resulting in lower image clarity after enhancement.
Using the U-Net network structure, feature maps are extracted through the encoder network and downsampled, and deblurred processing is used to use the bottleneck layer to perform upsampling and adaptive convolution, and different processing is performed for high-frequency and low-frequency information to improve image clarity.
Through adaptive convolution processing, the degree of detail blurring of underwater images is improved and the clarity of the image is improved.
Smart Images

Figure CN120544015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an underwater image enhancement method, device, electronic equipment and medium. Background Art
[0002] Underwater images are generally images taken in water. Underwater images will be blurred. This is mainly because in an underwater environment, light reflected from objects and scattered forward will cause a blurring effect, while light scattered backward from non-target objects will enhance the blurriness of the image.
[0003] Existing underwater image enhancement solutions all perform global enhancement on underwater images, which does not significantly improve the blurriness of underwater images, resulting in relatively low clarity of the enhanced underwater images. Summary of the Invention
[0004] The present invention provides an underwater image enhancement method and device, which utilizes a bottleneck layer for deblurring and utilizes the grayscale variation degree represented by features in a fused feature map to process each feature of the underwater image, so that features with high grayscale variation and features with low variation are processed differently, thereby improving the problem of detail blurring in the underwater image and enhancing the clarity of the underwater image.
[0005] In a first aspect, an embodiment of the present invention provides an underwater image enhancement method, comprising:
[0006] The first step is to extract an image feature map from the underwater image and downsample the image feature map to obtain a low-dimensional feature map;
[0007] The second step is to perform deblurring on the low-dimensional feature map to obtain a bottleneck feature map;
[0008] In the third step, the bottleneck feature map is upsampled to obtain an enhanced feature map, the enhanced feature map is fused with the image feature map to obtain a fused feature map, the fused feature map is adaptively convolved to obtain an adaptive feature map, and the adaptive feature map is feature extracted to obtain an enhanced underwater image.
[0009] The above method uses an encoder network to extract feature maps and perform downsampling processing, uses a decoder network to extract features and perform upsampling processing, and fuses the feature maps extracted by the encoder network with the feature maps extracted by the decoder, performs adaptive convolution processing on the fused feature maps, and uses the adaptive convolution-processed map to perform restoration processing to obtain an enhanced underwater image. Since the adaptive convolution processing is for the entire image, high-frequency information and low-frequency information are processed differently, thereby improving the detail blur of the underwater image and improving the clarity of the underwater image.
[0010] In one possible implementation, performing adaptive convolution processing on the fused feature map to obtain an adaptive feature map includes:
[0011] Performing convolution processing on the fused feature map through multiple convolution layers containing small convolution kernels to obtain a convolved fused feature map;
[0012] The convolved fusion feature map is adaptively convolved by an adaptive convolution layer including a small convolution kernel to obtain an adaptive feature map; wherein the size of the small convolution kernel of the adaptive convolution layer is smaller than the small convolution kernel of the convolution layer.
[0013] The above method can extract more local features of the fused feature map through multiple convolution layers containing small convolution kernels, obtain an enhanced underwater image through more local features, and improve the clarity of the enhanced underwater image.
[0014] In one possible implementation, a plurality of convolutional layers including small convolutional kernels are sequentially arranged, and the convolutional layer including the small convolutional kernel includes the following structure: a small convolutional kernel and an activation function;
[0015] The fused feature map is convolved through multiple convolutional layers containing small convolution kernels, including:
[0016] For each convolution layer containing a small convolution kernel, the first target feature map is processed in turn by the small convolution kernel and activation function in the current convolution layer containing a small convolution kernel;
[0017] Wherein, if the current convolution layer containing the small convolution kernel is the convolution layer that performs convolution processing for the first time, then the first target feature map is the fused feature map;
[0018] If the current convolution layer containing the small convolution kernel is not the convolution layer that performs convolution processing for the first time, the first target feature map is the feature map output by the previous convolution layer containing the small convolution kernel.
[0019] The above method can process the feature map through multiple convolutional layers in sequence, extract the local features of the underwater image through a small convolution kernel in each convolutional layer, and use an activation function to enhance the expression ability. This enables the network to obtain more local features, and when the enhanced underwater image is subsequently formed, the clarity of the underwater image is improved.
[0020] In one possible implementation, the first step is implemented using an encoder network;
[0021] The encoder network includes a plurality of first feature extraction modules and a plurality of downsampling modules, wherein the first feature extraction modules correspond to the downsampling modules in a one-to-one manner;
[0022] For each first feature extraction module, extract a first spatial domain feature and a first frequency domain feature from the target feature map through the current first feature extraction module, fuse the first spatial domain feature with the first frequency domain feature to obtain an image feature map of the current first feature extraction module, and input the image feature map of the current first feature extraction module into the decoder network;
[0023] For each downsampling module, downsampling the image feature map of the first feature extraction module corresponding to the current downsampling module is performed by the current downsampling module to obtain a low-dimensional feature map of the current downsampling module;
[0024] Wherein, if the current first feature extraction module is the first feature extraction module that performs feature extraction for the first time, the target feature map is the underwater image;
[0025] If the current first feature extraction module is not the first feature extraction module that performs feature extraction for the first time, the target feature map is the low-dimensional feature map of the previous downsampling module;
[0026] If the current downsampling module is the downsampling module that performed the downsampling process for the last time, the low-dimensional feature map of the current downsampling module is input to the bottleneck layer.
[0027] The above method can extract spatial domain features and frequency domain features through the feature extraction module. The spatial domain features are pixel-level processing, and the frequency domain features are Fourier change processing. The combination of the two features can obtain more features of the underwater image, thereby performing further image enhancement processing based on the features, thereby improving the clarity of the enhanced underwater image.
[0028] In one possible implementation, the third step is implemented using a decoder network;
[0029] The decoder network includes a plurality of second feature extraction modules, a plurality of adaptive filtering modules and a plurality of upsampling modules; the second feature extraction modules, the adaptive filtering modules and the downsampling modules correspond one to one; the downsampling modules correspond one to one to the upsampling modules;
[0030] For each upsampling module, upsample the target bottleneck feature map through the current upsampling module to obtain an enhanced feature map with the same dimension as the image feature map of the first feature extraction module corresponding to the downsampling module corresponding to the current upsampling module, and fuse the enhanced feature map corresponding to the current upsampling module with the image feature map of the first feature extraction module corresponding to the current upsampling module to obtain a fused feature map;
[0031] For each adaptive filtering module, performing adaptive convolution processing on the fused feature map through the current adaptive filtering module to obtain an adaptive feature map;
[0032] For each second feature extraction module, extracting a second spatial domain feature and a second frequency domain feature from the adaptive feature map of the adaptive filtering module corresponding to the current second feature extraction module through the current second feature extraction module, and fusing the second spatial domain feature and the second frequency domain feature to obtain a feature map of the current second feature extraction module;
[0033] Wherein, if the current upsampling module is an upsampling module that performs upsampling processing for the first time, the target bottleneck feature map is the bottleneck feature map output by the bottleneck layer;
[0034] If the order of the current upsampling module is not the upsampling module that performs upsampling processing for the first time, the target low-dimensional feature map is the feature map of the second feature extraction module corresponding to the previous upsampling module;
[0035] If the current second feature extraction module is the second feature extraction module that performs feature extraction for the last time, the feature map of the current second feature extraction module is the enhanced underwater image.
[0036] The above method can extract spatial domain features and frequency domain features through the feature extraction module. The spatial domain features are pixel-level processing, and the frequency domain features are Fourier change processing. The combination of the two features can obtain more features of the underwater image, thereby performing further image enhancement processing based on the features, thereby improving the clarity of the enhanced underwater image.
[0037] In one possible implementation, the first feature extraction module and the second feature extraction module each include the following structure: two normalization layers, first to fourth convolutional layers, a depthwise convolutional layer, two activation functions, a channel attention layer, and two Fourier processing layers;
[0038] The feature map input to the feature extraction module is processed sequentially through the first normalization layer, the first convolution layer, the depth convolution layer, the first activation function, the channel attention layer, and the second convolution layer to obtain spatial domain features;
[0039] The feature map input to the feature extraction module is processed in sequence through the second normalization layer, the first Fourier processing layer, the third convolution layer, the second activation function, the fourth convolution layer, and the second Fourier processing layer to obtain frequency domain features;
[0040] Wherein, if the feature extraction module is the first feature extraction module, the spatial domain feature is the first spatial domain feature, and the frequency domain feature is the first frequency domain feature;
[0041] If the feature extraction module is the second feature extraction module, the spatial domain feature is the second spatial domain feature, and the frequency domain feature is the second frequency domain feature.
[0042] The above method can obtain more frequency domain features and spatial domain features through normalization layer, convolution layer, activation function, channel attention, and Fourier processing.
[0043] In one possible implementation, spatial domain features and frequency domain features are fused in the following way:
[0044] Adding the spatial domain features and the frequency domain features of the same channel; or
[0045] The spatial domain features and the frequency domain features are concatenated.
[0046] The above method can fuse the spatial domain features and frequency domain features by adding or splicing. The combined features can obtain more features of the underwater image, and further image enhancement processing can be performed based on the features, thereby improving the clarity of the enhanced underwater image.
[0047] In a second aspect, an embodiment of the present invention provides an underwater image enhancement device, comprising:
[0048] An encoding module is used to extract an image feature map from the underwater image and downsample the image feature map to obtain a low-dimensional feature map;
[0049] A bottleneck processing module, configured to perform a deblurring process on the low-dimensional feature map to obtain a bottleneck feature map;
[0050] A decoding module is used to upsample the bottleneck feature map to obtain an enhanced feature map, fuse the enhanced feature map with the image feature map to obtain a fused feature map, perform adaptive convolution on the fused feature map to obtain an adaptive feature map, and perform feature extraction on the adaptive feature map to obtain an enhanced underwater image.
[0051] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0052] processor;
[0053] A processor is used to execute the computer program or instructions in the memory so that the underwater image enhancement method as described in any one of the first aspects is performed.
[0054] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which, when instructions in the storage medium are executed by a processor, enables the processor to execute any underwater image enhancement method as described in the first aspect.
[0055] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the underwater image enhancement method as described in any one of the first aspects.
[0056] In addition, the technical effects brought about by any implementation method in the second to fifth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here.
[0057] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A flowchart of an underwater image enhancement method provided by an embodiment of the present invention;
[0059] Figure 2 A structural diagram of an image enhancement network provided by an embodiment of the present invention;
[0060] Figure 3 A structural diagram of another image enhancement network provided by an embodiment of the present invention;
[0061] Figure 4 A structural diagram of a feature extraction module provided by an embodiment of the present invention;
[0062] Figure 5 A structural diagram of an adaptive filtering module provided by an embodiment of the present invention;
[0063] Figure 6 A structural diagram of an underwater image enhancement device provided by an embodiment of the present invention;
[0064] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0066] In the description of the embodiments of the present application, unless otherwise specified, in the description of the embodiments of the present application, "plurality" refers to two or more than two.
[0067] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0068] Glossary:
[0069] Dual domains: These include the spatial domain and the frequency domain. The spatial domain, also known as the air domain, refers to the space composed of image pixels. Directly processing pixel values using length as an independent variable in image space is called spatial domain processing, i.e., pixel-level processing. The frequency domain, also known as the frequency domain, is the signal obtained by Fourier transforming the spatial domain. The Fourier transform of an image yields an image spectrum, representing the image's energy gradient.
[0070] Since underwater images are usually blurred, an embodiment of the present invention provides an underwater image enhancement solution. The solution can process each feature of the underwater image by utilizing the grayscale variation degree represented by the features in the fusion feature map, so that features with high grayscale variation and features with low variation are processed differently, thereby improving the blur degree of the underwater image and enhancing the clarity of the underwater image.
[0071] Among them, the features with high grayscale changes are high-frequency information in the image, and the features with low grayscale changes are low-frequency information in the image.
[0072] In detail, an embodiment of the present invention proposes a structure based on a U-Net network, establishes an image enhancement network, inputs an underwater image into the image enhancement network, the image enhancement network processes the underwater image, and the image enhancement network outputs an enhanced underwater image.
[0073] Among them, the image enhancement network includes an encoder network, a bottleneck layer, and a decoder network; when the image enhancement network processes underwater images, it is processed in sequence through the encoder network, the bottleneck layer, and the decoder network to output the enhanced underwater image.
[0074] The present invention is described in detail below with reference to the accompanying drawings:
[0075] Combine Figure 1 As shown, an embodiment of the present invention provides an underwater image enhancement method, comprising:
[0076] S100: extracting an image feature map from the underwater image, and performing downsampling processing on the image feature map to obtain a low-dimensional feature map;
[0077] S101: Deblurring the low-dimensional feature map to obtain a bottleneck feature map;
[0078] S102: Upsampling the bottleneck feature map to obtain an enhanced feature map, fusing the enhanced feature map with the image feature map to obtain a fused feature map, performing adaptive convolution on the fused feature map to obtain an adaptive feature map, and reconstructing the adaptive feature map to obtain an enhanced underwater image.
[0079] In some embodiments, step 100 is implemented by an encoder network, step 101 is implemented by a bottleneck layer, and step 102 is implemented by a decoder network. Figure 2 As shown, the underwater image is input into the image enhancement network. First, the encoder network 200 in the image enhancement network extracts an image feature map from the underwater image, and the image feature map is input into the decoder network 202 in the image enhancement network. At the same time, the image feature map is downsampled to obtain a low-dimensional feature map, and the low-dimensional feature map is input into the bottleneck layer 201. The bottleneck layer 201 performs feature extraction on the low-dimensional feature map to obtain a bottleneck feature map, and the bottleneck feature map is input into the decoder network 202. The decoder network 202 upsamples the bottleneck feature map to obtain an enhanced feature map, and the enhanced feature map and the image feature map are fused to obtain a fused feature map. The enhanced target image is obtained according to the fused feature map, and the enhanced target image is output.
[0080] From the above, it can be seen that in dealing with the problem of underwater image blur, the embodiment of the present invention first deblurs the low-dimensional feature map, and at the same time performs adaptive convolution processing on the fused feature map to obtain an adaptive feature map. Since the underwater image includes high-frequency information and low-frequency information, the adaptive convolution processing, that is, adaptive processing of each feature in the fused feature map, indicates that high-frequency information and low-frequency information are processed differently, thereby improving the degree of detail blur of the underwater image. The embodiment of the present invention uses two deblurring operations to improve the clarity of the enhanced underwater image.
[0081] The encoder network includes multiple first feature extraction modules and multiple downsampling modules; the first feature extraction modules and the downsampling modules correspond one to one; that is, one first feature extraction module and one downsampling module are a group, and the encoder network includes multiple groups, for example, combined Figure 3 As shown, there are 4 groups, namely the first feature extraction module 1, downsampling module 1, first feature extraction module 2, downsampling module 2, first feature extraction module 3, downsampling module 3, first feature extraction module 4, and downsampling module 4.
[0082] For each first feature extraction module, extract the first spatial domain feature and the first frequency domain feature from the target feature map through the current first feature extraction module, fuse the first spatial domain feature and the first frequency domain feature to obtain the image feature map of the current first feature extraction module, and input the image feature map of the current first feature extraction module into the decoder network;
[0083] For each downsampling module, downsampling the image feature map of the first feature extraction module corresponding to the current downsampling module is performed by the current downsampling module to obtain a low-dimensional feature map of the current downsampling module;
[0084] Wherein, if the current first feature extraction module is the first feature extraction module that performs feature extraction for the first time, the target feature map is the underwater image;
[0085] If the current first feature extraction module is not the first feature extraction module that performs feature extraction for the first time, the target feature map is the low-dimensional feature map of the previous downsampling module;
[0086] If the current downsampling module is the downsampling module that performed the downsampling process for the last time, the low-dimensional feature map of the current downsampling module is input to the bottleneck layer.
[0087] by Figure 3 For example, the underwater image is used as the input of the first feature extraction module 1. The first feature extraction module 1 extracts an image feature map from the underwater image, and the image feature map is input to the downsampling module 1. The downsampling module 1 downsamples the image feature map and outputs a low-dimensional feature map of the downsampling module 1.
[0088] The low-dimensional feature map of the downsampling module 1 is used as the input of the first feature extraction module 2. The first feature extraction module 2 extracts the image feature map from the low-dimensional feature map of the downsampling module 1. The image feature map of the first feature extraction module 2 is input to the downsampling module 2. The downsampling module 2 downsamples the image feature map of the first feature extraction module 2 and outputs the low-dimensional feature map of the downsampling module 2.
[0089] The low-dimensional feature map of the downsampling module 2 is used as the input of the first feature extraction module 3. The first feature extraction module 3 extracts the image feature map from the low-dimensional feature map of the downsampling module 2. The image feature map of the first feature extraction module 3 is input to the downsampling module 3. The downsampling module 3 downsamples the image feature map of the first feature extraction module 3 and outputs the low-dimensional feature map of the downsampling module 3.
[0090] The low-dimensional feature map of the downsampling module 3 is used as the input of the first feature extraction module 4. The first feature extraction module 4 extracts the image feature map from the low-dimensional feature map of the downsampling module 3. The image feature map of the first feature extraction module 4 is input to the downsampling module 4. The downsampling module 4 downsamples the image feature map of the first feature extraction module 4 and outputs the low-dimensional feature map of the downsampling module 4.
[0091] The first feature extraction module includes the following structure: two normalization layers, the first to fourth convolutional layers, a depthwise convolutional layer, two activation functions, a channel attention layer, and two Fourier processing layers;
[0092] The feature map input to the first feature extraction module is processed sequentially through the first normalization layer, the first convolution layer, the depth convolution layer, the first activation function, the channel attention layer, and the second convolution layer to obtain the first spatial domain feature;
[0093] The feature map input to the feature extraction module is processed in sequence through the second normalization layer, the first Fourier processing layer, the third convolution layer, the second activation function, the fourth convolution layer, and the second Fourier processing layer to obtain the first frequency domain feature.
[0094] In detail, combined with Figure 3 As shown, when the first feature extraction module is the first feature extraction module 2, the low-dimensional feature map of the downsampling module 1 is normalized by the first normalization layer, the feature map output by the first normalization layer is convolved by the first convolution layer, the feature map output by the first convolution layer is deep convolutionally processed, the feature map output by the deep convolution layer is nonlinearly processed by the first activation function, the feature map output by the deep convolution layer is weighted by the channel attention layer, the feature map output by the channel attention layer is convolved by the second convolution layer, and the second convolution layer outputs the first spatial domain feature. At the same time, the second normalization layer is used to normalize the low-dimensional feature map of the downsampling module 1, the first Fourier processing layer is used to perform Fourier processing on the feature map output by the first normalization layer, the third convolutional layer is used to convolve the feature map output by the first Fourier processing layer, the second activation function is used to perform nonlinear processing on the feature map output by the third convolutional layer, the fourth convolutional layer is used to convolve the feature map output by the second activation function, the second Fourier processing layer is used to perform Fourier processing on the feature map output by the fourth convolutional layer, and the second Fourier processing layer outputs the first frequency domain feature.
[0095] For example, combined Figure 4 As shown, the first convolution layer includes a 1×1 convolution kernel, the depth convolution layer includes a 3×3 depth convolution kernel, the first activation function is the GELU function, the channel attention layer includes SCA, and the second convolution layer includes a 1×1 convolution kernel.
[0096] For the route of obtaining the first spatial domain feature, layer normalization is used to standardize the distribution of each input feature map to reduce internal covariate shift, which helps the network converge faster and improve training stability; 1×1 convolution is used to increase the dimension of the input feature, increase the nonlinearity of the network, and help the network better learn the relationship between features; 3×3 depth convolution is used to extract higher-level feature representations from the input feature map. By stacking multiple convolutional layers, the network can gradually learn more abstract features, thereby improving its ability to understand images; the GELU activation function is a nonlinear function that has a smoother shape than ReLU, which helps alleviate the gradient vanishing problem and improve the network's expressive power; channel attention is used to improve the effect of image enhancement; the last 1×1 convolution of this branch is used to reduce the dimension of the output feature of this branch, which helps the network better learn the final feature representation and output it to the next layer of the network.
[0097] For the route of obtaining the first frequency domain features, layer normalization is used to standardize the distribution of each input feature map; a Fourier transform is applied to each channel, converting the existing spatial features to the frequency domain through the Fourier transform, performing an efficient global update based on the frequency data, and then converting this data back to its original spatial domain. In frequency domain operations, 1×1 convolution retains more features of the input image and integrates information from different channels. Compared with the convolution kernel, it has fewer parameters, avoids overfitting, and introduces more nonlinearities to improve generalization. GELU enhances the model's nonlinear transformation capabilities, thereby enhancing its expressiveness. The 1×1 convolution layer after GELU can interact and reorganize the information of the tensor through the GELU function to further enhance the model's expressiveness. Since the Fourier transform processes complex numbers, it is crucial to ensure that the input and output of this module are real numbers for compatibility with other neural modules.
[0098] The first spatial domain feature and the first frequency domain feature are fused in the following manner:
[0099] Adding the features of the same channel of the first spatial domain feature and the first frequency domain feature; or
[0100] The first spatial domain feature and the first frequency domain feature are concatenated.
[0101] In detail, the first spatial domain feature may be spliced after the first frequency domain feature, or the first frequency domain feature may be spliced after the first spatial domain feature.
[0102] The bottleneck layer is implemented using a nonlinear non-activated network NAFNet. In order to strike a balance between enhancing the image network size and enhancing the image network performance, the bottleneck layer proposed in the embodiment of the present invention uses the component modules of the basic network of NAFNet as the bottleneck layer modules, and the number of them is set to 4. Figure 3 As shown in Figure 2, the NAF module uses the basic network components of NAFNet to deblur the low-dimensional feature map.
[0103] The decoder network includes multiple second feature extraction modules, multiple adaptive filtering modules and multiple up-sampling modules; the second feature extraction modules, the adaptive filtering modules and the down-sampling modules correspond one to one; the down-sampling modules correspond one to one to the up-sampling modules; that is, a second feature extraction module, an adaptive filtering module and an up-sampling module are a group, and the decoder network includes multiple groups, for example, combined with Figure 3 As shown, there are 4 groups, namely the second feature extraction module 11, the upsampling module 12, the adaptive filtering module 13, the second feature extraction module 21, the upsampling module 22, the adaptive filtering module 23, the second feature extraction module 31, the upsampling module 32, the adaptive filtering module 33, the second feature extraction module 41, the upsampling module 42, and the adaptive filtering module 43.
[0104] For each upsampling module, upsample the target bottleneck feature map through the current upsampling module to obtain an enhanced feature map with the same dimension as the image feature map of the first feature extraction module corresponding to the downsampling module corresponding to the current upsampling module, and fuse the enhanced feature map corresponding to the current upsampling module with the image feature map of the first feature extraction module corresponding to the current upsampling module to obtain a fused feature map;
[0105] For each adaptive filtering module, adaptive convolution processing is performed on the fusion feature map through the current adaptive filtering module to obtain an adaptive feature map;
[0106] For each second feature extraction module, extracting the second spatial domain feature and the second frequency domain feature from the adaptive feature map of the adaptive filtering module corresponding to the current second feature extraction module through the current second feature extraction module, and fusing the second spatial domain feature and the second frequency domain feature to obtain a feature map of the current second feature extraction module;
[0107] Wherein, if the current upsampling module is an upsampling module that performs upsampling processing for the first time, the target bottleneck feature map is the bottleneck feature map output by the bottleneck layer;
[0108] If the order of the current upsampling module is not the upsampling module that performs upsampling processing for the first time, the target low-dimensional feature map is the feature map of the second feature extraction module corresponding to the previous upsampling module;
[0109] If the current second feature extraction module is the second feature extraction module that performs feature extraction for the last time, the feature map of the current second feature extraction module is the enhanced underwater image.
[0110] by Figure 3 Taking the bottleneck feature map as the input of the upsampling module 12, the bottleneck feature map is upsampled to obtain an enhanced feature map of the upsampling module 12, the image feature map output by the first feature extraction module 4 and the enhanced feature map of the upsampling module 12 are fused to obtain a fused feature map, the fused feature map is used as the input of the adaptive filtering module 13, the adaptive filtering module 13 performs adaptive convolution processing on the fused feature map to obtain an adaptive feature map of the adaptive filtering module 13, the second feature extraction module 11 extracts the second spatial domain feature and the second frequency domain feature from the adaptive feature map of the adaptive filtering module 13, and fuses the second spatial domain feature and the second frequency domain feature to obtain the feature map of the second feature extraction module 11.
[0111] The feature map of the second feature extraction module 11 is used as the input of the upsampling module 22, and the feature map of the second feature extraction module 11 is upsampled to obtain an enhanced feature map of the upsampling module 22. The image feature map output by the first feature extraction module 3 and the enhanced feature map of the upsampling module 22 are fused to obtain a fused feature map. The fused feature map is used as the input of the adaptive filtering module 23, and the adaptive filtering module 23 performs adaptive convolution processing on the fused feature map to obtain an adaptive feature map of the adaptive filtering module 23. The second feature extraction module 21 extracts the second spatial domain feature and the second frequency domain feature from the adaptive feature map of the adaptive filtering module 23, and fuses the second spatial domain feature and the second frequency domain feature to obtain the feature map of the second feature extraction module 21.
[0112] The feature map of the second feature extraction module 21 is used as the input of the upsampling module 32, and the feature map of the second feature extraction module 21 is upsampled to obtain an enhanced feature map of the upsampling module 32. The image feature map output by the first feature extraction module 2 and the enhanced feature map of the upsampling module 32 are fused to obtain a fused feature map. The fused feature map is used as the input of the adaptive filtering module 33, and the adaptive filtering module 33 performs adaptive convolution processing on the fused feature map to obtain an adaptive feature map of the adaptive filtering module 33. The second feature extraction module 31 extracts the second spatial domain feature and the second frequency domain feature from the adaptive feature map of the adaptive filtering module 33, and fuses the second spatial domain feature and the second frequency domain feature to obtain the feature map of the second feature extraction module 31.
[0113] The feature map of the second feature extraction module 31 is used as the input of the upsampling module 42, and the feature map of the second feature extraction module 31 is upsampled to obtain an enhanced feature map of the upsampling module 42. The image feature map output by the first feature extraction module 1 and the enhanced feature map of the upsampling module 42 are fused to obtain a fused feature map. The fused feature map is used as the input of the adaptive filtering module 43, and the adaptive filtering module 43 performs adaptive convolution processing on the fused feature map to obtain an adaptive feature map of the adaptive filtering module 43. The second feature extraction module 41 extracts the second spatial domain feature and the second frequency domain feature from the adaptive feature map of the adaptive filtering module 43, and fuses the second spatial domain feature and the second frequency domain feature to obtain the feature map of the second feature extraction module 41, which is the enhanced underwater image.
[0114] The second feature extraction module includes the following structure: two normalization layers, the first to fourth convolutional layers, a depthwise convolutional layer, two activation functions, a channel attention layer, and two Fourier processing layers;
[0115] The feature map input to the second feature extraction module is processed in sequence through the first normalization layer, the first convolution layer, the depth convolution layer, the first activation function, the channel attention layer, and the second convolution layer to obtain the second spatial domain feature;
[0116] The feature map input to the second feature extraction module is processed in sequence through the second normalization layer, the first Fourier processing layer, the third convolution layer, the second activation function, the fourth convolution layer, and the second Fourier processing layer to obtain the second frequency domain feature.
[0117] The second spatial domain feature is fused with the second frequency domain feature in the following way:
[0118] Adding the features of the same channel of the second spatial domain feature and the second frequency domain feature; or
[0119] The second spatial domain feature and the second frequency domain feature are concatenated.
[0120] The specific implementation method is similar to that of the first feature extraction module, and you can refer to the part of the first feature extraction module.
[0121] In some embodiments, performing adaptive convolution processing on the fused feature map to obtain an adaptive feature map includes:
[0122] The fused feature map is convolved through multiple convolution layers containing small convolution kernels to obtain a convolved fused feature map;
[0123] The convolved fusion feature map is adaptively convolved by an adaptive convolution layer including a small convolution kernel to obtain an adaptive feature map; wherein the size of the small convolution kernel of the adaptive convolution layer is smaller than the small convolution kernel of the convolution layer.
[0124] Wherein, a plurality of convolutional layers containing small convolutional kernels are sequentially arranged, and the convolutional layer containing the small convolutional kernel includes the following structure: a small convolutional kernel and an activation function;
[0125] The fused feature map is convolved through multiple convolutional layers containing small convolution kernels, including:
[0126] For each convolution layer containing a small convolution kernel, the first target feature map is processed in turn by the small convolution kernel and activation function in the current convolution layer containing a small convolution kernel;
[0127] Wherein, if the current convolution layer containing the small convolution kernel is the convolution layer that performs convolution processing for the first time, then the first target feature map is the fused feature map;
[0128] If the current convolution layer containing the small convolution kernel is not the convolution layer that performs convolution processing for the first time, the first target feature map is the feature map output by the previous convolution layer containing the small convolution kernel.
[0129] In detail, the adaptive filtering module first convolves the fused feature map through multiple convolution layers containing small convolution kernels, determines the size of the convolution kernel corresponding to the features in the convolved fused feature map, and convolves the features in the convolved fused feature map according to the size of the convolution kernel corresponding to the features in the convolutional fused feature map to obtain an adaptive feature map.
[0130] Exemplarily, the filter adaptive convolutional (FAC) layer applies a dynamic convolution filter to each element in the feature. Building on the FAC layer, embodiments of the present invention propose a filter adaptive convolutional module (FACM) that leverages the enhanced information from the encoder network to address image blur.
[0131] like Figure 5 As shown, given the enhanced features F of different scales en ∈R H×W×C , the adaptive filtering module estimates the corresponding filter through three groups of 3×3 convolution layers followed by GELU functions and the last layer of 1×1 convolution kernel. Expand the feature dimension. Then, the FAC layer uses the filter f to decode the feature F de ∈R H×W×C For the feature F deFor each element of , FAC uses the corresponding d×d kernel (d represents the kernel size) in filter f to perform a convolution operator to obtain refined features. The final layer, composed of 1×1 convolution kernels, is an adaptive convolution layer that adaptively processes each feature in the feature map and combines the adaptively processed features into the final feature map.
[0132] Using multiple 3×3 convolutional layers to stack can effectively increase the depth of the network while expanding its receptive field, avoiding the use of large-size convolution kernels at the beginning, and allowing the network to gradually integrate a wider range of contextual information through multi-layer processing while maintaining sensitivity to details. In order to take into account computational complexity, model size and network performance, the kernel size is set to 3 at different scales in the embodiment of the present invention. Using a 3×3 convolution kernel can maintain a smaller receptive field, allowing the network to capture local features more finely. For certain tasks, such as image processing with rich details or complex textures, a smaller receptive field helps to maintain the clarity and accuracy of image details. Smaller convolution kernels can better process high-frequency details in images, such as fine textures and edges. This is because the 3×3 convolution kernel emphasizes the relationship between neighboring pixels and is very effective in capturing local feature changes in images.
[0133] The embodiment of the present invention is an image enhancement network built based on the U-Net network. The encoder and decoder in the U-Net network are significantly modified. It mainly includes two components: the feature extraction module and the adaptive convolution module. The network framework diagram is shown in Figure 3 . The embodiment of the present invention uses the feature extraction module as the encoder and decoder modules in the basic network, and the adaptive filtering module is applied after the skip connection. The entire network downsamples 4 times and upsamples 4 times. The size of the input image is H×W×3, which are height, width, and number of channels respectively. The network includes 4 skip connections, which skip some layers to directly connect the feature maps of specific layers of the downsampling path and the upsampling path. The skip connection combines the feature map with the feature map in the corresponding downsampling path to retain the detail information that may be lost during the upsampling process.
[0134] After the feature map is input into the network, the number of channels is first increased to 32 using 1×1 convolution, and then the number of channels is multiplied by 2 each time the feature map is downsampled. When the feature map reaches the bottleneck layer, the number of channels reaches 512. After passing through the basic modules of the 4 NAFNet networks, the feature map is upsampled. Each upsampling retains half of the original number of channels until the number of channels of the feature map is 32. Finally, the number of channels is restored to 3 using 1×1 convolution, and the resulting image is output, which has the same size as the input image, H×W×3. The embodiment of the present invention extracts the contextual features of the image through downsampling, and restores the detailed texture of the image through upsampling and jump connections, so that in the underwater image enhancement task, it can retain the precise edge position of the target and grasp the overall shape of the target.
[0135] Before using the image enhancement network to perform image enhancement processing on underwater images, the image enhancement network needs to be trained. During training, a training set is used, which includes multiple underwater images and corresponding reference images. During the training process, a loss function is used to adjust the parameters in the image enhancement network so that the enhanced image is closer to the reference image, while reducing the network training time and promoting the network to achieve better training results. In the embodiment of the present invention, a method combining Charbonnier loss (L Char ), perceptual loss (L Per ) and frequency reconstruction loss (L FR ) guides the network to learn more accurate image reconstruction results in both the spatial and frequency domains, thereby improving the performance and effectiveness of the underwater image enhancement network. See the following formula for details.
[0136] L=L Char +λ1L Per +λ2L FR
[0137] Among them, λ1 and λ2 are set to 0.04 and 0.01 respectively.
[0138] (1) Charbonnier loss
[0139] The embodiment of the present invention adopts Charbonnier loss (L Char ), which provides pixel-space regularization to ensure that the reconstructed image does not differ too much from the corresponding reference image. The formula of Charbonnier loss is as follows:
[0140]
[0141] Where ε is a constant value of 10 -3 .
[0142] (2) Perceptual loss
[0143] Referring to the VGG network, the embodiment of the present invention uses the perceptual loss (L Per ). To make it sensitive to color and semantics, the embodiment of the present invention selects Layer3_3 of VGG-16. The perceptual loss quantifies the difference between the feature representations of the predicted image and the reference image, and its formula is as follows:
[0144]
[0145] Where φ represents the j-th convolutional layer. The size of the j-th layer feature map in the VGG-16 network is represented by C j H j W j, where C j , H j and W j Represent the number, height and width of feature maps respectively.
[0146] (3) Frequency reconstruction loss
[0147] In the frequency domain, the frequency reconstruction loss (L FR ) to calculate the L1 loss, the formula is as follows:
[0148]
[0149] Where F represents the fast Fourier transform (FFT).
[0150] The deep learning underwater image enhancement method based on dual-domain information fusion is a U-Net-based method that combines a space-frequency dual-branch module and an adaptive convolution module. The space-frequency dual-branch module processes spatial and frequency domain features simultaneously, captures information in underwater images more comprehensively, improves the enhancement effect, and effectively improves image distortion, color cast, low contrast and other problems. At the same time, it enriches image details by processing high-frequency components. The adaptive filtering module aims to adaptively adjust the convolution kernel to adapt to the characteristics of different underwater images, helping the network to better adapt to underwater images in different scenes and conditions, improve the robustness and generalization ability of the enhancement effect, and enhance the deblurring ability of the network. The embodiment of the present invention has achieved significant improvements in color restoration and enhancement while retaining key details of underwater images. Comparative experimental results verify the superiority of this method in processing degraded underwater images.
[0151] like Figure 6 As shown, the present invention also provides an underwater image enhancement device, comprising:
[0152] The encoding module 600 is used to extract an image feature map from the underwater image and perform downsampling processing on the image feature map to obtain a low-dimensional feature map;
[0153] A bottleneck processing module 601 is used to perform a deblurring process on the low-dimensional feature map to obtain a bottleneck feature map;
[0154] The decoding module 602 is used to upsample the bottleneck feature map to obtain an enhanced feature map, fuse the enhanced feature map with the image feature map to obtain a fused feature map, perform adaptive convolution on the fused feature map to obtain an adaptive feature map, and reconstruct the adaptive feature map to obtain an enhanced underwater image.
[0155] Optionally, the decoding module 602 is further configured to: perform convolution processing on the fused feature map through a plurality of convolution layers including small convolution kernels to obtain a convolved fused feature map;
[0156] The convolved fusion feature map is adaptively convolved by an adaptive convolution layer including a small convolution kernel to obtain an adaptive feature map; wherein the size of the small convolution kernel of the adaptive convolution layer is smaller than the small convolution kernel of the convolution layer.
[0157] Optionally, multiple convolution layers containing small convolution kernels are sequentially arranged, and the convolution layer containing the small convolution kernel includes the following structure: a small convolution kernel and an activation function;
[0158] The decoding module 602 is further configured to: for each convolution layer including a small convolution kernel, sequentially process the first target feature map through the small convolution kernel and activation function in the current convolution layer including the small convolution kernel;
[0159] Wherein, if the current convolution layer containing the small convolution kernel is the convolution layer that performs convolution processing for the first time, then the first target feature map is the fused feature map;
[0160] If the current convolution layer containing the small convolution kernel is not the convolution layer that performs convolution processing for the first time, the first target feature map is the feature map output by the previous convolution layer containing the small convolution kernel.
[0161] Optionally, the encoder network includes a plurality of first feature extraction modules and a plurality of downsampling modules; the first feature extraction modules and the downsampling modules correspond one to one;
[0162] The encoding module 600 is specifically configured to:
[0163] For each first feature extraction module, extract a first spatial domain feature and a first frequency domain feature from the target feature map through the current first feature extraction module, fuse the first spatial domain feature with the first frequency domain feature to obtain an image feature map of the current first feature extraction module, and input the image feature map of the current first feature extraction module into the decoder network;
[0164] For each downsampling module, downsampling the image feature map of the first feature extraction module corresponding to the current downsampling module is performed by the current downsampling module to obtain a low-dimensional feature map of the current downsampling module;
[0165] Wherein, if the current first feature extraction module is the first feature extraction module that performs feature extraction for the first time, the target feature map is the underwater image;
[0166] If the current first feature extraction module is not the first feature extraction module that performs feature extraction for the first time, the target feature map is the low-dimensional feature map of the previous downsampling module;
[0167] If the current downsampling module is the downsampling module that performed the downsampling process for the last time, the low-dimensional feature map of the current downsampling module is input to the bottleneck layer.
[0168] Optionally, the decoder network includes multiple second feature extraction modules, multiple adaptive filtering modules and multiple upsampling modules; the second feature extraction modules, the adaptive filtering modules and the downsampling modules correspond one to one; the downsampling modules correspond one to one to the upsampling modules;
[0169] The decoding module 602 is specifically used for:
[0170] For each upsampling module, upsample the target bottleneck feature map through the current upsampling module to obtain an enhanced feature map with the same dimension as the image feature map of the first feature extraction module corresponding to the downsampling module corresponding to the current upsampling module, and fuse the enhanced feature map corresponding to the current upsampling module with the image feature map of the first feature extraction module corresponding to the current upsampling module to obtain a fused feature map;
[0171] For each adaptive filtering module, performing adaptive convolution processing on the fused feature map through the current adaptive filtering module to obtain an adaptive feature map;
[0172] For each second feature extraction module, extracting a second spatial domain feature and a second frequency domain feature from the adaptive feature map of the adaptive filtering module corresponding to the current second feature extraction module through the current second feature extraction module, and fusing the second spatial domain feature and the second frequency domain feature to obtain a feature map of the current second feature extraction module;
[0173] Wherein, if the current upsampling module is an upsampling module that performs upsampling processing for the first time, the target bottleneck feature map is the bottleneck feature map output by the bottleneck layer;
[0174] If the order of the current upsampling module is not the upsampling module that performs upsampling processing for the first time, the target low-dimensional feature map is the feature map of the second feature extraction module corresponding to the previous upsampling module;
[0175] If the current second feature extraction module is the second feature extraction module that performs feature extraction for the last time, the feature map of the current second feature extraction module is the enhanced underwater image.
[0176] Optionally, the first feature extraction module and the second feature extraction module both include the following structure: two normalization layers, first to fourth convolution layers, a depthwise convolution layer, two activation functions, a channel attention layer, and two Fourier processing layers;
[0177] The feature map input to the feature extraction module is processed sequentially through the first normalization layer, the first convolution layer, the depth convolution layer, the first activation function, the channel attention layer, and the second convolution layer to obtain spatial domain features;
[0178] The feature map input to the feature extraction module is processed in sequence through the second normalization layer, the first Fourier processing layer, the third convolution layer, the second activation function, the fourth convolution layer, and the second Fourier processing layer to obtain frequency domain features;
[0179] Wherein, if the feature extraction module is the first feature extraction module, the spatial domain feature is the first spatial domain feature, and the frequency domain feature is the first frequency domain feature;
[0180] If the feature extraction module is the second feature extraction module, the spatial domain feature is the second spatial domain feature, and the frequency domain feature is the second frequency domain feature.
[0181] Optionally, the encoding module 600 and the decoding module 602 are further configured to:
[0182] Adding the spatial domain features and the frequency domain features of the same channel; or
[0183] The spatial domain features and the frequency domain features are concatenated.
[0184] In addition, combined Figures 1-6 The underwater image enhancement method and apparatus described in the embodiments of the present invention may be implemented by an electronic device.
[0185] Electronic devices, including: processors;
[0186] a memory for storing instructions executable by the processor;
[0187] The processor is configured to execute the instructions to implement the underwater image enhancement method as described in any one of the above descriptions.
[0188] Based on the above introduction, for example, Figure 7 electronic equipment structure.
[0189] The electronic device may include a processor 710 and a memory 720 storing computer program instructions.
[0190] Specifically, the processor 710 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0191] The memory 720 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 720 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 720 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 720 may be inside or outside the data processing device. In a specific embodiment, the memory 720 is a non-volatile solid-state memory. In a specific embodiment, the memory 720 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0192] The processor 710 implements any one of the task execution methods in the above embodiments by reading and executing computer program instructions stored in the memory 720 .
[0193] In one example, the electronic device may further include a communication interface 730 and a bus 740. Figure 7 As shown, the processor 710 , the memory 720 , and the communication interface 730 are connected via a bus 740 and communicate with each other.
[0194] The communication interface 730 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present invention.
[0195] Bus 740 comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 740 can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.
[0196] The electronic device can execute the underwater image enhancement method in the embodiment of the present invention based on the received task, thereby realizing the combination of Figure 1-Figure 7 The invention describes a method and apparatus for underwater image enhancement.
[0197] In addition, in combination with the electronic device in the above embodiments, an embodiment of the present invention may provide a storage medium, which, when the instructions in the storage medium are executed by the processor of the electronic device, enables the electronic device to execute the underwater image enhancement method as described in any one of the above items.
[0198] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0199] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0201] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0202] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for underwater image enhancement, characterized in that: include: The first step is to extract an image feature map from the underwater image and downsample the image feature map to obtain a low-dimensional feature map; The second step is to perform deblurring on the low-dimensional feature map to obtain a bottleneck feature map; In the third step, the bottleneck feature map is upsampled to obtain an enhanced feature map, the enhanced feature map is fused with the image feature map to obtain a fused feature map, the fused feature map is adaptively convolved to obtain an adaptive feature map, and the adaptive feature map is reconstructed to obtain an enhanced underwater image.
2. The underwater image enhancement method according to claim 1, characterized in that: Performing adaptive convolution processing on the fused feature map to obtain an adaptive feature map, including: Performing convolution processing on the fused feature map through multiple convolution layers containing small convolution kernels to obtain a convolved fused feature map; The convolved fusion feature map is adaptively convolved by an adaptive convolution layer including a small convolution kernel to obtain an adaptive feature map; wherein the size of the small convolution kernel of the adaptive convolution layer is smaller than the small convolution kernel of the convolution layer.
3. The underwater image enhancement method according to claim 2, characterized in that: Multiple convolutional layers containing small convolutional kernels are set sequentially, and the convolutional layer containing small convolutional kernels includes the following structure: a small convolutional kernel and an activation function; The fused feature map is convolved through multiple convolutional layers containing small convolution kernels, including: For each convolution layer containing a small convolution kernel, the first target feature map is processed in turn by the small convolution kernel and activation function in the current convolution layer containing a small convolution kernel; Wherein, if the current convolution layer containing the small convolution kernel is the convolution layer that performs convolution processing for the first time, then the first target feature map is the fused feature map; If the current convolution layer containing the small convolution kernel is not the convolution layer that performs convolution processing for the first time, the first target feature map is the feature map output by the previous convolution layer containing the small convolution kernel.
4. The underwater image enhancement method according to any one of claims 1 to 3, characterized in that: The first step is implemented using an encoder network; The encoder network includes a plurality of first feature extraction modules and a plurality of downsampling modules, wherein the first feature extraction modules correspond to the downsampling modules in a one-to-one manner; For each first feature extraction module, extract a first spatial domain feature and a first frequency domain feature from the target feature map through the current first feature extraction module, fuse the first spatial domain feature with the first frequency domain feature to obtain an image feature map of the current first feature extraction module, and input the image feature map of the current first feature extraction module into the decoder network; For each downsampling module, downsampling the image feature map of the first feature extraction module corresponding to the current downsampling module is performed by the current downsampling module to obtain a low-dimensional feature map of the current downsampling module; Wherein, if the current first feature extraction module is the first feature extraction module that performs feature extraction for the first time, the target feature map is the underwater image; If the current first feature extraction module is not the first feature extraction module that performs feature extraction for the first time, the target feature map is the low-dimensional feature map of the previous downsampling module; If the current downsampling module is the downsampling module that performed the downsampling process for the last time, the low-dimensional feature map of the current downsampling module is input to the bottleneck layer.
5. The underwater image enhancement method according to claim 4, characterized in that: The third step is implemented using a decoder network; The decoder network includes a plurality of second feature extraction modules, a plurality of adaptive filtering modules and a plurality of upsampling modules; the second feature extraction modules, the adaptive filtering modules and the downsampling modules correspond one to one; the downsampling modules correspond one to one to the upsampling modules; For each upsampling module, upsample the target bottleneck feature map through the current upsampling module to obtain an enhanced feature map with the same dimension as the image feature map of the first feature extraction module corresponding to the downsampling module corresponding to the current upsampling module, and fuse the enhanced feature map corresponding to the current upsampling module with the image feature map of the first feature extraction module corresponding to the current upsampling module to obtain a fused feature map; For each adaptive filtering module, performing adaptive convolution processing on the fused feature map through the current adaptive filtering module to obtain an adaptive feature map; For each second feature extraction module, extracting a second spatial domain feature and a second frequency domain feature from the adaptive feature map of the adaptive filtering module corresponding to the current second feature extraction module through the current second feature extraction module, and fusing the second spatial domain feature and the second frequency domain feature to obtain a feature map of the current second feature extraction module; Wherein, if the current upsampling module is an upsampling module that performs upsampling processing for the first time, the target bottleneck feature map is the bottleneck feature map output by the bottleneck layer; If the order of the current upsampling module is not the upsampling module that performs upsampling processing for the first time, the target low-dimensional feature map is the feature map of the second feature extraction module corresponding to the previous upsampling module; If the current second feature extraction module is the second feature extraction module that performs feature extraction for the last time, the feature map of the current second feature extraction module is the enhanced underwater image.
6. The underwater image enhancement method according to claim 5, characterized in that: The first feature extraction module and the second feature extraction module each include the following structures: two normalization layers, the first to fourth convolutional layers, a depthwise convolutional layer, two activation functions, a channel attention layer, and two Fourier processing layers; The feature map input to the feature extraction module is processed sequentially through the first normalization layer, the first convolution layer, the depth convolution layer, the first activation function, the channel attention layer, and the second convolution layer to obtain spatial domain features; The feature map input to the feature extraction module is processed in sequence through the second normalization layer, the first Fourier processing layer, the third convolution layer, the second activation function, the fourth convolution layer, and the second Fourier processing layer to obtain frequency domain features; Wherein, if the feature extraction module is the first feature extraction module, the spatial domain feature is the first spatial domain feature, and the frequency domain feature is the first frequency domain feature; If the feature extraction module is the second feature extraction module, the spatial domain feature is the second spatial domain feature, and the frequency domain feature is the second frequency domain feature.
7. The underwater image enhancement method according to claim 6, characterized in that: The spatial domain features and frequency domain features are fused in the following ways: Adding the spatial domain features and the frequency domain features of the same channel; or The spatial domain features and the frequency domain features are concatenated.
8. An underwater image enhancement device, characterized in that: include: An encoding module is used to extract an image feature map from the underwater image and downsample the image feature map to obtain a low-dimensional feature map; A bottleneck processing module, configured to perform a deblurring process on the low-dimensional feature map to obtain a bottleneck feature map; A decoding module is used to upsample the bottleneck feature map to obtain an enhanced feature map, fuse the enhanced feature map with the image feature map to obtain a fused feature map, perform adaptive convolution on the fused feature map to obtain an adaptive feature map, and reconstruct the adaptive feature map to obtain an enhanced underwater image.
9. An electronic device, characterized in that: include: Memory, used to store computer programs or instructions; A processor is configured to execute the computer program or instructions in the memory so that the underwater image enhancement method according to any one of claims 1 to 7 is performed.
10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor, the processor is enabled to execute the underwater image enhancement method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Underwater image enhancement method based on adaptive multi-scale fusion and attention mechanism
CN117314787A
Cited By
Underwater image enhancement and restoration method based on double-domain adaptive fusion module
CN122023163A