Lightweight underwater image enhancement method and system based on deep learning

CN122492515BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610923993.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

但现有基于全通道轴向深度卷积的轻量级水下图像增强网络仍存在以下问题:其一,对所有通道执行相同形式的轴向空间建模,忽略不同通道对水下退化的响应差异,在低响应通道上产生冗余计算;其二,瓶颈层通常采用单尺度固定感受野,难以同时建模全局色偏、局部纹理模糊和对比度降低等多尺度退化特征;其三,跳跃连接通常直接拼接编码器浅层特征,未对水下环境中通道间非均匀光衰减造成的浅层特征偏差进行校正,影响解码阶段的边缘和纹理恢复质量

Benefits of technology

[0021]1、本基于深度学习的轻量级水下图像增强方法通过部分轴向深度卷积模块仅对部分通道执行轴向深度卷积,减少全通道空间建模带来的冗余计算,在保持空间上下文提取能力的同时降低模型参数量和计算量;通过多尺度轴向自适应空间注意力瓶颈模块对不同感受野下的退化特征进行并行提取和注意力加权,提高对全局色偏、局部纹理模糊和对比度下降的联合建模能力;通过通道空间注意力模块校正跳跃连接浅层特征,提高解码阶段的边缘和纹理恢复质量;通过多约束联合损失函数协同约束像素、结构、感知和颜色分布,使增强图像在色彩自然度、对比度和细节保真方面取得平衡;所述网络的参数量和计算量较低,适合在资源受限的水下作业平台中进行实时图像增强。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492515B_ABST
    Figure CN122492515B_ABST
Patent Text Reader

Abstract

The application discloses a kind of light weight underwater image enhancement method and system based on deep learning, comprising: S1, shallow feature is extracted;S2, shallow feature is gradually abstracted, and multiple layer coding features and deep features are obtained;S3, deep feature is modeled by multi-scale degenerative feature, and bottleneck enhancement feature is obtained;S4, multiple layer coding features are adaptively recalibrated, and recalibrated jump connection features are obtained;S5, bottleneck enhancement feature is gradually upsampled to obtain upsampled feature, and upsampled feature is spliced and fused with corresponding level recalibrated jump connection feature, to obtain fusion decoding feature;S6, fusion decoding feature is mapped to RGB space by reconstruction layer, to generate enhanced underwater image, while ensuring image quality, the application realizes real-time enhancement of underwater image under the condition of lower parameter amount and lower calculation amount, can reduce model parameter amount and calculation amount, improve the real-time image enhancement efficiency of resource-constrained underwater platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image enhancement technology, specifically to a lightweight underwater image enhancement method and system based on deep learning. Background Technology

[0002] During underwater optical imaging, light is affected by both absorption and scattering as it propagates through water. Different wavelengths of light attenuate at different rates; the red component decays rapidly with increasing distance, while the blue-green component is relatively more preserved, often resulting in a blue-green tint in underwater images. Suspended particles in the water cause backscattering, leading to image blurring, reduced contrast, blurred edges, and unclear texture details. These degradation phenomena reduce the perception quality of underwater visual sensing systems, further impacting the reliability of underwater navigation, resource exploration, environmental monitoring, and autonomous underwater robot operations.

[0003] Existing underwater image enhancement technologies broadly include traditional image processing methods, physical model methods, and deep learning methods. Traditional image processing methods improve brightness and contrast through histogram equalization, Retinex enhancement, image fusion, or color correction. While relatively simple to implement, they typically rely on statistical image characteristics and struggle to accurately adapt to the absorption, scattering, and non-uniform illumination inherent in complex aquatic environments. Physical model methods invert degradation processes by estimating transmissivity, background light, and water optical parameters, providing a clearer imaging mechanism. However, in real underwater environments, water type, suspended particle concentration, illumination conditions, and imaging distance constantly change, making it difficult to stably estimate relevant parameters. This can lead to under-recovery or over-correction during engineering deployment.

[0004] Deep learning methods have achieved good results in color restoration and detail enhancement by constructing end-to-end neural networks to learn the mapping relationship between degraded underwater images and clear images. Convolutional neural networks can extract local texture features, generative adversarial networks can improve the visual realism of enhanced images, Transformer structures can model long-range dependencies, and multi-color space fusion strategies can improve brightness and color balance. However, most deep learning underwater image enhancement models have a large number of parameters, high computational cost, and high inference latency. They usually rely on high-performance desktop GPUs for real-time processing, making them difficult to deploy directly on platforms with limited computing power, power consumption, and heat dissipation conditions, such as AUVs, ROVs, underwater imaging systems, and underwater environmental monitoring equipment.

[0005] Lightweight networks reduce model complexity through depthwise separable convolutions, channel rearrangements, axial convolutions, and U-shaped encoder-decoder structures. Axial depthwise convolutions decompose two-dimensional spatial convolutions into vertical and horizontal convolutions, reducing computational cost while preserving a large effective receptive field. However, existing lightweight underwater image enhancement networks based on full-channel axial depthwise convolutions still have the following problems: First, they perform the same form of axial spatial modeling on all channels, ignoring the differences in response to underwater degradation across different channels, resulting in redundant computation on low-response channels; second, bottleneck layers typically use a single-scale fixed receptive field, making it difficult to simultaneously model multi-scale degradation features such as global color shift, local texture blurring, and contrast reduction; third, skip connections typically directly stitch together shallow encoder features, failing to correct for shallow feature deviations caused by non-uniform light attenuation between channels in the underwater environment, affecting the edge and texture restoration quality during the decoding stage.

[0006] Therefore, there is an urgent need to develop a lightweight underwater image enhancement method based on deep learning. Under the premise of ensuring image quality, it can achieve real-time underwater image enhancement with a low number of parameters and a low computational cost, thereby reducing the number of model parameters and computational cost and improving the real-time image enhancement efficiency of resource-constrained underwater platforms. Summary of the Invention

[0007] This invention provides a lightweight underwater image enhancement method and system based on deep learning. Under the premise of ensuring image quality, it can achieve real-time underwater image enhancement with a low number of parameters and a low computational load. It can reduce the number of model parameters and computational load, and improve the real-time image enhancement efficiency of resource-constrained underwater platforms.

[0008] The present invention provides the following technical solution: a lightweight underwater image enhancement method based on deep learning, comprising: S1, receiving an RGB underwater degraded image and extracting shallow features through a primary convolutional layer;

[0009] S2. Input the shallow features into the encoder, and perform step-by-step feature abstraction through the partial axial depth convolution module contained in the encoder to obtain multi-layer encoded features and deep features.

[0010] S3. Input the deep features into the multi-scale axial adaptive spatial attention bottleneck module to perform multi-scale degradation feature modeling and obtain bottleneck enhancement features;

[0011] S4. The multi-layer coding features are input into the channel space attention module for adaptive recalibration to obtain the recalibrated skip connection features.

[0012] S5. Input the bottleneck enhancement feature into the decoder and upsample it step by step to obtain the upsampled feature. Then, concatenate and fuse the upsampled feature with the recalibrated skip connection feature of the corresponding level to obtain the fused decoding feature.

[0013] S6. The fused decoding features are mapped to the RGB space through the reconstruction layer to generate an enhanced underwater image.

[0014] This lightweight underwater image enhancement method based on deep learning reduces redundant computation caused by full-channel spatial modeling by performing axial depth convolution only on a subset of channels using a partial axial depth convolution module, thus reducing the number of model parameters and computational cost while maintaining spatial context extraction capabilities. A multi-scale axial adaptive spatial attention bottleneck module performs parallel extraction and attention weighting of degenerate features, improving joint modeling capabilities. A channel spatial attention module corrects shallow features in skip connections, improving restoration quality. A multi-constraint joint loss function provides collaborative constraints, achieving a balance across multiple aspects of the enhanced image. The network has low parameter and computational cost, making it suitable for real-time image enhancement on resource-constrained underwater platforms.

[0015] As an optional scheme of the lightweight underwater image enhancement method based on deep learning described in this invention, the following steps are taken: a partial axial depth convolution module divides shallow features into a spatial feature extraction branch and an identity mapping branch along the channel dimension; vertical and horizontal axial depth convolutions are performed on the spatial feature extraction branch, and the spatial branch output is obtained in the form of residuals; the identity mapping branch is directly used as the identity mapping branch output; the spatial branch output and the identity mapping branch output are concatenated along the channel dimension to obtain the module output. The primary convolutional layer enhances the input channels of the RGB underwater degraded image through convolutional layers and the GELU activation function, encoding low-level color and texture information to provide an initial feature representation with spatial context for the subsequent encoder. The encoder's step-by-step feature abstraction includes: extracting lightweight spatial features through a partial axial depth convolution module, adaptively adjusting channel weights through a channel attention module, obtaining the skip connection output of this layer through batch normalization, enhancing the channel dimension through convolution, performing spatial downsampling, and obtaining the input features of the next layer through GELU activation.

[0016] As an optional scheme of the lightweight underwater image enhancement method based on deep learning described in this invention, the multi-scale axial adaptive spatial attention bottleneck module includes a channel dimensionality reduction unit, a multi-scale partial axial convolution branch, a branch attention unit, and a channel restoration unit. The channel dimensionality reduction unit compresses the input deep features through convolution. The multi-scale partial axial convolution branch includes several partial axial convolution branches at different scales, corresponding to different receptive fields of the image. Through parallel multi-scale branches, it simultaneously represents the local blurring and global color attenuation features in the underwater image. The branch attention unit adaptively weights the features output by the partial axial convolution branches at several scales. The channel restoration unit concatenates the weighted features and restores the number of channels through convolution. The decoder consists of several symmetric decoding blocks and a reconstruction layer. The number of decoding blocks is consistent with the number of partial axial convolutional branches. Each decoding block first performs bilinear interpolation upsampling on the input features, then concatenates the upsampled input features with the recalibrated skip connection features of the corresponding level along the channel dimension. The number of channels is compressed by pointwise convolution, and the output features of the current decoding block are generated sequentially through batch normalization, partial axial depth convolution module, channel attention module, second pointwise convolutional layer and GELU activation function. The reconstruction layer maps the output features of the last level decoding block to the RGB space through convolution. The branch attention unit adaptively weights the output features of partial axial convolution branches at several scales, including: concatenating the output features of partial axial convolution branches at several scales along the channel dimension to obtain concatenated features; performing channel average pooling and channel max pooling on the concatenated features to obtain average pooling statistical features and max pooling statistical features; concatenating the average pooling statistical features and max pooling statistical features, and then generating spatial attention weights corresponding to several scales through convolution mapping and Sigmoid activation; and modulating the branch features at several scales using residual modulation and then concatenating them to obtain bottleneck enhancement features.

[0017] As an alternative to the lightweight underwater image enhancement method based on deep learning described in this invention, when the reconstruction layer generates the enhanced underwater image, a joint loss function is used for optimization during the training of the network used to generate the enhanced underwater image. The joint loss function consists of the Charbonnier pixel reconstruction loss L. Char CIELab color space constraint loss L LAB CIELCH color space constraint loss L LCH VGG perceived loss L VGG and structural similarity loss L SSIM Linear weighted composition.

[0018] A system applying any of the above-mentioned lightweight underwater image enhancement methods based on deep learning includes: a processor and a memory, wherein the processor is used for the lightweight underwater image enhancement method based on deep learning.

[0019] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a lightweight underwater image enhancement method based on deep learning.

[0020] The present invention has the following beneficial effects:

[0021] 1. This lightweight underwater image enhancement method based on deep learning reduces redundant computation caused by full-channel spatial modeling by performing axial depth convolution only on a portion of the channels using a partial axial depth convolution module, thus reducing the number of model parameters and computational cost while maintaining spatial context extraction capabilities. A multi-scale axial adaptive spatial attention bottleneck module performs parallel extraction and attention weighting of degenerate features under different receptive fields, improving the joint modeling capability for global color cast, local texture blurring, and contrast reduction. A channel spatial attention module corrects shallow features of skip connections, improving the edge and texture restoration quality during the decoding stage. A multi-constraint joint loss function collaboratively constrains pixel, structure, perception, and color distributions, achieving a balance between color naturalness, contrast, and detail fidelity in the enhanced image. The network has low parameter and computational cost, making it suitable for real-time image enhancement on resource-constrained underwater platforms.

[0022] 2. In this lightweight underwater image enhancement method based on deep learning, the structure of some axial depth convolution modules can reduce spatial convolution calculations on redundant channels while preserving the effective receptive field of axial convolution.

[0023] 3. In this lightweight underwater image enhancement method based on deep learning, the decoder receives corrected detail features to compensate for shallow feature deviations caused by non-uniform light attenuation between channels in the underwater environment.

[0024] 4. In this lightweight underwater image enhancement method based on deep learning, a joint loss function is used to train the network for generating enhanced underwater images, and the enhancement results are constrained from multiple aspects such as pixel fidelity, structural consistency, perceptual quality and color naturalness. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the overall network structure of a lightweight underwater image enhancement method based on deep learning, provided in an embodiment of the present invention.

[0026] Figure 2 This is a schematic diagram of the structure of a partial axial depth convolution module provided in an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram of the structure of the multi-scale axial adaptive spatial attention bottleneck module provided in an embodiment of the present invention.

[0028] Figure 4 This is a flowchart illustrating a lightweight underwater image enhancement method based on deep learning, provided in an embodiment of the present invention.

[0029] Figure 5 A schematic diagram of the training constraints and deployment process provided in the embodiments of the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Reference Figure 1 and Figure 4 In one embodiment, the lightweight underwater image enhancement method based on deep learning includes steps S101 to S107.

[0032] Step S101: Receive the underwater degradation image, wherein the underwater degradation image is an RGB image, denoted as... Where H represents the image height, W represents the image width, and 3 is the number of channels. The input image exhibits blue-green color cast, reduced contrast, fogging, and detail loss due to water absorption and scattering.

[0033] Step S102: Extract shallow features through a primary convolutional layer. The input image I undergoes preliminary feature extraction through a 5×5 convolutional layer and the GELU activation function, increasing the number of channels to 16 to obtain shallow features F0, whose expression is: .

[0034] Encoding low-level color and texture information through primary convolutional layers can provide initial feature representations with spatial context for subsequent encoders.

[0035] Step S103 involves progressive feature abstraction using an encoder. The encoder consists of three consecutive coding blocks with output channels of 32, 64, and 128 respectively. For the l-th coding block, its input feature is X. l First, lightweight spatial features are extracted using a partial axial depth convolution module. Then, a channel attention module adaptively adjusts the channel weights to suppress underwater noise channels and highlight key detail features. Finally, batch normalization is performed to obtain the skip connection output F of this layer. skip,l Its expression is: .

[0036] The encoding block continues to increase the channel dimension through 1×1 convolution, and performs spatial downsampling using max pooling with a stride of 2. The next layer's input feature X is obtained through GELU activation. l+1 Its expression is: .

[0037] After three encoding blocks, the feature map space size is progressively reduced from H×W to H / 8×W / 8, forming deep features. This progressive abstraction process enhances the network's robust representation of low-contrast structures and blurred textures in degraded underwater images.

[0038] Please refer to Figure 2 Partial axial depth convolution modules are used to reduce convolutional redundancy while maintaining spatial modeling capabilities. Let the input features be... The module divides the input features into spatial feature extraction branches along the channel dimension. and identity mapping branch Where C represents the total number of input channels, C p This indicates the number of channels participating in the axial depth convolution.

[0039] Spatial feature extraction branch I p The module performs vertical axial depth convolution and horizontal axial depth convolution respectively, and preserves the original branch features through residuals, resulting in: Where ⊗ represents the convolution operation, This represents the depthwise convolution kernel in the vertical direction. This represents the horizontal depthwise convolution kernel. For the identity mapping branch I... res The module outputs directly without introducing convolution calculations: .

[0040] The module concatenates the two branches along the channel dimension to obtain the output: When the same kernel size k is used in both the vertical and horizontal directions, the computational cost of some axial depth convolution modules is similar to that of C. p Proportional, which can be expressed as: .

[0041] In one specific implementation, C p When set to C / 2 and the kernel size k is set to 5, then: .

[0042] The computational cost of full-channel axial depthwise convolution is: .

[0043] Therefore, in At this time, the theoretical computational cost of the partial axial depth convolution module is reduced by half compared to the full-channel axial depth convolution. Since the remaining channels directly retain the original features through identity mapping, the module can retain the large receptive field advantage of axial convolution while reducing redundant spatial convolution computations.

[0044] Step S104 involves processing deep features using a multi-scale axial adaptive spatial attention bottleneck module. (Refer to...) Figure 3 After the deep features are input into the bottleneck module, they are first reduced in dimensionality through 1×1 convolution to obtain the dimensionality-reduced features Xr: .

[0045] The dimensionality-reduced feature Xr is fed into parallel partial axial convolution branches at three scales: 1×1, 3×3, and 5×5, yielding branch features F1, F3, and F5 at the three scales: ,

[0046] .

[0047] The three scales correspond to different receptive fields. The 1×1 branch preserves local channel response, the 3×3 branch extracts medium-range texture structure, and the 5×5 branch captures a wider range of color shift and contrast variations. Through parallel multi-scale branching, the bottleneck module can simultaneously characterize local blurring and global color attenuation features in underwater images.

[0048] Subsequently, the branch attention unit concatenates F1, F3, and F5 along the channel dimension to obtain Fcat: .

[0049] Performing channel-wise average pooling and channel-wise max pooling on Fcat yields statistical features Favg and Fmax. These two statistical features are concatenated and then subjected to convolutional mapping and sigmoid activation to generate spatial attention weights α, β, and γ corresponding to three scale branches. , , , .

[0050] Applying attention weights to the three scale branch features using residual modulation yields: , , .

[0051] Residual modulation preserves the basic representation of the original scale branches while enhancing the response in key regions, avoiding the low effective feature response caused by direct multiplicative suppression. Finally, the recalibrated multi-scale features are concatenated and the number of channels is restored through a 1×1 convolution to obtain the bottleneck enhancement feature Y: .

[0052] Step S105 involves recalibrating the skip connection features using a channel spatial attention module. The Fskip,l output from each layer of the encoder is input to the channel spatial attention module before being passed to the decoder. This module adaptively weights the shallow features from both the channel and spatial dimensions, outputting the recalibrated skip connection features F'skip,l. This process reduces shallow channel bias caused by non-uniform underwater light attenuation, allowing the subsequent decoder to fuse more stable edge and texture features.

[0053] Step S106 involves progressive upsampling and concatenation fusion using the decoder. The decoder comprises three symmetrical decoding blocks. For the l-th decoding block, its input feature X... dec,l First, upsampling is performed using bilinear interpolation, and then the recalibrated skip connection feature F' from the corresponding coding layer output is used. skip,l Stitching along the channel dimension: .

[0054] The concatenated features are compressed by pointwise convolution to reduce the number of channels, and then refined by batch normalization, partial axial depth convolution, and channel attention. .

[0055] The refined features are then used to generate the output features of the current decoded block through a second pointwise convolution and the GELU activation function: .

[0056] Bilinear interpolation upsampling does not introduce additional learnable parameters, enabling it to maintain lightweight design while restoring spatial resolution. By concatenating and fusing deep semantic features with recalibrated shallow detail features, the decoder is able to better recover target contours, texture edges, and local contrast in underwater images.

[0057] Step S107: Enhance the underwater image by outputting the reconstruction layer. After the spatial resolution is restored sequentially by the three decoding blocks, the final level of decoding features is mapped to the RGB space through a 1×1 convolution to obtain the enhanced underwater image. .

[0058] Please refer to Figure 5 During the training phase, the network uses a joint loss function L. total Optimization is performed. The joint loss function includes the Charbonnier pixel reconstruction loss L... Char CIELab color space constraint loss L LAB CIELCH color space constraint loss L LCH VGG perceived loss L VGG and structural similarity loss L SSIM Its expression is: .

[0059] The Charbonnier pixel reconstruction loss is used to constrain the pixel error between the augmented image and the reference image, and improves training stability through a smoothing constant. Its expression is: Among them, k i(p) Let k represent the pixel value at pixel position p in the i-th enhanced image. i *(p) represents the pixel value at pixel position p in the corresponding reference image, N represents the batch sample size, H and W represent the image height and image width, respectively, and ε represents the smoothing constant.

[0060] The CIELab color space constraint loss utilizes the decoupling property of luminance and chrominance in the Lab color space to constrain color shift; the CIELCH color space constraint loss constrains and enhances the color naturalness of an image from the perspectives of hue and saturation, and their expressions are as follows: , Where kᵢᴸᵃᵇ and kᵢ*ᴸᵃᵇ represent the representations of the i-th enhanced image and the reference image in the CIELab color space, respectively, and kᵢᴸᶜᴴ and kᵢ*ᴸᶜᴴ represent the representations of the i-th enhanced image and the reference image in the CIELCH color space, respectively. The VGG perceptual loss extracts high-dimensional features through a pre-trained VGG network and constrains the difference between the enhanced image and the reference image in the feature space; its expression is: Where φⱼ represents the feature map output by the j-th convolutional layer of the pre-trained VGG network, and Cⱼ, Hⱼ, and Wⱼ represent the number of channels, height, and width of the corresponding feature map, respectively. The structural similarity loss constrains the consistency between the enhanced image and the reference image in terms of brightness, contrast, and structure, satisfying: , , where μ x μᵧ and μᵧ represent the mean values ​​of the enhanced image and the reference image, respectively, and σ x ² and σᵧ² represent variance, respectively, and σ x ᵧ represents the covariance, and C1 and C2 represent the stability constants. Through a linear combination of the above loss terms, the network can simultaneously optimize pixel fidelity, structure preservation, perceptual quality, and color naturalness.

[0061] In one training implementation, the network is implemented based on the PyTorch deep learning framework. The training data uses the LSUI training set, input images are uniformly adjusted to 256×256 pixels, the batch size is set to 16, the number of training epochs is set to 100, the optimizer is AdamW, and the initial learning rate is set to... The weight decay coefficient is set to The learning rate scheduling strategy uses ReduceLROnPlateau. If the validation set loss does not decrease within five consecutive training epochs, the learning rate is reduced to 0.5 times its original value. The weight coefficients in the joint loss function are set to... , , , , .

[0062] In one test instance, the trained network was tested on three publicly available underwater image enhancement datasets: UIEB, EUVP, and LSUI, and also tested in a real-world embedded scenario within an underwater imaging system. The test results show that the network has 0.065M parameters, 0.493G FLOPs of computation, a single-frame processing latency of 3.46ms on an NVIDIA RTX 3090 platform, corresponding to a processing speed of 288.95 fps; and can stably achieve 75fps online enhancement under 256×256 input conditions on an embedded system. The enhancement results can reduce the blue-green color cast in the original underwater image, improve the discernibility of dark areas, enhance target edges and texture structure, and reduce the risks of oversaturation, local overexposure, and inter-frame flicker.

[0063] As can be seen from the above embodiments, the present invention combines partial axial depth convolution, multi-scale axial adaptive spatial attention, channel spatial attention skip connection recalibration, and multi-constraint joint loss function into the symmetric encoder-decoder structure, thereby achieving real-time underwater image enhancement under conditions of low parameter quantity and low computational quantity. It is suitable for AUVs, ROVs, underwater environmental monitoring equipment, underwater imaging systems, and underwater embedded vision platforms.

[0064] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0065] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A lightweight underwater image enhancement method based on deep learning, characterized in that, include: S1. Receive RGB underwater degraded images and extract shallow features through a primary convolutional layer; S2. Input the shallow features into the encoder, and perform step-by-step feature abstraction through the partial axial depth convolution module contained in the encoder to obtain multi-layer encoded features and deep features. S3. Input the deep features into the multi-scale axial adaptive spatial attention bottleneck module to perform multi-scale degradation feature modeling and obtain bottleneck enhancement features; S4. The multi-layer coding features are input into the channel space attention module for adaptive recalibration to obtain the recalibrated skip connection features. S5. Input the bottleneck enhancement feature into the decoder and upsample it step by step to obtain the upsampled feature. Then, concatenate and fuse the upsampled feature with the recalibrated skip connection feature of the corresponding level to obtain the fused decoding feature. S6. The fused decoded features are mapped to the RGB space through the reconstruction layer to generate an enhanced underwater image; Some axial depth convolution modules divide shallow features into spatial feature extraction branches and identity mapping branches in the channel dimension; Perform vertical axial depth convolution and horizontal axial depth convolution on the spatial feature extraction branch, and obtain the spatial branch output in the form of residuals; The identity mapping branch is directly output as the identity mapping branch; The spatial branch output and the identity mapping branch output are concatenated along the channel dimension to obtain the module output; The multi-scale axial adaptive spatial attention bottleneck module includes a channel dimensionality reduction unit, a multi-scale partial axial convolution branch, a branch attention unit, and a channel recovery unit. The channel reduction unit compresses the input deep features through convolution; The multi-scale partial axial convolution branch includes several scales of partial axial convolution branches, each corresponding to a different receptive field of the image. Through parallel multi-scale branches, it simultaneously characterizes the local blurring and global color attenuation features in underwater images. The branch attention unit adaptively weights the features output by partial axial convolution branches at several scales. The channel restoration unit concatenates the weighted features and restores the number of channels through convolution.

2. The lightweight underwater image enhancement method based on deep learning according to claim 1, characterized in that: The primary convolutional layer enhances the input channels of the RGB underwater degraded image through convolutional layers and the GELU activation function, encodes low-level color and texture information, and provides an initial feature representation with spatial context for the subsequent encoder.

3. The lightweight underwater image enhancement method based on deep learning according to claim 2, characterized in that: The encoder's stepwise feature abstraction includes: extracting lightweight spatial features through a partial axial depth convolution module, adaptively adjusting channel weights through a channel attention module, obtaining the skip connection output of the layer after batch normalization, increasing the channel dimension through convolution, performing spatial downsampling, and obtaining the input features of the next layer through GELU activation.

4. The lightweight underwater image enhancement method based on deep learning according to claim 1 or 3, characterized in that: The branch attention unit adaptively weights the features output by partial axial convolution branches at several scales, including: The spliced ​​features are obtained by concatenating the output features of partial axial convolution branches of several scales along the channel dimension. Perform channel average pooling and channel max pooling on the spliced ​​features respectively to obtain average pooling statistical features and max pooling statistical features; The average pooling statistical features and the max pooling statistical features are concatenated and then subjected to convolutional mapping and Sigmoid activation to generate spatial attention weights corresponding to several scales. Bottleneck enhancement features are obtained by modulating several scale branch features using residual modulation and then splicing them together.

5. The lightweight underwater image enhancement method based on deep learning according to claim 1 or 3, characterized in that: The decoder consists of several symmetric decoding blocks and a reconstruction layer. The number of decoding blocks is consistent with the number of partial axial convolutional branches. Each decoding block first performs bilinear interpolation upsampling on the input features, then concatenates the upsampled input features with the recalibrated skip connection features of the corresponding level along the channel dimension, compresses the number of channels through pointwise convolution, and generates the output features of the current decoding block through batch normalization, partial axial depth convolution module, channel attention module, second pointwise convolutional layer and GELU activation function in sequence. The reconstruction layer maps the output features of the last level decoding block to the RGB space through convolution.

6. The lightweight underwater image enhancement method based on deep learning according to claim 1, characterized in that: When the reconstruction layer generates enhanced underwater images, the network used to generate these images is optimized using a joint loss function during training. This joint loss function consists of the Charbonnier pixel reconstruction loss L. Char CIELab color space constraint loss L LAB CIELCH color space constraint loss L LCH VGG perceived loss L VGG and structural similarity loss L SSIM Linear weighted composition.

7. A lightweight underwater image enhancement system based on deep learning, characterized in that: It includes a processor and a memory, the processor being used to execute the lightweight underwater image enhancement method based on deep learning as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The medium stores a computer program, which, when executed by a processor, implements the lightweight underwater image enhancement method based on deep learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Semantic segmentation network and semantic segmentation method

    CN120635453A

  • Industrial part surface defect detection method and system with robustness

    CN121258875A