Underwater image enhancement method based on improved MobileNetV4
By improving the MobileNetV4 network and using a lightweight generative adversarial network (MWCA-GAN) with multi-module collaborative optimization, the problem of lightweight design and performance imbalance in underwater image enhancement technology is solved, achieving efficient and automatic color correction and detail restoration, suitable for resource-constrained devices.
Patent Information
- Application Number
- CN202511083751.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-28
AI Technical Summary
Existing underwater image enhancement technologies face challenges such as a trade-off between lightweight design and performance, insufficient multi-degradation collaborative processing capabilities, and limited high-frequency detail recovery, making it difficult to achieve efficient inference and high-quality image enhancement on resource-constrained devices.
We adopt an improved MobileNetV4 network architecture and combine generative adversarial networks (GANs) and multi-scale feature fusion mechanisms to design a lightweight generative adversarial network (MWCA-GAN). Through wavelet denoising, color correction and multi-query attention modules, we improve image quality and computational efficiency.
While maintaining high-efficiency computing, it significantly optimizes underwater image quality, reduces the number of parameters by 99.3%, reduces the inference time per image to 0.0039 seconds, improves color correction by 32.5%, and improves texture restoration by 15.1%, making it suitable for deployment on resource-constrained devices.
Smart Images

Figure CN121032828A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and image processing technology, and specifically relates to a lightweight generative adversarial network (GAN) enhancement method for underwater environment degradation characteristics. Background Technology
[0002] The complex optical characteristics of the underwater environment pose significant challenges to image acquisition and processing. The absorption and scattering of light by the water medium lead to widespread degradation in underwater images, including color distortion, low contrast, blurred details, and noise interference. These problems severely limit the application value of underwater images in fields such as ocean exploration, biological monitoring, and underwater robot navigation. With the increasing demand for underwater operations, developing efficient underwater image enhancement techniques has become an important research direction in the field of computer vision.
[0003] Underwater image enhancement aims to improve the visual quality of images through various signal processing techniques, including but not limited to contrast adjustment, color correction, and dehazing. Commonly used underwater image enhancement techniques mainly include non-physical models, physical models, and deep learning methods. Meanwhile, we conducted an in-depth analysis of real-world underwater target detection datasets to study the performance of different underwater image enhancement algorithms on underwater images of varying quality, and to explore the correlation between underwater image enhancement and underwater target detection tasks.
[0004] Currently, mainstream underwater image enhancement methods can be categorized into three main types: Non-physical model methods: Histogram equalization algorithms (such as HE and AHE) improve contrast by adjusting pixel distribution, but are prone to local over-enhancement or noise amplification. Retinex theory-derived methods (such as MSRCR) achieve color correction by simulating the human visual system, but have limited adaptability to complex underwater degradation scenarios. Physical model methods: Algorithms based on dark channel priors (DCP) or red channel priors (such as UDCP) recover images through inverse imaging models, but rely on prior assumptions and are prone to failure in turbid waters or under artificial lighting. Optical property modeling methods (such as illumination correction proposed by Galdran et al.) can alleviate color cast, but have limited effectiveness in noise and detail recovery. Deep learning methods: CNN-based models (such as UIE-Net and Water-Net) achieve enhancement through end-to-end learning, but require a large amount of paired data, and have a large number of model parameters and slow inference speed. GAN-based methods (such as FUnIE-GAN and CycleGAN) generate high-quality images through adversarial training, but traditional GAN structures are complex and difficult to deploy on resource-constrained devices. For example, although FUnIE-GAN is lightweight, it still suffers from color temperature inconsistency; UGAN-P optimizes stability through Wasserstein distance, but it has high computational costs.
[0005] Current underwater image enhancement technologies face core challenges such as the trade-off between lightweight design and performance, insufficient multi-degradation collaborative processing capabilities, and limited high-frequency detail recovery. There is an urgent need for a novel network architecture that balances efficient inference with high-quality enhancement. This architecture should overcome the limitations of existing models by integrating frequency domain denoising, dynamic color correction, and multi-scale feature fusion mechanisms, providing a reliable solution for underwater vision tasks. Summary of the Invention
[0006] The purpose of this invention is to overcome the severe color deviation often seen in underwater images due to the absorption and scattering of light of different wavelengths by water. This deviation manifests as an overall blue or green tint, with a significant loss of warm tones such as red. This color deviation not only affects the subjective perception of the human eye but also leads to a decrease in the accuracy of subsequent underwater target detection and recognition tasks. This invention proposes an efficient and automatic color correction module, which significantly improves the color restoration capability of underwater images.
[0007] This invention provides an underwater image enhancement method based on an improved MobileNetV4, comprising the following steps: Construct a lightweight generative adversarial network for underwater image enhancement, comprising a generator network and a discriminator network; The generative network includes an encoder and a decoder. The encoder consists of several layers of MobileNetV4 DownSample modules and Universal Inverted Bottleneck (UIB) modules, which alternate to achieve multi-scale feature extraction and spatial downsampling. The final layer of the MobileNetV4 DownSample module encoder is sequentially connected to the wavelet denoising module and the color correction module, which are used to remove high-frequency noise and adaptively correct color deviation, respectively. The color correction module is then connected to the decoder; The decoder includes several layers of MobileNetV4 UpSample modules for progressively restoring resolution. LayerScale and MultiQuery Attention modules are interspersed among the MobileNetV4 UpSample modules to improve the stability of feature flow and the ability to perceive key regions; The last MobileNetV4 UpSample module is followed by a transposed convolution and a Tanh activation function, which outputs the generated image obtained by the generator network. The generated image and the unenhanced real image are input into the discriminator network to obtain the enhanced underwater image; The MobileNetV4DownSample module includes depthwise separable convolution, batch normalization, and ReLU activation operations, and achieves spatial size reduction through depthwise separable convolution. The MobileNetV4UpSample module includes transposed convolution, batch normalization, and ReLU activation operations. Transposed convolution increases the spatial size.
[0008] Preferably, the wavelet denoising module removes high-frequency noise and preserves image structural details through multi-scale decomposition; the color correction module uses a channel attention mechanism to achieve adaptive correction of color cast in underwater images; and the multi-query attention module enhances the decoder's ability to focus on key regions and details, significantly improving the visual quality and information expression of the enhanced image. The wavelet denoising module performs frequency domain decomposition on the feature map of the input wavelet denoising module through discrete wavelet transform (DWT), decomposing each channel into low-frequency and high-frequency components to remove high-frequency noise interference. The wavelet denoising module uses Haar wavelets as the basis function to extract features from low-frequency components through a separable convolutional structure containing 3×3 depth convolution and 1×1 point convolution. Finally, the spatial structure of the image is restored through inverse discrete wavelet transform (IDWT).
[0009] Preferably, the color correction module extracts global channel statistics from the input feature map through adaptive average pooling to capture the overall color distribution trend. Then, it generates channel attention weights through several layers of convolutional networks and introduces a learnable scaling factor to adjust the amplitude of the channel response, thereby correcting the blue bias, green bias, and saturation imbalance in underwater images.
[0010] Preferably, the multi-query attention module generates a query vector through 1×1 convolution, generates a key vector and a value vector through depthwise separable convolution with dilation, then flattens the spatial dimension, calculates the correlation between query and key using a scaled dot product attention mechanism, obtains attention weights, performs weighted summation on the value vector to obtain context-enhanced features, and finally completes the projection output through 1×1 convolution.
[0011] Preferably, the LayerScale module is used to first initialize a scaling vector with the same number of input feature channels, and then expand the scaling vector to the same spatial dimension as the input features through a Broadcast operation; finally, the expanded scaling parameters are multiplied element-wise with the input features of the LayerScale module to achieve adaptive control of the activation intensity of each channel.
[0012] Preferably, a skip connection mechanism is adopted to directly transfer the shallow high-resolution features of the encoder to the corresponding layer of the decoder, thereby achieving effective fusion of multi-scale features.
[0013] Preferably, the generated image and the real image are stitched together and then input into the discriminator network; The discriminator comprises three parallel scale branches: a primary scale branch (Scale 1), a secondary scale branch (Scale 2), and a coarse scale branch (Scale 3), each branch used to capture image features at a specific level. The original scale branch Scale 1 is used to process the input image at the original resolution, including four cascaded downsampling modules, ZeroPad operation, and convolution operation; The medium-scale branch Scale 2 includes average pooling operations, four cascaded downsampling modules, ZeroPad operations, and convolution operations, used to process medium-resolution input images. The coarse-scale branch Scale 3 has the same structure as the medium-scale branch Scale 2, except that the average pooling operation is more extensive than that of the medium-scale branch Scale 2, and it is used to process input images with coarse resolution. Each downsampling module consists of: a 4×4 convolutional layer, a BN layer, and a LeakyReLU activation function; The feature maps obtained from the three parallel scale branches are fused using the Interpolate upsampling operation. The feature maps processed by each branch are restored to a uniform spatial size. Then, the channel dimensions are stitched together by the stitching operation to obtain a comprehensive feature representation, which in turn yields the enhanced underwater image.
[0014] Preferably, the underwater image enhancement method based on the improved MobileNetV4 further includes: alternating training of the generator and the discriminator, specifically including: When training the discriminator, the generator parameters are fixed, and the discriminator is fed real images labeled as true labels and generated images labeled as false labels generated by the generator. Huber loss is used as adversarial loss to optimize the discriminator and enhance its ability to distinguish between real and false images. When training the generator, the discriminator parameters are fixed, and the generator is jointly optimized using adversarial loss, similarity loss (Smooth L1), and content-aware loss (calculated based on the mean squared error of deep features extracted by VGG19 and ResNet50).
[0015] Preferably, optimizing the generator using the adversarial loss includes: inputting noise into the generator, then labeling the fake images generated by the generator as real labels and sending them to the discriminator, and using Huber loss as the adversarial loss to optimize the generator, so that the discriminator's discrimination of the fake images generated by the generator is closer to that of the real images; optimizing the generator using the similarity loss includes: using the Smooth L1 loss function as the similarity loss to measure the difference between the generated image and the real image at the pixel level, so that the generated image is close to the real image in terms of structure and detail; optimizing the generator using the content-aware loss includes: using a pre-trained image feature extraction model to extract image features of the generated image and the real image respectively, obtaining the generated image features and the real image features, calculating the mean square error of the generated image features and the real image features as the content-aware loss, and optimizing the generator; the image feature extraction model includes a VGG19 network and a ResNet50 network.
[0016] The beneficial effects of this invention are as follows: The MWCA-GAN designed in this invention achieves significant optimization of underwater image quality while maintaining high computational efficiency through a modular structure and multi-scale strategy. The model uses the MobileNetV4 backbone network with only 0.37M parameters, a 99.3% reduction compared to CycleGAN (54.42M). The inference time per image is as low as 0.0039 seconds (RTX 3080), making it suitable for deployment in resource-constrained scenarios such as ROV vision systems. To address multi-source degradation, the color correction module dynamically adjusts the RGB response weights, achieving a UCIQE score of 0.631 (a 32.5% improvement over UDCP); the wavelet denoising module effectively separates high-frequency interference, improving UIQM by 15.1%. In texture restoration, the MQA mechanism focuses on local features, achieving an SSIM score of 0.818 (surpassing FUnIE-GAN's 0.002), while the LayerScale mechanism stabilizes gradient propagation, improving edge sharpness by 28% (qualitative evaluation). Experiments show that this method maintains a PSNR ≥ 25 dB (standard deviation ≤ 0.85) in turbid waters and low-light environments, and its energy consumption per unit is reduced by 82% compared to UGAN-P. Its overall performance is superior to existing mainstream algorithms, and it achieves the best performance on the EUVP / UIEBD datasets, providing robust and reliable visual support for tasks such as marine exploration and ecological monitoring. Attached Figure Description
[0017] Figure 1 This is a diagram of the lightweight generative network structure based on MobileNetV4 constructed in this invention; Figure 2 This is a diagram of the structure of a generative adversarial network model; Figure 3These are the two key module structures in the generative network of this invention: (a) the MobileNetV4DownSample module; and (b) the MobileNetV4UpSample module. Figure 4 This is a diagram of the Universal Inverted Bottleneck Module (UIB) architecture; Figure 5 The network denoising and color correction modules are: (a) wavelet transform denoising module; (b) color correction module; Figure 6 These are two key enhancement modules in the decoder: (a) a multi-query attention mechanism; and (b) a LayerScale module. Figure 7 It is a generated network graph based on MobileNetV4 and the coordination enhancement module; Figure 8 It is a multi-scale discriminator; Figure 9 Comparison of the enhanced effects after adding key modules to the ablation experiment. Detailed Implementation
[0018] The technical solution of the present invention includes the following key steps: Step 1: Construct a lightweight generative network: The generative network uses MobileNetV4 as the encoder-decoder backbone, employs a Universal Inverted Bottleneck (UIB) module to optimize feature extraction, and enhances degenerate features through multi-module collaboration. Encoder path: The input degraded image (256×256×3) is downsampled layer by layer through a 5-level MobileNetV4DownSample module, compressing the output size to 8×8 and expanding the number of channels to 112, extracting multi-scale semantic features. Each downsampling module includes depthwise separable convolution, batch normalization, and ReLU activation.
[0019] A wavelet denoising module is embedded at the end of the encoding: Haar wavelet decomposition is performed on the feature map to separate the low-frequency component (LL) from the high-frequency components (LH, HL, HH). A 3×3 depthwise separable convolutional filter is applied to the low-frequency component to suppress noise interference. The denoised features are reconstructed through inverse wavelet transform (IDWT).
[0020] A color correction module is integrated at the end of the encoding path: channel statistical features are extracted through global average pooling to generate channel attention weights. Introducing a learnable scaling factor Dynamically adjust the response intensity of the RGB channels Decoder path: Upsampling is performed layer by layer through a 4-level MobileNetV4UpSample module. Each level includes deconvolution, batch normalization, and ReLU activation, restoring the output size to 256×256. A Multi-Query Attention (MQA) module is embedded in the upsampling path: a depthwise convolution with a dilation rate of 2 is used to generate multi-scale query vectors, expanding the receptive field. Region correlation is calculated through scaled dot-product attention, and key features are weighted and fused. A LayerScale module is introduced to apply a learnable scaling factor to each channel of the decoded features, dynamically adjusting the feature amplitude.
[0021] Step 2. Multi-scale Discriminator Design: The discriminator network adopts a three-branch parallel architecture to process the original resolution, mid-scale, and low-scale inputs respectively. Multi-level feature fusion enhances discriminative ability: The branch structure design includes a high-resolution branch (Scale H), which takes the original image (256×256×3) as input and downsamples it to 16×16 sequentially through a 4-level MSDiscBranch module. Each module contains 4 core convolutional layers (stride 2, padding 1), with an increasing channel count pattern: Layer 1 (128×128, 64), Layer 2 (64×64, 128), Layer 3 (32×32, 256), and Layer 4 (16×16, 512). The activation function is LeakyReLU (slope 0.2), and normalization is BatchNorm2d. The mid-scale branch (Scale M) takes the image (128×128×3) after 2×2 average pooling, and the processing flow is the same as Scale H. The output feature map dimension is 8×8×512. Low-scale branch (Scale L): Input is a 4×4 average pooled image (64×64×3), and the processing flow is the same as Scale H. Output feature map dimension: 4×4×512.
[0022] Feature alignment and fusion scale standardization: Bilinear interpolation is used to upsample the Scale M / L feature map to 256×256 to ensure uniform spatial dimensions of each branch; Multi-scale fusion: The three are stitched together along the channel dimension, and the feature dimension after stitching is: 256×256×(512×3)=256×256×1536 Discriminant branch design: Classification branch, concatenated features are processed through two convolutional layers. Layer 1: 3×3 convolution, the number of channels is reduced to 512, ReLU activation; Layer 2: Global mean pooling (GAP) → 1×1 fully connected layer, outputting scalar discriminant values.
[0023] The intermediate supervision branch adds an extra discriminant head after the Scale H / M / L branch, outputting intermediate discriminant values to assist in training stability.
[0024] Step 3: First, the introduction of adversarial loss enables the model to possess the capabilities of a Generative Adversarial Network (GAN). This is achieved through the generator... G and discriminator D In an adversarial game, the generator continuously learns to generate more realistic images. Secondly, similarity loss is used to improve the generated images. With real images At the pixel level, the generated image is closer to the real image in both local details and global structure; Step 4: Normalize the image to [-1,1], batch size 16, training epochs 200, output enhanced image is mapped to the range [0,255] by Tanh activation, and the resolution is kept at 256×256.
[0025] Preferably, the parameters of the MobileNetV4 network in the underwater image enhancement model are replaced through transfer learning. The replacement includes the following steps: Step 1: Obtain the network parameters of the MobileNetV4 pre-trained model. Obtain the pre-trained parameters of the MobileNetV4 network fine-tuned on the ImageNet dataset and import them into the underwater image enhancement generator model, covering the parameter paths of the MobileNetV4 encoder part in the generator model network that are the same as those in the pre-trained MobileNetV4.
[0026] Step 2: Optimize the BottleNeck module. Preferably, replace the standard convolutional layers in the wavelet denoising module inserted into the BottleNeck residual module with channel attention modules having hybrid convolutional kernels. For all BottleNeck residual modules in Layer 1, Layer 2, and Layer 3, use the output of the optimized hybrid convolutional kernel channel attention module as the input to the next BottleNeck residual module. For the last two layers of the BottleNeck residual module in Layer 4, input the output of the optimized channel attention module into the multi-query attention mechanism module set after the backbone network.
[0027] Step 3: Model Training Process The model training process includes the following steps: Step (3.1): Set model parameters. Set the maximum number of hidden layer nodes to 300, the maximum number of residual nodes to 6, the expected error to 1e-3, and the initial number of hidden nodes to 100. Generate the initial weight matrix and bias matrix through the training process of Generative Adversarial Network (GAN). Calculate the error by combining the Huber loss function, the Smooth L1 loss function, and the content-aware loss function. If the error is greater than the expected error, proceed to step (3.2); otherwise, proceed to the model evaluation stage.
[0028] Step (3.2): Optimize hidden nodes. If the error does not reach the expected value, the value of the hidden node node is increased by 1, and the network weight matrix and bias matrix are updated through the GAN process. If the error does not decrease after adding a new node, the counter S of the continuously invalidally added nodes is incremented by 1. When S equals the set threshold, the model is considered to have reached its limit and enters the evaluation stage. If the error decreases, S is cleared to zero and optimization continues until the error is less than the expected value.
[0029] Step 4: Dataset and Preprocessing. The dataset includes 20,000 underwater images, of which 16,000 are used for training and 4,000 for testing, with the training and testing data divided in an 8:2 ratio. The dimensions of all underwater images are adjusted to 256×256. The images are then regularized and mapped to a normal distribution function.
[0030] Preferably, to more effectively guide the generator to simultaneously consider image realism, content consistency, and structural detail reconstruction, this study ultimately integrates adversarial loss, similarity loss, and content-aware loss into a comprehensive loss function for the generator. Its definition is as follows: ; in: and These are the hyperparameters that balance the various loss terms, and we obtained them through experimental tuning. =0.7 and =0.3, to reasonably weigh the impact of different loss items. By optimizing the design of the global hybrid pooling layer, the feature representation capability and computational efficiency in the multi-layer network are significantly improved, effectively enhancing the color correction and detail restoration performance of underwater images.
[0031] The invention will now be further described with reference to the accompanying drawings.
[0032] like Figure 7 As shown, the purpose of this invention is to address the problems of poor performance and slow speed of existing underwater image enhancement algorithms by proposing a method based on MobileNetV4 and a lightweight generative adversarial network (MWCA-GAN) with multi-module collaborative optimization, and using the improved network for underwater image enhancement.
[0033] This invention includes the following steps: Step 1: Obtain the dataset 1.1 Dataset Acquisition The EUVP dataset mainly contains underwater images with various degradation types, covering typical underwater scenes. The dataset is divided into several subsets, including real images and synthetic degraded images. In this experiment, we selected a subset of paired images from the EUVP dataset for our experiments.
[0034] 1.2 Preparing the Dataset and Preprocessing: 20,000 underwater images were selected from the dataset, with 16,000 images designated for training and 4,000 images for testing, following an 8:2 ratio. All underwater images were resized to 256×256 pixels and normalized to a normal distribution function to suit the model training requirements.
[0035] Step 2: Feature Extraction and Generative Network Design The generative network employs an encoder-decoder structure. The encoder consists of alternating layers of MobileNetV4 DownSample modules and Universal Inverted Bottleneck (UIB) modules to achieve multi-scale feature extraction and spatial downsampling. At the end of the encoder, a wavelet denoising module and a color correction module are connected sequentially to remove high-frequency noise and adaptively correct color deviations, respectively. The decoder consists of multiple layers of MobileNetV4 UpSample modules, progressively restoring the spatial resolution of the feature maps. LayerScale modules and MultiQuery Attention modules are interspersed throughout the layers to improve feature flow stability and key region perception capabilities. Finally, the enhanced underwater image is output through transposed convolution and the Tanh activation function.
[0036] The MobileNetV4DownSample module reduces the spatial size of the feature map (e.g., with a stride of 2) through depthwise separable convolutions, effectively extracting deeper semantic features while reducing computational and parameter requirements. The module typically includes operations such as depthwise separable convolutions, batch normalization, and ReLU activation, which accelerate network convergence and improve training stability.
[0037] The MobileNetV4UpSample module typically employs deconvolution (transposed convolution) to progressively restore low-resolution feature maps to the same spatial size as the input image. Simultaneously, the module incorporates batch normalization and ReLU activation functions to ensure the stability of the feature distribution and its non-linear expressive power.
[0038] Multi-module collaborative enhancement mechanism: This invention innovatively integrates a wavelet denoising module, a color correction module, and a multi-query attention mechanism. The wavelet denoising module effectively removes high-frequency noise while preserving image structural details through multi-scale decomposition; the color correction module, combined with a channel attention mechanism, achieves adaptive correction of color cast in underwater images; and the multi-query attention mechanism enhances the network's ability to focus on key regions and details, significantly improving the visual quality and information representation of the enhanced image.
[0039] The wavelet denoising module performs frequency domain decomposition on the input feature map using Discrete Wavelet Transform (DWT), decomposing each channel into low-frequency and high-frequency components. It selectively retains the low-frequency components to effectively remove high-frequency noise interference to the image. This module uses Haar wavelets, which have good time-frequency locality and edge sensitivity, as the basis function. It applies separable convolutional structures (including 3×3 depthwise convolutions and 1×1 pointwise convolutions) to extract features from the low-frequency components. Finally, it recovers the image spatial structure using Inverse Discrete Wavelet Transform (IDWT).
[0040] The color correction module integrates a channel attention mechanism with a learnable Gamma correction strategy, specifically addressing the typical blue-green bias and saturation imbalance issues in underwater images through lightweight and interpretable color correction. This module first extracts global channel statistics from the input feature map using adaptive average pooling to capture the overall color distribution trend. Then, it generates channel attention weights through a multi-layer convolutional network and introduces a learnable scaling factor to adjust the amplitude of the channel responses, forming a weighted correction mechanism.
[0041] The multi-query attention mechanism is a key module designed to address feature drift and local information loss during the decoding stage, aiming to enhance the model's ability to focus on key regions and model context. This module receives the upsampled input feature map, generates a Query vector through 1×1 convolutions, and generates Key and Value vectors through depthwise separable convolutions with dilation. The dilated convolution design effectively expands the receptive field and enhances the contextual representation ability between features. Subsequently, the module flattens the spatial dimensions, calculates the correlation between the Query and Key using a scaled dot product attention mechanism, obtains the attention weights, and then performs a weighted summation of the Value features to obtain context-enhanced features. Finally, a 1×1 convolution is used to complete the projection output.
[0042] LayerScale module and feature flow control: The LayerScale module is a lightweight feature scaling component that dynamically adjusts feature amplitude by introducing a learnable channel scaling parameter γ. The module first initializes a scaling vector with the same number of input feature channels through a Learnable layer, typically with an initial value set to a small positive number (e.g., 0.01). Then, it expands the scaling vector to the same spatial dimension as the input features using a Broadcast operation. Finally, it performs element-wise multiplication of the expanded scaling parameter with the original feature map to achieve adaptive control of the activation intensity of each channel.
[0043] By interspersing LayerScale modules in each layer of the decoder, the feature amplitude is finely adjusted, which improves the stability of network training and the controllability of feature flow, effectively avoids problems such as gradient vanishing or exploding, and ensures the robustness of the model on different hardware platforms.
[0044] Skip connections and multi-scale feature fusion: By employing a skip connection mechanism, high-resolution features from the shallow layers of the encoder are directly passed to the corresponding layers of the decoder, achieving effective fusion of multi-scale features. This design not only supplements image detail information but also improves the global consistency and local detail restoration capability of the enhancement results.
[0045] Step 3: Discriminator Network Design and Training Employing an innovative multi-scale discriminator architecture, this design leverages the characteristic that images exhibit different levels of information at different resolutions. It achieves hierarchical quality supervision from global contours to local textures by processing features at multiple scales in parallel. The entire discriminator network receives channel-concatenated inputs from generated and real images, forming a tensor of dimensions (B, 2C, H, W), where B is the batch size, C is the number of channels per image, and H and W are the image height and width, respectively.
[0046] The discriminator is designed with three parallel scale branches, each specifically responsible for capturing image features at a particular level: Scale 1 (Original Scale Branch): This branch processes the input image at its original resolution, primarily capturing fine texture details, edge sharpness, and local structural information. It is highly sensitive to pixel-level quality differences and possesses a strong ability to detect subtle artifacts, texture distortion, and edge blurring in the generated image.
[0047] Scale 2 (Medium-scale branch): This branch downsamples the input image using AvgPool2d average pooling, focusing on mid-scale feature pattern recognition. It primarily focuses on mid-level semantic information of the image, such as object shape and contours, transitions between regions, and mid-scale texture patterns, effectively balancing the discriminative power between global structure and local details.
[0048] Scale 3 (Coarse Scale Branch): After further downsampling with AvgPool2d, this branch is specifically designed to process the global structural information of the image. It is primarily responsible for evaluating the overall layout rationality, large-scale semantic consistency, and global color distribution coordination of the image, ensuring that the generated image maintains consistency with the real image at a macroscopic level.
[0049] Each scale branch consists of multiple carefully designed downsampled blocks strung together. The specific structure of a single downsampled block includes: a 4×4 convolutional layer: using a large convolutional kernel (4×4) and a stride of 2, which achieves both feature extraction and reduction of spatial resolution. Compared to the traditional 3×3 convolution, the 4×4 convolution has a larger receptive field, which can better capture the spatial correlation and texture patterns of local regions.
[0050] Normalizing the convolutional output effectively alleviates the internal covariate shift problem, accelerates network convergence, and improves training stability. The introduction of Batch Normalization (BN) layers allows the network to use a larger learning rate while reducing its sensitivity to parameter initialization.
[0051] Instead of the traditional ReLU activation function, LeakyReLU is used to avoid neuron death by introducing a small negative slope (typically 0.01 or 0.02). LeakyReLU maintains a small gradient in the negative region, which helps to preserve the network's expressiveness and training dynamism, making it particularly suitable for adversarial training scenarios.
[0052] After being processed by multiple layers of Downsample Blocks, each scale branch undergoes final feature mapping via ZeroPad+Conv operations. The ZeroPad operation ensures the integrity of features at the convolutional boundaries, while subsequent convolutional operations map the multidimensional features into the representation space required for discrimination.
[0053] To achieve effective fusion of features at different scales, an Interpolate upsampling operation was designed to restore the feature maps processed by each branch to a uniform spatial size. This alignment mechanism ensures that information at different scales can be compared and fused within the same spatial framework.
[0054] The multi-scale features, after being aligned to their dimensions, are concatenated along their channel dimensions using a Concat operation to form a comprehensive feature representation with dimensions (B, num_scales, H', W'). This design allows the discriminator to consider information from multiple scales simultaneously, resulting in a more comprehensive and robust discriminative capability.
[0055] Step 4: Model Training 4.1 Joint training of the generator and discriminator: The EUVP dataset was used for training, with a total of 200 training rounds and a batch size of 16. The Adam optimizer was selected, with an initial learning rate of 0.0003.
[0056] 4.2 Model Optimization Strategy: In actual training, the discriminator is updated first, while the generator parameters remain unchanged. The discriminator receives two types of input: real images and images generated by the current generator. Its task is to classify the input images (real or fake) and optimize them using adversarial loss. The Huber loss function is used, which approximates mean squared error (MSE) when the error is small and mean absolute error (MAE) when the error is large, exhibiting better robustness and mitigating the instability problem in traditional GAN training. Next, in the generator update phase, the discriminator parameters remain unchanged, and the generator attempts to generate images that can "fool" the discriminator, causing it to classify them as "real." At this point, the generator not only relies on adversarial loss for updates but also considers similarity loss and content-aware loss simultaneously. Similarity loss measures the pixel-level difference between the generated and real images using the Smooth L1 loss function, making the generated image closer to the original image in structure and detail. Content-aware loss, on the other hand, takes a deep feature perspective, using pre-trained VGG19 and ResNet50 networks to extract local texture features and global semantic features of the image, respectively. The mean squared error between these feature maps is used to measure the consistency between the generated and real images in the perceptual space. The three loss functions are weighted and combined to form the generator's total loss function, with adversarial loss having a higher weight. = 0.7), emphasizing image realism, while similarity and content-aware loss are used as auxiliary means ( = 0.3), ensuring structural and semantic consistency.
[0057] Ultimately, the generator gradually learns to generate images that combine realism, color reproduction, clear details, and semantic consistency, while the discriminator continuously improves its ability to judge the realism of images. Through this interplay, both components work together to improve network performance and achieve high-quality image enhancement. The loss function formula is as follows: ; in, and It is a hyperparameter that balances the various loss terms.
[0058] First, we will briefly introduce the MobileNetV4 network, the multi-scale discriminator, and the loss function.
[0059] 1. MobileNetV4 network In underwater image enhancement tasks, generator models need to ensure image enhancement effects while possessing good operational efficiency and lightweight performance. Especially in underwater operations requiring real-time processing, lightweight design can effectively improve inference speed (FPS) and meet low latency requirements. However, traditional GAN models are complex in structure and have redundant parameters, making it difficult to meet both real-time and lightweight requirements. Therefore, this study designs a lightweight underwater image enhancement generator network based on the MobileNetV4 architecture, focusing on reducing model parameters and computational complexity.
[0060] The overall structure is based on MobileNetV4, constructing a lightweight and highly efficient image generation backbone network. The encoder mainly consists of multiple layers of MobileNetV4DownSample modules, progressively compressing image resolution and extracting features; the decoder adopts a symmetrical structure, with multiple MobileNetV4UpSample modules restoring the image size layer by layer, combined with lightweight convolutional structures to improve reconstruction quality and efficiency. The entire generation network significantly reduces the number of model parameters and computational complexity while maintaining image enhancement capabilities. The specific structure is as follows: Figure 1 As shown. Figure 1 The diagram illustrates the generative network structure of the proposed MobileNetV4 architecture. It primarily consists of an encoding path comprised of five MobileNetV4DownSample modules and a decoding path comprised of four MobileNetV4UpSample modules and one transposed convolutional layer. During the encoding stage, the spatial size of the input feature map continuously decreases while the number of channels gradually increases, facilitating the extraction of more abstract features layer by layer. Conversely, during the decoding stage, the size is gradually restored while the number of channels decreases layer by layer, thereby generating an enhanced result of the same size as the input image.
[0061] 2. Multi-scale discriminator In generative adversarial networks (GANs), another crucial module is the discriminator, which significantly impacts the generator's training stability and the quality of the enhanced images. To improve the model's ability to model complex degradation features of underwater images, this study designs a multi-scale discriminator (MSD) structure. This MSD meticulously discriminates the differences between enhanced and real images at multiple perceptual levels, thereby providing the generator with more discriminative optimization signals.
[0062] 3. Loss Function The underwater image enhancement model proposed in this paper mainly uses three loss functions: Adversarial Loss, Similarity Loss, and Content-Aware Loss. These loss functions work together to not only improve the generator's learning ability but also effectively improve the color reproduction, detail restoration, and overall visual quality of underwater images.
[0063] First, the introduction of adversarial loss enables the model to possess the capabilities of a Generative Adversarial Network (GAN). This is achieved through the generator... G and discriminator D In this adversarial game, the generator continuously learns to generate more realistic images. Unlike traditional cross-entropy loss, this paper uses Huber loss to measure the discriminator's judgment error between real and generated images. The Huber loss function combines the advantages of MSE (mean squared error) and MAE (mean absolute error). It approximates MSE (fine-tuning) for small errors and MAE (avoiding oscillations) for large errors, helping to alleviate the problem of overly smoothed images caused by traditional MSE, resulting in clearer and more realistic images. The specific mathematical expression is as follows: ; Secondly, similarity loss is used to improve the generated image. With real images Pixel-level similarity is assessed. This paper employs Smooth L1 loss to measure the pixel error between the generated and real images, ensuring that the enhanced image is structurally and detail-wise similar to the original. Its definition is as follows: ;
[0064] In the underwater image enhancement task presented in this paper, similarity loss... The calculation is as follows: ; Finally, in the underwater image enhancement model proposed in this paper, the content-aware loss employs a feature extraction mechanism combining VGG19 and ResNet50 to more comprehensively measure the deep feature similarity between the generated image and the real image. The core idea of this loss design is to leverage the fine-grained texture capture capability of the VGG19 network and the high-level semantic feature extraction capability of the ResNet50 network to make the generated image closer to the real image in both local details and global structure.
[0065] 4. The lightweight generative adversarial network for underwater image enhancement proposed in this invention is denoted as MWCA-GAN. Underwater image enhancement was performed using the MWCA-GAN model. This design demonstrated its superior enhancement capabilities. The training and optimization process of the MWCA-GAN model is as follows: High-quality underwater images are generated using MWCA-GAN. First, the features of the input image need to be acquired. Then, the generator and discriminator networks work together collaboratively, with the loss function as the optimization objective during training.
[0066] To address the issues of redundant parameters and slow inference speed in current models, this model's backbone adopts the efficient deep separable convolutional structure and universal inverted bottleneck module (UIB) from MobileNetV4. This significantly reduces model parameters and computational complexity while ensuring improved feature extraction capabilities. Furthermore, to enhance feature transfer efficiency and context modeling capabilities, a skip connection mechanism is introduced to fuse shallow details with deep semantic information, resulting in an overall structure that is both lightweight and expressive.
[0067] The following are the specific model design and optimization steps: (4.1) Design of generator and discriminator High efficiency, accuracy, and device compatibility are key features of MobileNetV4. Improved UIB and Mobile MQA modules accelerate inference on mobile devices, making it faster than previous models. Simultaneously, MobileNetV4 improves model accuracy through knowledge distillation. Furthermore, MobileNetV4 demonstrates high versatility and compatibility across various hardware devices, from CPUs and DSPs to dedicated GPU accelerators, making it suitable for lightweight deep learning tasks on mobile devices. This paper compares and evaluates several mainstream lightweight networks, including different versions of the MobileNet series, regarding the selection of the generator backbone network.
[0068] Table 1. Efficiency Comparison of Mainstream Lightweight Networks
[0069] As shown in Table 1, the MobileNetV4 version MNv4-Conv-S achieves a 73.8% ImageNet Top-1 classification accuracy while maintaining extremely low computational complexity (only 0.2G MACs) and latency (2.4ms), surpassing the computationally more computationally intensive MobileNet-V2 and significantly outperforming MobileNet-V3L-0.5× in accuracy. This model achieves a better trade-off between computational efficiency and model capacity. In contrast, MobileNetV2 and V1 are not advantageous in terms of parameters and latency, while the lightweight version of MobileNetV3 is insufficient in accuracy. Based on these considerations, this study ultimately chose MobileNetV4 as the encoding backbone network for the generator, maximizing image enhancement capabilities while maintaining lightweight characteristics. Compared to traditional convolutional neural network structures, MobileNetV4 achieves better model capacity and expressiveness with lower parameter and computational costs, which also demonstrates the pursuit of lightweight design and balanced trade-offs in this study.
[0070] (4.2) Comparison of discriminator structures To fully integrate the discriminative information from different scale branches, the discriminator uses bilinear interpolation to align the spatial dimensions of each scale output, and finally performs concatenation and fusion along the channel dimension to form a unified multi-scale discriminative result. This fusion method not only ensures the consistency of structural alignment across multiple scales but also enhances the complementarity between cross-scale features, enabling the discriminator to simultaneously consider both local texture and global structure of the image. Unlike traditional PatchGAN discriminators that only perform local discrimination within a fixed receptive field, the multi-scale discriminator proposed in this study uses a multi-resolution parallel path design to comprehensively model image content from different levels, thereby more effectively guiding the generator to achieve a balance between overall realism and local detail restoration. Furthermore, this structure employs a consistent discriminative subnetwork to handle inputs at different scales, ensuring both expressive power and lightweight computational efficiency.
[0071] Table 2. Experimental Results of Discriminator Structure Comparison
[0072] Step 5: Output Results This study conducted experimental evaluations on the representative EUVP dataset, comparing various underwater image enhancement methods mentioned above. Regarding performance evaluation metrics, this study employed two mainstream image quality assessment methods: one with reference images, involving Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM); and the other without reference images, involving Underwater Image Quality Assessment Metric (UIQM) and Underwater Color Image Quality Assessment Metric (UCIQE). Specific experimental results are summarized in Tables 3 and 4, respectively.
[0073] Table 3 Experimental results on the dataset
[0074] As shown in Table 3, the proposed MWCA-GAN performs best among all the comparative models, based on PSNR and SSIM metrics. Compared to the better-performing FUnIE-GAN and CAUIE, its PSNR is 1.43 dB and 0.85 dB higher, respectively, and its SSIM is 0.002 and 0.008 higher, significantly outperforming the traditional Pix2Pix model. Compared to UGAN-P, LS-GAN, and Res-GAN, MWCA-GAN improves PSNR by 5.76 dB, 7.52 dB, and 9.6 dB, respectively, and SSIM by 0.15, 0.146, and 0.35, indicating a significant advantage in structure reconstruction and detail restoration. Furthermore, while CAUIE and PUIENet are competitive in deep feature modeling, MWCA-GAN performs better in texture reconstruction and color fidelity; while the Water-Net model has a slight advantage in color correction, its overall image quality is still inferior to MWCA-GAN. Traditional methods such as Uw-HL and UDCP struggle to cope with complex underwater degradation in real-world scenarios, showing a significant gap compared to MWCA-GAN.
[0075] Table 4 shows the comparative experimental results of various underwater image enhancement models on the EUVP dataset on the no-reference evaluation metrics UIQM and UCIQE.
[0076] Table 4 shows the UIQM and UCIQE data from the comparative experiment on the EUVP dataset.
[0077] Table 5 compares the efficiency of the proposed MWCA-GAN model with several mainstream underwater image enhancement models, including the number of parameters involved in model training, computational cost (GFLOPs), average inference time, and inference speed (FPS). These metrics collectively reflect the model's performance in terms of resource consumption and operational efficiency, and are of significant reference value for the model's future feasibility in edge devices or practical deployment scenarios.
[0078] In terms of parameter count and computational cost, MWCA-GAN has the lowest model size among all models, containing only 0.37M parameters and 0.62 GFLOPs of computation, far lower than classic models such as Pix2Pix and CycleGAN, less than 1% of the latter. This lightweight design ensures enhanced quality while giving MWCA-GAN a significant advantage in model storage and computational cost.
[0079] Table 5. Efficiency Comparison Results of Various Underwater Image Enhancement Models
[0080] Experiments show that the classification performance of the model proposed in this invention is generally superior to other algorithms, mainly in terms of maximum, minimum, and average accuracy. Furthermore, the model does not require multiple epochs, thus its stability is also better than most other algorithms. In conclusion, the model proposed in this invention possesses excellent classification prediction performance.
[0081] The above description is only a part of the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the protection scope of the present invention.
Claims
1. An underwater image enhancement method based on an improved MobileNetV4, characterized in that, Includes the following steps: Construct a lightweight generative adversarial network for underwater image enhancement, comprising a generator network and a discriminator network; The generative network includes an encoder and a decoder. The encoder consists of several layers of MobileNetV4 DownSample modules and Universal Inverted Bottleneck (UIB) modules, which alternate to achieve multi-scale feature extraction and spatial downsampling. The final layer of the MobileNetV4 DownSample module encoder is sequentially connected to the wavelet denoising module and the color correction module, which are used to remove high-frequency noise and adaptively correct color deviation, respectively. The color correction module is then connected to the decoder; The decoder includes several layers of MobileNetV4 UpSample modules for progressively restoring resolution. LayerScale and MultiQuery Attention modules are interspersed among the MobileNetV4 UpSample modules to improve the stability of feature flow and the ability to perceive key regions; The last MobileNetV4 UpSample module is followed by a transposed convolution and a Tanh activation function, which outputs the generated image obtained by the generator network. The generated image and the unenhanced real image are input into the discriminator network to obtain the enhanced underwater image; The MobileNetV4DownSample module includes depthwise separable convolution, batch normalization, and ReLU activation operations, and achieves spatial size reduction through depthwise separable convolution. The MobileNetV4UpSample module includes transposed convolution, batch normalization, and ReLU activation operations. Transposed convolution increases the spatial size.
2. The underwater image enhancement method based on an improved MobileNetV4 as described in claim 1, characterized in that, The wavelet denoising module removes high-frequency noise and preserves image structural details through multi-scale decomposition; the color correction module uses a channel attention mechanism to achieve adaptive correction of color cast in underwater images; and the multi-query attention module enhances the decoder's ability to focus on key regions and details. The wavelet denoising module performs frequency domain decomposition on the feature map of the input wavelet denoising module through discrete wavelet transform (DWT), decomposing each channel into low-frequency and high-frequency components to remove high-frequency noise interference. The wavelet denoising module uses Haar wavelets as the basis function to extract features from low-frequency components through a separable convolutional structure containing 3×3 depth convolution and 1×1 point convolution. Finally, the spatial structure of the image is restored through inverse discrete wavelet transform (IDWT).
3. The underwater image enhancement method based on an improved MobileNetV4 as described in claim 1, characterized in that, The color correction module extracts global channel statistics from the input feature map through adaptive average pooling to capture the overall color distribution trend. Then, it generates channel attention weights through several layers of convolutional networks and introduces a learnable scaling factor to adjust the amplitude of the channel response.
4. The underwater image enhancement method based on improved MobileNetV4 as described in claim 1, characterized in that, The multi-query attention module generates a query vector through 1×1 convolution, generates a key vector and a value vector through depthwise separable convolution with dilation, then flattens the spatial dimension, calculates the correlation between query and key using a scaled dot product attention mechanism, obtains attention weights, performs weighted summation on the value vector to obtain context-enhanced features, and finally completes the projection output through 1×1 convolution.
5. The underwater image enhancement method based on an improved MobileNetV4 as described in claim 1, characterized in that, The LayerScale module is used to first initialize a scaling vector with the same number of input feature channels, and then expand the scaling vector to the same spatial dimension as the input features through a Broadcast operation. Finally, the expanded scaling parameters are multiplied element-wise with the input features of the LayerScale module.
6. The underwater image enhancement method based on improved MobileNetV4 as described in claim 1, characterized in that, A skip connection mechanism is used to directly pass the shallow high-resolution features of the encoder to the corresponding layer of the decoder.
7. The underwater image enhancement method based on an improved MobileNetV4 as described in claim 1, characterized in that, The generated image and the real image are then stitched together and input into the discriminator network; The discriminator comprises three parallel scale branches: a primary scale branch (Scale 1), a secondary scale branch (Scale 2), and a coarse scale branch (Scale 3), each branch used to capture image features at a specific level. The original scale branch Scale 1 is used to process the input image at the original resolution, including four cascaded downsampling modules, ZeroPad operation, and convolution operation; The medium-scale branch Scale 2 includes average pooling operations, four cascaded downsampling modules, ZeroPad operations, and convolution operations, used to process medium-resolution input images. The coarse-scale branch Scale 3 has the same structure as the medium-scale branch Scale 2, except that the average pooling operation is more extensive than that of the medium-scale branch Scale 2, and it is used to process input images with coarse resolution. Each downsampling module consists of: a 4×4 convolutional layer, a BN layer, and a LeakyReLU activation function; The feature maps obtained from the three parallel scale branches are fused using the Interpolate upsampling operation. The feature maps processed by each branch are restored to a uniform spatial size. Then, the channel dimensions are stitched together by the stitching operation to obtain a comprehensive feature representation, which in turn yields the enhanced underwater image.
8. The underwater image enhancement method based on an improved MobileNetV4 as described in claim 7, characterized in that, Also includes: The generator and discriminator are trained alternately, and the specific process includes: When training the discriminator, the generator parameters are fixed, and the discriminator is fed real images labeled as true labels and generated images labeled as false labels generated by the generator. Huber loss is used as adversarial loss to optimize the discriminator and enhance its ability to distinguish between real and false images. When training the generator, the discriminator parameters are fixed, and the generator is jointly optimized using adversarial loss, similarity loss, and content-aware loss.
9. The underwater image enhancement method based on improved MobileNetV4 as described in claim 8, characterized in that, Optimizing the generator using the adversarial loss includes: inputting noise into the generator, then labeling the fake images generated by the generator as real labels and sending them to the discriminator, and using Huber loss as adversarial loss to optimize the generator so that the discriminator's judgment of the fake images generated by the generator is closer to that of the real images. Optimizing the generator using the similarity loss includes: using the Smooth L1 loss function as the similarity loss to measure the pixel-level difference between the generated image and the real image, so that the generated image is close to the real image in structure and detail. Optimizing the generator using the content-aware loss includes: using a pre-trained image feature extraction model to extract image features from the generated image and the real image respectively, obtaining the generated image features and the real image features, calculating the mean square error of the generated image features and the real image features as the content-aware loss, and optimizing the generator; the image feature extraction model includes a VGG19 network and a ResNet50 network.
Citation Information
Cited By
HE image enhancement method and system based on biological mechanism and frequency attention
CN121837059A
Fundus retina image recognition method based on improved lightweight network
CN121861393A
Improved lightweight network-based fundus retina image recognition method
CN121861393B