A semi-supervised underwater image enhancement method based on multi-scale context perception
Through a multi-scale context-aware semi-supervised underwater image enhancement method, combining image recovery and detail repair branches, the color shift and blurring of underwater images are solved, image quality and model generalization capabilities are improved, and different underwater environments are adapted to different underwater environments.
Patent Information
- Application Number
- CN202510928768.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing underwater image enhancement technologies are difficult to accurately construct degraded models in complex and changeable underwater environments, resulting in problems such as color shift, low contrast and blurred details. The lack of high-quality samples based on deep learning methods leads to poor generalization capabilities of models.
Using a semi-supervised underwater image enhancement method based on multi-scale context perception, branches are repaired through image recovery branches and details, combining multi-scale input-level fusion module, multi-branch hybrid convolution attention residual module and pixel differential convolution, the network model is trained using semi-supervised learning strategy and joint loss function.
It significantly improves the visual quality of underwater images and the generalization performance of models in different underwater scenes, and can generate high-quality clear images without reference conditions, reducing the dependence on high-quality reference images.
Smart Images

Figure CN120430966B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater image processing, and in particular to a semi-supervised underwater image enhancement method based on multi-scale context perception. Background Art
[0002] Underwater optical images, as a key information carrier, contain rich color and texture features and are widely used in fields such as marine resource exploration, ecological protection, and underwater robot navigation. To accurately capture critical underwater information, common methods for improving underwater optical image quality include upgrading imaging equipment and using algorithms to enhance images. While directly upgrading underwater imaging equipment can improve imaging to some extent, it is costly and has limited effectiveness. In contrast, using algorithms to correct and enhance degraded underwater images is not only more cost-effective but also provides more comprehensive and accurate information support for various underwater operations. Underwater image enhancement technology aims to eliminate noise and color distortion in images to highlight the features of objects of interest while suppressing irrelevant background information, thereby obtaining high-quality underwater images. Currently, most traditional methods achieve underwater image enhancement by directly adjusting image pixels or simulating the underwater imaging process. Simulating the underwater imaging process is generally achieved by establishing a physical degradation model, then estimating the model parameters using various a priori assumptions, and finally reversing the degradation process to obtain a clear image. In some challenging underwater scenes, the color channels of optical images are significantly attenuated, making it difficult to accurately construct underwater degradation models and estimate model parameters. Therefore, this method is often only effective for specific types of underwater images, and is prone to introducing additional color casts and artifacts, resulting in under- or over-enhancement. Due to its reliance on artificially designed features and prior assumptions, traditional methods struggle to cope with the complex and ever-changing nature of real underwater environments. Deep learning-based methods, on the other hand, significantly improve underwater image enhancement and its adaptability to complex environments by constructing deep neural network models that automatically extract image features and learn the nonlinear mapping relationship from the original image to the reference image.
[0003] In existing technologies, due to the significant and complex differences in turbidity, depth, lighting and other conditions in different waters, and the selective attenuation and scattering effects of light in underwater environments, images are prone to degradation problems such as color cast, low contrast and blurred details, which are not conducive to subsequent image analysis and downstream visual tasks. Secondly, some image enhancement networks only focus on pixel-level color correction, ignore scene structure information, and have poor recovery effects on detail information such as texture and edges, resulting in unnatural enhanced images or the introduction of additional artifacts. In addition, deep learning-based methods are constrained by the number and quality of sample images, making it difficult to accurately fit nonlinear degradation processes such as underwater light attenuation-scattering coupling effects and dynamic turbidity changes, and the model has poor cross-domain generalization capabilities.
[0004] To address the blue-green cast, low contrast, and blurred details in underwater images, a multi-scale context-aware semi-supervised underwater image enhancement method is proposed. The model involved in this method consists of an image restoration branch and a detail restoration branch. In the image restoration branch, a multi-scale input-level fusion module is proposed to fuse shallow features of the original image at different scales to preserve sufficient background and context information. Considering the complex and diverse features of underwater images, a multi-branch hybrid convolutional attention residual module is designed to improve the model's ability to extract multi-scale features. In the detail restoration branch, multiple pixel differential convolutions are used to extract high-frequency features to preserve and enhance image details. A channel-aware feedforward neural network is then used for global integration to avoid the model's excessive focus on local details and neglect of overall semantic information. Due to the complex and changing underwater environment, the lack of high-quality paired training samples can hinder the performance of the underwater image enhancement network model. Therefore, a semi-supervised learning strategy is adopted to generate reliable pseudo-reference images for the original underwater images without reference. Then, a joint loss function consisting of a weighted combination of supervised loss and unsupervised loss is designed, so that the network can not only learn accurate pixel and color mapping on real paired data, but also maintain output consistency on pseudo reference data, thereby significantly improving the model's image enhancement effect and cross-scene generalization ability. Summary of the Invention
[0005] To address the severe degradation issues commonly found in underwater optical images, such as color cast, low contrast, and blurred details, this paper proposes a semi-supervised underwater image enhancement method based on multi-scale context awareness, overcoming the reliance on large amounts of paired training data. The constructed underwater image enhancement network model consists of two parts: an image restoration branch and a detail restoration branch, and the network is jointly trained using a semi-supervised learning strategy. This method not only improves the visual quality of underwater images but also enhances the model's generalization performance across diverse underwater scenarios.
[0006] The technical means adopted in the present invention are as follows:
[0007] A semi-supervised underwater image enhancement method based on multi-scale context perception includes the following steps:
[0008] Acquire an underwater image; establish an underwater image enhancement network model, wherein the underwater image enhancement network model includes a parallel image restoration branch and a detail restoration branch, wherein the image restoration branch includes a multi-scale input-level feature fusion module, a multi-branch hybrid convolution attention residual module, a downsampling module, and an upsampling module, wherein the image restoration branch is used to capture global and local contextual information in the image and restore the extracted feature map to the original resolution; the detail restoration branch includes a pixel differential convolution layer, a normalization layer, and a channel-aware feedforward neural network, wherein the detail restoration branch is used to calculate the gradient difference of pixels using differential convolution to highlight the detail features and edge information of the underwater image; the outputs of the image restoration branch and the detail restoration branch are added element by element to obtain the output of the underwater image enhancement network model; the underwater image enhancement network model is trained by a semi-supervised learning strategy and a joint loss function; the underwater image is input into the trained underwater image enhancement network model to obtain an enhanced underwater image.
[0009] Furthermore, the network architecture of the image restoration branch includes a multi-scale input-level feature fusion module, a first downsampling module, a first multi-branch mixed convolutional attention residual module, a second downsampling module, a second multi-branch mixed convolutional attention residual module, a third downsampling module, a third multi-branch mixed convolutional attention residual module, a first upsampling module, a second upsampling module, a third upsampling module, and an output layer connected in sequence; the first multi-branch mixed convolutional attention residual module, the second multi-branch mixed convolutional attention residual module and the third multi-branch mixed convolutional attention residual module have the same structure, the first downsampling module, the second downsampling module and the third downsampling module have the same structure, and the first upsampling module, the second upsampling module and the third upsampling module have the same structure.
[0010] Furthermore, the workflow of the image restoration branch is specifically as follows:
[0011] The underwater image is input into the multi-scale input level feature fusion module to obtain a fused feature map; the fused feature map is input into the first downsampling module and the first multi-branch mixed convolution attention residual module in series to obtain a first feature map; the first feature map is input into the second downsampling module and the second multi-branch mixed convolution attention residual module in series to obtain a second feature map; the second feature map is input into the third downsampling module and the third multi-branch mixed convolution attention residual module in series to obtain a third feature map; the third feature map is input into the first upsampling module, the third feature map is upsampled, and the upsampling operation is performed on the third feature map. The sampled third feature map is spliced with the second feature map to output a fourth feature map; the fourth feature map is input into the second upsampling module, the fourth feature map is upsampled, and the upsampled fourth feature map is spliced with the first feature map to output a fifth feature map; the fifth feature map is input into the third upsampling module, the fifth feature map is upsampled, and the upsampled fifth feature map is spliced with the fusion feature map to output a sixth feature map; the sixth feature map is passed through the output layer, which includes a first convolutional layer and a Tanh activation function, to obtain the output of the image restoration branch.
[0012] Furthermore, the workflow of the multi-scale input-level feature fusion module is specifically as follows:
[0013] The underwater image is input into three different receptive fields, and the outputs of the three different receptive fields are spliced to obtain the receptive field feature map; the receptive field feature map is subjected to convolution and Sigmoid function operations to obtain a fusion feature map.
[0014] Furthermore, the workflow of the first multi-branch hybrid convolutional attention residual module is as follows:
[0015] The feature map output by the first downsampling module is input into several parallel hybrid convolution modules, where the hybrid convolution module includes a second convolution layer, a spatial shift layer, an instance normalization layer and a GELU function connected in sequence. The output of the hybrid convolution module is added element by element and then input into a serially connected spatial attention module and a channel attention module to obtain a first feature map.
[0016] Furthermore, the network architecture of the detail restoration branch includes several detail restoration modules connected in series, and the detail restoration module includes a pixel difference convolution layer, a first normalization layer, a channel-aware feedforward neural network and a second normalization layer connected in sequence.
[0017] Furthermore, the workflow of the detail restoration module is specifically as follows:
[0018] The underwater image is input into the serial pixel difference convolution layer and the first normalization layer to obtain the seventh feature map; the underwater image and the seventh feature map are added element by element, and then input into the serial channel-aware feedforward neural network and the second normalization layer to obtain the eighth feature map; the seventh feature map and the eighth feature map are added element by element to obtain the output of the detail restoration module.
[0019] Furthermore, the pixel difference convolution layer includes a center pixel difference convolution layer, an angle pixel difference convolution layer, a vertical pixel difference convolution layer, a horizontal pixel difference convolution layer and a third convolution layer connected in parallel. The calculation method of the output feature map of the pixel difference convolution layer is as follows:
[0020] ,
[0021] Where i=1, 2, 3, 4, 5, - are the weights of the center pixel difference convolution layer, the angle pixel difference convolution layer, the vertical pixel difference convolution layer, the horizontal pixel difference convolution layer and the third convolution layer, is the feature map of the input pixel difference convolution layer, is the output feature map of the pixel difference convolution layer, is the combined convolution weight.
[0022] Furthermore, the channel-aware feedforward neural network includes a third convolutional layer, a depth-separable convolutional layer, a GELU function, a channel-aware module, and a fourth convolutional layer connected in series. The calculation method of the channel-aware feedforward neural network is as follows:
[0023] ,
[0024] in, is the input feature map, is the third convolutional layer, is a depth-wise separable convolutional layer, is the GELU function, is the channel perception module, is the fourth convolutional layer, is the output of the channel-aware feedforward neural network.
[0025] Furthermore, the underwater image enhancement network model is trained by a semi-supervised learning strategy and a joint loss function, including:
[0026] The labeled image in the underwater image is preprocessed, and the preprocessing step includes converting the underwater image into RGB format, adjusting the size of the RGB format image by Lanczos interpolation, and normalizing the resized image; the preprocessed labeled image is used as the input of the student model, and the output of the underwater image enhancement network model and the preprocessed labeled image are used to calculate the supervision loss. The calculation formula of the supervision loss is:
[0027] ,
[0028] in, For content loss, is the perceptual loss, is the underwater dark channel loss, is the edge loss, is the hyperparameter related to perceptual loss, is the hyperparameter related to underwater dark channel loss, is a hyperparameter related to edge loss, and the calculation formula of the content loss is as follows:
[0029] ,
[0030] in, For content loss, is the output image of the underwater image enhancement network model pixel values, The reference image pixel values, is the number of image pixels, and are the first weight coefficient and the second weight coefficient respectively. The perceptual loss is calculated by comparing the output of the underwater image enhancement network model and the features of the reference image through the pre-trained visual geometry group model. The perceptual loss is calculated as follows:
[0031] ,
[0032] in, is the perceptual loss, The first model of the pre-trained visual geometry group layer, is the number of layers of the selected visual geometry group model, is the number of pixels of the output feature map, is the output image of the underwater image enhancement network model, is the reference image, The output image of the underwater image enhancement network model is pre-trained visual geometry group model The first layer outputs the feature map pixel values, The reference image is pre-trained visual geometry group model The first layer outputs the feature map pixel values, the underwater dark channel loss is calculated by guiding the model to learn the prior knowledge of dark channels in underwater scenes. The calculation formula of the underwater dark channel loss is as follows:
[0033] ,
[0034] in, is the underwater dark channel loss, In pixels The dark channel value at In pixels The local area centered Indicates the color channel The pixel value on is the color channel, is the output image of the underwater image enhancement network model, is the reference image, Pixel The pixels in the local area centered at It is a green channel. The edge loss is a blue channel. The edge loss uses the Laplace operator to measure the difference in edge information between the enhanced image and the reference image, which is used to retain and enhance image detail information. The edge loss is calculated as follows:
[0035] ,
[0036] in, is the edge loss, is the output image of the underwater image enhancement network model, is the reference image; the teacher model weight is updated using the exponential moving average strategy, and the teacher model weight update process is as follows:
[0037] ,
[0038] in, is the teacher model weight to be updated in this round, is the teacher model weight of the previous round, is the student model weight, is the smoothing coefficient; the unlabeled image in the underwater image is strongly enhanced, the strongly enhanced unlabeled image is input into the student model, and the enhanced image is output; the unlabeled image in the underwater image is weakly enhanced, the weakly enhanced unlabeled image is input into the teacher model, and the pseudo reference image is output; the unsupervised loss is calculated using the enhanced image and the pseudo reference image:
[0039] ,
[0040] in, is the unsupervised loss, is the mean absolute error loss between the enhanced image and the pseudo reference image, is the third weight coefficient, The contrast loss is established by the pre-trained visual geometry group 19 model. The calculation process of the contrast loss is as follows:
[0041] ,
[0042] in, is the contrast loss, For the pre-trained visual geometry group 19 hidden layers, is the index of the 19 hidden layers of the pre-trained visual geometry group, For the The weight coefficient of the layer, is the total number of hidden layers, is the output image of the underwater image enhancement network model, is a pseudo reference image, is an exponential function with base e, To enhance the image, the supervised loss and unsupervised loss are combined and the adaptive moment estimation projection optimizer is used to update the student model parameters.
[0043] Compared with the prior art, the present invention has the following advantages:
[0044] This invention significantly improves the visual quality of underwater images through the synergistic enhancement of its two branches: image restoration and detail restoration. The image restoration branch corrects for color casts and reduced contrast found underwater, resulting in a more natural and richer overall image color distribution. The detail restoration branch extracts and integrates local high-frequency features, preserving and enhancing image details such as edges and textures, improving image clarity and the visibility of microscopic details.
[0045] The present invention makes full use of the original image without reference through a semi-supervised learning strategy and a joint loss function, which can reduce the dependence on high-quality reference images, improve the generalization ability of the underwater image enhancement network model in different complex underwater scenes, and enable the network to simultaneously adapt to degraded images under different water depths, water turbidity and lighting conditions.
[0046] Based on the above reasons, the present invention can be widely promoted in fields such as underwater image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0048] Figure 1 This is the network structure diagram of underwater image enhancement of the present invention.
[0049] Figure 2 This is a structural diagram of the multi-scale input-level feature fusion module of the present invention.
[0050] Figure 3 This is a structural diagram of the multi-branch hybrid convolutional attention residual module of the present invention.
[0051] Figure 4 This is a structural diagram of the pixel difference convolution layer of the present invention.
[0052] Figure 5 This is a structural diagram of the angle pixel difference convolution layer of the present invention.
[0053] Figure 6 This is a structural diagram of the channel-aware feedforward neural network of the present invention.
[0054] Figure 7 This is a diagram of the underwater image enhancement framework based on semi-supervised learning in the present invention.
[0055] Figure 8 Comparison of the visual effects of different image enhancement methods on UIEB.
[0056] Figure 9 A comparison chart of the performance of different image enhancement methods in feature point detection tasks.
[0057] Figure 10 Comparison chart of the performance of different image enhancement methods in feature point matching tasks. DETAILED DESCRIPTION
[0058] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0059] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0060] This paper proposes an underwater image enhancement network model that integrates image restoration and detail inpainting to address image degradation issues such as color distortion, low contrast, and blurred details in complex underwater environments. To fully utilize both reference and non-reference underwater images, the constructed network model is trained using a semi-supervised learning strategy and a joint loss function, improving its image enhancement performance and cross-scene generalization performance.
[0061] To address the degradation problems of underwater images such as color cast and low contrast, a multi-scale input-level feature fusion module is designed in the image restoration branch. By extracting and integrating shallow features at different scales, the network model's ability to express macroscopic structures and local semantics is improved. At the same time, a multi-branch hybrid convolutional attention residual module is proposed to enhance the network's feature extraction capabilities and focus on the illumination and color information in underwater images.
[0062] To address the problem of loss of original details caused by downsampling and upsampling operations in convolutional neural networks, the detail restoration branch introduces multiple pixel differential convolutions to extract high-frequency features from the original image, thereby enhancing image detail and texture features. Furthermore, to prevent the model from over-focusing on local details and neglecting overall semantic information, a channel-aware feedforward neural network is used to integrate global information.
[0063] like Figure 1-7 As shown, the present invention provides a semi-supervised underwater image enhancement method based on multi-scale context perception, the specific steps are as follows:
[0064] S1. Acquire underwater images.
[0065] S2. Establish an underwater image enhancement network model. The underwater image enhancement network model includes a parallel image restoration branch and a detail restoration branch. The image restoration branch adopts an encoder-decoder network architecture, including a multi-scale input-level feature fusion module, a multi-branch hybrid convolutional attention residual module, a downsampling module, an upsampling module, and a convolutional layer. The image restoration branch is used to capture global and local contextual information in the image and restore the extracted feature map to the original resolution. The detail restoration branch includes a pixel difference convolution layer, a normalization layer, and a channel-aware feedforward neural network. The detail restoration branch uses differential convolution to calculate pixel gradient differences, highlighting the detailed features and edge information of underwater images.
[0066] Figure 1 This is the overall structure diagram of the underwater image enhancement network model of the present invention. Figure 1 As shown in Figure 1, the underwater image enhancement network model consists of two parallel branches: image restoration and detail inpainting. The image restoration branch is responsible for capturing global and local contextual information in the image and restoring the extracted feature maps to the original resolution. The detail inpainting branch uses differential convolution to calculate pixel gradient differences, highlighting detailed features and edge information in the underwater image. Finally, the feature maps output by the image restoration and detail inpainting branches are fused through element-by-element addition to produce the desired high-quality, clear underwater image.
[0067] The image restoration branch consists of three main parts: a multi-scale input-level feature fusion module, an encoder, and a decoder. At the network input, shallow features of different scales of the original image are extracted separately, and multi-scale information is integrated through fusion operations to provide rich contextual representation for the subsequent encoder. The encoder contains multiple downsampling operations and a multi-branch hybrid convolutional attention residual module, which can capture deep semantic and contextual information in both global and local areas and pass key features to the decoder through skip connections. The decoder uses a series of upsampling layers to gradually map the low-resolution, high-semantic feature maps output by the encoder back to the original resolution, gradually restoring the details and local information of the image.
[0068] The detail restoration branch primarily consists of various pixel difference convolutions (PDConv), layer normalization (LN), and a channel attention feed-forward neural network (CA-FFN). By stacking these modules, this branch enhances the texture and edge structure features of underwater images, providing precise detail compensation for the final output fusion.
[0069] The network architecture of the image restoration branch consists of a sequentially connected multi-scale input-level feature fusion module, a first downsampling module, a first multi-branch hybrid convolutional attention residual module, a second downsampling module, a second multi-branch hybrid convolutional attention residual module, a third downsampling module, a third multi-branch hybrid convolutional attention residual module, a first upsampling module, a second upsampling module, a third upsampling module, and an output layer. The first, second, and third multi-branch hybrid convolutional attention residual modules have the same structure, as do the first, second, and third downsampling modules. The first, second, and third upsampling modules have the same structure. The downsampling module uses a convolution operation with a kernel size of 3×3, a stride of 2, and a padding of 1. Compared to traditional pooling layers, it preserves more complete image information and can extract more complex features, improving the model's expressiveness. The upsampling module uses transposed convolution with a convolution kernel size of 3×3, a stride of 2, a padding of 1, and an additional output padding of 1. Compared with traditional interpolation methods, it has trainable convolution kernel weights and can better adapt to different types of underwater image restoration.
[0070] Specifically, the workflow of the image restoration branch is as follows:
[0071] Step 1: Input the underwater image into the multi-scale input-level feature fusion module to obtain the fused feature map.
[0072] Step 2: Input the fused feature map into the first downsampling module and the first multi-branch hybrid convolutional attention residual module in series to obtain the first feature map with 128 channels.
[0073] Step 3: Input the first feature map into the second downsampling module and the second multi-branch mixed convolutional attention residual module in series to obtain a second feature map with 256 channels.
[0074] Step 4: Input the second feature map into the third downsampling module and the third multi-branch mixed convolution attention residual module in series to obtain a third feature map with a channel of 512.
[0075] Step 5: Input the third feature map into the first upsampling module, perform an upsampling operation on the third feature map, and concatenate the upsampled third feature map with the second feature map. After 1×1 convolution, the output channel is 256, which is a fourth feature map.
[0076] Step 6: Input the fourth feature map into the second upsampling module, perform an upsampling operation on the fourth feature map, and concatenate the upsampled fourth feature map with the first feature map. After 1×1 convolution, the output channel is 128, resulting in a fifth feature map.
[0077] Step 7: Input the fifth feature map into the third upsampling module, perform an upsampling operation on the fifth feature map, and concatenate the upsampled fifth feature map with the fusion feature map. After 1×1 convolution, the output channel is 64, which is the sixth feature map.
[0078] Step 8: Pass the sixth feature map through the output layer to obtain the output of the image restoration branch. The output has the same shape as the original underwater image. The output layer uses a first convolutional layer with a convolution kernel size of 3×3, a stride of 1, and padding of 1, and a Tanh activation function to adjust the output image channels and map pixel values to the range [-1, 1]. The output of the image restoration branch restores the true structure and color of the original image, providing complete and accurate semantic information for the detail restoration branch.
[0079] Figure 2 This is the architecture of the multi-scale input-level feature fusion module. The module's workflow is as follows: An underwater image is input into three different receptive fields. The three branches with different receptive fields learn local and global features from the input image. The outputs from the three different receptive fields are concatenated to generate a receptive field feature map. The receptive field feature map is then subjected to convolution and sigmoid function operations. The convolution operation automatically learns and generates a weight tensor for multi-scale features. The sigmoid function constrains the weights to the range [0, 1], enabling dynamic feature fusion. The final output is a fused feature map with 64 channels.
[0080] Figure 3 This is the architecture of a multi-branch hybrid convolutional attention residual module. The first multi-branch hybrid convolutional attention residual module works as follows: the feature map output by the first downsampling module is fed into several parallel hybrid convolutional modules, generating feature maps with both deep semantics and original information and shallow details from top to bottom. The hybrid convolutional module consists of a second convolutional layer, a spatial shift layer, an instance normalization (IN) layer, and a GELU function, connected in sequence. The second convolutional layer and the spatial shift layer are used to reduce model parameters and optimize inference speed. The second convolutional layer uses a 1×1 convolution. The output of the hybrid convolutional module is element-wise summed and then fed into the serially connected spatial attention module and the first channel attention module. This enhances the network model's feature extraction capabilities while improving its focus on color and illumination information. This module uses residual connections to prevent network gradient vanishing and maximize the preservation of important spatial structure and details in the original underwater image. Finally, the output is a first feature map with 128 channels.
[0081] Because the image restoration stage uses multiple consecutive downsampling and upsampling operations, the detailed features of the original image are lost, which affects the quality of the enhanced image. To preserve the details of the original image, a residual connection can be used to add the input and output images element by element. However, simply adding a severely degraded underwater image to the output image of the image restoration stage may result in poor final image enhancement. To address this issue, a detail restoration branch is added. By stacking multiple residual modules consisting of a pixel difference convolutional layer (PDConv), a normalization layer (LN), and a channel-aware feedforward neural network (CA-FFN), the detailed features of the underwater image are preserved and enhanced as much as possible.
[0082] The network architecture of the detail restoration branch includes several detail restoration modules connected in series, and the detail restoration module includes a pixel difference convolution layer, a first normalization layer, a channel-aware feedforward neural network and a second normalization layer connected in sequence.
[0083] The workflow of the detail repair module is as follows:
[0084] Step 1: Input the underwater image into a series of pixel-by-pixel difference convolutional layers and a first normalization layer to generate the seventh feature map. The pixel-by-pixel difference convolutional layer (PDConv) calculates the differences between adjacent pixels in the input feature map, extracting high-frequency features such as edges and textures to restore and enhance blurred details in the original underwater image. Furthermore, the first normalization layer (LN) prevents vanishing or exploding gradients, improving model training stability.
[0085] Step 2: After element-by-element addition of the underwater image and the seventh feature map to fully preserve details such as edges and textures, the image is fed into a series of channel-aware feedforward neural networks and a second normalization layer. The channel-aware feedforward neural network (CA-FFN) is used to globally integrate the extracted local detail features, preventing the network model from over-focusing on local information and neglecting overall semantic information. Similarly, a second normalization layer (LN) is used to stabilize the model training process. Finally, the eighth feature map is obtained.
[0086] Step 3: Add the seventh feature map and the eighth feature map element by element to obtain the output of the detail restoration module, which provides accurate texture and edge compensation for the subsequent feature fusion with the image restoration branch.
[0087] like Figure 4 As shown in the figure, the pixel difference convolution layer includes a parallel connection of the center pixel difference convolution layer, the angle pixel difference convolution layer, the vertical pixel difference convolution layer, the horizontal pixel difference convolution layer and the third convolution layer. The pixel difference convolution is mainly used to enhance the gradient change of the local area, while the standard convolution is responsible for global feature extraction, avoiding the loss of global information due to excessive focus on local differences, and providing semantic guidance for detail restoration. Taking the angle pixel difference convolution as an example, as shown in the figure, Figure 5 As shown in the figure, considering that connecting multiple convolutional layers in parallel will inevitably lead to an increase in parameters and an increase in inference time, the detail restoration branch takes advantage of the additivity of the convolutional layers and merges them into a standard convolution, maintaining the model's expressiveness while improving its inference efficiency. Overall, the calculation method for the output feature map of the pixel difference convolution layer is as follows:
[0088] ,
[0089] Where i=1, 2, 3, 4, 5, - are the weights of the center pixel difference convolution layer, the angle pixel difference convolution layer, the vertical pixel difference convolution layer, the horizontal pixel difference convolution layer and the third convolution layer, is the feature map of the input pixel difference convolution layer, is the output feature map of the pixel difference convolution layer, is the combined convolution weight.
[0090] Figure 6 This is the structural diagram of the channel-aware feedforward neural network (CA-FFN). The channel-aware feedforward neural network includes the third convolutional layer, the depth-separable convolutional layer, the GELU function, the channel-aware module and the fourth convolutional layer connected in series. First, the input feature channel is converted from C Expand to rC , increase the channel capacity to capture more complex features. Secondly, a depth-wise separable convolution with a convolution kernel size of 3×3 is used to aggregate the extracted local information in the spatial dimension while maintaining the independence of the channels. Then the GELU activation function is introduced to enhance the nonlinear expression ability of the model. Compared with other activation functions, the smooth gradient characteristics of GELU are more suitable for detail repair tasks. Then the channel perception module is used to dynamically aggregate cross-channel information, automatically adjust the channel weight according to the input content, suppress the noise generated by channel expansion, highlight the important feature information in the channel dimension, and improve its adaptability to complex underwater environments. Finally, the number of channels is increased from 1×1 through 1×1 convolution. rC Restore to C , the calculation method of the channel-aware feedforward neural network is as follows:
[0091] ,
[0092] in, is the input feature map, is the third convolutional layer, is a depth-wise separable convolutional layer, is the GELU function, is the channel perception module, is the fourth convolutional layer, is the output of the channel-aware feedforward neural network.
[0093] First, the number of channels is compressed to 1 through 1×1 convolution to generate a global spatial attention map and focus on the key areas. Then, the GELU activation function is used to perform nonlinear transformation on the compressed features. Finally, a learnable scaling factor is introduced to dynamically fuse the original features with the compressed features. The calculation formula of the channel perception module (CA) is as follows:
[0094] ,
[0095] in, is the input feature map, is the output of the channel perception module, is the GELU function, To learn the scaling factor, It is a 1×1 convolution.
[0096] S3. Add the outputs of the image restoration branch and the detail restoration branch element by element to obtain the output of the underwater image enhancement network model.
[0097] S4. Training the underwater image enhancement network model via a semi-supervised learning strategy and a joint loss function.
[0098] In order to solve the pain points of the above-mentioned network model's excessive reliance on paired data and insufficient cross-domain generalization ability, a training strategy based on semi-supervised learning is adopted to significantly improve the model's enhancement effect and generalization performance by making full use of reference and non-reference data. Figure 8 As shown in the figure, the underwater image enhancement framework based on semi-supervised learning consists of a student model and a teacher model, both of which use a network model with the same structure. Among them, the student model is responsible for learning the underwater image features, and the teacher model is responsible for generating high-quality pseudo reference images for underwater images without reference. Specifically, the network model training process based on semi-supervised learning is as follows: Step 1: Preprocess the labeled image in the underwater image. The preprocessing steps include converting the underwater image into RGB format, adjusting the size of the RGB format image by Lanczos interpolation, and normalizing the resized image. Step 2: Use the preprocessed labeled image as the input of the student model, and use the output of the underwater image enhancement network model and the preprocessed labeled image to calculate the supervision loss. The calculation formula of the supervision loss is:
[0099] ,
[0100] in, For content loss, is the perceptual loss, is the underwater dark channel loss, is the edge loss, is the hyperparameter related to perceptual loss, is the hyperparameter related to underwater dark channel loss, is a hyperparameter related to the marginal loss. While adjusting the loss weight, each loss function is kept at the same order of magnitude to ensure its effectiveness.
[0101] To ensure that the enhanced image is consistent with the reference image at the pixel level, a single mean absolute error loss ( ) and mean square error loss ( ) may cause the network to overfit the training data, combined with and As the content loss, the calculation formula of content loss is as follows:
[0102] ,
[0103] in, For content loss, is the output image of the underwater image enhancement network model pixel values, The reference image pixel values, is the number of image pixels, and are the first weight coefficient and the second weight coefficient respectively.
[0104] Perceptual loss compares the features of the enhanced image with those of the reference image using a pre-trained Visual Geometry Group (VGG)16 model to ensure that the enhanced image is semantically and visually close to the reference image. VGG is the name of a CNN network model. VGG stands for Visual Geometry Group. The perceptual loss is calculated as follows:
[0105] ,
[0106] in, is the perceptual loss, The first of 16 models in the pre-trained visual geometry group layer, is the number of layers of the selected visual geometry group 16 model, is the number of pixels of the output feature map, is the output image of the underwater image enhancement network model, is the reference image, The output image of the underwater image enhancement network model is pre-trained visual geometry group 16 model The first layer outputs the feature map pixel values, The reference image is pre-trained visual geometry group 16 models The first layer outputs the feature map pixel values.
[0107] To generate more realistic and clearer underwater images, the underwater dark channel loss removes the color distortion and blurring effects of underwater images by guiding the model to learn the prior knowledge of dark channels in underwater scenes. The calculation formula of the underwater dark channel loss is as follows:
[0108] ,
[0109] in, is the underwater dark channel loss, In pixels The dark channel value at In pixels The local area centered Indicates the color channel The pixel value on is the color channel, is the output image of the underwater image enhancement network model, is the reference image, Pixel The pixels in the local area centered at It is a green channel. For the blue channel.
[0110] Edge loss uses the Laplace operator to measure the difference in edge information between the enhanced image and the reference image, which is used to preserve and enhance image detail information. The edge loss is calculated as follows:
[0111] ,
[0112] in, is the edge loss, is the output image of the underwater image enhancement network model, is the reference image.
[0113] Step 3: Update the teacher model weights using the Exponential Moving Average (EMA) strategy. The EMA strategy can reduce sudden changes during the weight update process, thereby improving the stability of the model and helping the teacher model generate reliable pseudo reference images. The teacher model weight update process is as follows:
[0114] ,
[0115] in, is the teacher model weight to be updated in this round, is the teacher model weight of the previous round, is the student model weight, is the smoothing coefficient, which is responsible for controlling the retention ratio of historical parameters of the student and teacher models.
[0116] Step 4: Use strong enhancements such as Gaussian blur and color perturbation to strongly enhance the unlabeled image in the underwater image, input the strongly enhanced unlabeled image into the student model, and output the enhanced image.
[0117] Step 5: Perform weak augmentation with random flipping on the unlabeled image. This weakly augmented unlabeled image is fed into the teacher model, and the generated images are screened to output a pseudo-reference image. To ensure the reliability of the pseudo-reference images generated by the teacher model, a comprehensive pseudo-reference screening mechanism is established, and a no-reference image quality evaluation metric is selected to assess image quality. Only images output by the teacher model that are of better quality than the output image from Step 4 and previously generated pseudo-reference images are considered reliable pseudo-reference images.
[0118] Step 6: Using the enhanced image and the pseudo reference image, calculate the unsupervised loss:
[0119] ,
[0120] in, is the unsupervised loss, is the enhanced image and pseudo reference image loss, is the third weight coefficient, is the contrast loss.
[0121] The contrast loss is established by pre-training VGG19 model, which converts high-quality pseudo reference images As a positive sample, select the strongly enhanced underwater image As a negative sample, the calculation process of contrast loss is as follows:
[0122] ,
[0123] in, is the contrast loss, For the pre-trained visual geometry group 19 hidden layers, is the index of the 19 hidden layers of the pre-trained visual geometry group, For the The weight coefficient of the layer, is the total number of hidden layers, is the output image of the underwater image enhancement network model, is a pseudo reference image, is an exponential function with base e, is the enhanced image. Damage assessment enhanced image Minimizing the contrast loss between positive and negative samples in the high-level feature space can make the output image of the network model close to the positive sample and away from the negative sample.
[0124] Step 7: Combine the supervised loss and unsupervised loss and use the Adaptive Moment Estimation Projection (AdamP) optimizer to update the student model parameters to make full use of limited paired data and a large amount of unpaired data, balance training stability and convergence speed, and thus significantly improve the effect and generalization ability of the underwater image enhancement network model in complex scenes.
[0125] S5. Input the underwater image into the trained underwater image enhancement network model to obtain an enhanced underwater image.
[0126] To demonstrate the effectiveness of the underwater image enhancement method based on semi-supervised learning proposed in this invention, experiments were conducted on the real underwater dataset benchmark UIEB dataset. The experimental settings are as follows: the batch size is set to 2, and the number of training rounds is 200. The initial learning rate of the optimizer AdamP is 2e-4. When the model is trained to 100 and 150 rounds, the initial learning rate is reduced to 0.1 times the original to prevent the model from overfitting due to excessive learning rate in the later stage. In order to comprehensively verify the advantages of the method proposed in this invention, the two traditional methods Sea-Thru and MLLE are compared with the two deep learning-based methods FUnIE-GAN and HCLR:
[0127] 1. Sea-Thru method proposed by Akkaynak and Treibitz, reference "D. Akkaynak and T.Treibitz, 'Sea-thru: A method for removing water from underwater images,' CVPR 2019, pp. 1682–1691"
[0128] 2. MLLE method proposed by Zhang et al., reference “W. Zhang, P. Zhuang, HH Sun et al., "Underwater image enhancement via minimal color loss and locally adaptive contrast enhancement," IEEE Transactions on Image Processing, vol. 31, pp. 3997–4010, 2022.”
[0129] 3. FUnIE-GAN method proposed by Islam et al., referenced in “MJ Islam, Y. Xia, J. Sattar, "Fast underwater image enhancement for improved visual perception," IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3227-3234, 2020.”
[0130] 4. HCLR method proposed by Zhou et al., reference “J. Zhou, J. Sun, C. Li et al., "HCLR-Net: Hybrid contrastive learning regularization with locally randomized perturbation for underwater image enhancement," International Journal of Computer Vision, vol. 132, no. 10, pp. 4132–4156, 2024.”
[0131] The experimental results are as follows:
[0132] Table 1 shows the comparative experimental results of different underwater image enhancement methods on the UIEB dataset. In order to measure the similarity and quality difference between the enhanced image and the reference image, the full reference evaluation index is used, including PSNR. The larger the value, the closer the enhanced image is to the reference image, which means the higher the quality. For underwater images without reference images, the non-reference evaluation index UIQM is used to evaluate the color shift, blur and low contrast of underwater images. The larger the value, the clearer the image, the higher the contrast and the more balanced the color. As shown in Table 1, the method of the present invention is significantly better than other methods in terms of evaluation indicators, indicating that the enhanced image is closer to the reference image in color, contrast and brightness, and is more in line with human visual perception.
[0133] Table 1 Performance comparison of different underwater image enhancement methods on the UIEB dataset
[0134]
[0135] Figure 8 The image enhancement effects of the method of the present invention and four other methods in subjective evaluation are compared. Figure 8As shown in Figure 2, the proposed method corrects the blue-green cast in underwater images, improves overall contrast, and is suitable for challenging low-light underwater scenes. Furthermore, the enhanced images have clearer, richer details and texture information.
[0136] Figure 9 The results of feature point detection on the original underwater image and the image enhanced by different methods are shown. The feature points are mainly distributed at corners, edges and areas with complex textures. The more feature points detected, the more detailed information and areas with obvious gradient changes in the image. Figure 9 As shown, the method of the present invention not only detects the most feature points, but also distributes them more evenly, indicating that it performs outstandingly in retaining image structural information, improving contrast, and restoring details.
[0137] Figure 10 The results of feature point matching on the original underwater image and the image enhanced by different methods are shown. After the underwater image enhancement algorithm is processed, the quality of the original underwater image is significantly improved, and the feature point extraction and matching performance is also enhanced. Figure 10 As shown in the figure, the method of the present invention performs outstandingly in the feature point matching task, which not only improves the visual quality of underwater images, but also enhances the semantic information and structural features of the images, and can extract more stable and distinguishable feature points.
[0138] In summary, the underwater image enhancement method based on semi-supervised learning proposed in this paper makes full use of paired and unpaired underwater image data, achieves significant improvements in color restoration, contrast enhancement and detail restoration, has strong generalization performance, and is beneficial to subsequent underwater vision tasks.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised underwater image enhancement method based on multi-scale context perception, characterized in that: The following steps are involved: Acquire underwater images; An underwater image enhancement network model is established, comprising a parallel image restoration branch and a detail restoration branch. The image restoration branch comprises a multi-scale input-level feature fusion module, a multi-branch hybrid convolutional attention residual module, a downsampling module, and an upsampling module. The image restoration branch is used to capture global and local contextual information in the image and restore the extracted feature map to its original resolution. The detail restoration branch comprises a pixel differential convolution layer, a normalization layer, and a channel-aware feedforward neural network. The detail restoration branch is used to calculate pixel gradient differences using differential convolution to highlight the detail features and edge information of the underwater image. Adding the outputs of the image restoration branch and the detail restoration branch element by element to obtain the output of the underwater image enhancement network model; Training the underwater image enhancement network model through a semi-supervised learning strategy and a joint loss function; The underwater image is input into a trained underwater image enhancement network model to obtain an enhanced underwater image.
2. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 1, characterized in that: The network architecture of the image restoration branch includes a multi-scale input level feature fusion module, a first downsampling module, a first multi-branch hybrid convolutional attention residual module, a second downsampling module, a second multi-branch hybrid convolutional attention residual module, a third downsampling module, a third multi-branch hybrid convolutional attention residual module, a first upsampling module, a second upsampling module, a third upsampling module, and an output layer connected in sequence; The first multi-branch mixed convolutional attention residual module, the second multi-branch mixed convolutional attention residual module and the third multi-branch mixed convolutional attention residual module have the same structure, the first downsampling module, the second downsampling module and the third downsampling module have the same structure, and the first upsampling module, the second upsampling module and the third upsampling module have the same structure.
3. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 2, characterized in that: The workflow of the image restoration branch is specifically as follows: Inputting the underwater image into the multi-scale input-level feature fusion module to obtain a fused feature map; Inputting the fused feature map into a first downsampling module and a first multi-branch hybrid convolutional attention residual module connected in series to obtain a first feature map; Inputting the first feature map into a second downsampling module and a second multi-branch hybrid convolutional attention residual module in series to obtain a second feature map; Inputting the second feature map into a third downsampling module and a third multi-branch hybrid convolutional attention residual module in series to obtain a third feature map; Inputting the third feature map into a first upsampling module, performing an upsampling operation on the third feature map, and concatenating the upsampled third feature map with the second feature map to output a fourth feature map; Inputting the fourth feature map into a second upsampling module, performing an upsampling operation on the fourth feature map, and concatenating the upsampled fourth feature map with the first feature map to output a fifth feature map; Inputting the fifth feature map into a third upsampling module, performing an upsampling operation on the fifth feature map, and concatenating the upsampled fifth feature map with the fused feature map to output a sixth feature map; The sixth feature map is passed through the output layer, where the output layer includes a first convolutional layer and a Tanh activation function, to obtain an output of the image restoration branch.
4. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 3, characterized in that: The workflow of the multi-scale input-level feature fusion module is as follows: The underwater image is input into three different receptive fields, and the outputs of the three different receptive fields are spliced to obtain the receptive field feature map; The receptive field feature map is convolved and subjected to Sigmoid function operations to obtain a fused feature map.
5. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 3, characterized in that: The workflow of the first multi-branch hybrid convolutional attention residual module is as follows: Input the feature map output by the first downsampling module into several parallel hybrid convolution modules, where the hybrid convolution module includes a second convolution layer, a spatial shift layer, an instance normalization layer, and a GELU function connected in sequence; The outputs of the hybrid convolution module are added element by element and then input into the serially connected spatial attention module and channel attention module to obtain the first feature map.
6. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 1, characterized in that: The network architecture of the detail restoration branch includes several detail restoration modules connected in series, and the detail restoration module includes a pixel difference convolution layer, a first normalization layer, a channel-aware feedforward neural network and a second normalization layer connected in sequence.
7. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 6, characterized in that: The workflow of the detail repair module is as follows: Input the underwater image into the serial pixel difference convolution layer and the first normalization layer to obtain the seventh feature map; After element-by-element addition of the underwater image and the seventh feature map, the resultant addition is input into a channel-aware feedforward neural network and a second normalization layer connected in series to obtain an eighth feature map; The seventh feature map and the eighth feature map are added element by element to obtain the output of the detail restoration module.
8. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 7, characterized in that: The pixel difference convolution layer includes a central pixel difference convolution layer, an angle pixel difference convolution layer, a vertical pixel difference convolution layer, a horizontal pixel difference convolution layer and a third convolution layer connected in parallel. The calculation method of the output feature map of the pixel difference convolution layer is as follows: , Where i=1, 2, 3, 4, 5, - are the weights of the center pixel difference convolution layer, the angle pixel difference convolution layer, the vertical pixel difference convolution layer, the horizontal pixel difference convolution layer and the third convolution layer, is the feature map of the input pixel difference convolution layer, is the output feature map of the pixel difference convolution layer, is the combined convolution weight.
9. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 7, characterized in that: The channel-aware feedforward neural network includes a third convolutional layer, a depth-separable convolutional layer, a GELU function, a channel-aware module, and a fourth convolutional layer connected in series. The calculation method of the channel-aware feedforward neural network is as follows: , in, is the input feature map, is the third convolutional layer, is a depth-wise separable convolutional layer, is the GELU function, is the channel perception module, is the fourth convolutional layer, is the output of the channel-aware feedforward neural network.
10. The method for semi-supervised underwater image enhancement based on multi-scale context perception according to claim 1, characterized in that: The underwater image enhancement network model is trained by a semi-supervised learning strategy and a joint loss function, comprising: Preprocessing the labeled image in the underwater image, the preprocessing step comprising converting the underwater image into RGB format, resizing the image in the RGB format using a Lanczos interpolation method, and normalizing the resized image; The preprocessed labeled image is used as the input of the student model, and the output of the underwater image enhancement network model and the preprocessed labeled image are used to calculate the supervision loss. The calculation formula of the supervision loss is: , in, For content loss, is the perceptual loss, is the underwater dark channel loss, is the edge loss, is the hyperparameter related to perceptual loss, is the hyperparameter related to underwater dark channel loss, is the hyperparameter related to the marginal loss, The calculation formula of the content loss is as follows: , in, For content loss, is the output image of the underwater image enhancement network model pixel values, The reference image pixel values, is the number of image pixels, and are the first weight coefficient and the second weight coefficient respectively, The perceptual loss is calculated by comparing the output of the underwater image enhancement network model and the features of the reference image through a pre-trained visual geometry group model. The perceptual loss is calculated as follows: , in, is the perceptual loss, The first model of the pre-trained visual geometry group layer, is the number of layers of the selected visual geometry group model, is the number of pixels of the output feature map, is the output image of the underwater image enhancement network model, is the reference image, The output image of the underwater image enhancement network model is pre-trained visual geometry group model The first layer outputs the feature map pixel values, The reference image is pre-trained visual geometry group model The first layer outputs the feature map pixel values, The underwater dark channel loss is calculated by guiding the model to learn the prior knowledge of dark channels in underwater scenes. The calculation formula of the underwater dark channel loss is as follows: , in, is the underwater dark channel loss, In pixels The dark channel value at In pixels The local area centered Indicates the color channel The pixel value on is the color channel, is the output image of the underwater image enhancement network model, is the reference image, Pixel The pixels in the local area centered at It is a green channel. is the blue channel, The edge loss uses the Laplace operator to measure the difference in edge information between the enhanced image and the reference image, and is used to retain and enhance image detail information. The edge loss is calculated as follows: , in, is the edge loss, is the output image of the underwater image enhancement network model, is the reference image; The teacher model weight is updated using the exponential moving average strategy. The teacher model weight update process is as follows: , in, is the teacher model weight to be updated in this round, is the teacher model weight of the previous round, is the student model weight, is the smoothing coefficient; Strongly enhancing an unlabeled image in the underwater image, inputting the strongly enhanced unlabeled image into a student model, and outputting an enhanced image; Weakly enhancing an unlabeled image in the underwater image, inputting the weakly enhanced unlabeled image into a teacher model, and outputting a pseudo reference image; Using the enhanced image and the pseudo reference image, calculate the unsupervised loss: , in, is the unsupervised loss, is the mean absolute error loss between the enhanced image and the pseudo reference image, is the third weight coefficient, is the contrast loss, The contrast loss is established by a pre-trained visual geometry group model. The calculation process of the contrast loss is as follows: , in, is the contrast loss, The first one for the pre-trained visual geometry group hidden layers, is the index of the hidden layer of the pre-trained visual geometry group, For the The weight coefficient of the layer, is the total number of hidden layers, is the output image of the underwater image enhancement network model, is a pseudo reference image, is an exponential function with base e, is the enhanced image; The supervised loss and unsupervised loss are combined and an adaptive moment estimation projection optimizer is used to update the student model parameters.
Citation Information
Patent Citations
Semi-supervised medical image segmentation method and system based on visual language model
CN118115516A
Medical image automatic segmentation method based on semi-supervised learning
CN119205807A