Double-stage underwater image enhancement method based on deep learning
Through the two-stage underwater image enhancement method, combined with the encoder decoder architecture of CNN and Transformer and reasonable loss function, the underwater image clarity and color problems are solved, and a higher quality underwater image is generated, suitable for engineering applications.
Patent Information
- Application Number
- CN202510445117.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively restore the clarity and color authenticity of underwater images. Traditional methods are not effective under the complexity of underwater environments. Deep learning methods still have room for improvement in detail clarity and natural color restoration.
The two-stage underwater image enhancement method based on deep learning is adopted. In the first stage, the image quality is improved through the encoder decoder architecture with a parallel bidirectional architecture of CNN and Transformer. In the second stage, the details are enhanced through the encoder decoder structure composed of CNN, and the gradual optimization of the image is achieved by combining flexible network structure and reasonable loss functions.
Higher quality underwater images are generated, the object details are processed more finely and the color is more appropriate, which is suitable for practical engineering applications.
Smart Images

Figure CN120387932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a two-stage underwater image enhancement method based on deep learning. Background Art
[0002] The ocean is closely connected to human society and holds significant scientific and environmental value for the development and utilization of marine resources. Underwater imagery carries a wealth of marine information and plays a key role in resource exploration, ecological monitoring, and engineering applications. In recent years, underwater imaging technology has become increasingly widely used in marine research and engineering, becoming a crucial tool for advancing related fields.
[0003] However, due to the unique characteristics of the underwater environment, light propagation in water is affected by factors such as refraction, absorption, and scattering. This often results in images with low contrast, color shift, and blurred details. These issues pose significant challenges in fields such as underwater archaeology, target identification, and deep-sea exploration, and severely limit the development of underwater vision technology. Therefore, improving the quality of underwater images is crucial for advancing marine science research and applications.
[0004] To address this issue, researchers have proposed a variety of underwater image enhancement (UIE) methods. Traditional enhancement methods typically directly adjust pixel brightness, contrast, and saturation, but often struggle to effectively restore the image's true color and detail. Physical model-based methods attempt to correct images by accurately calculating medium transmission characteristics and underwater imaging parameters. However, due to the complexity of the underwater environment, the assumptions of such methods are not always applicable, and errors in parameter estimation can also affect the final enhancement effect.
[0005] In recent years, the development of deep learning technology has provided new insights into underwater image enhancement. Deep learning-based methods typically rely on large amounts of training data to learn the mapping relationships for image enhancement and have achieved significant progress in generalization and enhancement effectiveness. These methods typically utilize supervised learning, training neural networks on labeled datasets to automatically restore degraded underwater imagery without relying on complex physical imaging models. However, despite the success of deep learning methods, many current algorithms still have room for improvement in terms of detail clarity and natural color reproduction. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention proposes a two-stage underwater image enhancement method based on deep learning to solve problems such as unclear underwater images and color deviation. The distorted images are gradually enhanced through a two-stage network structure, and a step-by-step optimization strategy is adopted to achieve more refined and comprehensive image enhancement. This method mainly initially improves the quality of underwater images through the encoder-decoder architecture in the first stage, where the encoder consists of a parallel bidirectional architecture of CNN and Transformer. In the second stage, an encoder-decoder structure composed of CNN is used for detail enhancement. Then, by training the neural network, optimal parameters are obtained to achieve the enhancement of underwater images.
[0007] To solve the above technical problems, the technical solution of the present invention is as follows:
[0008] A two-stage underwater image enhancement method based on deep learning, comprising the following steps:
[0009] S1. Obtain an underwater image dataset.
[0010] S2. Based on CNN and Transformer, construct a two-stage hybrid underwater image enhancement model, process the underwater image dataset in two stages, and output an underwater image enhancement result map.
[0011] In the first stage, an encoder-decoder architecture is adopted, and the feature combination of CNN and Transformer is realized through bidirectional bridging to initially improve the quality of underwater images.
[0012] In the second stage, the features of the underwater image are refined through an encoder-decoder structure to complete the enhancement of the underwater image.
[0013] S3. Use the underwater image dataset to search for and test the two-stage hybrid underwater image enhancement model.
[0014] Preferably, in the first stage, an encoder-decoder architecture is adopted, where the convolutional structure of the original U-Net is replaced by Swin Transformer. In the first-stage network, each layer of each encoder and decoder part is respectively composed of two Swin Transformer blocks, and these two blocks are in a serial relationship. There is a bottleneck layer composed of two Swin Transformer blocks between the encoder and the decoder, and there are skip connections between the corresponding layers of the encoder and the decoder; each Swin Transformer block contains window multi-head self-attention W-MSA and a feed-forward network FFN, and is equipped with layer normalization and residual connections; in each encoder layer, a CNN and Transformer combination module is connected in parallel.
[0015] In each Swin Transformer block, the input features are first normalized, then feature extraction is performed through W-MSA, and they are fused with the original input through a residual connection. Next, the fused features are normalized again and processed by FFN, and finally, the calculation of the Swin Transformer block is completed through a residual connection.
[0016] Preferably, the CNN and Transformer combination module: First, on the one hand, the input I of the previous encoder is a two-dimensional Transformer feature; this two-dimensional feature is reshaped into a three-dimensional spatial structure and the data dimensions are adjusted to generate new features; first, the new features are subjected to a convolution operation through the CNN and Transformer combination module to generate a query feature Q. At the same time, on the other hand, for the Transformer feature Z after being transformed by Swin Transformer, it is converted into a key feature K and a value feature V through two different convolutional layers respectively. Then, the dot product of the query feature Q and the key feature K is calculated, and softmax normalization is used to generate attention weights. Finally, through weighted value feature V, the attention output is obtained. The output of the attention is used as the weight of the output feature map of the corresponding layer of the encoder on the one hand and input to the next encoder layer, and the output of the attention is used as the input feature of the next CNN and Transformer combination module on the other hand, and this is executed sequentially until the entire encoder part is completed.
[0017] Preferably, in the second stage, the U-Net or encoder-decoder architecture is also adopted. First, the original image I is concatenated with the output O1 of the first stage to generate an input feature I2, which is then passed into the network of the second stage to generate the final underwater image enhancement result map.
[0018] Specifically, the encoder part is composed of four convolutional blocks. Each convolutional block includes a 3×3 convolutional layer, an InstanceNorm normalization layer, and a ReLU activation function. In addition, three convolutional operation residual blocks are added to enhance the feature expression ability. In the decoder part, a structure symmetric to the encoder is adopted to achieve high-quality image reconstruction. At the same time, the encoder and the decoder are connected through skip connections to retain key information and optimize feature transfer. Finally, through this two-stage processing method, the network can generate a higher-quality underwater image O2.
[0019] The present invention has the following characteristics and beneficial effects:
[0020] With the above technical solution, the present invention proposes an end-to-end two-stage hybrid network for improving the quality of underwater images. The first stage is used to initially enhance the quality of underwater images, and the second stage is used to refine the details of underwater images. At the same time, a CNN-Transformer combined module is designed, which connects the CNN and the transformer through a bidirectional bridge. In particular, the bridging mechanism helps the bidirectional fusion of local CNN features and global transformer features. The flexible network structure combined with a reasonable loss function setting can obtain a better-quality clean underwater image. Compared with other underwater image enhancement methods, the present invention processes the details of objects more precisely and appropriately processes colors, resulting in better effects. At the same time, the model constructed by the present invention is easy for engineering practitioners to understand, so as to carry out engineering deployment faster and better. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a network method diagram of an embodiment of the present invention;
[0022] Figure 2 It is the structure of the first stage;
[0023] Figure 3 It is the structure of the CNN-Transformer combined module;
[0024] Figure 4 It is the structure of the second stage;
[0025] Figure 5 It is the test effect diagram of the present invention and the comparison of the effects with the original image and the GT image. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention provides a two-stage underwater image enhancement method based on deep learning, as Figure 1 shown, including the following steps:
[0027] S1. Obtain an underwater image dataset.
[0028] S2. Construct a two-stage hybrid underwater image enhancement method combining CNN and Transformer. The model includes a first stage and a second stage.
[0029] S2.1. Through the first stage, use the encoder-decoder architecture to initially enhance the quality of underwater images.
[0030] As Figure 2 shown, the first stage adopts the encoder-decoder architecture, in which the convolutional structure of the original U-Net is replaced by Swin Transformer. The proposed first-stage network (as Figure 2As shown in the figure, two Swin Transformer modules are integrated in the encoder part. Each module contains window multi-head self-attention (W-MSA) and a feed-forward network (FFN), and is equipped with layer normalization (LN) and residual connections.
[0031] In each Swin Transformer block (STB), the input features are first normalized, and then feature extraction is performed through W-MSA and fused with the original input through residual connections. Then, the features are normalized again and processed through FFN, and finally the entire STB calculation is completed through residual connections. In addition, in each encoder layer, a combined CNN and Transformer module designed by S2.2 is connected in parallel to enhance the feature extraction ability.
[0032] It can be understood that in step S3, an encoder-decoder composed of Swin Transformer is used, and the proposed encoder and decoder each contain 4 stages. In the encoder, a 2*2 transposed convolutional layer with a stride of 2 is used for upsampling, and the number of feature channels is halved in each layer.
[0033] S2.2. Design a combined CNN and Transformer module, as Figure 3 shown, to achieve feature combination of CNN and Transformer through bidirectional bridging.
[0034] First, the input I of the previous encoder is two-dimensional Transformer features; this two-dimensional feature is reshaped into a three-dimensional spatial structure and the data dimensions are adjusted to generate new features; first, the new features are convolved through the combined CNN and Transformer module to generate query features Q; for the Transformer features Z after being transformed by Swin Transformer, they are respectively converted into key features K and value features V through two different convolutional layers; then, the dot product of the query feature Q and the key feature K is calculated, and softmax normalization is used to generate attention weights; finally, the weighted value feature V is used to obtain the attention output. On the one hand, the output of the attention is used as the weight of the output feature map of the corresponding layer of the encoder and input to the next encoder layer. On the other hand, the output of the attention is used as the input feature of the next combined CNN and Transformer module, and this is executed sequentially until the entire encoder part is completed.
[0035] S2.3. As Figure 4 shown, further refinement of underwater image features is performed through the encoder-decoder structure in the second stage;
[0036] This network introduces a second stage, which also adopts a U-Net or encoder-decoder architecture, as Figure 2 and Figure 5As shown below. First, the original image I is concatenated with the output O1 of the first stage (i.e., S2.1) to generate the input feature I2, which is then passed into the network of the second stage.
[0037] Specifically, the encoder part consists of four convolutional blocks. Each convolutional block includes a 3×3 convolutional layer, an InstanceNorm normalization layer, and a ReLU activation function. In addition, three residual blocks are added in the bottleneck layer to enhance the feature expression ability. In the decoder part, a structure symmetric to the encoder is adopted to achieve high-quality image reconstruction. Meanwhile, the encoder and the decoder are connected by skip connections to retain key information and optimize feature transmission. Finally, through this two-stage processing method, the network can generate a higher-quality underwater image O2.
[0038] S3. Train the constructed underwater image enhancement network model based on cascade adaptation.
[0039] Specifically, the specific method of step S3: During the training process, the results of the overall network are supervised, and the GT map is also cropped to have the same size as the input data. By comparing the differences between the predicted map and the GT map, it is judged whether the sum of the loss values converges to determine the training process of the network.
[0040] The size of the input data is uniformly adjusted to 256×256×3, the batch size is set to 16, and the Adam optimizer is used to update the model parameters during the training process. The initial learning rate is set to 1e-4. The final loss value is the sum of two loss parts, namely calculating loss1 from the output of the first stage and the GT, and calculating loss2 from the output of the second stage and the GT.
[0041] In the above technical solution, the calculation of each part of the loss function is composed of the sum of the L1 loss, the SSIM loss, and the perceptual loss. The loss function calculates the L1 distance between the predicted image and the GT. The SSIM loss is used in the loss function to impose structural and texture similarity on the predicted image. The definition of pixel SSIM for pixel x is:
[0042]
[0043] Among them, the calculation of SSIM(x) requires the neighborhood of pixel x, that is, the patch. The mean and standard deviation of the patch of the predicted image are μ I , σ I , and are the mean and standard deviation of the GT. is the cross-covariance. Here, a 11x11 patch around pixel x is taken for calculation, and C1 = 0.01 and C2 = 0.03 are set.
[0044] The loss function of SSIM can be written as setting E(x) = 1 - SSIM(x):
[0045]
[0046] The perceptual loss is expressed as the distance between the feature representations of the predicted image and the ground truth image to measure the high-level perceptual and semantic differences between images:
[0047]
[0048] where φ j is the j-th convolutional layer of the pre-trained VGG19; N is the number of each batch during the training process; C j H j W j represents the dimension of the j-th layer of the VGG-19 network feature map, where C j 、H j 、W j represent the number, height, and width of the feature maps respectively.
[0049] During this training process, it is judged whether the training process of the network converges by observing whether the sum of the three loss values converges. If the value converges, the training of this network is completed.
[0050] Adopting the above technical solution, the present invention proposes an end-to-end two-stage hybrid network for improving the quality of underwater images. The first stage is used to initially improve the quality of underwater images, and the second stage is used to refine the details of underwater images. At the same time, a CNN and Transformer combined module is designed, and this module connects the CNN and the transformer through a bidirectional bridge. In particular, the bridging mechanism helps the bidirectional fusion of local CNN features and global transformer features. The flexible network structure combined with the reasonable setting of the loss function can obtain better clean underwater images. The experimental results are as Figure 5 shown. Compared with other underwater image enhancement methods, the present invention processes the details of objects more finely, appropriately processes colors, and can produce better effects. At the same time, the model constructed by the present invention is easy for engineering practical application personnel to understand, so as to carry out engineering deployment faster and better.
[0051] In addition, due to its ability to effectively remove color deviation and improve the clarity of pictures, it has wide application value in applications such as underwater archaeology, underwater target detection, and seabed exploration.
[0052] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations can be made to these embodiments including components, and still fall within the protection scope of the present invention.
Claims
1. A two-stage underwater image enhancement method based on deep learning, characterized in that It includes the following steps: S1. Obtain an underwater image dataset; S2. Based on CNN and Transformer, construct a two-stage hybrid underwater image enhancement model, process the underwater image dataset in two stages, and output an underwater image enhancement result map; S3. Use the underwater image dataset to train and test the two-stage hybrid underwater image enhancement model.
2. The two-stage underwater image enhancement method based on deep learning according to claim 1, wherein The two-stage hybrid underwater image enhancement model includes two stages. The first stage adopts an encoder-decoder architecture, realizes the feature combination of CNN and Transformer through bidirectional bridging, and initially improves the quality of underwater images; the second stage refines the underwater image features through an encoder-decoder structure to complete underwater image enhancement.
3. The two-stage underwater image enhancement method based on deep learning according to claim 2, wherein The specific implementation process of the two-stage hybrid underwater image enhancement model is as follows: S2.
1. The first stage adopts an encoder-decoder architecture, in which the convolutional structure of the original U-Net is replaced by Swin Transformer. Each layer of the encoder and decoder parts is serially composed of two Swin Transformer blocks. There is a bottleneck layer composed of two Swin Transformer blocks between the encoder and the decoder, and there are skip connections between the corresponding layers of the encoder and the decoder; each Swin Transformer block contains a window multi-head self-attention W-MSA and a feed-forward network FFN connected in sequence, and is equipped with layer normalization and residual connections; in each encoder layer, a CNN and Transformer combination module is connected in parallel; S2.
2. The input I of the previous encoder is a two-dimensional Transformer feature; this two-dimensional feature is reshaped into a three-dimensional spatial structure and the data dimension is adjusted to generate a new feature; first, the new feature is convolved through the CNN and Transformer combination module to generate a query feature Q; for the Transformer feature Z after being transformed by Swin Transformer, it is converted into a key feature K and a value feature V through two different convolutional layers respectively; Then, calculate the dot product of the query feature Q and the key feature K, and use softmax normalization to generate attention weights; finally, through weighted value feature V, an attention output is obtained. On the one hand, the output of the attention is used as the weight of the output feature map of the corresponding layer of the encoder and input to the next encoder layer. On the other hand, the output of the attention is used as the input feature of the next CNN and Transformer combination module, and this is executed sequentially until the entire encoder part is completed; S2.
3. In the second stage, the U-Net architecture is also adopted; First, the original image I is concatenated with the output O1 of the first stage to generate an input feature I2, and it is passed into the network of the second stage to generate the final underwater image enhancement result map.
4. The dual-stage underwater image enhancement method based on deep learning according to claim 3, wherein The specific implementation process of the Swin Transformer block is as follows: The input features are first normalized, and then feature extraction is performed through W-MSA and fused with the original input through a residual connection; Next, the fused features are normalized again, processed through FFN, and finally the calculation of the Swin Transformer block is completed through a residual connection.
5. The dual-stage underwater image enhancement method based on deep learning according to claim 4, wherein, In the second stage, the encoder part consists of four convolutional blocks, each convolutional block includes a convolutional layer, a normalization layer, and a ReLU activation function; In addition, three residual blocks including convolutional operations are added; In the decoder part, a structure symmetric to the encoder is adopted. At the same time, the encoder and the decoder are connected through skip connections; Finally, the underwater image O2 is generated.
Citation Information
Cited By
Deep-sea ecology monitoring method and system for deep optimization
CN121366346A