Underwater image restoration method based on multi-modal visual guidance and characteristic decomposition

By using multimodal visual guidance and feature decomposition methods in underwater image recovery, multimodal prompt information and high-frequency features are used to restore the color and details of underwater images, the image quality problems caused by light absorption and scattering are solved, and more efficient underwater image recovery is achieved.

CN120070257APending Publication Date: 2025-05-30CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411923530.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Underwater image recovery technology faces the problems of color distortion and reduced contrast caused by light absorption, as well as image blur and detail loss caused by light scattering.

Method used

The underwater image restoration method based on multimodal visual guidance and feature decomposition is adopted to form an image restoration model through an encoder, U-Transformer model and a decoder. The complementary characteristics between the multimodals are used to generate prompt information, guide the recovery of the color channel, and extract high-frequency information through learnable wavelets, and restore subtle details using the high-frequency hybrid attention module.

Benefits of technology

This method exhibits superior performance in color restoration and texture detail restoration of underwater images, improving the overall quality and robustness of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070257A_ABST
    Figure CN120070257A_ABST
Patent Text Reader

Abstract

The invention relates to an underwater image restoration method based on multi-modal visual guidance and characteristic decomposition, which is used for solving the problems in the prior art that color distortion and contrast reduction of an image are caused by light absorption due to gradual attenuation of light with different wavelengths in water, and the light is scattered when encountering particles in the water, so that the image cannot be restored. And the problems of image blurring and detail loss are solved. According to the scheme, multi-mode prompt learning is adopted, complementarity among multiple modes is utilized, and features from different modes are integrated into prompt information so as to support color and contrast recovery of underwater images; and extracting high-frequency information from the RGB modal features by using learnable wavelets, and focusing on fine and difficult-to-perceive detail areas in the high-frequency features by using a high-frequency mixed attention module so as to recover the definition of textures in the underwater image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing, and in particular to an underwater image restoration method based on multimodal visual guidance and feature decomposition. Background Art

[0002] Underwater image restoration technology has been widely applied in fields such as ocean exploration, biology, archaeology, and underwater robots. However, there are currently two major challenges in underwater image restoration. On the one hand, due to the light absorption caused by the gradual attenuation of light with different wavelengths in water, the image shows color distortion and reduced contrast. On the other hand, light scattering occurs when light encounters particles in water, resulting in image blurring and detail loss. Therefore, overcoming these two challenges is crucial for achieving efficient underwater image restoration.

[0003] Existing underwater image restoration methods can be roughly divided into vision-prior-based methods and data-driven methods. Vision-prior-based methods mainly restore the quality of underwater images by adjusting pixel attributes (such as contrast and saturation). These methods rely on explicit degradation models or prior knowledge to calibrate model parameters. Commonly used techniques include underwater dark channel prior, attenuation curve prior, transmission prior, and lumen flux prior to solve the unique color distortion and detail blurring problems in underwater images, thereby restoring the visual quality of the images. Although vision-prior-based underwater image restoration methods have achieved remarkable results, the variability and unpredictability of the underwater environment may affect the robustness of these priors, further limiting the performance of the model in color and detail restoration. Summary of the Invention

[0004] The purpose of this case is to propose an underwater image restoration method based on multimodal visual guidance and feature decomposition to solve the problems of color distortion and reduced contrast caused by light absorption due to the gradual attenuation of light with different wavelengths in water and the problems of image blurring and detail loss caused by light scattering when light encounters particles in water. The specific technical solutions are as follows.

[0005] In the first aspect, this case proposes an underwater image restoration method based on multimodal visual guidance and feature decomposition, which is characterized in that an image restoration model is composed of an encoder, a U-Transformer model, and a decoder, and the trained image restoration model is used to restore the underwater image. Let i ∈ [1, 4] be the output feature of the encoder. The obtaining steps of include: based on the original underwater RGB image successively obtain images of different scales Based on the original underwater depth image successively obtain images of different scales Based on the RGB image Modal features Depth image Modal features Fuse to generate prompt information i ∈ [1, 3]; Based on the RGB image Modal features Decompose the high-frequency components through learnable wavelet, and then generate high-frequency features based on these high-frequency components; The modal features of the RGB image Prompt information And high-frequency features are additively bound to obtain features The modal features of the RGB image And prompt information Are additively bound to obtain features i ∈ [2, 3]; The modal features of the RGB image Are used as features

[0006] In an implementation of the above technical solution, based on the RGB image Modal features Depth image Modal features Fuse to generate prompt information i ∈ [1, 3], the steps include: Concatenate the modal features And perform preliminary fusion enhancement; Split the features of the preliminary fusion enhancement into RGB features And depth features And project To reduce the dimension, and obtain the corresponding projected features as Obtain Enhanced modal information Finally And Are additively bound to obtain hybrid modal prompt information i ∈ [1, 3].

[0007] In an implementation of the above technical solution, a modal vision guidance module is adopted, based on the RGB image Modal features Depth image Modal features Fuse to generate prompt information i ∈ [1, 3]; The modal vision guidance module includes a fusion enhancement block, a projection block, and a spatial annotation block; The fusion enhancement block consists of multiple convolutional blocks, and each convolutional block consists of a 3×3 convolutional layer and a ReLU activation function; The projection block consists of a 1×1 convolutional layer; The spatial annotation block for Perform a spatial fixation operation to obtain a channel-level spatial attention mask

[0008] In one implementation of the above technical solution, multiply the attention mask with element-wise to generate enhanced modality information

[0009] In one implementation of the above technical solution, based on the RGB image modal features Decompose the high-frequency components through learnable wavelet transform, and then generate high-frequency features based on these high-frequency components. The steps include: Decompose the modal features through learnable wavelet transform to obtain high-frequency components lh, hl, hh, where lh represents the high-frequency details in the horizontal direction of the feature, hl represents the high-frequency details in the vertical direction of the feature, and hh represents the high-frequency details in the diagonal direction of the feature; Combine the high-frequency components in an element-wise addition manner to form initial high-frequency information Based on the initial high-frequency information obtain high-frequency features

[0010] In one implementation of the above technical solution, adopt a feature decomposition module to decompose the high-frequency components through learnable wavelet transform based on the modal features of the RGB image modal features and then generate high-frequency features based on these high-frequency components; The feature decomposition module includes a high-frequency filter and a high-frequency hybrid attention block; The high-frequency filter is configured to perform wavelet decomposition based on the input feature The processing of the high-frequency hybrid attention block can be expressed as: where CA(·) represents channel attention, SA(·) represents spatial attention, ResB represents a residual block operation with batch normalization, and the residual block consists of a convolutional layer, batch normalization, and a ReLU activation function, and there are multiple convolutional layers.

[0011] In a second aspect, this case proposes an underwater image restoration system based on multi-modal visual guidance and feature decomposition. The system uses a trained image restoration model to restore underwater images. The image restoration model consists of an encoder, a U-Transformer model, and a decoder; Among them: The encoder is configured to sequentially obtain images of different scales based on the original underwater RGB image sequentially obtain images of different scales Based on the original underwater depth image sequentially obtain images of different scales Based on the RGB image modal features depth image Modal features Fusion-generated prompt information i ∈ [1, 3]; Based on the RGB image Modal features Extract high-frequency components through learnable wavelet decomposition, and then generate high-frequency features based on these high-frequency components; Combine the modal features of the RGB image Prompt information And high-frequency features are additively bound to obtain features Combine the modal features of the RGB image And prompt information Are additively bound to obtain features i ∈ [2, 3]; Use the modal features of the RGB image As features Use i ∈ [1, 4] as output features

[0012] The beneficial technical effects of this case are as follows: By utilizing the complementary characteristics between multi-modalities to generate prompt information and guiding the restoration of color channels, the overall color restoration effect of underwater images can be enhanced. While using multi-modal prompts, learnable wavelets are used to extract high-frequency information from the RGB image at the largest scale, and the high-frequency hybrid attention module is used to focus on subtle and imperceptible detailed features, which helps to restore the clarity of textures in underwater images Brief description of the drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings

[0014] Figure 1 Comparison schematic diagram between the existing image restoration method and the implementation method of this case

[0015] Figure 2 、 One Overall framework schematic diagram in a certain implementation manner

[0016] Figure 3 、 One Overall structure schematic diagram of the U-Transformer model in a certain implementation manner

[0017] Figure 4 Overall structure schematic diagram of a certain MVG module

[0018] Figure 5 Overall structure schematic diagram of a certain FD module

[0019] Figure 6 Schematic diagram of the visual effect comparison of the method in this case with four vision-prior-based methods (ULAP, IBLA, WaterNet, and Ucolor) and five data-driven methods (FUnIE, UWCNN, U-Shape, UDA, and SMDR-IS) for underwater image restoration.

[0020] Figure 7 Schematic diagram of the visual effect comparison of UWCNN, U-Shape, SMDI-IS, and the method in this case on the full-reference dataset.

[0021] Figure 8 Schematic diagram of the visual effect comparison presented by the BL module, BL+FD module, BL+MVG module, and the complete model. Detailed implementation

[0022] For the explanation of the parameter meanings, see Table 1.

[0023] Table 1

[0024]

[0025]

[0026]

[0027]

[0028] With the progress of deep learning models and the continuous expansion of underwater image datasets, data-driven underwater image restoration methods have gradually become the mainstream, and these methods have shown superior performance in detail sharpening and color correction. Figure 1 In (a) of [reference], it schematically shows the processing flow of the mainstream underwater image restoration methods. Such methods usually focus on the features extracted from the RGB modality for training the basic model of underwater image restoration. However, these methods ignore the complementary and consistent information provided by other modalities, and this information can improve the performance of the basic model. Figure 1 In (b) of [reference], it schematically shows the processing flow of some existing multimodal underwater image processing methods. These existing methods have introduced additional modality information in the underwater image restoration task. For example, Li et al. adjusted the model's attention to severely degraded regions in the image by introducing depth information. However, most existing methods mainly use the additional modality information for model fine-tuning, rather than fully utilizing the complementary and consistent information between different modalities to train the basic model. Figure 1Among them, (c) schematically shows the processing flow of the underwater image restoration method based on multi-modal visual guidance and feature decomposition (MVG-FD) proposed in this case. The MVG-FD model adopts a different multi-modal information fusion method. In this method, by adopting multi-modal prompt learning and utilizing the complementarity between modalities, the features from different modalities are integrated into prompt information and incorporated into the training of the base model. In addition, MVG-FD also includes an additional high-frequency feature decomposition module.

[0029] In the method of this case, considering that the absorption rate of light waves in water changes with distance, a modal visual guidance (MVG) module is designed in this case. This module integrates the consistency information from the RGB modality and the depth modality to support the color and contrast restoration of underwater images. In addition, light scattering will cause texture blur and detail loss in underwater scenes. To restore these imperceptible textures and details, a feature decomposition (FD) module is designed in this case. The FD module not only extracts high-frequency information from RGB modality features using learnable wavelets, but also introduces a high-frequency hybrid attention (HFMA) module, which focuses on the part of high-frequency features that is most helpful for restoring the texture details of underwater images. This case also applies a multi-channel color loss function to guide the learning of the model, which combines the losses in the RGB, LAB, and LCH color spaces to enhance the color restoration ability of the model. Experimental results show that MVG-FD has strong robustness and exhibits superior performance in terms of image quality and visual effects compared with existing underwater image restoration methods.

[0030] Next, in combination with the accompanying drawings, how the technical solution of this case is implemented will be clearly and completely described. Obviously, the described implementation manners are only part of the implementation manners of this case, rather than all the implementation manners. Based on the implementation manners in this case, all other implementation manners obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0031] (1) Framework

[0032] See Figure 2 A method implementation framework shown schematically. MVG-FD includes a multi-scale encoder, a Transformer module, and a decoder.

[0033] (1.1) Encoder

[0034] In the encoder, the original RGB image and depth image are directly input into the network and downsampled multiple times respectively. Subsequently, they pass through the Channel Augmented Convolution Block (CACB) to expand the number of channels. The CACB includes multiple 3×3 convolutional layers, normalization operations, and LeakyReLU activation functions. Secondly, feature maps of different scales are input into the corresponding MVG modules to generate multi-modal visual cues at three scales and added to the corresponding RGB feature maps in the form of residuals.

[0035] Figure 2 in, i ∈ [1, 4], i ∈ [1, 3] represent the RGB and depth images at different scales respectively. The Modal Visual Guidance (MVG) module fuses multi-modal features to generate cue information, i ∈ [1, 3] represents the multi-modal visual cues at each scale. Specifically, based on the RGB image modal features of depth image modal features of are fused to generate cue information i ∈ [1, 3].

[0036] The RGB features at the original scale are then input into the Feature Decomposition module (FD), and the Feature Decomposition (FD) module is used to extract high-frequency information. The Feature Decomposition (FD) module uses Learnable Wavelets (LWD) to decompose the features and extract high-frequency components, and then generates high-frequency features based on these high-frequency components. The High-Frequency Mixing Attention module (HFMA) focuses on important regions and further extracts subtle and indistinguishable textures from the high-frequency features, and these textures are also added to the RGB feature maps in the form of residuals to obtain the final features at different scales i ∈ [1, 4]. Specifically, based on the RGB image modal features of high-frequency components are decomposed through learnable wavelets, and then high-frequency features are generated based on these high-frequency components; the modal features of the RGB image cue information and high-frequency features are additively bound to obtain the feature the modal features of the RGB image and cue information are additively bound to obtain the feature i ∈ [2, 3]; the modal features of the RGB image are used as the feature Finally, the feature maps at four scales are input into the Transformer module.

[0037] (1.2) Transformer module

[0038] See Figure 3 The schematic Transformer module. In this case, the U-Transformer model is preferably used as the Transformer module. The U-Transformer model consists of a multi-scale feature channel cross-fusion Transformer (MSFCCT) module and a global feature modeling Transformer (GFMT) module. The U-Transformer model inputs features of five different scales

[0039] The MSFCCT module takes the features of the first four scales as input. Then, a pooling operation with a kernel size of 2×2 and a stride of 2 is performed, and the resulting features are used as the input to the GFMT module.

[0040] The GFMT is placed in the bottleneck layer to model global information and enhance the model's attention to severely degraded regions in underwater images. The input feature F 5 has a size of Before inputting into the GFMT model, a linear projection is first performed on the feature map to flatten it into a two-dimensional feature sequence Then, learnable position encoding is added to retain the important position information of each region. Therefore, S in can be expressed as:

[0041] S in = W*F 5 + PE, (1)

[0042] W*F 5 represents the linear projection operation, and PE represents the position encoding operation. Next, the feature sequence S in is input into the SFMT module, which consists of 4 standard Transformer layers. Each Transformer layer contains a multi-head self-attention block (MHA) and a feed-forward network (FFN). The FFN contains a normalization layer and a fully connected layer. The output of the l-th layer (l ∈ 1, 2,..., 4) is expressed as:

[0043] S l = MHA(LN(S l-1 )) + S l-1 , (2)

[0044] S l = FFN(LN(S' l )) + S' l , (3)

[0045] The final output sequence is After performing feature back-projection, it is remapped back to the feature map

[0046] MSFCCT establishes skip connections between the encoder and the generator. The input of MSFCCT is four feature maps of different scales: Different from vision that directly processes images, MSFCCT uses convolution filters with a kernel size of and a stride of for linear projection on feature maps of different scales (i = 0, 1, 2). Here, P is taken as 32. Then, four feature sequences of different scales are obtained where After the above convolution operation, the channel dimension remains unchanged, and the feature map is divided into equally sized blocks. Then, four query vectors can be obtained using Equation 4 and and

[0047]

[0048] where, represents the learnable weight matrix; S is generated by concatenating on the channel dimension, where C = C 1 + C 2 + C 3 +... In this work, C 1 , C 2 , C 3 , C 4 are set to 32, 64, 128, 256 respectively.

[0049] The Channel Multi-Head Attention (CMHA) block has six inputs, which are respectively and the output of channel attention can be obtained in the following way:

[0050]

[0051] where, IN represents instance normalization, which is different from batch normalization and mainly maintains global consistency by removing the specific style information of a single image. The output of the i-th CMHA layer can be expressed as:

[0052] CMHA i =(CA i 1 + CA i 2 +………, + CAi N ) / N + Q i ,

[0053] Among them, N is the number of heads, which is set to 4 in the experiment.

[0054] The feed - forward network (FFN) performs a non - linear transformation on the input features, and its mathematical expression is:

[0055] O i = CMHA i + MLP(LN(CMHA i )),(6)

[0056] In this expression, where MLP represents a multi - layer perceptron.

[0057] Finally, four different output feature sequences are subjected to feature remapping and re - organized into four feature maps

[0058] (1.3) Decoder

[0059] In the decoder, a channel - enhanced convolutional block is composed of a channel - expansion convolutional block (CACB) and a feature - transformed RGB image (FTR). Four channel - enhanced convolutional blocks with different scales first receive four outputs ε i from the Transformer module, i ∈ [1, 4]. These channel - enhanced convolutional blocks are consistent with the design in the encoder and are used here to reduce the number of channels. Then, the low - scale features are upsampled in sequence and concatenated with the high - scale features. Finally, a 1×1 convolution is used to generate the output image of the corresponding scale. , i ∈ [1, 4] represents the output results of each scale of the decoder. Specifically, for two adjacent channels, the lower channel is of low scale and the upper channel is of high scale. The features output by the lower channel - expansion convolutional block are upsampled, the upsampled features are input into the upper channel - expansion convolutional block for fusion, and the upsampled features are channel - connected to the output of the upper channel - expansion convolutional block. The connected features are processed by the feature - transformed RGB image. The features output by the bottom - most channel - expansion convolutional block are directly processed by the feature - transformed RGB image.

[0060] (1.4) Multi - modal visual guidance module (MVG)

[0061] The above MVG module generates visual cues by taking advantage of the complementary information between the RGB modality and the depth modality to assist in the color restoration of underwater images. This is the first time to introduce the concept of multi - modal cue learning into the underwater image restoration task.

[0062] InFigure 2 In it, the proposed MVG module is inserted into multiple scales of the encoder. Specifically, in each of the first three scales, the MVG module receives the output modal features after the channel enhancement convolutional block for the images corresponding to the RGB modality and the depth modality, respectively denoted as where \(i\) represents the scale layer.

[0063] For the MVG module, assume that the two input modal features are respectively denoted as and The MVG module generates multi-modal visual cue information based on these input features, denoted as

[0064] In this way, the MVG module can effectively learn the complementarity between the two modalities and generate cue information for underwater image color restoration.

[0065] The detailed design of the MVG module is as Figure 4 shown, including two input modal features and Specifically, the features of the RGB modality and the depth modality are first concatenated, and then preliminarily fused and enhanced through a fusion addition block to achieve interaction between modalities. An example structure of the fusion addition block consists of multiple convolutional blocks, and each convolutional block consists of a 3×3 convolutional layer and a ReLU activation function.

[0066] Exemplarily, 3 convolutional blocks are used for preliminary fusion enhancement, and the preliminary fusion enhancement is represented by the enhancement function \(T(\cdot)\), i.e.:

[0067]

[0068] To handle the redundancy of information between modalities, we split the features of the preliminary fusion enhancement into RGB features and depth features and project for dimensionality reduction, specifically projecting to dimension to obtain the corresponding projected features as The formula is as follows:

[0069]

[0070] In this case, the input channel \(C\) varies with different layers of the MVG-FD model, and the reduction factor \(\eta\) in the MVG is designed as \(C / 8\). And the first projection function \(g\) 1 (·) and the second projection function \(g\) 2 (·) are simple 1×1 convolutional layers.

[0071] Then, the spatial fixation block (SFB) operates on Perform a spatial gaze operation, which first applies spatial softmax with λ smoothing in all spatial dimensions and then generates enhanced modal information by applying the channel-level spatial attention mask to to generate enhanced modal information The formula is as follows:

[0072]

[0073] where i = 1, 2,..., H, j = 1, 2,..., W, and λ is the learnable weight parameter for each block. Finally, the hybrid modal cue information is obtained through additive binding.

[0074] (1.5) Feature Decomposition Module

[0075] Underwater images are usually affected by light scattering, resulting in blurred textures and lost details. Current feature extractors are difficult to effectively capture fine texture features. To address this issue, a Feature Decomposition (FD) module is specifically designed. Since high-frequency components usually exist in the underlying layer of image features, the FD module is only inserted into the first layer of the model.

[0076] The overall structure of the FD module is as Figure 5 shown. The input feature is decomposed into high-frequency components (lh, hl, hh) by learnable wavelet decomposition (LWD). These components are combined by element-wise addition to form the initial high-frequency information Next, is input into the High-Frequency Mixing Attention (HFMA) module for further processing to generate high-frequency attention features

[0077]

[0078] Here represents 's high-frequency information. W HF represents the learnable high-frequency filter, whose coefficients are updated after the AWD operation and initialized as Haar wavelets. Compared with manually designed wavelets, this learnable wavelet filter can better adapt to underwater data, thus further promoting the extraction of detail features. Therefore, a High-Frequency Feature Mixing Attention (HFMA) module is designed to capture the features most beneficial for detail recovery in the decomposed high-frequency information.

[0079] (1.5.1) High-Frequency Feature Mixing Attention Module

[0080] In the High - Frequency Mixed Attention (HFMA) module, residual blocks are first applied to preserve the texture. These residual blocks (ResB) consist of multiple 3×3 convolutional layers, batch normalization (BN), and ReLU activation functions. Subsequently, spatial attention (SA) and channel attention (CA) mechanisms are adopted to highlight the significant parts in the spatial domain and channel domain respectively.

[0081] As Figure 5 shown, in the CA module, after the input features are processed by average pooling, a convolutional layer, and a ReLU activation function, channel weights are generated through the Sigmoid function. These weights are multiplied element - by - element with the original features to emphasize the key channels. In the SA module, the input features are first processed by a convolutional layer and a ReLU activation function, and then passed through a second convolutional layer and the Sigmoid function to generate spatial weights. These weights are multiplied element - by - element with the input features to enhance the important spatial regions. Thus, the high - frequency attention features can be expressed as:

[0082]

[0083] where ResB(·) represents the residual block with batch normalization (BN), CA(·) represents the channel attention, and SA(·) represents the spatial attention.

[0084] (1.5.2) Loss Function

[0085] During model training, a multi - color - space loss function is adopted, combining the RGB, LAB, and LCH color spaces to train our network. First, the images are converted from the RGB space to the LAB and LCH spaces.

[0086]

[0087] where x, y, and G(x) represent the original RGB image, the reference image, and the sharp image output by the generator respectively.

[0088] The loss functions in the LAB and LCH spaces are expressed as follows:

[0089]

[0090] where Q represents the quantization operator. By quantizing specific channels in each color space, the cross - entropy loss between the enhanced image and the reference image on these channels can be calculated. The subscripts i, j in equations (9) and (10) represent the pixel coordinates.

[0091] The final overall loss G * is expressed as:

[0092] G *= αLoss LAB (G(x), y) + βLoss LCH (G(x), y) + γLoss RGB (G(x), y) + μLoss per (G(x), y). (17)

[0094] where α, β, γ, and μ are hyperparameters, set to 0.001, 1, 0.1, and 100 respectively. The purpose of assigning these weights is to ensure that each component of the final loss function remains in the same order of magnitude after being multiplied by its corresponding weight.

[0095] (II) Experimental Section

[0096] To evaluate the enhanced performance of our proposed MVG-FD method, we selected three datasets: LSUI, UIEB, and UWCNN. The LSUI dataset contains 4,279 pairs of real underwater images, covering diverse water scenes, water types, lighting conditions, and target categories. The UIEB dataset includes 890 underwater images with different degradation scenes. The UWCNN dataset is based on the RGB-D, NYU-v2 indoor dataset and generates multiple degradation levels by applying the light attenuation coefficients of different water types to simulate a real underwater environment.

[0097] 2.1 Experimental Setup Details

[0098] The LSUI dataset was randomly divided into a training set Train-L (3,729 images) and a test set Test-L400 (400 images) for training and testing respectively. When inputting into the network, all images were resized to a fixed size of 256×256, and the pixel values were normalized to the range [0, 1]. We implemented the MVG-FD model using Python and the PyTorch framework on an NVIDIA RTX 3090 GPU. The model was trained for 800 epochs using the Adam optimizer with a batch size of 4. The initial learning rate was set to 0.0005 for the first 600 epochs and reduced to 0.0002 for the remaining 200 epochs. Additionally, the learning rate decayed by 20% every 40 epochs. Besides Train-L, we also used a second training set Train-U, which contains 800 pairs of underwater images from the UIEB dataset and 1,250 synthetic underwater images generated from the UWCNN dataset.

[0099] The test dataset consists of two parts. Full-reference test datasets: Test-L400 and Test-U90 (the remaining 90 pairs of images in UIEB); No-reference test dataset: Test-U60, including 60 real underwater images without reference images, from the UIEB dataset.

[0100] 2.2 Comparison Results

[0101] For the Test-L400 and Test-U90 datasets containing reference images, we conducted full-reference evaluations using the PSNR, SSIM, UIQM, and UCIQE metrics. For the images in the reference-free Test-U60 dataset, we adopted the UCIQE and UIQM reference-free evaluation metrics. A higher PSNR value indicates that the image content is closer to the reference image, a higher SSIM value reflects a higher similarity in structure and texture, and a higher UCIQE or UIQM score indicates a better human visual perception effect of the image.

[0102] MVG-FD was compared with nine state-of-the-art underwater image restoration methods to verify its performance advantages, see Figure 6 . The evaluation included four vision prior-based methods: ULAP, IBLA, WaterNet, and Ucolor, and five data-driven methods: FUnIE, UWCNN, U-Shape, UDA, and SMDR-IS.

[0103] Table 2

[0104]

[0105] As shown in Table 2, MVG-FD demonstrated strong performance on different datasets. On the Test-L400 and Test-U60 datasets, MVG-FD achieved the best performance in terms of the PSNR and SSIM metrics, and also showed excellent visual quality in terms of the UIQM and UCIQE metrics. In addition to individual performance metrics, we combined these metrics into a comprehensive metric called "ALL", representing the sum of all metrics. This measurement method provides an overall evaluation by balancing multiple criteria. Among the multiple metrics of each dataset, the MVG-FD method showed the best overall performance. It is worth noting that the SMDR-IS method performed best on the Test-U90 dataset but had poor performance on other datasets. We believe that its feature extraction method and loss function settings are too specific to a certain type of data, resulting in weak generalization performance.

[0106] Next, we will explore the problems existing in the existing methods. IBLA does not consider the inconsistent features of underwater images. The medium transmission map prior of Ucolor cannot effectively represent the attenuation degree of each region, resulting in unsatisfactory results in terms of contrast, brightness, and detailed texture. U-Shape performs well in color restoration of underwater images, but the MVG-FD method that combines depth information to generate multi-modal cues performs better in terms of contrast. Although the UDA and SMDR-IS methods focus on the refined restoration of multi-scale details, they still perform poorly in dealing with subtle and imperceptible details.

[0107] For the full-reference dataset, Figure 6 the visual comparison shows that our underwater image restoration results are closest to the reference image, with fewer color artifacts and higher target region fidelity, especially excelling in capturing details in underwater scenes. For the no-reference dataset, Figure 7 the visual comparison shows that the MVG-FD method provides higher contrast and exhibits excellent visual quality.

[0108] IBLA does not consider the inconsistent features of underwater images. The medium transmission map prior of Ucolor cannot effectively represent the attenuation degree of each region, resulting in unsatisfactory results in terms of contrast, brightness, and detailed texture. U-Shape performs well in color restoration of underwater images, but the MVG-FD method that combines depth information to generate multi-modal cues performs better in terms of contrast. Although the UDA and SMDR-IS methods focus on the refined restoration of multi-scale details, they still perform poorly in dealing with subtle and imperceptible details.

[0109] 2.3 Ablation Experiments

[0110] To verify the effectiveness of the multi-color space loss function, we conducted ablation experiments on the loss function. We trained the network on the Train-U dataset and then tested it on the Test-U90 dataset, and calculated SSIM and PSNR. The results are shown in Table 3.

[0111] Table 3

[0112]

[0113] Among them, BL represents the baseline model without adding additional color space loss, LAB represents adding LAB color space loss, and LCH represents adding LCH color space loss. The experiments show that adding LAB and LCH color space losses can improve the overall performance of the model.

[0114] To verify the effectiveness of the MVG and FD components, we conducted a series of ablation studies on the Test-L400 dataset. A single-modal U-Transformer++ model was designed by adding convolutional blocks to the U-Transformer model, serving as the baseline model (BL) for subsequent ablation studies.

[0115] Table 4

[0116]

[0117] The entire experiment was trained on the Train-L dataset, with Test-L400 as the test set. As shown in Table 4, BL performed worse than BL+MVG and BL+FD in terms of PSNR and SSIM metrics. This indicates that our MVG and FD components play an important role in underwater image restoration. Additionally, MVG-FD achieved the best quantization performance on the test dataset, highlighting the effectiveness of the combination of MVG and FD modules.

[0118] As Figure 8 shown, the complete model presented the best visual effects. Compared with the BL module, the results of BL+FD showed less noise and artifacts because high-frequency information helps reconstruct local details. The overall color of BL+MVG was closer to the reference image because depth information provides the color attenuation distance, thus promoting color restoration. The MVG and FD modules each have their specific functions during the enhancement process, and their combination can improve the overall performance of the entire network.

[0119] (III) Summary

[0120] In summary, this case first addresses the problem of underwater image restoration from the perspective of multi-modal prompt learning. By introducing the MVG module, this module effectively utilizes the complementarity between depth information and the original data to guide the color restoration of underwater images. Additionally, through the FD module, high-frequency information is efficiently utilized for the restoration of texture details. A large number of experiments have verified the excellent performance of MVG-FD. Finally, effectively utilizing low-frequency information and designing loss functions across multiple color spaces can further enhance the contrast and saturation of the output image.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that a corresponding system can be implemented according to the method of the present disclosure. Exemplarily, an underwater image restoration system based on multi-modal vision guidance and feature decomposition, characterized in that the system uses a trained image restoration model to restore underwater images, and the image restoration model is composed of an encoder, a U-Transformer model, and a decoder; wherein: the encoder is configured to sequentially obtain images of different scales based on the original underwater RGB image ​Based on the original underwater depth image Successively obtain images of different scales Based on the RGB image Modal features Depth image Modal features Fuse and generate prompt information i ∈ [1, 3]; Based on the RGB image Modal features , Decompose the high-frequency components through learnable wavelet decomposition, and then generate high-frequency features based on these high-frequency components; Combine the modal features of the RGB image Prompt information And high-frequency features are additively bound to obtain features Combine the modal features of the RGB image And prompt information Are additively bound to obtain features i ∈ [2, 3]; Use the modal features of the RGB image As features Use i ∈ [1, 4] as output features.

[0122] The above encoder includes a modal visual guidance module; The modal visual guidance module is configured to fuse and enhance the concatenated modal features; Split the fused and enhanced features into RGB features And depth features And Perform projection dimensionality reduction on To obtain the corresponding projection features as Obtain Enhanced modal information Finally, And Are additively bound to obtain mixed-modal prompt information i ∈ [1, 3].[[]END]]

[0123] The above encoder includes a feature decomposition module; The feature decomposition module is configured to perform learnable wavelet decomposition on the modal features to obtain high-frequency components lh, hl, hh, where lh represents the high-frequency details in the horizontal direction of the feature, hl represents the high-frequency details in the vertical direction of the feature, and hh represents the high-frequency details in the diagonal direction of the feature; Combine the high-frequency components in an element-wise addition manner to form initial high-frequency information , Based on the initial high-frequency information Obtain high-frequency features

[0124] The above encoder includes a high-frequency hybrid attention module; The high-frequency hybrid attention module is configured to implement the following arithmetic operations:​ CA(·) represents channel attention, SA(·) represents spatial attention, and ResB represents a residual block operation with batch normalization. The residual block is composed of a convolutional layer, batch normalization, and a ReLU activation function, and there are multiple convolutional layers.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the method or system of the present disclosure can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, in more cases for the present disclosure, software program implementation is a better implementation manner.

[0126] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. An underwater image restoration method based on multimodal vision guidance and feature decomposition, characterized in that: The encoder, U-Transformer model and decoder are used to form an image restoration model. The trained image restoration model is used to restore underwater images. i∈[1, 4] is the output feature of the encoder, The steps to obtain include: Based on the original underwater RGB image Get images of different scales in sequence Based on the original underwater depth image Get images of different scales in sequence Based on RGB image The modal characteristics of Depth Image The modal characteristics of Fusion generates prompt information Based on RGB image The modal characteristics of Decompose the high-frequency components through learnable wavelets, and then generate high-frequency features based on these high-frequency components; The modal features of the RGB image Tips Additive binding with high-frequency features to obtain features The modal features of the RGB image and prompt information Perform additive binding to obtain features The modal features of the RGB image As a feature 2. The method according to claim 1, characterized in that Based on RGB image The modal characteristics of Depth Image The modal characteristics of Fusion generates prompt information The steps include: Modal Features After splicing, preliminary fusion enhancement is performed; Separate the initial fusion enhanced features into RGB features and deep features And Perform projection dimensionality reduction and obtain the corresponding projection features: get Enhanced modal information Finally and Obtain mixed modal prompt information through additive binding 3. The method according to claim 2, characterized in that: Adopting modality vision guidance module, based on RGB image The modal characteristics of Depth Image The modal characteristics of Fusion generates prompt information The modality vision guidance module includes a fusion addition block, a projection block, and a space annotation block; The fusion enhancement block consists of multiple convolution blocks, each of which consists of a 3×3 convolution layer and a ReLU activation function; The projection block consists of a 1×1 convolutional layer; The space annotation block is Perform spatial fixation operations to obtain channel-level spatial attention masks 4. The method according to claim 3, characterized in that: Mask the attention and Perform element-wise multiplication to generate enhanced modal information 5. The method according to claim 1, characterized in that Based on RGB image The modal characteristics of The high-frequency components are decomposed by learnable wavelets, and then high-frequency features are generated based on these high-frequency components. The steps include: Modal Features After learnable wavelet decomposition, high-frequency components lh, h1, and hh are obtained, where lh represents the high-frequency details in the horizontal direction of the feature, h1 represents the high-frequency details in the vertical direction of the feature, and hh represents the high-frequency details in the diagonal direction of the feature; Combine the high-frequency components by element-wise addition to form the initial high-frequency information Based on initial high-frequency information Get high frequency features 6. The method according to claim 5, characterized in that: Using feature decomposition module, based on RGB image The modal characteristics of Decompose the high-frequency components through learnable wavelets, and then generate high-frequency features based on these high-frequency components; The feature decomposition module includes a high-frequency filter and a high-frequency mixed attention block; The high frequency filter is configured based on the input characteristics Perform wavelet decomposition; The processing of the high-frequency mixed attention block can be expressed as: Among them, CA(·) represents channel attention, SA(·) represents spatial attention, ResB represents a residual block operation with batch normalization, and the residual block is composed of a convolutional layer, batch normalization and a ReLU activation function, and the convolutional layer is multiple.

7. An underwater image restoration system based on multimodal vision guidance and feature decomposition, characterized in that: The system uses a trained image restoration model to restore underwater images. The image restoration model is composed of an encoder, a U-Transformer model and a decoder; wherein: The encoder is configured based on the original underwater RGB image Get images of different scales in sequence Based on the original underwater depth image Get images of different scales in sequence Based on RGB image The modal characteristics of Depth Image The modal characteristics of Fusion generates prompt information Based on RGB image The modal characteristics of The high-frequency components are decomposed by learnable wavelets, and then high-frequency features are generated based on these high-frequency components; the modal features of the RGB image are Tips Additive binding with high-frequency features to obtain features The modal features of the RGB image and prompt information Perform additive binding to obtain features The modal features of the RGB image As a feature Will as output features.

8. The system according to claim 7, characterized in that: The encoder includes a multimodal vision guidance module; The multimodal visual guidance module is configured to convert the modal features After splicing, preliminary fusion enhancement is performed; Then split the initial fusion enhanced features into RGB features and deep features And Perform projection dimensionality reduction and obtain the corresponding projection features: get Enhanced modal information Finally and Obtain mixed modal prompt information through additive binding 9. The system according to claim 7, characterized in that: The encoder includes a feature decomposition module, wherein the feature decomposition module is configured to decompose the modal features After learnable wavelet decomposition, high-frequency components lh, hl, and hh are obtained. lh represents the high-frequency details in the horizontal direction of the feature, h1 represents the high-frequency details in the vertical direction of the feature, and hh represents the high-frequency details in the diagonal direction of the feature; the high-frequency components are combined by element-wise addition to form the initial high-frequency information Based on initial high-frequency information Get high frequency features 10. The system according to claim 9, characterized in that: The encoder includes a high-frequency hybrid attention module; The high-frequency mixed attention module is configured to implement the following operations: Among them, CA(·) represents channel attention, SA(·) represents spatial attention, ResB represents a residual block operation with batch normalization, and the residual block is composed of a convolutional layer, batch normalization and a ReLU activation function, and the convolutional layer is multiple.

Citation Information

Cited By

  • Table area positioning correction method based on image edge detection

    CN120807564A

  • Layered hybrid network-based underlying visual color imaging learning method and device

    CN121353105A