An Underwater Robot Vision Clarification Method Based on a Multimodal Fusion Network

Through a multimodal fusion network combining RGB and polarization information, the contrast and texture details of underwater images are enhanced, and the problems of underwater image blur and color deviation are solved, and the performance of underwater visual tasks is improved.

CN119478648BActive Publication Date: 2025-07-18NAN TONG QI ZHI ZHI NENG KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411488920.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-18
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

Underwater images cause blur and color deviation due to light attenuation and scattering of suspended particles, affecting the performance of underwater visual tasks.

Method used

A multimodal fusion network is adopted, combining RGB mode and polarization mode, and a fusion network is enhanced through detail focusing differential convolution and polarization-guided fusion network, which enhances the contrast and texture details of underwater images, and uses polarization degree and polarization angle information to restore the real color.

Benefits of technology

It significantly improves the clarity and quality of underwater images, reduces interference in computer vision tasks, and improves the performance of subsequent visual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478648B_ABST
    Figure CN119478648B_ABST
Patent Text Reader

Abstract

The present invention discloses an underwater robot vision clarification method based on a multi-modal fusion network, belonging to the technical field of underwater image processing, mainly including: collecting turbid underwater images and corresponding clear underwater images, constructing an underwater polarization image data set, including underwater polarization images, polarization degree images, and polarization angle images at different angles; constructing an underwater robot vision clarification model based on a multi-modal fusion network, including a multi-modal fusion network and an image enhancement network; training the multi-modal fusion network based on the underwater polarization image data set, and using pixel multi-scale fusion to update RGB information and polarization information during training to generate fusion features; training the image enhancement network using the underwater polarization image data set to obtain an image enhancement model based on the network; obtaining the turbid fusion features to be processed and inputting them into the network-based image enhancement model to obtain clear underwater images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of marine environment perception, digital image processing, and image enhancement, and particularly relates to an underwater robot vision clarity method based on a multimodal fusion network. Background Art

[0002] With the development of human exploration of the ocean, underwater robots have become an important tool for obtaining seabed information. However, due to the strong absorption and scattering of light by water media and the complex underwater environment, there is energy attenuation during the propagation of light underwater, and impurities and suspended particles in the water will cause light to scatter during propagation, resulting in the images captured underwater being blurrier than those captured on land. These problems will lead to color deviation and low clarity in the obtained images, seriously affecting the visual quality and the performance of downstream underwater vision tasks. Underwater optical imaging is currently the core method for completing underwater environment perception and detection tasks, playing an indispensable role in scientific research fields such as underwater robots, ocean scientific investigations, and some downstream vision tasks (underwater target recognition or tracking). Due to the complexity and instability of the underwater environment, underwater images often face problems such as color cast, low contrast, and blurriness. Specifically, when light propagates in water, it is first affected by the water depth. As the depth increases, light of different wavelengths will continuously attenuate. Therefore, since red light disappears first, underwater images tend to be mainly blue or green, resulting in color cast in the images. At the same time, different from the air medium on land, water contains a large number of suspended particles, which will cause light to scatter in the water, further challenging underwater optical imaging. In addition, due to the uncertainty of the movement of imaging equipment underwater, the obtained images have problems such as blurred details and low contrast, affecting visual perception and posing more severe challenges to subsequent advanced vision tasks. At the same time, due to the lack of underwater scenes and high-quality images, UIE faces various challenges, such as over-enhancement and blurred detail features. These problems limit the performance of UIE methods and lead to poor performance of downstream tasks. Summary of the Invention

[0003] In view of the deficiencies of the prior art, the present invention provides an underwater robot vision clarity method based on a multimodal fusion network. Inspired by multimodal learning, it attempts to introduce polarization information as an additional modality to enhance the original underwater image. A novel detail-focusing and polarization-guided multimodal fusion network is proposed, which integrates the RGB modality and the polarization modality to enhance underwater images. Detail-focusing differential convolution is used to capture more detail and edge information, and the degree of polarization information and polarization angle information are used to enhance the contrast and texture details of different regions in the image, so as to more accurately restore the true color of the image, reduce the serious interference it brings to subsequent computer vision tasks, and significantly improve the quality of underwater images. The technical solutions are as follows:

[0004] An underwater robot vision clarity method based on a U-Net multimodal fusion network, characterized by comprising the following steps:

[0005] S1. Collect turbid underwater images and corresponding clear underwater images, and construct an underwater polarization image dataset, where the underwater polarization image dataset includes underwater polarization images, degree of polarization images, and polarization angle images at different angles;

[0006] S2. Construct an underwater robot vision clarity model based on a multimodal fusion network, including a multimodal fusion network and an image enhancement U-Net network;

[0007] S3. Train the multimodal fusion network based on the underwater polarization image dataset. During training, use pixel multi-scale fusion to update RGB information and polarization information and generate fusion features;

[0008] S4. Train the image enhancement U-Net network using the underwater polarization image dataset to obtain an image enhancement model based on the U-Net network;

[0009] S5. Obtain the turbid fusion features to be processed and input them into the image enhancement model based on the U-Net network, so as to obtain clear underwater images.

[0010] Further, the method for constructing the underwater polarization image dataset in step S1 is as follows: By making different-color water bodies with different turbidity levels in a water body scene, use a polarization camera to collect turbid underwater images of objects in different-color water bodies with different turbidity levels, collect clear underwater images of objects in pure water, and use the clear underwater images as label images; According to the turbid underwater images and the label images, construct a training set and a test set.

[0011] Further, in step S2, the network architecture of the underwater robot vision clarity model based on the multimodal fusion network is an end-to-end structure, generating feature maps of different sizes at each level, so that the network can capture features at different scales; Before the fusion network, preprocess the turbid underwater polarization images in the training set to obtain RGB modal information and polarization modal information, and the polarization modal information includes degree of polarization information DoLP and polarization angle information AoLP; Use the obtained RGB modal information and polarization modal information as the input of the multimodal fusion network.

[0012] Further, the multimodal fusion network includes two modules: a feature fusion module and a polarization-guided fusion module, where

[0013] The feature fusion module robustly fuses the DoLP and AoLP features from the polarization mode input domain by leveraging global and local information; two spatial attention maps are generated based on the two token embedding sequences provided by the two Conformers for the two input features DoLP and AoLP, and the extracted convolutional features are then weighted according to the spatial attention maps and fused together to obtain the polarization mode features;

[0014] The polarization-guided fusion module is used to process the modal deviation, enhance the input features of the RGB modal input feature X by using the operation of attention, and update and generate the fused feature X by guiding with the polarization mode feature M * , the polarization mode feature and the RGB mode feature are concatenated and projected through a multi-layer perceptron to generate key (k x ), query (q x ), and value (v x ); by reducing the embedding height and width H, W of the query and key in the spatial dimension, the channel statistics S q , S k are learned, so as to obtain the guided and updated channel relationship M * .

[0015] Furthermore, the method of concatenating and projecting the polarization mode feature and the RGB mode feature through a multi-layer perceptron is as follows: to generate key (k x ), query (q x ), and value (v x ), as learnable parameters:

[0016]

[0017] [X * = FC(softmax(FC([q x ; k x )) ⊙ v x )

[0018] where ⊙ represents element-wise multiplication and FC represents a fully-connected layer with filtering.

[0019] Furthermore, by reducing the embedding height and width H, W of the query and key in the spatial dimension, the channel statistics S q , S k are learned, so as to obtain the guided and updated channel relationship M * The formula is as follows:

[0020] K m , Q m , V m= X, M, k x

[0021]

[0022] M * = M x + FC((s q Q m + s k K m ) ⊙ V m )。

[0023] Furthermore, the image enhancement model based on the U-Net network consists of three parts: an encoder part, a feature transformation part, and a decoder part; in the image enhancement U-Net network, feature extraction blocks are deployed from the first layer to the third layer, that is, different blocks are used at different levels to extract corresponding features, and the third layer uses a Detail Enhancement Attention Block (DEAB) to capture more detail and edge features;

[0024] The Detail Enhancement Attention Block (DEAB) includes a Detail Focus Convolution Block and a Content-Guided Attention Block. The Detail Focus Convolution Block uses differential convolution to integrate prior information, supplement the convolutional layers for parallel processing operations, and enhance the representation ability; by using the reparameterization technique, the detail aggregation convolution is equivalently transformed into a convolution operation, thereby reducing the number of parameters and computational costs;

[0025] The Content-Guided Attention Block adopts a dynamic fusion method. By assigning a unique spatial importance map to each channel, more useful information encoded in the features is obtained: the low-dimensional features from the encoder part are fused with the high-dimensional features from the decoder part, and the features are modulated by the learned spatial weights, thereby adaptively fusing the low-dimensional features from the encoder part with the corresponding high-dimensional features from the decoder part; the input features are also added through skip connections to alleviate the problem of gradient disappearance and simplify the learning process; the fused features are mapped through a 3×3 convolutional layer to obtain the final sharpened result.

[0026] Furthermore, two downsampling and two upsampling operations are adopted between different layers to ensure dimensional consistency. The downsampling operation halves the spatial dimension and doubles the number of channels; it is implemented through a convolutional layer by setting the stride value to 2 and setting the number of output channels to twice the number of input channels; the upsampling operation is regarded as the inverse form of the downsampling operation, and the downsampling operation is implemented through a deconvolutional layer, where the sizes of the first, second, and third layers are C×H×W,

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] The present invention proposes an underwater robot vision clarification method based on a multi-modal fusion network. This method can effectively avoid the imaging defect problems caused by existing underwater image clarification methods in the face of high turbidity. While removing turbidity, it improves the quality of underwater images, and has effectiveness and robustness, thus laying a theoretical and technical foundation for subsequent vision tasks such as panoramic underwater observation. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic flowchart of an underwater robot vision clarification method based on a multi-modal fusion network in an embodiment of the present invention.

[0031] Figure 2 It is an architecture diagram of an underwater robot vision clarification network based on a multi-modal fusion network in an embodiment of the present invention.

[0032] Figure 3 They are the turbid image, polarization angle image, and polarization degree image obtained by the underwater robot vision clarification network based on the multi-modal fusion network in an embodiment of the present invention.

[0033] Figure 4 It is a clear underwater image output by the underwater robot vision clarification network based on the multi-modal fusion network in an embodiment of the present invention.

[0034] Figure 5 It is a comparison table of the underwater robot vision clarification network based on the multi-modal fusion network in an embodiment of the present invention and other existing networks in 5 common underwater image quality evaluation indicators. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0036] It should be noted that the terms in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] As Figure 1 shown, the present invention provides an underwater robot vision clarification method based on a multimodal fusion network, which mainly includes the following steps:

[0038] S1. Construct an underwater polarization image dataset. The training dataset includes underwater polarization images with four different angles, polarization degree images and polarization angle images. Specifically, by making different-color water bodies with different turbidity levels in an artificially built indoor water body scene, collecting turbid underwater images of objects in different-color water bodies with different turbidity levels, collecting clear underwater images of objects in pure water, and using the clear underwater images as label images.

[0039] For example, by building an indoor turbid underwater polarization image acquisition platform, which includes a glass water tank, a polarization camera, a computer, an illumination system and a camera tripod. The size of the glass water tank is 150 cm × 35 cm × 50 cm, and the illumination system is three floodlights of different colors, namely blue, green and white. The steps for collecting turbid underwater images are as follows:

[0040] First, collect object images in water bodies with different colors and different turbidity levels. For each group, obtain object images in 5 different turbidity levels and 3 different color scenes, that is, turbid underwater images, and also collect clear label images. During the shooting and collection process, use physical methods to fix the polarization camera and the objects in the water tank so that their relative positions do not change. Second, prepare various objects such as corals, starfishes, conches, and shells, and divide them into two groups during the data collection process. One group is to cover the bottom of the water tank with bottom sand and gravel to simulate the seabed or riverbed, and fix the objects on it as underwater images in complex scenes. The other group fixes the objects on a pure white background board as underwater images in simple scenes. A total of 1600 turbid underwater images and their corresponding label images are obtained through the turbid underwater image acquisition platform. Among them, there are 800 turbid underwater images in simple scenes and complex scenes respectively, and the image resolution is 1024 × 1224.

[0041] S2. Construct an underwater robot vision clarification model based on a multi-modal fusion network, including a multi-modal fusion network and an image enhancement U-Net network; the network architecture of the underwater robot vision clarification model based on the multi-modal fusion network is an end-to-end structure, generating feature maps of different sizes at each level, so that the network can capture features of different scales.

[0042] Before the fusion network, preprocess the turbid underwater polarization images in the training set to obtain RGB information and polarization information. The polarization camera measures the polarization angle φ through focal plane spectroscopy. pol The light intensity I passing through the linear polarizer pol , and its calculation formula is:

[0043] Ipol = Iun * (1 + ρ * cos(2φ - 2φpol))

[0044] S0 = I 0° +I 90° =I 45° +I 135°

[0045] S1 = I 0° -I 90°

[0046] S2 = I 45° -I 135°

[0047]

[0048] Among them, I un is the total incident light entering the camera, which is generally unpolarized light; ρ is the linear polarization degree, φ is the linear polarization angle, S0, S1, and S2 represent Stokes constants, where S0 is also used to represent RGB modal information, and the polarization modal information includes DoLP and AoLP, where DoLP represents the polarization degree information and AoLP represents the polarization angle information, I 0° I 90° I 45° I 135° respectively represent the images captured by the four linear polarization states of the light recorded by the polarization camera at angles 0°, 45°, 90°, and 135°; subsequently, the obtained RGB modal information and polarization modal information are used as the input of the multi-modal fusion network.

[0049] S3. The multi-modal fusion network in this application includes: a feature fusion module and a polarization-guided fusion module. Train the multi-modal fusion network based on the underwater polarization image dataset. During training, use the feature fusion module and polarization guidance to update the RGB modal information and polarization modal information and generate fusion features.

[0050] S301. The feature fusion module is adopted to robustly fuse the DoLP and AoLP features from the polarization modal input domain by leveraging global and local information. Two spatial attention maps are generated according to the two token embedding sequences provided by two Conformers for the two input features DoLP and AoLP. The extracted convolutional features are then weighted according to the attention maps and fused together:

[0051] M φ ,M ρ = softmax(Ω(T φ ), Ω(T ρ ))

[0052]

[0053] where C and T are the convolutional features and token embeddings generated by the conv and trans branches in the Conformer respectively, is element-wise multiplication. M are the attention maps generated by φ(AoLP) and ρ(DoLP) respectively, and Ω is a function that first reduces the dimension of each token embedding to 1 through a fully connected layer and then reshapes the generated embedding into a two-dimensional map. By using DoLP and AoLP, the features extracted by Conformers at different layers can capture more details and edge information, enhance the contrast and texture details in different regions of the image, and thus more accurately restore the true color of the image.

[0054] S302. Since the polarization modality deviates greatly from the RGB modality, the importance of the clues collected from the RGB modality and the polarization modality is scene-related. Simply combining these clues may dilute the influence of strong clues with weak signals and even amplify the adverse effects of confounding clues. To address this modality deviation, a polarization-guided fusion module is designed to enhance the input features of the RGB modality through attention operations, and the polarization modality features M are used to guide the update and generate the fused features X * , and the polarization modality features and RGB modality features are concatenated and projected through a multi-layer perceptron to generate key (k x ), query (q x ), and value (v x ), as learnable parameters:

[0055]

[0056] [X * = FC(softmax(FC([q x ; k x )) ⊙ vx )

[0057] Among them, ⊙ represents element-wise multiplication, and FC represents a fully connected layer with filtering.

[0058] By reducing the embedding height and width H and W of the query and key in the spatial dimension, the channel statistics S of the query and key are learned q 、S k , thereby obtaining the channel relationship for guiding the update.

[0059] K m ,Q m ,V m =X,M,k x

[0060]

[0061] M * =M x +FC((s q Q m +s k K m )⊙V m )

[0062] S4. Use the underwater polarization image dataset to train the enhancement network, obtain the output of the multi-modal fusion network: the fusion features with turbidity to be processed, and input them into the image enhancement network based on the U-Net network, so as to obtain a clear underwater image; the image enhancement network based on the U-Net network consists of three parts: an encoder part, a feature transformation part, and a decoder part.

[0063] For the fusion features with turbidity to be processed, the goal of the U-Net is to restore the corresponding clear image. However, for tasks such as de-turbidity that are sensitive to details, it is not possible to only consider transforming features in the low-resolution space, which will lead to information loss. Therefore, feature extraction blocks are deployed from the first layer to the third layer in the U-Net network, and different blocks are used at different levels to extract corresponding features. The first layer and the second layer use conventional feature extraction blocks (DEB), and the third layer uses detail enhancement attention blocks (DEAB) to capture more details and edge features. At the same time, in the U-Net, two downsampling and two upsampling operations are adopted. The downsampling operation halves the spatial dimension and doubles the number of channels; it is implemented through a common convolutional layer by setting the stride value to 2 and setting the number of output channels to twice the number of input channels; the upsampling operation can be regarded as the inverse form of the downsampling operation, and the downsampling operation is implemented through a deconvolutional layer, where the sizes of the first layer, the second layer, and the third layer are C×H×W,

[0064] Furthermore, the proposed detail enhancement attention block consists of a detail-focusing convolution block and a content-guided attention block, which is used to enhance feature learning and thus improve the dehazing performance.

[0065] The detail-focusing convolution block uses differential convolution to integrate prior information, supplement the information of ordinary convolution, and enhance the representation ability. Then, by using the reparameterization technique, the detail aggregation convolution can be equivalently transformed into ordinary convolution, thus reducing the number of parameters and computational costs.

[0066] The content-guided attention block consists of channel attention and spatial attention, which are used to calculate the attention weights in the channel and spatial dimensions in turn. The channel attention calculates a channel vector, that is to recalibrate the features. The spatial attention calculates the spatial importance map, that is to adaptively indicate the information region. The content-guided attention block processes different channels and pixels unequally, thus improving the denoising performance.

[0067]

[0068]

[0069] where max(0, x) represents the ReLU activation function, represents the convolutional layer with a kernel size of k×k, and [·] represents the channel concatenation operation. and represent the features processed by global average pooling operation in the spatial dimension, global average pooling operation in the channel dimension, and global max pooling operation in the channel dimension respectively. To reduce the number of parameters and limit the complexity of the model, the first 1×1 convolution reduces the channel dimension from C to The second 1×1 convolution expands the channel dimension back to C.

[0070] W coa = W c + W s

[0071] Using the content-guided attention block, the exclusive spatial importance map of each single channel of the input features is obtained in a coarse-to-fine manner, while fully mixing the channel attention weights and spatial attention weights to ensure information interaction. According to the broadcasting rule, W c and W s are fused together through a simple addition operation to obtain the coarse spatial importance map Since W c is channel-based, W coa and X are consistent in the channel. To obtain the final refined spatial importance map W, W coaEach channel is adjusted according to the corresponding input features. Guided by the content of the input features, the final specific channel spatial importance map W is generated. In particular, W coa and each channel of X are rearranged in an alternating manner through the channel shuffle operation. Combined with the subsequent grouped convolutional layer, the number of parameters can be greatly reduced.

[0072]

[0073] where σ represents the sigmoid operation, CS(·) represents the channel shuffle operation, represents the grouped convolutional layer with a kernel size of k×k. In the implementation, the number of groups is set to C. The content-guided attention mechanism assigns a unique spatial importance map to each channel, guiding the model to focus on the important regions of each channel. Therefore, more useful information encoded in the features can be emphasized, effectively improving the defogging performance.

[0074] Secondly, the dynamic fusion scheme based on the content-guided attention block can effectively fuse features and help the gradient flow. It fuses the features after the downsampling operation with the corresponding features before the upsampling operation, adopting an encoder-decoder-like architecture. The feature F low from the encoder part is high fused with the feature F fuse from the decoder part to obtain F

[0075] F fuse = C 1×1 (F low ·W + F high ·(1 - W) + F low + F high )

[0076]

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An underwater robot vision clarity method based on a multi-modal fusion network, characterized in that, It includes the following steps: S1. Collect turbid underwater images and corresponding clear underwater images, and construct an underwater polarization image dataset, where the underwater polarization image dataset includes underwater polarization images, degree of polarization images, and polarization angle images at different angles; S2. Construct an underwater robot vision clarification model based on a multi-modal fusion network, including a multi-modal fusion network and an image enhancement U-Net network; S3. Train the multi-modal fusion network based on the underwater polarization image dataset. During training, use pixel multi-scale fusion to update RGB information and polarization information and generate fusion features; S4. Train the image enhancement U-Net network using the underwater polarization image dataset to obtain an image enhancement model based on the U-Net network; S5. Obtain the turbid fusion features output by the multi-modal fusion network to be processed and input them into the image enhancement model based on the U-Net network, so as to obtain clear underwater images.

2. The underwater robot vision clarity method based on a multi-modal fusion network according to claim 1, wherein, The method for constructing the underwater polarization image dataset in step S1 is as follows: By creating different-color water bodies with different turbidity levels in a water body scene, use a polarization camera to collect turbid underwater images of objects in different-color water bodies with different turbidity levels, collect clear underwater images of objects in pure water, and use the clear underwater images as label images; According to the turbid underwater images and the label images, construct a training set and a test set.

3. The underwater robot vision clarity method based on a multi-modal fusion network according to claim 1, wherein In step S2, the network architecture of the underwater robot vision clarification model based on the multi-modal fusion network is an end-to-end structure, generating feature maps of different sizes at each level, so that the network can capture features at different scales; Before the fusion network, preprocess the turbid underwater polarization images in the training set to obtain RGB modal information and polarization modal information, where the polarization modal information includes degree of polarization information DoLP and polarization angle information AoLP; Use the obtained RGB modal information and polarization modal information as the input of the multi-modal fusion network.

4. A method for underwater robot vision clarity based on a multi-modal fusion network according to claim 1, characterized in that The multi-modal fusion network includes two modules: a feature fusion module and a polarization-guided fusion module, where The feature fusion module robustly fuses the DoLP and AoLP features from the polarization modal input domain by using global and local information; Generate two spatial attention maps according to the two token embedding sequences provided by two Conformers for the two input features DoLP and AoLP. The extracted convolutional features are then weighted according to the spatial attention maps and fused together to obtain polarization modal features; A polarization-guided fusion module is used to process modal deviation, enhance the input features of the RGB modal input feature X by using the operation of attention, and update and generate the fusion feature X by guiding with the polarization modal feature M * , the polarization modal feature and the RGB modal feature are concatenated and projected through a multi-layer perceptron to generate key (k x ), query (q x ), and value (v x ); by reducing the embedding height and width H, W of the query and the key in the spatial dimension, learn the channel statistics S q , S k , so as to obtain the channel relationship M * for guided update.

5. The underwater robot vision clarity method based on a multi-modal fusion network according to claim 4, characterized in that The method of concatenating and projecting polarization mode features and RGB mode features through a multi-layer perceptron is as follows: to generate key (k x ), query (q x ), and value (v x ), as learnable parameters: [X * = FC(softmax(FC([q x ; k x )) ⊙ v x ) where, ⊙ represents element-wise multiplication, and FC represents a fully connected layer with filtering.

6. A method for underwater robot vision clarity based on a multi-modal fusion network according to claim 5, characterized in that, Reduce the embedding height and width H, W of the query and key in the spatial dimension, and learn the channel statistics S of the query and key q , S k , so as to obtain the channel relationship M for guiding the update * The formula is as follows: K m ,Q m ,V m =X,M,k x M * = M x + FC((s q Q m + s k K m ) ⊙ V m )。 7. A method for underwater robot vision clarity based on a multi-modal fusion network according to claim 1, characterized in that The image enhancement model based on the U-Net network consists of three parts: an encoder part, a feature transformation part, and a decoder part; In the image enhancement U-Net network, feature extraction blocks are deployed from the first layer to the third layer, that is, different blocks are used at different levels to extract corresponding features, and the third layer uses a detail enhancement attention block DEAB to capture more details and edge features; The detailed enhancement attention block (DEAB) includes a detail focus convolution block and a content-guided attention block. The detail focus convolution block uses differential convolution to integrate prior information, supplementing the convolution layers of the parallel processing operation and enhancing the representation ability. By using the reparameterization technique, the detail aggregation convolution is equivalently transformed into a convolution operation, thereby reducing the parameters and computational cost. The content-guided attention block adopts a dynamic fusion method. By assigning a unique spatial importance map to each channel, more useful information encoded in the features is obtained: the low-dimensional features from the encoder part are fused with the high-dimensional features from the decoder part, and the features are modulated by the learned spatial weights, so as to adaptively fuse the low-dimensional features from the encoder part with the corresponding high-dimensional features from the decoder part. The input features are also added through skip connections to alleviate the gradient vanishing problem and simplify the learning process. The fused features are mapped through a 3×3 convolution layer to obtain the final sharpened result.

8. A method for underwater robot vision clarity based on a multi-modal fusion network according to claim 7, characterized in that Two downsampling and two upsampling operations are adopted between different layers to ensure dimensional consistency. The downsampling operation halves the spatial dimension and doubles the number of channels; it is implemented through a convolutional layer by setting the stride value to 2 and setting the number of output channels to twice the number of input channels. The upsampling operation is regarded as the inverse form of the downsampling operation, and the downsampling operation is implemented through a deconvolutional layer, where the sizes of the first, second, and third layers are C×H×W respectively,

Citation Information

Patent Citations

  • CNN-based intensity image and polarization image fusion enhancement method

    CN116740515A

  • High-flexibility underwater intelligent robot

    CN117048814A