Underwater polarization image enhancement method fusing CLIP pre-training large model
The CLIP pre-trained model and generative diffusion model with U-Net network enhance underwater polarized images, addressing the challenges of turbid underwater imaging by improving clarity, contrast, and color fidelity while adapting to diverse underwater conditions.
Patent Information
- Application Number
- CN202311643812.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-15
AI Technical Summary
Existing underwater image enhancement methods are not effective in turbid underwater environments and require a large number of paired data sets to be trained, making it difficult to optimize for different water environments.
The underwater polarization image enhancement method of fusion CLIP pre-trained large models is adopted. By combining the generative diffusion model with the U-Net network, the polarization image is encoded using the image encoder of the CLIP pre-trained large models to guide the diffusion model to generate clear underwater images, and combine the feature fusion module and the loss function to optimize the training process.
It improves the clarity, contrast and color reduction of underwater images, enhances the generalization ability and training efficiency of the model, and can generate high-quality images in different underwater environments.
Smart Images

Figure CN120318085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image enhancement, and particularly to an underwater polarization image enhancement method integrating a CLIP pre-trained large model. Background Art
[0002] Underwater image enhancement is to use optical technology and computer vision technology to convert the blurred and low-contrast images obtained by underwater shooting into clear and high-contrast images. It is one of the important application directions in the field of computer vision. Currently, the common method for underwater image enhancement is to use the method of deep neural network, encoding the original underwater image with a generative model, and then generating a high-quality underwater image through a decoder. These algorithms will adopt network models such as GAN network structure, transformer, and U-Net to learn the mapping relationship between underwater images and clear images. These algorithms have achieved good results in correcting the color cast problem of underwater images and improving the contrast of underwater images. Underwater polarization imaging uses the partial polarization property of light after passing through water to remove the scattered background light, and then obtains a clear underwater image. Underwater images with different polarization angles can reflect different characteristics of the target object. Using a deep learning network, learning the relationship between polarization images and clear underwater images, extracting the features of polarization images at different angles and fusing them, so as to obtain high-quality underwater images. However, when facing a relatively turbid underwater environment, due to the complexity of underwater imaging, these models cannot well restore clear underwater images. At the same time, most of these methods require a large number of paired data sets for training, and the acquisition of underwater image data sets is a very difficult thing. In addition, the content of suspended particles in the water body is different in different waters, resulting in different qualities of underwater captured images. Due to the lack of rich and diverse data sets, it is also very difficult to optimize the algorithm for a certain underwater environment. Summary of the Invention
[0003] The technical problem to be solved and the technical task proposed by the present invention are to improve and refine the existing technical solutions, and provide an underwater polarization image enhancement method integrating a CLIP pre-trained large model to achieve the purpose of improving the clarity of underwater image enhancement. To this end, the present invention adopts the following technical solutions.
[0004] An underwater polarization image enhancement method integrating a CLIP pre-trained large model includes the following steps:
[0005] 1) Obtain the underwater polarization image to be enhanced;
[0006] 2) Input the underwater polarization image to be enhanced into the trained underwater polarization image enhancement model; The training method of the underwater polarization image enhancement model is:
[0007] 21) Collect underwater polarization images, and form a large dataset containing different turbidity, light intensities, 0°, 45°, 90° linear underwater polarization images, circular polarization images, and clear underwater images by combining the collected underwater polarization image datasets.
[0008] 22) Use the image encoder of the CLIP pre-trained large model to encode the 0°, 45°, 90° linear underwater polarization images and circular polarization images to obtain feature maps of 4 underwater polarization images.
[0009] 23) Send the feature maps of the 4 underwater polarization images to the feature fusion module. The feature fusion module learns the feature maps, extracts image features, and fuses different features to output a feature image.
[0010] 24) Feed the feature maps output by the feature fusion module into the image generation module. The image generation module uses a diffusion model and combines the U-Net network and the transformer model to generate clear underwater images.
[0011] 25) According to the definition of the loss function, determine whether the value of the loss function reaches the set requirement. If so, the training of the underwater polarization image enhancement model ends.
[0012] 3) The underwater polarization image enhancement model outputs the enhanced underwater polarization image.
[0013] The underwater image enhancement method in the present invention does not simply train directly using a generative model with blurred and clear underwater images, nor does it directly fuse polarized images at different polarization angles using a neural network based on image fusion. Instead, it uses a generative diffusion model to input polarized images to generate clear underwater images. At the same time, the image encoder of the CLIP pre-trained large model is used to encode the polarized images to guide the diffusion model to generate, improving the clarity, contrast, color restoration degree, and texture details of underwater image enhancement, and being able to quickly extract the features of the images, thus accelerating the training process. At the same time, by combining the diffusion model and the U-Net network, clear underwater images can be efficiently generated, improving the training efficiency. By using a generative diffusion model, more features and patterns can be learned from the data, thus improving the generalization ability of the model. In addition, by using the image encoder of the CLIP pre-trained large model to encode the polarized images, the features and content of the images can be better understood, thus improving the adaptability of the model to different polarization angles and underwater environments. By combining the U-Net network and the transformer model, the color information of the original image can be better retained, thus improving the color restoration degree. In addition, through the learning and optimization of the feature fusion module, the feature information of different images can be better extracted and fused, thus improving the color richness and detail performance of the generated images.
[0014] As a preferred technical means: during training, freeze the image encoder of the CLIP pre-trained large model; use the feature map in step 23) as the prompt semantics of the diffusion model and send it into the forward process of the diffusion model as a prior for the diffusion model to learn; during the forward training process of the diffusion model, continuously add Gaussian noise fused with the prompt feature map to the clear picture:
[0015]
[0016]
[0017] where x0 represents the original input picture, x t represents the picture after adding noise at time t, q(.) represents the probability distribution, and {β} t=1:T represents a constant or a learned variance table for controlling the noise step size; in the forward process, sample x from x0 t
[0018]
[0019] where α t = 1 - β t and β t is the variance at time t,
[0020] In the process of generating clear underwater images by the diffusion model in the reverse direction:
[0021] q(x t-1 ∣x t ,y) = Zq(x t-1 ∣x t )q(y∣x t-1 )
[0022] where y represents the feature map of the prompt semantics, and Z is the normalization constant.
[0023] By freezing the image encoder of the CLIP pre-trained large model, repeated training can be avoided, saving computational resources and time; at the same time, using the feature map as the prompt semantics of the diffusion model can skip some unnecessary calculations and improve the training efficiency. By using the feature map as the prompt semantics, the diffusion model can be better guided in training, thereby improving the stability and reliability of the model; in addition, by controlling the step size and variance table of the noise, the addition of noise can be better controlled during the forward training process, thus better generating clear images.
[0024] As a preferred technical means: in step 22), the CLIP pre-trained large model uses Vision Transformer as the basic architecture of the image encoder; the input picture is divided into multiple parts, and then each part is projected into a vector of a fixed length and sent into the Transformer for image encoding, while position encoding is embedded; finally, the encoded feature map is output. When using Vision Transformer for image encoding, the input picture is divided into multiple parts, and then each part is projected into a vector of a fixed length and sent into the Transformer for encoding, which increases the robustness of the model, independently processes the image information of different parts, and reduces the impact of image noise and interference on the model performance. Vision Transformer has strong feature extraction capabilities and can effectively extract useful feature information from images. By dividing the input picture into multiple parts and projecting them, local details and features in the image can be better captured, improving the model's understanding and expression ability of the image content. In Vision Transformer, by embedding position encoding, the spatial information in the image can be better captured, helping the model better understand the object positions and relative relationships in the image and improving the model's parsing ability of the image.
[0025] As a preferred technical means: The diffusion model uses the U-Net network as the main architecture, adds a feature reconstruction module in the middle of the U-Net network, and adds a color reconstruction module in the residual connection of the U-Net network; the original feature map is downsampled three times and passed through a 1*1 convolutional layer, and then input into the feature reconstruction module. The color reconstruction module can better restore the color information of the image. By adding a color reconstruction module in the residual connection of the U-Net network, the color features of the original image can be better retained, and the color restoration degree can be improved. By downsampling the original feature map three times and passing through a 1*1 convolutional layer, the size and dimension of the input feature map can be better controlled, so that the feature reconstruction module can better process these feature information, and the stability and reliability of the model can be improved.
[0026] As a preferred technical means: During the generation process of the diffusion model, the U-Net network and the transformer model are combined to construct a generator;
[0027] 241) The image x with added noise T is directly input into the network, and at the same time the image x T is downsampled 3 times respectively; then after passing through a 1*1 convolution, the three-scale feature maps are input into the corresponding scale convolution blocks; the outputs of the four convolution blocks are the inputs of the feature reconstruction module and the color reconstruction module. After the feature remapping, the output of the feature reconstruction module is directly sent to the first convolution block; at the same time, the four convolution blocks of different scales will receive the four outputs of the color reconstruction module;
[0028] 242) The feature reconstruction module assists the network in modeling global information and strengthens the network's attention to severely degraded parts; assuming the size of the input feature map is X in ∈R H / 16*W / 16*C ;
[0029] For the one-dimensional sequence of the transformer, a linear projection is used to stretch the two-dimensional feature map into a feature sequence X in ∈R HW / 256*C ; In order to retain the valuable position information of each region, the learnable position embeddings are directly merged, expressed as:
[0030] X in = WX in + P E
[0031] where W*X i represents the linear projection operation, and P E represents the position encoding operation;
[0032] Then the feature sequence X inInput to the Transformer block, which contains 6 standard Transformer layers. Each Transformer layer contains a multi-head attention block (MTA) and a feed-forward network (FLN); the FLN includes a normalization layer and a fully-connected layer; the output of the l-th (l ∈ [1, 2,...., l]) layer in the Transformer block can be calculated as follows:
[0033] X l ′ = MTA(LN(X l-1 )) + X l-1
[0034] X l = FLN(LN(X′ l )) + X′ l
[0035] where LN represents layer normalization, and X l represents the output sequence of the l-th layer in the Transformer block;
[0036] The output feature sequence of the last Transformer block is X l ∈ R HW / 256*C , which is restored to the feature map of X out ∈ R H / 16*W / 16*C after feature remapping;
[0037] 243) The input to the color reconstruction module is the feature X of different sizes i ∈ R H / 2^i*W / 2^i*Ci , using a correlation filter with size P / 2 i *P / 2 i (i = 0, 1, 2, 3) and stride P / 2 i(i = 0, 1, 2, 3), perform linear projections on feature maps at different scales; then obtain four feature sequences; fuse and splice these four feature sequences and send them into 4 standard transformer modules to obtain feature vectors of color channels; through the feature mapping module, remap the feature vectors into 4 feature maps of different sizes and input them into the upsampling U-Net network. By combining the U-Net network and the transformer model, the advantages of both the convolutional neural network and the transformer can be utilized simultaneously to effectively capture multi-scale feature information. The downsampling and upsampling processes of the U-Net network can capture feature information at different scales, while the transformer model can effectively model the feature relationships between sequences. The feature reconstruction module can assist the network in modeling global information and strengthen the network's attention to severely degraded parts. The color reconstruction module can better restore the color information of the image, improving the color restoration degree and image quality. Utilizing the advantages of the U-Net network and the transformer model, efficient feature extraction and image generation are achieved. The convolutional layer of the U-Net network can quickly process the input image, while the transformer model can effectively model the feature relationships between sequences, accelerating the generation process.
[0038] As a preferred technical means: The loss function is defined as:
[0039] where E[] represents expectation, x0 represents the clear image, x’0 represents the generated image, and ‖.‖ 2 represents the L2 norm.
[0040] Beneficial effects: The present invention does not simply directly train using a generative model with blurred and clear underwater images, nor does it directly fuse polarized images at different polarization angles using a neural network based on image fusion. Instead, it uses a generative diffusion model to input polarized images to generate clear underwater images. At the same time, the image encoder of the CLIP pre-trained large model is used to encode the polarized images to guide the diffusion model to generate, improving the clarity, contrast, color restoration degree, and texture details of underwater image enhancement. Description of the Drawings
[0041] Figure 1 is the flowchart of the present invention. Detailed Embodiments
[0042] The technical solutions of the present invention will be further described in detail below in conjunction with the drawings in the specification.
[0043] As Figure 1As shown in the figure, the steps of underwater polarization image enhancement in the present invention are as follows: obtaining an underwater polarization image to be enhanced; inputting the underwater polarization image to be enhanced into an underwater polarization image enhancement model that fuses a pre-trained large CLIP model; and the enhancement model outputs the enhanced underwater polarization image.
[0044] Specifically, the training of the underwater polarization image enhancement model includes the steps:
[0045] S1: Construction of an underwater image dataset
[0046] S11: Adding substances such as different amounts of sediment and milk to the experimental water tank to simulate underwater environments with different turbidity and scattering degrees.
[0047] S12: Placing the water tank in a dark room and using an incandescent lamp with adjustable brightness to simulate natural light under different water depth conditions.
[0048] S13: Using a polarization camera to obtain 0°, 45°, 90° linear underwater polarization images and circular polarization images in the water tank.
[0049] S14: Replacing the water in the water tank with clear water, keeping other conditions unchanged, and taking clear underwater pictures.
[0050] S2: Using the image encoder of the pre-trained large CLIP model to encode the 0°, 45°, 90° linear underwater polarization images, circular polarization images and clear underwater images to obtain feature maps.
[0051] The pre-trained large CLIP model has tested multiple different backbone networks such as ResNet, EfficientNet, and Transformer in the image encoder. The CLIP pre-trained large model adopted in the present invention uses Vision Transformer (ViT) as the basic architecture of the image encoder. The input picture is divided into multiple parts, and then each part is projected into a vector with a fixed length and sent to the Transformer for image encoding, while position encoding is embedded. Finally, the encoded feature map is output. The image encoder of CLIP is pre-trained through large-scale self-supervised learning. During the pre-training process, the model has been exposed to hundreds of millions of images to learn to encode different image contents. This self-supervised learning method helps the model learn to capture high-level semantic features in images. At the same time, the image encoder of CLIP has the ability of zero-shot learning. This means that it can encode concepts or objects that have not been seen before and accurately extract the high-level semantic features therein.
[0052] S3: During the forward process of diffusion model training, add noise that combines the feature maps of linear underwater polarization images at 0°, 45°, and 90° and the feature map of circular polarization images to the original image.
[0053] The forward process of the diffusion model can be expressed as p(x0):=∫p(x 0:T )dx 1:T , where x1…x T are latent variables with the same dimension as the data x0~q(x0). The joint distribution pθ(x 0:T ) is called the inverse process, which is a Markov chain with a learned Gaussian transition starting from p(x T )=N(x T;0 ,I).
[0054] S4: During the generation process of the diffusion model, use a combination of the U-Net network and the transformer model to construct a generator.
[0055] S41: Directly input the image x T added with noise into the network, and at the same time, downsample the image x T three times respectively. Then, after 1*1 convolution, input the three-scale feature maps into the corresponding scale convolution blocks. The outputs of the four convolution blocks are the inputs of the feature reconstruction module and the color reconstruction module. After feature remapping, directly send the output of the feature reconstruction module to the first convolution block. At the same time, the four convolution blocks of different scales will receive the four outputs of the color reconstruction module.
[0056] S42: The feature reconstruction module is used to replace the original bottleneck layer of the U-Net, which can assist the network in modeling global information and strengthening the network's attention to severely degraded parts. Assume that the size of the input feature map is X in ∈R H / 16*W / 16*C .
[0057] For the one-dimensional sequence of the transformer, use linear projection to stretch the two-dimensional feature map into a feature sequence X in ∈R HW / 256*C . To retain the valuable position information of each region, directly merge the learnable position embeddings, which can be expressed as:
[0058] X in =WX in +P E
[0059] where W*X i represents the linear projection operation, and P E represents the position encoding operation.
[0060] Then, input the feature sequence X inInput to the Transformer block, which contains 6 standard Transformer layers. Each Transformer layer contains a multi-head attention block (MTA) and a feed-forward network (FLN). The FLN includes a normalization layer and a fully-connected layer. The output of the l-th (l ∈ [1, 2,...., l]) layer in the Transformer block can be calculated as follows:
[0061] X′ l = MTA(LN(X l-1 )) + X l-1
[0062] X l = FLN(LN(X′ l )) + X′ l
[0063] where LN represents layer normalization, and X l represents the output sequence of the l-th layer in the Transformer block.
[0064] The output feature sequence of the last Transformer block is X l ∈ R HW / 256*C , which is restored to the feature map of X out ∈ R H / 16*W / 16*C after feature remapping.
[0065] The input to the color reconstruction module is features X of different sizes i ∈ R H / 2^i*W / 2^i*Ci , and linear projections are performed on the feature maps of different scales using correlation filters of size P / 2 i * P / 2 i (i = 0, 1, 2, 3) and stride P / 2 i (i = 0, 1, 2, 3). Four feature sequences are obtained. These four feature sequences are fused and concatenated and fed into 4 standard Transformer modules to obtain the feature vectors of the color channels. Through the feature mapping module, the feature vectors are remapped into 4 feature maps of different sizes and input into the upsampled U-Net network.
[0066] S5: During the training process, the loss function is defined as:
[0067]
[0068] where E[] represents the expectation, x0 represents the clear image, x’0 represents the generated image, and ‖.‖ 2 represents the L2 norm.
[0069] According to the definition of the loss function, when it is determined whether the value of the loss function reaches the set requirements, if so, the training of the underwater polarization image enhancement model ends.
[0070] During training, freeze the image encoder of the CLIP pre-trained large model. Use the feature map as the prompt semantics of the diffusion model and feed it into the forward process of the diffusion model as a prior for the diffusion model to learn.
[0071] During the forward training process of the diffusion model, continuously add Gaussian noise fused with the prompt feature map to the clear picture.
[0072]
[0073]
[0074] Among them, x0 represents the original input picture, x t represents the picture after adding noise at time t, q(.) represents the probability distribution, {β} t=1:T represents a constant or a learned variance table that controls the noise step size. During the forward process, x can be sampled from x0 t
[0075]
[0076] Among them, α t = 1 - β t , β t is the variance at time t,
[0077] During the process of the diffusion model generating a clear underwater image backward:
[0078] q(x t-1 |x t ,y) = Zq(x t-1 |x t )q(y|x t-1 )
[0079] Among them, y represents the feature map of the prompt semantics, and Z is the normalization constant.
[0080] The diffusion model uses the U-Net network as the main architecture, adds a feature reconstruction module in the middle of the U-Net network, and adds a color reconstruction module in the residual connection of the U-Net network. First, the original feature map is downsampled three times and passed through a 1*1 convolutional layer, and then input into the feature reconstruction module.
[0081] The feature reconstruction module replaces the bottleneck layer of the original U-Net network, assists the entire network in modeling global information, and strengthens the network's learning of the degraded part of the underwater image.
[0082] In the feature reconstruction module, first, the two-dimensional feature map is converted into a feature vector through a linear layer, and positional encoding is added to retain the position information of each region in the feature map. Then, the feature sequence is input into 6 standard Transformer modules. Each Transformer contains a multi-head attention module, a normalization layer, and a fully connected layer. After the Transformer module, there is a feature mapping module that remaps the feature vector into a feature map.
[0083] Since water has different absorption capabilities for light of different wavelengths, it causes serious color cast in underwater images. The color reconstruction module enhances the network model's repair of the attenuated colors of the image. The input of the color reconstruction module is feature maps of different sizes. Convolutional networks are used for filtering, and then linear projection and feature position encoding operations are performed on the feature maps of different sizes to obtain four feature sequences. These four feature sequences are fused and concatenated and fed into 4 standard Transformer modules to obtain the feature vectors of the color channels. Through the feature mapping module, the feature vectors are remapped into 4 feature maps of different sizes and input into the upsampled U-Net network.
[0084] Compared with the existing technologies, the underwater image enhancement method in the present invention does not simply directly train with a generative model using blurred and clear underwater images, nor does it directly fuse polarized images at different polarization angles using a neural network based on image fusion. Instead, it uses a generative diffusion model to input polarized images to generate clear underwater images. At the same time, it utilizes the image encoder of the CLIP pre-trained large model to encode the polarized images to guide the diffusion model to generate, improving the clarity, contrast, color restoration degree, and texture details of underwater image enhancement.
[0085] The above-described underwater polarized image enhancement method integrating the CLIP pre-trained large model is a specific embodiment of the present invention, which already reflects the substantial features and progress of the present invention. According to actual usage needs, under the inspiration of the present invention, equivalent modifications can be made to its shape, structure, etc., and all are within the protection scope of this solution.
Claims
1. An underwater polarization image enhancement method integrating a pre-trained large CLIP model, characterized in that It includes the following steps: 1) Obtain the underwater polarized vibration image to be enhanced; 2) Input the underwater polarized vibration image to be enhanced into the trained underwater polarization image enhancement model; The training method of the underwater polarization image enhancement model is as follows: 21) Collect underwater polarization images, and form a large dataset containing different turbidity, light intensity, 0°, 45°, 90° linear underwater polarization images, circular polarization images, and clear underwater images by combining the collected underwater polarization image datasets; 22) Use the image encoder of the CLIP pre-trained large model to encode the 0°, 45°, 90° linear underwater polarization images and circular polarization images to obtain the feature maps of 4 underwater polarization images; 23) Send the feature maps of the 4 underwater polarization images to the feature fusion module. The feature fusion module learns the feature maps, extracts image features, and fuses different features to output a feature image; 24) Send the feature maps output by the feature fusion module into the image generation module. The image generation module uses a diffusion model to combine the U-Net network and the transformer model to generate clear underwater images; 25) According to the definition of the loss function, judge whether the loss function value reaches the set requirement. If so, the training of the underwater polarization image enhancement model ends; 3) The underwater polarization image enhancement model outputs the enhanced underwater polarized vibration image.
2. The underwater polarization image enhancement method integrating the CLIP pre-trained large model according to claim 1, characterized in that: During training, freeze the image encoder of the CLIP pre-trained large model; Use the feature maps in step 23) as the prompt semantics of the diffusion model and send them into the forward process of the diffusion model as priors for the diffusion model to learn; During the forward training process of the diffusion model, continuously add Gaussian noise fused with the prompt feature maps to the clear pictures: where \(x_0\) represents the original input image, \(x\) t represents the image after adding noise at time \(t\), \(q(.)\) represents the probability distribution, \(\{\beta\}\) t=1:T represents a constant or learning variance table that controls the noise step size, \(N(.)\) represents the normal distribution. In the forward process, sample \(x\) from \(x_0\) t where α t = 1 - β t , β t is the variance at time t, During the backward generation process of the diffusion model to generate clear underwater images: q(x t-1 |x t ,y) = Zq(x t-1 |x t )q(y|x t-1 ) where y represents the feature map of the prompt semantics and Z is the normalization constant.
3. An underwater polarization image enhancement method integrating a CLIP pre-trained large model according to claim 2, characterized in that: In step 22), the CLIP pre-trained large model uses Vision Transformer as the basic architecture of the image encoder; Divide the input image into multiple parts, project each part into a vector of a fixed length and send it into the Transformer for image encoding, and at the same time embed the position encoding; Finally, output the encoded feature map.
4. An underwater polarization image enhancement method integrating a CLIP pre-trained large model according to claim 3, characterized in that: The diffusion model uses the U-Net network as the main architecture, adds a feature reconstruction module in the middle of the U-Net network, and adds a color reconstruction module in the residual connection of the U-Net network; Perform three downsamplings on the original feature map and pass through a 1*1 convolutional layer, and then input it into the feature reconstruction module.
5. An underwater polarization image enhancement method integrating a CLIP pre-trained large model according to claim 4, characterized in that: During the generation process of the diffusion model, use the combination of the U-Net network and the transformer model to construct the generator; (241) Input the image x with added noise directly into the network. At the same time, perform downsampling on the image x three times respectively. Then, after 1×1 convolution, input the three-scale feature maps into the corresponding scale convolution blocks. The outputs of the four convolution blocks are the inputs to the feature reconstruction module and the color reconstruction module. After feature remapping, directly send the output of the feature reconstruction module to the first convolution block. At the same time, the four convolution blocks of different scales will receive the four outputs of the color reconstruction module. T Input the image x directly into the network. At the same time, perform downsampling on the image x three times respectively. T After 1×1 convolution, input the three-scale feature maps into the corresponding scale convolution blocks. The outputs of the four convolution blocks are the inputs to the feature reconstruction module and the color reconstruction module. After feature remapping, directly send the output of the feature reconstruction module to the first convolution block. At the same time, the four convolution blocks of different scales will receive the four outputs of the color reconstruction module. 242) The feature reconstruction module assists the network in modeling global information and enhances the network's attention to severely degraded parts; assume that the size of the input feature map is X in ∈R H / 16*W / 16*C ; For the one-dimensional sequence of the transformer, use linear projection to stretch the two-dimensional feature map into the feature sequence X in ∈R HW / 256*C ; To retain the valuable position information of each region, directly merge the learnable position embeddings, expressed as: X in = WX in + P E where W*X i represents a linear projection operation, P E represents a positional encoding operation; Then the feature sequence X in is input into the Transformer block, which contains 6 standard Transformer layers. Each Transformer layer contains a multi-head attention block (MTA) and a feed-forward network (FLN); the FLN includes a normalization layer and a fully connected layer; the output of the l-th (l ∈ [1, 2,...., l]) layer in the Transformer block can be calculated as follows: X′ l = MTA(LN(X l-1 )) + X l-1 X l = FLN(LN(X′ l )) + X′ l where LN represents layer normalization, and X l represents the output sequence of the l-th layer in the transformer block; The output feature sequence of the last transformer block is X l ∈R HW / 256*C , which is restored to X out ∈R H / 16*W / 16*C feature map; 243) The input of the color reconstruction module is special X of different sizes i ∈R H / 2^i*W / 2^i*Ci , using a correlation filter with size P / 2 i *P / 2 i (i = 0, 1, 2, 3) and a step size of P / 2 i (i = 0, 1, 2, 3), perform linear projection on the feature maps at different scales; then obtain four feature sequences; fuse and splice these four feature sequences and send them into 4 standard transformer modules to obtain the feature vectors of the color channels; through the feature mapping module, remap the feature vectors into 4 feature maps of different sizes and input them into the upsampled U-Net network.
6. The underwater polarization image enhancement method integrating a CLIP pre-trained large model according to claim 5, characterized in that: The loss function is defined as: where E[] represents the expectation, x0 represents the clear image, x,0 represents the generated image, and ‖.‖ 2 represents the L2 norm.
Citation Information
Cited By
Semi-supervised learning underwater degraded image restoration method and related device
CN121235955A