An underwater image enhancement system and method based on multi-head transposed spatial attention
By introducing multi-head transposed spatial attention mechanism, wavelet transformation and probability adaptive instance normalization modules into the underwater image enhancement system, the problem of poor robustness of the existing technology in complex environments is solved, and higher quality image enhancement and stronger model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510336137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing underwater image enhancement algorithms are poorly robust in complex environments and cannot adapt to factors such as lighting conditions, weather conditions and sensor noise in different underwater environments.
A multi-head transposed spatial attention mechanism is introduced, combining wavelet transformation and probability adaptive instance normalization module to improve image quality and improve the generalization ability of the model.
This significantly improves image quality, enhances the robustness and generalization capabilities of the model, allowing the system to adapt to different underwater environments.
Smart Images

Figure CN119850463B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater image enhancement based on computer vision, and particularly relates to an underwater image enhancement system and method based on multi-head transposed spatial attention. Background Art
[0002] In recent years, underwater image enhancement has been an important research direction in underwater deep learning. Due to the particularity of the underwater environment, underwater images suffer from degradation due to the harsh and complex lighting conditions in water. The underwater environment, due to factors such as light scattering, absorption, and refraction, causes problems such as low contrast, blurriness, color cast, and a large amount of image noise in underwater images. Therefore, underwater image enhancement technology can improve the image quality by adjusting image parameters, and can better observe the underwater environment, and can also further lay a foundation for underwater target detection, underwater drones and other work.
[0003] To achieve underwater image enhancement, there are currently many studies focusing on underwater image enhancement technology. It is mainly divided into traditional methods and deep learning-based methods. Traditional methods include histogram equalization, contrast limited adaptive histogram equalization (CLAHE), ultra-high contrast map (UCM), etc. Deep learning includes methods based on convolutional neural networks (CNNs), generative adversarial networks (GANs), autoencoders, etc. These methods learn a large amount of underwater image data to adaptively restore the details and colors of the images, while removing noise and distortion. However, the existing various algorithms for underwater image enhancement cannot adapt to different underwater environments, resulting in poor robustness of the model in complex environments, such as different lighting conditions, weather conditions, sensor noise, etc. How to improve the generalization ability, reduce the dependence on training data, and improve the interpretability of the model remains an important research direction in the future. Summary of the Invention
[0004] In view of the above problems, the present invention introduces a probability network, wavelet transform, and multi-head transposed spatial attention, which can greatly improve the image quality, obtain better results in underwater image enhancement, and have high robustness, improve the generalization ability of the model, and can adapt to different underwater environments.
[0005] In the first aspect of the present invention, an underwater image enhancement system based on multi-head transposed spatial attention is proposed, including a feature extraction and encoding module, a probability adaptive instance normalization module, and a decoding module;
[0006] The feature extraction and encoding module includes a processing sub-module, a feature refinement sub-module, and a feature perception sub-module; the processing sub-module extracts local features from the input image by scanning with a convolution kernel and downsamples the feature map; the feature refinement sub-module performs feature refinement to enhance valuable features while suppressing features with less information, obtaining a refined feature map; the feature perception sub-module uses the multi-head transposed self-attention mechanism to handle long-range dependencies and adjusts the spatial dimension of the feature map through a transposed operation to enhance the spatial perception ability. At the same time, spatial attention calculation is added to assign weights to each spatial position, obtaining an enhanced feature map;
[0007] The probability adaptive instance normalization module matches the image features with the desired enhancement effect by adaptively adjusting the mean and standard deviation of the input features. Based on the learned distribution, it adjusts the feature statistics of the image based on random samples, aligning the input mean and standard deviation to match the corresponding statistics;
[0008] The decoding module reconstructs the feature map through multiple convolutions and activation functions, restores the spatial information to improve the resolution, and then uses the multi-head transposed self-attention mechanism to make the decoder focus on the key regions, dynamically adjusting the spatial weights of the feature map to enhance the feature expression of the target region, obtaining the final enhanced image result.
[0009] Preferably, the feature extraction and encoding module is specifically:
[0010] After the input image undergoes local feature extraction by two groups of 3×3 two-dimensional convolutions and activation functions, the output feature map a is obtained; a undergoes max pooling, the feature refinement sub-module, and two groups of 3×3 two-dimensional convolutions and activation functions for feature map dimensionality reduction and feature refinement operations, enhancing valuable features while suppressing features with less information, enhancing the expression ability of the model, and obtaining the refined feature map b; b undergoes further local feature extraction by two groups of 3×3 two-dimensional convolutions and activation functions to obtain the feature map c; c undergoes max pooling, 2 feature perception sub-modules, and upsampling for feature map dimensionality reduction and feature weight assignment to achieve feature enhancement, focusing on key regions and suppressing redundant information, obtaining the feature map d, d is added to c to get e, e undergoes further feature extraction by two groups of 3×3 two-dimensional convolutions and activation functions and upsampling and is added to b to get f1, f1 undergoes two groups of 3×3 two-dimensional convolutions and activation functions and upsampling and is added to a to obtain the output feature map.
[0011] Preferably, the feature refinement sub-module is specifically:
[0012] For the input feature map E, its size is B C H W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions. The input is split into two parts X1 and X2 along the channel dimension. A convolution operation is performed on X1, and then the convolved X1 and the unprocessed X2 are re - concatenated and rearranged into the size B (H W) C, that is, the two - dimensional spatial information is flattened into one - dimensional; the flattened features are passed into a linear layer and then divided into two parts x_1 and x_2 along the last dimension; x_1 undergoes depth - wise convolution processing, and the result after convolution is rearranged back into the size B(H W) C; finally, x_1 and x_2 are multiplied to form a gating mechanism to control the information flow through the multiplication operation, and then a linear transformation is performed through a linear layer and restored to B C H W size, that is, restored to the form of an image, and the processed output and the original input are added for residual connection.
[0013] Preferably, the feature perception sub - module includes a two - dimensional convolution, a wavelet convolution, and a multi - head transposed spatial self - attention unit; the input feature map first undergoes a 3*3 two - dimensional convolution, then passes through an activation function, then passes through two consecutive wavelet convolutions, and finally passes through the multi - head transposed spatial self - attention unit. The result obtained is added to the input feature map and then output;
[0014] The wavelet convolution includes a wavelet transform filter. After the input passes through the filter, the image is divided into four parts, and then after depth - wise convolution with a stride of 2, it passes through a two - dimensional convolution and a transposed convolution, and is added to the feature map of the input after two - dimensional convolution to obtain the output;
[0015] The multi - head transposed spatial self - attention unit includes a multi - head transposed self - attention sub - unit and a spatial attention sub - unit; the input feature map F passes through the multi - head transposed self - attention sub - unit, and the output attention weights are multiplied by the input feature map to obtain a feature map ,and then after passing through the spatial attention sub - unit, the spatial attention weights are obtained. This weight is multiplied by to obtain the output feature map.
[0016] Preferably, the multi - head transposed spatial self - attention sub - unit is specifically:
[0017] First, for the input feature map F with the size B C H W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions; first, a 1x1 two - dimensional convolution is performed on the input to obtain a feature map with the size B C H The tensor of W, representing the query q, key k, and value v in the attention mechanism; then, through a depth convolution operation on this result, an output of size B is obtained. C H The tensor of W; then the output value is split into three parts along the channel dimension: the query q, key k, and value v, with sizes: the size of q is B C H W, the size of k is B C H W, the size of v is B C H W;
[0018] The query tensor q is rearranged from B C H W to B Head (C / head) (H W), where head represents the number of attention heads, c / head represents the number of channels per head, and H W is the result after flattening the spatial dimensions; similarly, the key tensor k is rearranged in a similar way to obtain a tensor of size B Head (C / head) (H W); the value tensor v is rearranged in the same way to obtain a tensor of size B Head (C / head) (H W); then q and k are normalized in the last feature dimension and the dot product of the query q and key k is calculated to obtain a tensor of shape B Head (C / head) (H W); and a scaling and softmax operation is performed on this tensor to obtain the attention weights at each position, ensuring that their values are in the range [0, 1] and sum to 1, obtaining the attention weights attn;
[0019] By multiplying the attention weights attn with the tensor v, the final output is obtained, with a size of B Head (C / head) (H W); Then, rearrange the output tensor to restore it to the original spatial dimensions B C H W, but at this time the number of channels becomes head * c / head;
[0020] Finally, pass through a 1 1 two-dimensional convolutional layer to adjust the number of channels to be the same as the input dimension, obtaining the multi-head transposed self-attention weight matrix , and apply the multi-head transposed self-attention weight matrix to the input feature map F, then the feature map of the image obtained through multi-head transposed self-attention can be obtained, that is:
[0021]
[0022] Among them, represents the multi-head transposed self-attention weight matrix obtained by the above operations , which is used to calculate the attention weights between each element in the input feature map and other elements, capture global dependencies, and determine the influence degree of each element on other elements, so as to achieve effective fusion of information, represents the input feature map, represents the output feature map calculated by multi-head transposed self-attention.
[0023] Preferably, the spatial attention sub-unit is specifically:
[0024] First, for the input feature map , whose size is B C H W, where B is the batch size, C is the number of channels, and H and W are the spatial dimensions (height and width), perform global max pooling and global average pooling operations to obtain two B 1 H W feature maps; Concatenate the results of global max pooling and global average pooling along the channels to obtain a feature map with a size of B 2 H W feature map; Perform a 7 7 two-dimensional convolutional operation on the concatenated feature map to obtain a feature map with a size of B 1 H W feature map; Pass this feature map through the Sigmoid activation function to obtain the spatial attention weight matrix , that is:
[0025]
[0026] Among them, the spatial attention weight matrix , is the global average pooling operation, which performs dimensionality reduction and feature aggregation of the feature map, aggregates spatial information through average pooling, retains the global information of each feature map, facilitates better calculation of the spatial attention matrix, and generates the feature map produced by global average pooling as , and this feature map retains the global information of each feature map; is the max pooling operation, which is mainly used to extract local significant features, enhance the model's attention to key regions, help the model pay more attention to the region with the strongest response in the feature map, thereby improving the sensitivity to key features. The feature map produced by global max pooling , and this feature map extracts local significant features, represents a 7 × 7 convolution operation, that is, after connecting the results of average pooling and max pooling, a convolution operation is performed for feature enhancement, which can not only retain the global context information but also enhance the attention to local significant features, further fuse global and local information, and generate the spatial attention weight;
[0027] Apply the spatial attention weight matrix to the input feature map , and then the feature map of the image obtained through spatial attention calculation can be obtained;
[0028] Apply the multi-head transposed self-attention weight matrix and the spatial attention weight matrix to the input feature map F in sequence, and then the feature map of the image obtained through multi-head transposed self-attention and spatial attention calculation, that is, the final refined feature map .
[0029] Preferably, the probability adaptive instance normalization module is specifically:
[0030] For the input feature map with size B×C×H×W, where B, C, H, and W represent the batch size, number of channels, height, and width respectively. First, calculate the mean value of each channel through spatial dimensional average pooling in terms of height and width, compress the feature map into a single tensor, and obtain the tensor B×C×1×1; then, use a 1×1 convolution to obtain the mean value and log standard deviation vector B×2N×1×1, and split it into a mean value vector and a log standard deviation vector, where the size of the mean value vector is B×N×1×1 and the size of the log standard deviation vector is B×N×1×1, and construct a Gaussian distribution with the mean value vector and the standard deviation vector obtained by log standard deviation transformation;
[0031] Meanwhile, for each channel, the standard deviation in the spatial dimensions of height and width is calculated, and the feature map is compressed into a single tensor, resulting in a tensor of B×C×1×1, representing the degree of variation of the input features in the spatial dimensions. Then, a 1×1 convolution is used to obtain the mean and log standard deviation vectors of the standard deviation, B×2N×1×1, which are split into a mean vector and a log standard deviation vector. The mean vector has a size of B×N×1×1, and the log standard deviation vector has a size of B×N×1×1. A Gaussian distribution is constructed using the mean vector and the standard deviation vector obtained from the log standard deviation transformation.
[0032] After the distribution is constructed, random samples are extracted from the Gaussian distribution, and random samples are generated using reparameterized sampling. The random samples of the mean and standard deviation obtained have a size of B×N×1×1.
[0033] The random samples are then injected into the Adaptive Instance Normalization (AdaIN) unit to transform the statistical information of the received features, and the mean and standard deviation of the received features are aligned based on random sampling by extracting random activation values from the learned distribution.
[0034] Preferably, the Adaptive Instance Normalization (AdaIN) unit is specifically:
[0035] Randomly sample the mean and standard deviation of the latent variables from the posterior distribution. After being mapped through convolution operations respectively, the sizes change from B×N×1×1 to B×C×1×1, obtaining the feature maps corresponding to the mean and standard deviation. After the original feature map is instance-normalized, it is fused with the feature maps corresponding to the sampled mean and standard deviation, that is, by applying the random sample mean g and standard deviation h to the normalized features, the mean and standard deviation of the content features are adjusted to new target values.
[0036] The Adaptive Instance Normalization (AdaIN) unit is implemented through the following operations:
[0037]
[0038] Among them, is the input feature, and are the mean and standard deviation of the features of the image to be enhanced, that is, the content features, and are the mean and standard deviation of the enhanced features, that is, the style features.
[0039] Preferably, the input is convolved by a 3 3 two-dimensional convolution, then passed through an activation function, and then convolved by another 3 3 two-dimensional convolution and passed through an activation function, then passed through two consecutive feature perception sub-modules and passed through a 1 The output image is obtained after two-dimensional convolution of 1, where the structure of the feature perception sub-module is the same as that of the feature perception sub-module in the feature extraction and encoding module.
[0040] In the second aspect of the present invention, an underwater image enhancement method based on multi-head transposed spatial attention is proposed. The underwater image enhancement system described in the first aspect is applied, and the following process is included:
[0041] Take an underwater real image.
[0042] Input the preprocessed image into the underwater image enhancement system.
[0043] Output the enhanced underwater image.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention proposes multi-head transposed spatial attention, which effectively combines multi-head transposed self-attention and spatial attention. Among them, multi-head transposed self-attention combines the advantages of multi-head attention and transposed operation, making the model have significant improvements in capturing complex dependencies, enhancing context modeling, improving parallel computing efficiency, etc., and improving the interpretability and generalization ability of the model. At the same time, on this basis, spatial attention is fused, and by calculating the importance of different spatial positions, the weights of each position of the feature map are dynamically adjusted, so as to better focus on key regions when processing visual data. At the same time, wavelet convolution is introduced, which can better capture detail information and analyze it at multiple scales, has better robustness to noise and signal interference, and can effectively extract useful features from complex or noisy input signals, especially underwater images, and can effectively retain and reconstruct details. And a feature refinement sub-module is added to further refine the features of the input image, which can refine the feature representation of the image, improve the feature expression ability of the model, and significantly improve the task performance. And the adaptive instance normalization module for image style transfer is applied to image enhancement, further improving the effect of image enhancement. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is the overall principle block diagram of the underwater image enhancement system of the present invention.
[0048] Figure 2Schematic diagram of the network structure of the feature extraction and encoding module of the present invention.
[0049] Figure 3 Schematic diagram of the network structure of the probability adaptive instance normalization module of the present invention.
[0050] Figure 4 Schematic diagram of the network structure of the decoding module of the present invention.
[0051] Figure 5 Schematic diagram of the network structure of the feature refinement sub-module of the present invention.
[0052] Figure 6 Schematic diagram of the network structure of the feature perception sub-module of the present invention.
[0053] Figure 7 Schematic diagram of the network structure of the multi-head transposed spatial self-attention unit of the present invention.
[0054] Figure 8 Schematic diagram of the network structure of the multi-head transposed spatial self-attention sub-unit of the present invention.
[0055] Figure 9 Schematic diagram of the network structure of the spatial attention sub-unit of the present invention.
[0056] Figure 10 Schematic diagram of the network structure of the wavelet convolution unit of the present invention.
[0057] Figure 11 Original experimental image in the embodiment of the present invention.
[0058] Figure 12 Enhanced image in the embodiment of the present invention. Detailed implementation manners
[0059] An underwater image enhancement system based on multi-head transposed spatial attention provided by the present invention has an overall structure as Figure 1 shown, including a feature extraction and encoding module, a probability adaptive instance normalization module, and a decoding module.
[0060] Among them, the structure of the feature extraction and encoding module is as Figure 2As shown, the features of the input image are extracted. The preprocessed underwater target image is used as the input. After passing through a combination of multiple convolutions, activation functions, max pooling, feature refinement sub-modules, feature perception sub-modules, and upsampling in sequence, the feature extraction is performed and the output is obtained. First, through the processing sub-module (including multiple convolutions, activation functions, and max pooling operations), local features are extracted from the input data by scanning the input image with convolutional kernels, and the feature map is downsampled through activation functions and pooling operations to reduce the size of the feature map, generating a higher-level and more abstract feature representation. Then, through the feature refinement sub-module, the features are refined, enhancing valuable features while suppressing features with less information, obtaining the refined feature map. The underwater image degradation problem has diversity and complexity (such as local area fogging, global color shift, noise interference, etc.). The deep convolutional network in the feature refinement sub-module has a hierarchical structure. The lower layer can capture subtle edge and texture information, while the higher layer focuses on global illumination, structure, and semantic information. The extraction of multi-scale features can capture different frequency degradation phenomena, and the linear transformation helps to integrate this information to form a unified and clear restored image. The refined and efficient feature fusion and reconstruction can significantly improve the image degradation phenomenon in underwater images caused by problems such as low contrast, color deviation, and noise. Then, through the feature perception sub-module, it is used to further enhance the feature extraction ability. The multi-head transposed self-attention mechanism is used to handle long-range dependencies and the spatial dimension of the feature map is adjusted through transposed operations to enhance the spatial perception ability and thus improve the feature expression ability. At the same time, spatial attention calculation is added to assign weights to each spatial position, suppressing irrelevant or unimportant information, thereby obtaining richer feature maps at multiple scales and multiple angles, that is, local features are obtained through this module and information is extracted from the global context, combined with spatial information processing, to obtain the enhanced feature map.
[0061] The feature perception sub-module includes wavelet convolution and multi-head transposed spatial self-attention units. Among them, wavelet convolution is mainly used for multi-scale feature extraction and signal analysis. It can perform multi-scale analysis on images for the problem that various degradation factors in the underwater environment show different performances at different scales, decompose signals or images into sub-bands of different frequencies, and can capture both low-frequency parts (global information, such as background and overall brightness) and extract high-frequency details (edges, textures, etc.), which is very crucial for the subsequent restoration of underwater images. The multi-head transposed spatial self-attention unit processes long-range dependencies through multi-head transposed self-attention and adjusts the spatial dimension of the feature map through transposed operations to enhance the spatial perception ability and thus improve the feature expression ability. At the same time, the subsequent spatial attention calculation can obtain more abundant feature maps at multiple scales and multiple angles, that is, local features are obtained through this module and information is extracted from the global context. Combining spatial information processing, the receptive field of the model is expanded through multi-head transposed self-attention and spatial attention. A larger receptive field enables the processing unit to observe information in a larger area, thereby more accurately estimating the global underwater illumination and color distribution. This global underwater information helps to compensate for color deviation problems caused by water quality and light changes. And the increase in the receptive field can capture both local details and overall structural information at the same time, realizing the fusion of multi-scale features. In this way, both details can be retained, and the overall tone and brightness can be effectively balanced to achieve a more natural enhancement effect, reducing the interference of anomalies in a single local area on the overall result, thereby improving the stability and robustness of the enhancement effect. The global information helps to more accurately restore the contrast and color of the image. A larger receptive field can better capture the overall light change, thus achieving a balance among dynamic range compression, color consistency, and detail enhancement, and finally realizing better underwater image restoration.
[0062] Among them, the structure of the probability adaptive instance normalization module is as Figure 3 shown. Based on the feature map output by the feature extraction and encoding module, by adaptively adjusting the mean and standard deviation of the input features, the image features are matched with the desired enhancement effect. Through the learned distribution, the feature statistical information of the image is adjusted based on random samples, and the mean and standard deviation of the input are aligned to match the statistical information, that is, the mean and standard deviation are adjusted to obtain the adjusted feature map.
[0063] Among them, the structure of the decoding module is as Figure 4 shown. Based on the feature map adjusted by the mean and standard deviation output by the probability adaptive instance normalization module, the feature map is reconstructed through multiple convolutions and activation functions to restore spatial information and improve the resolution. Then, through the multi-head transposed spatial attention unit, the decoder focuses on key regions to suppress irrelevant information, dynamically adjusts the spatial weights of the feature map to enhance the feature expression of the target region, and obtains the final enhanced image result.
[0064] Underwater images usually have global problems such as uneven illumination, color shift, and haze, and there are also phenomena and problems such as uneven illumination in the whole picture. Ordinary image enhancement algorithms, due to their limited perception range and inability to estimate the global illumination situation, cannot accurately estimate the illumination and color deviation of underwater images. Moreover, various degradation factors in the underwater environment often show different characteristics at different scales, making ordinary enhancement algorithms unable to be effectively applied underwater. Therefore, in this invention, both the feature extraction and encoding module and the decoding module adopt a feature refinement sub-module and a feature perception sub-module. The feature perception sub-module includes wavelet convolution and a multi-head transposed spatial self-attention unit, where: Wavelet convolution is mainly used for multi-scale feature extraction and signal analysis. It can perform multi-scale analysis on images to address the problem that various degradation factors in the underwater environment show different characteristics at different scales, decompose signals or images into sub-bands of different frequencies, and can capture both low-frequency parts (global information, such as background and overall brightness) and extract high-frequency details (edges, textures, etc.), which is crucial for the subsequent restoration of underwater images; The multi-head transposed spatial self-attention mechanism processes long-range dependencies through multi-head transposed self-attention and adjusts the spatial dimension of the feature map through transposed operations to enhance spatial perception ability and thus improve feature expression ability. At the same time, the subsequent spatial attention calculation can obtain more abundant feature maps at multiple scales and multiple angles, that is, local features are obtained through this module and information is extracted from the global context, combined with spatial information processing. Through multi-head transposed self-attention and spatial attention, the receptive field of the model is expanded. A larger receptive field enables the processing unit to observe information in a larger area, thereby more accurately estimating the global illumination and color distribution underwater. This global information underwater helps to compensate for the color deviation problem caused by water quality and illumination changes. And the increase in the receptive field can capture both local details and overall structural information at the same time, realizing the fusion of multi-scale features. In this way, both details can be retained, and the overall tone and brightness can be effectively balanced, achieving a more natural enhancement effect, reducing the interference of abnormal single local areas on the overall result, and thus improving the stability and robustness of the enhancement effect. Global information helps to more accurately restore the contrast and color of the image. A larger receptive field can better capture the overall illumination change, thus achieving a balance among dynamic range compression, color consistency, and detail enhancement, and ultimately realizing better underwater image restoration.
[0065] In summary, this invention enables the image processing algorithm to integrate more global and local detail information, conduct more comprehensive optimization for the global degradation problems in underwater images, and thus improve the overall quality of image enhancement.
[0066] The following further elaborates on the specific structures of the three modules.
[0067] I. Feature Extraction and Encoding Module
[0068] The feature extraction and encoding module includes multiple convolutional layers, activation functions, max pooling, a feature refinement sub-module, a multi-head transposed spatial attention sub-module, and an upsampling unit; the input image undergoes local feature extraction through two sets of 3×3 two-dimensional convolutional layers and activation functions to obtain the output feature map a. Feature map a undergoes max pooling, the feature refinement sub-module, and two sets of 3×3 two-dimensional convolutional layers and activation functions for feature map dimensionality reduction and feature refinement operations, enhancing valuable features while suppressing features with less information, enhancing the model's expressive ability, and obtaining the refined feature map b. Feature map b undergoes further local feature extraction through two sets of 3×3 two-dimensional convolutional layers and activation functions to obtain feature map c. Feature map c undergoes max pooling, two multi-head transposed spatial attention sub-modules, and upsampling for feature map dimensionality reduction and feature weight allocation to achieve feature enhancement, focusing on key regions and suppressing redundant information, obtaining feature map d. Feature map d is added to c to obtain e. After e undergoes two sets of 3×3 two-dimensional convolutional layers and activation functions and upsampling, further feature extraction is performed and added to b to obtain f1. After f1 undergoes two sets of 3×3 two-dimensional convolutional layers and activation functions and upsampling, it is added to a to obtain the output feature map.
[0069] 1. Feature Refinement Sub-module
[0070] The feature refinement sub-module performs refined processing on the input image features through depth convolution and linear transformation. It uses depth convolution to extract rich multi-scale features and performs efficient feature fusion and reconstruction through linear transformation. The depth convolutional network has a hierarchical structure. The lower layers can capture fine edges and texture information, while the higher layers focus on global illumination, structure, and semantic information. Forming the extraction of multi-scale features can capture different frequency degradation phenomena, and linear transformation helps integrate this information to form a unified and clear restored image. The refined and efficient feature fusion and reconstruction can significantly improve the image degradation phenomena caused by problems such as low contrast, color deviation, and noise in underwater images. The structure is as Figure 5 shown, and mainly includes the following steps: Split the input features and perform partial convolution processing. Expand the feature dimension through linear transformation and apply a gating mechanism. Optimize the features using depthwise separable convolution and then combine with the original features for output. Finally, through residual connection, add the optimized features to the original input to enhance the model's expressive ability. The detailed steps are:
[0071] For the input feature map E with size B C H W, where B is the batch size, C is the number of channels, and H and W are the spatial dimensions (height and width). Split the input along the channel dimension into two parts X1 and X2, perform a convolution operation on X1, and then re-concatenate the convolved X1 and the unprocessed X2 and re-arrange them into size B (H W) C, that is, flatten the two-dimensional spatial information into one dimension. After passing the flattened features into the linear layer, they are divided into two parts, x_1 and x_2, along the last dimension. x_1 undergoes depth convolution processing, and the result after convolution is rearranged back to the size B (H W) C. Finally, x_1 and x_2 are multiplied to form a gating mechanism that controls the information flow through the multiplication operation. Then, it undergoes a linear transformation through the linear layer and is restored to B C H W size, that is, restored to the form of an image. The processed output and the original input are added together for residual connection.
[0072] 2. Feature perception sub-module
[0073] The feature perception sub-module in the feature extraction and encoding module is as Figure 6 shown, including two-dimensional convolution, wavelet convolution, and multi-head transposed spatial self-attention unit. The input feature map is convolved by a 3 3 two-dimensional convolution, then passed through an activation function (LeakyRelu), followed by two consecutive wavelet convolutions, and then through a multi-head transposed spatial self-attention unit. The resulting output is added to the input feature map and then output.
[0074] The feature perception sub-module includes two-dimensional convolution, wavelet convolution, and multi-head transposed spatial self-attention unit. The input feature map is convolved by a 3 3 two-dimensional convolution, then passed through an activation function (LeakyRelu), followed by two consecutive wavelet convolutions, and then through a multi-head transposed spatial self-attention unit. The resulting output is added to the input feature map and then output.
[0075] (1) Multi-head transposed spatial self-attention unit
[0076] Among them, the multi-head transposed spatial self-attention unit is as Figure 7 shown, including a multi-head transposed self-attention sub-unit and a spatial attention sub-unit. The input feature map F is multiplied by the attention weights output by the multi-head transposed self-attention sub-unit to obtain the feature map , which is used as the input to the spatial attention sub-unit to obtain the spatial attention weights. These weights are multiplied by to obtain the output feature map.
[0077] 1) Multi-head transposed self-attention sub-unit
[0078] Among them, the structure of the multi-head transposed self-attention sub-unit is as Figure 8 shown. First, for the input feature map F with the size of B C H W, where B is the batch size, C is the number of channels, and H and W are the spatial dimensions (height and width). First, perform a 1x1 2D convolution on the input to obtain a tensor of size B C H W, representing the query (q), key (k), and value (v) in the attention mechanism. Then, perform a depth convolution operation on this result to output a tensor of size B C H W. Next, split the output values into three parts along the channel dimension: query (q), key (k), and value (v), with sizes: the size of q is B C H W, the size of k is B C H W, and the size of v is B C H W.
[0079] Rearrange the query tensor q from B C H W to a shape of B Head (C / head) (H W), where head represents the number of attention heads, c / head represents the number of channels per head, and H W is the result of flattening the spatial dimensions. Similarly, perform a similar rearrangement on the key tensor k to obtain a tensor of size B Head (C / head) (H W). Perform a similar rearrangement on the value tensor v to obtain a tensor of size B Head (C / head) (H W). Then, normalize q and k in the last dimension (i.e., the feature dimension) and calculate the dot product of query q and key k to obtain a shape of B Head (C / head) (H The tensor of (W). Then, scale and perform a softmax operation on this tensor to obtain the attention weights at each position, ensuring that their values are within the range [0,1] and sum to 1, thus obtaining the attention weights (attn). Multiply the attention weights (attn) by the tensor v to obtain the final output, whose size is B Head (C / head) (H W). Then rearrange the output tensor to restore it to the original spatial dimensions B C H W, but at this time the number of channels becomes head * c / head, that is, the multi-head information is restored. Finally, through a 1 ×1 two-dimensional convolutional layer, adjust the number of channels to be consistent with the input dimension to obtain the multi-head transposed self-attention weight matrix , and apply the multi-head transposed self-attention weight matrix to the input feature map F, then the feature map of the image obtained through multi-head transposed self-attention can be obtained, that is:
[0080]
[0081] Among them, represents the multi-head transposed self-attention weight matrix obtained by the above operations , which is used to calculate the attention weights between each element in the input feature map and other elements, capture global dependencies, and determine the influence degree of each element on other elements, so as to achieve effective information fusion, represents the input feature map, represents the output feature map calculated by multi-head transposed self-attention.
[0082] This multi-head transposed self-attention weight matrix handles long-range dependencies and adjusts the spatial dimensions of the feature map through transpose operations to enhance the spatial perception ability, thereby improving the feature expression ability, expanding the receptive field of the model. A larger receptive field enables the processing unit to observe information in a larger area. Through this matrix, the underwater global illumination and color distribution can be estimated more accurately. This underwater global information helps to compensate for the color deviation problems caused by water quality and light changes. And the increase in the receptive field can capture both local details and overall structural information at the same time, realizing the fusion of multi-scale features, being able to retain details, effectively balance the overall tone and brightness, achieve a more natural enhancement effect, reduce the interference of anomalies in a single local area on the overall result, and thus improve the stability and robustness of the enhancement effect.
[0083] 2) Spatial attention sub-unit
[0084] Among them, the spatial attention sub-unit is as follows Figure 9 As shown, first, for the input feature map with the size of B C H W, where B is the batch size, C is the number of channels, and H and W are the spatial dimensions (height and width), global max pooling and global average pooling operations are performed to obtain two feature maps of B 1 H W; the results of global max pooling and global average pooling are concatenated along the channels to obtain a feature map with the size of B 2 H W; a 7 7 two-dimensional convolution operation is performed on the concatenated feature map to obtain a feature map with the size of B 1 H W; passing this feature map through the Sigmoid activation function can obtain the spatial attention weight matrix , that is:
[0085]
[0086] Among them, the spatial attention weight matrix , is the global average pooling operation, which is used for dimensionality reduction and feature aggregation of the feature map. By average pooling, spatial information is aggregated, and the global information of each feature map is retained, which is convenient for better calculation of the spatial attention matrix. The feature map generated by global average pooling is , and this feature map retains the global information of each feature map; it can converge the statistical information of the entire image, thereby helping the network understand the illumination and color characteristics of the global distribution of the image, and is used to solve problems such as uneven illumination, color deviation, and haze often existing in underwater images is the max pooling operation, which is mainly used to extract local significant features, enhance the model's attention to key regions, help the model pay more attention to the region with the strongest response in the feature map, thereby improving the sensitivity to key features, and is used for local detail focusing, mainly to solve problems such as local scattered light, low contrast, and noise existing in underwater images. Through local detail focusing, the regions with serious damage or insufficient details in the underwater image can be automatically identified and given higher weights, so that the enhancement algorithm can perform more effective detail restoration in these key regions. The feature map generated by global max pooling , and this feature map extracts local significant features represents 7 The convolution operation of 7, that is, after connecting the results of average pooling and max pooling and performing convolution operation for feature enhancement, can not only retain the global context information but also enhance the attention to local significant features, further fuse the global and local information, and generate the spatial attention weights.
[0087] Apply the spatial attention weight matrix to the input feature map to obtain the feature map of the image calculated by spatial attention, that is:
[0088]
[0089] where represents the calculation of the spatial attention weight matrix, which is used to calculate the relationship between each spatial position in the feature map and other positions, capture the global spatial dependence in the feature map, converge the statistical information of the entire image, help the network understand the illumination and color characteristics of the global distribution of the image, and is used to solve problems such as uneven illumination, color deviation, and haze often existing in underwater images. At the same time, through local detail focusing, it can automatically identify the areas in the underwater image with serious damage or insufficient details and give higher weights, so that the enhancement algorithm can perform more effective detail restoration in these key areas, and capture the global spatial dependence in the feature map. represents the input feature map, represents the output feature map after spatial attention calculation.
[0090] Apply the multi-head transposed self-attention weight matrix and the spatial attention weight matrix to the input feature map F in sequence to obtain the feature map of the image calculated by multi-head transposed self-attention and spatial attention, that is, the final refined feature map :
[0091]
[0092]
[0093] (2) Wavelet convolution unit
[0094] where the wavelet convolution is as Figure 10 shown; the wavelet convolution includes wavelet transform filters:
[0095]
[0096]
[0097]
[0098]
[0099] After passing through the filter, the image is divided into four parts:
[0100] The (Low-Low) represents the filter for the low-frequency part. The part obtained through this filter contains the main information and energy of the image and is used for image compression and feature extraction;
[0101] The (Low-High) represents the filter for the part from low frequency to high frequency. The part obtained through this filter mainly contains horizontal features;
[0102] The (High-Low) represents the filter for the part from high frequency to low frequency. The part obtained through this filter mainly contains vertical features;
[0103] The (High-High) represents the filter for the high-frequency part. The part obtained through this filter contains the least amount of information and is used for noise processing;
[0104] Through wavelet convolution, multi-scale features can be extracted. It can perform multi-scale analysis on the image for the problem that various degradation factors in the underwater environment show differently at different scales, decompose the signal or image into sub-bands of different frequencies, and can capture both the low-frequency part (global information, such as background, overall brightness) and extract high-frequency details (edges, textures, etc.), which is very crucial for the subsequent restoration of underwater images;
[0105] Then, after depth convolution with a stride of 2, through two-dimensional convolution and transposed convolution, and adding it to the feature map obtained by two-dimensional convolution of the input, the output is obtained.
[0106] II. Probability Adaptive Instance Normalization Module
[0107] The probability adaptive instance normalization module can effectively match the content features with the desired enhancement effect by adaptively adjusting the mean and standard deviation of the input features.
[0108] For the input feature map with dimensions B×C×H×W, where B, C, H, and W represent the batch size, number of channels, height, and width respectively. First, calculate the mean of each channel through spatial average pooling over height and width, compressing the feature map into a single tensor to obtain a tensor of B×C×1×1. Then, use a 1×1 convolution to obtain the mean and log standard deviation vectors of B×2N×1×1, and split them into a mean vector and a log standard deviation vector. The mean vector has dimensions B×N×1×1 and the log standard deviation vector has dimensions B×N×1×1. Construct a Gaussian distribution with the mean vector and the standard deviation vector obtained from the log standard deviation transformation, that is, the Gaussian distribution of the mean distribution of the input features. (N can take any value, depending on the enhancement effect, and here N is taken as 20).
[0109] Meanwhile, calculate the standard deviation of each channel through spatial calculation over height and width, compressing the feature map into a single tensor to obtain a tensor of B×C×1×1, which represents the degree of variation of the input features in the spatial dimension. Then, use a 1×1 convolution to obtain the mean and log standard deviation vectors of the standard deviation of B×2N×1×1, and split them into a mean vector and a log standard deviation vector. The mean vector has dimensions B×N×1×1 and the log standard deviation vector has dimensions B×N×1×1. Construct a Gaussian distribution with the mean vector and the standard deviation vector obtained from the log standard deviation transformation, that is, the Gaussian distribution of the standard deviation distribution of the input features.
[0110] After the distribution is constructed, extract random samples from the Gaussian distribution and generate random samples using reparameterized sampling.
[0111]
[0112]
[0113] where g and h are two random samples from the prior distributions of the mean and standard deviation of the input feature map respectively, where and are the N-dimensional Gaussian distributions of the mean and standard deviation of the input feature map respectively, is the mean of the mean Gaussian distribution of the input feature map, is the standard deviation of the mean Gaussian distribution of the input feature map, is the mean of the Gaussian distribution of the standard deviation of the input feature map, is the standard deviation of the Gaussian distribution of the standard deviation of the input feature map, and x is the original input feature map.
[0114] Note that during the inference phase, g and h only depend on the input image. During the training phase, the posterior distribution is learned through the input image and the corresponding reference image. That is, the main difference between the training and testing phases is that during the training phase, it is necessary to additionally calculate the posterior distribution (i.e., the enhanced reference image) and perform random sampling, while during the testing phase, the prior distribution is directly used for image generation. The random activation values extracted from the distribution learned during the training phase are used to align the mean and standard deviation of the received features. The random sample size of the mean and standard deviation obtained through this step is B×N×1×1.
[0115] The random sample is then injected into the Adaptive Instance Normalization (AdaIN) unit to transform the statistics of the received features. A typical AdaIN receives a content input and a style input and simply aligns the mean and standard deviation of the content input to those of the style input to match their statistics. In this method, the mean and standard deviation of the received features are aligned based on random sampling by extracting random activation values from the learned distribution. Latent variables (mean and standard deviation) are randomly sampled from the posterior distribution. After the mean and standard deviation are respectively mapped through convolutional operations, their sizes change from B×N×1×1 to B×C×1×1, obtaining the feature maps corresponding to the mean and standard deviation. After the original feature map is instance-normalized, it is fused with the feature maps corresponding to the sampled mean and standard deviation. That is, by applying the random samples g (mean) and h (standard deviation) to the normalized features, the mean and standard deviation of the content features are adjusted to new target values. This means that the statistics of the input features, i.e., the content features, will be adjusted to conform to the desired enhancement effect (i.e., the style features).
[0116] Specifically, the Adaptive Instance Normalization (AdaIN) unit is implemented through the following operations:
[0117]
[0118] where is the input feature, and are respectively the mean and standard deviation of the features of the image to be enhanced, i.e., the content features, that is, the features of the original underwater image, and are the mean and standard deviation of the enhanced features, i.e., the style features, that is, the features of the enhanced image.
[0119] The input is convolved by a 3 3 two-dimensional convolution, then passed through an activation function (LeakyRelu), and then convolved by a 3 3 two-dimensional convolution and passed through an activation function, then passed through two consecutive multi-head transposed spatial attention sub-modules, and passed through a 1 The output image is obtained after two-dimensional convolution of 1. The structure of the feature perception sub-module is the same as that of the feature perception sub-module in the feature extraction and encoding module.
[0120] IV. Explanation of Experimental Results
[0121] The original experimental image used in this embodiment is as Figure 11 shown, and the enhanced experimental results Figure 12 are shown. The model effect data uses two common image quality evaluation metrics, PSNR and SSIM, which are often used to compare the quality differences between the original image and the compressed or processed image. Among them, PSNR (Peak Signal-to-Noise Ratio) is a metric for measuring image quality and is often used in tasks such as image compression and denoising, representing the ratio between the signal intensity and the noise intensity of the image. The larger the PSNR, the better the image quality and the smaller the error. SSIM (Structural Similarity Index) is an image quality evaluation standard based on human visual perception. It takes into account not only the brightness and contrast of the image but also the structural information of the image. The goal of SSIM is to evaluate the degradation of image quality by simulating the perception mechanism of the human eye. As shown in Table 1.
[0122] Table 1 Experimental Results Table
[0123]
[0124] Among them, Retinex represents the enhancement algorithm of the retinex theory, CLAHE is the contrast-limited adaptive histogram equalization algorithm, and other models are all the names of conventional image enhancement algorithms in this field.
[0125] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0126] Although the specific implementation manners of the present invention have been described above, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative labor by those skilled in the art are still within the protection scope of the present invention.
Claims
1. An underwater image enhancement system based on multi-head transposed spatial attention, characterized by: It includes a feature extraction encoding module, a probability adaptive instance normalization module and a decoding module; The feature extraction and encoding module includes a processing submodule, a feature extraction submodule and a feature perception submodule; The processing submodule extracts local features from the input image scanned by the convolution kernel and downsamples the feature map; the feature extraction submodule refines the features, enhances the valuable features and suppresses the features with less information, and obtains the refined feature map; The feature perception submodule uses a multi-head transposition self-attention mechanism to process long-distance dependencies and adjusts the spatial dimension of the feature map through a transposition operation to enhance spatial perception capabilities. At the same time, it increases spatial attention calculation to assign weights to each spatial position to obtain an enhanced feature map. The probabilistic adaptive instance normalization module matches the image features with the desired enhancement effect by adaptively adjusting the mean and standard deviation of the input features, and adjusts the feature statistics of the image based on random samples through the learned distribution, aligning the input mean and standard deviation to match the corresponding statistics; The decoding module reconstructs the feature map through multiple convolutions and activation functions, restores spatial information and improves resolution, and then uses a multi-head transposed self-attention mechanism to focus the decoder on the key area, dynamically adjusts the spatial weight of the feature map to enhance the feature expression of the target area, and obtains the final enhanced image result.
2. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 1, characterized in that: The feature extraction coding module is specifically: The input image is subjected to two sets of 3×3 two-dimensional convolutions and activation functions for local feature extraction to obtain the output feature map a; a is subjected to maximum pooling, feature refinement submodules and two sets of 3×3 two-dimensional convolutions and activation functions to perform feature map dimensionality reduction and feature refinement operations, enhance valuable features while suppressing features with less information, enhance the expressiveness of the model, and obtain the refined feature map b; b is subjected to two sets of 3×3 two-dimensional convolutions and activation functions for further local feature extraction to obtain feature map c; c is subjected to maximum pooling, two feature perception submodules and upsampling for feature map dimensionality reduction and feature weight allocation to achieve feature enhancement, focus on key areas and suppress redundant information, and obtain feature map d, d is added to c to obtain e, e is subjected to two sets of 3×3 two-dimensional convolutions and activation functions and upsampling for further feature extraction and addition to b to obtain f1, f1 is subjected to two sets of 3×3 two-dimensional convolutions and activation functions and upsampling and then added to a to obtain the output feature map.
3. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 1, characterized in that: The feature extraction submodule is specifically: For the input feature map, its size is B C H W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions, split the input into two parts X1 and X2 along the channel dimension, perform a convolution operation on X1, and then reconcatenate the convolved X1 and the unprocessed X2 and rearrange them to size B (H W) C, that is, flattening the two-dimensional spatial information into one dimension; passing the flattened features into the linear layer and dividing them into two parts x_1 and x_2 along the last dimension; x_1 is processed by deep convolution, and the convolution result is rearranged back to size B (H W) C; Finally, x_1 and x_2 are multiplied to form a gating mechanism, which controls the information flow through the multiplication operation, and then undergoes a linear transformation through the linear layer, and then recovers to B C H The W size is restored to the image form, and the processed output is added to the original input for residual connection.
4. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 1, characterized in that: The feature perception submodule includes two-dimensional convolution, wavelet convolution and multi-head transposed spatial self-attention unit; its input feature map is first subjected to 3*3 two-dimensional convolution, then to activation function, then to two consecutive wavelet convolutions, and finally to multi-head transposed spatial self-attention unit, and the result is added to the input feature map and then output; The wavelet convolution includes a wavelet transform filter, and its input is divided into four parts after passing through the filter, and then after a depth convolution with a step size of 2, a two-dimensional convolution and a transposed convolution are performed, and the output is added to the feature map of the input after the two-dimensional convolution; The multi-head transposition spatial self-attention unit includes a multi-head transposition self-attention subunit and a spatial attention subunit; the input feature map F is obtained by multiplying the input feature map by the attention weight output by the multi-head transposition self-attention subunit. , and then input the spatial attention subunit to get the spatial attention weight, which is After dot multiplication, the output feature map is obtained.
5. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 4, characterized in that: The multi-head transposed spatial self-attention subunit is specifically: First, for the input feature map F, its size is B C H W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions; first perform a 1x1 two-dimensional convolution on the input to obtain a B C H The tensor W represents the query q, key k and value v in the attention mechanism; then the result is deep convolution, and the output size is B C H W tensor; then split the output value into three parts according to the channel dimension: query q, key k and value v, with sizes: q has a size of B C H The size of W, k is B C H The dimensions of W and v are B C H W; Rearrange the query tensor q from B C H W to B head (C / head) (H W) shape, where head represents the number of attention heads, C / head represents the number of channels per head, and H W is the result of flattening the spatial dimensions; similarly, a similar rearrangement is performed on the key tensor k to obtain a size of B head (C / head) (H W); perform the same rearrangement on the value tensor v to get a tensor of size B head (C / head) (H W); then q and k are normalized in the last feature dimension and the dot product of query q and key k is calculated to obtain a tensor of shape B head (C / head) (H W) tensor; and scale and softmax the tensor to obtain the attention weight of each position, ensuring that its value is in the range of [0, 1] and the sum is 1, and obtain the attention weight attn; By multiplying the attention weight attn with the tensor v, we get the final output, whose size is B head (C / head) (H W); then rearrange the output tensor to restore it to its original spatial dimension B C H W, that is, the number of channels is restored to C; Last passed 1 1 two-dimensional convolution layer, adjust the number of channels to make the number of channels consistent with the input dimension, and obtain the multi-head transposed self-attention weight matrix , transpose the multi-head self-attention weight matrix Applied to the input feature map F, we can get the feature map of the image obtained by multi-head transposed self-attention, that is: in, Represents the multi-head transposed self-attention weight matrix obtained by the above operation , used to calculate the attention weights between each element in the input feature map and other elements, capturing global dependencies, represents the input feature map, Represents the output feature map after multi-head transposed self-attention calculation.
6. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 4, characterized in that: The spatial attention subunit is specifically: First, for the input feature map , whose size is B C H W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions, perform global maximum pooling and global average pooling operations, and get two B 1 H The feature map of W; the results of global maximum pooling and global average pooling are spliced according to the channel to obtain a size of B 2 H The feature map of W; perform 7 on the concatenated feature map 7 two-dimensional convolution operation, resulting in a size of B 1 H The feature map of W; the spatial attention weight matrix can be obtained by passing the feature map through the Sigmoid activation function ,Right now: Among them, the spatial attention weight matrix , The global average pooling operation is used to reduce the dimension and aggregate the features of the feature map. The spatial information is aggregated by average pooling, and the global information of each feature map is retained to facilitate better calculation of the spatial attention matrix. The feature map generated by global average pooling is , this feature map retains the global information of each feature map; It is a maximum pooling operation, which is used to extract local significant features, enhance the model's attention to key areas, and help the model pay more attention to the areas with the strongest response in the feature map. The feature map generated by the global maximum pooling , the feature map extracts local salient features, Indicates 7 The convolution operation of 7, that is, the results of average pooling and maximum pooling are connected and then subjected to convolution operation for feature enhancement, which can not only retain the global context information, but also enhance the attention to local significant features, fuse global and local information, and generate spatial attention weights; The spatial attention weight matrix Applied to the input feature map On the above, we can get the feature map of the image calculated by spatial attention; In the above way, the multi-head transposed self-attention weight matrix and the spatial attention weight matrix Apply it to the input feature map F in sequence, and you can get the feature map of the image calculated by multi-head transposed self-attention and spatial attention, that is, the final refined feature map .
7. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 1, characterized in that: The probability adaptive instance normalization module is specifically: For the input feature map, the size is B×C×H×W, where B, C, H, and W represent the batch size, number of channels, height, and width, respectively. First, the average pooling of each channel in terms of height and width is calculated, and the feature map is compressed into a single tensor to obtain a tensor of B×C×1×1. Then, a 1×1 convolution is used to obtain the mean and logarithmic standard deviation vectors B×2N×1×1, which are divided into the mean vector and the logarithmic standard deviation vector, where the mean vector size is B×N×1×1 and the logarithmic standard deviation vector size is B×N×1×1. The Gaussian distribution is constructed using the standard deviation vector obtained by transforming the mean vector and the logarithmic standard deviation. At the same time, each channel calculates the standard deviation of the spatial dimension in terms of height and width, compresses the feature map into a single tensor, and obtains a tensor of B×C×1×1, which represents the degree of change of the input feature in the spatial dimension; then, a 1×1 convolution is used to obtain the mean of the standard deviation and the logarithmic standard deviation vector B×2N×1×1, which is divided into a mean vector and a logarithmic standard deviation vector, where the size of the mean vector is B×N×1×1 and the size of the logarithmic standard deviation vector is B×N×1×1, and the Gaussian distribution is constructed with the standard deviation vector obtained by transforming the mean vector and the logarithmic standard deviation; When the distribution is constructed, random samples are drawn from the Gaussian distribution, and random samples are generated using reparameterized sampling. The size of the random samples of the mean and standard deviation is B×N×1×1; The random samples are then injected into the adaptive instance normalization unit AdaIN to transform the statistics of the received features by extracting random activation values from the learned distribution to align the mean and standard deviation of the received features based on random sampling.
8. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 7, characterized in that: The adaptive instance normalization unit AdaIN is specifically: The mean and standard deviation of the latent variables are randomly sampled from the posterior distribution. After the mean and standard deviation are mapped by the convolution operation, the size changes from B×N×1×1 to B×C×1×1, and the feature maps corresponding to the mean and standard deviation are obtained; After the original feature map is instance normalized, it is fused with the feature map corresponding to the mean and standard deviation obtained by sampling, that is, the mean and standard deviation of the content feature are adjusted to the new target value by applying the random sample mean g and standard deviation h to the standardized feature; The adaptive instance normalization unit AdaIN is implemented by the following operations: in, are input features, and are the mean and standard deviation of the features of the image to be enhanced, i.e., the content features, i.e., the features of the original underwater image, and It is the mean and standard deviation of the enhanced features, namely the style features, that is, the features of the enhanced image.
9. The underwater image enhancement system based on multi-head transposed spatial attention as claimed in claim 1, characterized in that: The decoding module is specifically: Input 3 3 two-dimensional convolution after the activation function and then a 3 3 two-dimensional convolution after the activation function, and then through two consecutive feature perception sub-modules and 1 1, the output image is obtained after the two-dimensional convolution, in which the feature perception submodule structure is the same as the feature perception submodule of the feature extraction and encoding module.
10. An underwater image enhancement method based on multi-head transposed spatial attention, characterized in that: The underwater image enhancement system according to any one of claims 1 to 9 is applied, and includes the following process: Shoot and obtain real underwater images; Inputting the preprocessed image into the underwater image enhancement system; Output enhanced underwater image.
Citation Information
Patent Citations
Infrared image denoising method based on convolution transpose self-attention
CN118229572A
Apparatus and method of performing matrix multiplication operation of neural network
EP3832498A1