Self-adaptive underwater image enhancement method and system based on Retinex theory and Mamba
By combining Retinex theory and the adaptive underwater image enhancement method of Mamba architecture, the problem of underwater image enhancement in the prior art has been solved, high-quality underwater image enhancement is achieved, and the details and contrast of the image are significantly improved.
Patent Information
- Application Number
- CN202510525153.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Existing underwater image enhancement technology will lead to image information loss and tampering after color correction, which cannot effectively solve the problems of low contrast and color offset of underwater images.
Using the adaptive underwater image enhancement method based on Retinex theory and Mamba, the light estimation component and degradation characteristics of underwater scenes are extracted to perform adaptive enhancement, and high-quality underwater images are generated.
It effectively solves the problems of low contrast and color shift in underwater images, improves the detailed texture and contrast of the image, optimizes the model inference efficiency, and achieves significant performance improvements on multiple evaluation indicators.
Smart Images

Figure CN120047337A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of underwater image enhancement technology. More specifically, it relates to an adaptive underwater image enhancement method and system based on the Retinex theory and Mamba. Background Art
[0002] When light propagates in water, the scattering and attenuation effects cause common degradation problems in underwater images, such as low contrast and color shift. In addition, the noise pollution and artifact superposition generated during the exposure process of underwater imaging devices further exacerbate the blurring and distortion of image details. For underwater image processing problems, the technical solutions adopted in the prior art are mainly color correction, and there are the following several types: (1) The Chinese invention patent with the publication number CN119648588A discloses an image enhancement method based on high-frequency and low-frequency adaptive fusion, including: performing contrast enhancement on the original image in each color channel to obtain intermediate feature maps; calculating the pixel means of the intermediate feature maps in each color channel; performing color compensation on the intermediate feature maps according to the pixel means to obtain a color reconstruction image; decomposing the color reconstruction image into high-frequency image features and low-frequency image features; performing normalization processing and gamma correction on the high-frequency image features to obtain high-frequency corrected images; performing sharpening processing on the high-frequency corrected images to obtain high-frequency restored images; performing haze removal processing on the weights of the low-frequency image features to obtain low-frequency restored images, and adaptively fusing the high-frequency restored images and the low-frequency restored images to obtain an output image.
[0003] (2) The Chinese invention patent with the publication number CN119741245A discloses an image enhancement algorithm based on adaptive filtering. By analyzing the pixel values of the r, g, and b channels of the image, the attenuation degree of each channel of the low-quality image is evaluated, and an adaptive adjustment coefficient is designed to balance the color difference to obtain a color-corrected image. Then, a backscattered light estimation method is introduced to evaluate the scattering distribution in the color-corrected image. Finally, the transmittance is deduced by combining the dark channel prior theory, and the backscattered light and the transmittance are combined using the atmospheric scattering model to generate a final enhanced image with rich contrast and details.
[0004] It can be seen that in the existing schemes, the image enhancement methods first perform color correction (or color compensation), and then perform image attenuation repair. The images enhanced in this way can show good color effects, but there will be losses and tampering with image information. Summary of the Invention
[0005] To solve the above problems, the technical solution adopted in this application is an adaptive underwater image enhancement method based on the Retinex theory and Mamba, including: Optical information estimation sub-network: Taking the original underwater image and its corresponding illumination prior map as inputs, it successively passes through feature concatenation, the first convolutional layer, the multi-scale channel aggregation module layer, and the second convolutional layer to obtain the illumination factor map. The illumination estimation component of the underwater scene is extracted by the multi-scale channel aggregation module layer. The illumination factor map is subjected to Hadamard product operation with the original image to obtain the basic enhanced image; Degradation restoration sub-network: The basic enhanced image extracts multi-scale degradation features step by step through the encoder, restores the spatial resolution through decoder skip connections, and enhances the illumination-texture correlation through selective state space modeling of the WLMamba bottleneck module, and finally generates a high-quality underwater image enhancement result.
[0006] Optionally, the multi-scale channel aggregation module layer includes adjusting the detailed features initially extracted by the first convolutional layer simultaneously through the first upper branch, the first middle branch, and the first lower branch, and then performing element-wise multiplication on the results of the three branches to obtain the final illumination estimation component of the underwater scene; The first upper branch retains the original input features; The first middle branch is used for multi-scale aggregation on multiple channels; The first lower branch performs a pooling operation in the spatial direction to obtain a spatial descriptor representation, then learns the correlation between channels through the lower branch convolutional layer, and generates a channel weight representation through the Sigmoid function.
[0007] Optionally, the first middle branch evenly splits the input features according to the number of channels. Subsequently, a multi-scale downsampling strategy is implemented on the split input features through the feature channels of Channel 1 to 4. The specific operations are as follows: Channel 1: Extracts global semantic features by using an 8-fold downsampling operation; Channel 2: Balances local details and context information through a 4-fold downsampling operation; Channel 3: Performs a 2-fold downsampling operation to retain medium-scale features; Channel 4: Maintains the original underwater resolution to retain high-frequency details; Performs 3×3 depth convolution on the features after downsampling through Channel 1 to 3, and restores the dimension through an upsampling operation. Concatenates the features of the feature channels of Channel 1 to 4, and then performs an aggregation operation through the first middle branch convolutional layer to finally obtain multi-scale aggregated features.
[0008] Optionally, the WLMamba bottleneck module adjusts the features it receives simultaneously through the second upper branch, the second middle branch, and the second lower branch; The second upper branch retains the original input features; The second intermediate branch takes the input feature channels and first linearly increases them to , which is a predefined channel expansion factor, then performs depth convolution, and then passes through SiLU activation, SSM-2D layer, and layer normalization in sequence; The second lower branch first cascades and combines the input features through a channel attention module and a spatial attention module modulated by the optical information estimation component, and then passes through SiLU activation.
[0009] Optionally, the input of the channel attention module consists of the illumination estimation component and the intermediate feature map of image processing. First, the intermediate feature map captures the global statistical features of each channel through the global average pooling layer, and then the fully connected layer group performs a non-linear transformation on the compressed channel feature vector to learn the complex relationships between channels. Then, it is activated through the Sigmoid function to output the channel weights. In the feature recalibration stage, the above channel weights are multiplied element-wise with the channels corresponding to the illumination estimation component to achieve adaptive modulation of the feature map, and the final output is the weighted feature map.
[0010] Optionally, using the weighted feature map output by the channel attention module as the input of the spatial attention module, first perform average pooling and max pooling operations on the input feature map respectively, and then concatenate the average feature map and the max feature map along the channel dimension to generate a composite feature that combines global and local features. Perform spatial information fusion and non-linear modeling on the composite feature through a 3×3 convolutional kernel, generate spatial attention weights after activation through the Sigmoid function, and finally multiply the spatial attention weights element-wise with the illumination estimation component to achieve illumination-spatial characteristic modulation and obtain the feature map optimized by spatial attention.
[0011] Optionally, capture the long-range dependencies in the underwater image through the SSM-2D layer, including converting the input feature map into one-dimensional sequences in four directions through a flattening operation, arranging the two-dimensional features into one-dimensional sequences according to the scanning sequence of the directions, applying a dynamically parameterized state space model to each direction's sequence, restoring the output sequences in the four directions to a two-dimensional feature map, and performing attention-weighted fusion.
[0012] Optionally, the overall architecture of the degradation repair sub-network adopts a four-layer encoder-decoder structure, including cross-level skip connections and WLMamba bottleneck modules; the encoders in the 1-3 levels are respectively fused with the decoders in their corresponding levels through skip connections for multi-scale feature fusion, and the encoder features in the 4th level are adjusted by the WLMamba bottleneck module and then transmitted to the decoder in the 4th level.
[0013] Optionally, the skip connection consists of a feature alignment operation and an adaptive fusion operation. The feature alignment operation includes spatial alignment and channel alignment. The spatial alignment includes upsampling the encoder features to the decoder size through bilinear interpolation and then adjusting the number of channels through a 1×1 convolutional kernel to obtain the aligned features. ; The input of the adaptive fusion operation is composed of concatenating the output features of the feature alignment operation and the decoder output features in channels, and two learnable parameters are set. , where , let the decoder feature of the -th layer be , and the adaptively fused features are obtained; the adaptively fused features are input into the decoder of the -th layer to complete the adaptive fusion operation.
[0014] The present application also provides an adaptive underwater image enhancement system based on the Retinex theory and Mamba, including: Optical information estimation sub-network module: used to take the original underwater image and its corresponding illumination prior map as inputs, and sequentially pass through feature concatenation, the first convolutional layer, the multi-scale channel aggregation module layer, and the second convolutional layer to obtain the illumination factor map. The multi-scale channel aggregation module layer extracts the illumination estimation component of the underwater scene, and performs Hadamard product operation on the illumination factor map and the original image to obtain the basic enhanced image; Degradation restoration sub-network module: used to extract multi-scale degradation features of the basic enhanced image through the encoder step by step, restore the spatial resolution through the decoder skip connection, and enhance the illumination-texture correlation through the WLMamba bottleneck module for selective state space modeling, and finally generate a high-quality underwater image enhancement result.
[0015] The beneficial effects of the adaptive underwater image enhancement method and system based on the Retinex theory and Mamba provided by the present application are as follows: (1) The present application provides an underwater image enhancement model that integrates the Retinex theory, the Mamba architecture, and the U-shaped network. By establishing a mathematical representation system of the optical properties of water bodies, it innovatively combines the Retinex theory with the deep learning framework, effectively solving the deficiency of traditional neural networks in terms of interpretability. The model uses an encoder-decoder topological structure to achieve multi-scale feature interaction. By combining the underwater optical attenuation characteristics with the Retinex theory, the model can adaptively adjust the chromaticity compensation intensity to achieve image restoration under the physical characteristics constraints of the underwater scene.
[0016] (2)This study innovatively improves the traditional Retinex theory. By integrating the underwater light scattering-attenuation characteristics and the device exposure attenuation mechanism, an improved Retinex theory framework for the underwater environment is constructed, and the theoretical mapping is realized through a neural network architecture. It includes two sub-networks: the light information estimation sub-network and the degradation restoration sub-network. The light information estimation sub-network is based on the improved Retinex theory. Through the multi-scale channel aggregation module layer, a multi-scale feature decoupling mechanism is adopted to accurately extract the illumination estimation component and the basic enhanced image of the underwater scene. The degradation restoration sub-network innovatively designs the WLMamba module. The WLMamba module deeply integrates the spatio-temporal modeling ability of the Mamba architecture and the illumination estimation component interaction mechanism, and combines the skip connections of the U-shaped network to construct a dual-path feature interaction architecture, forming an adaptive degradation restoration network. Experiments show that this method has achieved significant performance improvements on the LSUI, UIEB standard datasets and the no-reference SQUID dataset. While maintaining high image quality indicators, it has greatly optimized the model inference efficiency, verifying the effectiveness of this theoretical framework in complex underwater environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0018] Figure 1 is the overall network structure diagram of the adaptive underwater image enhancement method based on the Retinex theory and Mamba provided by the embodiments of the present application; Figure 2 is the structure diagram of the light information estimation sub-network provided by the embodiments of the present application; Figure 3 is the structure diagram of the multi-scale channel aggregation module provided by the embodiments of the present application; Figure 4 is the structure diagram of the degradation restoration sub-network provided by the embodiments of the present application; Figure 5 is the structure diagram of the WLMamba module provided by the embodiments of the present application; Figure 6 is the schematic diagram of the working principle of SSM-2D provided by the embodiments of the present application; Figure 7 is the training-loss diagram provided by the embodiments of the present application; Figure 8 is the real underwater image provided by the embodiments of the present application; Figure 9 is Figure 8 the three-dimensional RGB chromaticity distribution diagram of Figure 10 is Figure 8 the RGB channel distribution histogram of Figure 11 is Figure 8 the RGB intensity curve graph (middle row); Figure 12 is the enhanced underwater image provided by the embodiment of the present application; Figure 13 is Figure 12 the three-dimensional RGB chromaticity distribution graph; Figure 14 is Figure 12 the RGB channel distribution histogram; Figure 15 is Figure 12 the RGB intensity curve graph (middle row); Figure 16 is the label map provided by the embodiment of the present application; Figure 17 is Figure 16 the three-dimensional RGB chromaticity distribution graph; Figure 18 is Figure 16 the RGB channel distribution histogram; Figure 19 is Figure 16 the RGB intensity curve graph (middle row). Detailed implementation manners
[0019] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] Embodiment 1 As Figure 1 shown, the present application provides an adaptive underwater image enhancement method based on the Retinex theory and Mamba, including: Optical information estimation sub-network: Taking the original underwater image and its corresponding illumination prior map as inputs, passing through feature splicing, the first convolutional layer, the multi-scale channel aggregation module layer and the second convolutional layer in sequence to obtain an illumination factor map, extracting the illumination estimation component of the underwater scene by the multi-scale channel aggregation module layer, and performing Hadamard product operation on the illumination factor map and the original image to obtain a basic enhanced image; Degradation repair sub-network: The basic enhanced image extracts multi-scale degradation features step by step through the encoder, restores the spatial resolution through the decoder skip connection, enhances the illumination-texture correlation through the WLMamba bottleneck module for selective state space modeling, and finally generates a high-quality underwater image enhancement result.
[0021] According to the traditional Retinex theory. Any natural image can be decomposed into a reflection component and an illumination component , and the expression is as follows: ; In the formula, ⊙ represents the Hadamard product operation.
[0022] However, this theory is based on the assumption of non-degraded images, which is significantly different from the actual underwater imaging scenario. The degradation of underwater images mainly stems from the following mechanisms: The medium scattering effect leads to a decrease in imaging quality, specifically including the forward scattering and backward scattering processes: Forward scattering (the reflected light deviates from the original path due to medium interference and is captured by the sensor) usually has a relatively small impact.
[0023] Backward scattering (the incident light is reflected by the water body or suspended particles and enters the imaging system before reaching the target) constitutes the dominant factor of scattering.
[0024] The wavelength-dependent attenuation characteristics cause color deviation. The water body absorbs the red light band most significantly, while the absorption of the blue-green light band is relatively weak.
[0025] The imaging process will significantly amplify the sensor noise and image artifacts, thereby causing under-exposure or color distortion.
[0026] To address the above degradation mechanisms, we introduce a perturbation term into the original Retinex model to establish an improved equation: ; ; In the formula represents the illumination component perturbation term, represents the perturbation term of the reflection component. The reflection component is regarded as the reference image under ideal exposure conditions. To obtain an image with better quality, a light factor mapping processing is applied to both sides of the improved formula, where : ; In the formula, characterizes the degradation process in which the noise artifacts generated during the imaging process are amplified by the light mapping , comprehensively reflects the coupling effect of water body scattering attenuation and exposure color deviation. The simplified equation can be expressed as: ; In the formula, represents the pure image after degraded illumination, characterizes the composite degradation function.
[0027] Furthermore, a Degradation Optimization Processing (ROP) framework is constructed: ; In the formula, As the light feature estimation module, it is responsible for extracting the underwater light prior features . The original input and its corresponding light prior map (obtained by calculating the pixel mean in the channel dimension) are input into this module, and the output includes the basic enhanced image and the light estimation component extracted based on the underwater scene features .
[0028] Subsequently, and are input into the degradation restoration module, and joint optimization is performed to achieve degradation compensation, and finally an enhanced image is generated.
[0029] Light Information Estimation Sub-network The traditional Retinex algorithm has limitations because it does not consider the underwater light propagation characteristics. As Figure 2 shown, this study innovatively designs a light information estimation sub-network. By mining the semantic correlation of underwater images, this sub-network constructs a multi-scale feature fusion mechanism, which can effectively extract the deep features closely related to the optical properties of water bodies. Specifically, the network uses a cross-channel attention mechanism to capture the color attenuation pattern, thereby significantly improving the representation ability of the feature map. The specific structure is as follows: According to the Degradation Optimization Processing (ROP) framework, the input of the entire light information estimation is the original underwater image input and its corresponding light prior map (obtained by calculating the pixel mean in the channel dimension) are concatenated in the channel dimension and enter the feature extraction. The specific structure is as follows: conv1x1 Fusion: First, a convolutional layer with a kernel size of 1×1 is used to fuse the concatenated result for preliminary detailed feature extraction.
[0030] Multi-scale Channel Aggregation Module (MSCA) Structure In view of the strong absorption of the red band and the weak absorption of the blue-green band by water bodies, this study innovatively proposes a multi-scale channel aggregation module (MSCA). This module dynamically integrates the multi-channel information of underwater images through an adaptive feature aggregation strategy, effectively compensates for the wavelength-dependent attenuation, and thus approaches the spectral equilibrium requirement of the gray world hypothesis (the ratio of the R, G, and B channels is 1:1:1). The specific structure is as Figure 3 shown, MSCA has three branches.
[0031] The first intermediate branch is mainly used for multi-scale aggregation on multiple channels. First, the input feature is evenly split into four parts according to the number of channels, that is, . Subsequently, for Figure 3 the characteristic channels of Channel 1-4 in perform a multi-scale downsampling strategy, and the specific operations are as follows: Channel 1: Extract global semantic features by performing an 8-fold downsampling operation; Channel 2: Balance local details and context information through a 4-fold downsampling operation; Channel 3: Perform a 2-fold downsampling operation to retain medium-scale features;
[0032] This operation generates a multi-scale feature representation through progressive scale compression. The design basis is that in the underwater environment, the red light band (corresponding to Channel 1) requires a larger range of context compensation due to strong absorption, while the blue-green bands (corresponding to Channel 2 and Channel 3) can retain more original details due to weak absorption, and the part based on the prior illumination map needs to be completely retained (corresponding to Channel 4), obtaining features.
[0033] For ease of calculation, the features after three multi-scale downsamplings are ; In the formula, upsample (·) represents the upsampling process, DWConv represents the depth convolution process.
[0034] The current features are concatenated on the channels and then aggregated through a 1×1 convolutional layer to obtain multi-scale aggregated features , and then is applied with the GELU activation function to obtain the final output features .
[0035] ; The first lower branch adjusts the channel importance of the entire input feature through channel attention. First, a pooling operation is performed in the spatial direction to obtain a spatial descriptor representation, then the correlation between channels is learned through a 1×1 convolutional layer, and a channel weight representation is generated through the Sigmoid function: ; In the formula, AvgPlool (·) represents average pooling, Conv represents the convolution operation, Sigmoid(·) represents the operation of generating channel weight representations through the Sigmoid function, and this equation corresponds to Figure 3 the first lower branch in Figure 5 (to distinguish it from the lower branch in the following text, the lower branch in Figure 3 is called the first lower branch, and the naming rules for the first upper branch and the first middle branch are the same). The ReLU function in the first lower branch is not reflected in the equation. The ReLU function is used between two 1×1 convolutional layers to capture non-linear relationships; The first upper branch represents the original input features. Since multi-scale operations may cause loss of some detailed information in the spatial dimension, retaining the original input can effectively supplement spatial details, thereby enhancing the integrity of feature representation.
[0036] Finally, the results of the three branches are multiplied element-wise to obtain the final illumination estimation component of the underwater scene : ; (3) 1x1 convolution: Integrate the illumination estimation component through a 1×1 convolution kernel to generate an illumination factor map . We set as a three-channel RGB tensor to better achieve the color enhancement effect. Finally, perform the Hadamard product operation on and the original image to obtain the basic enhanced image .
[0037] Degradation repair sub-network This application adopts a degradation repair sub-network based on a U-shaped adaptive network, aiming to correct the residual errors in the detail texture and channel response of the image processed by the optical information estimation sub-network, so as to generate high-quality underwater scene reconstruction results. The network structure is as shown in Figure 4 .
[0038] The overall framework of the degradation repair sub-network adopts a four-layer encoder-decoder structure, including cross-level skip connections and the WLMamba bottleneck module. This architecture takes the design of the WLMamba module as the core, and through the deep fusion of the channel-spatial domain characteristics of the illumination estimation component and the feature map, realizes the precise repair of underwater image details. Among them, the encoder extracts multi-scale degradation features step by step, the decoder restores the spatial resolution through skip connections, and the WLMamba bottleneck module enhances the illumination-texture correlation through selective state space modeling, and finally generates high-quality underwater image enhancement results. The design scheme of WLMamba is as follows: Use the WLMamba module, and the module architecture is as shown in Figure 5As shown, the input feature enters three branches.
[0039] The second upper branch retains the original input feature; The second middle branch: The feature channel is first increased to , where is a predefined channel expansion factor, set to 2 here, and then depthwise convolution (DWCNN) is performed, followed by the SiLU activation function, the SSM-2D layer, and layer normalization (LN: LayerNorm).
[0040] The second lower branch: The feature is first modulated by the optical information estimation component and cascaded with the spatial attention module, and finally activated by SiLU.
[0041] The overall fusion formula is as follows: ; where X 1 represents the process of the second middle branch, X 2 represents the process of the second lower branch, and X out represents the combination and output process of the second upper branch, the second middle branch, and the second lower branch. Linear(·) represents linear transformation, DWConv(·) represents depthwise convolution, SiLU represents activation through the SiLU function, SSM2D(·) represents processing through the SSM-2D layer, and LN(·) represents layer normalization. The second middle branch of the WL Mamba structure in Encoder 1 may not include the layer normalization process in the initial stage of feature processing. The second middle branch of the WL Mamba structure in Encoders 2-4 and the WL Mamba structure between Encoder 4 and Decoder 4 need to include the layer normalization process. Therefore, the structure in the above formula can represent the WL Mamba structure in Encoder 1, while Figure 5 the structure shown can represent the WL Mamba structure in Encoders 2-4, Figure 5 and the SiLU activation function is not shown in.
[0042] The input of the channel attention module (CA) consists of the illumination estimation component and the intermediate feature map of image processing . For the WL Mamba structure of Encoder 1, this intermediate feature map of image processing refers to the basic enhanced image , for the WL Mamba structure in Encoders 2 - 4 and the WL Mamba structure between Encoder 4 and Decoder 4, the intermediate image - processing feature map refers to the feature map output after being processed by the previous encoder. First, the feature map passes through a global average pooling layer to compress the spatial dimension, eliminating the dependence on spatial positions, thereby capturing the global statistical features of each channel. Subsequently, a fully - connected layer group performs a non - linear transformation on the compressed channel feature vectors to learn the complex relationships between channels, and then is activated through a Sigmoid function for output, with the output being a weight vector. In the feature recalibration stage, the weight vector is multiplied element - by - element with the corresponding channels of the illumination estimation component to achieve the adaptive modulation of the feature map. Through this mechanism, the responses of channels with strong pixel correlations are significantly enhanced, redundant channel information is suppressed, and the model focuses on the key feature dimensions. The final output is the weighted feature map .
[0043] The spatial attention module (SA) first performs average pooling and max - pooling operations on the input feature map respectively. Average pooling obtains global statistical features by calculating the mean of each channel's feature map, while max - pooling extracts the local maximum of each channel to retain the significant response regions. These two pooling operations capture spatial features from two dimensions: global statistics and local saliency. Average pooling emphasizes the overall distribution characteristics, and max - pooling focuses on the strongly activated regions. Subsequently, the average feature map and the max - feature map are concatenated along the channel dimension to generate a composite representation that combines global and local features. Spatial information fusion and non - linear modeling are performed on the composite features through a 3×3 convolutional kernel, and after being activated by a Sigmoid function, the spatial attention weights are generated. Finally, the feature map in this module is multiplied element - by - element with the illumination component to achieve illumination - spatial characteristic modulation, strengthen regional responses, and suppress redundant information, obtaining the feature map optimized by spatial attention.
[0044] (3) SSM - 2D (State Space Model - 2D, two - dimensional state - space model) module: According to the Mamba principle, SSM - 2D is used to capture long - range dependencies in underwater images. The principle is as Figure 6 shown, and its mathematical principle is as follows: ① Input feature serialization: The input feature map is converted into one - dimensional sequences in four directions through a flattening operation. S d represents the feature sequence flattened according to direction d: ; where, represents arranging the two - dimensional features into a one - dimensional sequence according to the scanning sequence in direction , with each sequence length being , and the element dimension being Among them here the four directions represented by the numbers are as follows: Main diagonal scan: from top left to bottom right direction, traversing in row-major order; Sub-diagonal scan: from bottom right to top left direction, traversing in reverse row-major order; Vertical scan: from bottom left to top right direction, traversing in column-major order; Anti-vertical scan: from top right to bottom left direction, traversing in reverse column-major order.
[0045] ② Multi-directional selective state space model: For each direction's sequence , apply a dynamically parameterized state space model: Parameter dynamic generation: Generate a direction-related parameter matrix from the input sequence through linear projection: ; Among them, represents a learnable projection matrix, is a discretization parameter related to the time step, represents the input feature at time step t, represents the input matrix at time step t, represents the output matrix at time step t, and softplus(·) represents the softplus activation function.
[0046] State space equation discretization: Perform zero-order hold (ZOH) discretization on the continuous state space parameter : ; In the formula, represents the discretized state transition matrix, represents the discretized input matrix, represents the matrix exponential, which is used to discretize the continuous state space parameter, where represents the identity matrix.
[0047] The discretized hidden state update equation is: ; In the formula, represents the hidden state at time step t, represents the hidden state at time step t−1.
[0048] The output feature is: ; Among them is the skip connection parameter matrix.
[0049] ③Multi-directional feature fusion, restoring the output sequences in four directions to a two-dimensional feature map and performing attention-weighted fusion: ; In the formula, Y d represents the two-dimensional feature map restored in the direction d, y (d) represents the output sequence in the direction d, Y represents the finally fused feature map, represents the feature of restoring the tensor to by the inverse operation in the direction , is the direction attention weight, generated by lightweight convolution: ; where represents the Sigmoid function, represents element-wise multiplication, [Y 1 ; Y 2 ; Y 3 ; Y 4 : represents concatenating the feature maps in four directions along the channels.
[0050] The U-shaped network area is composed as follows: Encoder: The encoder consists of a WLMamba and a downsampling convolutional layer (convolution kernel size is 3×3, stride is 2, padding is 1); Decoder: The decoder consists of a WLMamba and an upsampling transposed convolutional layer (convolution kernel size is 3×3, stride is 2, padding is 1, additional padding is 1); Skip Connection: The Skip Connection in the U-shaped network is used to bridge the corresponding levels of the encoder and the decoder to achieve multi-scale feature fusion. The specific implementation consists of two steps: feature alignment operation and adaptive fusion operation. The principle is as follows: Feature alignment operation: To achieve effective fusion of the encoder and decoder features at the th layer, it is necessary to ensure that their sizes match. The main process includes two aspects: spatial alignment and channel alignment. Spatial alignment is to upsample the encoder features to the decoder size through bilinear interpolation, and then adjust the number of channels through a 1×1 convolution kernel to obtain the aligned features .
[0051] Adaptive fusion operation: The input of the adaptive fusion is composed of concatenating the output features of the feature alignment operation and the decoder output features along the channels. Here, two learnable parameters are set, where . Let the The decoder features of the layer are , and the adaptively fused features are obtained . The result of the adaptively fused features is input into the decoder of the th layer to complete the adaptive operation.
[0052] Neural network training In terms of dataset selection, two reference datasets, UIEB (Underwater Image Enhancement Benchmark) and LSUI (Large Scale Underwater Image Dataset), are used for training. The UIEB dataset contains 890 real underwater images with corresponding labels and is randomly divided into UIEB-Train (800 images) and UIEB-Test (90 images) as the training set and test set. The LSUI dataset contains 4,279 real underwater images with corresponding labels. Here, we use the dataset officially divided by LSUI, including LSUI-Train (3,879 images) as the training set and LSUI-Test (400 images) as the test set. In addition, we also use the C60, U45, and UCCS benchmark datasets for no-reference dataset quality assessment. The C60 dataset is derived from the 60 challenge sets provided by the UIEB dataset official website. U45 contains 45 no-reference images, covering underwater fog scenes and underwater green scenes. UCCS contains 300 no-reference images with blue-green, green, and blue hues.
[0053] Training details This method is implemented using an NVIDIA RTX4090 graphics card under the PyTorch framework. The MAE (Mean Absolute Error) loss is used as the loss. The total number of training epochs is set to 100, and the batch size is set to 8. The Adam (Adaptive Moment Estimation) optimizer is used, and the initial learning rate is set to , and it steadily decreases to during the training process through the cosine annealing method.
[0054] In terms of data preprocessing: The input images are adjusted to a fixed size (224 × 224), the mean and variance of the overall dataset are calculated for normalization, and the training data is augmented by random rotation and flipping. The training-loss image is as Figure 7 shown.
[0055] Observing the above training process is sufficient to verify the convergence of this method during training.
[0056] Model result verification On the reference dataset, the peak signal-to-noise ratio (PSNR), structural similarity (SSIM), UCIQE (Underwater Color Image Quality Evaluation), and UIQM (Underwater Image Quality Measure) are used as evaluation metrics. For the non-reference dataset, UCIQE and UIQM are used as evaluation metrics. The above evaluation metrics are defined as follows: Peak signal-to-noise ratio: The peak signal-to-noise ratio (PSNR) is a commonly used image quality evaluation criterion. It calculates the mean squared error (MSE) between the original image and the processed image, and then converts it into decibels to represent the image quality. The formula is as follows: ; where and are the original image and the processed image respectively, and are the height and width of the image, and are the pixel values of the image at position respectively.
[0057] The higher the PSNR value, the smaller the difference between the processed image and the original image, and the better the image quality.
[0058] Structural similarity (SSIM) is a metric used to measure the similarity between two images. It not only considers the gray difference of the images, but also considers the structural information of the images, which is more in line with the perception of image quality by the human visual system. The formula is as follows: ; where are the means of images and respectively, represent the variances of images and respectively, is the covariance of the image, C 1 and C 2 are constants used to stabilize the formula and prevent the denominator from being zero.
[0059] The value range of SSIM is usually between [-1, 1]. When SSIM approaches 1, it means that the correlation between the two images is greater; when it approaches -1, it means that the correlation between the two images is smaller; when the SSIM value is 0, it means that there is no linear correlation between the two images.
[0060] UCIQE is an objective metric for evaluating the quality of underwater color images. It comprehensively considers multiple features of underwater images, including brightness, contrast, and saturation, etc., to provide a comprehensive image quality assessment. The formula is as follows: ; In the formula, represents the image contrast, represents the image saturation, represents the image brightness, represent the weight coefficients respectively, which are used to adjust the contributions of contrast, saturation, and brightness in the overall assessment.
[0061] Contrast can be measured by calculating the standard deviation of the grayscale histogram of the image. The larger the standard deviation, the higher the contrast of the image. The calculation formula is as follows: ; In the formula, represents the grayscale value of the th pixel in the image, represents the average grayscale value of the image, represents the total number of pixels in the image.
[0062] Saturation can be measured by calculating the average value of the saturation channel in the HSV color space of the image. The higher the average value, the higher the saturation of the image. The formula is as follows: ; In the formula, is the saturation of the th pixel in the image.
[0063] Brightness is measured by calculating the average value of the grayscale values of the image. The formula is as follows: ; UIQM is a comprehensive metric for evaluating the quality of underwater images. UIQM is designed to be closer to the human visual system's perception of image quality, thus providing an objective and comprehensive image quality assessment. The formula is as follows: ; In the formula, represents the weight parameter, represent the image contrast, image saturation, and image brightness respectively. The specific formulas can be found in UCIQE.
[0064] Experimental results Through systematic experimental analysis, we conducted a horizontal comparative study on nine methods among the SOTA (State of the Art) and the method provided in this application. The nine methods among the SOTA specifically include: 2 traditional methods, 3 Convolutional Neural Network (CNN) - based methods, 3 Transformer - based methods, and 1 emerging Mamba - based method.
[0065] Table 1 Experimental Results of Traditional Methods and CNN - based Methods
[0066] ↑: An increase in the value of this metric represents an improvement in performance.
[0067] GDCP: Gradient Domain Color Prior; MLLE: Minimum Local Learning Enhancement; CWR: Color Weighted Reconstruction; SCNet: Spatial Convolutional Network; FiveA+: FiveA+ (a deep - learning - based image enhancement method); TIP'18: IEEE Transactions on Image Processing, 2018; TIP'22: IEEE Transactions on Image Processing, 2022; IGASS'22: International Geoscience and Remote Sensing Symposium, 2022; ICASSP'22: IEEE International Conference on Acoustics, Speech, and Signal Processing, 2022; BMVC'23: British Machine Vision Conference, 2023.
[0068] Table 2 Experimental Results Based on Transformer and Mamba
[0069] ↑: An increase in the value of this metric represents an improvement in performance.
[0070] U-Trans: U-Transformer (a U-shaped network based on Transformer); NU2Net: Nested U-Net (a nested U-shaped network); X-CAUNET: X-CAU Network (an improved U-shaped network); UWMamba: Underwater Mamba (an underwater Mamba network); TIP'23: IEEE Transactions on Image Processing, 2023. AAAI'23: Association for the Advancement of Artificial Intelligence, 2023. ICASSP'24: IEEE International Conference on Acoustics, Speech, and Signal Processing, 2024. SPL'24: IEEE Signal Processing Letters, 2024.
[0071] As shown in Table 1 and Table 2, the experimental results indicate that the adaptive underwater image enhancement method based on Retinex theory and Mamba provided by this application has achieved remarkable results in multiple evaluation metrics. At the same time, this method is superior to other methods in terms of texture details and contrast enhancement, fully demonstrating the superiority of this method in enhancing underwater scenes.
[0072] Color Channel Comparison Results The performance of the underwater image after enhancement on the RGB channels shows a significant improvement: Before enhancement, due to the selective light attenuation of water, the responses of the G / B channels were too strong and the R channel was significantly attenuated. As Figures 8 - 11 shown, the histogram distribution is concentrated and there is color bias; the result after enhancement, as Figures 12 - 15 shown, compensates for the channel attenuation difference, the intensity of the R channel is significantly increased, the distribution of the G / B channels is more uniform, the dynamic range of the overall histogram is expanded, and the contrast is improved. Compared with the label map (true value) Figures 16 - 19 shown, the enhanced image is roughly close to the label map in terms of channel intensity distribution, and the channel situation is closer to the color balance of the real scene.
[0073] Figure 11 、 Figure 15 and Figure 19 The so-called middle row refers to Figure 8 、 Figure 12 and Figure 16 the rows at the middle position of the horizontal pictures, and the sampling positions of the three RGB intensity curve graphs are the same.
[0074] Verified through systematic experiments, the method proposed in this study shows significant effectiveness, robustness, and generalization ability in the underwater image enhancement task, providing a reliable technical path for visual restoration in complex underwater environments.
[0075] Performance Analysis of the Enhancement Method In the evaluation of computational efficiency, this study conducted a standardized test based on 100 standard test images with a resolution of 224×224. The quantitative analysis results shown in Table 3 indicate that the method provided in this application significantly outperforms the comparative methods with the lowest computational complexity of 15.20G FLOPs and a single-frame processing time of 9ms, providing important technical support for real-time underwater vision systems.
[0076] Table 3 Experimental Results of Running Efficiency
[0077] ↓: It indicates that a decrease in the value of this indicator represents an improvement in performance.
[0078] FLOPs(G): Floating-point operations per second (in billions), representing computational complexity; TIME(S): Time (in seconds), representing the processing time of a single-frame image; Params(M): Number of parameters (in millions), representing the number of parameters of the model.
[0079] UIEC2: Underwater Image Enhancement CNN-based using two Color Spaces; X-CAUNET: A cross-color channel attention mechanism based on an underwater image enhancement transformer.
[0080] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. An adaptive underwater image enhancement method based on Retinex theory and Mamba, characterized in that: include: Light information estimation subnetwork: The original underwater image and its corresponding illumination prior map are used as input, and the illumination factor map is obtained through feature concatenation, the first convolutional layer, the multi-scale channel aggregation module layer, and the second convolutional layer. The illumination estimation component of the underwater scene is extracted by the multi-scale channel aggregation module layer, and the illumination factor map is subjected to Hadamard product operation with the original image to obtain the basic enhanced image. Degraded restoration sub-network: The encoder extracts multi-scale degradation features of the base enhanced image step by step, restores the spatial resolution through the decoder jump connection, and enhances the illumination-texture correlation through the selective state space modeling of the WLMamba bottleneck module, finally generating high-quality underwater image enhancement results.
2. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 1, characterized in that: The multi-scale channel aggregation module layer includes adjusting the detail features initially extracted by the first convolution layer through the first upper branch, the first middle branch and the first lower branch, and then performing element-by-element multiplication on the results of the three branches to obtain the final illumination estimation component of the underwater scene; The first upper branch retains the original input features; The first intermediate branch is used for multi-scale aggregation on multiple channels; The first lower branch performs a pooling operation in the spatial direction to obtain a spatial descriptor representation, then learns the correlation between channels through the lower branch convolution layer, and generates a channel weight representation through the Sigmoid function.
3. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 2, characterized in that: The first intermediate branch evenly splits the input features into four parts according to the number of channels, and then implements a multi-scale downsampling strategy on the split input features through the feature channels of Channel 1 to 4. The specific operations are as follows: Channel 1: Use 8 times downsampling to extract global semantic features; Channel 2: Balances local details and context information through 4x downsampling; Channel 3: Perform a 2x downsampling operation to retain medium-scale features; Channel 4: Maintains original underwater resolution to preserve high-frequency details; A 3×3 depth convolution is performed on the features after downsampling of Channel 1~3, and the dimension is restored through upsampling operation. The features of the feature channels of Channel 1~4 are concatenated, and then an aggregation operation is performed through the first intermediate branch convolution layer to finally obtain multi-scale aggregated features.
4. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 1, characterized in that: The WLMamba bottleneck module adjusts the features it receives through the second upper branch, the second middle branch, and the second lower branch at the same time; The second upper branch retains the original input features; The second intermediate branch will input the feature channel First, by linearly increasing , is a predefined channel expansion factor, followed by depthwise convolution, followed by SiLU activation, SSM-2D layer, and layer normalization; The second lower branch combines the input features first through a channel attention module modulated by the light information estimation component and a cascaded spatial attention module, and then through SiLU activation.
5. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 4, characterized in that: The input of the channel attention module consists of an illumination estimation component and an intermediate feature map of image processing. First, the intermediate feature map is passed through a global average pooling layer to capture the global statistical features of each channel. Subsequently, the compressed channel feature vector is subjected to a nonlinear transformation through a fully connected layer group to learn the complex relationship between channels. The channel weight is then activated and outputted through a Sigmoid function. In the feature recalibration stage, the above channel weight is multiplied element-by-element with the channel corresponding to the illumination estimation component to achieve adaptive modulation of the feature map, and the final output is a weighted feature map.
6. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 5, characterized in that: The weighted feature map output by the channel attention module is used as the input of the spatial attention module. First, the input feature map is subjected to average pooling and maximum pooling operations respectively. Then, the average feature map and the maximum feature map are concatenated along the channel dimension to generate a composite feature that integrates global and local features. The composite feature is subjected to spatial information fusion and nonlinear modeling through a 3×3 convolution kernel. After being activated by a Sigmoid function, a spatial attention weight is generated. Finally, the spatial attention weight is element-wise multiplied with the illumination estimation component to realize illumination-spatial characteristic modulation, and the feature map after spatial attention optimization is obtained.
7. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 4, characterized in that: The SSM-2D layer is used to capture long-range dependencies in underwater images, which involves converting the input feature map into a one-dimensional sequence in four directions through a flattening operation, arranging the two-dimensional features into a one-dimensional sequence according to the scanning sequence of the directions, and applying a dynamically parameterized state-space model to the sequence in each direction. The output sequences in the four directions are restored to a two-dimensional feature map and fused through attention weighting.
8. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 1, characterized in that: The overall architecture of the degradation repair sub-network adopts a four-layer encoder-decoder structure, including cross-layer skip connections and WLMamba bottleneck modules; The encoders of layers 1-3 perform multi-scale feature fusion with the decoders of the corresponding layers through skip connections. The encoder features of the 4th layer are adjusted by the WLMamba bottleneck module and then passed to the decoder of the 4th layer.
9. The adaptive underwater image enhancement method based on Retinex theory and Mamba according to claim 8, characterized in that: The jump connection consists of a feature alignment operation and an adaptive fusion operation, wherein the feature alignment operation includes spatial alignment and channel alignment. The spatial alignment includes upsampling the encoder features to the decoder size through bilinear difference, and then adjusting the number of channels through a convolution kernel of size 1×1 to obtain the aligned features. ; The input of the adaptive fusion operation consists of the output features of the feature alignment operation and the decoder output features concatenated on the channel, setting two learnable parameters ,in , set The decoder feature of the layer is , obtain the adaptive fusion features ; The adaptive fusion features Enter the The decoder of the layer completes the adaptive fusion operation.
10. An adaptive underwater image enhancement system based on Retinex theory and Mamba, characterized in that: include: Light information estimation subnetwork module: It is used to take the original underwater image and its corresponding illumination prior map as input, and obtain the illumination factor map through feature splicing, the first convolution layer, the multi-scale channel aggregation module layer and the second convolution layer in sequence. The illumination estimation component of the underwater scene is extracted by the multi-scale channel aggregation module layer, and the illumination factor map is subjected to Hadamard product operation with the original image to obtain the basic enhanced image; Degraded restoration sub-network module: It is used to extract multi-scale degradation features of the basic enhanced image step by step through the encoder, restore the spatial resolution through the decoder jump connection, enhance the illumination-texture correlation through the WLMamba bottleneck module selective state space modeling, and finally generate high-quality underwater image enhancement results.
Citation Information
Patent Citations
Image enhancement method based on high-frequency and low-frequency adaptive fusion
CN119648588A
Image enhancement algorithm based on adaptive filtering
CN119741245A
Low-light image enhancement method and system
CN119379551A
Lightweight underwater image enhancement method based on adaptive feature fusion
CN119540109A
Underwater image enhancement method based on frequency domain analysis and visual Mama
CN119784598A
Cited By
Image restoration method, system and device based on prior features and linear scanning
CN120219245A
Image Restoration Method, System, and Device Based on Prior Features and Linear Scanning
CN120219245B
Dense multiplexing and jump connection underwater image enhancement method for multilayer color features
CN120495149A
Remote sensing image building extraction method and system based on visual Mama model
CN120599504A
Lightweight low-light image enhancement method based on illumination iterative adjustment
CN120634936A