Underwater image enhancement method based on relation-driven dynamic state propagation

By using a relation-driven state-space modeling method to dynamically adjust the image enhancement path, the problem of insufficient modeling flexibility and computational efficiency in existing technologies is solved, achieving efficient and structure-aware enhancement of underwater images, and improving image quality and consistency.

CN120976043APending Publication Date: 2025-11-18HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510902611.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing underwater image enhancement techniques struggle to balance modeling flexibility, computational efficiency, and semantic adaptability in complex underwater environments, resulting in a lack of global consistency and insufficient restoration of structural details in the image enhancement results.

Method used

We adopt a relation-driven state-space modeling method, which dynamically samples the Mamba branch and the input-dependent convolution branch, and combines a cross-feature fusion bridge module and a composite loss function to adaptively adjust the image enhancement path, thereby improving the modeling ability of key semantic regions and the modeling of global context and local low-frequency information.

Benefits of technology

It significantly improves image color reproduction, texture clarity, and overall perceived quality, enhancing image enhancement effects and making it suitable for practical deployment in complex underwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976043A_ABST
    Figure CN120976043A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater image enhancement method based on relation-driven state space modeling, belongs to the technical field of computer vision and image processing, and aims to solve the problems of color distortion, detail blurring and the like of an underwater image caused by water attenuation and scattering. Carrying out structure perception enhancement modeling; and image reconstruction and decoding output. Wherein the structure sensing module extracts spatial continuity information through an offset generation network, adaptively rearranges scanning paths, preferentially focuses on a semantic rich region and executes dynamic state propagation, so that the accuracy and interpretability of global modeling are improved; and meanwhile, a local convolution kernel is dynamically generated according to global channel statistical characteristics by inputting a dependent convolution branch, so that the adaptability to a background region is enhanced. In order to further improve the feature fusion effect, a cross-feature fusion bridge module is provided, multi-level features are guided and fused through bidirectional attention of a structural path and a semantic path, and detail information and context semantics are effectively integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and digital image processing technology, specifically relating to an underwater image enhancement method based on relation-driven dynamic state propagation. This enhancement method is used to improve the visual quality of underwater images and can be widely applied in related fields such as marine resource exploration, marine monitoring, and intelligent underwater vehicle vision systems. Background Technology

[0002] Underwater image enhancement (UIE) is an important preprocessing task in computer vision, aiming to improve image quality degradation caused by the propagation characteristics of water. Light propagation in the underwater environment is affected by various factors such as absorption, scattering, and suspended particles, resulting in common problems in acquired images, including color distortion, decreased contrast, uneven brightness, and blurred details. These degradation effects severely impact the reliability of subsequent tasks such as underwater target detection, recognition, and navigation, becoming one of the key bottlenecks restricting the development of marine intelligent systems.

[0003] Existing underwater image enhancement techniques mainly include traditional methods based on physical priors and data-driven methods based on deep learning. The former typically restores the real image by modeling the physical attenuation model during underwater imaging and combining it with image restoration or compensation algorithms, such as algorithms based on dark channels, illumination estimation, or color compensation. These methods have the advantage of strong interpretability, but suffer from high dependence on assumptions about the underwater environment and poor generalization ability, making them difficult to adapt to diverse and complex real-world scenarios. In recent years, with the development of deep learning technology, more and more research has attempted to use end-to-end neural network methods for underwater image enhancement. Typical methods are based on convolutional neural networks (CNNs), which can automatically learn degradation patterns and restoration maps from a large number of samples, significantly improving the enhancement effect. However, due to the limited receptive field of CNNs, it is difficult to capture the dependencies between distant regions, and the enhancement results often lack global consistency. The Transformer architecture, due to its excellent long-range modeling capabilities due to its self-attention mechanism, has been gradually introduced into low-level vision tasks. However, the computational complexity of Transformers increases quadratically with image resolution, making it difficult to balance performance and efficiency when processing high-resolution underwater images, limiting its applicability to real-time deployment and edge computing scenarios.

[0004] To achieve a balance between modeling capability and computational efficiency, State Space Models (SSMs) have emerged as an emerging alternative. The representative architecture, Mamba, achieves long-distance modeling with linear complexity by decoupling state propagation from feature interaction, demonstrating good performance across multiple vision tasks. However, existing Mamba architectures propagate state information in a fixed linear scan order, failing to flexibly adjust the propagation path based on image content. This makes them ill-suited to the significant spatial heterogeneity of underwater images, such as structurally rich target regions and widely distributed but semantically sparse background regions. Furthermore, while multi-directional scanning strategies can improve spatial modeling capabilities, they still suffer from redundant computation and insufficient structural semantic modeling.

[0005] Therefore, there is an urgent need for an underwater image enhancement method that can balance modeling flexibility, computational efficiency, and semantic adaptability, in order to more effectively recover color and structural information in images, improve visual quality, and meet the deployment requirements in practical applications. Summary of the Invention

[0006] This invention provides an underwater image enhancement method based on relation-driven state-space modeling, which improves color reproduction, detail restoration, and structural clarity of images in complex underwater environments. The method is applicable to image enhancement tasks, especially image preprocessing in underwater scenes. The method mainly includes the following steps: Step 1, Shallow Feature Extraction and Initial Encoding: Convolutional embedding is performed on the input underwater image to extract basic image features and obtain a latent representation, which serves as the input for subsequent enhancement modeling.

[0007] Step 2, Structure-Aware Enhancement Modeling: A structure-aware enhancement module is introduced to perform dual-branch modeling of the features. One branch uses dynamically sampled Mamba state space modeling to adapt to semantically key regions; the other branch uses input-dependent dynamic convolution to supplement the response to low-frequency background. The outputs of the two branches are integrated through a fusion module to construct enhanced features.

[0008] Step 3: Image Reconstruction and Decoding Output: The enhanced features are fed into the decoding module to reconstruct the image, and the enhanced image result is output.

[0009] The following are the specific implementation steps.

[0010] Step 1: Feature Extraction The input underwater degraded image is acquired, and shallow feature extraction is performed on the original image using a convolutional embedding module to obtain a dimension of [dimensionality missing]. The image feature map is used as the basic input for subsequent state modeling.

[0011] Step 2: Structure-Aware Feature Enhancement The extracted feature map is divided into two parts along the channel using the structure-aware enhancement module, and then processed in a dual-branch manner, including a dynamic sampling Mamba branch and an input-dependent convolution branch.

[0012] Specifically, in the dynamic sampling Mamba branch, based on the spatial structural similarity of the input features, a deformable structure offset generation network is used to adaptively determine the sampling coordinate offset for each spatial location. The offset generation network consists of two main components: depthwise convolution and pointwise convolution. Depthwise convolution captures spatial structural changes within local regions of the input features, while pointwise convolution integrates local structural information with channel dimension information. The implementation is as follows: given an input feature map... The offset calculation formula is defined as follows:

[0013] In the formula, This represents a 3×3 depthwise convolution. This represents a 1×1 pointwise convolution. This represents the number of sampling neighbors for each spatial location. This indicates the number of groups into which the features are divided. Through this convolutional structure, the offset network can automatically learn regional features related to feature continuity, texture boundaries, and local structure, thereby generating a spatially sensitive offset.

[0014] After obtaining the spatial position offset, the spatial position The dynamic sampling coordinates are defined as follows:

[0015] in, Indicates the basic spatial location. Indicates the first Group 1 Each sampling point has a coordinate. For the calculated non-integer coordinates, bilinear interpolation is further used to accurately estimate the feature values ​​of the sampling points, ensuring the accuracy and smoothness of the spatial sampling process.

[0016] Next, the module counts the frequency of sampling for each spatial location within each group, using the following formula:

[0017] In the formula, For indicator functions, when position In the A value of 1 is assigned if the sampled location appears at any of the sampled locations, and a value of 0 otherwise. This sampling frequency statistical mechanism is used to measure the importance or significance of each location in the propagation of spatial structures.

[0018] Furthermore, this module divides the input feature map into multiple groups. Within each group, the sampling frequencies obtained from the above statistics are sorted to form a dynamic feature sequence based on structural saliency:

[0019] This dynamic sorting enables the starting position and path of state propagation to adaptively prioritize structurally and semantically significant regions, thereby significantly improving the effectiveness of subsequent state updates.

[0020] Finally, this module will sort the above-mentioned structure-aware dynamic sequences. As input, it is fed into the Mamba state-space model for an iterative state update process. The specific update formula is:

[0021] in, For a moment The hidden state, and The parameters are updated for the state space. Through the above method, this invention realizes a dynamic state update path based on structural relationships, which can significantly improve the accuracy and efficiency of spatial feature propagation.

[0022] On another input-dependent convolution branch, for large background areas (such as homogeneous water bodies) in underwater images, content-dependent dynamic convolution is used to achieve content-sensitive responses for spatial local regions.

[0023] Specifically, for a given input feature map First, Averaging pooling is performed to extract channel-level statistical features. Then, the pooling results are sequentially input into two layers. The convolution operator is used to generate a G-group attention matrix, where :

[0024] Next, the attention matrix is ​​reshaped and normalized using the Softmax function to obtain the attention distribution for each group:

[0025] Ultimately, the module learns by using learnable parameters. Perform element-wise multiplication to generate the final dynamic convolution kernel for each channel:

[0026] Step 3: The image reconstruction module aggregates shallow and deep features and restores them through the reconstruction module.

[0027] Furthermore, to achieve more accurate semantic guidance and detail preservation, this invention introduces a cross-feature fusion bridge module for content-aware fusion of multi-scale features between the encoder and decoder.

[0028] This module integrates high-resolution encoder features (structural details) with low-resolution decoder features (semantic abstraction). The specific integration process is as follows: encoder feature map Local window self-attention modeling is performed to preserve local detail structure. Its attention mechanism takes the following form:

[0029] in, These are the query, key, and value generated through linear transformation, respectively. This is the scaling factor.

[0030] Decoder feature map after adjusting the number of channels The encoder features are then used as the key and the query features are used to construct a heteroscale cross-attention mechanism:

[0031] This operation enables the encoder features to actively absorb semantic information, thereby achieving semantic guidance in the decoding path.

[0032] Finally, the attention features output from the two paths are adaptively fused, introducing a learnable gating mechanism along the channel dimension:

[0033]

[0034] Here, ⊙ represents element-wise multiplication. For upsampling operation, This is a channel-space joint attention map used to dynamically adjust the fusion ratio of structural and semantic information.

[0035] Furthermore, to improve the content consistency, structural restoration, and detail preservation capabilities of underwater image enhancement, this invention proposes a composite loss function design method. The loss function consists of three parts: pixel-level reconstruction loss, structural similarity loss, and edge preservation loss, with the following combined optimization objective: (1) Pixel-level reconstruction loss This loss term measures the pixel-level difference between the enhanced and original images to ensure consistent image content restoration and suppress over-enhancement. Its mathematical expression is:

[0036] (2) Structural similarity loss This loss function calculates the similarity between the enhanced image and the original image in terms of brightness, contrast, and structural layout based on the Structural Similarity Index (SSIM), thereby improving the overall structural consistency of the image. Its expression is:

[0037] (3) Edge preservation loss To enhance the recovery of image edge texture and high-frequency information, a Laplacian-based edge loss function is introduced. This loss metric improves the edge image difference between the original and the current image, and its calculation formula is as follows:

[0038] Here, Edges(•) represents the edge extraction operation based on the Laplacian operator, and ϵ is a small constant to prevent gradient explosion or division by zero errors.

[0039] (4) Total loss function The final optimization objective is a weighted combination of the three factors, expressed as:

[0040] The weight parameters are empirically set as follows: λ1=8, λ2=1, λ3=4 to balance pixel fidelity and structural edge preservation.

[0041] This invention also proposes an underwater image enhancement system for implementing an underwater image enhancement method based on relation-driven state-space modeling, the system comprising: The shallow feature extraction module receives the input underwater image and maps the original image into a feature map of dimension B×C×H×W through convolution operation, generating the potential representation required for subsequent enhancement modeling. The structure-aware enhancement module is used to perform dual-branch modeling on the feature image obtained by the shallow feature extraction module. The dual branches include a dynamic sampling state propagation branch and an input-dependent convolution branch.

[0042] The image reconstruction module is used to reconstruct the enhanced image from the enhanced feature map through a decoding operation. A cross-feature fusion module is used to integrate multi-scale features from the encoder and decoder during the image reconstruction stage.

[0043] The present invention also proposes a computer device, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the method of any one of claims 1 to 6.

[0044] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the method described in any one of claims 1 to 6.

[0045] Compared with the prior art, the present invention has the following beneficial technical effects: The relation-driven state propagation mechanism proposed in this invention can adaptively adjust the scanning order of the Mamba model, realize the priority propagation of state information from structurally significant regions, and significantly enhance the modeling ability of key semantic regions in images. (1) The dual-branch structure perception enhancement module designed in this invention can simultaneously model global context and local low-frequency information, improving the limitations of existing methods in structural detail recovery; (2) The cross-feature fusion bridge module adopts a dual-path attention mechanism and learnable channel weights, which effectively enhances the semantic and spatial consistency between features at different levels; (3) The overall method shows better image enhancement on multiple real and synthetic underwater image datasets, and can significantly improve the color reproduction, texture clarity and overall perceived quality of the image; (4) The structure of this invention can be modularly integrated into existing neural network architectures, and has good scalability and practical deployment value.

[0046] In summary, this invention provides an efficient, structure-aware image enhancement scheme suitable for complex underwater environments, with good engineering feasibility and promising prospects for widespread application. Attached Figure Description

[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0048] Figure 1 This is a flowchart of the underwater image enhancement method of the present invention.

[0049] Figure 2 This is a branch structure diagram of the dynamic sampling state propagation for underwater image enhancement according to the present invention.

[0050] Figure 3 This is a diagram of the input-dependent convolutional branch structure for underwater image enhancement according to the present invention.

[0051] Figure 4 This is a diagram of the cross-feature fusion bridge structure for underwater image enhancement according to the present invention.

[0052] Figure 5 This is the degraded input image to be processed in the underwater image enhancement method of the present invention.

[0053] Figure 6This is a processed result image from the underwater image enhancement method of the present invention. Detailed Implementation

[0054] The following combination Figures 1 to 3 The specific embodiments of the present invention will be described in detail below.

[0055] The following description and accompanying drawings are intended to clearly and completely illustrate implementations of the present invention, enabling those skilled in the art to practice the invention accordingly. Technical features in some embodiments may be combined in other embodiments unless explicitly contradictory. The terms "first," "second," etc., used herein are only for distinguishing elements of the same kind and do not represent a specific order. The terms "comprising," "including," etc., should be understood as open-ended limitations, meaning that they include the stated elements or features but do not exclude the presence of other unlisted technical elements.

[0056] like Figure 1 As shown, this embodiment provides an underwater image enhancement method based on relation-driven state-space modeling, including the following steps: Step 1: Shallow Feature Extraction The input underwater degraded image is fed into a convolutional embedding module, which extracts low-level image feature information through a series of convolutional layers, obtaining a dimension of [missing information]. The shallow feature map is used as the basic input for subsequent modeling.

[0057] Step 2: Structure-Aware Feature Enhancement This step utilizes the structure-aware enhancement module to perform channel segmentation on the shallow feature map, forming two branches: (1) Dynamic sampling state propagation branch The core of this branch lies in using a deformable sampling strategy to determine the sampling offset of each spatial location based on the local structural similarity of the image, thereby achieving priority state propagation for structurally significant regions.

[0058] Specifically, it includes the following sub-steps: 1. Given an input feature map Through a 3x3 The offset calculation formula for depthwise convolution is defined as follows:

[0059] In the formula, This represents a 3×3 depthwise convolution. This represents a 1×1 pointwise convolution. This represents the number of sampling neighbors for each spatial location. This indicates the number of groups into which the features are divided. Through this convolutional structure, the offset network can automatically learn regional features related to feature continuity, texture boundaries, and local structure, thereby generating a spatially sensitive offset.

[0060] 2. After obtaining the spatial position offset, the spatial position... The dynamic sampling coordinates are defined as follows:

[0061] in, Indicates the basic spatial location. Indicates the first Group 1 Each sampling point has a coordinate. For the calculated non-integer coordinates, bilinear interpolation is further used to accurately estimate the feature values ​​of the sampling points, ensuring the accuracy and smoothness of the spatial sampling process.

[0062] 3. Next, the module counts the frequency of sampling for each spatial location within each group, using the following formula:

[0063] In the formula, For indicator functions, when position In the A value of 1 is assigned if the sampled location appears at any of the sampled locations, and a value of 0 otherwise. This sampling frequency statistical mechanism is used to measure the importance or significance of each location in the propagation of spatial structures.

[0064] 4. Furthermore, this module divides the input feature map into multiple groups. Within each group, the sampling frequencies obtained from the above statistics are sorted to form a dynamic feature sequence based on structural saliency:

[0065] This dynamic sorting enables the starting position and path of state propagation to adaptively prioritize structurally and semantically significant regions, thereby significantly improving the effectiveness of subsequent state updates.

[0066] 5. Finally, this module will sort the above-mentioned structure-aware dynamic sequences. As input, it is fed into the Mamba state-space model for an iterative state update process. The specific update formula is:

[0067] in, For a moment The hidden state, and The parameters are updated for the state space. Through the above method, this invention realizes a dynamic state update path based on structural relationships, which can significantly improve the accuracy and efficiency of spatial feature propagation.

[0068] (2) Input-dependent convolutional branch This branch is used to locally enhance low-texture backgrounds (such as water bodies) that occupy a large area of ​​the image, thereby improving content adaptability.

[0069] Specifically, it includes the following sub-steps: 1. For a given input feature map First, Averaging pooling is performed to extract channel-level statistical features. Then, the pooling results are sequentially input into two layers. The convolution operator is used to generate a G-group attention matrix, where :

[0070] 2. Next, the attention matrix is ​​reshaped and normalized using the Softmax function to obtain the attention distribution for each group:

[0071] 3. Finally, the module works by comparing with learnable parameters. Perform element-wise multiplication to generate the final dynamic convolution kernel for each channel:

[0072] Step 3: The image reconstruction module aggregates shallow and deep features and restores them through the reconstruction module.

[0073] Furthermore, to achieve more accurate semantic guidance and detail preservation, this invention introduces a cross-feature fusion bridge module for content-aware fusion of multi-scale features between the encoder and decoder.

[0074] This module integrates high-resolution encoder features (structural details) with low-resolution decoder features (semantic abstraction). The specific integration process is as follows: encoder feature map Local window self-attention modeling is performed to preserve local detail structure. Its attention mechanism takes the following form:

[0075] in, These are the query, key, and value generated through linear transformation, respectively. This is the scaling factor.

[0076] Decoder feature map after adjusting the number of channels The encoder features are then used as the key and the query features are used to construct a heteroscale cross-attention mechanism:

[0077] This operation enables the encoder features to actively absorb semantic information, thereby achieving semantic guidance in the decoding path.

[0078] Finally, the attention features output from the two paths are adaptively fused, introducing a learnable gating mechanism along the channel dimension:

[0079]

[0080] Here, ⊙ represents element-wise multiplication. For upsampling operation, This is a channel-space joint attention map used to dynamically adjust the fusion ratio of structural and semantic information.

[0081] Furthermore, to improve the content consistency, structural restoration, and detail preservation capabilities of underwater image enhancement, this invention proposes a composite loss function design method. The loss function consists of three parts: pixel-level reconstruction loss, structural similarity loss, and edge preservation loss, with the following combined optimization objective: (1) Pixel-level reconstruction loss This loss term measures the pixel-level difference between the enhanced and original images to ensure consistent image content restoration and suppress over-enhancement. Its mathematical expression is:

[0082] (2) Structural similarity loss This loss function calculates the similarity between the enhanced image and the original image in terms of brightness, contrast, and structural layout based on the Structural Similarity Index (SSIM), thereby improving the overall structural consistency of the image. Its expression is:

[0083] (3) Edge preservation loss To enhance the recovery of image edge texture and high-frequency information, a Laplacian-based edge loss function is introduced. This loss metric improves the edge image difference between the original and the current image, and its calculation formula is as follows:

[0084] Here, Edges(•) represents the edge extraction operation based on the Laplacian operator, and ϵ is a small constant to prevent gradient explosion or division by zero errors.

[0085] (4) Total loss function The final optimization objective is a weighted combination of the three factors, expressed as:

[0086] The weight parameters are empirically set as follows: λ1=8, λ2=1, λ3=4 to balance pixel fidelity and structural edge preservation.

[0087] Effect verification To verify the effectiveness of the algorithm proposed in this invention, comparisons were made on multiple underwater datasets in two aspects: first, a comparison was made with the existing best algorithm on full-reference metrics; second, a comparison was made with the existing best algorithm on non-full-reference metrics.

[0088] Verification 1: Full Reference Indicators Data sets and metrics The model was trained on UIEB and LSUI, with training data consisting of 4130 underwater degradation images for each platform. During training, the input images were cropped to... Size. Testing is divided into full-reference metrics and non-full-reference metrics.

[0089] Results Analysis Multiple publicly available test datasets of natural and synthetic images were selected, including UIEB, LSUI, and EUVP datasets with reference images, and U45, C60, and Color7 datasets without reference images. The results were compared with several existing underwater image enhancement algorithms, specifically U-Trans, Semi-UIR, UIE-DM, CECF, HCLR, and WMamba methods. Detailed results with reference metrics are shown in Table 1.

[0090] As shown in Table 1, compared with the current state-of-the-art underwater image enhancement methods, the method proposed in this application achieves varying degrees of improvement in metrics (PSNR, SSIM, MSE) on multiple test datasets. Particularly on the UIEB dataset, the PSNR is improved by 1.03 dB compared to the current state-of-the-art method WMamba.

[0091] Table 1 compares objective results. Among the full-reference metrics PSNR, SSIM, and MSE, the method proposed in this invention achieves the best results on all publicly available test datasets compared to existing methods.

[0092] Table 1

[0093] To further verify the effectiveness of the algorithm of this invention, we also conducted tests on non-reference metrics, including underwater-specific metrics: Uciqe and Uiqm, as well as the natural visual metric Niqe used by the human eye. The results are shown in Table 2 below: As shown in Table 2, the algorithm of this invention remains highly competitive. It still ranks first in terms of mean performance for specific underwater indicators.

[0094] Table 2

[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.

Claims

1. An underwater image enhancement method based on relation-driven state-space modeling, characterized in that, Includes the following steps: Step 1, Shallow Feature Extraction and Initial Encoding: Convolutional embedding is performed on the input underwater image to extract basic image features and obtain a latent representation, which serves as the input for subsequent enhancement modeling; Step 2, Structure-Aware Enhancement Modeling: Introduce a structure-aware enhancement module to perform dual-branch processing on the image features obtained in Step 1, including a dynamic sampling state propagation branch and an input-dependent convolution branch; Dynamic sampling state propagation branches are used to adapt to semantic key regions; input-dependent dynamic convolution is used to supplement the response to low-frequency background; the dual-branch outputs are integrated through a fusion module to construct enhanced features; Step 3: Image Reconstruction and Decoding Output: The enhanced features are fed into the decoding module to reconstruct the image, and the enhanced image result is output.

2. The underwater image enhancement method based on relation-driven state-space modeling according to claim 1, characterized in that, In step one, the input underwater degraded image is acquired, and shallow feature extraction is performed on the original image using a convolutional embedding module to obtain a dimension of [missing information]. The image feature map is used as the basic input for subsequent state modeling.

3. The underwater image enhancement method based on relation-driven state-space modeling according to claim 1, characterized in that, In step two, the dynamic sampling state propagation branch includes: (1) Based on the spatial structural similarity of the input features, a deformable structural offset generation network is used to adaptively determine the sampling coordinate offset of each spatial location. The offset generation network consists of two main components: depthwise convolution and pointwise convolution. The depthwise convolution is used to capture the spatial structural change features in the local region of the input features, while the pointwise convolution is used to integrate local structural information and channel dimension information. Specifically, given the input feature map... The offset calculation formula is defined as follows: In the formula, This represents a 3×3 depthwise convolution. This represents a 1×1 pointwise convolution. This represents the number of sampling neighbors for each spatial location. Indicates the number of groups into which the feature is divided; (2) After obtaining the spatial position offset, calculate the sampling position and spatial position. The dynamic sampling coordinates are defined as follows: , in, Indicates the basic spatial location. Indicates the first Group 1 Each sampling position coordinate; for the calculated non-integer coordinate positions, a bilinear interpolation method is further used; (3) Statistical sampling frequency: The module counts the frequency of sampling for each spatial location within each group. The statistical formula is defined as: In the formula, For indicator functions, when position In the The value is 1 if it appears in any of the sampling positions, and 0 otherwise. (4) Sort by frequency: Divide the input feature map into multiple groups Within each group, the sampling frequencies obtained from the above statistics are sorted to form a dynamic feature sequence based on structural saliency: This dynamic sorting enables the starting position and path of state propagation to adaptively prioritize structurally and semantically significant regions, thereby improving the effectiveness of subsequent state updates. This module will process the above-mentioned sorted structure-aware dynamic sequence. As input, it is fed into the Mamba state-space model for an iterative state update process; the specific update formula is: in, For a moment The hidden state, and Update parameters for the state space.

4. The underwater image enhancement method based on relation-driven state-space modeling according to claim 1, characterized in that, The input-dependent convolutional branch generates dynamic convolutional kernels using the following method: (1) For a given input feature map First, Perform average pooling to extract channel-level statistical features; (2) Input the pooling results into the two layers in sequence. The convolution operator is used to generate the attention matrix grouped into G groups: (3) Reshape the attention matrix and normalize it using the Softmax function to obtain the attention distribution for each group: (4) The module uses learnable parameters Element-wise multiplication is performed to generate the final dynamic convolution kernel for each channel.

5. The underwater image enhancement method based on relation-driven state-space modeling according to claim 1, characterized in that, In step three, the image reconstruction module includes a cross-feature fusion bridge (CFB) module, whose fusion methods include: (1) Encoder feature map Local window self-attention modeling is performed to preserve local detailed structure; its attention mechanism takes the following form: in, These are the query, key, and value generated through linear transformation, respectively. This is the scaling factor; (2) The decoder feature map after adjusting the number of channels Using encoder features as keys and queries as queries, a heteroscale cross-attention mechanism is constructed: This operation enables encoder features to actively absorb semantic information, thereby achieving semantic guidance in the decoding path; (3) Adaptively fuse the attention features output from the two paths above, and introduce a learnable gating mechanism in the channel dimension: Here, ⊙ represents element-wise multiplication. For upsampling operation, This is a channel-space joint attention map used to dynamically adjust the fusion ratio of structural and semantic information.

6. An underwater image enhancement method based on relation-driven state-space modeling according to any one of claims 1 to 5, characterized in that, It also includes a composite loss function, which comprises pixel-level reconstruction loss, structural similarity loss, and edge preservation loss, with the following combined optimization objective: (1) Pixel-level reconstruction loss This loss is used to measure the difference between the enhanced image and the original image at the pixel level, in order to ensure the consistency of image content restoration and suppress over-enhancement. Its mathematical expression is: (2) Structural similarity loss This loss is based on the Structural Similarity Index (SSIM) to calculate the similarity between the enhanced image and the original image in terms of brightness, contrast, and structural layout, thereby improving the overall structural consistency of the image. Its expression is: (3) Edge preservation loss To enhance the recovery of image edge texture and high-frequency information, a Laplacian-based edge loss function is introduced. This loss metric improves the edge image difference between the original and the current image. The calculation formula is as follows: Where Edges(•) represents the edge extraction operation based on the Laplacian operator, and ϵ is a small constant to prevent gradient explosion or division by zero error; (4) Total loss function The final optimization objective is a weighted combination of the three factors, expressed as: The weight parameters are empirically set as follows: λ1=8, λ2=1, λ3=4 to balance pixel fidelity and structural edge preservation.

7. An underwater image enhancement system for implementing the underwater image enhancement method based on relation-driven state-space modeling as described in any one of claims 1 to 6, characterized in that, The system includes: The shallow feature extraction module receives the input underwater image and maps the original image into a feature map of dimension B×C×H×W through convolution operation, generating the potential representation required for subsequent enhancement modeling. The structure-aware enhancement module is used to perform bi-branch modeling on the feature images obtained by the shallow feature extraction module. The image reconstruction module is used to reconstruct the enhanced image from the enhanced feature map through a decoding operation. A cross-feature fusion module is used to integrate multi-scale features from the encoder and decoder during the image reconstruction stage.

8. The underwater image enhancement system according to claim 7, characterized in that: The structure-aware enhancement module includes: a dynamic sampling state propagation branch and an input-dependent convolution branch.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, The processor executes the computer program to implement the method described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it is used to implement the method described in any one of claims 1 to 6.

Citation Information

Cited By

  • Underwater image enhancement method based on wavelet Mama

    CN121353107A

  • An underwater image enhancement method based on wavelet Mamba

    CN121353107B

  • Real-time underwater image enhancement method based on attention fusion and histogram stretching

    CN121707834A