Underwater multi-color space image enhancement method based on depth information guidance
By combining the depth information generation module and the multicolor space feature fusion module with the depth enhancement module, the problems of spectral distortion and brightness degradation in underwater images were solved, achieving image detail restoration and color improvement, thus enhancing the quality of underwater images.
Patent Information
- Application Number
- CN202510943554.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
AI Technical Summary
Existing underwater image enhancement methods suffer from quality degradation phenomena such as spectral distortion, global contrast attenuation, and nonlinear brightness degradation when processing images acquired by underwater optical imaging systems, resulting in reduced image quality. In addition, deep learning-based methods cannot effectively enhance contrast, saturation, and brightness when processing in the RGB color space.
We adopt a depth-information-guided underwater multi-color space image enhancement method. Through a depth information generation module, a multi-color space feature fusion module, and a depth enhancement module, we utilize physical prior knowledge and multi-color information, combined with attention estimation and scene reconstruction training, to improve feature representation capabilities and model scene adaptability.
It effectively restores underwater image details, improves color difference and color cast issues, enhances image PSNR, SSIM and UCIQE performance, and improves image color fidelity and structural integrity, making it suitable for image processing in complex underwater environments.
Smart Images

Figure CN120852256A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for enhancing underwater multi-color space images based on depth information guidance. Background Technology
[0002] With the deepening of global marine resource development, underwater environmental research has become a key area of focus at the forefront of international science and technology. As a strategic resource for the sustainable development of the national economy, the in-depth exploration and scientific development of marine resources have significant economic value and practical implications. In the process of underwater environmental detection and investigation, underwater optical imaging, with its intuitiveness and rich information, has become a key means of acquiring information about complex underwater environments. With the innovation and breakthroughs in computer vision technology, underwater image enhancement technology has shown broad application prospects in many fields such as marine exploration, marine biological research, underwater archaeology, and intelligent underwater robots, providing strong technical support for a deeper understanding and precise analysis of the underwater world.
[0003] However, images acquired by underwater optical imaging systems generally suffer from quality degradation, including spectral distortion, global contrast reduction, and nonlinear brightness deterioration. The reason for this decline in underwater image quality lies in the color shift caused by the selective attenuation effect of the non-uniform spectral absorption characteristics of water on the visible light band. Simultaneously, discrete scattering by suspended particles in the water leads to disordered light field energy distribution and the formation of background noise fields, resulting in decreased image edge sharpness, blurred details, and a significant reduction in global contrast. Furthermore, the underwater environment, especially after leaving shallow water, severely weakens natural light conditions. Artificial light sources are needed to improve illumination to obtain clearer images. Since artificial light sources are similar to imaging equipment and are typically white light, illuminated areas exhibit higher brightness and less color distortion, while unilluminated areas tend to be dark. This leads to severe and uneven degradation of underwater images, further complicating the problem.
[0004] Existing underwater image enhancement methods mainly include physical model-less methods, physical model-based methods, and deep learning-based methods. Physical model-less methods, lacking a realistic underwater model, yield unsatisfactory results. Physical model-based methods, however, are limited in applicability and stability due to the complexity and dynamics of the underwater environment, and the established physical models are often ill-conditioned, meaning even small errors can lead to severe degradation. Most deep learning-based methods only process within the RGB color space, failing to reflect some crucial details of underwater images. For example, while the deep learning-based UIE method can solve color cast problems caused by dispersion, it cannot further enhance other image attributes (contrast, saturation, and brightness). Summary of the Invention
[0005] This invention proposes an underwater multi-color space image enhancement method based on depth information guidance, which improves feature representation capability and enhances the scene adaptability of the model.
[0006] The present invention adopts the following technical solution.
[0007] A depth-information-guided underwater multi-color space image enhancement method is proposed. The method's model is based on a training model framework that can utilize physical prior knowledge and multi-color information. It includes a depth information generation module, a multi-color space feature fusion module, and a depth enhancement module. By strengthening the connection between the depth image and the underwater scene through attention estimation, scene reconstruction training, and multi-color information fusion, the method can compensate for the insensitivity of the single RGB color space to image attributes such as brightness and saturation. This improves the feature representation capability and enhances the model's scene adaptability.
[0008] The depth information generation module takes the original underwater image as input and generates a monocular depth image through encoding-decoding and scaling transformation. At the same time, a depth supervision network is introduced during the training phase to generate a regression map, which jointly supervises the training and loss of the depth generation module, further promoting the rationality of the overall scene restoration.
[0009] The depth information generation module is based on an encoder and decoder structure, and its operation includes the following specific steps: Step A1, the depth information generation module uses a dual attention module structure that includes spatial information and channel information (such as...). Figure 2 As shown, a depth image is generated. In the channel attention part, an average pooling layer is used to reduce the feature size, and then spatial dependencies are captured by query vector q, key vector k, and value vector v. This process introduces a learnable parameter γ to dynamically adjust the learning intensity, expressed by the formula:
[0010] q = k = v = (AvgP(f)) Formula 1;
[0011] f s =γ×((q×k) soft ×v) Formula 2;
[0012] in,(·) soft It is the softmax function, γ is a learnable parameter used to adjust the intensity, and f s The weights are marked as input to the next stage;
[0013] Step A2: In the channel module, the module learns the interrelationships and importance between different channels through different pooling operations, providing a more comprehensive feature description;
[0014] f c =f s +f×(AvgP(f)clc +MaxP(f) clc ) sig Formula 3;
[0015] Where AvgP(·) and MaxP(·) represent the average pooling layer and the max pooling layer, respectively. sig f represents the sigmoid function. c The features of the labeled channel module are input into the weights of the next stage; subsequently, the two attention features are fused and residually connected with the original input to capture hidden features with significant regional degradation.
[0016] f h =f+f c ×(f s ) dc Formula 4;
[0017] Step A3: Use the decoder to generate the corresponding depth map, guide the network to perceive different degradation regions, introduce a deep supervision network during the training phase to generate regression mapping maps, and jointly supervise the training and loss of the deep generation module to further promote the rationality of the overall scene restoration.
[0018] The multi-color space feature fusion module uses a joint training strategy across color spaces. Specifically, it extracts complementary feature information from multiple color spaces, including color information, brightness, lightness, and saturation, and integrates them through a feature fusion strategy. This joint training strategy overcomes the problem of insufficient sensitivity of a single RGB space to image attributes such as brightness, saturation, and color fidelity, thereby enhancing the model's ability to perceive different color characteristics.
[0019] The multi-color space feature fusion module extracts and integrates complementary features from the RGB, LAB, and HSV color spaces during operation, including the following steps;
[0020] Step B1: Associate the three RGB channels using the A and B channels in the LAB space, so that the training model can learn the association between the three RGB channels in the underwater natural scene. When reconstructing the image, the constraints of the A / B channels on the color make the color fidelity of the reconstructed image higher.
[0021] Step B2: The H channel of the HSV space is used to represent the hue information of different regions of the image. Based on the threshold segmentation of hue, the blur noise introduced by the scattering body is removed, thereby enhancing the target information. By reconstructing the L channel (brightness), V channel (luminance), and S channel (saturation) of the image color space, the brightness and darkness changes of the reconstructed image are more natural (neither too dark nor too exposed) and the colors are richer.
[0022] Step B3: After extracting key information from the above images, the multi-color space feature fusion module integrates the information using a feature fusion strategy; its mathematical expression is as follows:
[0023]
[0024] Among them, F ori F represents the basic features extracted from the original RGB input. RGB ,F HSV and F LAB F represents the features extracted from the corresponding color space. fusion This is the feature fusion function.
[0025] Step B4: After feature fusion, a channel-adaptive attention mechanism is introduced to dynamically adjust the importance of features in each channel, making the network pay more attention to task-related information channels. The network structure is as follows: Figure 3 As shown, its mathematical expression can be represented as:
[0026]
[0027] Where CAAM(·) represents the channel attention function, This indicates element-wise multiplication.
[0028] In step B1, the original color information is preserved in the RGB space, where the three channels are independent of each other. For example, in underwater image reconstruction, if the red channel is not adequately compensated, a blue-green bias will appear.
[0029] The depth enhancement module is used to selectively guide the enhancement network by combining multi-color space information with the depth image obtained in the first part, making non-local features and local features more reasonable. The depth enhancement module also introduces surrogate attention, which further optimizes the feature fusion process through a more lightweight attention model, enabling the network to focus on key regions more efficiently.
[0030] The depth enhancement module is based on an encoder-decoder structure and includes the following steps during operation;
[0031] Step C1: The input underwater image and multidimensional color space information are first extracted stepwise through residual blocks RB and combined with residual skip connections to ensure effective information transmission;
[0032] Step C2: Set up a depth guidance module in the center of the network of the depth enhancement module, with the structure as follows: Figure 4 As shown, it accepts feature information and depth information from the encoder and consists of two feature processing paths: the upper path focuses on high-level semantic feature extraction, while the lower path focuses on multi-scale feature learning and integration, achieving effective processing of feature information, expressed by the formula:
[0033] f ck =d k ×f ck ×β k k = 1, 3, 5 (Formula 7)
[0034] f o =f r +cat(f c1 ,f c3 ,f c5 ) clc Formula 8;
[0035] Where cat() represents a cascading operation along the channel dimension, () clc This represents a convolution-activation-convolution structure, where β represents the feature fusion node;
[0036] Step C3: In the upper-layer path design, a proxy attention mechanism is introduced. This mechanism is an innovative attention computation paradigm, its design concept originating from the proxy concept in reinforcement learning. It dynamically adjusts the feature weight distribution through adaptive learning. Unlike traditional attention mechanisms, it introduces a more flexible context-aware capability, automatically adjusting the focus of attention for different tasks and input features.
[0037] Specifically, it includes four parts: feature transformation projection, attention score calculation, agent decision adjustment, feature aggregation and output, and its mathematical expression is as follows:
[0038] Q = W Q X,K=W K X,V=W V Formula 9;
[0039] Where X is the input feature, W Q W K W V This is the weight matrix for learning.
[0040]
[0041] Where, d k The scaling factor is the square root of the feature dimension.
[0042] A agent =A agent (A raw C) Formula 11;
[0043] Where f agent Let C be the proxy decision function, and C be the context feature matrix used to guide the attention allocation strategy. The proxy decision function is implemented in the following way:
[0044] fagent (A,C)=σ(W a A+W c C+b) Formula 12;
[0045] Where σ is the activation function, W a W c Here is the weight matrix, and b is the bias term;
[0046] Y = softmax(A agent Formula 13;
[0047] The final output feature Y is obtained by weighting and aggregating the value feature V using attention scores normalized by softmax, thereby achieving stronger feature extraction capabilities.
[0048] Step C4: In the design of the lower-level path, a multi-kernel convolution structure is adopted to obtain more perceptual fields, thereby enhancing the details of different sizes in the multi-scale encoded features. The β1, β2 and β3 nodes in the model network represent feature fusion points.
[0049] At these locations, features from different processing paths are integrated through a carefully designed fusion strategy, achieving an effective combination of multi-scale and multi-level features and significantly improving the model's representational capabilities.
[0050] Step C5: Using the feature results obtained in the previous two steps, and combining them with depth information, the underwater image is enhanced in a guided manner. The original size of the image is gradually restored through upsampling, and residual blocks (RB) are introduced to achieve progressively fine reconstruction of features, and finally the enhanced result is output.
[0051] The model training strategy of the method is as follows:
[0052] The model training process is supervised in two stages. The first stage calculates the loss of the deep generation module and the deep enhancement module respectively, which is the offline pre-training process. The second stage uses the deep generation module with fixed parameters to assist the deep enhancement module in generating enhanced images, thereby obtaining the corresponding loss.
[0053] The model training uses the EUVP and UIEB datasets, which are commonly used and publicly available for underwater image augmentation tasks. The first stage of training involves generating depth images of corresponding images using a deep image generation module. Image pairs are randomly selected from the public datasets to train the deep image generation module and the supervision module for 100 epochs, and the L1 loss is calculated.
[0054]
[0055] Where d and d GTThese represent the estimated depth map and the depth map of the ground truth GT, respectively. X and d represent the regression image and the underwater image, respectively; the hyperparameter λ1 = 3 is set to balance the loss; by incorporating the regression process into the training, d and X will be more tightly structured.
[0056] In the second stage, image pairs are randomly selected from the UIEB and EUVP datasets for training. To ensure that the augmented images have satisfactory color and detail, Charbonnier loss and SSIM loss are used to determine the total loss rate of the depth augmentation module; the former measures pixel similarity, and the latter measures structural similarity; the formula is as follows:
[0057]
[0058] Where x and y represent the enhanced and sharpened images, respectively, and λ² and ε are the hyperparameters of the balancing loss, set to 0.5 and 10, respectively. -3 .
[0059] The raw underwater images input to the depth information generation module are based on two public datasets, UIEB and EUVP. The underwater scenes include both natural and artificial lighting, with different conditions such as insufficient lighting (overall dark) and excessive lighting (local overexposure). In terms of water turbidity, the range covers water bodies from clear to severely turbid. Suspended particles in turbid water can cause light scattering, resulting in a fogging effect in the image and reducing contrast and clarity.
[0060] In terms of underwater objects, including marine life and underwater plants, these objects often suffer from problems such as blurred outlines, distorted colors, and difficulty in distinguishing details under complex lighting and turbid conditions.
[0061] During model training, all images are resized to a resolution of 256×256 before processing to ensure consistency of the input data.
[0062] This invention utilizes prior physical knowledge and multicolor information. Through attention estimation, scene reconstruction training, and the fusion of multicolor information, this framework effectively strengthens the connection between depth maps and underwater scenes, compensating for the insensitivity of the single RGB color space to image attributes such as brightness and saturation. It improves feature representation capabilities and enhances the model's scene adaptability.
[0063] like Figure 5 , 6 7, 8 and Figure 9As shown in Table 1, compared to existing underwater image enhancement methods, this invention combines depth information and multicolor space information, which not only effectively recovers image details visually but also effectively improves color difference and color cast issues. Objectively, the proposed method demonstrates superior performance in PSNR, SSIM, and UCIQE metrics, surpassing state-of-the-art methods. Specifically, PSNR primarily evaluates noise level and detail preservation; the multicolor space design of this invention effectively preserves color information, brightness, and saturation. Furthermore, the dual attention mechanism successfully simulates spatial correlation and enhances discriminative features, thereby improving signal quality. SSIM evaluates image brightness, contrast, and structural fidelity; the method of this invention utilizes complementary features of the three color spaces to preserve image information while maintaining scene spatial structure through deep integration. With the help of surrogate attention, key enhancement regions are identified to ensure structural integrity. UCIQE focuses on evaluating the balance between color, saturation, and contrast in the image; a higher UCIQ value indicates that our method achieves a more balanced relationship. Despite its slightly lower UIQM score, it remains competitive because UIQM emphasizes color richness, sharpness, and contrast, with a particular focus on color saturation.
[0064] The framework of this invention prioritizes natural color reproduction and scene realism over excessive saturation or contrast enhancement, which aligns with the invention's objective: perceptual natural enhancement, not merely numerical optimization. Furthermore, while achieving good enhancement results, this invention also demonstrates significant value in other related fields, such as semantic segmentation, where the enhancement achieved achieves better segmentation results compared to the same segmentation algorithm. In summary, this invention can effectively enhance underwater images while maintaining good performance in other tasks (such as semantic segmentation). Future development could consider combining it with semantic segmentation modules to achieve joint enhancement and habitat classification, potentially advancing automated solutions for large-scale marine resource monitoring.
[0065] This invention, as an image processing algorithm, can be combined with existing equipment and deployed on its computing backend to perform secondary processing and enhancement on images. Specifically, it can be used in the following devices or venues:
[0066] Underwater exploration equipment: Applicable to marine resource exploration equipment, effectively improving image clarity in turbid waters, enhancing underwater target identification capabilities, and significantly improving exploration efficiency and accuracy.
[0067] Underwater robots: can be integrated into various underwater robot systems to achieve real-time image quality optimization in complex underwater environments, ensuring the reliable execution of underwater operations, inspections, and maintenance tasks.
[0068] Marine scientific research equipment: can be used for marine scientific research and exploration, overcoming problems such as insufficient underwater lighting and scattering interference, and providing high-quality image support for marine ecological surveys, seabed topography surveys, etc.
[0069] Underwater monitoring system: Applicable to offshore facility monitoring systems, providing clear monitoring images under various water quality conditions, assisting in the safety inspection and fault early warning of offshore facilities.
[0070] Underwater archaeology equipment: Equipment that can be used for underwater cultural relic archaeology, restoring the original appearance of underwater cultural relics through intelligent image enhancement technology, thereby improving the efficiency and quality of archaeological work.
[0071] Port security facilities: can be deployed in the port's underwater security monitoring system to ensure clear imaging and real-time monitoring of underwater targets around the clock, thereby enhancing the port's security capabilities. Attached Figure Description
[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0073] Appendix Figure 1 This is a schematic diagram of the overall principle framework of the present invention;
[0074] Appendix Figure 2 This is a schematic diagram of a dual-channel attention module;
[0075] Appendix Figure 3 This is a schematic diagram of the multi-color space feature fusion module;
[0076] Appendix Figure 4 This is a schematic diagram of the deep guidance module of the present invention;
[0077] Appendix Figure 5 This is a schematic diagram comparing the present invention with other underwater enhancement methods (UIEB);
[0078] Appendix Figure 6 This is a schematic diagram comparing the pixel values of the red channel with other underwater image enhancement methods of the present invention;
[0079] Appendix Figure 7 This is a schematic diagram comparing the distribution of other underwater image enhancement methods of the present invention in RGB color (UIEB);
[0080] Appendix Figure 8 This is a schematic diagram comparing the present invention with other underwater enhancement methods (EUVP);
[0081] Appendix Figure 9 This is a schematic representation of the comparison data between the present invention and other underwater image enhancement methods. Detailed Implementation
[0082] like Figure 1As shown, an underwater multi-color space image enhancement method based on depth information is proposed. The model of the method is based on a training model framework that can utilize physical prior knowledge and multi-color information. It includes a depth information generation module, a multi-color space feature fusion module, and a depth enhancement module. By strengthening the connection between the depth image and the underwater scene through attention estimation, scene reconstruction training, and multi-color information fusion, the method can compensate for the insensitivity of the single RGB color space to image attributes such as brightness and saturation, improve feature representation capabilities, and enhance the scene adaptability of the model.
[0083] The depth information generation module takes the original underwater image as input and generates a monocular depth image through encoding-decoding and scaling transformation. At the same time, a depth supervision network is introduced during the training phase to generate a regression map, which jointly supervises the training and loss of the depth generation module, further promoting the rationality of the overall scene restoration.
[0084] The depth information generation module is based on an encoder and decoder structure, and its operation includes the following specific steps: Step A1, the depth information generation module uses a dual attention module structure that includes spatial information and channel information (such as...). Figure 2 As shown, a depth image is generated. In the channel attention part, an average pooling layer is used to reduce the feature size, and then spatial dependencies are captured by query vector q, key vector k, and value vector v. This process introduces a learnable parameter γ to dynamically adjust the learning intensity, expressed by the formula:
[0085] q = k = v = (AvgP(f)) Formula 1;
[0086] f s =γ×((q×k) soft ×v) Formula 2;
[0087] in,(·) soft It is the softmax function, γ is a learnable parameter used to adjust the intensity, and f s The weights are marked as input to the next stage;
[0088] Step A2: In the channel module, the module learns the interrelationships and importance between different channels through different pooling operations, providing a more comprehensive feature description;
[0089] f c =f s +f×(AvgP(f) clc +MaxP(f) clc ) sig Formula 3;
[0090] Where AvgP(·) and MaxP(·) represent the average pooling layer and the max pooling layer, respectively. sigf represents the sigmoid function. c The features of the labeled channel module are input into the weights of the next stage; subsequently, the two attention features are fused and residually connected with the original input to capture hidden features with significant regional degradation.
[0091] f h =f+f c ×(f s ) dc Formula 4;
[0092] Step A3: Use the decoder to generate the corresponding depth map, guide the network to perceive different degradation regions, introduce a deep supervision network during the training phase to generate regression mapping maps, and jointly supervise the training and loss of the deep generation module to further promote the rationality of the overall scene restoration.
[0093] The multi-color space feature fusion module uses a joint training strategy across color spaces. Specifically, it extracts complementary feature information from multiple color spaces, including color information, brightness, lightness, and saturation, and integrates them through a feature fusion strategy. This joint training strategy overcomes the problem of insufficient sensitivity of a single RGB space to image attributes such as brightness, saturation, and color fidelity, thereby enhancing the model's ability to perceive different color characteristics.
[0094] The multi-color space feature fusion module extracts and integrates complementary features from the RGB, LAB, and HSV color spaces during operation, including the following steps;
[0095] Step B1: Associate the three RGB channels using the A and B channels in the LAB space, so that the training model can learn the association between the three RGB channels in the underwater natural scene. When reconstructing the image, the constraints of the A / B channels on the color make the color fidelity of the reconstructed image higher.
[0096] Step B2: The H channel of the HSV space is used to represent the hue information of different regions of the image. Based on the threshold segmentation of hue, the blur noise introduced by the scattering body is removed, thereby enhancing the target information. By reconstructing the L channel (brightness), V channel (luminance), and S channel (saturation) of the image color space, the brightness and darkness changes of the reconstructed image are more natural (neither too dark nor too exposed) and the colors are richer.
[0097] Step B3: After extracting key information from the above images, the multi-color space feature fusion module integrates the information using a feature fusion strategy; its mathematical expression is as follows:
[0098]
[0099] Among them, F oriF represents the basic features extracted from the original RGB input. RGB ,F HSV and F LAB F represents the features extracted from the corresponding color space. fusion This is the feature fusion function.
[0100] Step B4: After feature fusion, a channel-adaptive attention mechanism is introduced to dynamically adjust the importance of features in each channel, making the network pay more attention to task-related information channels. The network structure is as follows: Figure 3 As shown, its mathematical expression can be represented as:
[0101]
[0102] Where CAAM(·) represents the channel attention function, This indicates element-wise multiplication.
[0103] In step B1, the original color information is preserved in the RGB space, where the three channels are independent of each other. For example, in underwater image reconstruction, if the red channel is not adequately compensated, a blue-green bias will appear.
[0104] The depth enhancement module is used to selectively guide the enhancement network by combining multi-color space information with the depth image obtained in the first part, making non-local features and local features more reasonable. The depth enhancement module also introduces surrogate attention, which further optimizes the feature fusion process through a more lightweight attention model, enabling the network to focus on key regions more efficiently.
[0105] The depth enhancement module is based on an encoder-decoder structure and includes the following steps during operation;
[0106] Step C1: The input underwater image and multidimensional color space information are first extracted stepwise through residual blocks RB and combined with residual skip connections to ensure effective information transmission;
[0107] Step C2: Set up a depth guidance module in the center of the network of the depth enhancement module, with the structure as follows: Figure 4 As shown, it accepts feature information and depth information from the encoder and consists of two feature processing paths: the upper path focuses on high-level semantic feature extraction, while the lower path focuses on multi-scale feature learning and integration, achieving effective processing of feature information, expressed by the formula:
[0108] f ck =d k ×f ck ×β k k = 1, 3, 5 (Formula 7)
[0109] f o =fr +cat(f c1 ,f c3 ,f c5 ) clc Formula 8;
[0110] Where cat() represents a cascading operation along the channel dimension, () clc This represents a convolution-activation-convolution structure, where β represents the feature fusion node;
[0111] Step C3: In the upper-layer path design, a proxy attention mechanism is introduced. This mechanism is an innovative attention computation paradigm, its design concept originating from the proxy concept in reinforcement learning. It dynamically adjusts the feature weight distribution through adaptive learning. Unlike traditional attention mechanisms, it introduces a more flexible context-aware capability, automatically adjusting the focus of attention for different tasks and input features.
[0112] Specifically, it includes four parts: feature transformation projection, attention score calculation, agent decision adjustment, feature aggregation and output, and its mathematical expression is as follows:
[0113] Q = W Q X,K=W K X,V=W V Formula 9;
[0114] Where X is the input feature, W Q W K W V This is the weight matrix for learning.
[0115]
[0116] Where, d k The scaling factor is the square root of the feature dimension.
[0117] A agent =f agent (A raw C) Formula 11;
[0118] Where f agent Let C be the proxy decision function, and C be the context feature matrix used to guide the attention allocation strategy. The proxy decision function is implemented in the following way:
[0119] f agent (A,C)=σ(W a A+W c C+b) Formula 12;
[0120] Where σ is the activation function, W a W c Here is the weight matrix, and b is the bias term;
[0121] Y = softmax(A agent Formula 13;
[0122] The final output feature Y is obtained by weighting and aggregating the value feature V using attention scores normalized by softmax, thereby achieving stronger feature extraction capabilities.
[0123] Step C4: In the design of the lower-level path, a multi-kernel convolution structure is adopted to obtain more perceptual fields, thereby enhancing the details of different sizes in the multi-scale encoded features. The β1, β2 and β3 nodes in the model network represent feature fusion points.
[0124] At these locations, features from different processing paths are integrated through a carefully designed fusion strategy, achieving an effective combination of multi-scale and multi-level features and significantly improving the model's representational capabilities.
[0125] Step C5: Using the feature results obtained in the previous two steps, and combining them with depth information, the underwater image is enhanced in a guided manner. The original size of the image is gradually restored through upsampling, and residual blocks (RB) are introduced to achieve progressively fine reconstruction of features, and finally the enhanced result is output.
[0126] The model training strategy of the method is as follows:
[0127] The model training process is supervised in two stages. The first stage calculates the loss of the deep generation module and the deep enhancement module respectively, which is the offline pre-training process. The second stage uses the deep generation module with fixed parameters to assist the deep enhancement module in generating enhanced images, thereby obtaining the corresponding loss.
[0128] The model training uses the EUVP and UIEB datasets, which are commonly used and publicly available for underwater image augmentation tasks. The first stage of training involves generating depth images of corresponding images using a deep image generation module. Image pairs are randomly selected from the public datasets to train the deep image generation module and the supervision module for 100 epochs, and the L1 loss is calculated.
[0129]
[0130] Where d and d GT These represent the estimated depth map and the depth map of the ground truth GT, respectively. X and d represent the regression image and the underwater image, respectively; the hyperparameter λ1 = 3 is set to balance the loss; by incorporating the regression process into the training, d and X will be more tightly structured.
[0131] In the second stage, image pairs are randomly selected from the UIEB and EUVP datasets for training. To ensure that the augmented images have satisfactory color and detail, Charbonnier loss and SSIM loss are used to determine the total loss rate of the depth augmentation module; the former measures pixel similarity, and the latter measures structural similarity; the formula is as follows:
[0132]
[0133] Where x and y represent the enhanced and sharpened images, respectively, and λ² and ε are the hyperparameters of the balancing loss, set to 0.5 and 10, respectively. -3 .
[0134] The raw underwater images input to the depth information generation module are based on two public datasets, UIEB and EUVP. The underwater scenes include both natural and artificial lighting, with different conditions such as insufficient lighting (overall dark) and excessive lighting (local overexposure). In terms of water turbidity, the range covers water bodies from clear to severely turbid. Suspended particles in turbid water can cause light scattering, resulting in a fogging effect in the image and reducing contrast and clarity.
[0135] In terms of underwater objects, including marine life and underwater plants, these objects often suffer from problems such as blurred outlines, distorted colors, and difficulty in distinguishing details under complex lighting and turbid conditions.
[0136] During model training, all images are resized to a resolution of 256×256 before processing to ensure consistency of the input data.
[0137] Example:
[0138] In this example, to evaluate the performance of the proposed model, experiments were conducted using two benchmark datasets: the UIEB dataset and the EUVP dataset. The UIEB dataset consists of 950 real-world underwater images, divided into two subsets: 890 pairs of original underwater images with corresponding high-quality reference images and 60 pairs of challenging images without references.
[0139] In the UIEB dataset, this example randomly selected 790 image pairs as the training set and 100 image pairs as the test set; however, the EUVP dataset contains both paired and unpaired images, including not only real underwater images but also simulated images, emphasizing the diversity and breadth of the data.
[0140] In the EUVP dataset, this example selects Underwater ImageNet in the paired dataset and randomly selects 3330 image pairs as the training set and 370 image pairs as the test set.
[0141] To ensure the consistency of the input data, all images were resized to a resolution of 256×256 before processing. The model was trained and tested in a PyTorch framework on eight A100 80GB PCIe chips. In this example, the learning rate was initialized to 0.0002 and gradually adjusted using the Adam optimizer.
Claims
1. A depth-information-guided underwater multi-color space image enhancement method, characterized in that: The proposed method is based on a training model framework that can utilize prior physical knowledge and multicolor information. It includes a depth information generation module, a multicolor space feature fusion module, and a depth enhancement module. Through attention estimation, scene reconstruction training, and multicolor information fusion, the connection between depth images and underwater scenes is strengthened.
2. The underwater multi-color space image enhancement method based on depth information guidance according to claim 1, characterized in that: The depth information generation module takes the original underwater image as input and obtains the generated monocular depth image through encoding-decoding and scaling transformation. At the same time, a depth supervision network is introduced during the training phase to generate a regression map, which jointly supervises the training of the depth generation module and its loss.
3. The underwater multi-color space image enhancement method based on depth information guidance according to claim 2, characterized in that: The depth information generation module is based on an encoder and decoder structure, and its operation includes the following specific steps; Step A1: The depth information generation module generates depth images through a dual attention module structure that includes both spatial and channel information. In the channel attention part, an average pooling layer is used to reduce the feature size, and then spatial dependencies are captured by query vector q, key vector k, and value vector v. This process introduces a learnable parameter γ to dynamically adjust the learning intensity, expressed by the formula: q = k = v = (AvgP(f)) Formula 1; f s =γ×((q×k) soft ×v) Formula 2; in,(·) soft It is the softmax function, γ is a learnable parameter used to adjust the intensity, and f s The weights are marked as input to the next stage; Step A2: In the channel module, the module learns the interrelationships and importance between different channels through different pooling operations; f c =f s +f×(AvgP(f) clc +MaxP(f) clc ) sig Formula 3; Where AvgP(·) and MaxP(·) represent the average pooling layer and the max pooling layer, respectively. sig f represents the sigmoid function. c The features of the labeled channel module are input into the weights of the next stage; subsequently, the two attention features are fused and residually connected with the original input to capture hidden features with significant regional degradation. f h =f+f c ×(f s ) dc Formula 4; Step A3: Use the decoder to generate the corresponding depth map, guide the network to perceive different degradation regions, introduce a deep supervision network during the training phase to generate regression mapping maps, and jointly supervise the training and loss of the deep generation module.
4. The underwater multi-color space image enhancement method based on depth information guidance according to claim 1, characterized in that: The multi-color space feature fusion module uses a joint training strategy across color spaces, specifically: extracting complementary feature information from multiple color spaces, including color information, brightness, lightness, and saturation, and integrating them through a feature fusion strategy.
5. The underwater multi-color space image enhancement method based on depth information guidance according to claim 4, characterized in that: The multi-color space feature fusion module extracts and integrates complementary features from the RGB, LAB, and HSV color spaces during operation. Includes the following steps; Step B1: Associate the three RGB channels using the A and B channels in the LAB space, so that the training model can learn the association between the three RGB channels in the underwater natural scene. When reconstructing the image, the constraints of the A / B channels on the color make the color fidelity of the reconstructed image higher. Step B2: Characterize the hue information of different regions of the image using the H channel of the HSV space, and remove the blur noise introduced by the scattering body according to the threshold segmentation of hue, thereby enhancing the target information; By reconstructing the L, V, and S channels of the image color space, the brightness variations of the reconstructed image are made more natural and the colors are richer. Step B3: After extracting key information from the above images, the multi-color space feature fusion module integrates the information using a feature fusion strategy; its mathematical expression is as follows: Among them, F ori F represents the basic features extracted from the original RGB input. RGB ,F HSV and F LAB F represents the features extracted from the corresponding color space. fusion For feature fusion function; Step B4: After feature fusion, a channel-adaptive attention mechanism is introduced to dynamically adjust the importance of features in each channel, making the network pay more attention to task-related information channels. The mathematical expression of the network structure is: Where CAAM(·) represents the channel attention function, This indicates element-wise multiplication.
6. The underwater multi-color space image enhancement method based on depth information guidance according to claim 5, characterized in that: In step B1, the original color information is preserved in the RGB space, where the three channels are independent of each other.
7. The underwater multi-color space image enhancement method based on depth information guidance according to claim 1, characterized in that: The depth enhancement module is used to selectively guide the enhancement network by combining multi-color space information with the depth image obtained in the first part, making non-local and local features more reasonable. The deep enhancement module introduces proxy attention, which optimizes the feature fusion process through attention modeling, enabling the network to focus on key regions.
8. The underwater multi-color space image enhancement method based on depth information guidance according to claim 7, characterized in that: The depth enhancement module is based on an encoder-decoder structure and includes the following steps during operation; Step C1: The input underwater image and multidimensional color space information are first extracted stepwise through residual blocks RB and combined with residual skip connections to ensure effective information transmission; Step C2: A deep guidance module is set up in the center of the deep enhancement module's network. This module receives feature information and depth information from the encoder and consists of two feature processing paths: the upper path focuses on high-level semantic feature extraction, while the lower path focuses on multi-scale feature learning and integration, achieving effective processing of feature information. This can be expressed as a formula: f ck =d k ×f ck ×β k k = 1, 3, 5 (Formula 7) f o =f r +cat(f c1 ,f c3 ,f c5 ) clc Formula 8; Where cat() represents a cascading operation along the channel dimension, () clc This represents a convolution-activation-convolution structure, where β represents the feature fusion node; Step C3: In the upper-level path design, an agent attention mechanism is introduced, which includes four parts: feature transformation projection, attention score calculation, agent decision adjustment, feature aggregation and output. Its mathematical expression is as follows: Q = W Q X,K=W K X,V=W V Formula 9; Where X is the input feature, W Q W K W V This is the weight matrix for learning. Where, d k The scaling factor is the square root of the feature dimension. A agent = f agent (A raw , C) Formula 11; Where f agent Let C be the proxy decision function, and C be the context feature matrix used to guide the attention allocation strategy. The proxy decision function is implemented in the following way: f agent (A, C) = σ(W a A + W c C + b) Equation 12; Where σ is the activation function, W a W c Here is the weight matrix, and b is the bias term; Y = softmax(A agent )V Equation 13; The final output feature Y is obtained by weighting and aggregating the value features V using attention scores normalized by softmax. Step C4: In the design of the lower-level path, a multi-kernel convolution structure is adopted to obtain more perceptual fields, thereby enhancing the details of different sizes in the multi-scale encoded features. The β1, β2 and β3 nodes in the model network represent feature fusion points. Step C5: Using the feature results obtained in the previous two steps, and combining them with depth information, the underwater image is enhanced in a guided manner. The original size of the image is gradually restored through upsampling, and residual blocks (RB) are introduced to achieve progressively fine reconstruction of features, and finally the enhanced result is output.
9. The underwater multi-color space image enhancement method based on depth information guidance according to claim 2, characterized in that: The model training strategy of the method is as follows: The model training process is supervised in two stages. The first stage is to calculate the loss of the deep generation module and the deep augmentation module respectively, which is the offline pre-training process. The second stage uses a depth generation module with fixed parameters to assist the depth enhancement module in generating an enhanced image, thereby obtaining the corresponding loss. The datasets used for model training are the EUVP and UIEB datasets from the underwater image augmentation task; The first stage of training involves generating depth images of corresponding images using a deep generation module. Image pairs are randomly selected from a public dataset to train the deep generation module and the supervision module, and the L1 loss is calculated. Where d and d GT These represent the estimated depth map and the depth map of the ground truth GT, respectively. X and d represent the regression image and the underwater image, respectively; the hyperparameter λ1 is set to balance the loss; by incorporating the regression process into the training, d and X will be more tightly structured. In the second stage, image pairs are randomly selected from the UIEB and EUVP datasets for training; Charbonnier loss and SSIM loss are used to determine the total loss rate of the depth augmentation module; the former measures pixel similarity, and the latter measures structural similarity; the formula is as follows: Where x and y represent the enhanced image and the sharpened image, respectively, and λ2 and ε are the hyperparameters of the balancing loss.
10. The underwater multi-color space image enhancement method based on depth information guidance according to claim 2, characterized in that: The raw underwater images input to the depth information generation module are based on two public datasets, UIEB and EUVP. Underwater scenes: In terms of lighting, there are both natural and artificial lights, with some areas experiencing insufficient or excessive lighting; in terms of water turbidity, the range covers waters from clear to severely turbid. Suspended particles in turbid water can cause light scattering, resulting in a fogging effect in the image and reducing contrast and clarity. In terms of targets in underwater scenarios, these include marine life and underwater plants; During model training, all images are resized to a resolution of 256×256 before processing to ensure consistency of the input data.
Citation Information
Cited By
Degree quantization underwater depth estimation method based on confidence guiding fusion and monocular backspacing mechanism
CN121527149A
Underwater image depth estimation method oriented to real-time underwater robot and based on adaptive physical information and significance guidance
CN121640258A
Image processing method and system for identifying and positioning consumables of thrombelastography instrument
CN121746760A