Underwater degraded image restoration method based on consistency model and related device
By constructing a lightweight underwater depth estimation network based on a consistency model and a space-depth guided denoising network, the problems of inconsistent generation results and insufficient real-time performance in underwater degraded image restoration are solved, achieving high-precision and physically consistent underwater image restoration results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI UNIV OF SCI & TECH
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing underwater image restoration methods suffer from problems such as lack of physical consistency in the generated results, incomplete color correction, insignificant brightness enhancement, and unsatisfactory contrast improvement, making it difficult to meet the needs of accurate observation in extreme environments. Furthermore, deep learning methods struggle to meet real-time requirements during the sampling process.
An underwater degraded image restoration method based on a consistency model is adopted. By constructing a lightweight underwater depth estimation network LUDEN and a spatial-depth guided denoising network SDGDN, multi-scale feature extraction and depth feature fusion are combined. Feature selection and fusion are performed using depth-separable convolution and pixel attention modules to construct a collaborative denoising network and directly restore the image in the consistency model.
It achieves high-precision and physically consistent restoration of underwater degraded images, improves the color correction effect and brightness of the images, can obtain clear underwater images in a single sampling, adapts to complex and ever-changing underwater environments, and meets the requirements of real-time and precise observation.
Smart Images

Figure CN122048731A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to a method and related apparatus for underwater degraded image restoration based on a consistency model. Background Technology
[0002] During the underwater image acquisition phase, the absorption and scattering of incident light by the water body are simultaneously affected, and the intensity of these two effects fluctuates regularly with changes in water turbidity. Specifically, the selective absorption of long-wavelength light by water inevitably causes a blue-green tint in images acquired by underwater passive imaging systems, directly interfering with the accuracy of target identification. Furthermore, the light scattering effect induced by suspended particles and water molecules in the water not only blurs image details and reduces contrast but also significantly compresses the effective detection range of underwater observation systems, drastically reducing detection efficiency. Simultaneously, the inherent low-light conditions of the underwater environment severely limit the light throughput capture efficiency of detection devices in passive imaging mode, leading to an overall darker image and further exacerbating the loss of image information. To address these technical challenges, there is an urgent need to develop an underwater image correction and compensation technology system to simultaneously correct for issues such as color cast, blurring, and insufficient brightness during the imaging process, thereby ensuring the stable operation of the underwater imaging system and its high-precision target detection capabilities.
[0003] Currently, the methods used for underwater degraded image restoration in existing technologies are mostly: 1) using mathematical modeling and inversion methods to restore underwater degraded images; 2) using deep learning methods to restore underwater degraded images.
[0004] Mathematical modeling inversion methods restore underwater images by constructing a physical model and performing inversion operations. The restoration performance is highly dependent on the accuracy of the physical model, as well as the estimation accuracy of core parameters such as transmittance and atmospheric light intensity. However, accurately solving for these unknown parameters is typically quite challenging.
[0005] When using deep learning methods to restore degraded underwater images, a dedicated network architecture must first be built, and then the mapping relationship between degraded and clear images must be learned based on sample data. This method not only has better robustness but also adapts well to complex and changing underwater environments, and has now become the mainstream technology in the field of underwater degraded image restoration.
[0006] For example, patent document CN119887552B discloses a lightweight underwater image enhancement method based on a step-sampling diffusion model. This method achieves independent parallel processing of temporal step encoding and color information by designing a parallel structure of an attention-driven Transformer (AP-Trans) module. At the same time, it introduces a spatial attention mechanism to enhance detail recovery capabilities and uses a global channel interaction module to maintain color fidelity, thus achieving high-quality enhancement of underwater images. Furthermore, by replacing the large-parameter self-attention module in the traditional diffusion model with a lightweight channel attention mechanism, the number of model parameters is significantly reduced. In addition, a dynamic step-sampling strategy is adopted to reduce the 20-50 step sampling process of the traditional step-sampling diffusion model to 5 steps. This method significantly improves the enhancement effect while maintaining the excellent generation capability of the diffusion model, effectively solves the problem of interaction interference between temporal step encoding and color information, and balances the restoration of color and detail in underwater images.
[0007] While deep learning-based methods can achieve good underwater image restoration, they still have the following obvious limitations: 1. Although deep learning methods can solve problems such as color cast and blur in underwater images, the generated results often lack physical consistency and are prone to deviating from the optical characteristics of real underwater scenes. Specifically, this manifests as uneven restoration, incomplete color correction, insignificant brightness improvement, and unsatisfactory improvement in contrast and sharpness, making it difficult to meet the needs of accurate observation in extreme environments.
[0008] 2. In deep learning methods, in order to accurately fit the data distribution from degraded image to clear image, network models often need to rely on more advanced models, such as underwater image restoration methods based on diffusion models (DDPM, Denoising Diffusion Probabilistic Model), which can further improve the generalization of the model. However, these methods are limited by multiple iterations of sampling during the sampling process, making it difficult to meet the needs of application scenarios with strict real-time requirements. Summary of the Invention
[0009] The purpose of this invention is to provide a method and related apparatus for underwater degraded image restoration based on a consistency model, in order to solve the problem of poor performance in underwater degraded image restoration in the prior art.
[0010] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for underwater degraded image restoration based on a consistency model, comprising the following steps: The original underwater degraded image is acquired, and noise is added to the original underwater degraded image to obtain a noisy image; The original underwater degraded image is input into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The depth pseudo-labels were filtered to obtain the RRDB (Reliable RGB-Depth Bank) dataset; Based on the RRDB dataset, the LUDEN (Lightweight Underwater Depth Estimation Network) network was trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv (Mobile Inverted Bottleneck Convolution) module and a decoder network based on the DSConvB (Depthwise Separable Convolution Block) module and the FFB (Feature Fusion Block) module. The depth information, the original underwater degraded image, and the noisy image are input into the trained SDGDN (Spatially and Depth-guided Denoising Network) network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB (Multi-scale feature extraction Block) module and the SDFFB (Spatial-Depth Features Fusion Block) module, the MSDAB (Multi-Scale Spatial-Depth Attention Block) module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on the depth-separable convolutional module and the pixel attention module, an LDFSB (Lightweight depthfeature selection Block) module is constructed. The LDFSB module is used to extract features from the depth information. SDGDN network is constructed based on the MSDAB and LDFSB modules; The intermediate result image is input into the parameterization formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
[0011] A further improvement of this invention is that, when training the LUDEN network based on the RRDB dataset, a composite loss function is used, the expression of which is:
[0012] in, Represents the composite loss function. express loss, This represents the depth estimation result of the LUDEN network on the original underwater degraded image. This represents the depth pseudo-label corresponding to the original underwater degraded image. Represents the structural similarity index. and This represents the weighting coefficient that balances the two types of losses.
[0013] A further improvement of the present invention is that the MSFEB module adopts a multi-scale convolution parallel processing mechanism.
[0014] A further improvement of the present invention is that the LDFSB module adopts a serial processing mechanism of depth-separable convolution module and pixel attention module.
[0015] A further improvement of this invention is that the parameterization formula of the consistency model is:
[0016] in, Representing a consistency model, This indicates the parameters of the SDGDN network. This represents the parameters of the LUDEN network after pre-training and freezing. This represents a noisy image. Indicates the boundaries of the time range. This represents the original underwater degraded image. and All are differentiable functions. This represents an intermediate result image. This represents a pre-trained LUDEN network with frozen parameters, used to obtain depth information.
[0017] A further improvement of the present invention is that the trained SDGDN network is obtained through the following steps: Obtain the trained LUDEN network and the untrained SDGDN network, and freeze the parameters of the trained LUDEN network; Based on the isolation training framework of the consistency model, a collaborative denoising network is constructed, which includes a LUDEN network and an SDGDN network. Obtain the ground truth image corresponding to the original underwater degradation image; A training set is constructed based on the original underwater degraded image and the ground truth image corresponding to the original underwater degraded image; The collaborative denoising network is trained based on the training set to obtain the trained SDGDN network.
[0018] A further improvement of this invention is that, when training the collaborative denoising network based on the training set, the loss function used is a consistency constraint loss function, the expression of which is:
[0019] in, Representing a consistency model, Indicates to Adding noise , This represents the model parameters updated using the exponential moving average strategy. Indicates about The weighting function is used to balance the loss weights at different time boundaries. , Indicates the minimum weight. Indicates the maximum weight. Represents the metric function. Indicates the distance from the walk. This represents a discrete interval index.
[0020] Secondly, the present invention provides an underwater degraded image restoration system based on a consistency model, comprising: The data acquisition module is used to acquire the original underwater degraded image and add noise to the original underwater degraded image to obtain a noisy image; The depth information determination module is used to input the original underwater degraded image into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. The intermediate result image determination module is used to input depth information, the original underwater degraded image, and the noisy image into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; The image restoration module is used to input the intermediate result image into the parameterized formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
[0021] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the underwater degraded image restoration method based on the consistency model described above.
[0022] Fourthly, the present invention provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the underwater degraded image restoration method based on the consistency model described above.
[0023] Compared with the prior art, the present invention has the following beneficial effects: The underwater degraded image restoration method proposed in this invention, based on a consistency model, firstly filters underwater depth pseudo-labels to obtain the RRDB dataset. Based on the RRDB dataset, a LUDEN network is trained to obtain a depth estimator. This operation can quickly and effectively obtain the depth information corresponding to the original underwater degraded image, laying the foundation for improving the sampling speed of the consistency model in the restoration system. Secondly, this invention constructs an SDGDN network based on the MSDAB and LDFSB modules. The SDGDN network can accurately fuse depth information during the encoding stage using the LDFSB module, filtering out redundant depth features while ensuring the encoder's lightweight nature. Simultaneously, the SDGDN network can accurately perceive depth features in the target scene using the MSDAB module, performing regional restorations of varying intensities to address the uneven degradation problem in underwater degraded images. By implicitly modeling underwater degradation factors, the restored image is made more consistent with the real physical scene, thus solving the problem of poor underwater degraded image restoration performance in existing technologies. Furthermore, this invention inputs the intermediate result image into the parameterized formula of the trained consistency model to directly obtain the restored underwater degraded image. This demonstrates that while maintaining the high generalization ability of the generative model, this invention only requires one sampling to obtain a clear restored underwater image (the restored underwater degraded image). Simultaneously, addressing the complex degradation characteristics of underwater images, this invention employs a dual-guidance approach by fusing spatial information (the original underwater degraded image) and depth prior information (also called depth information), thereby improving the accuracy and fidelity of underwater degraded image restoration.
[0024] Furthermore, this invention discloses that the LDFSB module adopts a serial processing mechanism of depth-separable convolution module and pixel attention module. This design can accurately filter depth information (e.g., filter out areas with blurred scene depth), assist the SDGDN network in accurately perceiving depth, and indirectly characterize and utilize underwater transmittance to further improve the robustness of image restoration in complex scenes. Attached Figure Description
[0025] Figure 1 This is a flowchart of the underwater degraded image restoration method based on the consistency model of the present invention; Figure 2 This is a schematic diagram of the underwater degraded image restoration system based on the consistency model of the present invention; Figure 3 This is a block diagram illustrating the overall principle of the underwater degraded image restoration method based on the consistency model in Embodiment 3 of the present invention. Figure 4 This is a schematic diagram of underwater depth pseudo-tag generation in Embodiment 3 of the present invention and a schematic diagram of the network structure of LUDEN; Figure 5 This is a qualitative evaluation comparison chart of LUDEN's depth estimation of underwater degraded images in Embodiment 3 of the present invention; Figure 6 This is a network structure diagram of the MSDAB module in Embodiment 3 of the present invention; Figure 7 This is a network structure diagram of the LDFSB module in Embodiment 3 of the present invention; Figure 8 This is a comparison of the qualitative evaluation results of the present invention with mainstream image restoration methods on the UIEB, LSUI, and EUVP datasets in Embodiment 3 of the present invention. Figure 9 This is a comparison of the quantitative evaluation results of the method with mainstream image restoration methods on the LSUI dataset in Embodiment 3 of the present invention; Figure 10 This is a qualitative result diagram of the generalization experiment performed on the O-Haze and LOL-Dataset datasets in Embodiment 3 of the present invention; Figure 11 This is a comparison of feature point matching results before and after underwater degraded image restoration using the method of the present invention in Example 3; Figure 12 This is a comparison of edge detection results before and after underwater degraded image restoration using the method of the present invention in Example 3; Figure 13 This is a schematic diagram of the structure of the electronic device of the present invention. Detailed Implementation
[0026] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.
[0027] Example 1: The flowchart of the underwater degraded image restoration method based on the consistency model of this invention is as follows: Figure 1 As shown, the underwater degraded image restoration method based on the consistency model of the present invention includes the following steps: S1. Acquire the original underwater degraded image and add noise to the original underwater degraded image to obtain a noisy image; S2. Input the original underwater degraded image into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. S3. Input the depth information, the original underwater degraded image, and the noisy image into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; S4. Input the intermediate result image into the parameterization formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
[0028] Example 2: A schematic diagram of the underwater degraded image restoration system based on the consistency model of this invention is shown below. Figure 2 As shown, the underwater degraded image restoration system based on the consistency model of the present invention includes: The data acquisition module is used to acquire the original underwater degraded image and add noise to the original underwater degraded image to obtain a noisy image; The depth information determination module is used to input the original underwater degraded image into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. The intermediate result image determination module is used to input depth information, the original underwater degraded image, and the noisy image into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; The image restoration module is used to input the intermediate result image into the parameterized formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
[0029] Example 3: The overall principle block diagram of the underwater degraded image restoration method based on the consistency model of this invention is as follows: Figure 3 As shown, the underwater degraded image restoration method based on the consistency model of the present invention includes the following steps: S1. Acquire the original underwater degraded image and add noise to the original underwater degraded image to obtain a noisy image.
[0030] S2. Input the original underwater degraded image into the depth estimator to obtain depth information.
[0031] The method for constructing the depth estimator in this step includes: A. Perform underwater scene depth estimation on the original degraded underwater image to obtain underwater depth pseudo-labels; B. Filter the pseudo-labels of underwater depth to obtain the RRDB dataset; C. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module.
[0032] Step A will be explained in detail below: In step A, such as Figure 4 As shown in (a), the MiDas model and the Depth AnythingV2 model are two large-scale depth estimation models, which can output reliable depth pseudo-labels. This invention utilizes these two large-scale models to perform depth estimation on five widely used underwater datasets (UIEB, LSUI, EUVP, SUIM, and UFO120), resulting in two underwater depth pseudo-label libraries totaling 36k images.
[0033] In step B, the two underwater depth pseudo-label libraries obtained in step A are filtered twice (also called two depth label selection operations). The three indicators GM (Gradient-Magnitude), LV (Laplacian-Variance), and EPR (Edge-Pixel-Ratio) are used to filter the underwater depth map labels that have clear edges, rich details, and are more in line with the characteristics of human eye perception.
[0034] The two depth label selection operations in step B specifically include: Based on two underwater depth pseudo-label libraries obtained from the large depth estimation model, the thresholds of GM, LV and EPR are dynamically set according to the characteristics of human eye perception. Through the joint constraints of the three conditions, a preliminary depth pseudo-label library is obtained.
[0035] Based on the initial pseudo-deep label library, depth maps from the same source in the two label libraries are assigned different weights using three filtering criteria, and the corresponding scores of the two datasets are calculated. The depth map with the higher score is selected and placed into the initial trusted label library, along with non-overlapping labels from the two label libraries. Then, the initial trusted label library is filtered again to obtain the RRDB dataset.
[0036] The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB and FFB modules. The LUDEN network is as follows: Figure 4 As shown in (b), the encoder network and decoder network are described below: The encoder network is built based on the MBConv module of EfficientNet. The encoder network uses group normalization instead of batch normalization to improve the stability and performance of the model under small batch training.
[0037] The decoder network aims to balance feature reconstruction quality with model lightweighting. First, it uses a depthwise separable convolutional module (DSConvB) to map the encoder output to a high-dimensional space to enhance feature representation capabilities. Then, based on the fusion module (FFB) in DPT (VisionTransformers for Dense Prediction), it replaces all the standard 3×3 convolutions with DSConvB modules, significantly reducing the number of model parameters and computational complexity.
[0038] In this step, when training the LUDEN network based on the RRDB dataset, the loss function used is a composite loss function, the expression of which is:
[0039] in, Represents the composite loss function. express loss, This represents the depth estimation result of the LUDEN network on the original underwater degraded image. This represents the depth pseudo-label corresponding to the original underwater degraded image. Represents the structural similarity index. and In this embodiment, the weighting coefficients represent the factors that balance the two types of losses. and The values are 0.15 and 0.85, respectively.
[0040] The trained LUDEN network has only 4.7M parameters. Qualitative results of the trained LUDEN network for underwater depth estimation are as follows: Figure 5 As shown, the first row represents underwater images of different degradation scenarios, and the second row represents the corresponding depth images. Figure 5 (a) to Figure 5 (c) represents a depth image of an underwater scene with a greenish tint. Figure 5 (d) and Figure 5 (e) represents a depth image of an underwater scene with a bluish tint. Figure 5 (f) represents a depth image of an underwater low-light scene. From Figure 5 As can be seen, the LUDEN network can estimate the depth information of images in different underwater degradation scenarios. The estimated depth information is rich in layers and retains complete details, laying the foundation for the effective use of the depth information in the future.
[0041] The training process of the LUDEN network is explained in detail below: First, initialize the training image pairs based on the RRDB pseudo-label library (also called the RRDB dataset). ,in, Indicates underwater degraded images, This represents the depth pseudo-label corresponding to the degraded image. This indicates the number of training images. During training, from... Medium-sampled degraded image and corresponding depth pseudo-tags and will The input model, after convolution, group normalization, and nonlinear activation operations, yields shallow high-resolution features. Then, high-resolution features The corresponding low-resolution features are obtained after four different layers of MBConv operations. ,in To ensure feature representation capability, the four low-resolution features are first mapped to a high-dimensional feature space using DSConvB, and then fed into the DSCResB module and three FFB units of the decoder via skip connections. Simultaneously, to reduce the number of parameters and ensure depth estimation performance, 1×1 pointwise convolutions are used for dimensionality reduction and channel information integration. In the last layer of the decoder, the fourth feature is processed in the same dimensional space, and then channel information is integrated using pointwise convolutions. Finally, grouped convolutions are used for dimensionality reduction to obtain deep, high-resolution features. Finally, after upsampling to restore the input image size, the feature is output as depth information through the output head. To ensure that the depth information estimated by LUDEN has accurate numerical values and good visual structural integrity, this invention uses... The network is optimized by combining loss and structural similarity index loss.
[0042] S3. Input the depth information, the original underwater degraded image, and the noisy image into the trained SDGDN network to obtain the intermediate result image.
[0043] The SDGDN network in this step specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; An SDGDN network is constructed based on the MSDAB and LDFSB modules.
[0044] In this embodiment, the MSFEB module employs a multi-scale convolution parallel processing mechanism. The LDFSB module employs a depthwise separable convolution module and a pixel attention module serial processing mechanism.
[0045] The following provides a detailed explanation of the MSDAB, MSFEB, and SDFFB modules: The network structure of the MSDAB module is as follows: Figure 6 As shown, using Figure 6 (a) The two MSFEB modules extract spatial and depth features respectively, and then... Figure 6(b) uses the SDFFB module to fuse spatial and depth features. The input features are connected through a residual structure and further optimized using a pre-selected residual block (PRB). Through this depth feature perception mechanism, the model can effectively perceive depth features while integrating multi-scale features, strengthening the weight of effective information, suppressing redundant noise, and achieving spatially important feature region representation guided by depth features.
[0046] The MSFEB module employs grouped convolutions with different kernel sizes for feature representation, inputting spatial and depth features into two parallel MSFEB modules with identical structures. Each MSFEB module uses grouped convolutions with a dilation rate of 2 and a kernel size of 5 to represent the output features of grouped convolutions with a kernel size of 3. Then, the two output features are passed through 1×1 pointwise convolutions for inter-channel information exchange. The specific operation process of the MSFEB module is as follows:
[0047] in, This represents the features (spatial features or depth features) input into MSDAB. The features represented by small kernel convolution are... This represents the characteristics of large kernel convolution. This indicates a group normalization operation. Represents a non-linear activation function. This represents pointwise convolution. This represents a grouped convolution with a kernel size of 3. This represents a grouped convolution with a kernel size of 5 and an inflation rate of 2. After passing through MSFEB, the four different features from the two branches are concatenated to obtain... .
[0048] The SDFFB module will After global average pooling and capturing global dependencies along the channel dimension through a fully connected layer, a channel attention matrix is generated using a sigmoid activation function. This matrix is then compared with the initial features. Matrix multiplication is performed to effectively model the channel dependencies between spatial and depth features and enhance the feature information of key channels. Then, the SDFFB module divides the enhanced features into four features based on different scales. The spatial and depth features are then multiplied element-wise according to their scales, resulting in two scales of perceptual features after applying depth features to spatial features. These two perceptual features are then added element-wise to achieve feature fusion at different scales. Finally, 1×1 convolutions are used to further filter channel information. SDFFB can be represented as:
[0049] in, Indicates to Features after channel information enhancement This represents spatial features after depth perception enhancement and fusion of multi-scale information. Indicates global average pooling. Indicates a fully connected layer. This represents the Sigmoid activation function. This indicates segmentation by channel.
[0050] The characteristic input-output relationship of the MSDAB module can be represented as follows:
[0051] in, This indicates the characteristics of multi-scale spatial deep fusion. Represents the spatial features of the input. This represents the deep features of the input. This indicates that the components are assembled according to the channels.
[0052] The network structure of the LDFSB module is as follows: Figure 7 As shown, the LDFSB module employs a residual structure. Deep features are input into the LDFSB module, first undergoing group normalization and non-linear activation functions to enhance the model's expressive power, then passing through 3×3 depthwise separable convolutions and pointwise convolutions to capture spatial and channel features. Subsequently, the features pass through two branches. In the first branch, the features undergo pointwise convolutions to reduce dimensionality, followed by a non-linear activation function, and then further pointwise convolutions to reduce dimensionality to a single channel. A sigmoid activation function is used to calculate the pixel attention matrix, which is then element-wise multiplied with the features from the second branch to filter important deep features. Finally, the residual connections are added to the initial input to output the selected features.
[0053] The process of adding noise to the original degraded underwater image and the process of sampling the noisy image are as follows: Figure 3 As shown in (a), the noise addition process is based on the probability flow ordinary differential equation (PF-ODE) to add noise to the ground image. The added noise conforms to a standard normal distribution. .
[0054] The specific process of sampling a noisy image once is as follows: Figure 3 As shown in (b) and (c), the LUDEN and SDGDN networks are used for collaborative inference to first process the original underwater degraded image. The LUDEN network is used as a guiding condition to estimate single-channel depth information; then, the number of channels of this depth information is expanded to 3 channels to obtain... and will Input into the SDGDN network; then set the boot conditions. Noisy images (also called images with added noise) After concatenation, the data is input into the SDGDN encoder. The LDFSB module is used to extract depth features and integrate them into each coding layer. With the assistance and guidance of depth features, spatial text features are selectively fused with depth features for encoding, indirectly representing and utilizing the underwater scattering environment transmittance, and finally obtaining intermediate results.
[0055] The trained SDGDN network in this step is obtained through the following steps: Obtain the trained LUDEN network and the untrained SDGDN network, and freeze the parameters of the trained LUDEN network; Based on the isolation training framework of the consistency model, a collaborative denoising network is constructed, which includes a LUDEN network and an SDGDN network. Obtain the ground truth image corresponding to the original underwater degradation image; A training set is constructed based on the original underwater degraded image and the ground truth image corresponding to the original underwater degraded image; The collaborative denoising network is trained based on the training set to obtain the trained SDGDN network.
[0056] In this step, when training the collaborative denoising network based on the training set, the loss function used is the consistency constraint loss function, the expression of which is:
[0057] in, Representing a consistency model, Indicates to Adding noise , This represents the model parameters updated using the exponential moving average strategy. Indicates about The weighting function is used to balance the loss weights at different time boundaries. , Indicates the minimum weight. Indicates the maximum weight. The present invention uses L2 and LPIPS loss functions as the metrics, with the weights of L2 loss and LPIPS loss set to 1.0 and 0.3, respectively. Indicates the distance from the walk. This represents a discrete interval index, in this embodiment. , .
[0058] S4. Input the intermediate result image into the parameterization formula of the trained consistency model to directly obtain the restored underwater degradation image.
[0059] The consensus model in this step includes the LUDEN network and the SDGDN network.
[0060] The parameterization formula for the consistency model in this step is:
[0061] in, Representing a consistency model, This indicates the parameters of the SDGDN network. This represents the parameters of the LUDEN network after pre-training and freezing. This represents a noisy image. Indicates the boundaries of the time range. This represents the original underwater degraded image. and All are differentiable functions. This represents an intermediate result image. This represents a pre-trained LUDEN network with frozen parameters, used to obtain depth information.
[0062] To verify the effectiveness of the proposed underwater degraded image restoration method based on a consistency model, this embodiment uses three large publicly available underwater image datasets for training and validation. LSUI includes 4279 paired images covering different water types, lighting conditions, and target subjects. 3779 images were used for training, and 500 images served as the validation set. UIEB contains 950 real underwater images; 800 images with reference images were used as the training set, 90 as the validation set, and 60 unlabeled but challenging images were used as the generalization performance test dataset. EUVP-Scence contains 2185 natural and synthetic underwater images; 1960 were used as the training set, and 225 as the validation set. To evaluate the generalization performance of this invention, the C60 dataset and the OceanDark dataset with 183 underwater images under artificial lighting were used for generalization performance testing. Furthermore, the O-Haze dataset and the LOL-Dataset dataset were used to test the generalization performance of this invention in other domains. The training process consisted of two phases: Phase 1 involved pre-training LUDEN using the Adam optimizer with a learning rate of 1e-4. Data augmentation employed color shift, resizing, and rotation strategies, and the training lasted for 200 epochs. Phase 2 involved training SDGDN, resizing images to 256*256 pixels and constructing a corresponding dataset to ensure consistency with other comparative methods. The RAdam optimizer was used with a batch size of 8, an initial learning rate of 4e-4, a minimum learning rate of 4e-5, weight decay of 4e-5, and random rotation for data augmentation. Based on the SDGDN training configuration, the network was trained for 1000 epochs. The initial training phase included a 2000-iteration warm-up of the learning rate, followed by a cosine annealing learning rate decay strategy until training was complete. The hardware environment for network training consisted of an Intel Xeon Silver 4210R @ 2.4GHz processor, 128GB of RAM, and three GPUs: an NVIDIA GeForce RTX 3090 with 24GB of RAM.
[0063] This embodiment restores the test images from the above three datasets and compares them with four mainstream traditional methods (MLLE, SPDF, WWPF, WFAC) and eight deep learning-based methods (URSCT, U-shape, Diffwater, Pixmamba, WF-Diff, HCLR-Net, SeaDiff, and GHS-UIR).
[0064] The qualitative evaluation results of this invention on three different datasets compared with mainstream image restoration methods are as follows: Figure 8 As shown, from Figure 8 As can be seen from the image restored by the present invention ( Figure 8 The images marked Ours not only exhibit significant color correction effects but also effectively preserve image details. In some blurred or complex degraded scenes, this invention significantly outperforms other methods in restoring texture clarity and edge sharpness, and the restored images are closer to the reference images.
[0065] To further quantitatively evaluate the quality of underwater degraded image restoration, this embodiment performs statistical averaging analysis on 500 test images from LSUI and uses full-reference evaluation metrics, including peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
[0066] PSNR is used to measure the difference between two images, such as an underwater degraded image and a real image. The minimum PSNR value is 0, and the larger the PSNR value, the smaller the difference between the two images. SSIM is based on the assumption that the human eye can extract structured information from an image, which is more in line with human visual perception than traditional methods. SSIM is less than or equal to 1, and the larger the SSIM value, the more similar the two images are.
[0067] The quantitative evaluation results of this invention on the LSUI dataset, compared with mainstream image restoration methods, are as follows: Figure 9 As shown, from Figure 9 As can be seen, the PSNR and SSIM scores of this invention are the highest, indicating that the underwater degraded image restored by the method of this invention has the best effect in terms of color balance, clarity and contrast, and is most similar to the ground truth image corresponding to the underwater degraded image, indicating that this invention has achieved good underwater degraded image restoration quality.
[0068] To further verify the generalization performance of this invention in other related fields (natural environment dehazing and low-light enhancement), it was first trained on an 800-image paired dataset from UIEB, and then representative scene images from the O-Haze and LOL-Dataset datasets were restored. The results are as follows. Figure 10 As shown, from Figure 10 As can be seen, this invention not only demonstrates excellent performance in underwater environments but also exhibits superior generalization results in foggy and low-light scenes. For foggy images, the fog is effectively eliminated. For low-light images, not only is image brightness significantly improved, but the restored image also displays richer colors.
[0069] To further verify the effectiveness of this invention in computer vision applications, this embodiment experimentally tests the effectiveness of the underwater degraded image restoration method proposed in this invention in feature point detection and edge detection.
[0070] This embodiment uses the SIFT feature point detection algorithm to evaluate the feature point detection effect before and after underwater degraded image restoration. Figure 11 As shown. Figure 11 In the images marked (a), (b), and (c), the left side of the first row represents the underwater degradation image. Figure 11 In the images marked (a), (b), and (c), the right side of the first row represents the ground truth image corresponding to the underwater degradation image. Figure 11 In the images marked (a), (b), and (c), the left side of the second row represents the underwater degradation image restored by the present invention. Figure 11 The right side of the second row in each of the images marked (a), (b), and (c) represents the ground truth image corresponding to the restored underwater degradation image. Figure 11 (d) is a comparison chart of the number of feature points matched before and after underwater degraded image restoration. Figure 11 (d) It can be seen that after the underwater degraded image is restored by the present invention, the number of paired feature points detected by the SIFT feature point detection algorithm is significantly increased.
[0071] This embodiment uses the Canny edge detection algorithm to evaluate the impact of underwater degraded image restoration on edge detection before and after restoration. Figure 12 As shown. Figure 12 The row of images labeled (a) represents the original underwater degraded image, and the row of images labeled (b) represents the underwater degraded image restored using the method of this invention. From Figure 12 As can be seen in (b), the restored image has richer edge details and the number of detected edges is significantly increased, which shows that the image restoration process of the present invention significantly improves the image edges.
[0072] In summary, the underwater degraded image restoration method based on the consistency model proposed in this invention can effectively achieve high-quality restoration of underwater degraded images. This method not only accurately corrects color distortion in underwater images, significantly improving image contrast and clarity, but also adaptively compensates for image brightness, further enhancing the perception and recognition performance of target scenes in underwater environments.
[0073] Example 4: Please see Figure 13 As shown, the present invention also provides an electronic device 100 for underwater degraded image restoration based on a consistency model; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0074] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the underwater degraded image restoration method based on the consistency model described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0075] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.
[0076] The memory 101 in the electronic device 100 stores multiple instructions to implement an underwater degraded image restoration method based on a consistency model, and the processor 102 can execute the multiple instructions to achieve the following: The original underwater degraded image is acquired, and noise is added to the original underwater degraded image to obtain a noisy image; The original underwater degraded image is input into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. The depth information, the original underwater degraded image, and the noisy image are input into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; The intermediate result image is input into the parameterization formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
[0077] Example 5: If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0078] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0080] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0081] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for underwater degraded image restoration based on a consistency model, characterized in that, Includes the following steps: The original underwater degraded image is acquired, and noise is added to the original underwater degraded image to obtain a noisy image; The original underwater degraded image is input into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. The depth information, the original underwater degraded image, and the noisy image are input into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; The intermediate result image is input into the parameterization formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
2. The underwater degraded image restoration method based on a consistency model according to claim 1, characterized in that, When training the LUDEN network based on the RRDB dataset, a composite loss function is used. The expression for the composite loss function is as follows: in, Represents the composite loss function. express loss, This represents the depth estimation result of the LUDEN network on the original underwater degraded image. This represents the depth pseudo-label corresponding to the original underwater degraded image. Represents the structural similarity index. and This represents the weighting coefficient that balances the two types of losses.
3. The underwater degraded image restoration method based on a consistency model according to claim 1, characterized in that, The MSFEB module employs a multi-scale convolution parallel processing mechanism.
4. The underwater degraded image restoration method based on a consistency model according to claim 1, characterized in that, The LDFSB module employs a serial processing mechanism of depth-separable convolutional module and pixel attention module.
5. The underwater degraded image restoration method based on a consistency model according to claim 1, characterized in that, The parameterization formula for the consistency model is as follows: in, Representing a consistency model, This indicates the parameters of the SDGDN network. This represents the parameters of the LUDEN network after pre-training and freezing. This represents a noisy image. Indicates the boundaries of the time range. This represents the original underwater degraded image. and All are differentiable functions. This represents an intermediate result image. This represents a pre-trained LUDEN network with frozen parameters, used to obtain depth information.
6. The underwater degraded image restoration method based on a consistency model according to claim 1, characterized in that, The trained SDGDN network is obtained through the following steps: Obtain the trained LUDEN network and the untrained SDGDN network, and freeze the parameters of the trained LUDEN network; Based on the isolation training framework of the consistency model, a collaborative denoising network is constructed, which includes a LUDEN network and an SDGDN network. Obtain the ground truth image corresponding to the original underwater degradation image; A training set is constructed based on the original underwater degraded image and the ground truth image corresponding to the original underwater degraded image; The collaborative denoising network is trained based on the training set to obtain the trained SDGDN network.
7. The underwater degraded image restoration method based on a consistency model according to claim 6, characterized in that, When training the collaborative denoising network based on the training set, the loss function used is the consistency constraint loss function, which is expressed as follows: in, Representing a consistency model, Indicates to Adding noise , This represents the model parameters updated using the exponential moving average strategy. Indicates about The weighting function is used to balance the loss weights at different time boundaries. , Indicates the minimum weight. Indicates the maximum weight. Represents the metric function. Indicates the distance from the walk. This represents a discrete interval index.
8. An underwater degraded image restoration system based on a consistency model, characterized in that, include: The data acquisition module is used to acquire the original underwater degraded image and add noise to the original underwater degraded image to obtain a noisy image; The depth information determination module is used to input the original underwater degraded image into the depth estimator to obtain depth information; The method for constructing the depth estimator includes: Underwater scene depth estimation is performed on the original degraded underwater image to obtain underwater depth pseudo-labels; The RRDB dataset was obtained by filtering the pseudo-labels for underwater depth. Based on the RRDB dataset, the LUDEN network is trained to obtain a depth estimator. The LUDEN network includes an encoder network based on the MBConv module and a decoder network based on the DSConvB module and the FFB module. The intermediate result image determination module is used to input depth information, the original underwater degraded image, and the noisy image into the trained SDGDN network to obtain the intermediate result image; The SDGDN network specifically includes: Based on the MSFEB and SDFFB modules, an MSDAB module is constructed. The MSFEB module is used to extract spatial and depth features, and the SDFFB module is used to fuse the extracted spatial and depth features. Based on a depthwise separable convolutional module and a pixel attention module, an LDFSB module is constructed, which is used to extract features from the depth information; SDGDN network is constructed based on the MSDAB and LDFSB modules; The image restoration module is used to input the intermediate result image and the noisy image into the parameterized formula of the trained consistency model to directly obtain the restored underwater degradation image. The consistency model includes the LUDEN network and the SDGDN network.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the underwater degraded image restoration method based on the consistency model as described in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater degraded image restoration method based on the consistency model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A lightweight underwater image enhancement method based on skip-sampling diffusion model
CN119887552B