Interpretable underwater image enhancement method based on physical prior guidance

By constructing a diffusion model based on physical prior guidance, high-frequency information of underwater images is extracted and fused, and combined with cross attention and self-attention mechanisms, the problem of poor underwater image enhancement effect in the prior art is solved, and high-quality underwater image enhancement is achieved.

CN120495107APending Publication Date: 2025-08-15CHANGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510589394.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture and utilize high-frequency details in underwater image enhancement, and fails to effectively combine the physical mechanism of global background light map, transmission map and clear underwater images after defogging, resulting in poor enhancement effect.

Method used

A diffusion model based on physical prior guidance is constructed, including high-frequency prior generation branches and diffusion model enhancement branches that perceive physical information. Multi-scale high-frequency feature maps are extracted through Laplace pyramid decomposition, and a transmission map and global background light map are generated by combining the atmospheric scattering model, and a cross-attention and self-attention mechanism are used to guide the diffusion process.

Benefits of technology

It significantly improves the high-frequency detail recovery ability of underwater images, corrects color distortion, generates high-quality clear underwater images, and enhances the semantic consistency and image quality of feature representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495107A_ABST
    Figure CN120495107A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a physical prior guidance-based interpretable underwater image enhancement method, which comprises the following steps of: constructing a physical prior guidance-based diffusion model which comprises a high-frequency prior generation branch and a physical information perception diffusion model enhancement branch; in the high-frequency prior generation branch, firstly, an input underwater image is decomposed by using a Laplacian pyramid to obtain a multi-scale high-frequency feature map, the multi-scale high-frequency feature map is integrated with an original feature map extracted from the underwater image, and then a transmission map, a global background light map and a defogged clear underwater image are generated by combining an atmospheric scattering model; the diffusion model enhancement branch of physical information perception comprises a cross attention module and a diffusion model, and the diffusion model guides the diffusion process through the cross attention module by using a physical attribute graph to obtain an enhanced image. According to the method, the high-frequency information of the underwater image and the original image features are fused, and the internal relation between different scale features is enhanced, so that the semantic consistency of feature representation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to an interpretable underwater image enhancement method based on physical prior guidance. Background Art

[0002] Underwater clear imaging technology has crucial application value in areas such as marine target detection, underwater robot navigation, and exploration. However, due to the complex and diverse underwater environment, light propagation in water is subject to refraction, scattering, and absorption. Coupled with the inherent low illumination conditions underwater, images captured by underwater imaging systems often exhibit severe degradation. These degradations manifest themselves primarily as blurred details, reduced contrast, color distortion, and increased noise, which not only severely impact visual quality but also significantly hinder the successful execution of subsequent underwater missions. Underwater image enhancement technology aims to improve image clarity, contrast, and color fidelity by compensating for or suppressing effects such as scattering and absorption during underwater light propagation, while also reducing noise in the image. Therefore, research on underwater image enhancement technology has important theoretical significance and practical application value.

[0003] In recent years, deep learning-based underwater image enhancement methods have made significant progress. These methods typically rely on a large amount of paired training data (i.e., degraded underwater images and their corresponding clear reference images) to learn the mapping relationship between the two and demonstrate excellent performance in subsequent image enhancement tasks. However, existing deep learning-based methods still face several key challenges:

[0004] (1) Underwater image enhancement methods based on Generative Adversarial Networks (GANs) sometimes introduce artifacts that do not exist in the training set into the generated images, resulting in a decrease in the quality of the enhanced results. This is mainly due to the fact that existing GANs are not able to capture high-frequency details. At the same time, the training effect depends on the loss balance between the generator and the discriminator, and the image processing ability is poor for images with a lot of noise and obvious underwater degradation.

[0005] (2) Traditional convolutional neural network (CNN)-based methods usually directly learn the nonlinear mapping from degraded images to clear images in the pixel space of the image. However, they fail to extract and utilize the high-frequency features of the image (such as edges and textures), which limits the generation of high-quality underwater images.

[0006] (3) The existing denoising diffusion probabilistic models (DDPM) and denoising diffusion implicit models (DDIM) gradually add noise to the image through a forward diffusion process for model training, and then gradually restore the clear image from the noisy image through a reverse denoising process, providing a new approach for underwater image enhancement. However, such models currently cannot well integrate the physical mechanism of underwater imaging. Specifically, they fail to effectively apply the global background light map, transmission map, and the clear underwater image after defogging, combined with the scattering formula to reconstruct the enhanced image. Therefore, during the image enhancement process, they lack the ability to perceive the difficulty of enhancing different areas of the image, resulting in limited improvement in the details of the enhanced underwater image. Summary of the Invention

[0007] The purpose of the present invention is to provide an interpretable underwater image enhancement method based on physical prior guidance, which is used to address the shortcomings of the existing technology in improving the capture and utilization efficiency of high-frequency details, as well as the technical problems of failing to effectively utilize the physical mechanisms such as the global background light map, the transmission map, the clear underwater image after dehazing, and the scattering formula in the application of the diffusion model, resulting in poor improvement of the details of the enhanced image.

[0008] The interpretable underwater image enhancement method based on physical prior guidance constructs a diffusion model based on physical prior guidance. The diffusion model based on physical prior guidance includes a high-frequency prior generation branch and a physical information-aware diffusion model enhancement branch. In the high-frequency prior generation branch, the high- and low-frequency feature fusion enhancement module first uses a Laplace pyramid to decompose the input underwater image to obtain a multi-scale high-frequency feature map of the underwater image, and integrates the multi-scale high-frequency feature map with the original feature map extracted from the underwater image. Then, three physical attribute maps are generated in combination with the atmospheric scattering model, namely, a transmission map, a global background light map, and a clear underwater image after defogging. The physical information-aware diffusion model enhancement branch adopts a U-Net architecture, including a cross-attention module and a diffusion model. The diffusion model uses the physical attribute map to guide the diffusion process through the cross-attention module to obtain an enhanced image. After training the constructed diffusion model based on physical prior guidance, the trained model is used to enhance the underwater image.

[0009] Preferably, the process of integrating the multi-scale high-frequency feature map and the original feature map by the high- and low-frequency feature fusion enhancement module includes the following steps:

[0010] 1) Perform high-frequency decomposition on the input image;

[0011] 2) Two encoding modules are used to extract corresponding high-frequency features and original features from the multi-scale high-frequency feature map and the input underwater image respectively;

[0012] 3) Adaptively fuse multi-scale high-frequency feature maps and corresponding original features;

[0013] Step 2) obtains the original features of the highest layer, and step 3) obtains the final fusion enhancement result. The original features of the highest layer and the final fusion enhancement result are used to generate three physical attribute maps.

[0014] Preferably, in step 3), the expression of adaptive fusion is:

[0015]

[0016] Among them, the high-frequency features of the i-th layer and original features Obtain high-frequency feature maps of the same dimension through convolution or deconvolution and the original feature map Conv1(·) represents convolution, L represents the Leaky ReLu activation function, Res(·) represents the residual block, and the superscript 2 of the residual block represents that the residual block is stacked twice. is the fused feature of the corresponding i-th layer obtained by adaptive fusion; the fused features of each layer are then fused layer by layer from low to high based on the adaptive fusion method of the above expression to obtain the final fusion enhancement result.

[0017] Preferably, the high-frequency prior generation branch also includes a multi-branch decoding module, which includes two parallel decoding branches; one branch is applied to the final fusion enhancement result: the final fusion enhancement result output by the high- and low-frequency feature fusion enhancement module is subjected to a series of upsampling, and the original features of each layer obtained by registering the downsampling of the original feature map are jump-connected with the upsampling results of the same layer of the final fusion enhancement result, and finally the decoding input features containing high-frequency information are obtained; the other branch is applied to the original features of the highest layer: the original features of the highest layer obtained by downsampling convolution are directly decoded.

[0018] Preferably, the original features of the highest layer are used as low-frequency input features f r , the decoded input features containing high-frequency information are used as high-frequency input features f n , each of the two decoding branches has a residual group, which is input from the low-frequency feature f r and high-frequency input features f n Processing to obtain background light residual features and transmission map residual features Then, the estimated value B of the global background light map and the estimated value T of the penetration map are obtained through processing by the corresponding decoding modules.

[0019] Preferably, the estimated value B of the global background light map and the estimated value T of the penetration map are input into the physical model constraint module, and the estimated value J of the clear underwater image is generated in combination with the atmospheric scattering formula. By calculating the loss between the estimated value J of the clear underwater image and the true result J0 determined in the training set, the optimization of the high-frequency prior generation branch is achieved; at the same time, the predicted value I' of the underwater image before defogging is obtained by calculating the true result J0 and the atmospheric scattering formula, and the loss between the predicted value I' of the underwater image and the input underwater image I is calculated to optimize the high-frequency prior generation branch.

[0020] Preferably, the diffusion model enhancement branch of physical information perception adopts the U-Net architecture. The noise image xt at the t-th sampling time and the guidance condition formed by the physical prior information are spliced in the channel dimension to obtain the estimated feature F of the fused noise. The estimated feature F of the fused noise and the estimated value T of the penetration map are used as input to construct the query matrix Q, key matrix K and value matrix V of the cross attention module. The corresponding expressions are as follows:

[0021] Q = Conv 1×1 (T)

[0022] K = Conv 1×1 (F)

[0023] V=Conv 1×1 (T)

[0024] Then, the cross attention module obtains the output feature T through the following formula out :

[0025]

[0026] Among them, d k Represents the dimension of the K matrix and outputs the feature T out is used to guide the diffusion process.

[0027] Preferably, the physical information perception diffusion model enhancement branch also introduces a physical perception self-attention mechanism, including: first, incorporating the time step t into the estimated feature F of the fused noise:

[0028]

[0029] Among them, Norm is layer normalization, To incorporate the features of step length;

[0030] Features based on integration step size Calculate the new query matrix Q separately f , final bond matrix and the final value matrix

[0031]

[0032] Where W d and W p Represents 1×1 point-by-point convolution and 3×3 depth-separable convolution respectively; at the same time based on the output feature T out Generate another new query matrix Q t ;Q t =W d W p T out ; Fusion of two new query matrices Q f and Q t To obtain the final query matrix The formula is as follows:

[0033]

[0034] Among them, concat means connection, Conv 1×1 Represents 1×1 convolution; the output obtained by the self-attention mechanism module As shown below:

[0035]

[0036] Among them, α is a learnable parameter.

[0037] The present invention has the following advantages: On the one hand, the present invention focuses on extracting and fusing high-frequency information of underwater images through high-frequency prior generation branches, and uses the decoding output results containing high-frequency features and the original features of the highest level in the input image extraction features through a multi-branch decoding module to generate corresponding physical property maps: a transmission map and a global background light map. The transmission map obtained in this way fully extracts and fuses the high-frequency features and effectively utilizes the high-frequency information. Therefore, the obtained transmission map is of high quality, and further high-quality clear underwater images can be obtained based on physical properties. The improvement of the present invention enhances the intrinsic connection between features of different scales by fusing the high-frequency information (such as edges and textures) of underwater images with the original image features, thereby improving the semantic consistency of feature representation; it overcomes the technical problems of the generative adversarial network properties in the prior art and the artifacts and poor enhancement effects caused by the inability to effectively utilize high-frequency information.

[0038] This method, on the other hand, introduces a physical property map containing physical prior information into a diffusion model based on a U-net architecture through cross-attention and self-attention mechanisms. This physical prior information provided by the high-frequency prior generation branch is used to constrain and guide the diffusion model's inverse denoising process, thereby achieving more accurate and physically accurate underwater image enhancement. The self-attention mechanism models the transmission map's features, guiding the diffusion model to more effectively restore image detail, suppress noise, and correct color distortion. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of the model of the interpretable underwater image enhancement method guided by physical priors in the present invention.

[0040] Figure 2 It is a schematic diagram of the high- and low-frequency feature fusion enhancement module in the present invention. DETAILED DESCRIPTION

[0041] The specific implementation methods of the present invention will be further explained in detail below through the description of embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0042] like Figure 1-Figure 2As shown, the present invention provides an interpretable underwater image enhancement method based on physical prior guidance, constructs a diffusion model guided by physical prior (Physics-Guided Diffusion Framework, PGDF), and the diffusion model guided by physical prior includes a high-frequency prior generation (HG) branch and an enhancement branch of diffusion model for physical information perception (PPDF) branch. Among them, in the high-frequency prior generation (HG) branch, the high- and low-frequency feature fusion enhancement module first uses the Laplacian pyramid to decompose the input underwater image to obtain a multi-scale high-frequency feature map of the underwater image, and integrates the multi-scale high-frequency feature map with the original feature map extracted from the underwater image. Then, it combines the atmospheric scattering model (ASM) to generate three physical attribute maps, namely the transmission map, the global background light map, and the clear underwater image after defogging. The physical information perception diffusion model enhancement branch adopts the U-Net architecture, including the Cross-Attention Module (CAM) and the Diffusion Module (DM). The diffusion model uses the physical property graph (i.e., the physical prior features generated by the high-frequency prior generation branch) through the cross-attention module to guide the diffusion process, introduces physical prior information, cross-attention mechanism and self-attention mechanism to process the image, and obtains an enhanced image.

[0043] Existing underwater image enhancement methods typically achieve image enhancement by fusing high- and low-frequency information extracted from a network to compensate for information loss caused by water scattering and absorption. However, these methods primarily rely on indirect high-frequency features generated during the network encoding and decoding process. These indirect high-frequency features tend to decay with increasing network depth, making it difficult to effectively recover high-frequency details in the image, such as edges and textures.

[0044] Therefore, the high-frequency prior generation (HG) branch in the present invention includes: a high-and low-frequency feature fusion enhancement module (High-and Low-frequency Feature Enhancement, HFE) and a multi-branch decoding module (Multi-branch Decoding Module, MDM).

[0045] High- and low-frequency feature fusion enhancement module: This module simultaneously encodes the input original underwater image and its corresponding multi-scale high-frequency feature map. By explicitly introducing high-frequency features, it can extract a richer and more direct multi-scale high-frequency feature map. The extracted multi-scale high-frequency feature map is then adaptively fused with the original feature map (i.e., the encoded features of the input underwater image), effectively restoring the image's fine structure and texture details, improving the overall image restoration quality. This direct and effective high-frequency feature compensation mechanism significantly enhances underwater image restoration performance.

[0046] Specifically, the processing steps of the high- and low-frequency feature fusion enhancement module are as follows:

[0047] 1) Perform high-frequency decomposition on the input image, the expression is: I h =ψ(I), where ψ(·) represents a high-frequency decoding method. In the embodiment, the Laplacian pyramid (LP) method is used to extract high-frequency features. is the input underwater image, It is a multi-scale high-frequency feature map obtained by Laplacian pyramid decoding.

[0048] 2) Two encoding modules are used to extract the corresponding high-frequency features and original features from the multi-scale high-frequency feature map and the input underwater image respectively. The expression is: Among them, E h (·) represents the extraction of high-frequency features from multi-scale high-frequency feature maps through the corresponding encoding module; E r (·) indicates that the corresponding encoding module extracts original features from the input underwater image; and The high-frequency features and original features of the i-th layer are respectively high-frequency features and original features.

[0049] 3) Adaptively fuse the multi-scale high-frequency feature map and the corresponding original features, and the expression is: Among them, the high-frequency features of the i-th layer and original features Obtain high-frequency feature maps of the same dimension through convolution or deconvolution and the original feature map These are the two features being fused. Conv1(·) represents convolution, L represents the Leaky ReLu activation function, Res(·) represents the residual block, and the superscript 2 of the residual block represents that the residual block is stacked twice. is the fused feature of the corresponding layer i obtained by adaptive fusion. The fused features of each layer are then fused layer by layer from low to high based on the adaptive fusion method of the above expression to obtain the final fusion enhancement result.

[0050] In the embodiment, f1, f2, and f3 are used to represent the fused features of the first, second, and third layers respectively. The above-mentioned feature fusion method is used to first fuse the fused features of the first and second layers to obtain the secondary fused feature F1, and then the secondary fused feature F1 is fused with the fused feature f3 of the third layer to obtain the final fused enhancement result F. eh .

[0051] This fusion method can more effectively utilize high-frequency features and enhance the texture details of the reconstructed image. rgb and F hf Represents two encoding modules E at different layers r and E h The output result (including high-frequency features and original features of each layer) is used as the input of the high- and low-frequency feature fusion enhancement module. The expression of the fusion process of the high- and low-frequency feature fusion enhancement module is: eh =HFE(F rgb , F hf ), HFE(·) represents the processing of high- and low-frequency feature fusion enhancement module, F eh Represents the final fusion enhancement result.

[0052] Multi-branch decoding module: This module contains two parallel decoding branches, which aims to recover the three key components of the underwater image from the results processed by the HFE module: the global background light map B, the transmission map T, and the dehazed underwater image J.

[0053] One branch is applied to the final fusion enhancement result: the final fusion enhancement result F output by the high and low frequency feature fusion enhancement module eh After a series of upsampling, the original features of each layer obtained by registering and downsampling the original feature map are jump-connected with the upsampling results of the final fusion enhancement result at the same layer, and finally the decoding input features containing high-frequency information are obtained.

[0054] The other branch is applied to the original features of the highest layer: for the original features of the highest layer obtained by downsampling convolution Decoding is then performed directly, where k represents the maximum value of i, which is 3 in this embodiment.

[0055] Due to the original characteristics of the top layer It does not have high-frequency features, but the decoding input features containing high-frequency information are integrated with high-frequency features. Therefore, for the multi-branch decoding module, the former is a low-frequency input feature f r , the latter is the high-frequency input feature f n .

[0056] The two decoding branches of the multi-branch decoding module each have one residual group, and both residual groups contain six residual blocks. rand high-frequency input features f n The branch features of different branches are extracted by processing the residual groups of the corresponding decoding branches. Specifically, the corresponding expressions are expressed as:

[0057]

[0058]

[0059] Among them, RES B Represents the residual group of the decoding branch corresponding to the global background light map B, RES T represents the residual group of the decoding branch corresponding to the transmission map T; represents the background light residual feature obtained by the residual group processing in the decoding branch corresponding to the global background light map B, Represents the residual features of the transmission map obtained by processing the residual group in the decoding branch corresponding to the transmission map T.

[0060] After that, the background light residual features and the transmission map residual features are processed by the decoding modules of different decoding branches respectively; the expressions for calculating the global background light map and transmission map through the decoding module are:

[0061]

[0062] Among them, D B Denotes the decoding module for estimating the global background light map, D T Represents the decoding module of the estimated transmission map. In this way, in the multi-branch decoding module, the low-frequency input feature f r and high-frequency input features f n The corresponding global background light map estimation value B and the transmission map estimation value T are processed in the two branches of the multi-branch decoding module (MDM) respectively.

[0063] Physical Model Constraint Module (PM): The high-frequency prior branch also includes a physical model constraint module. The estimated value B of the global background light map and the estimated value T of the transmission map are input into the physical model constraint module. In the physical model constraint module, the atmospheric scattering formula is rewritten as follows:

[0064]

[0065] Where I is the input underwater image, and J is the estimated value of the clear underwater image, which is calculated based on the estimated value B of the global background light map and the estimated value T of the penetration map. By calculating the loss between the estimated value J of the clear underwater image and the true result J0 determined in the training set, the high-frequency prior generation branch is optimized. At the same time, the calculation formula for the underwater image prediction value I' before dehazing after the atmospheric scattering formula is:

[0066] I'=J0T+(1-T)B

[0067] Here, I' is the predicted value of the underwater image before dehazing, calculated using the ground truth result J0 and the atmospheric scattering formula. The loss between the predicted underwater image value I' and the input underwater image I is further calculated to optimize the high-frequency prior generation branch. After optimization, the estimated values of the global background light map B and the estimated value of the transmission map T, as well as the further estimated value of the clear underwater image J, serve as the input to the physical information-aware diffusion model enhancement (PPDF) branch.

[0068] The diffusion model enhancement branch for physical information perception in this paper adopts a U-Net architecture and includes a Cross-Attention Module (CAM) and a Diffusion Module (DM). The Diffusion Module is built based on the Denoising Diffusion Probabilistic Model (DDPM) and consists of two main stages: forward diffusion and backward diffusion.

[0069] Forward Diffusion Process: The forward diffusion process follows the properties of a Markov chain and gradually adds Gaussian noise to the clear image until the image is completely degraded into a pure noise image. Given a clear underwater image x0, Gaussian noise is gradually introduced into the clear underwater image x0 based on the time step t. The mathematical expression of this process is as follows:

[0070]

[0071] Among them, β t is a variable used to control the size of the noise added, q(x t |x t-1 ) represents the single-step transition probability of the noise-adding process, Represents Gaussian distribution. By setting α t =1-β t , this process can be defined as:

[0072]

[0073] Combining the Gaussian distribution, we can get:

[0074]

[0075] Backward Diffusion Process: The backward diffusion process aims to recover clear data from Gaussian noise images. The backward diffusion process can be expressed as:

[0076]

[0077] Among them, p θ (xt-1 |x t , x c ) represents the probability of neural network fitting, x c represents the input underwater image, μ θ (x t , x c , t) and The above diffusion process is the basic principle of the diffusion module in the prior art, so it will not be described in detail.

[0078] In this approach, the physical information-aware diffusion model enhancement branch first generates a noise image xt corresponding to sampling time step t based on the forward process of the diffusion model. The estimated value J of the clear underwater image, the estimated value B of the global background light map, and the estimated value T of the penetration map output by the high-frequency prior generation branch are used as guidance conditions. These guidance conditions are then concatenated with the noise image xt at sampling time t in the channel dimension. After convolutional fusion, the estimated features F of the fused noise are obtained.

[0079] Cross-Attention Module: This module uses the physical prior features of the input underwater image to guide the diffusion process. Specifically, the estimated features F of the fused noise and the estimated value T of the penetration map are taken as input. This method uses different 1x1 convolution kernels to construct the query matrix Q, key matrix K, and value matrix V. The corresponding expressions are as follows:

[0080] Q = Conv 1×1 (T)

[0081] K = Conv 1×1 (F)

[0082] V=Conv 1×1 (T)

[0083] Then, the cross attention module obtains the output feature T through the following formula out :

[0084]

[0085] Among them, d k Represents the dimension of the K matrix.

[0086] Next, a physically aware self-attention mechanism is introduced into the diffusion module DM. Specifically, the time step t is incorporated into the estimated feature F of the fused noise:

[0087]

[0088] Among them, Norm is layer normalization, To incorporate the features of the step length.

[0089] Cross-Attention Module: Based on the feature of incorporating stride length Calculate the query matrix separately key Sum Where W a and W p Represents 1×1 point-by-point convolution and 3×3 depth-separable convolution respectively. At the same time, the output feature T based on the cross attention mechanism out Generate another query Q t (Q t =W d W p T out ). The present invention applies the attention strategy of physical perception and finally integrates Q f and Q t To obtain the final query matrix The formula is as follows:

[0090]

[0091] Among them, concat means connection, Conv 1×1 represents 1×1 convolution, represents the final query matrix.

[0092] Finally, the output obtained by the self-attention mechanism module is As shown below:

[0093]

[0094] Among them, α is a learnable parameter.

[0095] The above method introduces physical prior information, cross-attention mechanism and self-attention mechanism into the diffusion module DM, which can fully utilize the physical information at the feature level and establish the association between features and transmission maps, thereby effectively helping the diffusion model to recover the lost detail features in the reverse process and correct the loss of image color, ultimately obtaining clearer underwater images.

[0096] After training the constructed diffusion model based on physical prior guidance, the trained model is used to enhance underwater images.

[0097] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.

Claims

1. An interpretable underwater image enhancement method based on physical prior guidance, characterized by: A diffusion model guided by physical priors is constructed. This model consists of a high-frequency prior generation branch and a physical information-aware diffusion model enhancement branch. Within the high-frequency prior generation branch, the high- and low-frequency feature fusion enhancement module first decomposes the input underwater image using a Laplacian pyramid to obtain a multi-scale high-frequency feature map of the underwater image. This multi-scale high-frequency feature map is then integrated with the original feature map extracted from the underwater image. Finally, three physical property maps, namely a transmission map, a global background light map, and a clear underwater image after dehazing, are generated by combining this with an atmospheric scattering model. The diffusion model enhancement branch of physical information perception includes a cross-attention module and a diffusion model. The diffusion model uses the physical property graph to guide the diffusion process through the cross-attention module to obtain an enhanced image. After training the constructed diffusion model based on physical prior guidance, the trained model is used to enhance underwater images.

2. The interpretable underwater image enhancement method based on physical prior guidance according to claim 1, characterized in that: The process of integrating multi-scale high-frequency feature maps and original feature maps in the high- and low-frequency feature fusion enhancement module includes the following steps: 1) Perform high-frequency decomposition on the input image; 2) Two encoding modules are used to extract corresponding high-frequency features and original features from the multi-scale high-frequency feature map and the input underwater image respectively; 3) Adaptively fuse multi-scale high-frequency feature maps and corresponding original features; Step 2) obtains the original features of the highest layer, and step 3) obtains the final fusion enhancement result. The original features of the highest layer and the final fusion enhancement result are used to generate three physical attribute maps.

3. The interpretable underwater image enhancement method based on physical prior guidance according to claim 2, characterized in that: In step 3), the expression of adaptive fusion is: Among them, the high-frequency features of the i-th layer and original features Obtain high-frequency feature maps of the same dimension through convolution or deconvolution and the original feature map Conv1(·) represents convolution, L represents the Leaky ReLu activation function, Res(·) represents the residual block, and the superscript 2 of the residual block represents that the residual block is stacked twice. is the fused feature of the corresponding i-th layer obtained by adaptive fusion; the fused features of each layer are then fused layer by layer from low to high based on the adaptive fusion method of the above expression to obtain the final fusion enhancement result.

4. The interpretable underwater image enhancement method based on physical prior guidance according to claim 3, characterized in that: The high-frequency prior generation branch also includes a multi-branch decoding module, which includes two parallel decoding branches; One branch is applied to the final fusion enhancement result: the final fusion enhancement result output by the high- and low-frequency feature fusion enhancement module undergoes a series of upsampling, and the original features of each layer obtained by registering the downsampling of the original feature map are jump-connected with the upsampling results of the same layer of the final fusion enhancement result, and finally the decoding input features containing high-frequency information are obtained; the other branch is applied to the original features of the highest layer: the original features of the highest layer obtained by downsampling convolution are directly decoded.

5. The interpretable underwater image enhancement method based on physical prior guidance according to claim 4, characterized in that: The original features of the highest layer are used as low-frequency input features f r , the decoded input features containing high-frequency information are used as high-frequency input features f n , each of the two decoding branches has a residual group, which is input from the low-frequency feature f r and high-frequency input features f n Processing to obtain background light residual features and transmission map residual features Then, the estimated value B of the global background light map and the estimated value T of the penetration map are obtained through processing by the corresponding decoding modules.

6. The interpretable underwater image enhancement method based on physical prior guidance according to claim 5, characterized in that: The estimated value B of the global background light map and the estimated value T of the penetration map are input into the physical model constraint module. Combined with the atmospheric scattering formula, the estimated value J of the clear underwater image is generated. By calculating the loss between the estimated value J of the clear underwater image and the true result J0 determined in the training set, the high-frequency prior generation branch is optimized. At the same time, the predicted value I' of the underwater image before dehazing is calculated using the true result J0 and the atmospheric scattering formula. The loss between the predicted value I' and the input underwater image I is calculated to optimize the high-frequency prior generation branch.

7. The interpretable underwater image enhancement method based on physical prior guidance according to claim 1, characterized in that: The diffusion model enhancement branch of physical information perception adopts the U-Net architecture. The noise image xt at the t-th sampling time and the guidance condition formed by the physical prior information are spliced in the channel dimension to obtain the estimated feature F of the fused noise. The estimated feature F of the fused noise and the estimated value T of the penetration map are used as input to construct the query matrix Q, key matrix K and value matrix V of the cross attention module. The corresponding expressions are as follows: Q=Conv 1×1 (T) K=Conv 1×1 (F) V=Conv 1×1 (T) Then, the cross attention module obtains the output feature T through the following formula out : Among them, d k Represents the dimension of the K matrix and outputs the feature T out is used to guide the diffusion process.

8. The interpretable underwater image enhancement method based on physical prior guidance according to claim 7, characterized in that: The diffusion model enhancement branch of physical information perception also introduces a physical perception self-attention mechanism, including: first, incorporating the time step t into the estimated feature F of the fused noise: Among them, Norm is layer normalization, To incorporate the features of step length; Features based on integration step size Calculate the new query matrix Q separately f , final bond matrix and the final value matrix Where W d and W p Represents 1×1 point-by-point convolution and 3×3 depth-separable convolution respectively; at the same time based on the output feature T out Generate another new query matrix Q t ;Q t =W d W p T out ; Fusion of two new query matrices Q f and Q t To obtain the final query matrix The formula is as follows: Among them, concat means connection, Conv 1×1 Represents 1×1 convolution; the output obtained by the self-attention mechanism module As shown below: Among them, α is a learnable parameter.

Citation Information

Cited By

  • Fusion segmentation AI product graph enhancement method

    CN121280297A