Underwater vision sharpening method based on polarization information guide diffusion model
By introducing polarization information and diffusion models into the underwater image enhancement technology, an underwater visual clarification network including conditional diffusion processes, cross-modal enhancement networks and conditional denoising networks have been constructed, which has solved the problem of improving underwater image quality in the existing technology and achieved more efficient image clarification and detail recovery effects.
Patent Information
- Application Number
- CN202510263058.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-01
AI Technical Summary
Existing underwater image enhancement technologies are difficult to effectively improve the sharpness, contrast and color accuracy of images in complex underwater environments, and are prone to excessive enhancement and blurred details.
Using a diffusion model guided by polarization information, a diffusion model is constructed and an underwater visual clarification network containing conditional diffusion processes, cross-modal enhancement networks and conditional denoising networks are designed, and the polarization information is used to compensate and denoise RGB images to improve the quality of the image.
It significantly improves the clarity, contrast and color accuracy of underwater images, reduces interference to subsequent computer vision tasks, and improves the visual quality and processing capabilities of the images.
Smart Images

Figure CN120235780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of marine environment perception, digital image processing, and image enhancement, and specifically relates to an underwater vision clarification method based on a polarization information-guided diffusion model. Background Art
[0002] Approximately 70% of the Earth's surface is covered by the ocean, which contains unknown biological species and huge energy resources and plays a crucial role in the continuation of life on Earth. With the development of human exploration of the ocean, underwater vision has become an important tool for obtaining seabed information. Underwater imaging vision technology has been widely and deeply applied in many fields such as seabed topography detection, marine biology monitoring, underwater resource exploration, and underwater archaeology, and has attracted much attention because of its ability to carry high information density. However, due to the strong absorption and scattering of light by the water medium and the complex underwater environment, the propagation distance of natural light in the water body is very limited and shows exponential decay, resulting in a decrease in the clarity and contrast of the image; secondly, research shows that the water body has different absorption degrees for light of different wavelengths, and only light of specific wavelengths can propagate in the water body at a lower attenuation rate, thus causing color distortion in underwater images; in addition, substances such as particles contained in the water body will not only make the underwater image have large noise, but also absorb and scatter light. The scattering effect is mainly divided into forward scattering and backward scattering. The former causes the light to deviate from the original direction, resulting in a decrease in image resolution; the latter will transmit the particle information in the water body to the detector, resulting in a decrease in image contrast. Therefore, an excellent underwater image restoration method is expected to improve the contrast of underwater images, eliminate color offsets in the images, and enhance their details, thereby effectively improving the visual quality of the images.
[0003] Currently, the main method for underwater image enhancement relies on the RGB color gamut. However, due to factors such as light scattering, absorption, and color distortion, RGB images are limited in terms of reliability. Traditional imaging detection techniques mainly operate by obtaining the intensity information of the target. However, when the intensity difference between targets is small, it becomes difficult to distinguish between them using intensity information. Especially in a scattering medium, due to the strong scattering and absorption of light by the scattering medium, the intensity of the target's reflected light gradually decreases as the detection distance increases, while the backscattered light caused by direct scattering continuously increases as the detection distance increases. As a result, the acquired images exhibit characteristics such as poor quality, low contrast, and the drowning of target detail information. It is difficult to effectively detect targets in a timely manner using traditional intensity imaging techniques. To overcome these problems, polarization optical imaging technology, as a new imaging method, has received extensive attention and research. Polarization optical imaging technology detects the target by obtaining the polarization information of the target, and promotes underwater image enhancement by integrating linear polarization imaging (described by the angle of linear polarization (AoLP) and the degree of linear polarization (DoLP)). According to the inconsistency of the polarization states of the target's reflected light and the backscattered light, the target's reflected light and the backscattered light are separated, reducing the impact of scattered light on target imaging and showing stronger ability to highlight the target in a complex background, demonstrating its unique advantages different from intensity imaging. However, due to the lack of underwater scenes and high-quality images, underwater image enhancement (UIE) faces various challenges, such as over-enhancement and blurred detail features. These problems limit the performance of UIE methods and lead to poor performance in downstream tasks. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides an underwater vision clarification method based on a polarization information-guided diffusion model. Based on the idea of multi-modal learning, polarization information is introduced as an additional modality to enhance the original underwater image. The polarization information-guided diffusion network model used in this method can use linear polarization as an additional clue to enhance the clarity, contrast, and color accuracy of the image under different underwater conditions. The degree of polarization information and the angle of polarization information are used to enhance the contrast and texture details of different regions in the image, thereby more accurately restoring the true color of the image, reducing the serious interference it brings to subsequent computer vision tasks, and significantly improving the quality of underwater images.
[0005] The technical means adopted by the present invention are as follows:
[0006] An underwater vision clarification method based on a polarization information-guided diffusion model, comprising the following steps:
[0007] S1. Obtain underwater polarization images under natural conditions and laboratory conditions respectively, and construct an underwater polarization image dataset. The construction of the underwater polarization image dataset includes underwater polarization images, degree of polarization images, and angle of polarization images at different angles.
[0008] S2. Construct an underwater vision clarification network based on a polarization information-guided diffusion model. The underwater vision clarification network includes a serial conditional diffusion process, a cross-modal enhancement network, and a conditional denoising network. The conditional diffusion process is used to simulate the degradation process of the image underwater and estimate the noise information. The cross-modal enhancement network is used to compensate the RGB image information with the polarization image information to generate a pre-enhanced result, and update the polarization information with the pre-enhanced result to produce a fused feature. The conditional denoising network is used to predict the noise information at the current time information based on the noise, the fused feature, and the time information, and obtain a clear underwater image using the diffusion reverse process and the noise information at the current time information.
[0009] S3. Train the underwater vision clarification network using the underwater polarization image dataset, and output a clear underwater image based on the trained underwater vision clarification network.
[0010] Further, the steps of obtaining underwater polarization images under laboratory conditions include:
[0011] By making different-color water bodies with different turbidity levels in a laboratory scene, using a polarization camera to collect the turbid underwater polarization images of objects in different-color water bodies with different turbidity levels, collecting the clear underwater polarization images of objects in pure water, and using the clear underwater polarization images as label images. Construct a training set and a test set according to the turbid underwater polarization images and the label images.
[0012] The steps of obtaining underwater polarization images in a real-world scene include: making polarization images of different water body environments and different underwater depth environments as a test set.
[0013] Further, the degree of polarization image and the angle of polarization image are obtained according to the following method:
[0014]
[0015] Among them, DoLP represents the degree of polarization image, AoLP represents the angle of polarization image, S0 represents the total light intensity of the polarization image, S1 represents the ratio of 0° linear polarization to vertical polarization, and S2 represents the ratio of 45° linear polarization to vertical polarization.
[0016] Further, the conditional diffusion process adopts a mean restoration diffusion framework based on SDE, including:
[0017] Forward process: Use the SDE to diffuse the image into a degraded image with a noise distribution;
[0018] Backward process: Generate samples by learning and simulating the reverse SDE, thereby training a neural network to estimate the score function of the noise data distribution.
[0019] Furthermore, the forward process is described as follows:
[0020] dx = θ t (μ - x)dt + σ t dw
[0021] where θ t and σ t are two time-varying parameters, λ is a fixed noise level, dx represents the sampling state varying with time, and w represents the standard Wiener process;
[0022]
[0023] where p t (x) represents the Gaussian probability distribution, ▽ x logp t (x) represents the ground truth score in the inference stage, represents the reverse Wiener process.
[0024] Furthermore, the cross-modal enhancement network includes:
[0025] Pre-enhancement branch: Encode the polarization modality and RGB modality images through two convolutional encoders to obtain the polarization feature F PoL and the RGB feature F RGB , merge F PoL and F RGB and send them into a general attention mechanism. Through self-attention and cross-attention schemes, obtain a preliminary enhanced feature;
[0026] Fusion branch: Use the preliminary enhanced feature as spatial guidance to fuse the polarization information, then extract local features from the fused information using residual learning, and subsequently use the preliminary enhanced feature as guidance to compensate the polarization information, and obtain the aggregated feature through spatial attention operation.
[0027] Furthermore, the process of obtaining the preliminary enhanced feature is described as follows:
[0028] CA = σ(MLP(AvgPool(F PoL , F RGB ) + MLP(MaxPool(F PoL , F RGB )))
[0029] SA = σ(f 7×7 ([AvgPool(CA); MaxPool(CA)]))
[0030] where σ is the sigmoid operation, and f 7×7 is a convolutional kernel of size 7×7, and the output after the SA operation is the preliminarily enhanced feature;
[0031] The feature aggregation process of the fusion branch is described as:
[0032] Fuse pol = H RDB (PixelUnshuffle(Concat(DoLP, AoLP)))
[0033] Pol = MLP(AvgPool(Fuse pol + I*)) × (Fuse pol + I*)
[0034] where I* represents the preliminarily enhanced feature, Pol represents the aggregated feature, DoLP represents the degree of polarization image, AoLP represents the angle of polarization image, and H RDB represents the dense residual connection block.
[0035] Furthermore, the process of the conditional denoising network using NAFNet to output a clear image is described as:
[0036]
[0037] F3 = Conv 1×1 (SCA(SimpleGate(DEConv 3×3 (F2))))
[0038] F2 = Conv 1×1 (γ1 ⊙ Norm(F1) + β1)
[0039] F1 = Conv 1×1 (EME(Pol) ⊙ Norm(X) + EME(Pol))
[0040] where Pol represents the aggregated feature, X represents the result of adding noise to the input, γ1 represents the first feature scaling parameter, β1 represents the first feature shift parameter, γ2 represents the second feature scaling parameter, and β2 represents the second feature shift parameter.
[0041] Compared with the prior art, the present invention has the following advantages:
[0042] 1. The present invention constructs a dataset containing underwater polarization images, polarization degree images, and polarization angle images at different angles, providing rich training samples for the underwater vision clarification network, thereby realizing the adaptive modeling of complex underwater environments, and further improving the network's processing ability for underwater images with high turbidity.
[0043] 2. The present invention simulates the degradation process of images underwater through a conditional diffusion process and estimates noise information, thereby realizing the accurate prediction of underwater image noise, and further improving the denoising effect.
[0044] 3. The present invention uses the cross-modal enhancement network to compensate the RGB image information with polarization image information and update the polarization information to generate fused features, thereby realizing the effective restoration of turbid regions, and further improving the color restoration and detail clarity of underwater images.
[0045] 4. The present invention combines noise, fused features, and time information through a conditional denoising network to predict and utilize the diffusion reverse process to obtain clear underwater images, thereby realizing efficient and stable image clarification, and further improving the effectiveness and robustness of the entire method, laying a solid theoretical and technical foundation for subsequent vision tasks such as underwater panoramic observation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic flowchart of an underwater vision clarification method based on a polarization information-guided diffusion model in an embodiment of the present invention.
[0048] Figure 2 It is an architecture diagram of an underwater vision clarification network based on a polarization information-guided diffusion model in an embodiment of the present invention.
[0049] Figure 3 They are low-quality images, linear polarization angle images, and linear polarization degree images obtained by an underwater vision clarification network based on a polarization information-guided diffusion model in an embodiment of the present invention.
[0050] Figure 4 It is a clear underwater image output by an underwater vision clarification network based on a polarization information-guided diffusion model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0052] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0053] As Figure 1 shown, the present invention provides an underwater vision clarification method based on a polarization information-guided diffusion model, which mainly includes the following steps:
[0054] S1. Underwater polarization images are respectively acquired under natural conditions and laboratory conditions to construct an underwater polarization image dataset, and the construction of the underwater polarization image dataset includes underwater polarization images, polarization degree images, and polarization angle images at different angles.
[0055] Specifically, the dataset includes two parts: indoor experimental data and real underwater data. A polarization camera (LUCID, TRI050S) is used, and the camera is equipped with a Sony IMX250MZR CMOS polarization sensor for data acquisition. This enables four images with different polarizations to be obtained in one shot in a scene with obvious dynamic changes. The camera performs real-time processing through the filters in these four directions and outputs the light intensity and polarization angle information of each pixel. In one shot, the camera can simultaneously capture four polarization images, which are respectively denoted as I 0° 、I 45° 、I 90° and I 135° , and the spatial resolution is 1224×1024.
[0056] The indoor experimental data is an underwater scene photographed in a glass tank. During the experiment, a clear ground truth image I gt, then gradually add sand or milk to the ground truth environment to simulate different turbidity conditions, for example: in each scene, add 3-5ml of milk to the water tank multiple times to simulate the turbidity under different conditions. Secondly, add filters of different colors in front of the imaging lens or add inks of different colors to the water body to simulate the color cast caused by light absorption in different depths of the sea. The indoor dataset has a total of 445 sets of data, including 97 sets of mild turbidity, 134 sets of moderate turbidity, 164 sets of severe turbidity, and 50 sets that only show color cast. It contains images of the same scene at different turbidities, as well as images of different scenes at different turbidities.
[0057] The real underwater data was collected by an underwater robot platform in the real offshore environment and the real diving lake environment. There is no ground truth map for the real underwater data, so it is only used to test the performance of our model. The real underwater data set has a total of 200 sets of data, including 100 sets of offshore environments and 100 sets of diving lake environments.
[0058] Since each pixel of the image captured by the polarization camera is represented by four pixels, and each pixel corresponds to the light intensity at four different polarization angles. The polarization state of light can be described by the Stokes vector S = [S0, S1, S2, S3], where S0 represents the total light intensity, S1 and S2 describe the ratio of 0° / 45° linear polarization to vertical polarization, and S3 is the circular polarization ratio. The Stokes elements S0, S1, and S2 can be expressed based on the measured value I 0° ,I 45° ,I 90° and I 135° The calculations show that:
[0059] S0=(I 0° +I 90° +I 45° +I 135° ) / 4 (1)
[0060] S1=(I 0° -I 90° ) / twenty two)
[0061] S2=(I 45° -I 135° ) / twenty three)
[0062]
[0063] In order to generate the ground truth image, this application introduces Stokes parameters. According to this physical model, the first Stoke parameter S0 describes the total intensity of the light beam and can be calculated by the formula. Then, we can calculate a ground truth I gt ,Right now
[0064]
[0065] Among them, I 0° 、I 45° 、I 90° and I 135。 are polarization images at different angles. G(-) is a gamma correction function, and we empirically set the gamma value γ to 2.2.
[0066] S2. Construct an underwater vision clarity network based on a polarization information-guided diffusion model. The underwater vision clarity network includes a serial conditional diffusion process, a cross-modal enhancement network, and a conditional denoising network. The conditional diffusion process is used to simulate the degradation process of the image underwater and estimate the noise information. The cross-modal enhancement network is used to compensate the RGB image information with the polarization image information to generate a pre-enhanced result, and update the polarization information with the pre-enhanced result to produce a fused feature. The conditional denoising network is used to predict the noise information at the current time based on the noise, the fused feature, and the time information, and obtain a clear underwater image using the diffusion reverse process and the noise information at the current time.
[0067] Specifically, the underwater vision clarity model network architecture based on the polarization information-guided diffusion model consists of three key modules:
[0068] (1) Conditional diffusion process. An SDE-based mean restoration diffusion framework that is widely recognized is adopted. This framework consists of two stages:
[0069] 1) Forward process: Use SDE to diffuse the image into a degraded image with a noise distribution;
[0070] dx = θ t (μ - x)dt + σ t dw (7)
[0071] 2) Reverse process: Generate samples by learning and simulating the corresponding reverse SDE. The goal is to train a neural network to estimate the score function of the noise data distribution.
[0072]
[0073] (2) Cross-modal enhancement network. RGB provides basic color and texture features, while polarization cues are not affected by scattering. Therefore, these two modalities can be enhanced in a complementary way to improve the clarity and fidelity of underwater image enhancement. An auto-domain and cross-domain attention scheme is adopted to enable one modality to contribute to the feature expression of the other modality.
[0074] (3) Conditional denoising network. The polarization-guided diffusion process is utilized to focus on introducing polarization cues into the reverse diffusion process in a simple and efficient manner, thereby enhancing the network's generation ability.
[0075] S3. Use the underwater polarization image dataset to train the underwater vision clarity network, and output clear underwater images based on the trained underwater vision clarity network.
[0076] Training the conditional diffusion process based on the underwater polarization image dataset includes:
[0077] S311. The forward process converts the clear image x0 into a degraded image x T . Specifically, the forward diffusion process is described as follows:
[0078] dx = θ t (μ - x)dt + σ t dw (9)
[0079] where θ t and σ t are two time-varying parameters that respectively control the speed of mean regression and random fluctuation. By setting where λ is a fixed noise level. Given x0 and t ∈ [0, T], for the state x corresponding to the intermediate time t T can be strictly represented by the closed-form solution of (9):
[0080]
[0081] In this case, x T follows a Gaussian probability distribution p t (x), and the expression is as follows: where is the mean, is the variance. It can be found that as t increases, the mean m t and the variance v t converge to μ and λ respectively 2 . Therefore, x T can be approximated as a combination of a degraded image and pure Gaussian noise.
[0082] S312. The purpose of the reverse process is to recover the HQ image from the noisy low-quality LQ image. Its reverse diffusion process is expressed as:
[0083]
[0084] where, represents the reverse-time Wiener process. ▽ x logp t(x) is the ground truth score in the inference stage. Since the HQ image is available during training, the ground truth score can be calculated.
[0085]
[0086] In addition, reparameterize x T as where is standard Gaussian noise. Then Since m t (x) and v t are known, it is only necessary to use the conditional denoising network f ∈ to estimate the noise.
[0087] Furthermore, use the underwater polarization image dataset to train the cross-modal enhancement network to obtain the output of the cross-modal enhancement network: the initial enhancement output and the fused polarization features. This initial enhancement output captures the low-frequency and structural information of the final inpainted image. On the other hand, use this initial enhancement output as spatial guidance to fuse the polarization information, so as to obtain the final guidance information to better control and guide the denoising process of the diffusion model, ensuring the accuracy and effectiveness of the denoising direction; the cross-modal enhancement network consists of two branches: the pre-enhancement branch and the fusion branch. The specific steps include:
[0088] S401. Pre-enhancement branch. First, encode each modality through two convolutional encoders to obtain F PoL and F RGB , merge them and send them into a general attention mechanism. Through self-attention and cross-attention schemes, one modality can contribute to the feature representation of the other modality. Simply directly fusing different modalities may lead to the loss or conflict of important information. Specifically, channel and spatial attention mechanisms are adopted to perform preliminary enhancement on the RGB modality: use polarization information to enhance the low-quality image, aggregate the different information contained in the feature maps of DoLP, AoLP, and LQ in terms of channels and pixel positions, and obtain a preliminarily enhanced feature I* after CA and SA operations.
[0089] CA = σ(MLP(AvgPool(F PoL , F RGB ) + MLP(MaxPool(F PoL , F RGB ))) (13)
[0090] SA = σ(f 7×7 ([AvgPool(CA); MaxPool(CA)])) (14)
[0091] where σ is the sigmoid operation, f 7×7According to past experience, a convolutional kernel of size 7*7 is adopted.
[0092] S402. Fusion branch: Using this initial enhanced output as spatial guidance, fuse the polarization information. DoLP can obtain the contrast and details of different regions, and AoLP can distinguish the characteristics of the background and target objects. First, merge DoLP and AoLP together, and use the PixelUnshuffle operation to downsample them to obtain the input of the residual dense block and reduce the computational requirements. Then, use residual learning to extract robust local features, and subsequently use the previously obtained global feature I* as guidance to compensate for the polarization information, and aggregate the features through spatial attention operation.
[0093] Fuse pol =H RDB (PixelUnshuffle(Concat(DoLP,AoLP))) (15)
[0094] Pol=MLP(AvgPool(Fuse pol +I*))×(Fuse pol +I*) (16)
[0095] Furthermore, the conditional denoising network described in this application adopts NAFNet as the noise prediction network, which will reduce the computational cost compared with the previous Transformer-based denoising network. Similar to most restoration tasks, the U-shaped encoder-decoder structure is also adopted, where NAFNet is the main structure in the encoder and decoder. By inputting I*, the inferior LQ image x T and the auxiliary information Pol, Time, predict the noise at the current moment:
[0096] F1=Conv 1×1 (EME(Pol)⊙Norm(X)+EME(Pol)) (17)
[0097] For the time information Time, γ and β are obtained through MLP operations. The role of γ is to perform scaling operations, and the role of β is to achieve feature movement. In this way, the time step t can be embedded into NAFNet:
[0098] F2=Conv 1×1 (γ1⊙Norm(F1)+β1) (18)
[0099] Subsequently, in order to better restore image details, a differential convolutional layer DEConv is used to explore more detailed information. The 3×3 conventional convolution is replaced with a convolution block composed of four different differential convolutions and conventional convolutions, and it is optimized using the reparameterization technique without introducing additional time complexity. Secondly, in order to introduce more abundant scene information and more detailed textures in the image, we embed Pol as a modulation parameter in each layer of NAFNet, and restore the image by adding the repaired details to the feature map. By using simple channel attention operations and simple gate operations, additional non-linear representations are incorporated:
[0100] F3 = Conv 1×1 (SCA(SimpleGate(DEConv 3×3 (F2)))) (19)
[0101] After layer normalization, scaling and shifting operations are performed again for modulation:
[0102]
[0103] Finally, the output can be obtained through the following formula:
[0104]
[0105] Next, through specific application examples, the solutions and effects of the present invention will be further described.
[0106] As Figure 4 shows the comparison of the actual application effects of the method of the present invention and other existing methods. Table 1 gives the comparison data. It can be seen from Table 1 that in this embodiment, the performance of different models on two datasets, UCPD and RGBP-UIE(Real), is compared. The table lists the performances of multiple models under different quality evaluation metrics (UIQM, UCIQE, NIQE, EME).
[0107] On the UCPD dataset: In terms of the UIQM metric, the method of the present invention (1.68) performs better than most other models, second only to PIPFNet (0.72). In terms of the UCIQE metric, the method of the present invention (0.51) performs best, significantly better than all other models. In terms of the NIQE metric, the method of the present invention (5.65) performs better than most models, second only to PIPFNet (3.75). In terms of the EME metric, the method of the present invention (10.18) performs best, significantly better than all other models.
[0108] Table 1 Performance comparison of different models
[0109]
[0110] On the RGBP-UIE (Real) dataset: In terms of the UIQM metric, the method (1.51) of the present invention performs best, significantly outperforming all other models. In terms of the UCIQE metric, the method (0.55) of the present invention performs best, significantly outperforming all other models. In terms of the NIQE metric, the method (3.09) of the present invention performs better than most models, second only to PIPFNet (4.36). In terms of the EME metric, the method (15.96) of the present invention performs best, significantly outperforming all other models.
[0111] From the above analysis, it can be seen that the method of the present invention performs excellently in multiple quality assessment metrics, especially in metrics such as UCIQE and EME. This indicates that the method of the present invention has significant advantages in underwater image enhancement, can effectively improve the image quality, and has high effectiveness and robustness. Through these technical features, a significant improvement in the quality of underwater images is achieved, thus providing a solid theoretical and technical foundation for subsequent visual tasks such as panoramic underwater observation.
[0112] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An underwater visual clarity method based on a polarization information guided diffusion model, characterized in that: The following steps are involved: S1. Acquire underwater polarization images under natural conditions and laboratory conditions respectively, and construct an underwater polarization image dataset, wherein the constructed underwater polarization image dataset includes underwater polarization images, polarization degree images, and polarization angle images at different angles; S2. constructing an underwater visual clarity network based on a polarization information guided diffusion model, wherein the underwater visual clarity network includes a serial conditional diffusion process, a cross-modal enhancement network, and a conditional denoising network; The conditional diffusion process is used to simulate the degradation process of the image underwater and estimate the noise information. The cross-modal enhancement network is used to compensate the RGB image information using the polarization image information to produce a pre-enhancement result, and the polarization information is updated using the pre-enhancement result to produce a fusion feature. The conditional denoising network is used to predict the noise information under the current time information based on the noise, the fusion feature and the time information, and obtain a clear underwater image using the diffusion reverse process and the noise information under the current time information. S3. Using the underwater polarization image dataset to train the underwater vision clarity network, and outputting clear underwater images based on the trained underwater vision clarity network.
2. The underwater visual clarity method based on polarization information guided diffusion model according to claim 1, characterized in that: The steps to acquire underwater polarization images under laboratory conditions include: By creating water bodies of different colors with different turbidity levels in a laboratory scene, using a polarization camera to collect turbid underwater polarization images corresponding to objects in water bodies of different colors with different turbidity levels, collecting clear underwater polarization images corresponding to objects in pure water, and using the clear underwater polarization images as label images; constructing a training set and a test set based on the turbid underwater polarization images and the label images; The steps of acquiring underwater polarization images in a real-world scenario include: preparing polarization images of different water environments and different underwater depth environments as test sets.
3. The underwater visual clarity method based on polarization information guided diffusion model according to claim 2, characterized in that: The polarization degree image and the polarization angle image are obtained according to the following method: Among them, DoLP represents the degree of polarization image, AoLP represents the angle of polarization image, S0 represents the total light intensity of the polarization image, S1 represents the ratio of 0° linear polarization to vertical polarization, and S2 represents the ratio of 45° linear polarization to vertical polarization.
4. The underwater visual clarity method based on polarization information guided diffusion model according to claim 1, characterized in that: The conditional diffusion process adopts a mean-restoration diffusion framework based on SDE, including: Forward process: Use SDE to diffuse the image into a degraded image with noise distribution; Reverse process: Generate samples by learning and simulating the inverse SDE, thereby training a neural network to estimate the score function of the noisy data distribution.
5. The underwater visual clarity method based on polarization information guided diffusion model according to claim 1, characterized in that: The forward process is described as follows: dx=θ t (μ-x)dt+σ t dw Among them, θ t and σ t are two time-varying parameters, λ is a fixed noise level, dx represents the sampling state that changes with time, and w represents the standard Wiener process; Among them, p t (x) represents Gaussian probability distribution, ▽ x logp t (x) represents the ground truth score at the inference stage, Represents the reverse Wiener process.
6. The underwater visual clarity method based on polarization information guided diffusion model according to claim 1, characterized in that: The cross-modal enhancement network includes: Pre-enhancement branch: The polarization mode and RGB mode images are encoded by two convolutional encoders to obtain the polarization feature F PoL and RGB feature F RGB , F PoL and F RGB After merging, they are fed into a general attention mechanism to obtain a preliminary enhanced feature through self-attention and cross-attention schemes; Fusion branch: Use the initial enhanced features as spatial guidance to fuse the polarization information, then use residual learning to extract local features from the fused information, then use the initial enhanced features as guidance to compensate for the polarization information, and obtain aggregated features through spatial attention operations.
7. The underwater visual clarity method based on polarization information guided diffusion model according to claim 6, characterized in that: The feature acquisition process of the preliminary enhancement is described as follows: CA=σ(MLP(AvgPool(F PoL ,F RGB )+MLP(MaxPool(F PoL ,F RGB ))) SA=σ(f 7×7 ([AvgPool(CA);MaxPool(CA)])) Among them, σ is the sigmoid operation, f 7×7 It is a convolution kernel of size 7*7, which outputs preliminary enhanced features after SA operation; The feature aggregation process of the fusion branch is described as follows: Fuse pol =H RDB (PixelUnshuffle(Concat(DoLP,AoLP))) Pol=MLP(AvgPool(Fuse pol +I*))×(Fuse pol +I*) Among them, I* represents the initial enhanced features, Pol represents the aggregated features, DoLP represents the degree of polarization image, AoLP represents the angle of polarization image, and H RDB Denotes a dense residual connection block.
8. The underwater visual clarity method based on polarization information guided diffusion model according to claim 1, characterized in that: The process of the conditional denoising network using NAFNet to output a clear image is described as: F3=Conv 1×1 (SCA(SimpleGate(DEConv 3×3 (F2)))) F2=Conv 1×1 (γ1⊙Norm(F1)+β1) F1=Conv 1×1 (EME(Pol)⊙Norm(X)+EME(Pol)) Among them, Pol represents the aggregated feature, X represents the result after adding noise to the input, γ1 represents the first feature scaling parameter, β1 represents the first feature shift parameter, γ2 represents the second feature scaling parameter, and β2 represents the second feature shift parameter.
Citation Information
Cited By
Underwater image de-scattering model construction method based on polarization imaging
CN121616478A
An underwater image despeckling model construction method based on polarization imaging
CN121616478B
Method and system for removing highlight based on polarization guidance
CN122243826A