Unsupervised underwater depth estimation method based on modeling of underwater light transmission guided by physics
Patent Information
- Application Number
- CN202611290555.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-25
AI Technical Summary
因此,这类方法在深度估计任务中精度有限
[0022]1)本发明提出一种新的水下深度估计框架,该框架结合位姿-场景联合优化NeRF、水下散射模型以及颜色先验,实现了高精度的深度估计并在一定程度上解决尺度不一致性问题;
Smart Images

Figure CN122820792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater depth estimation technology, and in particular to an unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling. Background Technology
[0002] Currently, terrestrial 3D reconstruction and depth estimation technologies have formed a mature system, but the unique characteristics of the underwater environment bring many severe challenges. Water attenuates light much more strongly than air and exhibits wavelength-selective absorption, leading to color distortion and decreased contrast in underwater images. Scattering effects caused by suspended particles in the water blur scene details and reduce visibility, creating a visual degradation similar to underwater haze. Furthermore, natural light attenuates exponentially with depth, and artificial light sources have limited illumination range, resulting in uneven brightness distribution in images. These problems severely restrict the accuracy of underwater 3D reconstruction and depth estimation.
[0003] As one of the most widely used techniques, monocular depth estimation (MDE) aims to infer pixel-level depth from a single RGB image. In recent years, advanced land-based methods such as DPT and Depth Anything have made significant progress in generalization and prediction quality. However, MDE relies on scene priors, and the fundamental differences between underwater and land imaging domains lead to performance degradation across domains. Currently, most mainstream underwater depth estimation algorithms are monocular, and researchers have improved performance through methods such as dark channel priors and combining depth and ambiguity relationships, while also incorporating strategies like style transfer and color correction to address data limitations. However, MDE lacks baseline information found in stereo vision, resulting in a lack of absolute distance references and thus inconsistencies across viewpoint scales. The non-uniformity and layering effects of the underwater medium disrupt the medium assumptions and perspective geometry upon which MDE relies, leading to systematic depth estimation biases. Furthermore, MDE is a highly ill-posed problem, heavily reliant on high-quality data constraints, while severely degraded underwater images result in scarce training data.
[0004] To compensate for the scale limitations of depth estimation (MDE), binocular depth estimation (BDE) uses disparity calculation between the left and right views as its core logic. Existing methods often employ fully convolutional networks trained on labeled data to achieve depth estimation. These methods possess absolute scale consistency and utilize the disparity redundancy information of the left and right views for depth calculation, making them more robust to local image degradation. However, water scattering and light absorption cause image blurring, and inconsistent degradation between binocular images makes reliable feature matching between the left and right views difficult. Underwater pressure changes and temperature fluctuations can easily cause slight deformations in the rigid connection structure of the binocular camera, leading to calibration parameter drift. Disparity calculation for binocular matching is typically at the sub-pixel level; errors caused by parameter drift can directly overwrite the effective disparity signal, causing depth estimation to completely fail. Furthermore, the unique underwater environment makes recalibration extremely difficult, and system errors are hard to control. Simultaneously, large areas of low-texture underwater and interference from suspended particles further reduce the algorithm's robustness. Coupled with the scarcity of high-quality underwater stereo annotation data, BDE-like solutions still face numerous challenges in underwater scenarios.
[0005] In recent years, Neural Radiation Field (NeRF) and 3D Gaussian Spray (3DGS) have attracted widespread attention in the field of 3D reconstruction and rendering due to their ability to synthesize high-quality new views. They achieve reconstruction and rendering through multi-view images and corresponding camera poses. However, underwater light absorption and scattering cause image degradation, making it difficult for traditional Structure in Motion (SFM) methods to accurately estimate camera pose. Image degradation caused by water scattering triggers a haze-like effect, violating the clean air medium assumption. Scattered light from non-object areas participates in the imaging process, causing ambiguity in depth estimation, i.e., incorrectly identifying the scattering medium as the foreground.
[0006] While NeRF-based methods have achieved good rendering results, they do not accurately represent the 3D structure of the scene. Instead, they treat underwater scatterers as foreground objects, leading to incorrect depth estimations. They prioritize rendering quality over geometric accuracy. As Yu et al. pointed out, good 2D rendering results do not equate to high 3D mesh accuracy. Existing methods such as Seathru-NeRF and WaterSplatting primarily model scattering based on simplified surface-guided assumptions or unquantized scattering approximations, neglecting the volumetric continuous transmission nature of scattering, wavelength-selective attenuation characteristics, and the need for precise geometric reconstruction. Specifically, scattering has core characteristics such as continuous interaction between light and suspended particles, significant differences in attenuation between different colors of light, and the participation of scattered light in imaging. Existing methods have failed to fully integrate NeRF or 3DGS with the physical transmission laws of scattering to construct a joint transmission correction model for light intensity and color attenuation during light propagation, thus failing to achieve more accurate light transmission modeling. Therefore, these methods have limited accuracy in depth estimation tasks. Summary of the Invention
[0007] The purpose of this invention is to provide an unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling. Even when underwater images suffer from poor image quality due to scattering or other reasons, this method can still recover high-quality 3D scene structures and obtain accurate depth estimation results, avoiding scale inconsistencies and depth ambiguities.
[0008] The technical solution to achieve the objective of this invention is: an unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling, comprising:
[0009] Step 1: Acquire raw underwater images from multiple perspectives;
[0010] Step 2: Jointly optimize camera pose and NeRF underwater implicit scene model;
[0011] Step 3: Construct a physically constrained underwater scattering model to predict water optical parameters and generate physical depth;
[0012] Step 4: Construct a priori color depth based on the underwater color attenuation law, and design multiple sets of joint constraint losses;
[0013] Step 5: Generate an adaptive progressive background mask and optimize foreground and background differences by applying constraints.
[0014] Step 6: Output a high-precision underwater dense depth map with a globally uniform scale.
[0015] Furthermore, step 1 specifically involves: acquiring raw underwater images from different shooting angles and distances using a single camera. The acquisition process does not require pre-calibration of the camera or calculation of the camera pose using the SFM algorithm. All raw underwater images are directly retained as input materials for the model. The underwater scene includes objects such as corals, reefs, and standard spheres, covering various water scattering environments, including clear, moderately turbid, and highly turbid waters.
[0016] Furthermore, step 2 specifically involves: directly inputting all the raw underwater images acquired in step 1 into the network to build a NeRF implicit radiation field framework, and simultaneously optimizing the camera's intrinsic and extrinsic parameters (rotation and translation pose) and the implicit geometric field of the underwater 3D scene; learning the true color and spatial density of clean, scatter-free underwater objects through a multilayer perceptron (MLP), and outputting a clean scene map free from water scattering interference and a NeRF initial geometric depth map to solve the problems of feature matching failure and inaccurate camera pose estimation caused by underwater scattering.
[0017] Furthermore, step 3 specifically involves: constructing a physically constrained underwater scattering model that conforms to the real underwater optical laws; decoupling three types of optical parameters—water attenuation coefficient, scattering coefficient, and background light at infinity—based on the physical laws of light absorption and volume scattering in water; using a CNN network (ResNet34) to learn and predict the corresponding water area optical parameters; fully simulating the entire imaging process of light attenuation and scattering in water; and calculating the physical depth map based on the physical imaging formula.
[0018] Furthermore, step 4 specifically involves: utilizing the different attenuation rates of red, green, and blue light underwater, constructing a nonlinear color prior depth calculation function to generate a color prior depth that conforms to the underwater distant scene pattern; simultaneously designing three types of loss constraints: image reconstruction loss, background physical constraint loss, and depth regularization loss, aligning the NeRF geometric depth, physical depth, and color prior depth with each other to eliminate the defects of misjudging scattering media as foreground and inconsistent depth scales across viewing angles.
[0019] Furthermore, step 5 specifically involves: dynamically generating background masks in stages: in the early stages of training, the initial background region is divided based on the red-blue and blue-green color ratios of the image plus the local brightness variance of the image; in the mid-to-late stages of training, depth reliability is calculated by combining depth gradient and local depth variance, and the mask is iteratively updated using a depth statistical threshold; then, temporal smoothing is performed on the masks of adjacent frames to avoid abrupt changes in the mask. Optical depth constraints are applied only to the background region marked by the mask, while the foreground region is optimized through NeRF geometric autonomous modeling; when the number of iterations is less than 0.4 times the total number of iterations, it is considered to be in the early stages of training, and then it enters the mid-to-late stages of training.
[0020] Furthermore, step 6 specifically involves: after completing the above multi-module collaborative iterative training, relying on the jointly optimized camera pose, the trained physical scattering model, the multi-loss constraint network, and the adaptive background mask, calculating the depth value of all pixels in the underwater scene at a uniform scale, and finally outputting an unambiguous, scale-consistent, and adaptable underwater dense depth map that is suitable for high and low turbidity water bodies, for use in underwater 3D reconstruction, ocean exploration, and downstream tasks of underwater robot environmental perception.
[0021] Compared with the prior art, the significant advantages of the present invention are:
[0022] 1) This invention proposes a novel underwater depth estimation framework that combines pose-scene joint optimization NeRF, underwater scattering model and color prior, to achieve high-precision depth estimation and solve the scale inconsistency problem to a certain extent.
[0023] 2) To ensure the accuracy of the scattering coefficient and process simulation, this invention proposes a physically constrained underwater scattering model to assist NeRF modeling and estimate the precise depth;
[0024] 3) This invention proposes an optical constraint optimization mechanism based on color prior, which solves the ambiguity problem in depth estimation through physical constraints, geometric regularization and photometric reconstruction collaborative supervision;
[0025] 4) This invention designs an adaptive background masking strategy, which dynamically updates the mask based on optical characteristics, ensuring the physical necessity of the scattering process while helping to eliminate depth ambiguity. Attached Figure Description
[0026] Figure 1 A schematic diagram illustrating the overall principle of underwater depth estimation.
[0027] Figure 2 This is a schematic diagram of scattering and imaging.
[0028] Figure 3 This provides qualitative results from ablation experiments on the standard sphere dataset regarding the elimination of ambiguity in depth estimation.
[0029] Figure 4 The results show the comparison of different scattering intensities.
[0030] Figure 5 The qualitative results of the experiment were compared with those of the MDE-type scheme.
[0031] Figure 6 The qualitative results of the experiment were compared with those of the BDE-type scheme.
[0032] Figure 7 The results are qualitative from the comparative experiments with NeRF-type schemes.
[0033] Figure 8 Qualitative experimental results (Kitty dataset) were compared with NeRF-like schemes. Detailed Implementation
[0034] Underwater images suffer from scattering and attenuation, leading to reduced contrast and color distortion, directly constraining depth estimation accuracy. Currently, monocular depth estimation (MDE) and binocular depth estimation (BDE) are the mainstream methods. MDE relies on visual cues from a single image to estimate depth, while BDE infers depth through binocular parallax and geometric relationships. Problems such as blurred visual features in underwater images, the disruption of geometric assumptions due to medium inhomogeneity, and equipment drift caused by water parameters result in poor depth estimation quality. Furthermore, both methods are highly dependent on the true underwater depth value and scene priors. In addition, MDE suffers from cross-viewpoint scale inconsistencies, further degrading estimation quality. Neural radiation field (NeRF) schemes, which rely on new view synthesis, can only estimate coarse depth and are prone to background depth ambiguity underwater, misidentifying scatterers as foreground.
[0035] To address the problem of underwater depth estimation, this invention proposes a Physically Guided Underwater Neural Radiation Field (UW-PhyNeRF) method. This method uses only the original underwater image as input, eliminating the need for underwater scattering priors and true depth values, and achieves high-precision depth estimation through unsupervised learning. Specifically, this invention constructs a physically constrained underwater scattering model, following the laws of underwater light transmission, incorporating scattering, absorption, and background light into a unified framework to simulate the complete underwater imaging process. The water scattering coefficient is jointly optimized with the NeRF implicit geometric representation to ensure geometric and scale consistency in depth estimation. Secondly, to address the ambiguity in reconstruction and depth estimation, an optical constraint optimization mechanism based on color priors is designed. A color prior is constructed based on the underwater attenuation differences of different wavelengths of light, enhancing the ability to distinguish between background and foreground. Furthermore, this invention proposes an adaptive progressive masking strategy to ensure the resolution of each region and the effectiveness of each module. Experiments show that the proposed method outperforms mainstream methods in both NeRF reconstruction and depth estimation, and it remains functional even in highly turbid environments with water scattering coefficients exceeding 0.5.
[0036] This invention makes the following assumptions regarding the underwater dataset: First, it is assumed that all images within the dataset were captured within a forward setting with a certain degree of rotation and translation flexibility to facilitate proper training of the Neural Radiation Field (NeRF). Second, it is assumed that the underwater images were acquired under natural lighting conditions. The principle of the method of this invention is as follows: Figure 1 As shown, this method proposes a physically constrained underwater scattering model to accurately predict water parameters. It leverages the multi-view consistency and high-quality modeling process of NeRF to assist in depth estimation. An optical constraint optimization mechanism based on color prior is designed to simultaneously eliminate scale uncertainty and depth ambiguity issues that are difficult to address in underwater depth estimation from multiple angles. To ensure that each module has accurate physical meaning, an adaptive progressive masking strategy is designed to impose additional constraints on the background region, thereby achieving high-precision depth estimation with scale consistency across the entire scene.
[0037] (1) NeRF and pose joint optimization
[0038] NeRF uses a multilayer perceptron (MLP) to represent a scene as a continuous 5D function. The input is the 3D spatial location. and 2D perspective direction ,in Polar angle, The direction angle is given, and the output is the color c and the volume density. .
[0039] First, define a ray. o is the origin of the ray, and d is the ray direction vector; for each distance It calculates the corresponding 3D position along the ray and performs position encoding, while optimizing by minimizing photometric error. The loss function is: (1)
[0040] in Original color, The closest point to the ray. At the farthest point of the ray, To predict color, The predicted color of the rays rendered for the foreground object, where R is the set of all rays. Given an input image I and a camera pose Π, NeRF renders a novel synthetic attempt. And use formula (2) for training:
[0041] (2)
[0042] Where L is the loss function. For network parameters, These are the initial network parameters that have not been predicted.
[0043] In underwater environments, light scattering and attenuation often cause haze and color distortion in images. Blurred images and reduced contrast can lead to serious errors or even failures in image feature extraction and matching, making it difficult to accurately estimate camera pose using Structure of Motion (SFM) algorithms. NeRFmm proposes a joint optimization scheme for pose and scene. Using only RGB images as input, for camera intrinsics, it defaults to (W / 2, H / 2) as the camera center point based on a pinhole camera model, where W and H are the width and height of the image, respectively, and only the focal length needs to be evaluated. For camera extrinsic parameters, it calculates the rotation matrix using the Rodrigues formula: (3) in Let W be the identity matrix and W be the axis angle. α is the rotation angle, and θ is the rotation axis. It is the tilt operator that converts the vector "·" into a tilt matrix. The translation matrix is defined in Euclidean space and can be directly set as the training parameter. In actual training, formula (4) is used for training:
[0044] (4)
[0045] in, This represents the camera parameters updated during the optimization process, where I is the input image. To predict the image, These are the pose parameters.
[0046] To better address the underwater depth estimation problem, this invention also compares NeRF and 3DGS. NeRF, based on a continuous implicit field representation of fine geometry, provides more continuous surface geometry modeling and theoretically has infinite resolution. 3DGS, on the other hand, relies on discrete Gaussian ellipsoids, which, while approximating complex geometry with a large number of Gaussians, has limitations in representing fine structures and semi-transparent objects. Underwater images suffer from edge blurring and structural degradation, making it difficult for 3DGS to obtain dense depth information based on continuum density integrals. NeRF's continuous geometric representation is more adaptable to degraded observations. Furthermore, NeRF has lower dependence on the initial point cloud; even with sparse views or occlusion, it can better generalize and fill in unobserved areas through the inductive bias of the MLP. 3DGS, however, relies more on accurate initial point clouds and visible areas, and its expansion capability is inferior to NeRF, making it prone to holes or geometric errors. Underwater scattering images are difficult to initialize effectively using traditional SFM algorithms for high quality; therefore, NeRF has better versatility than 3DGS.
[0047] (2) Physically constrained underwater scattering model
[0048] Compared to air, scattering and attenuation in underwater environments often lead to problems such as blurring, color distortion, and decreased contrast in underwater images, which to some extent exacerbates the difficulty of NeRF reconstruction. Accurate simulation of the underwater scattering process is crucial. Currently, the underwater scattering process is typically simulated using a modified model, i.e., equation (5):
[0049] (5)
[0050] Where I represents the image scattered by the water body, i.e., the image captured by the camera at a distance d, and C represents the clean image of the scene at a distance d if there were no scattering medium. The attenuation coefficient is... The scattering coefficient is... Background light at infinity and Let be the wavelength correlation function, representing respectively and Dependence on ambient light spectrum, camera spectral response, object reflectivity, distance, wavelength attenuation coefficient, and water physical scattering coefficient.
[0051] Currently, most solutions, such as Seathru-NeRF and WaterSplatting, directly assume... Furthermore, the background light is uniform, which facilitates calculation. However, simple scattering models cannot accurately simulate the scattering process, leading to depth estimation failure. Therefore, this invention proposes a Physically Constrained Underwater Scattering Model (PCUS), which optimizes the simulation of the scattering process and then combines it with NeRF to achieve accurate underwater depth estimation and eliminate depth ambiguity.
[0052] According to Lambert-Beer's law, light decays exponentially when propagating in water, and its transfer function is: ,in, Let d(x) be the attenuation coefficient, d(x) be the distance, and the scattering coefficient s be expressed as the volume scattering function. Integrals: (6) Assuming that all data were collected under natural lighting conditions, the background light is all ambient light within the imaging field of view that is scattered into the lens from various angles by water molecules and suspended particles, and is unrelated to the subject.
[0053] The background light scattered into the lens is calculated along the camera ray; it can be considered as a truncated cone of infinitesimal thickness, such as... Figure 2 As shown. The scattering produced by the frustum at a distance l from the camera is as shown in equation (7):
[0054] (7)
[0055] in, Indicates the scattering direction. The scattering angle is... The angle between the object and the ray. For ambient light, , and For camera-related parameters, The background light is scattered by infinitesimal elements. The attenuation coefficient is... The angle is the rotation angle.
[0056] Since NeRF samples and integrates along the ray, the frustum-shaped ray is parameterized to better adapt to NeRF, as shown in Equation (8):
[0057] (8)
[0058] in, This represents the cross-sectional area of the frustum. Indicates the effective receiving area of a pixel. As the normalized infinitesimal scattered background light that can be directly used for NeRF ray integration after correction, we have:
[0059] (9)
[0060] Considering only natural light conditions, the underwater light intensity does not vary significantly within a certain range; for ease of further calculation, it is assumed to be a constant value E. Furthermore, since the object is usually much farther from the camera than the camera's focal length, i.e. ,think .make Therefore:
[0061] (10)
[0062] Background light at a distance z from the camera:
[0063] (11)
[0064] in, Total cumulative background light;
[0065] For the background light at infinity, i.e. At that time, we can obtain:
[0066] (12)
[0067] in, The scattering coefficient s can be expressed as the integral of the volume scattering function, therefore it can be deduced that... That is, the background light is directly proportional to the scattering coefficient and inversely proportional to the attenuation coefficient.
[0068] This invention uses CNN for parameter prediction instead of Transformer structure because Transformer is better at capturing low-frequency information, while underwater images suffer from fog effects, low contrast, etc., which lead to a lack of high-frequency information. CNN has a stronger ability to construct high-frequency representations.
[0069] (3) Optical constraint optimization mechanism based on color prior
[0070] In atmospheric depth estimation, objects are in a good medium, allowing for accurate depth estimation of both the object and the background. However, in underwater scenes, especially in highly turbid water, objects can be understood as being enveloped by a scattering medium. As water turbidity increases, objects become increasingly invisible. In this case, severe scattering can cause NeRF to incorrectly learn the scattering medium as a foreground covering the entire space. The more severe the scattering, the more difficult it is to accurately estimate the depth, and the more serious the depth ambiguity problem becomes.
[0071] According to Lambert-Beer's law, the attenuation coefficient of light propagating underwater satisfies... ,in The attenuation coefficient of the R channel. The attenuation coefficient of the G channel. Let B be the attenuation coefficient. Therefore, the relative attenuation of the red channel and the blue-green channel satisfies: Where M is the maximum intensity value of the blue-green channel, and R is the intensity value of the red channel. The attenuation coefficient corresponding to the channel with the largest blue-green channel. The prior depth calculation based on color prior design is shown in formula (13):
[0072] (13)
[0073] in, , , Here, x is the spatial coordinate of a single pixel in the image plane, R(x) is the grayscale value of the red channel corresponding to pixel x, M(x) is the maximum value of the green and blue channel grayscale values of pixel x, and ε is a numerical stability constant. This is a pixel-level color prior depth constructed based on the underwater wavelength-selective attenuation law. The above scheme integrates optical attenuation characteristic embedding and nonlinear depth response to improve the model's expressive ability in turbid water through a dual-modal mechanism. It is worth noting that this prior depth is more accurate in the background region, while the depth estimation in the foreground is prone to errors due to the influence of object color, lighting conditions, etc.
[0074] For the background area, i.e. At this time there is And thus k is a fixed geometric constant derived from the camera optical parameters and water scattering integral, s is the water volume scattering coefficient, and a is the attenuation coefficient; the background light A is the upper limit of I, therefore, after simplifying the above formula, the physical depth can be obtained as: At this point, the background physical constraint loss is proposed:
[0075] (14)
[0076] Where B represents the set of pixels in the background region. Let B be the total number of pixels in the background region. This constraint is calculated only in the background region to ensure that the depth calculation conforms to physical reality. It assists in the calculation of scattering parameters, refines the simulation of the scattering process, and solves the problem of scale inconsistency to some extent.
[0077] Based on underwater optical properties and the principle of geometric consistency, a depth-canonical loss is proposed:
[0078] (15)
[0079] in , where is the weighting coefficient and ε is the numerical stability constant. This loss aligns geometric rendering with color optical properties, forcing them to be consistent in the background region, ensuring consistency across multiple viewpoints, and providing an absolute scale reference to resolve scale uncertainty issues. Adding distance-related weighting ensures adaptive focusing of the scene, strictly constrains accuracy for distant views, and eliminates background ambiguity in depth estimation.
[0080] The predicted water parameters are combined with a clean NeRF to simulate the scattering process and obtain the scattered image, thereby constructing the reconstruction loss: (16) N is the total number of pixels. To predict the image, The input image serves as the basis for this process; this ensures the accuracy of NeRF modeling and depth estimation, and guides the scattering process correctly.
[0081] In summary, the overall loss of the multi-scale optimization mechanism for color perception is as follows:
[0082] (17)
[0083] in, and All of these are hyperparameters.
[0084] (4) Adaptive progressive background mask
[0085] Physical depth constraints only hold true in the background region and are difficult to implement in the foreground object region. Therefore, to ensure the effectiveness of physical constraints and as an auxiliary training strategy, an adaptive progressive background mask is designed. Furthermore, since the depth of the foreground region is mainly determined by object geometry, while the depth of the background region is mainly determined by the scatterer, this strategy can also achieve optimized directional separation.
[0086] Since blue light attenuates the least and red light attenuates the most in underwater scenes, an initial mask is accurately constructed based on underwater optical characteristics during the initial training phase. Specifically, the blue-red ratio threshold (BR) and blue-green ratio threshold (BG) are first calculated, and the base mask is determined based on color priors and empirical values. (18) Choosing a dual threshold can enhance generalization across different water areas to some extent. Since the underwater background region has a uniform brightness distribution, brightness consistency analysis is performed, and the background region is selected by calculating the local brightness variance. ,in, This is a background binary mask obtained based on local brightness variance filtering. Let be the local brightness variance of the input image I within the x-neighborhood of pixel x, and thre be a preset threshold for the brightness variance; the threshold is determined based on the typical variance values of different water bodies. Combining these two ensures the accuracy of the initial mask selection from a multi-dimensional perspective:
[0087] (19)
[0088] This serves as the initial mask. To ensure mask continuity, morphological closing operations are used to smooth the edges, making the boundaries more natural and eliminating holes in the background area.
[0089] As training progresses, the depth becomes increasingly reliable. At this point, the background region has a relatively large depth, a relatively uniform distribution, and is correlated with color features. In the later stages of training, a dynamic threshold mask is constructed based on depth statistics. First, a reliability function is built relying on the depth gradient and local variance, and then the depth reliability is evaluated.
[0090] (20)
[0091] in, For gradient, Let R(x,y) be the local variance, and let R(x,y) be the reliability score of the depth map at pixel (x,y).
[0092] Subsequently, background depth sets are extracted based on reliable depth and the background mask of the previous frame. With foreground depth set , This is the background mask for the previous frame. To ensure accuracy, a statistic is constructed using the quantile method:
[0093] (twenty one)
[0094] (twenty two)
[0095] in, The statistical mean of the background depth set. The standard deviation of the background depth set. This is the set of depths corresponding to all background pixels selected from the mask of the previous frame. Let D be the total number of elements in the background depth set, and D(x,y) be the depth value predicted by the network at pixel (x,y). To sum the depth of all background pixels, percentile(D,80) is the 80th percentile for the entire image depth. The sum of squares of the differences between all background depths and the mean, 100 is the threshold for determining whether the sample size of background depth is sufficient, and 0.1 is a fixed scaling factor.
[0096] To ensure that the foreground is not misclassified as background and cause ambiguity, an additional upper bound for the foreground is constructed:
[0097] (twenty three)
[0098] in, The upper bound threshold for foreground depth. This is the depth set corresponding to all foreground pixels selected from the mask of the previous frame. The total number of elements in the foreground depth set. The maximum value in the foreground depth set is denoted by percentile(D,50), which is the 50th percentile of the total image depth, and 100 is the threshold for determining whether the foreground depth sample size is sufficient.
[0099] Based on the above, construct a dynamic depth threshold:
[0100] (twenty four)
[0101] Where T is the dynamic depth threshold. The statistical mean of the background depth set. The standard deviation of the background depth set;
[0102] Based on the dynamic depth threshold adjustment mask of formula (24), an empirical value is set to ensure color consistency. As an additional constraint, a final mask is constructed jointly, where B is the blue channel value and R is the red channel value.
[0103] When fusing with the mask from the previous frame, temporal consistency should be maintained; that is, the updated mask should not abruptly change. Therefore, to address the issue of abrupt mask changes in consecutive frame iterations, a time-smooth update mechanism is designed. First, the overlap rate of the two frames' masks is calculated:
[0104] (25)
[0105] in, Let be the overlap ratio of the background masks of frame t and frame (t-1), where W and H are the image width and height, respectively, and i and j are the pixel row and column coordinates. For the original binary mask of frame t, The original binary mask of the previous frame is used. The higher the overlap rate, the more consistent the background distribution of the current frame is with that of the previous frame. If the overlap rate is too low, it means that the scene has changed significantly or there is an error in the optimization.
[0106] Set the adaptive blending factor: This achieves an exponentially moving smooth mask:
[0107] (26)
[0108] in, Let be the mask value of the t-th frame after time-smoothing. For the adaptive masking mixing weight coefficients of frame t, For the unsmoothed original mask of frame t, Let i and j be the original mask for the (t-1)th frame, where i and j are the row and column coordinates of the image pixels;
[0109] Binarize formula (26) to obtain the final mask:
[0110] (27)
[0111] in, is the final binary background mask for frame t.
[0112] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0113] Example
[0114] This embodiment verifies the effectiveness of each module through ablation experiments. Through comparative experiments, the scheme of the present invention is qualitatively and quantitatively compared with the current mainstream schemes, and the method of the present invention is comprehensively evaluated, thereby proving the practicality and reliability of the method of the present invention.
[0115] For publicly available datasets, the Seathru-NeRF multi-scene underwater dataset was selected, covering multiple different sea areas (such as the Red Sea, Caribbean Sea, and Pacific Ocean), encompassing various water conditions. To ensure the diversity of water hues caused by different scattering conditions, additional data from the South China Sea was collected to verify the effectiveness of the proposed method. Simultaneously, an underwater standard sphere (radius 254 mm) dataset was collected to assist in verifying the accuracy of reconstruction and depth estimation. None of the data required pre-calculation of camera pose using SFM algorithms (such as Colmap).
[0116] This invention evaluates depth estimation from multiple perspectives. Since true underwater color data is unavailable and true underwater depth data is lacking, peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and learned perceptual patch similarity (LPIPS) are used to evaluate rendering accuracy for open water datasets, providing another angle to demonstrate depth accuracy. Furthermore, the absolute error and percentage error are calculated using an underwater standard sphere dataset to quantitatively evaluate the accuracy of depth estimation.
[0117] The framework of this invention is implemented based on PyTorch, with the NeRF part following the NeRFmm architecture. For the water parameters, ResNet34 is used as the backbone for prediction due to its stronger high-frequency representation capability compared to MLP and Transformer. For the color prior depth, multiple sets of refined depth maps of different water areas and hues are manually selected for training, and 10 cross-validations are used to obtain the desired result. , , .
[0118] This embodiment systematically verifies the role and effectiveness of the proposed UW-PhyNeRF core module. Ablation experiments were conducted on multiple modules and constraints in the framework, including the scattering-attenuation coupled physics-optimized scattering model and the color-aware multi-scale optimization mechanism. These modules were then removed and tested on an underwater dataset.
[0119] This embodiment records the absolute accuracy and visualization effects of standard sphere reconstruction, as well as the quantitative indicators between the synthesized new view and the real viewpoint from the public dataset, as shown in Table 1. Figure 3 As shown in Table 2.
[0120] Table 1 Quantitative results of ablation experiments using the standard sphere dataset
[0121] Table 2 Quantitative Results of the Curasao Dataset
[0122] Experimental results show that the PCUS module of this invention plays a crucial role in NeRF modeling and depth estimation. Without PCUS, the framework can hardly complete the depth estimation task. Furthermore, the color prior-based optical constraint optimization mechanism also plays a significant role in improving reconstruction and depth estimation. Specifically, the depth regularization loss assists in reconstruction accuracy and depth estimation results in the object region, while the background physical constraint loss guides NeRF to learn the background region more accurately, eliminating ambiguity issues in the depth estimation process. Figure 3 The study also demonstrates the qualitative effects of depth estimation. In the depth estimation results without introducing background physical constraints, the object accuracy is relatively good, but significant errors occur in the background region. As can be seen from the figure, it incorrectly treats the scattering objects in the background region as the foreground, resulting in serious errors in depth estimation.
[0123] In real-world scenarios, scattering varies significantly depending on factors such as water flow velocity and temperature. To demonstrate that this method functions correctly under various scattering conditions, we additionally simulated underwater environments with different scattering intensities using a simulation dataset and reconstructed them accordingly. The results are as follows: Figure 4 As shown.
[0124] As can be seen from the above results, the reconstruction results of Baseline NeRF are greatly affected by the increase of scattering, until finally it is even impossible to reconstruct. The method of the present invention can guarantee the reconstruction accuracy as the scattering intensity gradually increases, and the PSNR and SSIM are both satisfactory, thus proving the accuracy of the method of the present invention in depth estimation. The experimental results show that the model of the present invention has strong robustness to different levels of underwater scattering.
[0125] (1) Comparison with monocular depth estimation (MDE) schemes
[0126] Monocular depth estimation is a common method for underwater depth estimation tasks. Although monocular depth estimation methods offer clearer and cleaner visualizations, the single view often results in the lack of baseline length required for triangulation and the absence of absolute distance references, leading to scale inconsistencies. Using the underwater standard sphere dataset, we tested and compared several high-performance monocular depth estimation methods. Specific qualitative results are as follows: Figure 5 As shown in Table 3, the quantitative results are as follows.
[0127] Table 3. Quantitative results of the experiment compared with the MDE-type scheme.
[0128] The experimental results above demonstrate that the method of this invention significantly outperforms monocular depth estimation schemes in terms of accuracy. While MDE methods offer superior visualization with clearer contours and higher contrast, their scale uncertainty is severe. Changes in image perspective and other factors within the same scene can lead to significant deviations in predicted depth. For example, Depth Anything v2 has an average absolute error of 238.58 mm, but experiments revealed a maximum error of 1032.4 mm and a minimum error of 15.42 mm. This indicates that MDE schemes completely lack scale consistency and can only provide relative depth. In contrast, the method of this invention, relying on the 3D geometry provided by NeRF, eliminates scale inconsistency and achieves more stable absolute accuracy in depth estimation tasks.
[0129] (2) Comparison with binocular depth estimation (BDE) schemes
[0130] Binocular depth estimation has certain limitations in underwater depth estimation tasks, including difficulty in data acquisition, calibration parameter drift, and feature matching difficulties, resulting in a limited number of BDE-based solutions for underwater depth estimation. This paper uses an underwater standard sphere dataset for testing and compares the results with current state-of-the-art binocular depth estimation methods. Specific qualitative results are as follows: Figure 6 As shown in Table 4, the quantitative results are as follows.
[0131] Table 4. Quantitative results of the experiment compared with the BDE-type scheme.
[0132] Since the dataset includes standard spheres from multiple viewpoints, epipolar correction is prioritized to ensure fairness in the comparative experiments before binocular depth estimation. The experimental results demonstrate that the method described in this invention outperforms binocular depth estimation schemes in terms of accuracy. BDE methods often suffer from unstable underwater lighting and water flow disturbances, leading to feature matching failures between consecutive frames underwater and making it difficult to acquire training data. Therefore, although BDE methods also offer relatively high accuracy in depth estimation, they still have limitations in practical applications. In contrast, the method described in this invention uses only a single camera to capture underwater data, avoiding the difficulty in fixing the attitude of the left and right cameras due to water flow disturbances and pressure changes in the underwater environment, and also avoiding the impact of dynamic drift caused by temperature fluctuations on depth estimation. Furthermore, the method described in this invention also outperforms BDE-type schemes in terms of depth estimation accuracy.
[0133] (3) Comparison with NeRF-type schemes
[0134] Current mainstream NeRF-based approaches primarily focus on novel view synthesis. In underwater scenes, some advanced methods have achieved good novel view synthesis results in the foreground region, but artifacts often exist in the background region. However, when it comes to depth estimation, these methods lack sufficient constraints, relying entirely on volume density calculations for depth estimation, resulting in poor depth estimation quality. 3DGS-based approaches even exhibit hole artifacts. In this task, the method of this invention does not require pre-calculation of pose using SFM algorithms (such as Colmap), while other approaches require accurate pose. Standard sphere datasets are difficult to estimate accurately due to underwater scattering, so we use the Kitty dataset for absolute error comparison. Real underwater datasets lack accurate depth data, so we provide PSNR, SSIM, and LPIPS as quantitative evaluations. The Coral scene is introduced to demonstrate the effectiveness of this method under various water conditions and color tones. Specific qualitative and quantitative results are as follows: Figure 7 , Figure 8As shown in Tables 5 and 6, the three different colored points in the figure represent the maximum and minimum depths of the selected spheres, and the absolute error is calculated based on these values.
[0135] Table 5. Quantitative results of experiments compared with NeRF-like schemes
[0136]
[0137] Table 6. Quantitative results (absolute error) of the experiment compared with NeRF-type schemes.
[0138] Based on the experimental results above, it is evident that the method of this invention outperforms mainstream underwater solutions in terms of new view synthesis and radiation field reconstruction. Particularly regarding the background issue, the method of this invention can significantly alleviate background artifacts while achieving clearer background region reconstruction and obtaining more accurate background depth. Furthermore, the absolute error results on the Kitty dataset demonstrate that our advantage in depth estimation accuracy is even more significant.
Claims
1. A method for unsupervised underwater depth estimation based on physical-guided underwater optical transmission modeling, characterized in that, include: Step 1: Acquire raw underwater images from multiple perspectives; Step 2: Jointly optimize camera pose and NeRF underwater implicit scene model; Step 3: Construct a physically constrained underwater scattering model to predict the optical parameters of the water body and generate the physical depth; Step 4: Construct a priori color depth based on the underwater color attenuation law, and design multiple sets of joint constraint losses; Step 5: Generate an adaptive progressive background mask and optimize foreground and background differences by applying constraints. Step 6: Output a high-precision underwater dense depth map with a globally uniform scale.
2. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling as described in claim 1, characterized in that, In step 1, raw underwater images are captured from different shooting angles and distances using a single camera.
3. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling as described in claim 1, characterized in that, In step 2, all the raw underwater images acquired in step 1 are directly input into the network to build the NeRF implicit radiation field framework. The camera intrinsic and extrinsic parameters and the implicit geometric field of the underwater 3D scene are optimized simultaneously. The true color and spatial density of clean, scatter-free underwater objects are learned through a multilayer perceptron (MLP), and a clean scene map without water scattering interference and the NeRF initial geometric depth map are output.
4. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling as described in claim 3, characterized in that, Step 2 is as follows: First, define a ray. o is the origin of the ray, and d is the ray direction vector; for each distance It calculates the corresponding 3D position along the ray and performs position encoding, while optimizing by minimizing photometric error. The loss function is: (1) in Original color, To predict color, The predicted color of the rays rendered for the foreground object, where R is the set of all rays; given the input image I and the camera pose Π, NeRF renders the predicted image. And use formula (2) for training: (2) Where L is the loss function. For network parameters, These are the initial network parameters that have not been predicted. Using an RGB image as input, for camera intrinsics, the camera center point is defaulted to (W / 2, H / 2) based on the pinhole camera model, where W and H are the width and height of the image, respectively; for camera extrinsics, the rotation matrix is calculated using the Rodrigues formula: (3) in Let W be the identity matrix and W be the axis angle. , Let θ be the rotation angle, and θ be the axis of rotation. It is the tilt operator that converts the vector "·" into a tilt matrix; in actual training, formula (4) is used for training: (4) in, This represents the camera parameters updated during the optimization process. These are the pose parameters.
5. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 4, characterized in that, Step 3 specifically involves: A physically constrained underwater scattering model is constructed. Based on the physical laws of light absorption and volume scattering in water, three types of optical parameters are decoupled: water attenuation coefficient, scattering coefficient, and background light at infinity. A CNN network is used to learn and predict the corresponding optical parameters of the water area, and the entire imaging process of light attenuation and scattering in water is fully simulated. The physical depth map is calculated based on the physical imaging formula.
6. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 5, characterized in that, Construct a physically constrained underwater scattering model, specifically as follows: The underwater scattering process is simulated using a modified model, namely formula (5): (5) Where I represents the image scattered by the water body, i.e., the image captured by the camera at a distance d, and C represents the clean image of the scene at a distance d if there were no scattering medium. The attenuation coefficient is... The scattering coefficient is... Background light at infinity and Let be the wavelength correlation function, representing respectively and Dependence on ambient light spectrum, camera spectral response, object reflectivity, distance, wavelength attenuation coefficient, and water physical scattering coefficient; According to Lambert-Beer's law, light decays exponentially when propagating in water, and its transfer function is: Where d(x) is the distance and the scattering coefficient is... Expressed as volume scattering function Integrals: (6) Assuming that all data were collected under natural lighting conditions, and that the background light was all ambient light within the imaging field of view that was scattered into the lens from various angles by water molecules and suspended particles; The background light scattered into the lens is calculated along the camera ray. The scattering produced by the frustum at a distance l from the camera is shown in formula (7): (7) in, Indicates the scattering direction. The scattering angle is... The angle between the object and the ray. For ambient light, , and For camera-related parameters, Background light is scattered by infinitesimal elements; Since NeRF samples and integrates along the ray, the ray along the frustum is parameterized as shown in Equation (8): (8) in, This represents the cross-sectional area of the frustum. Indicates the effective receiving area of a pixel. This is a standardized infinitesimal scattering background light that can be directly used for NeRF ray integration after correction; (9) Under natural light conditions, the underwater light intensity is a constant value E; because ,think ;make Therefore: (10) Background light at a distance z from the camera: (11) in, This represents the total cumulative background light; for background light at infinity, i.e. At that time, we obtained: (12) in, The scattering coefficient s is expressed as the integral of the volume scattering function, therefore it can be deduced that That is, the background light is directly proportional to the scattering coefficient and inversely proportional to the attenuation coefficient.
7. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 1, characterized in that, In step 4, the optical characteristics of different attenuation rates of red, green and blue light underwater are utilized to build a nonlinear color prior depth calculation function to generate color prior depth that conforms to the underwater scene rules. Simultaneously, three types of loss constraints are designed: image reconstruction loss, background physical constraint loss and depth regularization loss. The NeRF geometric depth, physical depth and color prior depth are aligned and constrained to eliminate the defects of misjudging the scattering medium as the foreground and inconsistent depth scales across viewpoints.
8. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 7, characterized in that, Step 4 is as follows: According to Lambert-Beer's law, the attenuation coefficient of light propagating underwater satisfies... ,in The attenuation coefficient of the R channel. The attenuation coefficient of the G channel. Let B be the attenuation coefficient. Therefore, the relative attenuation of the red channel and the blue-green channel satisfies: Where M is the maximum intensity value of the blue-green channel, and R is the intensity value of the red channel. The attenuation coefficient corresponding to the channel with the largest blue-green channel; the prior depth is calculated according to the color prior design as shown in formula (13): (13) in, , , Here, x is the spatial coordinate of a single pixel in the image plane, R(x) is the grayscale value of the red channel corresponding to pixel x, M(x) is the maximum value of the green and blue channel grayscale values of pixel x, and ε is a numerical stability constant. The pixel-level color prior depth is constructed based on the underwater wavelength selective attenuation law; For the background area, i.e. At this time there is And thus , k is a fixed geometric constant derived from the camera optical parameters and the water scattering integral; Background light A is the upper limit of I, physical depth for: Background physical constraint loss for: (14) Where B represents the set of pixels in the background region. The total number of pixels in the background region pixel set B; Based on underwater optical properties and the principle of geometric consistency, a depth canonical loss is proposed. : (15) in , where is the weighting coefficient, and ε is the numerical stability constant; The predicted water parameters are combined with a clean NeRF to simulate the scattering process and obtain the scattered image, thereby constructing the reconstruction loss: N is the total number of pixels. To predict the image, For the input image; Design a multi-scale optimization mechanism for color perception, with its comprehensive loss for: (16) in, and All of these are hyperparameters.
9. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 1, characterized in that, Step 5 specifically involves: The training process dynamically generates adaptive progressive background masks in stages: In the early stage of training, the initial background region is divided based on the red-blue and blue-green color ratios of the image and the local brightness variance of the image; in the middle and late stages of training, the depth reliability is calculated by combining the depth gradient and the local depth variance, and the mask is iteratively updated by using a depth statistical threshold; then, temporal smoothing is performed on the masks of adjacent frames; optical depth constraints are applied only to the background region marked by the mask, and the foreground region is optimized by NeRF geometric autonomous modeling; when the number of iterations is less than 0.4 times the total number of iterations, it is considered to be in the early stage of training, and then it enters the middle and late stages of training.
10. The unsupervised underwater depth estimation method based on physical-guided underwater optical transmission modeling according to claim 9, characterized in that, Step 5 specifically involves: In the initial training phase, an initial mask is constructed based on underwater optical characteristics. First, the blue-red ratio threshold BR and the blue-green ratio threshold BG are calculated, and the basic mask is determined based on color priors and empirical values. A brightness consistency analysis was performed, and the background area was selected by calculating the local brightness variance. ,in, This is a background binary mask obtained based on local brightness variance filtering. Let x be the local brightness variance of the input image I within the neighborhood of pixel x, and thre be a preset threshold for the brightness variance; construct an initial mask. : (17) To ensure mask continuity, morphological closing operations are used to smooth the edges, making the boundaries more natural and eliminating holes in the background area; In the later stages of training, a dynamic thresholding mask is constructed based on depth statistics. First, a reliability function is constructed using depth gradients and local variances, and then depth reliability is evaluated. (18) in, For gradient, Let R(x,y) be the local variance, and let R(x,y) be the reliability score of the depth map at pixel (x,y). Subsequently, background depth sets are extracted based on reliable depth and the background mask of the previous frame. With foreground depth set , Use the background mask of the previous frame; construct statistics using the quantile method: (19) (20) in, The statistical mean of the background depth set. The standard deviation of the background depth set. This is the set of depths corresponding to all background pixels selected from the mask of the previous frame. Let D be the total number of elements in the background depth set, and D(x,y) be the depth value predicted by the network at pixel (x,y). To sum all background depth pixels, percentile(D,80) is the 80th percentile for the depth of the entire image; This is the sum of squares of the differences between all background depths and the mean. To ensure that the foreground is not misclassified as background and cause ambiguity, an additional upper bound for the foreground is constructed: (21) in, The upper bound threshold for foreground depth. This is the depth set corresponding to all foreground pixels selected from the mask of the previous frame. The total number of elements in the foreground depth set. The maximum value in the foreground depth set is denoted by percentile(D,50), which is the 50th percentile of the total image depth. Constructing a dynamic depth threshold: (22) Where T is the dynamic depth threshold. The statistical mean of the background depth set. The standard deviation of the background depth set; Based on the dynamic depth threshold adjustment mask of formula (22), an empirical value is set to ensure color consistency. As an additional constraint, a final mask is jointly constructed, where B is the blue channel value and R is the red channel value; When fusing with the mask of the previous frame, temporal consistency should be maintained, meaning the updated mask should not be abrupt; a smooth temporal update should be designed; first, the overlap rate of the two frames' masks should be calculated: (23) in, Let be the overlap ratio of the background masks of frame t and frame (t-1), where W and H are the image width and height, respectively, and i and j are the pixel row and column coordinates. For the original binary mask of frame t, The original binary mask for frame t-1; i,j are the row and column coordinates of the image pixels; Set the adaptive blending factor: This achieves an exponentially moving smooth mask: (24) in, Let be the mask value of the t-th frame after time-smoothing. The adaptive mask mixing weight coefficients for frame t; Binarize formula (24) to obtain the final mask: (25) in, is the final binary background mask for frame t.