Underwater depth estimation method based on confidence guided fusion and monocular fallback mechanism
By adopting an underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism, the problem of unstable depth estimation in underwater environment is solved, and real-time and accurate depth perception is achieved under harsh conditions, thereby improving the operational safety and positioning accuracy of autonomous underwater vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DALIAN MARITIME UNIVERSITY
- Filing Date
- 2025-11-27
- Publication Date
- 2026-06-19
AI Technical Summary
Existing underwater depth estimation methods are unstable and inaccurate in complex underwater environments, and cannot maintain the continuity and reliability of the measurement scale under conditions of reduced visibility, which limits the operational safety and positioning accuracy of autonomous underwater vehicles.
A quantized underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism is adopted. A dense depth and confidence map is generated by a GPU-accelerated nearest neighbor filling module. The method combines a confidence-guided fusion network and a multi-stage deformable refinement network, uses a temporal memory backoff module to maintain scale consistency, and is trained by supervised regression loss and confidence-weighted smoothing constraints.
It achieves highly robust and efficient depth perception in complex underwater environments, maintains consistency and stability of depth estimation metrics, and can provide real-time and accurate depth information under harsh conditions, making it suitable for platforms such as autonomous underwater vehicles.
Smart Images

Figure CN121527149B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically, to a quantitative underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism. Background Technology
[0002] Underwater robots rely on accurate depth estimation for tasks such as navigation, detection, and mapping. However, the underwater environment presents numerous challenges, such as optical scattering, turbidity, and color decay, all of which diminish the effectiveness of traditional vision-based depth estimation systems. Binocular cameras, sonar, and structured light sensors typically provide only sparse or noisy measurements, while monocular depth estimation only predicts relative depth without absolute depth information, thus inherently suffering from scale ambiguity and drift problems.
[0003] Existing fusion methods (such as TRUDepth and SBES) typically assume that all depth priors have uniform reliability. However, this assumption does not hold true in real underwater environments because sensor degradation or failure is often unpredictable. These issues lead to unstable and inaccurate depth predictions, thus limiting the operational safety and positioning accuracy of autonomous underwater vehicles (AUVs).
[0004] Therefore, there is an urgent need for a robust, efficient, and real-time underwater depth estimation method that can maintain the continuity and reliability of the measurement scale even when visibility decreases. Summary of the Invention
[0005] In view of the shortcomings of existing technologies, this invention provides a metric underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism. This invention mainly integrates confidence-guided fusion and a memory-based backoff mechanism to maintain the consistency and stability of the metric scale of depth estimation under harsh underwater conditions.
[0006] The technical means employed in this invention are as follows:
[0007] A quantized underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism includes the following steps:
[0008] Acquire an underwater image dataset and construct a training set based on a portion of the underwater image data, wherein the underwater images include synchronously extracted RGB images and geometric priors;
[0009] A metric underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism is constructed. This model is trained on a training set and includes a GPU-accelerated nearest neighbor filling module, a confidence-guided fusion network, a memory checkpoint-based temporal memory backoff and recovery module, and a multi-stage deformable refinement network. The GPU-accelerated nearest neighbor filling module performs completion and confidence estimation from sparse to dense depths. The confidence-guided fusion network dynamically integrates RGB images and geometric priors based on local reliability. The memory checkpoint-based temporal memory backoff and recovery module tracks scale and offset parameters. The multi-stage deformable refinement network enhances structural consistency and compensates for underwater refraction and scattering effects.
[0010] The underwater RGB image to be estimated is acquired and input into the trained quantized underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism to perform depth estimation.
[0011] Furthermore, the GPU-accelerated nearest neighbor filling module finds the nearest valid sparse point in the set of valid sparse points for each pixel based on the image pixel coordinates, thereby generating a dense depth prior.
[0012] Furthermore, the GPU-accelerated nearest neighbor filling module generates a geometric confidence map based on the following calculations. :
[0013]
[0014] in, For pixels, To and The nearest effective sparse point For an effective sparse point set, To effectively measure the maximum distance between confidence scores of any pixel, the formula is as follows:
[0015]
[0016] in, and These represent the height and width of the image, respectively.
[0017] Furthermore, the confidence-guided fusion network, based on the input monocular depth prediction result and the completed depth prior, and combined with the geometric confidence map, uses an adaptive weighting function to determine the contribution ratio of each data source according to local reliability, thereby obtaining the fused output depth. The fused output depth is calculated according to the following formula:
[0018]
[0019] in, Indicates the output depth after fusion. For the geometric confidence of the depth set, For dense depth results, This is the result of monocular depth prediction.
[0020] Furthermore, the memory-based rollback and recovery module maintains the exponential smoothing of the scale and offset parameters based on the memory rollback mechanism according to the following formula:
[0021]
[0022] in, This represents the smoothing scale estimate for the current frame. This indicates the smoothed scale estimate of the previous frame. This represents the smooth update scale calculated based on the current frame. .2 indicates a smoothing factor used to control the update rate; the maintenance process is subject to sound standards. Constraints, when When the value reaches the upper limit of the constraint, the system performs a smooth scaling update. The memory rollback mechanism depends on the system reliability, which is defined by the memory confidence metric.
[0023]
[0024] in, Indicates the number of successful updates. This refers to the amount of memory cache.
[0025] Furthermore, the multi-stage deformable refinement network learns spatial offsets through the adaptive spatial deformation process of the deformation fusion module, and aligns the depth prediction with the underlying geometry of the underwater scene based on the spatial offsets.
[0026] Furthermore, during the training of the quantified underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism, supervised regression loss, confidence-weighted smoothing constraint, and temporal consistency objective are used as the loss function. The loss function is:
[0027]
[0028] The scale-invariant logarithmic loss is:
[0029]
[0030] in, Indicates the number of valid pixels. and To predict the actual depth value, Weight for scale invariance control;
[0031] The depth-weighted loss is:
[0032]
[0033] Where Ω represents the set of valid pixels;
[0034] The gradient loss function is:
[0035]
[0036] in, This represents the spatial gradient of the depth map in the x-direction. This represents the spatial gradient of the depth map in the y-direction. , , These are the weighting coefficients for each loss term.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] This invention discloses a real-time underwater depth estimation framework for underwater robot perception, aiming to achieve highly robust and efficient depth perception in complex underwater environments. This application integrates a confidence-guided fusion module and a memory-based backoff mechanism to maintain the consistency and stability of the depth estimation metric under harsh underwater conditions.
[0039] The GPU-accelerated nearest neighbor fill module (NNFILLGPU) introduced in this invention can quickly complete the filling from sparse depth to dense depth and generate pixel-by-pixel confidence maps. These confidence maps are used to guide the adaptive fusion network to dynamically integrate monocular and geometric cues. Meanwhile, the temporal memory module maintains scale consistency when the sensor degrades or fails.
[0040] Furthermore, the multi-stage deformable refinement network in this invention can correct underwater optical distortion, thereby improving the clarity and detail accuracy of depth boundaries.
[0041] Experimental results show that the method of this invention can achieve a real-time performance of 25 frames per second on a standard RTX3070 GPU and can run in real time on embedded hardware such as Jetson Orin. This invention provides an accurate, scalable, and real-time underwater depth estimation solution suitable for platforms such as autonomous underwater vehicles (AUVs). Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This invention presents a quantitative underwater depth estimation model architecture based on confidence-guided fusion and monocular backoff mechanism in an embodiment of the present invention.
[0044] Figure 2 This is the algorithm flow for the confidence-guided fusion and monocular backoff mechanism of this invention.
[0045] Figure 3 This study compares the performance of confidence-guided fusion and monocular backoff mechanism on the FLSea dataset with other methods in predicting depth map structural details under different depth sparse conditions.
[0046] Figure 4 This study compares the temporal stability and robustness of confidence-guided fusion and monocular backoff mechanisms with other methods on the FLSea dataset for quantized underwater depth estimation models under sparse priors. Detailed Implementation
[0047] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0048] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0049] This invention provides a quantitative underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism, which mainly includes the following steps.
[0050] S1. Obtain an underwater image dataset and construct a training set based on a portion of the underwater image data. The underwater images include synchronously extracted RGB images and geometric priors.
[0051] The dataset used for training and evaluation in this application consists of underwater scenes captured under different visibility and lighting conditions. This embodiment preferably uses the FLSea dataset, which contains twelve independent underwater scenes from two different marine environments, serving as the basis for all experiments. Its main subset, FLSea-Inertia, contains 22,451 images covering the following twelve scenes.
[0052] Canyon environments (Canyons): U Canyon, Horse Canyon, Tiny Canyon, and Flatiron, characterized by complex geological structures;
[0053] The Red Sea environment includes the Big Dice Loop, Coral Table Loop, Cross Pyramid Loop, Dice Path, Northeast Path, Landward Path, Pier Path, and Sub Pier, featuring rich coral morphology and biotextural characteristics.
[0054] Each frame provides a synchronized RGB image and a geometric prior generated via inertial-assisted structure-from-motion reconstruction to ensure consistent metric supervision for monocular-geometric hybrid training. The dataset is divided into eleven scene groups, with 90% of the data used for training and 10% for validation. The SubPier scene is reserved specifically for testing and evaluation in unknown environments.
[0055] To evaluate cross-domain generalization capability, this embodiment also uses the FLSea-Stereo Canyon1 subset as an external test set. This subset exhibits significant domain shift due to different ocean conditions and imaging characteristics, providing a rigorous benchmark for verifying the robustness and adaptability of the invention in diverse underwater environments.
[0056] S2. Construct a metric underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism, and train the metric underwater depth estimation model based on the training set. For example... Figure 2As shown, the quantized underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism includes a GPU-accelerated nearest neighbor filling module, a confidence-guided fusion network, a memory checkpoint-based temporal memory backoff and recovery module, and a multi-stage deformable refinement network.
[0057] A GPU-accelerated nearest neighbor fill module is used for completion and confidence estimation from sparse to dense depth. In existing technologies, sparse depth priors obtained from underwater sensors such as SLAM, binoculars, or sonar are often irregular and spatially discontinuous, making them unsuitable for direct input into convolutional networks. Therefore, this invention introduces the NNFILLGPU module—a deterministic, GPU-accelerated sparse-to-dense depth completion module that can convert sparse input into a continuous depth field and corresponding geometric confidence map in real-time. Traditional CPU-based nearest neighbor fill methods rely on sequential search loops, resulting in high computational latency (approximately 2-3 frames / second), making them unsuitable for real-time applications.
[0058] Specifically, the GPU-accelerated nearest neighbor filling module NNFILLGPU in this application completely eliminates iterative loops by fully utilizing the massively parallel computing capabilities of the GPU. All image pixel coordinates and effective sparse points are simultaneously broadcast to the GPU tensor, thereby achieving fully parallel Euclidean distance calculation across the entire image range. For each pixel... In the effective point set Find its nearest effective sparse point The calculation method is as follows:
[0059]
[0060] Among them, Indicates density depth. This represents the sparse depth, and the equation can be expressed as for each pixel. to densify its depth The value is assigned to its nearest valid neighbor. The calculation for sparse points is as follows:
[0061]
[0062] This design enables a comprehensive yet parallel search, achieving a speedup of approximately 338 times compared to CPU-based implementations, and generating dense deep priors with sub-millisecond latency.
[0063] At the same time, NNFILLGPU will generate a geometric confidence map. Used to quantify space reliability:
[0064]
[0065] in
[0066]
[0067] In relation to In the formula, and represents the height and width of the image, respectively. The confidence level is 1 at the original sparse points and decreases linearly with distance. The generated dense depth prior and the confidence map together form a robust geometric support for subsequent confidence-guided fusion, enabling the network to adaptively trust reliable geometric regions while minimizing uncertainty.
[0068] Furthermore, a confidence-guided fusion network is used to dynamically integrate RGB images and geometric priors based on local reliability. The fusion network takes monocular depth prediction results and the completed depth priors as input, and combines them with a pixel-wise confidence map. An adaptive weighting function determines the contribution ratio of each data source based on local reliability, thereby obtaining the fused output depth, which is calculated according to the following formula:
[0069]
[0070] in, Indicates the output depth after fusion. For the geometric confidence of the depth set, For dense depth results, This is the result of monocular depth prediction.
[0071] Furthermore, the timing memory rollback and recovery module based on memory checkpoints maintains the exponential smoothing (EMA) of the scale and offset parameters according to the following formula:
[0072]
[0073] in, This represents the smoothing scale estimate for the current frame. This indicates the smoothed scale estimate of the previous frame. This represents the smooth update scale calculated based on the current frame. The smoothing factor is used to control the update rate; this maintenance process is subject to skeletal standards. Constraints, when When the value reaches the upper limit of the constraint, the system performs a smooth scaling update. The memory rollback mechanism depends on the system reliability, which is defined by the memory confidence metric.
[0074]
[0075] in, Indicates the number of successful updates. This refers to the amount of memory cache.
[0076] When sensor data is lost, the system rescales the monocular prediction results based on stored parameters, thereby maintaining consistency of the measurement scale without additional computation.
[0077] Furthermore, a multi-stage deformable refinement network is used to enhance structural consistency and compensate for underwater refraction and scattering effects. The fused depth map is refined through a multi-stage deformable convolutional network to compensate for optical distortion and improve boundary accuracy and visual consistency. This network learns spatial offsets to align the depth prediction results with the geometry of the underwater scene.
[0078] The model in this application is trained end-to-end, comprehensively utilizing supervised regression loss, confidence-weighted smoothing constraints, and a temporal consistency objective. The training loss function is:
[0079]
[0080] The scale-invariant logarithmic loss is:
[0081]
[0082] in, Indicates the number of valid pixels. and To predict the actual depth value, The weight is the control weight for scale invariance.
[0083] The depth-weighted loss is:
[0084]
[0085] Where Ω represents the set of valid pixels.
[0086] The gradient loss function is:
[0087]
[0088] in, This represents the spatial gradient of the depth map in the x-direction. This represents the spatial gradient of the depth map in the y-direction. , , The weighting coefficients for each loss term are determined through cross-validation to balance the contribution of each component to the overall loss.
[0089] During training, sparse priors are randomly discarded to simulate real sensor failures, thereby improving the model's robustness under degraded input conditions.
[0090] The trained model achieved real-time inference performance of 25 frames per second on RTX3070 and 16 frames per second on JetsonOrin, outperforming TRUDepth and UDepth in various accuracy metrics. This invention demonstrates excellent generalization ability and strong adaptability to environmental changes, providing reliable perception capabilities for underwater robots.
[0091] S3. Obtain the underwater RGB image to be estimated, input it into the trained metric underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism, and perform depth estimation.
[0092] The following specific application examples will further illustrate the solution and effects of the present invention.
[0093] like Figure 3 and Figure 4 As shown, this embodiment evaluates the model of this application on the FLSea underwater benchmark dataset, which covers environments with different turbidity and illumination conditions. Experimental results show that the method performs excellently in terms of both accuracy and real-time inference performance.
[0094] As shown in Table 1, under extremely sparse conditions (only 10 points, similar to sonar input), the absolute relative error (Abs Rel) of HI-UDepth reaches 0.1061, which is 42.6% higher than that of TRUDepth. Even when using only monocular backoff, this method still improves upon previous baseline methods by 77.8%.
[0095] Table 1. Comparison of the effects of the present invention with various methods
[0096]
[0097] As shown in Table 2, the model runs at 25.1 frames per second on an RTX 3070 GPU and 16.8 frames per second on an NVIDIA Jetson Orin, achieving better performance than existing algorithms while still maintaining the system's real-time performance.
[0098] Table 2 Comparison of computational efficiency between the present invention and other methods
[0099]
[0100] As shown in Tables 3 and 4, the ablation experiment results indicate that confidence-guided fusion, deformable refinement, and memory rollback mechanisms all play a significant role in improving accuracy and temporal stability.
[0101] Table 3 Ablation Experiment Results Based on Confidence-Based Fusion Algorithm
[0102]
[0103] Table 4. Experimental results of the ablation algorithm based on the depth densification and checkpoint rollback of this invention.
[0104]
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for metric underwater depth estimation based on confidence guided fusion and monocular fallback mechanism, characterized in that, Includes the following steps: Acquire an underwater image dataset and construct a training set based on a portion of the underwater image data, wherein the underwater image data includes synchronously extracted RGB images and geometric priors; A metric underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism is constructed. This model is trained on a training set and includes a GPU-accelerated nearest neighbor filling module, a confidence-guided fusion network, a memory checkpoint-based temporal memory backoff and recovery module, and a multi-stage deformable refinement network. The GPU-accelerated nearest neighbor filling module performs completion and confidence estimation from sparse to dense depths. The confidence-guided fusion network dynamically integrates RGB images and geometric priors based on local reliability. The memory checkpoint-based temporal memory backoff and recovery module tracks scale and offset parameters. The multi-stage deformable refinement network enhances structural consistency and compensates for underwater refraction and scattering effects. The underwater RGB image to be estimated is acquired and input into the trained quantized underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism to perform depth estimation.
2. The method of claim 1, wherein, The GPU-accelerated nearest neighbor filling module generates a dense depth prior by finding the nearest valid sparse point in the set of valid sparse points for each pixel based on the image pixel coordinates.
3. The method of claim 2, wherein, The GPU-accelerated nearest neighbor filling module also generates a geometric confidence map according to : in, For pixels, To and The nearest effective sparse point For an effective sparse point set, To effectively measure the maximum distance between confidence scores of any pixel, the formula is as follows: wherein and denote the height and width of the image, respectively.
4. The quantized underwater depth estimation method based on confidence-guided fusion and monocular backoff mechanism according to claim 1, characterized in that, The confidence-guided fusion network, based on the input monocular depth prediction results and the completed depth prior, combined with the geometric confidence map, uses an adaptive weighting function to determine the contribution ratio of each data source according to local reliability, thereby obtaining the fused output depth. The fused output depth is calculated according to the following formula: in, Indicates the output depth after fusion. For the geometric confidence of the depth set, For dense depth results, This is the result of monocular depth prediction.
5. The method of claim 1, wherein, The memory-based checkpoint-based temporal memory rollback and recovery module maintains exponential smoothing of the scale and offset parameters based on the memory-based rollback mechanism according to the following formula: in, This represents the smoothing scale estimate for the current frame. This indicates the smoothed scale estimate of the previous frame. This represents the smooth update scale calculated based on the current frame. .2 indicates a smoothing factor used to control the update rate; the maintenance process is subject to sound standards. Constraints, when When the value reaches the upper limit of the constraint, the system performs a smooth scaling update. The memory rollback mechanism depends on the system reliability, which is defined by the memory confidence metric. wherein, represents the number of successful updates, is the amount of memory cache.
6. The method of claim 1, wherein, The multi-stage deformable refinement network learns spatial offsets through the adaptive spatial deformation process of the deformation fusion module, and aligns the depth prediction with the underlying geometry of the underwater scene based on the spatial offsets.
7. The method of claim 1, wherein, The loss function for training the quantized underwater depth estimation model based on confidence-guided fusion and monocular backoff mechanism is: The scale-invariant logarithmic loss is: in, Indicates the number of valid pixels. and To predict the actual depth value, Weight for scale invariance control; The depth-weighted loss is: Where Ω represents the set of valid pixels; The gradient loss function is: in, This represents the spatial gradient of the depth map in the x-direction. This represents the spatial gradient of the depth map in the y-direction. , , These are the weighting coefficients for each loss term.
Citation Information
Patent Citations
Monocular depth map completion method introducing scattering rate
CN116071416A
IMAGE ENHANCEMENT AND OBJECT DETECTION SYSTEM FOR DEGRADED UNDERWATER IMAGES USING ZERO-REFERENCE DEEP CURVE ESTIMATION and G-UNET
US20250045891A1