Endoscope image depth estimation method based on deep learning
By introducing a registration module into the depth estimation algorithm, and automatically registering and generating visibility masks using optical flow networks, the accuracy problems of traditional depth estimation algorithms in complex situations such as brightness changes and occlusion in endoscopic scenes are solved, and a higher precision depth and posture estimation is achieved.
Patent Information
- Application Number
- CN202510343557.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional depth estimation algorithms perform poorly in endoscopic scenarios when facing complex situations such as highlights, reflections, and occlusions, resulting in increased depth estimation errors, affecting the accuracy of surgical navigation, and being difficult to meet the sub-mm-level accuracy requirements.
The registration module is introduced to predict forward and backward optical flows through the optical flow network, perform automatic registration steps, generate visibility masks, filter pixel points that are blocked or exceeded the field of view, thereby improving the estimation accuracy of the appearance flow network.
Effectively respond to complex brightness changes in the endoscopic environment, improve the accuracy of depth and posture estimation, meet the sub-mm-level accuracy requirements, and significantly reduce the depth estimation error.
Smart Images

Figure CN120182346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning, computer vision, and depth estimation, aiming to achieve monocular depth estimation in an endoscopic scenario to ensure high-quality three-dimensional reconstruction and surgical navigation of the endoscopic scenario. Background Art
[0002] With the rapid development of modern medical technology, endoscopic technology, as a minimally invasive diagnosis and treatment method, plays an increasingly important role in the medical field. According to the statistical data of the World Health Organization (WHO), there are approximately 50 million endoscopic examinations and surgeries worldwide every year, and its application scope has expanded from the initial gastrointestinal tract examination to multiple fields such as respiration, urology, and gynecology. However, traditional endoscopic technology still has many limitations in terms of image quality, operation accuracy, and intelligent assistance, which prompts researchers to continuously explore new technological breakthroughs.
[0003] Research is carried out on key issues such as the robustness, accuracy, and illumination changes of depth estimation in the endoscopic scenario. With the popularization of minimally invasive surgery, endoscopic surgery plays an increasingly important role in modern medicine. However, due to the particularity of the endoscopic environment, traditional depth estimation algorithms are often unstable in the face of complex situations such as specular reflection, occlusion, and tissue deformation, seriously affecting surgical safety. According to statistics, approximately 25% of surgical complications are directly related to depth perception errors. Especially in key surgical operations, inaccurate depth estimation may lead to serious medical accidents.
[0004] In terms of the robustness of depth estimation, existing algorithms perform poorly in dealing with complex situations such as high highlights, reflections, and occlusions unique to the endoscopic scenario. Research shows that these visual interferences can increase the depth estimation error by 50% - 80%, seriously affecting the accuracy of surgical navigation. At the same time, the perspective changes brought about by tissue deformation and camera pose also pose great challenges to depth estimation, and the stability of existing algorithms in dynamic scenarios needs to be improved urgently. In terms of accuracy, traditional depth estimation algorithms often fail to meet clinical requirements in the endoscopic scenario. Statistical data shows that the depth estimation error of existing algorithms in the actual surgical environment is usually in the range of 3 - 5 millimeters, far exceeding the sub-millimeter accuracy requirements for surgery. This lack of accuracy not only affects the accuracy of surgical operations but also may increase medical risks. In addition, the illumination changes in the endoscopic scenario seriously violate the light constancy assumption relied on by traditional depth estimation algorithms. Research shows that illumination changes during the surgical process can cause a 40% - 60% decrease in the depth estimation accuracy. In a dynamic surgical environment, this problem is particularly prominent, and traditional algorithms often fail to estimate when dealing with drastic illumination changes.
[0005] Spencer et al. proposed enhancing the supervision signal by learning dense visual representations and using a robust feature space to improve feature consistency when the brightness constancy assumption fails. However, the overall texture of endoscopic images is sparser and more homogeneous than that of natural scenes, which is not conducive to visual representation learning and may even lead to a decline in depth estimation performance. Yang et al. and Ozyoruk et al. respectively introduced affine brightness transformers and tried to align the target frame and its corresponding frame to similar brightness conditions. However, by applying the same transformation parameters to all pixel points in the target frame, these affine brightness transformers are inefficient in dealing with complex local brightness changes caused by non-Lambertian reflection and mutual reflection. Although these methods alleviate the brightness inconsistency problem to some extent, none of them fundamentally rethinks the theoretical basis of self-supervised depth estimation. They either avoid the problem by introducing additional supervision signals or try to meet the brightness constancy assumption through simple brightness correction, lacking active modeling and representation of complex brightness changes in the endoscopic environment.
[0006] Shao et al. proposed introducing appearance flow into the self-supervised monocular depth estimation framework and actively modeling brightness changes by estimating the appearance flow network, effectively compensating for the depth estimation interference caused by inter-frame brightness changes. However, this method overly relies on the prediction accuracy of the appearance flow. Once feature extraction becomes difficult due to the sparse texture of organs and tissues in the endoscopic scene, combined with problems such as camera movement and pixel occlusion between adjacent frames, it will lead to poor prediction results of the final appearance flow, resulting in the failure of brightness compensation and thus poor depth estimation accuracy. Summary of the Invention
[0007] Aiming at the problem of drastic inter-frame brightness changes in the endoscopic scene, based on the existing self-supervised monocular depth estimation framework, a method of using optical flow to compensate for appearance flow prediction is proposed. This method makes the brightness change more prominent between adjacent frames by minimizing the pose component, which helps to more accurately extract the appearance flow. Specifically, the proposed solution of the present invention designs a registration module that introduces an optical flow network. By estimating the forward optical flow and the backward optical flow, it assists the estimation of the appearance flow network, thereby improving the estimation accuracy of the appearance flow network and further enhancing the accuracy of the final depth estimation of endoscopic images. The present invention solves the core problem of the failure of the brightness constancy assumption in the endoscopic scene from the theoretical basis. Different from traditional optical flow that only represents pixel displacement, appearance flow simultaneously considers the effects of geometric transformation and radiometric transformation and uses the optical flow network to assist in training, thus being able to effectively describe the complex brightness change patterns in the endoscopic scene caused by light source movement, non-Lambertian reflection, and mutual reflection, etc.
[0008] To address the above problems, the present invention proposes a method for endoscopic image depth estimation based on deep learning. This method assists the training of the appearance flow network by introducing a registration module. The registration module predicts the optical flow between frames through the optical flow network OFNet and performs an automatic registration step. The registration module also generates a visibility mask to filter out occluded or out-of-view pixels. Through an improved self-supervised framework, OFNet does not require ground-truth optical flow labels and learns high-quality motion information to align adjacent frames, reduce the impact of endoscopic camera motion, improve the accuracy of appearance flow estimation, and thus effectively cope with complex brightness changes in the endoscopic environment, providing more accurate depth and pose estimation for endoscopic images.
[0009] The core of the registration module is the optical flow network OFNet, which is mainly used to predict forward and backward optical flows. OFNet has the same architecture as AFNet. In the prediction layer of OFNet: OFNet outputs a two-channel optical flow, namely horizontal and vertical displacements; OFNet receives adjacent frame pairs of the spliced endoscopic images as inputs and outputs optical flow fields at four different resolutions. The optical flow field prediction layer uses a linear activation function, allowing displacement representations without range limitations:
[0010] F of (p) = x
[0011] where F of (p) is the optical flow vector at pixel p, and x is the original output of the network.
[0012] Furthermore, the automatic registration process performed by the registration module is as follows: Through an optical flow-guided registration process, the pose component is minimized, highlighting the brightness change. The specific process is as follows: OFNet predicts the forward optical flow and the backward optical flow
[0013]
[0014] where represents OFNet; using the forward optical flow and the spatial transformation network, the target frame is reconstructed:
[0015]
[0016] where represents the optical flow-based image deformation operation.
[0017] Furthermore, the registration module generates a mask V(p) by range map checking to filter out unreliable regions:
[0018]
[0019] V(u, v) = [R(u, v) > 0.95]
[0020] Among them, R(u,v) is the range map, which indicates the possibility that the pixel (u,v) is not blocked. and They represent the horizontal and vertical components of the backward optical flow, respectively, and W and H represent the width and height of the endoscope image.
[0021] Furthermore, the optical flow network OFNet is used to separate the endoscope posture component and brightness change. Including rigid motion F rigid and the non-rigid deformation F non-rigid :
[0022]
[0023] Through optical flow registration, the rigid component is approximately eliminated:
[0024]
[0025] The remaining difference is caused by brightness changes and provides an ideal input for appearance flow prediction:
[0026] C δ (p) = I t (p)-I recon (p)
[0027] OFNet achieves the function of separating posture and brightness changes by predicting optical flow, and minimizes the rigid motion component through registration, and the remaining difference C δ (p) Responds to brightness changes and provides clean input for AFNet.
[0028] Furthermore, the workflow of the improved self-supervised framework for endoscopic images includes the following steps:
[0029] (1) Input processing: receiving the endoscope image target frame I t and source frame I s As input, in the training phase, the source frame includes the adjacent frames before and after the target frame, i.e., I s ∈{I t-1 ,I t+1};
[0030] (2) Geometric transformation estimation: The depth module predicts the target frame I of the endoscopic image t The depth map D t , the pose module estimates the relative pose M between the target frame and the source frame t→s , the key supervisory signal comes from the warp-based view synthesis. Once the per-pixel depth value of the target frame is estimated, the pixels on the image plane are back-projected to the 3D camera space using the known camera intrinsic parameters. Using the estimated self-motion, the 3D point cloud is projected onto another image plane. The view synthesis is expressed as:
[0031]
[0032] Among them, K represents the camera internal parameters; h(p s→t ) and h(p t ) represent the uniform pixel coordinates in the target view t and the source view s respectively; D t represents the target frame depth map;
[0033] (3) View synthesis: Using the rigid flow and the spatial transformation network, synthesize the target frame from the source frame to obtain the initial synthesized frame I s→t ;
[0034] (4) Appearance flow prediction: The appearance module predicts the appearance flow C δ based on the target frame and the source frame, and aligns the brightness conditions of different frames through the brightness correction program:
[0035]
[0036] Among them, I t (p) represents the target frame; C δ (p) represents the appearance flow; I s (p) represents the source frame;
[0037] (5) Correspondence enhancement: The registration module predicts the forward and backward optical flows through the optical flow network OFNet, performs the automatic registration step, and generates the visibility mask V(p) of the endoscopic image to filter the occluded or out-of-view pixels;
[0038] (6) Loss calculation and optimization: Based on the synthesized frame, the corrected target frame and the visibility mask, calculate the data fidelity loss:
[0039]
[0040] Among them, represents the data fidelity loss; V(p) represents the visibility mask; Φ represents the image similarity metric function.
[0041] Edge-aware smoothing loss encourages the depth to be discontinuous at the color edges and continuous in the smooth regions:
[0042]
[0043] Among them, represents the first-order gradient of the depth map.
[0044] The photometric error loss between the synthesized frame and the real target frame
[0045]
[0046] Among them, α represents the weight coefficient; SSIM(I a ,I b ) represents the structural similarity coefficient; I a represents the target frame; I b represents the synthesized frame.
[0047] Residual-based smoothing loss
[0048]
[0049] Among them, represents the first-order gradient of the appearance flow; |I t (p)-I s→t (p)| approximately reflects the brightness change degree of each pixel in I t (p);
[0050] The complete self-supervised depth estimation loss function is:
[0051]
[0052] Compared with the self-supervised framework that only uses the appearance flow, the depth estimation framework for endoscopic images has unique advantages in many aspects: introducing an automatic registration enhancement mechanism, the corresponding module can minimize the pose component, making the brightness change more prominent, thereby improving the accuracy of appearance flow extraction, and thus can actively model the brightness change, accurately capture the complex local brightness changes in the endoscopic scene through the appearance flow, and fundamentally solve the problem of the failure of the brightness constancy assumption; adopting a modular design, four core modules (depth, pose, appearance, registration) each perform their own functions and work together to provide a flexible and scalable architecture; designing a multi-level loss function system, jointly constraining the network learning through the data fidelity loss and various regularization losses to ensure the physical rationality of the depth, pose, appearance flow, and optical flow prediction results; implementing visibility-aware training, filtering unreliable pixels that are occluded or out of the field of view through an accurate visibility mask, and significantly improving the training stability and prediction accuracy.
[0053] When the registration module collaborates with the depth module and the pose module, the optical flow information provided by the registration module helps to evaluate the accuracy of geometric transformation, especially when dealing with non-rigid deformations; then, when collaborating with the appearance module, the automatic registration step minimizes the pose component, highlights the brightness change, and creates favorable conditions for appearance flow prediction. At the same time, the mask ensures that the appearance flow is only predicted and optimized in reliable areas; in addition, in the role of the loss function, the mask generated by the registration module is used to weight each loss, filter the contributions of unreliable areas, and improve the training efficiency and model performance.
[0054] Generally speaking, through the design of automatic registration and masking, the registration module significantly enhances the framework's ability to handle complex conditions in endoscopic scenarios, providing important support for accurate depth and pose estimation. Overall, the depth module (DepthNet) is responsible for predicting depth maps from single-frame images; the pose module (PoseNet) estimates the relative camera pose between adjacent frames; the appearance module (AFNet) predicts the appearance flow and corrects the brightness conditions; the registration module (OFNet) performs automatic registration and generates visibility masks. These four modules work together to jointly construct a unified self-supervised framework that can effectively handle the particularities of endoscopic scenarios. The design of these modules fully considers the characteristics of endoscopic scenarios, especially complex illumination changes and non-rigid deformation problems. By introducing the automatic registration step, the framework can accurately predict the appearance flow, thereby accurately modeling geometric and radiometric transformations, fundamentally solving the problem of the failure of the brightness constancy assumption in traditional self-supervised methods in endoscopic scenarios. Description of the Drawings
[0055] Figure 1 is a monocular depth estimation framework based on appearance flow.
[0056] Figure 2 is the architecture of the registration module (OFNet).
[0057] Figure 3 is the curve graph of the original framework training loss function.
[0058] Figure 4 is the curve graph of the improved framework training loss function.
[0059] Figure 5 is the intermediate feature analysis graph of photometry.
[0060] Figure 6 is the visualization comparison of depth maps. Detailed Implementation Manner
[0061] The present invention will be described in detail below with reference to the accompanying drawings.
[0062] The self-supervised framework of the present invention improves on the existing monocular depth estimation framework based on appearance flow. The original framework mainly includes three core modules as Figure 1 shown: the depth module, the pose module, and the appearance module.
[0063] The depth module is responsible for predicting depth maps from single-frame RGB images, corresponding to the geometric transformation part in the Generalized Dynamic Image Constraint (GDIC). This module uses a depth estimation network (DepthNet) to predict the dense depth map of the target frame;
[0064] The pose module is used to estimate the relative camera pose between adjacent frames, and together with the depth module, it constitutes a geometric transformation. This module predicts the 6DOF relative pose parameterized by Euler angles and translation vectors through a pose estimation network (PoseNet);
[0065] The appearance module is used to predict the appearance flow and correct the brightness conditions, corresponding to the radiometric transformation part in the dynamic image constraint. This module includes an appearance flow network (AFNet) and a brightness correction program, which can effectively compensate for the inter-frame brightness differences in the endoscopic scene;
[0066] The depth module, pose module, and appearance module work together to jointly achieve the unified modeling of geometric transformation and radiometric transformation in the endoscopic scene. However, this method overly relies on the prediction accuracy of the appearance flow. Once it is difficult to extract features due to the sparse-textured organ tissues in the endoscopic scene, combined with problems such as camera movement and pixel occlusion between adjacent frames, it will lead to poor prediction results of the final appearance flow, resulting in the failure of brightness compensation, and further causing poor depth estimation accuracy.
[0067] To address the above problems, the present invention proposes a method for endoscopic image depth estimation based on deep learning. By introducing a registration module to assist in the training of the appearance flow network. Specifically, the registration module predicts the optical flow between frames through an optical flow network (OFNet), performs an automatic registration step to make the brightness changes more prominent between adjacent frames, which helps to accurately extract the appearance flow. The registration module also generates a visibility mask to filter out the occluded or out-of-view pixel points. Through an improved self-supervised framework, OFNet can learn high-quality motion information without real optical flow labels, thus aligning adjacent frames, reducing the impact of camera movement, improving the accuracy of appearance flow estimation, and being able to effectively handle the complex brightness changes in the endoscopic environment, providing more accurate depth and pose estimation.
[0068] The core of the registration module is the optical flow network (OFNet), which is used to predict the forward and backward optical flows. OFNet has the same architecture as AFNet (as Figure 2 shown), but the final prediction layer is different: OFNet outputs a two-channel optical flow (horizontal and vertical displacements), while AFNet outputs a three-channel appearance flow (RGB brightness changes). OFNet receives the concatenated pair of adjacent frames as input and outputs the optical flow field at four different resolutions. Different from the appearance flow representation, the optical flow field prediction layer uses a linear activation function, allowing displacement representations without range limitations:
[0069] F of (p) = x
[0070] where F of (p) is the optical flow vector at pixel p, and x is the original output of the network.
[0071] The automatic registration process performed by the registration module is as follows: Through the registration process guided by optical flow, the pose component is minimized, highlighting the brightness change. The specific process is as follows: OFNet predicts the forward optical flow and the backward optical flow
[0072]
[0073] Among them, denotes OFNet; using the forward optical flow and the spatial transformation network, the target frame is reconstructed:
[0074]
[0075] Among them, denotes the optical-flow-based image deformation operation.
[0076] This optical-flow-based reconstruction framework has two important significances: It can capture non-rigid deformations and complex poses in the endoscopic scene, which are difficult to handle by methods based on rigid transformations; this registration step minimizes the pose component between the source frame and the target frame, making the brightness change the main difference between the two frames, creating ideal conditions for appearance flow prediction. This registration effect provides key support for the accurate extraction of appearance flow.
[0077] In self-supervised depth estimation, occlusion and out-of-view regions are important interference factors. These regions have no effective pixel correspondence relationships and may introduce incorrect supervision signals. The registration module generates a mask V(p) based on range map checking to filter these unreliable regions:
[0078]
[0079] V(u,v) = [R(u,v) > 0.95]
[0080] Among them, R(u,v) is the range map, indicating the possibility that the pixel (u,v) is not occluded, and respectively represent the horizontal and vertical components of the backward optical flow, and W and H represent the image width and height. The threshold 0.95 is a value set according to the reference literature and fine-tuned on the SCARED dataset.
[0081] The mask V(p) plays an important role in the loss function calculation, ensuring that only reliable pixels participate in the generation of the supervision signal. The visibility mask can effectively identify occlusion regions and out-of-view regions, improving the training stability and prediction accuracy.
[0082] Traditional self-supervised depth estimation is based on the brightness constancy assumption: The pixel brightness between adjacent frames is only caused by geometric transformation, that is:
[0083] I t I(p) = I s (p + F rigid (p))
[0084] where F rigid (p) is a rigid flow, and the rigid flow is determined by depth and pose. However, in the endoscope scenario, the brightness changes due to the movement of the light source, tissue reflection, etc., and a radiation transformation needs to be introduced:
[0085] I t (p) = I s (p + F rigid (p)) + C δ (p)
[0086] where C δ (p) represents the radiation change, i.e., the appearance flow. Traditional methods ignore C δ (p), resulting in the failure of photometric error.
[0087] Separate the pose component and the brightness change through the optical flow network OFNet. The forward optical flow includes the rigid motion F rigid and the non-rigid deformation F non-rigid :
[0088]
[0089] Approximately eliminate the rigid component through optical flow registration:
[0090]
[0091] At this time, the remaining difference is caused by the brightness change, providing an ideal input for appearance flow prediction:
[0092] C δ (p) = I t (p) - I recon (p)
[0093] In summary, OFNet realizes the function of separating the pose and the brightness change by predicting the optical flow, minimizes the rigid motion component through registration, and the remaining difference C δ (p) reflects the brightness change and provides a clean input for AFNet.
[0094] The workflow of the improved framework can be summarized as the following steps:
[0095] (1) Input processing: Receive the target frame I t and the source frame I s as inputs. In the training stage, the source frame usually includes the adjacent frames before and after the target frame, i.e., I s ∈ {I t-1 , It+1};
[0096] (2) Geometric transformation estimation: The depth module predicts the depth map D t of the target frame I t , and the pose module estimates the relative pose M t→s between the target frame and the source frame. The key supervision signal comes from warping-based view synthesis. Once the per-pixel depth value of the target frame is estimated, the pixel points on the image plane are back-projected into the 3D camera space using the known camera intrinsics. Using the estimated self-motion, the 3D point cloud is projected onto another image plane. View synthesis is expressed as:
[0097]
[0098] where K represents the camera intrinsics; h(p s→t ) and h(p t ) represent the homogeneous pixel coordinates in the target view t and the source view s respectively; D t represents the depth map of the target frame;
[0099] (3) View synthesis: Using the rigid flow and the spatial transformation network, synthesize the target frame from the source frame to obtain the initial synthesized frame I s→t ;
[0100] (4) Appearance flow prediction: The appearance module predicts the appearance flow C δ based on the target frame and the source frame, and aligns the brightness conditions of different frames through a brightness correction procedure:
[0101]
[0102] where I t (p) represents the target frame; C δ (p) represents the appearance flow; I s (p) represents the source frame;
[0103] (5) Correspondence enhancement: The registration module predicts the forward and backward optical flows through the optical flow network (OFNet), performs an automatic registration step, and generates a visibility mask V(p) to filter out occluded or out-of-view pixels;
[0104] (6) Loss calculation and optimization: Based on the synthesized frame, the corrected target frame, and the visibility mask, calculate the data fidelity loss:
[0105]
[0106] where, represents the data fidelity loss; V(p) represents the visibility mask; Φ represents the image similarity metric function.
[0107] Edge-aware smoothing loss The encouraging depth is discontinuous at the color edges while remaining continuous in smooth regions:
[0108]
[0109] where, represents the first-order gradient of the depth map.
[0110] The photometric error loss between the synthesized frame and the real target frame
[0111]
[0112] where, α represents the weight coefficient; SSIM(I a ,I b ) represents the structural similarity coefficient; I a represents the target frame; I b represents the synthesized frame.
[0113] represents the residual-based smoothing loss
[0114]
[0115] where, represents the first-order gradient of the appearance flow; |I t (p)-I s→t (p)| approximately reflects the degree of brightness change of each pixel in I t (p);
[0116] The complete self-supervised depth estimation loss function is:
[0117]
[0118] The training strategy of this improved self-supervised framework is elaborated in detail below, including a two-stage training scheme, a parameter update strategy, a learning rate adjustment method, and model optimization techniques. In a complex multi-module network architecture, a scientific and reasonable training strategy has a decisive impact on the stability and performance of the framework, especially when there are close dependencies between modules.
[0119] This improved self-supervised framework integrates four interdependent core modules. Directly training all modules in an end-to-end manner may lead to unstable optimization processes, gradient conflicts, or getting stuck in local optimal solutions. To overcome these challenges, this study designed an efficient two-stage training scheme to gradually optimize each module in a reasonable dependency order.
[0120] The first stage focuses on the training of the optical flow network (OFNet) in the registration module, aiming to establish stable and reliable inter-frame pixel correspondences, laying the foundation for the geometric and radiometric transformation learning in the second stage. The training process of this stage includes: inputting adjacent frame pairs collected continuously; OFNet predicting the forward optical flow and backward optical flow; reconstructing the target frame based on the forward optical flow; generating the visibility mask through range map checking; calculating the loss function; and only updating the parameters of OFNet. The loss function of the first stage can be expressed as:
[0121]
[0122] where Φ is the image similarity metric function, is the edge-aware smoothing loss applied to the optical flow, and λ es is the weight coefficient.
[0123] The training in the first stage does not rely on depth or pose estimation, which enables OFNet to independently learn accurate pixel correspondences. By focusing on the training of a single module, the gradient interference that may be brought about by the simultaneous optimization of multiple modules is avoided, ensuring that OFNet obtains reliable registration capabilities.
[0124] After obtaining a stable OFNet, the parameters of OFNet are frozen in the second stage, while the depth module (DepthNet), pose module (PoseNet), and appearance module (AFNet) are trained simultaneously. The training process of this stage includes: inputting adjacent frame pairs; DepthNet predicting the depth map, and PoseNet estimating the relative pose; synthesizing the source frame based on the depth and pose; AFNet predicting the appearance flow and correcting the brightness condition; the frozen OFNet predicting the optical flow, generating the visibility mask and reconstructing the frame; calculating the loss function; and updating the parameters of DepthNet, PoseNet, and AFNet. The complete loss function of the second stage is:
[0125]
[0126] This two-stage training scheme has multiple advantages: the separate training reduces the interference between modules, enabling each module to focus on its own task; the pre-trained OFNet provides stable registration and visibility information, creating favorable conditions for the joint optimization in the second stage; this scheme significantly improves the training efficiency, accelerates the convergence process, and at the same time reduces the risk of falling into local optimal solutions.
[0127] The parameter update strategy is a key link in the training process, directly affecting the convergence quality and speed of the model. In the two-stage training scheme, different parameter update strategies are designed for different stages and different modules.
[0128] In the first stage, only the parameters of OFNet are updated, and the remaining modules have not participated in the training yet. In the second stage, a strategy of freezing OFNet and jointly updating the remaining modules is adopted, which specifically includes: fixing the parameters of OFNet trained in the first stage to serve as a stable "anchor point" to provide reliable pixel correspondence and visibility information for other modules; jointly updating DepthNet, PoseNet, and AFNet, where the three modules learn simultaneously and jointly adapt to the self-supervised signals. This joint update allows for co-adaptation among modules. For example, the appearance flow learned by AFNet can compensate for the inaccuracies of DepthNet and PoseNet, and vice versa.
[0129] To ensure that each module can work efficiently in cooperation and avoid unnecessary interference, specific gradient propagation and truncation strategies are adopted: allowing the loss function to generate gradients for all active modules through the synthetic image and the corrected image to achieve end-to-end joint optimization; applying gradient clipping in specific cases to limit the gradient norm within a preset threshold to ensure training stability; carefully adjusting the weight coefficients of the loss function to balance the gradients received by each module and prevent a certain module from dominating the optimization process.
[0130] The initialization of the model weights has an important impact on the training process and the final performance. The following strategies are adopted: the DepthNet encoder is initialized with the pre-trained ResNet-18 weights on ImageNet. This transfer learning strategy can effectively improve the network's feature extraction ability and accelerate the training convergence; the decoder and other modules are initialized with He initialization, which adaptively adjusts the weight variance according to the input dimension, contributing to the stable training of the deep network; during the training process, the encoder and the decoder can use different learning rates. Usually, the encoder uses a smaller learning rate for fine-tuning, while the decoder uses a larger learning rate to train from scratch.
[0131] The learning rate is one of the most crucial hyperparameters in the training of deep learning models, and its setting and adjustment directly affect the convergence speed and the final performance of the model. Considering the complex visual characteristics in the endoscopic scenario and the characteristics of the multi-module architecture, a systematic learning rate adjustment method is designed.
[0132] The two-stage training scheme corresponds to different learning rate settings: in the first stage (OFNet training), an initial learning rate of 1e-4 is adopted. This value has been verified through experiments to balance the training speed and stability and is suitable for the feature learning of the optical flow network; in the second stage (joint training), an initial learning rate of 1e-4 is also adopted, but different learning rate multipliers can be set according to different modules. For example, the pre-trained encoder part can use a smaller learning rate (such as 0.1 times the base learning rate).
[0133] As the training progresses, the learning rate needs to be gradually reduced to achieve fine optimization. A step decay strategy is adopted: after training for 10 epochs in each stage, the learning rate is reduced to 0.1 times the original; 0.1 is selected as the decay factor, which is a value with proven good effect and can avoid drastic oscillations during the optimization process while maintaining the training momentum; the lower limit of the learning rate is set to 1e-6 to prevent the training from stagnating due to an overly small learning rate. This step decay strategy can be expressed as:
[0134]
[0135] where η t is the learning rate at time step t, η0 is the initial learning rate, γ is the decay factor (0.1), t step is the decay step (10 epochs), represents the floor operation.
[0136] To avoid instability caused by a large learning rate in the initial stage of training, a learning rate warm-up strategy is adopted: in the first 500 iterations of each stage, the learning rate linearly increases from 0.1 times the initial value to the full value; this strategy helps the network find a good initial direction in the parameter space and then optimize using the full learning rate. The formula for the learning rate during the warm-up stage is:
[0137]
[0138] where t warmup is the warm-up iteration number (500).
[0139] Considering the learning characteristics and pre-training status of different modules, different learning rates are set for each module: The DepthNet encoder uses a smaller learning rate (such as 0.1 times the base learning rate) for fine-tuning due to using pre-trained weights; the DepthNet decoder is trained from scratch and uses the full learning rate; PoseNet, as a lightweight network, uses the full learning rate; AFNet and for novel tasks that require sufficient learning, use the full learning rate. This module-specific learning rate strategy can take into account the different needs of each module and improve the overall training efficiency.
[0140] In addition to the basic training strategy, a series of optimization techniques are also adopted to further improve the model performance and training efficiency. These techniques are aimed at the particularity of the endoscopic scenario and the complexity of the multi-module network, and their effectiveness has been verified through a large number of experiments.
[0141] An appropriate batch size is crucial for model training. Considering the complexity of the network and GPU memory limitations, the following strategies are adopted: Set the batch size to 12 to balance training stability and computational efficiency; for memory-constrained situations, use gradient accumulation techniques. Through multiple small-batch forward and backward propagations, accumulate gradients and then update the model, which is equivalent to using a larger batch size; utilize mixed-precision training with FP16 (half-precision floating-point) and FP32 (single-precision floating-point) to reduce memory occupancy and accelerate calculations while maintaining numerical stability.
[0142] Data augmentation is an important means to improve the generalization ability of the model. For the endoscopic scenario, specific data augmentation strategies are designed: Geometric transformations include random cropping, horizontal flipping, and random rotation (±10°). These transformations keep the spatial structure unchanged and help improve the robustness to perspective changes; photometric transformations include random brightness, contrast, and hue adjustments to simulate the lighting changes in the endoscopic scenario and enhance the model's adaptability to brightness changes; ensure that the same data augmentation is applied to adjacent frames within the same batch to maintain the temporal consistency between frames.
[0143] To prevent overfitting and improve the generalization ability of the model, multiple regularization techniques are adopted: Set an appropriate weight decay coefficient (such as 1e-4) in the optimizer to introduce L2 regularization for model parameters and suppress overly large weight values; apply Dropout (such as a rate of 0.3) to the key layers of the decoder. By randomly turning off some neurons, prevent the network from relying too much on specific features; apply batch normalization after the convolutional layer to accelerate training and improve model stability.
[0144] Effective training monitoring is crucial for timely detecting problems and avoiding overfitting: Evaluate the model performance on the validation set after each epoch and track the changes in key metrics such as Abs Rel, Sq Rel, and RMSE; when the performance metrics on the validation set do not improve for 5 consecutive epochs, automatically stop training to prevent overfitting and save computational resources; save the model weights with the best performance on the validation set for final evaluation and application.
[0145] Complex lighting and occlusions in the endoscopic scenario may lead to unstable training. The following measures are taken to ensure training stability: Limit the gradient norm within a reasonable range (such as 5.0) to prevent gradient explosion; monitor the loss value and gradients in real-time. When NaN values are detected, automatically skip the current batch or roll back to the previous stable checkpoint; identify and filter out abnormal samples in the training set, such as black screen frames, severely blurred frames, or frames containing large areas of highlights.
[0146] Optimizing the training time is crucial for rapid iteration and experimentation. Make full use of the GPU parallel computing power to optimize the data loading and preprocessing pipeline. For large-scale datasets, adopt a multi-GPU parallel training strategy to linearly accelerate the training process. Reduce redundant computations, optimize the memory access pattern, and improve the computing efficiency.
[0147] In summary, a scientific and reasonable training strategy is crucial for this improved self-supervised framework. The proposed two-stage training scheme, combined with a carefully designed parameter update strategy, learning rate adjustment method, and multiple optimization techniques, jointly ensures the excellent performance of the model in the endoscopic scenario, providing a solid foundation for high-precision depth estimation and three-dimensional reconstruction.
[0148] Compared with the prior art, the present invention uses the SCARED (Stereo Correspondence And Reconstruction of Endoscopic Data) dataset, which is a high-quality endoscopic stereo vision dataset designed specifically for evaluating stereo matching and three-dimensional reconstruction algorithms in the endoscopic scenario. The data is from fresh porcine abdominal anatomical specimens collected using the Da Vinci Xi surgical system, and it contains 35 endoscopic video sequences with real annotations of point clouds and camera poses. The original image resolution is 1280×1024 pixels and is adjusted to 320×256 pixels in the experiment. Referring to the Eigen-Zhou evaluation protocol, the dataset is divided into a training set (15,351 frames), a validation set (1,705 frames), and a test set (551 frames). The effective depth range is from 10 mm to 150 mm, covering the common operating distances in endoscopic surgery.
[0149] The framework is implemented based on PyTorch and consists of four core modules: The Depth Module (DepthNet) adopts an encoder-decoder architecture. The encoder uses ResNet-18 with the fully connected layers removed, and the decoder uses a part of Monodepth2. The Pose Module (PoseNet) adopts a lightweight CNN network and also refers to Monodepth2. The Appearance Module (AFNet) has a structure similar to DepthNet but outputs a three-channel appearance flow field. The Registration Module (OFNet) shares the same architecture as AFNet but the prediction layer outputs a two-channel optical flow field.
[0150] The training adopts a two-stage scheme, using the Adam optimizer with parameters set as β1 = 0.9, β2 = 0.99, a batch size of 12, and an initial learning rate of 1e-4, which is reduced to 0.1 times the original after 10 epochs of training. The first stage (OFNet training) is 20 epochs, and the second stage (joint training) is 20 epochs. The experiments are conducted on a server equipped with an NVIDIA 3090 graphics card.
[0151] To verify the feasibility of the method proposed in the present invention, the calculation results of the loss function in the training process are visualized as follows Figure 3 as shown; comparing Figure 3 and Figure 4 it can be seen that after adding the registration module, although the initial loss value is high, it drops rapidly and the drop amplitude is larger, indicating that the optimization direction is correct. Then, in the later stage, the oscillation amplitude is smaller, showing a more stable convergence process, and finally reaching a stable state with a low loss.
[0152] From Figure 5 it can be seen that in the regions significantly affected by light intensity, the intermediate feature maps of the network show higher interest, especially being more sensitive to changes in light intensity at small scales, and at the same time also showing good capture effects on tissue regions such as blood vessels. It is proved that the optical flow network can indeed assist the appearance flow network in capturing regions with obvious brightness change differences, and also has obvious recognition of different texture features and organ tissues. Table 1 shows the quantitative comparison results on the SCARED dataset.
[0153] Table 1 Quantitative comparison of depth estimation on the SCARED dataset
[0154]
[0155] It can be observed from Table 1 that the appearance flow-based method proposed in the present invention is significantly superior to the existing self-supervised depth estimation methods in all evaluation metrics. Particularly noteworthy is the significant improvement in the Sq Rel metric, indicating that the method of the present invention has obvious advantages in dealing with large error regions (usually regions with drastic brightness changes), which fully proves that after introducing the registration module, it can well assist the appearance flow in dealing with complex brightness changes in the endoscopic scene.
[0156] The qualitative results are as Figure 5 shown. The method of the present invention successfully restores the depth details of complex anatomical structures, while the original method produces over-smoothed or discontinuous results. This advantage is mainly attributed to the introduction of the registration module, which effectively compensates for the influence of brightness changes, enabling the network to focus on learning the true geometric structure.
[0157] To verify the effectiveness and contribution of each component in the improved self-supervised framework, a series of ablation experiments were carried out to systematically analyze the influence of different modules and design choices on the model performance.
[0158] Finally, the effects of different training strategies on performance were analyzed. Table 2 shows the results of the relevant ablation experiments. DPNet represents the combination of DepthNet and PoseNet; DPANet represents the combination of DepthNet, PoseNet, and AFNet; F represents freezing the parameters of OFNet in the second stage; J represents the joint update method; A represents the alternating update method.
[0159] Table 2 Ablation Experiments on Training Strategies
[0160]
[0161] As can be seen from Table 2, the two-stage training scheme (ID 4) is significantly better than the baseline method (ID 1) and other training strategies. This indicates that the strategy of training OFNet in the first stage and freezing the parameters of OFNet while training DepthNet, PoseNet, and AFNet in the second stage is the key to achieving the best performance. The staged training reduces the interference between modules, enabling each module to focus on its own task; the pre-trained OFNet provides stable registration and visibility information, creating favorable conditions for the joint optimization in the second stage.
Claims
1. A method for estimating the depth of an endoscopic image based on deep learning, characterized in that: The registration module is introduced to assist the training of the appearance flow network. The registration module predicts the optical flow between frames through the optical flow network OFNet and performs the automatic registration step. The registration module also generates a visibility mask to filter out pixels that are blocked or out of view. Through the improved self-supervisory framework, OFNet does not require real optical flow labels and learns high-quality motion information to align adjacent frames, reduce the impact of endoscopic camera motion, and improve the accuracy of appearance flow estimation, so that it can effectively cope with complex brightness changes in the endoscopic environment and provide more accurate depth and posture estimation for endoscopic images. The core of the registration module is the optical flow network OFNet, which is used to predict forward and backward optical flows. OFNet maintains the same architecture as AFNet. In the OFNet prediction layer: OFNet outputs two-channel optical flow, namely horizontal and vertical displacements; OFNet receives adjacent frame pairs of spliced endoscopic images as input and outputs optical flow fields at four different resolutions; the optical flow field prediction layer uses a linear activation function, allowing displacement representation without range restrictions: F of (p)=x Among them, F of (p) is the optical flow vector at pixel p and x is the raw output of the network.
2. The method for estimating the depth of an endoscope image based on deep learning according to claim 1, characterized in that: The automatic registration process performed by the registration module is as follows: Through the registration process guided by optical flow, the posture component is minimized and the brightness change is highlighted; the specific process is as follows: OFNet predicts the forward optical flow and backward optical flow in, Represents OFNet; using forward optical flow and spatial transformation network, reconstruct the target frame: in, Represents an image deformation operation based on optical flow.
3. The method for estimating the depth of an endoscope image based on deep learning according to claim 1, characterized in that: The registration module generates a mask V(p) based on the range map check to filter out unreliable areas: V(u,v)=[R(u,v)>0.95] Among them, R(u,v) is the range map, which indicates the possibility that the pixel (u,v) is not blocked. and They represent the horizontal and vertical components of the backward optical flow, respectively, and W and H represent the width and height of the endoscope image.
4. The method for estimating depth of an endoscope image based on deep learning according to claim 1, characterized in that: The endoscope posture component and brightness change are separated by the optical flow network OFNet; forward optical flow Including rigid motion F rigid and the non-rigid deformation F non-rigid : Through optical flow registration, the rigid component is approximately eliminated: The remaining difference is caused by brightness changes and provides an ideal input for appearance flow prediction: C δ (p)=I t (p)-I recon (p) OFNet achieves the function of separating posture and brightness changes by predicting optical flow, and minimizes the rigid motion component through registration, and the remaining difference C δ (p) Responds to brightness changes and provides clean input for AFNet.
5. The method for estimating depth of an endoscope image based on deep learning according to claim 1, characterized in that: The workflow of the improved self-supervised framework for endoscopic images includes the following steps: (1) Input processing: receiving the endoscope image target frame I t and source frame I s As input, in the training phase, the source frame includes the adjacent frames before and after the target frame, i.e., I s ∈{I t-1 ,I t+1 }; (2) Geometric transformation estimation: The depth module predicts the target frame I of the endoscopic image t The depth map D t , the pose module estimates the relative pose M between the target frame and the source frame t→s , the key supervisory signal comes from the distortion-based view synthesis; once the per-pixel depth value of the target frame is estimated, the pixels on the image plane are back-projected to the 3D camera space using the known camera intrinsic parameters; using the estimated self-motion, the 3D point cloud is projected onto another image plane; the view synthesis is expressed as: Among them, K represents the camera internal parameter; and h(p t ) represent the uniform pixel coordinates in the target view t and the source view s respectively; D t represents the target frame depth map; (3) View synthesis: Using rigid flow and spatial transformation network, synthesize the target frame from the source frame to obtain the initial synthesized frame I s→t ; (4) Appearance flow prediction: The appearance module predicts the appearance flow C based on the target frame and the source frame. δ , and align the brightness conditions of different frames through the brightness correction procedure: Among them, I t (p) represents the target frame; C δ (p) represents the appearance flow; I s (p) represents the source frame; (5) Correspondence enhancement: The registration module predicts the forward and backward optical flows through the optical flow network OFNet, performs the automatic registration step, and generates the visibility mask V(p) of the endoscopic image to filter out pixels that are occluded or beyond the field of view; (6) Loss calculation and optimization: Based on the synthetic frame, the corrected target frame, and the visibility mask, the data fidelity loss is calculated: in, represents data fidelity loss; V(p) represents visibility mask; Φ represents image similarity metric function; Edge-aware smoothing loss Encourage depth to be discontinuous at color edges and continuous in smooth areas: in, Represents the first-order gradient of the depth map; Photometric error loss between synthetic frame and true target frame Among them, α represents the weight coefficient; SSIM(I a ,I b ) represents the structural similarity coefficient; I a I represents the target frame; b represents a synthetic frame; Residual-based smoothing loss in, represents the first-order gradient of the appearance flow; |I t (p)-I s→t (p)|approximately reflects I t (p) The degree of brightness change of each pixel; The complete self-supervised depth estimation loss function is:
Citation Information
Cited By
Data processing method, equipment and program product for renal artery FFR evaluation
CN121147168A
Model training method, task processing method based on optical flow estimation and related products
CN121147668A