Unsupervised endoscope depth estimation method and system
Through the light correction network and uncertainty mask combined with the three-dimensional dynamic convolution module, the problem of degradation of depth estimation due to abnormal light in the endoscopic environment is solved, and the depth estimation accuracy and stability is achieved, which is suitable for medical imaging processing and surgical navigation.
Patent Information
- Application Number
- CN202510363336.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-15
AI Technical Summary
The existing unsupervised monocular depth estimation method has reduced depth estimation accuracy due to abnormal light in the endoscopic environment, especially in local abnormal light areas (overexposed or underexposed), and it is impossible to effectively estimate depth information.
Light correction network (ICN) is used for lighting correction, combining uncertainty mask (UM) and three-dimensional dynamic convolution module (TDC), the uncertainty mask is calculated through the lighting correction ratio, blocking the interference information of the abnormal lighting area, and improving the local feature extraction capability through three-dimensional dynamic convolution.
The robustness and accuracy of depth estimation are improved, especially on the SCARED dataset, the accuracy is improved by 5.1%, and the error (AbsRel) is reduced by 11.5%, adapting to complex lighting environments, and improving the extraction ability of local features.
Smart Images

Figure CN120495400A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to an unsupervised endoscope depth estimation method and system. Background Art
[0002] Most existing endoscopes use unsupervised monocular depth estimation (MDE) methods, which infer the depth information of a scene based on a single RGB image frame, usually relying on reprojection error optimization. Unsupervised methods can adapt to different endoscopic scenarios and equipment, have strong generalization capabilities, and are computationally efficient, making them suitable for real-time applications such as surgical navigation and robotic-assisted surgery. They can also generate 3D models of endoscopic scenes to assist doctors in diagnosis and surgical planning.
[0003] Unsupervised endoscopic depth estimation methods are of great significance in reducing annotation dependence, improving generalization ability, real-time applications, 3D reconstruction, surgical navigation, robot-assisted surgery, medical image analysis and academic research, and have broad application prospects and important practical value.
[0004] Current unsupervised monocular depth estimation methods primarily rely on the photometric consistency assumption, inferring depth information through image reconstruction. However, in endoscopic images, the lighting conditions are extremely unstable due to the synchronous motion of the camera and light source, making this assumption difficult to hold. Furthermore, existing methods (such as AF-SfmLearner and IID-SfmLearner) perform poorly in areas with abnormal lighting, making them unable to effectively estimate depth information.
[0005] Some studies have attempted to compensate for illumination variations through optical flow correction or intrinsic image decomposition, but these methods still fail to address the problem of localized abnormal illumination (overexposure or underexposure). Furthermore, existing methods fail to fully integrate local and global features, resulting in inaccurate depth estimation in detailed areas (such as tissue edges). Summary of the Invention
[0006] The technical problem to be solved by the embodiments of the present invention is to provide an unsupervised endoscopic depth estimation method and system to solve the problem of decreased depth estimation accuracy caused by abnormal lighting in an endoscopic environment in traditional methods.
[0007] In order to solve the above technical problems, an embodiment of the present invention proposes an unsupervised endoscope depth estimation method, comprising: Step 1: Acquire an endoscopic video sequence to obtain continuous frame images; select the target frame image and adjacent reference frame images to obtain the corresponding camera pose; Step 2: Calculate the illumination correction ratio, perform illumination correction on the target frame image, and obtain a corrected image; Step 3: Calculate the uncertainty mask based on the illumination correction ratio; Step 4: performing depth estimation on the remaining parts of the corrected image except the mask to obtain the depth map and camera pose of the target frame; Step 5: Project the reference frame to the target frame perspective to generate a reconstructed target frame image; compare the pixel difference between the reconstructed image and the target frame, calculate the final loss function, and optimize and output the reconstructed target frame image based on the final loss function.
[0008] Accordingly, an embodiment of the present invention further provides an unsupervised endoscope depth estimation system, comprising: Acquisition module: collects endoscopic video sequences and obtains continuous frame images; selects target frame images and adjacent reference frame images to obtain the corresponding camera pose; ICN module: calculates the illumination correction ratio, performs illumination correction on the target frame image, and obtains a corrected image; UM module: calculates the uncertainty mask based on the illumination correction ratio; TDC module: performs depth estimation on the remaining parts of the corrected image except the mask to obtain the depth map and camera pose of the target frame; Optimization module: Projects the reference frame to the target frame perspective to generate a reconstructed target frame image; compares the pixel difference between the reconstructed image and the target frame, calculates the final loss function, and optimizes and outputs the reconstructed target frame image based on the final loss function.
[0009] The beneficial effects of the present invention are: The present invention combines an illumination calibration network (ICN), an uncertainty mask, and a three-dimensional dynamic convolution module (TDC), effectively solving the problem of decreased depth estimation accuracy caused by abnormal illumination (overexposure and underexposure) in traditional methods under endoscopic conditions. By introducing an illumination correction network (ICN), the present invention can adapt to complex lighting environments and perform illumination equalization on endoscopic images, thereby improving the robustness of depth estimation. The uncertainty mask further improves the accuracy of depth estimation and can intelligently shield interference information in areas with abnormal illumination, making the final predicted depth map more accurate. In addition, the three-dimensional dynamic convolution module (TDC) combines the advantages of CNN and Transformer, enabling the model to better capture depth information of tissue edges and detail areas, improving the ability to extract local features. Experimental results demonstrate that this method surpasses the state-of-the-art (SOTA) methods in multiple depth estimation metrics tested on the SCARED and Hamlyn datasets. On the SCARED dataset, accuracy is improved by 5.1% and the error (AbsRel) is reduced by 11.5%. This technical solution has broad application prospects in fields such as medical image processing and surgical navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 4 is a flow chart of an unsupervised endoscopic depth estimation method according to an embodiment of the present invention.
[0011] Figure 2 FIG. 4 is a schematic diagram of a processing flow of a TDC module according to an embodiment of the present invention.
[0012] Figure 3 4 is a schematic diagram of the processing flow of the unsupervised endoscopic depth estimation system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0013] It should be noted that, unless there is a conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention is further described in detail below with reference to the drawings and specific embodiments.
[0014] In the embodiments of the present invention, if there are directional indications (such as up, down, left, right, front, back, etc.), they are only used to explain the relative position relationship and movement status of the various components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0015] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.
[0016] Please refer to Figure 1 The unsupervised endoscope depth estimation method of the embodiment of the present invention includes steps 1 to 5.
[0017] Step 1: Acquire the endoscopic video sequence to obtain continuous frame images; select the target frame image and adjacent reference frame images to obtain the corresponding camera pose.
[0018] Step 2: Calculate the illumination correction ratio, perform illumination correction on the target frame image, and obtain a corrected image.
[0019] As an implementation method, in step 2, an illumination correction network is used to perform illumination correction, and the illumination correction ratio is predicted by the illumination correction network. The illumination correction network adopts a multi-stage illumination self-correction strategy to gradually optimize the illumination correction ratio to make the corrected image more stable.
[0020] Assume that the input endoscopic image is , ICN performs illumination correction through the following steps: Calculate the lighting correction ratio: ; Among them, Ray represents the illumination correction ratio, which is predicted by the ICN network.
[0021] A multi-stage illumination self-correction strategy is used to gradually optimize the illumination correction ratio, making the corrected image more stable: ; Among them, ϕ is responsible for illumination estimation and ψ is responsible for self-correction mapping.
[0022] Step 3: Calculate the uncertainty mask based on the illumination correction ratio.
[0023] Based on the illumination correction ratio Ray generated by ICN, calculate the uncertainty mask: ; Where: α and β are illumination boundaries, used to distinguish normal and abnormal illumination areas; k1 and k2 are attenuation coefficients, which control the smoothness of the mask; r is the actual calculated illumination correction ratio.
[0024] Step 4: Depth estimation is performed on the remaining parts of the rectified image except the mask to obtain the depth map and camera pose of the target frame.
[0025] Step 5: Project the reference frame to the target frame perspective to generate a reconstructed target frame image; compare the pixel difference between the reconstructed image and the target frame, calculate the final loss function, and optimize and output the reconstructed target frame image based on the final loss function.
[0026] The illumination-corrected image I′ is fed into the depth estimation network , generate depth map:
[0027] Using pose estimation network Calculate camera motion and generate reprojected images based on geometric transformations : ; Combined with the uncertainty mask to calculate the final loss function:
[0028] Among them, SSIM represents the structural similarity loss, λ1 and λ2 are scale parameters, and I s is the target frame image, I s→t is the reconstructed image, I t is the selected target frame image.
[0029] As an implementation method, in step 4, a preset three-dimensional dynamic convolution module is used for depth estimation. The preset three-dimensional dynamic convolution module includes a depth separation convolution layer, a three-dimensional dynamic convolution layer, and a feature fusion layer, wherein: The depth separation convolution layer convolves the input image to obtain the initial feature map; The three-dimensional dynamic convolution layer performs convolution operations on the initial feature map to obtain feature maps at different levels; The feature fusion layer fuses the feature maps of different levels to obtain a depth map; finally, the depth map is output.
[0030] In practice, a graph neural network (GNN)-based approach can replace the 3D dynamic convolution module (TDC). GNNs have stronger local information modeling capabilities, enabling more accurate extraction of tissue edges and texture information, further improving depth estimation accuracy.
[0031] As an implementation method, the three-dimensional dynamic convolution module uses three-dimensional dynamic convolution to calculate cross-level attention, and the calculation formula is as follows: ; Among them, E1 is the input channel attention, E2 is the convolution kernel attention, E3 is the output channel attention, I is the output feature map, I* is the input feature map, n is the number of convolution kernels, i∈n.
[0032] The unsupervised endoscopic depth estimation system of the embodiment of the present invention includes an acquisition module, an ICN module, a UM module, a TDC module, and an optimization module. The core of the unsupervised endoscopic depth estimation system of the present invention includes three key modules. The specific process is as follows: Figure 3 .
[0033] Acquisition module: acquires endoscopic video sequences and obtains continuous frame images; selects target frame images and adjacent reference frame images to obtain the corresponding camera pose.
[0034] ICN module: calculates the illumination correction ratio, performs illumination correction on the target frame image, and obtains a corrected image.
[0035] The ICN module, or Illumination Calibration Network (ICN), decomposes the image's illumination components and performs illumination correction on the original endoscopic image, resulting in a more uniform illumination distribution in the input image and thus improving the stability of depth estimation.
[0036] The ICN module uses an image enhancement method based on Retinex theory to extract and adjust the illumination component.
[0037] The ICN module is combined with a self-supervised training strategy to enable the network to automatically adapt to different lighting environments.
[0038] UM module: Calculates the uncertainty mask based on the illumination correction ratio.
[0039] The UM module is an uncertainty mask. The present invention addresses the depth estimation problem of overexposed and underexposed areas. The UM module performs an uncertainty mask method based on the illumination correction ratio.
[0040] The UM module calculates the illumination ratio of the corrected image to generate a soft mask, suppressing misleading information in areas of abnormal illumination. Compared to traditional hard masks (such as grayscale thresholding), this method can more flexibly integrate effective information and improve the accuracy of depth estimation.
[0041] TDC module: performs depth estimation on the parts of the corrected image except the mask to obtain the depth map and camera pose of the target frame.
[0042] The TDC module is the Three-Dimensional Dynamic Convolution (TDC). To improve the model's ability to capture features in local areas, the TDC module of this invention combines the advantages of CNN and Vision Transformer (ViT). The specific implementation details are as follows: Figure 2 shown.
[0043] The present invention enhances the model's ability to perceive local details through dynamic convolution. It embeds TDC at multiple levels of the Transformer, enabling a more efficient fusion of global information and local features.
[0044] The TDC module of the present invention includes a depth separation convolution layer, a three-dimensional dynamic convolution layer, and a feature fusion layer. The depth separation convolution layer convolves the input image to obtain an initial feature map; the three-dimensional dynamic convolution layer performs a convolution operation on the initial feature map to obtain feature maps of different levels; the feature fusion layer fuses the feature maps of different levels to obtain a depth map; and finally, the obtained depth map is output.
[0045] The TDC module uses three-dimensional dynamic convolution to calculate cross-level attention, and the calculation formula is as follows: ; Among them, E1 is the input channel attention, E2 is the convolution kernel attention, E3 is the output channel attention, I is the output feature map, I* is the input feature map, n is the number of convolution kernels, i∈n.
[0046] Optimization module: Projects the reference frame to the target frame perspective to generate a reconstructed target frame image; compares the pixel difference between the reconstructed image and the target frame, calculates the final loss function, and optimizes and outputs the reconstructed target frame image based on the final loss function.
[0047] As an implementation method, the ICN module uses an illumination correction network to perform illumination correction. The illumination correction ratio is predicted by the illumination correction network. The illumination correction network adopts a multi-stage illumination self-correction strategy to gradually optimize the illumination correction ratio to make the corrected image more stable.
[0048] The illumination correction network (ICN) of the present invention can also use a generative adversarial network (GAN)-based approach to enhance low-light areas, thereby improving the adaptability of illumination correction. This approach can utilize adversarial training mechanisms to enable the model to generate more natural and clear corrected images.
[0049] The UM module calculates the uncertainty mask according to the following formula: ; Among them, α and β are illumination boundaries, which are used to distinguish normal and abnormal illumination areas; k1 and k2 are attenuation coefficients, which control the smoothness of the mask; r is the actual calculated illumination correction ratio, n uc is the uncertainty mask; Ray min 、Ray max They are the minimum and maximum values of the illumination correction ratio, respectively.
[0050] During specific implementation, it is also possible to combine the dynamic mask strategy based on the attention mechanism to optimize the uncertainty mask, making the processing of abnormal lighting areas more intelligent, thereby reducing errors and improving the adaptability of the model.
[0051] As an implementation method, the optimization module calculates the final loss function according to the following formula: ; Among them, SSIM represents the structural similarity loss, λ1 and λ2 are scale parameters, and I s is the target frame image, I s→t is the reconstructed image, I t is the selected target frame image.
[0052] The present invention continuously improves the accuracy of depth estimation by optimizing the loss function, and has important application value in endoscopic surgery. It can provide doctors with richer spatial information and assist in surgical navigation and decision-making.
[0053] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An unsupervised endoscopic depth estimation method, characterized in that include: Step 1: Acquire an endoscopic video sequence to obtain continuous frame images; select the target frame image and adjacent reference frame images to obtain the corresponding camera pose; Step 2: Calculate the illumination correction ratio, perform illumination correction on the target frame image, and obtain a corrected image; Step 3: Calculate the uncertainty mask based on the illumination correction ratio; Step 4: performing depth estimation on the remaining parts of the corrected image except the mask to obtain the depth map and camera pose of the target frame; Step 5: Project the reference frame to the target frame perspective to generate a reconstructed target frame image; compare the pixel difference between the reconstructed image and the target frame, calculate the final loss function, and optimize and output the reconstructed target frame image based on the final loss function.
2. The unsupervised endoscopic depth estimation method according to claim 1, wherein: In step 4, a preset three-dimensional dynamic convolution module is used for depth estimation. The preset three-dimensional dynamic convolution module includes a depth separation convolution layer, a three-dimensional dynamic convolution layer, and a feature fusion layer, wherein: The depth separation convolution layer convolves the input image to obtain the initial feature map; The three-dimensional dynamic convolution layer performs convolution operations on the initial feature map to obtain feature maps at different levels; The feature fusion layer fuses the feature maps of different levels to obtain a depth map; finally, the depth map is output.
3. The unsupervised endoscopic depth estimation method according to claim 2, wherein: The three-dimensional dynamic convolution module uses three-dimensional dynamic convolution to calculate cross-level attention. The calculation formula is as follows: ; Among them, E1 is the input channel attention, E2 is the convolution kernel attention, E3 is the output channel attention, I is the output feature map, I* is the input feature map, n is the number of convolution kernels, i∈n.
4. The unsupervised endoscopic depth estimation method according to claim 1, wherein: In step 2, the illumination correction network is used to perform illumination correction. The illumination correction ratio is predicted by the illumination correction network. The illumination correction network adopts a multi-stage illumination self-correction strategy to gradually optimize the illumination correction ratio to make the corrected image more stable. In step 3, the uncertainty mask is calculated according to the following formula: ; Among them, α and β are illumination boundaries, which are used to distinguish normal and abnormal illumination areas; k1 and k2 are attenuation coefficients, which control the smoothness of the mask; r is the actual calculated illumination correction ratio, n uc is the uncertainty mask; Ray min 、Ray max They are the minimum and maximum values of the illumination correction ratio, respectively.
5. The unsupervised endoscopic depth estimation method according to claim 1, wherein: In step 5, the final loss function is calculated according to the following formula: ; Among them, SSIM represents the structural similarity loss, λ1 and λ2 are scale parameters, and I s is the target frame image, I s→t is the reconstructed image, I t is the selected target frame image.
6. An unsupervised endoscopic depth estimation system, characterized in that include: Acquisition module: collects endoscopic video sequences and obtains continuous frame images; selects target frame images and adjacent reference frame images to obtain the corresponding camera pose; ICN module: calculates the illumination correction ratio, performs illumination correction on the target frame image, and obtains a corrected image; UM module: calculates the uncertainty mask based on the illumination correction ratio; TDC module: performs depth estimation on the remaining parts of the corrected image except the mask to obtain the depth map and camera pose of the target frame; Optimization module: Projects the reference frame to the target frame perspective to generate a reconstructed target frame image; compares the pixel difference between the reconstructed image and the target frame, calculates the final loss function, and optimizes and outputs the reconstructed target frame image based on the final loss function.
7. The unsupervised endoscopic depth estimation system according to claim 6, wherein: The TDC module includes a depth separation convolution layer, a three-dimensional dynamic convolution layer, and a feature fusion layer, among which, The depth separation convolution layer convolves the input image to obtain the initial feature map; The three-dimensional dynamic convolution layer performs convolution operations on the initial feature map to obtain feature maps at different levels; The feature fusion layer fuses the feature maps of different levels to obtain a depth map; finally, the depth map is output.
8. The unsupervised endoscopic depth estimation system according to claim 7, wherein: The TDC module uses three-dimensional dynamic convolution to calculate cross-level attention, and the calculation formula is as follows: ; Among them, E1 is the input channel attention, E2 is the convolution kernel attention, E3 is the output channel attention, I is the output feature map, I* is the input feature map, n is the number of convolution kernels, i∈n.
9. The unsupervised endoscopic depth estimation system according to claim 6, wherein: The ICN module uses an illumination correction network to perform illumination correction. The illumination correction ratio is predicted by the illumination correction network. The illumination correction network adopts a multi-stage illumination self-correction strategy to gradually optimize the illumination correction ratio, making the corrected image more stable. The UM module calculates the uncertainty mask according to the following formula: ; Among them, α and β are illumination boundaries, which are used to distinguish normal and abnormal illumination areas; k1 and k2 are attenuation coefficients, which control the smoothness of the mask; r is the actual calculated illumination correction ratio, n uc is the uncertainty mask; Ray min 、Ray max They are the minimum and maximum values of the illumination correction ratio, respectively.
10. The unsupervised endoscopic depth estimation system according to claim 6, wherein: The optimization module calculates the final loss function according to the following formula: ; Among them, SSIM represents the structural similarity loss, λ1 and λ2 are scale parameters, and I s is the target frame image, I s→t is the reconstructed image, I t is the selected target frame image.