Monocular endoscope real-time depth estimation method and system based on self-supervised learning
Through the depth map generation network and the flow map generation network based on self-supervised learning, combining multi-scale feature fusion and camera internal reference information, the problem of difficulty in obtaining depth information in endoscopic surgery is solved, and high-precision and real-time monocular endoscopic depth estimation is achieved, which enhances the accuracy of surgical operations and reduces costs.
Patent Information
- Application Number
- CN202510380967.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
Smart Images

Figure CN120219364A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and medical image processing, and particularly relates to a monocular endoscope real-time depth estimation method and system based on self-supervised learning. Background Art
[0002] Endoscope technology transmits the images inside the closed and narrow natural cavity to an external display through a high-definition camera, enabling the operator to clearly observe the target area and perform precise operations. The application of robot-assisted endoscope technology can greatly improve the stability and efficiency of endoscope technology. However, it is difficult for the two-dimensional images obtained by the endoscope to reflect the spatial depth information, so it is relatively difficult to locate and identify the three-dimensional anatomical scene in robot-assisted endoscope surgery. Depth information is the distance between the three-dimensional space points corresponding to each pixel in the image obtained mainly by depth estimation technology and the camera lens, which can provide rich three-dimensional scene information for endoscope surgery. Currently, depth estimation technology can be divided into three categories: depth estimation based on structured light devices, depth estimation based on binocular images, and depth estimation based on monocular images.
[0003] The structured light device projects light rays with an intensity pattern of spatial variation or color coding onto the scene, and then uses a monocular camera to analyze the projected pattern. If the camera detects a plane, the pattern observed by the camera will be similar to the projected structured light pattern. In the case of surface deformation, the camera records a deformed projected pattern, and based on this deformed pattern, the depth of each point can be determined and the three-dimensional shape of the projected surface can be reconstructed. Although the structured light device can perform depth estimation of the endoscope-assisted surgery scene at high speed and accurately, the structured light device often requires external hardware support and is limited by narrow spaces.
[0004] Considerable achievements have been made in endoscope depth estimation based on stereo matching methods, but this method still faces many challenges. First, the use of binocular endoscopes will increase the surgical cost. Moreover, the generalization performance of the depth estimation method based on binocular image stereo matching has not been proven yet. Currently, the inference speed of the most efficient stereo matching method is only 11 frames per second, and the real-time performance is still insufficient.
[0005] Considering the cost and the characteristics of the narrow endoscope surgery scene, the depth estimation technology based on monocular images has become the main way to obtain depth information in the endoscope surgery scene. However, the depth estimation technology based on monocular images often can only obtain limited characterization information from monocular images, resulting in limited technical accuracy. And the artificial intelligence model depends on data-driven, so the three-dimensional perception model based on the artificial intelligence model often requires a large amount of data sets. And due to the diversity of different types of endoscope scenes, deep learning models often face the challenge of limited generalization.
[0006] Therefore, there is an urgent need to propose a monocular endoscope real-time depth estimation method and system based on self-supervised learning to achieve high-precision and generalizable real-time pose tracking and depth estimation based on monocular images. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a monocular endoscope real-time depth estimation method and system based on self-supervised learning to solve the problems existing in the above prior art.
[0008] To achieve the above object, the present invention provides a monocular endoscope real-time depth estimation method based on self-supervised learning, including the following steps:
[0009] Input a reference image and a target image obtained by the endoscope;
[0010] Perform downsampling, image block segmentation and stitching on the input target image to obtain an image block sequence;
[0011] Construct a depth map generation network and input the image block sequence into the depth map generation network;
[0012] Convert the image block sequence into corresponding feature encoding blocks through a visual encoder, perform image block fusion to obtain a multi-scale feature map, and use a decoder to output a corresponding depth map based on the multi-scale feature map;
[0013] Use an encoder to extract features from the input reference image and target image to obtain a feature map;
[0014] Construct an appearance flow map generation network and an optical flow map generation network, and input the feature map into the appearance flow map generation network and the optical flow map generation network respectively to obtain optical flow maps and appearance flow maps at corresponding scales;
[0015] Construct an ego-motion estimation network, input the feature map into the ego-motion estimation network, and obtain the relative motion transformation of the endoscope in two frames of images;
[0016] Input the feature map into a camera intrinsic decoder to obtain the endoscope focal length parameter and the endoscope compensation parameter, and further obtain the camera intrinsic matrix;
[0017] Reconstruct the target image based on the depth map, relative motion transformation and camera intrinsic matrix to obtain a reconstructed image;
[0018] Output a brightness-calibrated image based on the brightness registration of the reference image and the appearance flow map;
[0019] Obtain an optical flow generation map and a visibility mask based on the optical flow map and the reference image;
[0020] Combined with the visibility mask, the depth map generation network is optimized by supervising the consistency between the generated optical flow map and the reconstructed map and the image after brightness calibration;
[0021] Based on the optimized depth map generation network, real-time depth estimation of the monocular endoscope is completed.
[0022] Optionally, the process of obtaining the image patch sequence by performing downsampling, image patch segmentation, and stitching on the input target map includes:
[0023] Downsample the input target map to obtain low-resolution images with different resolutions;
[0024] Based on a sliding window, segment the input images with different resolutions into image patches, obtain multiple groups of image patches with the same size, and stitch them to obtain an image patch sequence.
[0025] Optionally, the process of converting the image patch sequence into corresponding feature encoding blocks through a visual encoder and performing image patch fusion to obtain a multi-scale feature map:
[0026] Convert the image patch sequence into corresponding feature encoding blocks through a visual encoder composed of 18 VMamba encoding modules, batch them into the first encoding block, the second encoding block, and the third encoding block, and perform image patch fusion respectively to obtain the corresponding first multi-scale feature map, second multi-scale feature map, and third multi-scale feature map; at the same time, let the first 25 encoding blocks output by the 6th VMamba encoding module and the 12th VMamba encoding module of the visual encoder be the fourth encoding block and the fifth encoding block respectively, and fuse them into the corresponding fourth multi-scale feature map and fifth multi-scale feature map; select the lowest-resolution low-resolution image from the obtained low-resolution images and process it based on the visual encoder to obtain a global feature map.
[0027] Optionally, the formulas for obtaining the optical flow map and appearance flow map at the corresponding scale are as follows:
[0028]
[0029] Among them, is a 3×3 convolution operation, and are the ELU and tanh activation functions respectively, A r→t is the appearance flow map, O r→t is the optical flow map.
[0030] Optionally, the formula for obtaining the relative motion transformation of the endoscope in two frames of images is as follows:
[0031]
[0032] Among them, is an n×n convolution operation, is the ReLU activation function, T r→t is the relative motion transformation of the endoscope.
[0033] Optionally, input the feature map into the camera intrinsic parameter decoder, and the formula for obtaining the endoscope focal length parameter is as follows:
[0034]
[0035] The formula for obtaining the endoscope compensation parameter is as follows:
[0036]
[0037] where * represents element-wise multiplication of vectors, is a 1×1 convolution operation with an output channel number of 2, W and H are the width and height of the original image obtained by the endoscope, respectively, [f x , f y is the endoscope focal length parameter, [c x , c y is the endoscope compensation parameter.
[0038] The present invention also provides a monocular endoscope real-time depth estimation system based on self-supervised learning for implementing the above method, including: an image input module, an image preprocessing module, a depth map generation module, an optical flow map and appearance flow map generation module, a relative motion transformation generation module, a camera intrinsic matrix generation module, a network optimization module, and a depth estimation module;
[0039] The image input module is used to input a reference image and a target image obtained by the endoscope;
[0040] The image preprocessing module is used to perform downsampling, image block segmentation, and splicing on the input target image to obtain an image block sequence;
[0041] The depth map generation module is used to construct a depth map generation network, input the image block sequence into the depth map generation network, convert it into corresponding feature encoding blocks through a visual encoder, and perform image block fusion to obtain a multi-scale feature map, and output a corresponding depth map based on the multi-scale feature map using a decoder;
[0042] The optical flow map and appearance flow map generation module is used to extract features from the input reference image and target image using an encoder to obtain a feature map; construct an appearance flow map generation network and an optical flow map generation network, and input the feature map into the appearance flow map generation network and the optical flow map generation network respectively to obtain optical flow maps and appearance flow maps at corresponding scales;
[0043] The relative motion transformation generation module is used to construct an endoscopic motion estimation network, input the feature map into the endoscopic motion estimation network, and obtain the relative motion transformation of the endoscope in two frames of images;
[0044] The camera intrinsic matrix generation module is used to input the feature map into the camera intrinsic decoder, obtain the endoscope focal length parameter and the endoscope compensation parameter, and further obtain the camera intrinsic matrix;
[0045] The network optimization module is used to reconstruct the target map based on the depth map, relative motion transformation and camera intrinsic matrix to obtain a reconstructed map; output a brightness-calibrated image based on the brightness registration of the reference map and the appearance flow map; obtain an optical flow generation map and a visibility mask based on the optical flow map and the reference map; combine the visibility mask, and optimize the depth map generation network by supervising the consistency between the optical flow generation map and the reconstructed map and the brightness-calibrated image;
[0046] The depth estimation module is used to complete real-time depth estimation of a monocular endoscope based on the optimized depth map generation network.
[0047] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method.
[0048] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.
[0049] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.
[0050] Compared with the prior art, the present invention has the following advantages and technical effects:
[0051] Through the method of self-supervised learning and multi-network collaboration, the present invention realizes real-time depth estimation of a monocular endoscope. It utilizes multi-scale feature fusion, appearance flow and optical flow information, and the camera intrinsic matrix, significantly improving the accuracy and robustness of depth estimation. At the same time, by combining brightness calibration and visibility mask to optimize the network, the image quality and the accuracy of depth estimation are further improved. In addition, this method adopts an efficient network architecture, which can quickly output depth information in medical scenarios with high real-time requirements, provide accurate three-dimensional vision assistance, enhance the precision of surgical operations, and at the same time reduce the equipment cost and complexity, and has good application prospects. Description of the Drawings
[0052] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the accompanying drawings:
[0053] Figure 1 It is a framework diagram of a monocular endoscope depth estimation network according to an embodiment of the present invention;
[0054] Figure 2 It is a network structure diagram of body motion estimation and camera internal parameter estimation according to an embodiment of the present invention. Detailed implementation manners
[0055] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the accompanying drawings and in combination with the embodiments.
[0056] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0057] Embodiment 1
[0058] As Figure 1 shown, this embodiment provides a method for real-time monocular endoscope depth estimation based on self-supervised learning, including the following steps:
[0059] Input the reference image and the target image obtained by the endoscope;
[0060] Perform downsampling, image block segmentation and splicing processing on the input target image to obtain an image block sequence;
[0061] Construct a depth map generation network and input the image block sequence into the depth map generation network;
[0062] Convert the image block sequence into corresponding feature encoding blocks through a visual encoder, and perform image block fusion to obtain a multi-scale feature map, and use a decoder to output a corresponding depth map based on the multi-scale feature map;
[0063] Use an encoder to extract features from the input reference image and target image to obtain a feature map;
[0064] Construct an appearance flow map generation network and an optical flow map generation network, and input the feature map into the appearance flow map generation network and the optical flow map generation network respectively to obtain optical flow maps and appearance flow maps of corresponding scales;
[0065] Construct a body motion estimation network, input the feature map into the body motion estimation network, and obtain the relative motion transformation of the endoscope in two frames of images;
[0066] Input the feature map into the camera intrinsic decoder to obtain the endoscope focal length parameter and the endoscope compensation parameter, and then obtain the camera intrinsic matrix;
[0067] Based on the depth map, relative motion transformation, and camera intrinsic matrix, reconstruct the target map to obtain the reconstructed map;
[0068] Based on the luminance registration of the reference map and the appearance flow map, output the luminance-calibrated image;
[0069] Based on the optical flow map and the reference map, obtain the optical flow generation map and the visibility mask;
[0070] Combined with the visibility mask, optimize the depth map generation network by supervising the consistency between the optical flow generation map, the reconstructed map, and the luminance-calibrated image;
[0071] Based on the optimized depth map generation network, complete the real-time depth estimation of the monocular endoscope.
[0072] Implementable, the process of downsampling, image block segmentation, and splicing the input target map to obtain the image block sequence includes:
[0073] Downsample the input target map to obtain low-resolution images with different resolutions; based on a sliding window, segment the input images with different resolutions into image blocks, obtain multiple groups of image blocks with the same size, and splice them to obtain the image block sequence.
[0074] Implementable, the process of converting the image block sequence into corresponding feature encoding blocks through a visual encoder and performing image block fusion to obtain the multi-scale feature map:
[0075] Convert the image block sequence into corresponding feature encoding blocks through a visual encoder composed of 18 VMamba encoding modules, batch them into the first encoding block, the second encoding block, and the third encoding block, and perform image block fusion respectively to obtain the corresponding first multi-scale feature map, second multi-scale feature map, and third multi-scale feature map; at the same time, let the first 25 encoding blocks output by the 6th VMamba encoding module and the 12th VMamba encoding module of the visual encoder be the fourth encoding block and the fifth encoding block respectively, and fuse them into the corresponding fourth multi-scale feature map and fifth multi-scale feature map; select the lowest-resolution low-resolution image from the obtained low-resolution images and process it based on the visual encoder to obtain the global feature map.
[0076] As an implementable way, input image I ×1 First, obtain the low-resolution images I ×0.5 and I ×0.25 which are reduced to one-half and one-fourth of the original size by downsampling.. Then, based on three input images with different resolutions, the method performs image patch segmentation based on a sliding window to obtain three sets of image patches P ×1 , P ×0.5 and P ×0.25 , and stitches them together to obtain an image patch sequence P. A depth map generation network is constructed, and the image patch sequence P is input into the depth map generation network. The image patch sequence P is transformed into corresponding feature encoding blocks f through a visual encoder based on Vision Mamba (VMamba). The method batches the feature encoding blocks into f3, f2, and f1, and performs image patch fusion respectively to obtain corresponding multi-scale feature maps F3, F2, and F1. At the same time, this method fuses the first 25 encoding blocks f4 and f5 of the outputs of the 6th VMamba encoder and the 12th VMamba encoder of ViT into corresponding feature maps F4 and F5 respectively. To provide global correlation features, the proposed network also directly uses the low-resolution image P ×0.25 to obtain a global feature map F0 based on the VMamba visual encoder. Then, the dimensions of the multi-scale feature maps change to F0, F1, F2, F3, F4, and F5 after resampling. Finally, based on the input of the multi-scale feature maps F0, F1, F2, F3, F4, and F5, the DPT decoder outputs the corresponding depth map D t .
[0077] Implementable, construct an appearance flow map generation network and an optical flow map generation network, and input the feature maps into the appearance flow map generation network and the optical flow map generation network respectively to obtain optical flow maps and appearance flow maps of corresponding scales.
[0078] As an implementable approach, based on the input reference image I r and the target image I t , the ConvNext encoder outputs the feature map F. In the decoder part, the network uses the processing shown in formula (1) to output the optical flow map and the shape flow map of the corresponding scale of the feature map. is a 3×3 convolution operation, and are the ELU and tanh activation functions respectively.
[0079]
[0080] Implementable, construct an endoscopic motion estimation network, and input the feature map into the endoscopic motion estimation network to obtain the relative motion transformation of the endoscope in two frames of images.
[0081] As a specific implementation, input the feature map F, and the pose decoder obtains the relative motion transformation T of the endoscope in two frames of images through the convolutional neural network operation shown in formula (2) r→t . Among them, is an n×n convolution operation, is the ReLU activation function.
[0082]
[0083] Furthermore, as Figure 2 shown, in this embodiment, the feature map F is input to the camera intrinsic decoder. In this decoder, the input features are subjected to global average pooling and the operation shown in formula (3) to obtain the endoscope focal length parameters [f x , f y . At the same time, the input features are subjected to global average pooling and the operation shown in formula (4) to obtain the endoscope compensation parameters [c x , c y . Wherein, * represents element-wise multiplication of vectors, is a 1×1 convolution operation with an output channel number of 2, and W and H are the width and height of the original image obtained by the endoscope respectively.
[0084]
[0085] Based on novel view synthesis, this embodiment can utilize the depth map, the relative motion transformation of the endoscope, and the camera intrinsic matrix K = [f x , f y , c x , c y to reconstruct the target map and obtain the reconstructed map I r→t . Based on the luminance registration of the reference map and the appearance flow map, the luminance-calibrated image I t +A r→t can be output. And based on the optical flow map and the reference map, this embodiment can obtain the corresponding optical flow generation map and the visibility mask M t . Finally, by combining the visibility mask and supervising the consistency between the optical flow generation map and the reconstructed map I r→t and the calibrated target image I t +A r→t , this embodiment optimizes the depth map generation network.
[0086] Furthermore, after optimizing the depth map generation network, the following steps are further included:
[0087] Training and verification: Apply the optimized depth map generation network to the image datasets of multiple endoscopic surgery scenarios, evaluate the performance of the network through the validation set, and adjust the network parameters to improve its generalization ability and accuracy;
[0088] Iterative update: According to the verification results, further adjust the structure or parameters of the depth map generation network, and re-train and verify until the network performance meets the preset accuracy and real-time requirements;
[0089] Performance testing: In the actual endoscopic surgery simulation environment, conduct real-time performance testing on the optimized depth map generation network to ensure that it can quickly and accurately output depth maps in actual applications and be seamlessly integrated with the endoscopic surgery system.
[0090] Furthermore, after optimizing the depth map generation network, it also includes the steps of evaluating and feedback on the optimization results:
[0091] Evaluate the depth map quality: By comparing with the annotated data with known depths, calculate the error metrics between the optimized depth map and the real depth, such as root mean square error (RMSE), mean absolute error (MAE), etc., to evaluate the accuracy of the depth map;
[0092] Real-time evaluation: Measure the frame rate and latency of the optimized depth map generation network during actual operation to ensure that it can meet the requirements of real-time depth estimation in endoscopic surgery;
[0093] Feedback and adjustment: According to the evaluation results, adjust the training strategy, network structure or hyperparameters of the depth map generation network to further improve the network performance.
[0094] Furthermore, after optimizing the depth map generation network, it also includes the steps of applying the optimized network to the actual endoscopic surgery system:
[0095] System integration: Integrate the optimized depth map generation network into the endoscopic surgery navigation system so that it can receive endoscopic images in real-time and output depth information to provide three-dimensional scene perception support for surgical operations;
[0096] User interaction: Provide a user interface for surgical operators to display the depth map and depth-related surgical navigation information in real-time, helping the operators better understand and operate the endoscopic surgical instruments;
[0097] System testing and optimization: Conduct comprehensive testing on the integrated system in the actual surgical scenario, collect user feedback, and further optimize the system performance and user experience.
[0098] Embodiment 2
[0099] This embodiment also provides a monocular endoscope real-time depth estimation system based on self-supervised learning for implementing the above method, including: an image input module, an image preprocessing module, a depth map generation module, an optical flow map and appearance flow map generation module, a relative motion transformation generation module, a camera intrinsic matrix generation module, a network optimization module, and a depth estimation module;
[0100] The image input module is used to input a reference image and a target image acquired by the endoscope;
[0101] The image preprocessing module is used to perform downsampling, image block segmentation, and stitching on the input target image to obtain an image block sequence;
[0102] The depth map generation module is used to construct a depth map generation network, input the image block sequence into the depth map generation network, convert it into corresponding feature encoding blocks through a visual encoder, and perform image block fusion to obtain a multi-scale feature map, and use a decoder to output a corresponding depth map based on the multi-scale feature map;
[0103] The optical flow map and appearance flow map generation module is used to extract features from the input reference image and target image using an encoder to obtain feature maps; construct an appearance flow map generation network and an optical flow map generation network, and input the feature maps into the appearance flow map generation network and the optical flow map generation network respectively to obtain optical flow maps and appearance flow maps of corresponding scales;
[0104] The relative motion transformation generation module is used to construct an ego-motion estimation network, input the feature map into the ego-motion estimation network, and obtain the relative motion transformation of the endoscope in two frames of images;
[0105] The camera intrinsic matrix generation module is used to input the feature map into a camera intrinsic decoder to obtain the endoscope focal length parameter and the endoscope compensation parameter, and further obtain the camera intrinsic matrix;
[0106] The network optimization module is used to reconstruct the target image based on the depth map, relative motion transformation, and camera intrinsic matrix to obtain a reconstructed image; output a brightness-calibrated image based on the brightness registration of the reference image and the appearance flow map; obtain an optical flow generation map and a visibility mask based on the optical flow map and the reference image; combine the visibility mask, and optimize the depth map generation network by supervising the consistency between the optical flow generation map and the reconstructed image and the brightness-calibrated image;
[0107] The depth estimation module is used to complete the real-time depth estimation of the monocular endoscope based on the optimized depth map generation network.
[0108] Embodiment III
[0109] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method.
[0110] Embodiment 4
[0111] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.
[0112] Embodiment 5
[0113] This embodiment also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.
[0114] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A real-time depth estimation method for monocular endoscope based on self-supervised learning, characterized in that: The following steps are involved: Input the reference image and target image acquired by the endoscope; Down-sample, segment and splice the input target image to obtain an image block sequence; Construct a deep map generation network and input the image block sequence into the deep map generation network; The image block sequence is converted into corresponding feature coding blocks through a visual encoder, and image blocks are fused to obtain a multi-scale feature map. The decoder is used to output the corresponding depth map based on the multi-scale feature map. Using the encoder to extract features from the input reference image and target image to obtain a feature map; Construct an appearance flow map generation network and an optical flow map generation network, input the feature map into the appearance flow map generation network and the optical flow map generation network respectively, and obtain the optical flow map and appearance flow map of the corresponding scale; Construct a proprioceptive motion estimation network, input the feature map into the proprioceptive motion estimation network, and obtain the relative motion transformation of the endoscope in the two frames of images; The feature map is input into the camera intrinsic parameter decoder to obtain the endoscope focal length parameter and the endoscope compensation parameter, and then the camera intrinsic parameter matrix is obtained; Reconstructing the target image based on the depth image, relative motion transformation and camera intrinsic parameter matrix to obtain a reconstructed image; Outputting a brightness-calibrated image based on brightness registration of the reference reference and the appearance flow map; Obtaining an optical flow generation map and a visibility mask based on the optical flow map and the reference; In combination with the visibility mask, the depth map generation network is optimized by supervising the consistency of the optical flow generated map and the reconstructed map with the brightness calibrated image; Real-time depth estimation of monocular endoscope is completed based on the optimized depth map generation network.
2. The method according to claim 1, characterized in that The process of downsampling, image block segmentation and splicing the input target image to obtain an image block sequence includes: Downsample the input target image to obtain low-resolution images of different resolutions; Based on the sliding window, the input images of different resolutions are divided into image blocks to obtain multiple groups of image blocks of the same size, which are then spliced to obtain image block sequences.
3. The method according to claim 2, characterized in that The process of converting the image block sequence into the corresponding feature coding block through the visual encoder, and fusing the image blocks to obtain the multi-scale feature map: The image block sequence is converted into corresponding feature coding blocks through a visual encoder composed of 18 VMamba coding modules, and is divided into the first coding block, the second coding block and the third coding block in batches, and the image blocks are fused respectively to obtain the corresponding first multi-scale feature map, the second multi-scale feature map and the third multi-scale feature map; at the same time, the first 25 coding blocks output by the 6th VMamba coding module and the 12th VMamba coding module of the visual encoder are respectively made the fourth coding block and the fifth coding block, and are fused into the corresponding fourth multi-scale feature map and the fifth multi-scale feature map respectively; the low-resolution image with the lowest resolution is selected from the obtained low-resolution images for processing based on the visual encoder to obtain a global feature map.
4. The method according to claim 1, characterized in that: The formula for obtaining the optical flow map and appearance flow map of the corresponding scale is as follows: in, is a 3×3 convolution operation, and are ELU and tanh activation functions respectively, A r→t is the appearance flow graph, O r→t It is an optical flow graph.
5. The method according to claim 1, characterized in that The formula for obtaining the relative motion transformation of the endoscope in two frames of images is as follows: in, is an n×n convolution operation, is the ReLU activation function, T r→t is the relative motion transformation of the endoscope.
6. The method according to claim 1, characterized in that The feature map is input into the camera intrinsic decoder, and the formula for obtaining the endoscope focal length parameter is as follows: The formula for obtaining the endoscope compensation parameters is as follows: Among them, * means vector multiplication by element, is a 1×1 convolution operation with an output channel number of 2, W and H are the width and height of the original image obtained by the endoscope, respectively, [f x ,f y ] is the focal length parameter of the endoscope, [c x ,c y ] is the endoscope compensation parameter.
7. A monocular endoscope real-time depth estimation system based on self-supervised learning, characterized in that: Used to implement the method described in any one of claims 1 to 6, comprising: an image input module, an image preprocessing module, a depth map generation module, an optical flow map and appearance flow map generation module, a relative motion transformation generation module, a camera intrinsic parameter matrix generation module, a network optimization module and a depth estimation module; The image input module is used to input a reference image and a target image acquired by the endoscope; The image preprocessing module is used to perform downsampling, image block segmentation and splicing processing on the input target image to obtain an image block sequence; The depth map generation module is used to construct a depth map generation network, input the image block sequence into the depth map generation network, convert it into a corresponding feature coding block through a visual encoder, and perform image block fusion to obtain a multi-scale feature map, and use a decoder to output a corresponding depth map based on the multi-scale feature map; The optical flow map and appearance flow map generation module is used to use the encoder to extract features from the input reference image and target image to obtain a feature map; construct an appearance flow map generation network and an optical flow map generation network, input the feature map into the appearance flow map generation network and the optical flow map generation network respectively, and obtain an optical flow map and an appearance flow map of corresponding scales; The relative motion transformation generation module is used to construct a body motion estimation network, input the feature map into the body motion estimation network, and obtain the relative motion transformation of the endoscope in two frames of images; The camera intrinsic parameter matrix generation module is used to input the feature map into the camera intrinsic parameter decoder to obtain the endoscope focal length parameter and the endoscope compensation parameter, and then obtain the camera intrinsic parameter matrix; The network optimization module is used to reconstruct the target image based on the depth map, relative motion transformation and camera intrinsic parameter matrix to obtain a reconstructed image; output a brightness-calibrated image based on brightness registration of the reference map and the appearance flow map; obtain an optical flow generation map and a visibility mask based on the optical flow map and the reference map; and optimize the depth map generation network by supervising the consistency of the optical flow generation map and the reconstructed map with the brightness-calibrated image in combination with the visibility mask; The depth estimation module is used to complete the real-time depth estimation of the monocular endoscope based on the optimized depth map generation network.
8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Training method of optical flow estimation model, optical flow estimation method, device and equipment
CN121259486A
Training method of optical flow estimation model, optical flow estimation method, device and equipment
CN121259486B