Monocular endoscope three-dimensional pose real-time tracking method and system based on deep learning

Through the three-dimensional pose real-time tracking method of monocular endoscopic endoscopic 3D poses, the optical flow estimation network and deep learning framework are used to solve the problem of three-dimensional scene perception and navigation difficulties in endoscopic surgery, achieving high-precision and real-time pose tracking, adapting to complex environments, and reducing dependence on traditional positioning technology.

CN120219477AActive Publication Date: 2025-06-27JIANGTAI INTELLIGENT TECHNOLOGY (SUZHOU) CO LTD

Patent Information

Application Number
CN202510296446.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In existing endoscopic surgery, the natural cavity space is narrow, the instrument operation space and freedom are limited, and the two-dimensional images acquired by the endoscopy are difficult to feedback spatial depth and hierarchical information, resulting in difficulty in three-dimensional scene perception and navigation.

Method used

Using a three-dimensional pose real-time tracking method of monocular endoscopic 3D poses, a deep learning framework is used to estimate the optical flow estimation network and deep learning framework, including graph convolution neural network, attention mechanism and multi-scale depth separable convolution, scene features, motion features and joint features are extracted, relative pose transformation vectors of the endoscopic, and absolute pose transformation vectors are estimated, and absolute 3D poses are iteratively determined.

Benefits of technology

It improves the accuracy of endoscopic pose tracking, realizes real-time pose updates, adapts to complex cavity environments, reduces dependence on traditional positioning technology, and enhances the robustness and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219477A_ABST
    Figure CN120219477A_ABST
Patent Text Reader

Abstract

The invention discloses a monocular endoscope three-dimensional pose real-time tracking method and system based on deep learning, and the method comprises the steps: splicing two input images to obtain a spliced image, and outputting a corresponding optical flow graph based on the two input images through an optical flow estimation network; extracting scene features of the two input images based on a feature extractor, and extracting motion features of the optical flow graph; meanwhile, joint features of the spliced images are extracted based on a multi-dimensional feature extractor; after splicing all the extracted features, estimating a relative pose transformation vector of the endoscope at corresponding moments of the two frames of images based on a pose decoder; and iteratively determining the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, thereby realizing real-time tracking of the endoscope. According to the invention, through deep learning and multi-feature fusion technologies, the precision and real-time performance of three-dimensional pose tracking of the monocular endoscope are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of medical imaging technology and endoscopic pose tracking technology, and particularly relates to a method and system for real-time tracking of the three-dimensional pose of a monocular endoscope based on deep learning. Background Art

[0002] Endoscopic technology transmits in-vivo images to an external display through a high-definition camera, enabling operators to clearly observe the target area and perform operations. In recent years, with the rapid development of technology, new endoscopic surgical robot systems based on endoscopic technology, such as Medtronic's Hugo RAS and CMR Surgical's Versius, have emerged continuously. Robot-assisted endoscopic surgery has become the mainstream method of minimally invasive endoscopic surgery. At the same time, the rise of embodied intelligence technology has also promoted the application of artificial intelligence models in robot systems.

[0003] However, endoscopic surgery faces many challenges. The natural cavity is narrow, the operating space and degrees of freedom of instruments are limited, and the two-dimensional images obtained by the endoscope are difficult to reflect spatial depth and hierarchical information, often accompanied by image distortion. Therefore, three-dimensional scene perception and navigation of the endoscope have become a key task for achieving precise and safe surgery. In recent years, the application of extended reality (XR) technology in surgical navigation has gradually increased. By registering preoperative imaging data with intraoperative endoscopic images, more intuitive three-dimensional visual information is provided. However, the accuracy of registering endoscopic videos with preoperative scan data highly depends on intraoperative three-dimensional scene perception.

[0004] Endoscopic pose tracking technology is an important means to achieve intraoperative three-dimensional scene perception. The endoscopic pose refers to the three-dimensional spatial coordinates and posture of the endoscope in the reference system of the surgical object. Currently, pose tracking technology is mainly divided into three categories: pose tracking based on optical positioning, pose tracking based on electromagnetic positioning, and pose tracking based on vision. Optical positioning technology often fails to track due to occlusion by personnel and instruments; electromagnetic positioning technology is easily interfered by metal objects and electromagnetic devices in the scene. Therefore, the vision-based endoscopic pose tracking technology has become a feasible way to obtain endoscopic pose information due to its adaptability in complex surgical scenarios.

[0005] In summary, there is an urgent need to propose a method and system for real-time tracking of the three-dimensional pose of a monocular endoscope based on deep learning to accurately estimate the endoscopic pose through deep learning technology. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a method and system for real-time tracking of the three-dimensional pose of a monocular endoscope based on deep learning to solve the problems existing in the above prior art.

[0007] To achieve the above object, the present invention provides a method for real-time tracking of the three-dimensional pose of a monocular endoscope based on deep learning, including the following steps:

[0008] Set the initial pose of the endoscope and collect the endoscope image sequence in real time;

[0009] For the current frame, splice two input images to obtain a spliced image, and use an optical flow estimation network to output a corresponding optical flow map based on the two input images;

[0010] Construct a deep learning framework, which includes a feature extractor based on a graph convolutional neural network, a multi-dimensional feature extractor based on an attention mechanism, and a pose decoder based on multi-scale depthwise separable convolution;

[0011] The feature extractor based on the graph convolutional neural network extracts the scene features of the two input images and the motion features of the optical flow map; at the same time, the multi-dimensional feature extractor based on the attention mechanism extracts the joint features of the spliced image;

[0012] After splicing all the extracted features, the pose decoder based on multi-scale depthwise separable convolution estimates the relative pose transformation vector of the endoscope at the corresponding moments of the two frames of images;

[0013] Based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, iteratively determine the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image to achieve real-time tracking of the endoscope.

[0014] Optionally, the optical flow estimation network adopts a FlowNet network, and the architecture is an encoder-decoder structure.

[0015] Optionally, the process of using the optical flow estimation network to output a corresponding optical flow map based on the two input images includes:

[0016] The encoder performs multi-layer convolution and pooling processing on the two input images to obtain corresponding features, and the decoder performs feature fusion and upsampling processing on the obtained features and then maps them back to the optical flow field to obtain a corresponding optical flow map.

[0017] Optionally, the graph convolutional neural network uses a model pre-trained on the ImageNet dataset.

[0018] Optionally, the process of the multi-dimensional feature extractor based on the attention mechanism extracting the joint features of the spliced image includes:

[0019] Input the spliced image into the self-attention module to obtain multiple feature maps with different channel dimension information through flipping operations; input the flipped feature maps into the channel attention mechanism module, and the channel attention mechanism module performs global average pooling and global max pooling operations on the flipped feature maps simultaneously, then calculates the attention weights for the results of global average pooling and global max pooling, and finally obtains the joint features of the spliced image.

[0020] Optionally, based on the relative pose transformation vector and the absolute three-dimensional pose of the previous frame The iterative formula for iteratively determining the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame image is as follows:

[0021]

[0022] Where, is the relative translation transformation vector, is the relative rotation transformation vector, is the pose vector at time i, is the known initial pose vector, is the position coordinate vector at time i-k, is the rotation vector at time i-k.

[0023] The present invention also provides a real-time monocular endoscope three-dimensional pose tracking system based on deep learning. Based on the above method, it includes: an image acquisition module, an image preprocessing module, a learning framework construction module, a feature extraction module, a pose decoding module, and a pose tracking module;

[0024] The image acquisition module is used to set the initial pose of the endoscope and collect the endoscope image sequence in real time;

[0025] The image preprocessing module is used for the current frame to splice two input images to obtain a spliced image, and use the optical flow estimation network to output the corresponding optical flow map based on the two input images;

[0026] The learning framework construction module is used to construct a deep learning framework based on the feature extractor of the graph convolutional neural network, the multi-dimensional feature extractor based on the attention mechanism, and the pose decoder based on the multi-scale depthwise separable convolution;

[0027] The feature extraction module is used to extract the scene features of the two input images based on the feature extractor of the graph convolutional neural network, and extract the motion features of the optical flow map; at the same time, the multi-dimensional feature extractor based on the attention mechanism extracts the joint features of the spliced image;

[0028] The pose decoding module is used to splice all the extracted features, and then estimate the relative pose transformation vector of the endoscope at the corresponding moments of two frames of images based on the pose decoder of multi-scale depthwise separable convolution;

[0029] The pose tracking module is used to iteratively determine the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, so as to realize the real-time tracking of the endoscope.

[0030] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method.

[0031] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.

[0032] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.

[0033] Compared with the prior art, the present invention has the following advantages and technical effects:

[0034] 1. Improve the accuracy of endoscope pose tracking:

[0035] By combining the optical flow estimation network and the deep learning framework, this method can extract rich scene features and motion features from endoscope images. The optical flow map provides pixel-level motion information, while the graph convolutional neural network and the attention mechanism further enhance the feature representation ability. This multi-feature fusion method can significantly improve the accuracy of endoscope pose tracking, especially in complex human body cavity environments.

[0036] 2. Realize real-time pose update:

[0037] This method updates the absolute three-dimensional pose of the endoscope in real time through an iterative formula, and can quickly respond to the motion changes of the endoscope. This real-time performance is crucial for endoscope navigation, can provide stable and reliable pose information for the endoscope navigation system, can operate the endoscope more accurately, and reduce the risk of misoperation.

[0038] 3. Adapt to complex cavity environments:

[0039] This method automatically learns the features of endoscope images through a deep learning network, and can adapt to different cavity environments and tissue types. In addition, the optical flow estimation network and the attention mechanism can process the noise and distortion in the images, further enhancing the robustness of the system in complex environments.

[0040] 4. Reduce the dependence on traditional positioning technologies:

[0041] Compared with pose tracking technologies based on optical or electromagnetic positioning, this method relies entirely on the endoscopic images themselves and is not affected by occlusion or electromagnetic interference in the examination scene. This makes this method more advantageous in the actual examination environment and can provide more stable and reliable pose tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0043] Figure 1 It is a schematic diagram of absolute pose tracking based on the initial pose and the trained deep learning network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0045] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0046] Embodiment 1

[0047] The traditional methods of endoscopic navigation can be divided into optical positioning tracking and electromagnetic positioning tracking, both of which have been widely used in the navigation system. However, optical positioning tracking is easily affected by the occlusion of the human body and equipment, while electromagnetic positioning tracking is easily interfered by the electromagnetic fields of certain medical instruments. At the same time, the deployment of optical positioning tracking or electromagnetic positioning tracking systems with multiple sensors and computing devices is expensive and complex in the actual scenario. Therefore, vision-based real-time endoscopic pose tracking with lower cost and more convenient configuration has become an important task for endoscopic navigation. The goal of pose estimation is to estimate the pose of an object from the scene image containing the object, which is impossible for closed and narrow human body cavities.

[0048] To solve this problem, as Figure 1 shown, this embodiment provides a method for real-time tracking of the three-dimensional pose of a monocular endoscope based on deep learning, including the following steps:

[0049] Set the initial pose of the endoscope and collect the endoscopic image sequence in real time;

[0050] For the current frame, two input images are spliced to obtain a spliced image, and an optical flow estimation network is used to output the corresponding optical flow map based on the two input images;

[0051] A deep learning framework is constructed, and the deep learning framework includes a feature extractor based on a graph convolutional neural network, a multi-dimensional feature extractor based on an attention mechanism, and a pose decoder based on multi-scale depthwise separable convolution;

[0052] The feature extractor based on the graph convolutional neural network extracts the scene features of the two input images and the motion features of the optical flow map; meanwhile, the multi-dimensional feature extractor based on the attention mechanism extracts the joint features of the spliced image;

[0053] After splicing all the extracted features, the pose decoder based on multi-scale depthwise separable convolution estimates the relative pose transformation vector of the endoscope at the corresponding moments of the two frames of images;

[0054] Based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image is iteratively determined to achieve real-time tracking of the endoscope.

[0055] As a specific implementation manner, first, the initial position of the endoscope p0 = [t0, q0] is given. For the image I obtained in the i-th frame i (i = 1, 2,...), the deep learning network framework is applied, and the image I obtained by the endoscope in the (i - k)-th frame sampled last time is used i-k and the corresponding estimated pose The relative pose transformation of the endoscope corresponding to the two frames of images is estimated Then, according to the iterative formula (1), the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image can be determined The effect of real-time tracking is achieved.

[0056]

[0057] Among them, is the relative translation transformation vector, is the relative rotation transformation vector, is the pose vector at the i-th moment, is the known initial pose vector, is the position coordinate vector at the (i - k)-th moment, is the rotation vector at the (i - k)-th moment.

[0058] As a specific implementation manner, the constructed deep learning framework includes multiple feature extractors and a pose decoder, which are used to, according to two adjacent frames I i-k and Ii to predict the relative pose transformation of the endoscope First, in the preprocessing stage, two input images are stitched to obtain the stitched image I c , and at the same time, the efficient optical flow estimation network outputs the corresponding optical flow map I based on the two input images o . In the feature extraction stage, the feature extractor based on the graph convolutional neural network extracts the corresponding scene features f i-k and f i in the two input images, as well as the motion feature f o contained in the optical flow map. At the same time, based on the stitched image I c , the multi-dimensional feature extractor based on the attention mechanism extracts the joint feature f c of the two frame observation images. After all the features are stitched, they are input into the pose decoder based on the multi-scale depthwise separable convolution , and the decoder outputs the relative pose transformation vector of the endoscope at the corresponding moments of the two frames of images

[0059] Furthermore, the optical flow estimation network in the preprocessing stage uses the FlowNet network, which adopts the encoder-decoder architecture to estimate the optical flow. The specific process includes: inputting adjacent image frames I1 and I2, the encoder extracts features through multi-layer convolution and pooling to obtain F1 and F2, and the decoder fuses F1 and F2 and maps the features back to the optical flow field through upsampling and other operations, and outputs the corresponding optical flow map

[0060] Furthermore, the feature extractor based on the graph convolutional neural network uses the model pre-trained on the ImageNet dataset

[0061] Furthermore, the process of the multi-dimensional feature extractor based on the attention mechanism extracting the joint feature of the two frame observation images includes:

[0062] Input the extracted features into the self-attention mechanism to calculate the relationship between the features. First, the input feature map obtains three feature maps with different channel dimension information through the flipping operation and The three flipped feature maps respectively pass through the channel attention mechanism module to calculate the attention weights. Specifically, the proposed channel attention mechanism module first performs global average pooling and global max pooling operations on the input feature map, and then calculates the attention weights of the results of global average pooling and global max pooling to obtain the final joint feature

[0063] Embodiment 2

[0064] This embodiment also provides a real-time monocular endoscope three-dimensional pose tracking system based on deep learning. Based on the above method, it includes: an image acquisition module, an image preprocessing module, a learning framework construction module, a feature extraction module, a pose decoding module, and a pose tracking module;

[0065] The image acquisition module is used to set the initial pose of the endoscope and collect the endoscope image sequence in real time;

[0066] The image preprocessing module is used for the current frame to splice two input images to obtain a spliced image, and output the corresponding optical flow map based on the two input images by using an optical flow estimation network;

[0067] The learning framework construction module is used to construct a deep learning framework based on a feature extractor of a graph convolutional neural network, a multi-dimensional feature extractor based on an attention mechanism, and a pose decoder based on a multi-scale depthwise separable convolution;

[0068] The feature extraction module is used to extract the scene features of two input images based on the feature extractor of the graph convolutional neural network, and extract the motion features of the optical flow map; at the same time, the multi-dimensional feature extractor based on the attention mechanism extracts the joint features of the spliced image;

[0069] The pose decoding module is used to splice all the extracted features, and then estimate the relative pose transformation vector of the endoscope at the corresponding moments of two frames of images based on the pose decoder based on the multi-scale depthwise separable convolution;

[0070] The pose tracking module is used to iteratively determine the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame of image based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, so as to realize the real-time tracking of the endoscope.

[0071] Embodiment III

[0072] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0073] Embodiment IV

[0074] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above method.

[0075] Embodiment V

[0076] This embodiment also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the above method.

[0077] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for real-time tracking of three-dimensional pose of a monocular endoscope based on deep learning, characterized in that: The following steps are involved: Set the initial posture of the endoscope and collect the endoscope image sequence in real time; For the current frame, the two input images are spliced ​​to obtain a spliced ​​image, and the optical flow estimation network is used to output the corresponding optical flow map based on the two input images; Constructing a deep learning framework, the deep learning framework includes a feature extractor based on a graph convolutional neural network, a multi-dimensional feature extractor based on an attention mechanism, and a pose decoder based on multi-scale deep separable convolution; The feature extractor based on graph convolutional neural network extracts scene features of the two input images and motion features of the optical flow map. At the same time, the multi-dimensional feature extractor based on the attention mechanism extracts joint features of the spliced ​​images. After concatenating all the extracted features, the pose decoder based on multi-scale depthwise separable convolution estimates the relative pose transformation vector of the endoscope at the corresponding moment of the two frames of images. Based on the estimated relative pose transformation vector and the absolute three-dimensional pose of the previous frame, the absolute three-dimensional pose of the endoscope at the corresponding moment of each frame image is iteratively determined to achieve real-time tracking of the endoscope.

2. The method according to claim 1, characterized in that The optical flow estimation network adopts the FlowNet network, and the architecture is an encoder-decoder structure.

3. The method according to claim 2, characterized in that The process of using the optical flow estimation network to output the corresponding optical flow map based on two input images includes: The encoder performs multi-layer convolution and pooling processing on the two input images to obtain the corresponding features. The decoder performs feature fusion and upsampling on the obtained features and maps them back to the optical flow field to obtain the corresponding optical flow map.

4. The method according to claim 1, characterized in that: The graph convolutional neural network uses a model pre-trained on the ImageNet dataset.

5. The method according to claim 1, characterized in that The process of extracting joint features of spliced ​​images by the multi-dimensional feature extractor based on the attention mechanism includes: The spliced ​​image is input into the self-attention module, and multiple feature maps with different channel dimension information are obtained through flipping operation; the flipped feature map is input into the channel attention mechanism module, and the channel attention mechanism module simultaneously performs global average pooling and global maximum pooling operations on the flipped feature map, and then calculates the attention weights of the results of global average pooling and global maximum pooling, and finally obtains the joint features of the spliced ​​image.

6. The method according to claim 1, characterized in that Based on the relative posture transformation vector and the absolute three-dimensional posture of the previous frame, the iterative formula for iteratively determining the absolute three-dimensional posture of the endoscope at the corresponding moment of each frame image is as follows: in, is the relative translation transformation vector, is the relative rotation transformation vector, is the pose vector at time i, is the known initial pose vector, is the position coordinate vector at time ik, is the rotation vector at time ik, and the relative posture transformation vector is The absolute 3D pose of the previous frame is 7. A real-time tracking system for three-dimensional posture of a monocular endoscope based on deep learning, characterized in that: The method according to any one of claims 1 to 6 comprises: an image acquisition module, an image preprocessing module, a learning framework construction module, a feature extraction module, a posture decoding module and a posture tracking module; The image acquisition module is used to set the initial posture of the endoscope and acquire an endoscopic image sequence in real time; The image preprocessing module is used to splice the two input images to obtain a spliced ​​image for the current frame, and output a corresponding optical flow map based on the two input images using an optical flow estimation network; The learning framework building module is used to build a deep learning framework based on a feature extractor of a graph convolutional neural network, a multidimensional feature extractor based on an attention mechanism, and a pose decoder based on multi-scale deep separable convolution; The feature extraction module is used to extract scene features of two input images based on the feature extractor of the graph convolutional neural network, and extract motion features of the optical flow map; at the same time, the multi-dimensional feature extractor based on the attention mechanism extracts the joint features of the spliced ​​images; The posture decoding module is used to estimate the relative posture transformation vector of the endoscope at the corresponding time of two frames of images based on a posture decoder of multi-scale depth-separable convolution after splicing all the extracted features; The posture tracking module is used to iteratively determine the absolute three-dimensional posture of the endoscope at the corresponding moment of each frame image based on the estimated relative posture transformation vector and the absolute three-dimensional posture of the previous frame, so as to realize real-time tracking of the endoscope.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Focus tracking method under digestive endoscope based on sequential feature learning

    CN111915573A

  • Cavity three-dimensional reconstruction method based on colorectum monocular endoscope video

    CN116363318A

  • Target object tracking method, device and system, medium and computing equipment

    CN116999163A

  • Method, device and equipment for calculating absolute and relative poses of endoscope in operation

    CN117671012A

  • Nasal cavity endoscopic surgery navigation system based on scene reconstruction

    CN118766589A

Cited By

  • Endoscope rotation image processing method and endoscope

    CN122089840A