Trachea cannula navigation method and device based on image segmentation and three-dimensional reconstruction

Through navigation methods based on image segmentation and three-dimensional reconstruction, three-dimensional point cloud model and digital twin images inside the trachea are generated, which solves the problems of high operating risks and low accuracy of traditional tracheal intubation technology, and improves the safety and accuracy of the intubation process.

CN120182545APending Publication Date: 2025-06-20INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510389758.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional tracheal intubation technology has high operating risks and is difficult to accurately locate the intubation head, resulting in collateral damage and operational complexity.

Method used

Using a navigation method based on image segmentation and three-dimensional reconstruction, the three-dimensional point cloud model inside the trachea is generated through endoscopic image segmentation and monocular scene three-dimensional reconstruction, and a digital twin image of the intubation process is generated by the Unreal Engine virtual engine to provide positioning navigation prompts.

Benefits of technology

It improves the accuracy and safety of the intubation process, provides doctors with three-dimensional intuitive information and effective auxiliary guidance, and reduces operational risks. It is especially suitable for young doctors and students to learn intubation skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182545A_ABST
    Figure CN120182545A_ABST
Patent Text Reader

Abstract

The invention relates to a tracheal intubation navigation method and device based on image segmentation and three-dimensional reconstruction, and belongs to the field of wisdom medicines.The tracheal intubation navigation method comprises the steps that an endoscope image frame sequence in the intubation process is collected, and image segmentation is conducted on the endoscope image frame sequence to obtain an intubation head segmentation profile diagram; the method comprises the following steps: performing three-dimensional reconstruction on the interior of a trachea by using an endoscope image frame sequence and adopting three network branches of monocular scene three-dimensional reconstruction, inter-frame multi-view geometric pose estimation and optical flow matching three-dimensional reconstruction to obtain a three-dimensional point cloud model of the interior of the trachea; performing pose estimation on the head of the trachea cannula through the segmentation contour of the instrument at the head of the trachea cannula in the endoscope image to obtain a three-dimensional point cloud model of the head of the trachea cannula and instrument estimation pose sequence information; and inputting the trachea internal three-dimensional point cloud model, the intubation head three-dimensional point cloud model and instrument estimation pose sequence information into a Unreal Engine virtual engine to generate a digital twin dynamic navigation image in the intubation process.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Endotracheal intubation is a common medical procedure during anesthesia and first aid. However, traditional endotracheal intubation techniques have problems such as high operation risks caused by the operator's non-intuitive judgment of the position of the intubation head in the trachea and potential collateral injuries. Although some auxiliary techniques have been used to improve the accuracy of endotracheal intubation, they still need to be improved in terms of operation simplicity, real-time performance, and safety. In addition, due to possible occlusion of key tissues and the intubation head during the intubation process, as well as large changes in light reflection conditions, it is difficult to provide stable auxiliary information. Summary of the Invention

[0003] To solve the problem of positioning and navigation of the intubation head during endotracheal intubation, a tracheal intubation navigation method based on image segmentation and three-dimensional reconstruction is proposed. This method uses an in-loop active learning method to segment key tissues and the intubation head in the endoscopic images during the intubation process; for a series of image frames during the intubation process, single-frame depth estimation, multi-view geometric relationships between different frames, and optical flow estimation between adjacent frames are respectively used to perform three-dimensional reconstruction on the tracheal internal tissue structure part; the contour of the intubation head in the endoscopic image and the virtual data generated by the Unreal Engine three-dimensional engine are used to semi-supervisedly estimate the pose of the intubation head, and the pose of the intubation head is registered with the three-dimensional scene pose inside the trachea to generate a digital twin image of the intubation process and display positioning and navigation prompts for the doctor during the endotracheal intubation process.

[0004] The key point of this method is to design an intelligent navigation of the endotracheal intubation process using image segmentation and vision-based three-dimensional reconstruction techniques to provide auxiliary intelligent prompts for young doctors.

[0005] The present invention proposes a tracheal intubation navigation method based on image segmentation and three-dimensional reconstruction, including the following steps:

[0006] Step S1, collect a sequence of endoscopic image frames during the intubation process, and perform image segmentation on the sequence of endoscopic image frames to obtain a segmentation contour map of the intubation head;

[0007] Step S2, use the sequence of endoscopic image frames during the intubation process, and perform three-dimensional reconstruction on the tracheal interior using three network branches: monocular scene three-dimensional reconstruction, inter-frame multi-view geometric pose estimation, and optical flow matching three-dimensional reconstruction, to obtain a three-dimensional point cloud model of the tracheal interior;

[0008] Step S3, estimate the pose of the intubation head through the segmentation contour of the instrument of the tracheal intubation head in the endoscopic image, to obtain a three-dimensional point cloud model of the intubation head and sequence information of the estimated pose of the instrument;

[0009] Step S4: Input the three-dimensional point cloud model of the trachea interior generated in Step S2, the three-dimensional point cloud model of the intubation head obtained in Step S3, and the instrument estimated pose sequence information into the Unreal Engine virtual engine to generate a digital twin dynamic navigation image of the intubation process.

[0010] The present invention also provides a tracheal intubation navigation device based on image segmentation and three-dimensional reconstruction, comprising:

[0011] An endoscope image segmentation module, which collects the sequence of endoscope image frames during the intubation process and performs image segmentation on the sequence of endoscope image frames to obtain a segmentation contour map of the intubation head;

[0012] A three-dimensional point cloud model construction module, which uses the sequence of endoscope image frames during the intubation process and performs three-dimensional reconstruction of the trachea interior through three network branches of monocular scene three-dimensional reconstruction, inter-frame multi-view geometric pose estimation, and optical flow matching three-dimensional reconstruction to obtain a three-dimensional point cloud model of the trachea interior;

[0013] A pose sequence information acquisition module, which estimates the pose of the intubation head through the segmentation contour of the instrument at the tracheal intubation head in the endoscope image to obtain a three-dimensional point cloud model of the intubation head and instrument estimated pose sequence information;

[0014] A digital twin dynamic navigation image generation module, which inputs the three-dimensional point cloud model of the trachea interior generated in Step S2, the three-dimensional point cloud model of the intubation head obtained in Step S3, and the instrument estimated pose sequence information into the Unreal Engine virtual engine to generate a digital twin dynamic navigation image of the intubation process.

[0015] The present invention has the following beneficial effects: Through the present invention, three-dimensional intuitive information and effective auxiliary guidance can be provided for doctors during the intubation process, helping doctors to deal with difficult tracheal intubation cases, and at the same time contributing to the learning of intubation skills by students and young doctors. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:

[0017] Figure 1 is a schematic flow chart of the tracheal intubation navigation method based on image segmentation and three-dimensional reconstruction of the present invention.

[0018] Figure 2 is a schematic flow chart of semantic segmentation under the doctor-in-the-loop active learning framework.

[0019] Figure 3 is a schematic diagram of self-supervised monocular three-dimensional reconstruction of the airway internal structure.

[0020] Figure 4Schematic diagram of a robust estimation framework for the tracheal intubation head based on image segmentation and virtual data generation. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, the present invention adopts the following technical solutions.

[0022] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and not to limit the invention. In addition, it should be noted that only the parts related to the relevant invention are shown in the drawings for the convenience of description.

[0023] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0024] The tracheal intubation navigation method based on image segmentation and three-dimensional reconstruction proposed by the present invention has a general flowchart as Figure 1 shown, including:

[0025] Step S1, as Figure 2 shown, first collect the endoscopic image frame sequence during the intubation process. Segment and label the key tissues such as the tracheopharyngeal epiglottis and the intubation head instrument in a small number of image frames among them, finely tune and train an image segmentation network with the SAM2 architecture, use the finely tuned and trained SAM2 image segmentation network to segment the key tissues and instruments of all frames, and give the confidence of the segmentation result, and introduce doctors to manually annotate and correct the segmentation images with low confidence. And cycle through the above steps. Through this effective annotation strategy of the doctor-in-the-loop actively targeting difficult samples, the endoscopic images can be efficiently segmented on the premise of investing a small amount of medical staff labor costs.

[0026] Step S2, as Figure 3 shown, use the image sequence under the endoscope during the intubation process, and perform three-dimensional reconstruction of the trachea interior using three network branches: monocular scene three-dimensional reconstruction, inter-frame multi-view geometric pose estimation, and optical flow matching three-dimensional reconstruction, to obtain a three-dimensional point cloud model of the trachea interior. The monocular scene three-dimensional reconstruction branch inputs a single-frame endoscopic image , and uses the UNet network architecture to output the depth map estimated from a single image ; the inter-frame multi-view geometric pose estimation branch inputs an image pair composed of adjacent frames of multiple endoscopic images , local features of the images are matched through SuperGlue feature points, and the fundamental matrix of the image pair is estimated by minimizing the epipolar geometry constraint through feature point matching. , the relative pose rotation matrix between the image pair is obtained by using non-linear optimization and QR decomposition. and ; the optical flow matching 3D reconstruction branch inputs adjacent image frames, and the optical flow field between adjacent frames is estimated by FlowNet using the UNet network architecture. . By optimizing the following formula, the depth map output by the monocular scene 3D reconstruction branch , the image pair pose relationship output by the inter-frame multi-view geometry pose estimation branch, and the optical flow field output by the optical flow matching 3D reconstruction branch are optimized to obtain the 3D point cloud model of the trachea interior;

[0027] ,

[0028] where , is the weight coefficient, is the camera projection function that projects 3D space points onto a 2D image, and SSIM is the 2D image structural similarity metric.

[0029] Step S3, as Figure 4 shown, the pose of the intubation head is estimated through the segmentation contour of the tracheal intubation head instrument in the endoscopic image. The 3D point cloud model of the tracheal intubation head is obtained through offline scanning, and the 3D point cloud model of the intubation head is input into the virtual engine Unreal Engine. Virtual cameras with different poses are evenly distributed in the space around the intubation head, and instrument contour images in different poses are automatically generated through virtual camera rendering. Using these automatically generated virtual images and the corresponding pose information as inputs, a UNet network structure-based intubation head instrument pose estimation network is trained. In the application stage, the intubation head segmentation contour map obtained in Step S1 is input into the trained instrument estimation network to output the instrument estimation pose sequence information.

[0030] Step S4, the time endoscopic sequence is processed sequentially according to Steps S1, S2, and S3. The 3D point cloud model of the trachea interior generated in Step S2, the 3D point cloud model of the intubation head obtained in Step S3, and the pose sequence information are input into the Unreal Engine virtual engine. Different colors are assigned to different tissues inside the trachea in the virtual engine, and the intubation head is moved and rotated according to the intubation head instrument pose sequence information, and this scene is rendered to generate a digital twin dynamic navigation image of the intubation process, enabling the doctor to obtain an intuitive display result of the intubation process.

[0031] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described storage device and processing device can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0032] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable EPROM, electrically erasable programmable EEPROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0033] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.

[0034] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, so that a process, method, article, or device / equipment including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to these processes, methods, articles, or devices / equipment.

[0035] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A tracheal intubation navigation method based on image segmentation and three-dimensional reconstruction, characterized in that: The following steps are involved: Step S1, collecting an endoscopic image frame sequence during the intubation process, and performing image segmentation on the endoscopic image frame sequence to obtain an intubation head segmentation contour map; Step S2, using the endoscopic image frame sequence during the intubation process, three network branches including monocular scene 3D reconstruction, inter-frame multi-view geometric pose estimation, and optical flow matching 3D reconstruction are used to perform 3D reconstruction of the interior of the trachea to obtain a 3D point cloud model of the interior of the trachea; Step S3, estimating the position and posture of the endotracheal tube head by segmenting the contour of the instrument of the endotracheal tube head in the endoscopic image, and obtaining a three-dimensional point cloud model of the endotracheal tube head and the estimated position and posture sequence information of the instrument; Step S4: input the 3D point cloud model of the trachea generated in step S2, the 3D point cloud model of the intubation head obtained in step S3, and the estimated instrument posture sequence information into the Unreal Engine virtual engine to generate a digital twin dynamic navigation image of the intubation process.

2. The endotracheal intubation navigation method based on image segmentation and three-dimensional reconstruction according to claim 1, characterized in that: In step 1, the key tissues and cannula head instruments contained in the image frames are segmented and annotated, the image segmentation network is fine-tuned and trained, the key tissues and instruments of all frames are segmented using the fine-tuned and trained image segmentation network, and the confidence of the segmentation results is given. Manual annotation correction is introduced for the segmented images with low confidence, and the above steps are repeated to achieve the segmentation of the endoscopic image frames.

3. The endotracheal intubation navigation method based on image segmentation and three-dimensional reconstruction according to claim 2, characterized in that: In step 2, the monocular scene 3D reconstruction network inputs a single frame of endoscopic image , using the UNet network architecture to output a single image estimated depth map ; The inter-frame multi-view geometric pose estimation network inputs an image pair consisting of adjacent frames of multiple endoscopic images. , match the image pairs through SuperGlue feature points, and use feature point matching to estimate the basic matrix of the image pair by minimizing the epipolar geometry constraint , using nonlinear optimization and QR decomposition to obtain image pairs The relative pose rotation matrix between and ; Optical flow matching 3D reconstruction network inputs adjacent image frames, and uses FlowNet based on UNet network architecture to estimate the optical flow field between adjacent frames ; The depth map output by the monocular scene 3D reconstruction network is obtained through the following formula: , the image pair pose relationship output by the multi-view geometric pose estimation network between frames and the optical flow field output by the optical flow matching 3D reconstruction Optimize and obtain the three-dimensional point cloud model inside the tube; , in, , is the weight coefficient, is the camera projection function, which projects a 3D space point to a 2D image, and SSIM is the 2D image structure similarity measure.

4. The endotracheal intubation navigation method based on image segmentation and three-dimensional reconstruction according to claim 3, characterized in that: In step 3, a three-dimensional point cloud model of the endotracheal tube is obtained through offline scanning, and the three-dimensional point cloud model of the tube head is input into the virtual engine Unreal Engine. Virtual cameras with different postures are evenly distributed in the space around the tube head, and the instrument contour images in different postures are automatically generated through virtual camera rendering. These automatically generated virtual images and corresponding posture information are used as input to train a tube head instrument posture estimation network with a UNet network structure, and the tube head segmentation contour map obtained in step S1 is input into the trained instrument estimation network to output the instrument estimated posture sequence information.

5. The endotracheal intubation navigation method based on image segmentation and three-dimensional reconstruction according to claim 4, characterized in that: In step 4, the 3D point cloud model of the inside of the trachea generated in step S2 and the 3D point cloud model of the intubation head and the estimated instrument pose sequence information obtained in step S3 are input into the Unreal Engine virtual engine. Different colors are assigned to different tissues inside the trachea in the virtual engine, and the intubation head is moved and rotated according to the instrument pose sequence information of the intubation head. The scene is rendered to generate a digital twin dynamic navigation image of the intubation process.

6. The endotracheal intubation navigation method based on image segmentation and three-dimensional reconstruction according to claim 2, characterized in that: The image segmentation network is based on SAM2 architecture.

7. A tracheal intubation navigation device based on image segmentation and three-dimensional reconstruction, characterized in that: include: An endoscope image segmentation module collects an endoscope image frame sequence during the intubation process, and performs image segmentation on the endoscope image frame sequence to obtain an intubation head segmentation contour map; The 3D point cloud model construction module uses the endoscopic image frame sequence during the intubation process to perform 3D reconstruction of the interior of the trachea using three network branches: monocular scene 3D reconstruction, inter-frame multi-view geometric pose estimation, and optical flow matching 3D reconstruction, to obtain a 3D point cloud model of the interior of the trachea. A posture sequence information acquisition module estimates the posture of the endotracheal tube head through the segmentation contour of the instrument on the endotracheal tube head in the endoscopic image, and obtains a three-dimensional point cloud model of the tube head and the estimated posture sequence information of the instrument; The digital twin dynamic navigation image generation module inputs the tracheal internal 3D point cloud model generated in step S2, the intubation head 3D point cloud model obtained in step S3, and the instrument estimated posture sequence information into the Unreal Engine virtual engine to generate a digital twin dynamic navigation image of the intubation process.

8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor implements the method according to any one of claims 1 to 6.

10. An intelligent terminal, comprising a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, the memory is used to store a computer program, the computer program comprises program instructions, and is characterized in that: The processor is configured to call the program instructions to execute the method according to any one of claims 1 to 5.

Citation Information

Cited By

  • Endoscope-catheter cooperative control and anchoring optimization system

    CN121667601A

  • Endoscope-catheter collaborative control and anchoring optimization system

    CN121667601B