Medical visualization and assisted guidance intubation systems
The GIS system uses AI and CNNs for real-time anatomical landmark classification to provide quantitative guidance, improving intubation accuracy and safety by confirming tube placement and depth on external video displays.
Patent Information
- Application Number
- PCT/US2025/041446
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-26
- Filing Date
- 2025-08-11
- Publication Date
- 2026-03-05
AI Technical Summary
Current intubation technologies lack real-time, quantitative guidance for accurate endotracheal tube placement, relying heavily on clinician expertise and subjective judgment, leading to potential misplacement and complications such as airway trauma and delayed verification methods.
A guided intubation system (GIS) using artificial intelligence (AI) with convolutional neural networks (CNNs) for real-time anatomical landmark classification, providing color-coded graphical overlays on external video display devices to confirm tube placement and depth, integrating with existing laryngoscopy or bronchoscopy setups.
Enhances intubation accuracy by reducing misplacement risks and airway trauma through real-time feedback, ensuring precise tube placement and depth, even in challenging anatomical scenarios.
Smart Images

Figure US2025041446_05032026_PF_FP_ABST
Abstract
Description
MEDICAL VISUALIZATION AND ASSISTED GUIDANCE INTUBATION SYSTEMSPRIORITY CLAIM
[0001] This application claims the benefit of U.S. Provisional Patent Application Nos. 63 / 687,044 (filed on August 26, 2024), the contents of which is incorporated herein by reference in their entireties.FIELD OF THE DISCLOSURE
[0002] The present disclosure relates to software applications. More particularly, the present disclosure relates to a medical visualization software application and assisted guidance intubation systems.BACKGROUND
[0003] Endotracheal intubation is a significant medical procedure often performed in emergency and surgical settings to secure a patient’s airway. However, the process is fraught with challenges, particularly in cases involving difficult airway anatomy, such as anterior vocal cords, obesity, large tongue, or other irregular features.
[0004] Bronchoscopy or Laryngoscopy may be performed for many reasons, such as, for example, to view the interior structure of the throat such as the larynx, to facilitate endotracheal intubation using an endotracheal tube (ETT), to perform a biopsy procedure, etc. Generally, direct laryngoscopy refers to the use of a handheld laryngoscope to view the interior structure of the throat along a direct line of sight, while indirect laryngoscopy, video laryngoscopy, or bronchoscopy refers to the use of a laryngoscope or bronchoscope in combination with an optical device to visualize the larynx or trachea along an indirect line of sight, such as for example, a mirror or prism, a fiberoptic stylet, a video laryngoscope, etc.
[0005] However, for certain medical procedures, direct laryngoscopy often presents challenges that make a direct line of sight view difficult due to planned or unplanned scenarios such as anterior vocal cords, obesity, large tongue, or other irregular airway anatomical features. In addition, factors such as prehospital or hospital setting have been linked to wider ranges in successful endotracheal tube placement. Technological advances such as the indirect methods described have aided clinicians in successful intubations. The widespread use of such devices hasincreased and have been incorporated into guidelines for intubation of difficult airways, with many clinicians advocating their standard use in everyday practice. Despite these advances, however, current technology is lacking in providing realtime, quantitative guided feedback for successful intubation and appropriate depth placement. A current gold standard for verification of intubation in the trachea requires the timely placement of an additional monitor such as a colorimeter or end- tidal carbon dioxide monitor after intubation. However, desaturation can occur within one minute of the onset of tracheal intubation.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 depicts software application video image guidance with all image classifiers of tube placement, in accordance with embodiments of the present disclosure.
[0007] FIG. 2 depicts in-situ operation of video image guidance of an intubation tube into glottic structures during intubation, in accordance with embodiments of the present disclosure.
[0008] FIG. 3 depicts in-situ operation of video image guidance of an intubation tube into tracheal structures during intubation, in accordance with embodiments of the present disclosure.
[0009] FIG. 4 depicts in-situ operation of video image guidance of an intubation tube into the vicinity of the carina or bronchus structures during intubation, in accordance with embodiments of the present disclosure.
[0010] FIG. 5 depicts in-situ operation of video image guidance of an intubation tube into the vicinity of the esophagus during intubation, in accordance with embodiments of the present disclosure.
[0011] FIG. 6 depicts a schematic diagram of an example guided intubation system (GIS) with an image feature extraction network and an image classifier network for training a machine learning model for key anatomical landmark classification of the artificial intelligence algorithm, in accordance with embodiments of the present disclosure.
[0012] FIG. 7 is a flow chart diagram that illustrates pretraining a machine learning model for key anatomical landmark classification and subsequent use of the trained algorithm to construct an anatomical structure classifier, in accordance with embodiments of the present disclosure.
[0013] FIG. 8 is a flow chart diagram that illustrates a more specific example of pretraining a machine learning model for key anatomical landmark classification and subsequent use of the trained algorithm to construct an anatomical structure classifier, in accordance with embodiments of the present disclosure.
[0014] FIG. 9 is a schematic diagram of the Cormack-Lehane classification system.
[0015] FIG. 10 depicts a schematic diagram of glottic structures, in accordance with embodiments of the present disclosure.
[0016] FIG. 11 depicts a schematic diagram of glottic structures and esophageal entry, in accordance with embodiments of the present disclosure.
[0017] FIG. 12 depicts in-situ operation of an example GIS during real-time intubation, in accordance with embodiments of the present disclosure.
[0018] FIG. 13 illustrates a system block diagram including an example of a computing device that may be used in implementing one or more features of the disclosure, in accordance with embodiments of the present disclosure.
[0019] FIG. 14 is a flow chart diagram that illustrates extraction of anatomical features of an image, subsequent processing by a classification network, and communication of feedback guidance to assist intubation in real-time, in accordance with embodiments of the present disclosure.
[0020] FIG. 15 is a flow chart diagram that illustrates guidance information provided by a video display of a video display device to assist intubation in real-time, in accordance with embodiments of the present disclosure.DETAILED DESCRIPTION
[0021] Embodiments of the present disclosure will now be described with reference to the drawing figures, in which like reference numerals refer to like parts throughout.
[0022] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0023] Conventional intubation techniques, including direct laryngoscopy and video laryngoscopy, have somewhat improved visualization but still rely heavily on the clinician’s expertise and subjective judgment. This reliance can lead to procedural errors, such as misplacement of the endotracheal tube into the esophagus or incorrect depth placement, which may result in airway trauma, hypoxia, esophageal intubation, or other complications. Furthermore, current verification methods, such as end-tidal CO2monitoring, are often delayed, leaving clinicians without real-time feedback during the pivotal moments of tube placement. Despite advancements in imaging devices and indirect visualization methods, existing systems lack the capability to provide quantitative, real-time guidance and predictive feedback to ensure accurate tube placement and depth.
[0024] Quantitative guided feedback from artificial intelligence algorithms during real-time intubation has the potential to help clinicians prevent erroneous tube manipulation and airway trauma. The present disclosure addresses the foregoing limitations by introducing a guided intubation system (GIS), device, method and non- transitory computer readable media (NTCRM) that leverages artificial intelligence (Al) for real-time anatomical landmark classification and feedback during intubation. The system incorporates a trained (pretrained) convolutional neural network (CNN) and deep learning algorithms to analyze live video streams from an external video display device, such as a smartphone or tablet. By identifying anatomical structures — including glottic openings, tracheal rings, carina, bronchial bifurcations, and esophageal inlets — the system, device, method and NTCRM embodiments described herein provide clinicians with visual feedback through color-coded graphical overlays. These overlays delineate the passage of the endotracheal tube during intubation and confirm the placement and depth of the tube in real-time, providing a feedback loop to the clinician performing the intubation and significantly reducing the risk of misplacement of the location and placement depth of intubation tube and other procedural errors. Specialized algorithms for feature extraction, image segmentation, and predictive modeling, enable effective intubation by aclinician of the system and device, even in scenarios involving obscured or abnormal anatomical structures.
[0025] Unlike conventional approaches, the GIS operates locally on the external video display device without reliance on cloud-based computing, ensuring rapid processing and deployment in time-sensitive medical environments. The system’s modular architecture integrates seamlessly with existing video laryngoscopy or bronchoscopy setups, enhancing their functionality with Al-driven guidance. Additionally, the intuitive user interface includes tools for video rotation, image capture, and algorithm toggling, allowing clinicians to adapt the system to their specific needs during the procedure. By combining advanced machine learning techniques with real-time video processing, the GIS represents a significant improvement over prior methods, enhancing intubation accuracy, reducing airway trauma, and improving patient outcomes.
[0026] The following detailed description provides illustrative embodiments of the disclosed subject matter, which pertains to systems and methods for guided intubation utilizing artificial intelligence for real-time anatomical landmark classification and feedback. The described subject matter is presented in the context of medical visualization and guidance systems, particularly those aimed at enhancing the accuracy, safety, and efficiency of endotracheal intubation procedures. While specific embodiments, configurations, and examples are provided, these are intended to facilitate understanding and are not to be construed as limiting the scope of the disclosed subject matter.
[0027] Those skilled in the art will understand that certain widely recognized elements, processes, or techniques may not be described in detail to prevent unnecessary complexity in the disclosure. Additionally, various modifications, substitutions, or rearrangements of components, methods, or steps may be implemented without deviating from the principles and scope of the subject matter as defined by the claims. The disclosed subject matter is intended to include all such variations and adaptations that would be apparent to a person of ordinary skill in the art.
[0028] In one configuration, a system is provided for guiding an endotracheal tube in the appropriate tracheal structure. The system includes an application softwarethat connects wirelessly or wired to an external video display device to provide video guidance of endotracheal tube, such as shown in FIG. 12. The application software includes instructions to extract and detect images compatible to its pretrained neural network from live feed video during endotracheal tube intubation.
[0029] The system also includes imbedded instructions storable on the memory drive of a connected external video display device which gives an output of markings or grids delineating glottic, laryngeal, and esophageal structures during real-time intubation. These stored instructions are executed by the processor of the video display device (i.e. smartphone, tablet, computer, touchscreen, a television, or other video display screen).
[0030] Algorithmic instructions via neural network and deep machine learning interpret images for guided, real-time decision making of successful intubation, appropriate placement and positioning of endotracheal tube, and lastly, identification of placement of tube into the esophagus via visible output indicators on an external video display device, as referenced in FIG. 8.
[0031] The algorithmic instructions are based on deep learning feature extraction network with convoluted neural network layers, pooling layer, transposed convolution layer, classification layer, or other various layers to encode the final output images (markings, grids, visual tracheal or esophagus tube placement indicators) during real-time tube placement, as referenced in FIG. 8.
[0032] The final image output serves as real-time assistance during intubation of whether tube is placed in the trachea or esophagus. A classifier output image has a color-coded column, displayed and represents with its own corresponding indicator a sequential passing of the tube into the vicinity of 1 ) the vocal cords opening, 2) trachea, and 3) just before the carina or bronchus. All three successfully classified images with its corresponding output image suggests proper tube placement in the trachea, as indicated in FIG. 6. On a separate color-coded column is displayed and represents with its own corresponding indicator passing of the tube into the vicinity of the esophagus, as shown in FIG. 5.
[0033] The foregoing and other aspects and advantages of the present disclosure will become apparent appear from the detailed description. In the description, reference is made to the accompanying drawings that form a part hereof, and inwhich there is shown by way of illustration a preferred embodiment. This embodiment does not necessarily represent the full scope of the invention, however, and reference is therefore made to the claims and herein for interpreting the scope of the invention. The detailed description and examples are given by illustration only as various changes and alterations within the scope of the present disclosure will become apparent to those skilled in the art from the following description.
[0034] Embodiments of the present disclosure will now be described with reference to the drawing figures, in which like reference numerals refer to like parts throughout. Embodiments of the present disclosure advantageously provide a system that broadens the use of standard laryngoscopy, video laryngoscopy or bronchoscopy beyond current available technology, as referenced in FIG. 11 .
[0035] The present disclosure recognizes the unsolved need for a method and software which incorporates electronic, real-time guidance and predictive value of correct or incorrect placement of endotracheal tubes using mathematical pattern recognition / predictive modeling as well as convoluted neural network algorithms for assistive decision-making of endotracheal tube placement, as shown in FIG. 8.
[0036] One of more embodiments of the disclosure include a method comprising use of images captured from an imaging device to create an image recognition and predictive model to identify glottic, laryngeal, and esophageal structures in real-time during active placement of endotracheal tubes, as shown in FIGs. 1-5. The system can comprise a software that integrates artificial intelligence with pattern recognition, video frame manipulation, segmentation, deep learning, or transfer learning and algorithms such as Fast Fourier Transform (FFT) and one or more neural networks, such as convoluted neural networks (CNNs) or generative adversarial networks (GANs) that are types of deep learning algorithms, in order to supply a predictive value of correct or incorrect placement of endotracheal tubes using the image data and working memory of a network connected external video display device via wired or wireless connection; reference FIG. 6. In particular, the usage of CNN or GAN networks in imbedded instructions of software that can be downloaded to a nonproprietary external display device, such as the external video display device described herein, and assist in providing visual output guidance and assistance to a clinician of the external video display device in the proper placement of an intubation tube during an intubation procedure itself is considered to be novel and quite helpful.
[0037] In some embodiments of the present disclosure, as shown in FIG. 6, the graphical representation of the attribute of the glottis, trachea, carina / bronchus, or esophagus, i.e. , the calculated value of the attribute (for example, color coded) of anatomical landmarks, can be displayed on the video display.
[0038] FIG. 1 depicts software application video image guidance with all image classifiers 136, 138, 140 of tube placement on a display screen of a video display device 100, in accordance with embodiments of the present disclosure. An image 110 is displayed together with a graphical user interface (GUI) 120 having a number of icons 122, 124, 126, 128, 132, 134 for executing commands and accessing portions of the computer application intubation hardware that runs on display device, and guidance indicators 136, 138, 140 that provide information to a clinician of the intubation. The purple color of indicators 136, 138, 140 indicates placement of the intubation tube during intubation. On and off icon 128, shown in red and therefore not on, is made available to turn on and off the Al algorithm image identification, extraction, and classification system, thereby enabling a bounding box to be overlaid the image when icon 128 is turned on.
[0039] FIG. 2 depicts in-situ operation of video image guidance of an intubation tube into glottic structures during intubation, in accordance with embodiments of the present disclosure. On a display screen of a video display device 200 an image 210 with a bounding box 215 (in purple) is displayed together with a graphical user interface (GUI) 220 having a number of icons 222, 224, 226, 228, 232, 234 for executing commands and accessing portions of the computer application intubation hardware that runs on display device, and guidance indicator 236 that provides information to a clinician of the intubation about the position of the intubation tube.
[0040] The purple color of indicator 236 indicates placement of the intubation tube during intubation. On and off icon 228 is made available to turn on and off the Al algorithm image identification, extraction, and classification system, thereby enabling a bounding box to be overlaid the image when turned on. In this case, icon 228 is shown in green and therefore on, allowing bounding box 215 to be displayed over the image 210.
[0041] As shown in FIGs. 2-4, the color-coded identification of the glottic structures is displayed on the right-bottom corner of the image. Note, the first graphical displayin a column corresponds with identification of the glottic structures, as shown in FIG.2. By displaying the calculated value of the attribute of the glottic structures on the display, the clinician may read the calculated value while intubating without interrupting the procedure.
[0042] The present disclosure addresses the aforementioned drawbacks by providing new systems and methods for guided airway intubation. The systems and methods provide for image recognition and analysis for segmentation of airway passages of interest from image data. The image analysis provides guidance for insertion of endotracheal tubes into a patient and may be accomplished by clinicians assisted by the guidance described.
[0043] Pretraining
[0044] FIG. 7 is a flow chart diagram 700 that illustrates pretraining a machine learning model for key anatomical landmark classification and subsequent use of the trained algorithm to construct an anatomical structure classifier, in accordance with embodiments of the present disclosure. Referring to flowchart diagram 700 of FIG.7, at block 710 annotated images of key anatomical structures / landmarks are provided to serve as pretraining image dataset. At block 720, a feature extraction network based on key anatomical structure recognition and extract feature data of the anatomical images through a trained feature extraction network is constructed. At block 730, a bounding box is provided during real-time intubation during extraction of feature data. At block 740, an anatomical structure classifier based on machine learning algorithm(s), CNN is constructed and visual classification on the extracted feature data of the anatomical images are performed by a classifier network to acquire depth and placement of the intubation tube.
[0045] Pretrained images include fiberoptic bronchoscope and endoscopic images (e.g., videos or images) of anatomical structures of the glottis, trachea, carina, bronchus, and esophagus in various scenarios of normal structures and abnormal structures (i.e. blood, vomitus, secretions, blurriness, diseased structures) (FIG. 8). The images may or may not be obtained from various media to include but not limited to the internet, frames from captured videos, open-sourced images, open web sources, medical or non-medical ImageNet, Lab View, application programming interface (API) with TensorFlow or image data platform, CNN, GAN, or other platformfor training large scale object recognition algorithms, machine learning, and pretraining purposes. Images were then labelled by a medical expert using a free version of roboflow. Corresponding XML annotations were curated manually.Images were resized to 416x416 pixels for YOLOv8 compatibility. YOLOv8 internally determined loss function and not explicitly modified in the code.
[0046] Referring to flowchart diagram 800 of FIG. 8, at block 810 endoscopic annotated images of key anatomical landmarks are provided. At block 820, a feature extraction network, such as that shown in FIG. 6, is constructed based on key anatomical landmark recognition, and feature data of the anatomical images are extracted through a trained feature extraction network (such as 630 of FIG. 6). At block 830, a YOLO bounding box is provided during real-time intubation during traction of feature data. At block 840, an anatomical landmark classifier based on a trained machine learning algorithm, a CNN and a performing classification using a visual classifier (such as 640 of FIG. 6) on the extracted feature data to acquire an evaluation result that includes an anatomical landmark and tube placement information.
[0047] In accordance with this description of FIG. 8, consider the following example process for training a machine learning model for anatomical landmark classification in accordance with an embodiment of the present disclosure. The process begins with the acquisition of a diverse dataset comprising images and videos of anatomical structures, including glottic openings, tracheal rings, carina, bronchial bifurcations, and esophageal inlets. These images are preprocessed through resizing, normalization, and data augmentation techniques such as flipping, rotation, and jittering to enhance the robustness of the model. The annotations for the dataset, stored in Pascal VOC XML format, are converted into YOLO-compatible format to facilitate training. A pretrained YOLOv8s model, optimized for real-time object detection, is fine-tuned using the augmented dataset, which includes thousands of expert-labeled medical images. The training pipeline incorporates feature extraction, bounding box generation, and classification layers to enable precise identification of anatomical landmarks. The final model is evaluated using metrics such as precision, recall, F1 score, and mean average precision (mAP) at multiple Intersection over Union (loU) thresholds, ensuring high accuracy and reliability. This trained model is subsequently optimized for mobile deployment by quantizing neural network weightsand exporting the model as a TorchScript library for integration into the guided intubation system.
[0048] In one embodiment, an image labelling system to include but not limited to Github open sources, or other system may be incorporated to label images according to parameters described here. The images are labelled as glottis, trachea, carina, bronchus, and esophagus. Each image is transformed via feature extraction (FIG. 8). The real-time images are compared to the pretrained algorithm, labelled with box bounding, and then inserted to a classification system (FIG. 6).
[0049] Referring now to FIG. 6, a block diagram 600 of an example guided intubation system (GIS) with an image feature extraction network and an image classifier network for training a machine learning model for key anatomical landmark classification of the artificial intelligence algorithm is shown. More specifically, guided intubation system (GIS) block diagram 600 of FIG. 6 includes a video capture element 610 that captures images 612 for input to a pre-training CNN 620 as shown. The pre-trained artificial intelligence model is used by image feature extraction network 630 and image classifier network 640, which provide processed real-time feedback output usable by a video display device 650 as shown. Image feature extraction network 630 has a CNN 632 and a feature extractor 634; image classifier network 640 has CNN 642 and classifier 644. The video display device 650 has a graphical user interface (GUI) that reflects the progress of the intubation vis-a-vis various structures, include esophageal 652, glottic 654, tracheal 656, and carina 658. Guidance indicators 660, shown here in yellow, 662, 664, 666, shown here in purple, are each associated with a particular endotracheal structure as shown.
[0050] FIG. 6 depicts a schematic diagram of an example guided intubation system (GIS) with an image feature extraction network and an image classifier network for training a machine learning model for key anatomical landmark classification of the artificial intelligence algorithm
[0051] Figure 6 illustrates a schematic diagram of the guided intubation system’s feature extraction network, which plays a significant role in training the machine learning model for anatomical landmark classification. The network employs a convolutional neural network (CNN) architecture (620, 632, 642), incorporating layers such as convolutional, pooling, and transposed convolution layers, to processand analyze input images. The training process begins with the acquisition of a diverse dataset of labeled medical images 612, including glottic openings, tracheal rings, carina, bronchial bifurcations, and esophageal inlets. These images undergo preprocessing steps such as resizing, normalization, and data augmentation to enhance the robustness of the model. The feature extraction network 630 identifies anatomical structures by segmenting and classifying the input images, generating bounding boxes and graphical overlays that delineate the anatomical regions of interest. The output of this network is used to train the system’s predictive model, supporting accurate real-time identification and classification of anatomical landmarks by image classifier network 640 during endotracheal intubation.
[0052] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes the first important anatomical structures during endotracheal intubation: glottic structures. The Cormack-Lehane Grading system 900 shown in FIG. 9, or modified form, is applied first to incorporate the 4 various views of glottic structures to include Grade 1 full view of glottis, Grade 2 partial view of the glottis or arytenoids, Grade 3 view of only the glottis, and Grade 4 neither glottis nor epiglottis visible (FIG. 9). The Cormack-Lehane classification system, as illustrated in FIG. 9, provides a standardized framework for assessing glottic visibility during laryngoscopy, which plays an important role in guiding endotracheal intubation. This system categorizes glottic views into four grades: Grade 1 represents a full view of the glottis, Grade 2 indicates a partial view of the glottis or arytenoids, Grade 3 shows only the epiglottis without glottic visibility, and Grade 4 reflects no visible glottis or epiglottis.
[0053] FIGs. 10 and 11 further expands on this classification by incorporating scenarios involving obscured anatomical structures, such as the presence of blood, vomitus, or secretions, which may impede visualization. These scenarios are categorized based on their location relative to the epiglottis and arytenoids, enabling the system to adapt to real-world complexities. By integrating these classifications into the guided intubation system, the software can accurately identify and label glottic structures under varying conditions, improving the reliability of real-time feedback during intubation procedures. FIG. 10 depicts a schematic diagram of glottic structures 1000, in accordance with embodiments of the present disclosure. Thus, in addition to glottic views, various scenarios are incorporated as a subset toeach corresponding view to include blood, closed vocal cords, secretions, vomit either above the epiglottis (A), in between the epiglottis and arytenoids (B), or below the arytenoids (C), as shown in FIG. 10. Lastly, common disease process is incorporated to pretraining the glottic structure view including laryngitis, and nodules or polyps on the vocal cords. FIG. 11 is a schematic diagram 1100 that highlights regions 1 and 2 of glottic structures and esophageal entry, in accordance with embodiments of the present disclosure.
[0054] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes the second important anatomical structures during endotracheal intubation: tracheal structures. Here, anterior tracheal rings, trachealis muscle wall posteriorly, and the carina distally are the tracheal structures for image recognition, as shown in FIG. 3. Pretraining the software involves the ability for the software to identify and count the number of tracheal rings from the carina.
[0055] Referring now to FIG. 3, on a display screen of a video display device 300 an image 310 with a bounding box 315 (in purple) is displayed together with a graphical user interface (GUI) 320 having a number of icons 322, 324, 326, 328, 332, 334 for executing commands and accessing portions of the computer application intubation hardware that runs on display device, and guidance indicators 336, 338, 340 that provide information to a clinician of the intubation tube on the placement and depth of the intubation tube during the intubation procedure. The purple color of indicator 338 indicates that the intubation tube is passing through the trachea. On and off icon 328 is made available to turn on and off the Al algorithm image identification, extraction, and classification system, thereby enabling a bounding box to be overlaid the image when turned on. In this case, icon 328 is shown in green and therefore on, allowing bounding box 315 to be displayed over the image 310. The other icons will be described further below.
[0056] In FIG. 3, the in-situ operation of the guided intubation system (GIS) during the passage of the endotracheal tube through the tracheal structures is shown. The figure highlights the system’s ability to identify and delineate anatomical features of significance, including the anterior tracheal rings and the posterior trachealis muscle wall, using real-time video analysis. The artificial intelligence algorithm processes the live video feed and overlays color-coded bounding boxes on the tracheal rings,providing visual confirmation of the tube’s progression through the trachea. This graphical feedback is displayed on the external video device, enabling the clinician to monitor the tube’s position and ensure proper alignment within the trachea. The system also counts the number of tracheal rings visible above the carina, assisting in the determination of appropriate tube depth placement. By providing this real-time guidance, the GIS reduces the risk of misplacement and enhances the safety and accuracy of the intubation procedure.
[0057] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes the third important anatomical structures during endotracheal intubation: carina and bronchial structures, FIG. 4. Here, one left and right opening to signify the right and left major bronchus separated by the carina in-between are the two structures for image recognition.
[0058] In one embodiment, two image variations of carina and bronchial structures are learned in order to assist in determining appropriate endotracheal tube depth (FIG. 4). The first image variation focuses on identifying three to four tracheal rings above the carina to signify appropriate endotracheal depth placement (FIG. 4). The second image variation focuses on identifying deep endotracheal tube placement whereby either is seen: 1 ) only one major bronchus (after carina already identified), or 2) both left and right bronchus with a smaller opening adjacent to the right main bronchus suggesting the right upper bronchus outlet (after carina already identified).
[0059] In one embodiment, pretraining the software involves the ability to recognize appropriate endotracheal depth placement, which is approximately 4-6 centimeters above the carina, or three to four tracheal rings above the carina on image analysis (FIG. 4).
[0060] Referring now to FIG. 4, on a display screen of a video display device 400 an image 410 with a bounding box 415 (in purple) is displayed together with a graphical user interface 420 having a number of icons 422, 424, 426, 428, 432, 434 for executing commands and accessing portions of the computer application intubation hardware that runs on display device, and guidance indicators 436, 438, 440 that provide information to a clinician of the intubation tube on the placement and depth of the intubation tube during the intubation procedure. The purple color of indicator 440 indicates that the intubation tube is passing through the carina andbronchial structure. On and off icon 428 is made available to turn on and off the Al algorithm image identification, extraction, and classification system, thereby enabling a bounding box to be overlaid the image when turned on. In this case, icon 428 is shown in green and therefore on, allowing bounding box 415 to be displayed over the image 410. The other icons will be described further below.
[0061] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes important anatomical structures during esophageal intubation: esophageal opening and esophagus (FIG. 5). Two particular views are learned: A) glottic structures observed with a posterior opening suggesting the esophageal inlet (FIG. 11 ), and B) dark center surrounded by smooth muscle throughout (FIG. 5).
[0062] Referring now to FIG. 5, on a display screen of a video display device 500 an image 510 with a bounding box 515 (in yellow) is displayed together with a graphical user interface 520 having a number of icons 522, 524, 526, 528, 530, 532, 534 for executing commands and accessing portions of the computer application intubation hardware that runs on display device. On and off icon 528 is made available to turn on and off the Al algorithm image identification, extraction, and classification system, thereby enabling a bounding box to be overlaid the image when turned on. In this case, icon 528 is shown in green and therefore on, allowing bounding box 515 to be displayed over the image 510. The other icons will be described further below.
[0063] FIG. 5 illustrates the in-situ operation of the guided intubation system (GIS) during the identification of esophageal structures, an important step in preventing unintentional esophageal intubation. The figure depicts the esophageal opening and the esophagus itself, characterized by a dark central region surrounded by smooth muscle tissue. The artificial intelligence algorithm processes the live video feed and overlays color-coded graphical indicators (in yellow) to delineate these anatomical features, providing real-time feedback to the clinician. This visual confirmation assists in distinguishing the esophagus from the trachea, ensuring that the endotracheal tube is not inadvertently placed in the esophagus. The system’s ability to identify esophageal structures under various conditions, including obscured views caused by secretions or blood, enhances the reliability of the intubation process and reduces the risk of complications associated with improper tube placement.
[0064] Furthermore, the image recognition algorithms via CNN, GAN, Viola Jones object detection framework or other system, may adopt a simple logical reasoning:If A is both esophagus and glottic structuresIf B is circular all around and dark centerIf C =Esophagus*Then A +B =C
[0065] Furthermore, in scenarios where structures are obscured, the image recognition algorithms may adopt a simple logical reasoning to include but not limited to:If 1 is glottic structures onlyIf 2 is tracheal structuresIf 3 is carina and bronchus*lf B - 2, then C; or*lf B-3, then C; or*lf A -2, then C; or*lf A-3, then C
[0066] The training inference pipeline utilized the following tools: PyTorch (via Ultralytics YOLOv8 library), Open CV (for image I / O and visualization), and Pandas (for logging and metadata processing). The base YOLOv8s model was trained on the Common Objects in Context (COCO) dataset before fine tuning. Fine tuning involved further training the model on thousands of medical expert labelled images of internal anatomy that would be encountered during endotracheal intubation. For training, images are in JPEG or PNG format. Annotations are stored in Pascal VOC XML files. In the deployed model, the input is a live stream video.
[0067] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes training to distinguish false positives from true positives based on manual verifications. This may involve manuallyannotated training sets with labelling images that are not the intended anatomical structure for the purposes of pretraining algorithm to identify false positives.
[0068] In one embodiment, pretraining the software for image recognition and algorithmic output decision-making includes training to distinguish where not to label based on a rule set of anatomical positioning or sequence of anatomical positioning identified. This may involve encoding a rule set of commands restricting the labelling of a structure because the sequence does not follow an anatomical possibility for the purposes of pretraining algorithm to identify false positives. For example, the rule set may restrict labelling a structure esophagus once it has already previously a combination labelled structures vocal cords, trachea, and tracheal rings.
[0069] Identification and Extraction
[0070] In some embodiments of the present disclosure, a detection and image processing algorithm may function by applying known imaging processing techniques to separate those parts of the images of the video frames arising during intubation. By this means it becomes possible to identify, extract, and compare images to the pretrained algorithm described prior for computer vision and machine learning. Such image processing techniques can include Fast Fourier Transform (FFT) algorithm, feature extraction network, Convoluted Neural Network (CNN), Generative Adversarial Network (GAN), Siamese Neural Networks, Deep Neural Networks (e.g. Inception modeling), transfer learning, Simultaneous Localization and Mapping (SLAM), Spatial Transformer Module (STM), Image Segmentation (e.g. U- Net, ResNet, region growing algorithms, edge-based segmentation algorithms), Data Augmentation (e.g. flipping, jittering, rotation), pattern recognition, heat mapping, and video frame manipulation (e.g. adding and subtracting image frames). One such common processing technique uses a Fast Fourier Transform (FFT) algorithm to extract components of the original images detected, and creates a separate image which can be then used as extracted features overlaid for the imaged frames detected by the external video display device. The FFT algorithm provides signal processing to the performed in real-time of each frame of the video stream.
[0071] In some embodiments of the present disclosure, a YOLO (You-only-look- once) bounding box detection system is employed to demonstrate identification of key anatomical landmarks according to the pretrained dataset, such as shown inFIGs. 2-5. For example, a YOLOv8s, the small variant, was selected, which uses a CSPDarknet inspired architecture with ELAN and C2f modules. It is anchor-free and optimized for real-time inference. Inference is performed using a YOLOv8’s pipeline. Each frame is passed through the model to produce bounding box predictions. Postprocessing involves rendering boxes into the livestream video feed. A bounding box is visible during real-time intubation for clinician feedback of key structures such as glottic structures, tracheal structures, bronchial structures, and esophageal structures. In summary, the Al processing algorithm analyzes multiple captured image frames, labels and displays a bounding box over the recognized high priority anatomical landmarks, analyzes an expected location within said landmarks, and provides real-time confirmation of tube placement and depth on display screen (FIG. 8).
[0072] In some embodiments the artificial intelligence algorithm uses objection detection to identify and label anatomical structures. It may or may not perform segmentation. The methodology uses detection using bounding boxes, and the output uses image coordinates and class labels per frame. Training involved converting bounding boxes from XML to YOLO format.
[0073] Classification
[0074] Once real-time images are extracted and compared to a pretrained algorithm for image identification, the identified images are processed through a classification system. Once again, a convolutional neural network 620, 632, 642 is used specifically as a pretrained image classifier for deep learning, machine learning of key anatomical landmarks, FIG. 6. Referring to FIGs. 2-5, the image classifier 644 of image classifier network 640 may be implemented as follows.
[0075] In some embodiments of the present disclosure, such as shown in FIGs. 2- 5, the graphical representation of the attribute of the glottis, trachea, carina / bronchus, or esophagus, i.e. , the calculated value of the attribute (for example, color coded) of anatomical landmarks, can be displayed on the video display. As shown in FIG. 3, the color-coded identification of the glottic structures displayed on the right-bottom corner of the image 310 by purple indicator 338. Note, the second graphical display 338 in a column corresponds with identification of the tracheal structures in FIG. 3. By displaying the calculated value of the attribute ofthe tracheal structures on the display, the clinician may read the calculated value while intubating without interrupting the procedure.
[0076] In some embodiments of the present disclosure, as shown in FIGs. 2-5, the graphical representation of the attribute of the glottis, trachea, carina / bronchus, or esophagus, i.e. , the calculated value of the attribute (for example, color coded) of anatomical landmarks, can be displayed on the video display. As shown in FIG. 4, for example, the color-coded identification of the carinal / bronchus structures is displayed on the right-bottom corner of the image (purple indicator 440). Note, the third graphical display in a column corresponds with identification of the carinal / bronchial structures. Note, the third display 440 in a column will remain visible when appropriate endotracheal tube depth placement is confirmed. By displaying the calculated value of the attribute of the carinal / bronchial structures on the display, the clinician may read the calculated value while intubating without interrupting the procedure. As surmised, each graphical display corresponding to the three structures identified are sequential to reflect the natural passing of the tube into these structures as they are met in real-time (FIG. 6).
[0077] In some embodiments of the present disclosure, as shown in FIG. 2-5, the graphical representation of the attribute of the glottis, trachea, carina / bronchus, or esophagus, i.e., the calculated value of the attribute (for example, color coded) of anatomical landmarks, can be displayed on the video display. As shown in FIG. 5, for example, the color-coded identification 530 of the esophageal structures is displayed on the left-bottom corner of the graphical user interface (GUI) displayed with image 510; in this case it is yellow as indicated by the yellow box . By displaying the calculated value of the attribute of the esophageal structures on the display, the clinician may read the calculated value while intubating and make corrections without interrupting the procedure.
[0078] The use of color-coded Indicators 130, 230, 330, 430, 530, and associated bounding boxes 215, 315, 415, 515 depicted as yellow or purple in FIGs. 2-5, respectively, provide real-time visual feedback by overlaying graphical elements on the live video stream. These indicators delineate anatomical landmarks such as glottic openings, tracheal rings, and esophageal inlets, by bounding boxes, thereby assisting clinicians in confirming tube placement and depth. The indicators are generated by the Al algorithm based on the analysis of the live video stream and aredisplayed sequentially to represent the passage of the endotracheal tube through various anatomical structures. This feature enhances the accuracy and safety of the intubation procedure by reducing the risk of misplacement.
[0079] Evaluation of the algorithm model was performed using a separate holdout dataset, distinct from training and validation sets. YOLOv8 returns the following object detection metrics:Precision, Recall, F1 , Mean average precision scores (mAP) with Intersection over Union (loU) thresholds, mAP@0.5 (mean average precision at loU threshold 0.5), mAP@0.5:0.95 (averaged over multiple loU thresholds).Class-specific precision, recall, F1 score, accuracy, and mAP values were also logged per anatomical structure.
[0080] Software elements
[0081] In the embodiments of the present disclosure, several icons serve as tools to execute commands and access portions of the software and are shown in various of the drawings including in FIGs. 1-5. For example, a rotate icon 124, 224, 324, 324, 524 of the drawings is made available to perform 90-degree rotations of the video image in order to assist the clinician in real-time to preferred angle. A settings icon 126, 226, 326, 426, 526 is made available to access key items such as saved images or videos gallery. A picture icon 132, 232, 332, 432, 532 and video icon 134, 234, 334, 434, 534 are made available to perform a snapshot image or record a video, respectively. An ‘eye icon’ 122, 222, 322, 422, 522 is made available to orient and serve as a compass for the clinician to quickly identify the anterior view in realtime. An on and off icon 128, 228, 328, 428, 528 is made available to turn on and off the Al algorithm image identification, extraction, and classification system.
[0082] In the embodiments of the present disclosure, the video display may comprise any machine configured to perform processing and / or calculations, may be but is not limited to a work station, a server, a desktop computer, a tablet computer, computing devices, a server farm, remote or wired machine, a personal data assistant, a smart phone, or any combination thereof; see FIG. 12 and also FIG. 13. As the software utilizes the processor and memory of an external video display device to function, the algorithm can be executed online or offline. The algorithms used herein has the built-in ability to configure video processing of the graphing reprocessing unit (GPU) usage along with the computer processing unit (CPU) usage to utilize machine resources to execute the tasks described. It is understood that the described platform can be implemented using any computing technique, e.g., as a stand-alone system, a distributed system, within a network environment, etc. All processing of the algorithm model is preferably processed on the external screen device. In these embodiments, the software application does not rely on cloud services for image detection and deployment of algorithm.
[0083] A server may be any server type such as, for example: a file server; an application server; web server; proxy server; an appliance; a network appliance; a gateway; a gateway server; a virtualization server; a deployment server; a Secure Sockets Layer Virtual Private Network (SSL VPN) server; a firewall; a web server; a server executing an active directory; a cloud server; or a server executing an application acceleration program that provides firewall functionality, application functionality, or balancing functionality. Reference also FIG. 13.
[0084] The memory may be any storage devices that are non-transitory and can implement data stores, and may compromise but are not limited to an optical storage device, a solid-state storage, hard disk drive, or any other magnetic medium, a ROM (Read Only Memory), a RAM (Random Access Memory), a cache memory and / or any other memory chip or cartridge, and / or any other medium from which a computer may read data, instructions, and / or code. Reference also FIG. 13.
[0085] The communications device may be any kind of device or system (e.g., hardware, software, firmware) that can enable communication with external apparatuses and / or with a network, and may comprise but are not limited to a modem, a network card, an infrared communication device, a wireless communication device, WiFi device, WiMax device, Near Field Communication (NFC), Wide Area Network (WAN), a Metropolitan Area Network (MAN), Wireless Local-Area Network (WLAN), 802.11 , Bluetooth, cellular communication facilities and / or the like. In a more particular example, hardware, software, firmware can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, etc.), WiFi connection, a Bluetooth connection, a cellular connection, and so on.
[0086] The software architecture is modular in dataset preparation, preprocessing and label formatting (via helper functions), and YOLOv8 training and evaluation (model conversion and export). The detection model is integrated with the software application or video system directly within the Flutter application using the pytorch ite plugin. A pretrained TorchScript library is included in the application to identify airway anatomy-related objects.
[0087] The software application processes live video during endotracheal intubation in a specific manner. When live camera is activated on the external video screen device, the software application streams frames to the integrated TorchScript model in real-time. The model analyzes each frame, returns bounding rectangle data for detected objects, and the software application overlays these rectangles on the video feed.
[0088] Therefore, as described herein, a guided intubation system (GIS), device, method and computer readable media for real-time anatomical landmark classification and feedback during endotracheal intubation is disclosed. Artificial intelligence-based software is configured to analyze live video streams from an external video display device, utilizing a pretrained convolutional neural network (CNN) to identify anatomical structures such as glottic openings, tracheal rings, carina, bronchial bifurcations, and esophageal inlets. The software provides real-time visual feedback through color-coded graphical overlays, assisting clinicians in confirming tube placement and depth. The system operates via wired or wireless connection to external devices, such as smartphones or tablets, and features an intuitive user interface, such as a GUI, with tools for video rotation, image capture, and algorithm toggling. The GIS enhances intubation accuracy, reduces procedural errors, and minimizes airway trauma by leveraging deep learning algorithms for predictive modeling and decision-making. The system is optimized for local deployment without reliance on cloud-based computing.
[0089] Referring now to FIG. 12, an in-situ operation of an example GIS 1200 during a real-time intubation is shown, in accordance with embodiments of the present disclosure. Intubation tube 1210, external video display device 1220 and the display screen 1230 with graphical user interface 1240 are shown.
[0090] FIG. 13 illustrates a system block diagram 1300 including an example of a computing device 1310, such as the external video display device discussed above, that may be used in implementing one or more features of the disclosure, in accordance with embodiments of the present disclosure. Computing device 1310 may, in some embodiments, implement one or more aspects of the disclosure by reading and / or executing instructions and performing one or more actions based on the instructions. In some embodiments, computing device 1310 may represent, be incorporated in, and / or include various devices such as a desktop computer, a computer server, a mobile device (e.g., a laptop computer, a tablet computer, a smart phone, any other types of mobile computing devices, and the like), and / or any other type of data processing device. Further, as discussed previously, the video display device may comprise any machine configured to perform processing and / or calculations, may be but is not limited to a work station, a server, a desktop computer, a tablet computer, computing devices, a server farm, remote or wired machine, a personal data assistant, a smart phone, or any combination thereof; see FIG. 12 and FIG. 13. Moreover, as previously described, a server may be any server type such as, for example: a file server; an application server; web server; proxy server; an appliance; a network appliance; a gateway; a gateway server; a virtualization server; a deployment server; a Secure Sockets Layer Virtual Private Network (SSL VPN) server; a firewall; a web server; a server executing an active directory; a cloud server; or a server executing an application acceleration program that provides firewall functionality, application functionality, or balancing functionality.
[0091] As the software utilizes the processor and memory of an external video display device to function, the algorithm can be executed online or offline. The algorithms used herein has the built-in ability to configure video processing of the graphing processing unit (GPU) usage along with the computer processing unit (CPU) usage to utilize machine resources to execute the tasks described. It is understood that the described platform can be implemented using any computing technique, e.g., as a stand-alone system, a distributed system, within a network environment, etc. All processing of the algorithm model is preferably processed on the external screen device. In these embodiments, the software application does not rely on cloud services for image detection and deployment of algorithm.
[0092] Computing device 1310 may, in some embodiments, operate in a standalone environment. In others, computing device 1310 may operate in a networked environment. As shown in FIG. 13, computing devices 1310, 1370, 1380, and 1390 may be interconnected via a network 1350, such as the Internet. Other networks may also or alternatively be used, including private intranets, corporate networks, LANs, wireless networks, personal networks (PAN), and the like. Network 1350 is for illustration purposes and may be replaced with fewer or additional computer networks. A local area network (LAN) may have one or more of any known LAN topologies and may use one or more of a variety of different protocols, such as Ethernet. Devices 1310, 1370, 1380, and 1390 and other devices (not shown) may be connected to one or more of the networks via twisted pair wires, coaxial cable, fiber optics, radio waves or other communication media.
[0093] As seen in FIG. 13, computing device 1310 may include a processor 1312, RAM 1314, ROM 1315, network interface 1316, input / output interfaces 1318 (e.g., keyboard, mouse, display, printer, etc.), and memory 1340. Processor 1312 may include one or more computer processing units (CPUs), graphical processing units (GPUs), and / or other processing units such as a processor adapted to perform computations associated with machine learning. I / O 1318 may include a variety of interface units and drives for reading, writing, displaying, and / or printing data or files. I / O 1318 may be coupled with a display such as display 1360.
[0094] With regard to the graphical user interface (GUI) by which a clinician of an external video display device 1310, Rotate icons 124, 224, 324, 324, 524 in the respective GUIs 120, 220, 320, 420, 520 of FIGs. 1-5, this feature enables clinicians to perform 90-degree rotations of the video image displayed on the external video device 201 . Such functionality proves beneficial during intubation procedures, as anatomical structures can be observed from various angles to improve visualization. A Rotate Icon 124, 224, 324, 324, 524 communicates with the software system 1348 by sending commands to the processor 1312 of the external video display device 1310, which modifies the orientation of the live video stream accordingly. This capability allows clinicians to adjust the view to their preferred angle without disrupting the procedure, thereby enhancing the usability and precision of the guided intubation system 600, 1310.
[0095] With regard to the Settings Icons 126, 226, 326, 426, 526, shown in FIG. 1 through FIG. 5, access to important configuration options within the guided intubation system 600, 1310 is provided. Through this icon, clinicians can navigate to settings menus to manage saved images or videos, adjust system preferences, and access additional tools. The Settings Icon interfaces with the software system 1348 to retrieve and display stored data from the memory 1340 of the external video display device and allows clinicians to customize the system to their specific requirements, facilitating a seamless and efficient workflow during intubation procedures.
[0096] With regard to the On / Off Icons 128, 228, 328, 428, 528, illustrated in FIGs. 1-5, a clinician may turn on and off the Al algorithm image identification, extraction, and classification system 1320, 1330. The On / Off Icon serves as a toggle for activating or deactivating the artificial intelligence (Al) algorithm 1340 embedded in the guided intubation system. When activated, the Al algorithm 1340 begins analyzing the live video stream to identify anatomical landmarks and provide realtime feedback. Conversely, deactivating the algorithm halts these processes, allowing the clinician to use the system in a manual mode if desired. The On / Off Icon communicates directly with the processor 1312 of the external video display device to execute these commands, ensuring that the system operates according to the clinician’s preferences.
[0097] The Eye Icons 122, 222, 322, 422, 522 of FIGs. 1-5 functions as an orientation tool within the user interface for the clinician. This graphical element acts as a compass for clinicians, aiding in the identification of the anterior view 217 during real-time intubation. Such functionality proves advantageous in situations where anatomical structures may be obscured or challenging to distinguish. The Eye Icons integrate with the software system 1348 to overlay directional indicators on the live video stream, supporting clinicians in preserving accurate orientation during the procedure.
[0098] The Picture Icons 132, 232, 332, 432, 532, shown in FIGs. 1-5, enable clinicians to capture snapshot images of the live video stream during intubation. These images are stored in the memory 1312 of the external video display device for later review or documentation purposes. The Picture Icon interfaces with the software system 1348 to execute image capture commands, ensuring that high- resolution still images are saved without disrupting the real-time video feed.
[0099] Video Icons 134, 234, 334, 434, 534 of FIGs. 1-5 allows clinicians to record video clips of the live video stream during intubation. These recordings can be used for training, documentation, or post-procedure analysis. The Video Icon interacts with the software system 1348 to initiate and terminate video recording commands, storing the captured footage in the memory 1340 of the external video display device.
[0100] Memory 1340 may store software for configuring computing device 1310 into a special purpose computing device in order to perform one or more of the various functions discussed herein. Memory 1340 may store operating system software 1342 for controlling overall operation of computing device 1310, control logic 1344 for instructing computing device 1310 to perform aspects discussed herein, machine learning software 1348, training set data 1349, and other applications 1346. Control logic 1344 may be incorporated in and may be a part of machine learning software 1348. In other embodiments, computing device 1310 may include two or more of any and / or all of these components (e.g., two or more processors, two or more memories, etc.) and / or other components and / or subsystems not illustrated here. Moreover, as described above and shown in connection with block diagram of FIG. 6, the device has a feature extraction network 1320 and a classifier network 1330.
[0101] As previously described, the memory may be any storage devices that are non-transitory and can implement data stores, and may compromise but are not limited to an optical storage device, a solid-state storage, hard disk drive, or any other magnetic medium, a ROM (Read Only Memory), a RAM (Random Access Memory), a cache memory and / or any other memory chip or cartridge, and / or any other medium from which a computer may read data, instructions, and / or code.
[0102] Devices 1370, 1380, and 1390 may have similar or different architecture as described with respect to computing device 1310. Those of skill in the art will appreciate that the functionality of computing device 1310 (or device 1370, 1380, and 1390) as described herein may be spread across multiple data processing devices, for example, to distribute processing load across multiple computers, to segregate transactions based on geographic location, clinician access level, quality of service (QoS), etc. For example, computing devices 1310, 1370, 1380, 1390, andothers may operate in concert to provide parallel computing features in support of the operation of control logic 1344 and / or machine learning software 1348.
[0103] One or more aspects discussed herein may be embodied in computer- usable or readable data and / or computer-executable instructions, such as in one or more program modules, executed by one or more computers or other devices as described herein. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types when executed by a processor in a computer or other device. The modules may be written in a source code programming language that is subsequently compiled for execution, or may be written in a scripting language such as (but not limited to) HTML or XML. The computer executable instructions may be stored on a computer readable medium such as a hard disk, optical disk, removable storage media, solid state memory, RAM, etc. As will be appreciated by one of skill in the art, the functionality of the program modules may be combined or distributed as desired in various embodiments. In addition, the functionality may be embodied in whole or in part in firmware or hardware equivalents such as integrated circuits, field programmable gate arrays (FPGA), and the like. Particular data structures may be used to more effectively implement one or more aspects discussed herein, and such data structures are contemplated within the scope of computer executable instructions and computer-usable data described herein. Various aspects discussed herein may be embodied as a method, a computing device, a data processing system, or a computer program product.
[0104] FIG. 14 is a flow chart diagram 1400 that illustrates extraction of anatomical features of an image, subsequent processing by a classification network, and communication of feedback guidance to assist intubation in real-time, in accordance with embodiments of the present disclosure. At block 1410, an anatomical feature of an image is detected and extracted and compared to a trained algorithm to determine an identified anatomical feature of the image by a feature extraction network of a guided intubation system (GIS). At block 1420, a classification network of the GIS processes the determined identified anatomical feature of the image in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature. Finally, at block 1430, a display device, such as the external video display device shown in FIGs. 6 and 13,displays the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device.
[0105] FIG. 15 is a flow chart diagram 1500 that illustrates guidance information provided by a video display of a video display device to assist intubation in real-time, in accordance with embodiments of the present disclosure. This flow shows the operation of the external video display device itself. At block 1510, responsive to processing in accordance with a trained artificial intelligence algorithm and classification in accordance with a trained image classifier of images of a video stream captured during intubation, guidance information is provided (displayed) on a video display of a video display device, said guidance information including intubation placement and / or depth of placement information of an intubation tube vis- a-vis one or more internal anatomical structures. As described previously, unlike conventional approaches, a GIS system having such a video display device provides for imbedded instructions storable on the memory drive of a connected external video display device which gives an output of markings or grids delineating glottic, laryngeal, and esophageal structures during real-time intubation. These stored instructions are executed by the processor of the video display device. Moreover, all processing of the algorithm model may be processed on the external video display device, in which case the software application does not rely on cloud services for image detection and deployment of the artificial intelligence algorithm.
[0106] The following embodiments are combinable.
[0107] In one system embodiment, a guided intubation system (GIS) for guiding placement of an intubation tube during intubation of a patient includes: a feature extraction network configured, for one or more images of a plurality of images of a video stream captured during intubation of the patient, to detect and extract an anatomical feature of an image and compare the extracted anatomical feature to a trained artificial intelligence algorithm to determine an identified anatomical feature of the image by a feature extraction network of the guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of two or more interior anatomical structures of the patient sequentially passed during the intubation of the patient; an image classification network configured to process the determined identified anatomical feature of theimage in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and a display device configured to display the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image, where the guidance information displayed on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
[0108] In another embodiment of the system, the plurality of images of the video stream provided to the feature extraction network by a video capture device coupled to the display device.
[0109] In another embodiment of the system, where the extracted anatomical feature is identified using a prediction feature of the trained artificial intelligence algorithm, the feature extraction network using a Fast Fourier Transform algorithm to extract the anatomical feature of the image prior to the processing of the determined identified anatomical feature of the image by the image classification network of the GIS.
[0110] In another embodiment of the system, where the extracted anatomical feature is identified using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network of the GIS renders a bounding box around the extracted anatomical feature and overlays extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted to delineate the extracted anatomical feature.
[0111] In another embodiment of the system, the display device including a processor operable to execute instructions stored on one or more of a memory of the display device or on a memory drive of the display device to provide the guidance information viewable as guidance indicators on the display of the display device.
[0112] In another embodiment of the system, where the graphical representation includes an output of one or more of placement indicators markings and gridsdisplayed on the display device together with an image of a current position of the intubation tube.
[0113] In another embodiment of the system, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina, bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
[0114] In another embodiment of the system, where the graphical representation of the determined identified anatomical feature of the image includes a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.
[0115] In another embodiment of the system, where color-coded indicators of extracted anatomical features of two or more interior anatomical structures are sequentially displayed on the display device include a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
[0116] In another embodiment of the system, the display device including a video display device configurable to store and execute instructions by a processor of the video display device to display the image and the graphical representation of the attribute of the determined identified anatomical feature of the image, the video display device coupled to the image classification network of the GIS and external to the patient.
[0117] In another embodiment of the system, where the instructions are instructions stored in a memory of the display device or on a memory drive of the display device.
[0118] In another embodiment of the system, where the display device is a video display device coupled to one or more of the image classification network and the intubation tube.
[0119] In another embodiment of the system, where the intubation tube is an endotracheal intubation tube and the two or more internal anatomical structures are tracheal structures of the patient.
[0120] In another embodiment of the system, further including a video capture device configured to capture the plurality of images of the video stream duringintubation of the patient, the video capture device having a wired or wireless connection to the display device.
[0121] In another embodiment of the system, where the video capture device is coupled to one or more of the intubation tube and the display device, external to the two or more internal anatomical structures of the patient, configured to process video and images and perform calculations and may be a remote or a wired machine.
[0122] In another embodiment of the system, where the display device includes one or more of a work station, a server or a server farm, a smartphone, a tablet computer, a computing device, a desktop computer, a touchscreen, a television, a video display screen, and a personal data assistant.
[0123] In another embodiment of the system, where the image and the graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface the display device is further configured to perform one or more of: rotation of the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, access by the clinician to saved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function to orient a view, and access to an on and off function of the trained artificial intelligence algorithm.
[0124] In another embodiment of the system, the user interface including an iconbased user interface including icons accessible by the clinician.
[0125] In another embodiment of the system, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN) trained to identify.
[0126] In another embodiment of the system, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of the feature extraction network and the image classification network of the GIS and deployed locally on the display device without reliance on cloud-based computing.
[0127] In another embodiment of the system, where the trained artificial intelligence algorithm is exported as a pre-trained library for integration into the display device.
[0128] In another embodiment of the system, where the pre-trained library is a TorchScript library that is integrated into an application run by the display device using a plugin for real-time inference on the display device.
[0129] In another embodiment of the system, the feature extraction network further configured to perform one or more of: conversion of annotation files in Pascal Visual Object Classes (VOC) Extensible Markup Language (XML) formation into a You Only Look Once (YOLO) annotation format, and perform data augmentation on the plurality of images to include one or more of flipping images, jittering images, and rotating images of the plurality of images.
[0130] In another embodiment of the system, where the feature extraction network is configured for YOLO annotation format to resize one or more of the plurality of images to a resolution of 416x416 pixels for compatibility with YOLOv8.
[0131] In another embodiment of the system, the feature extraction network further configured to fine-tune a base YOLOv8 pre-trained on a Common Objects in Context (COCO) dataset using a plurality of labeled endotracheal intubation images.
[0132] In one device embodiment, a video display device for guiding placement during intubation of a patient includes: the video display device configured to be coupled to an endotracheal intubation tube and a video capture device, external to the patient, and configured to store and execute instructions by a processor of the video display device to: responsive to processing in accordance with a trained artificial intelligence algorithm and classification in accordance with a trained image classifier of one or more images of a plurality of images of a video stream captured during the intubation of the patient, provide guidance information viewable as guidance indicators on a video display of the video display device in a feedback loop to aid placement of the endotracheal intubation tube during the intubation, said guidance information providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis an internal anatomical structure of the two or more interior anatomical structures during the intubation, where the guidance information viewable as guidance indicators on the video displayof the video display device is updated in the feedback loop responsive to movement of the endotracheal intubation tube during the intubation.
[0133] In another embodiment of the device, where the graphical representation includes an output of one or more of placement indicators markings and grids displayed on the display device together with an image of a current position of the intubation tube.
[0134] In another embodiment of the device, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina, bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
[0135] In another embodiment of the device, where the graphical representation of the determined identified anatomical feature of the image includes a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.
[0136] In another embodiment of the device, where color-coded indicators of the determined identified anatomical features of the two or more interior anatomical structures are sequentially displayed on the display device includes a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
[0137] In another embodiment of the device, the video display device configured to store and execute instructions by the processor of the video display device to: display an image and a graphical representation of an attribute of a determined identified anatomical feature of the image as guidance information on a display of the display device, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image.
[0138] In another embodiment of the device, where an image and a graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface the display device is further configured to perform one or more of: rotation of the image or the graphical representation of the attribute of the determined identified anatomical feature of theimage, access by the clinician to saved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function to orient a view, and access to an on and off function of the trained artificial intelligence algorithm.
[0139] In another embodiment of the device, the user interface including an iconbased user interface including icons accessible by the clinician.
[0140] In another embodiment of the device, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
[0141] In another embodiment of the device, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of a feature extraction network and an image classification network and deployed locally on the display device without reliance on cloud-based computing.
[0142] In one method embodiment, a method for guiding placement of an intubation tube during intubation of a patient, including for one or more images of a plurality of images of a video stream captured during intubation: detecting and extracting an anatomical feature of an image and comparing the extracted anatomical feature to a trained artificial intelligence algorithm to determine an identified anatomical feature of the image by a feature extraction network of a guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of the two or more interior anatomical structures of a patient being intubated; processing by an image classification network of the GIS the determined identified anatomical feature of the image in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and displaying by a display device the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a- vis the determined identified anatomical feature of the image, where the guidanceinformation viewable on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
[0143] In another embodiment of the method, further including identifying the extracted anatomical feature using a prediction feature of the trained artificial intelligence algorithm, the extracting the anatomical feature using a Fast Fourier Transform algorithm prior to processing by the image classification network of the GIS.
[0144] In another embodiment of the method, where identifying the extracted anatomical feature includes using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network rendering a bounding box around the extracted anatomical feature in accordance with the trained image classifier and overlaying the extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted.
[0145] In another embodiment of the method, further including displaying by the display device one or more of placement indicators together with an image of a current position of the intubation tube, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina, bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
[0146] In another embodiment of the method, further including displaying the graphical representation of the determined identified anatomical feature of the image as a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.
[0147] In another embodiment of the method, further including sequentially displaying the color-coded indicators of determined identified anatomical features of two or more interior anatomical structures on the display device as a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
[0148] In another embodiment of the method, further including a processor of the display device executing instructions stored on one or more of a memory of the display device or on a memory drive of the display device to provide the guidance information viewable as guidance indicators on the display of the display device,where executing instructions may be performed by the display device connected to a communication network or offline.
[0149] In another embodiment of the method, displaying the image and the graphical representation of the attribute of the determined identified anatomical feature of the image in a user interface displayed by the display device and where responsive to input from a clinician of the user interface execution of the instructions by the one or more processors further cause one or more of: rotating the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, accessing by the clinician saved images or videos, accessing by the clinician a picture or a video function, accessing by the clinician a compass function to orient a view, and accessing by the clinician an on and off function of the trained artificial intelligence algorithm.
[0150] In another embodiment of the method, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
[0151] In another embodiment of the method, further including downloading one or more of the trained artificial intelligence algorithm and the trained image classifier to one or more of the feature extraction network and the image classification network of the GIS for local deployment on the display device without reliance on cloud-based computing.
[0152] In one computer readable media (CRM) embodiment, one or more non- transitory computer readable media (NTCRM) comprising instructions for guiding placement of an intubation tube during intubation of a patient, where execution of the instructions by one or more processors cause a display device to: detect and extract an anatomical feature of an image and compare the extracted anatomical feature to a trained algorithm to determine an identified anatomical feature of the image by a feature extraction network of a guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of the two or more interior anatomical structures of a patient being intubated; process by an image classification network of the GIS the determined identified anatomical feature of the image in accordance with a trained imageclassifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and display by a display device the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image, where the guidance information displayed on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
[0153] In another embodiment of the NTCRM, where execution of the instructions by the one or more processors further cause the display device to identify the extracted anatomical feature using a prediction feature of the trained artificial intelligence algorithm.
[0154] In another embodiment of the NTCRM, where identification of the extracted anatomical feature includes using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network renders a bounding box around the extracted anatomical feature and overlays the extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted in accordance with the trained image classifier.
[0155] In another embodiment of the NTCRM, where execution of the instructions by the one or more processors further cause the display device to sequentially display the color-coded indicators of extracted anatomical features of two or more interior anatomical structures on the display device as a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
[0156] In another embodiment of the NTCRM, where the image and the graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface execution of the instructions by the one or more processors further cause one or more of: rotation of the image or the graphical representation of the attribute of the determined identified anatomicalfeature of the image, access by the clinician to saved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function, and access by the clinician to an on and off function of the trained artificial intelligence algorithm.
[0157] In another embodiment of the NTCRM, the user interface including an iconbased user interface including icons accessible by the clinician.
[0158] In another embodiment of the NTCRM, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
[0159] In another embodiment of the NTCRM, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of the feature extraction network and the image classification network of the GIS and deployed locally on the display device without reliance on cloud-based computing.
[0160] While implementations of the disclosure are susceptible to embodiment in many different forms, there is shown in the drawings and will herein be described in detail specific embodiments, with the understanding that the present disclosure is to be considered as an example of the principles of the disclosure and not intended to limit the disclosure to the specific embodiments shown and described. In the description above, like reference numerals may be used to describe the same, similar or corresponding parts in the several views of the drawings.
[0161] In this document, relational terms such as first and second, top and bottom, and the like may be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element preceded by “comprises ...a” does not, withoutmore constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0162] Reference throughout this document to “one embodiment,” “certain embodiments,” “an embodiment,” “implementation(s),” “aspect(s),” or similar terms means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases or in various places throughout this specification are not necessarily all referring to the same embodiment.Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments without limitation.
[0163] The term “or” as used herein is to be interpreted as an inclusive or meaning any one or any combination. Therefore, “A, B or C” means “any of the following: A; B; C; A and B; A and C; B and C; A, B and C.” An exception to this definition will occur only when a combination of elements, functions, steps or acts are in some way inherently mutually exclusive. Also, grammatical conjunctions are intended to express any and all disjunctive and conjunctive combinations of conjoined clauses, sentences, words, and the like, unless otherwise stated or clear from the context. Thus, the term “or” should generally be understood to mean “and / or” and so forth. References to items in the singular should be understood to include items in the plural, and vice versa, unless explicitly stated otherwise or clear from the text.
[0164] Recitation of ranges of values herein are not intended to be limiting, referring instead individually to any and all values falling within the range, unless otherwise indicated, and each separate value within such a range is incorporated into the specification as if it were individually recited herein. The words “about,” “approximately,” or the like, when accompanying a numerical value, are to be construed as indicating a deviation as would be appreciated by one of ordinary skill in the art to operate satisfactorily for an intended purpose. Ranges of values and / or numeric values are provided herein as examples only, and do not constitute a limitation on the scope of the described embodiments. The use of any and all examples, or exemplary language (“e.g.,” “such as,” “for example,” or the like) provided herein, is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the embodiments. No language in thespecification should be construed as indicating any unclaimed element as essential to the practice of the embodiments.
[0165] For simplicity and clarity of illustration, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. Numerous details are set forth to provide an understanding of the embodiments described herein. The embodiments may be practiced without these details. In other instances, well-known methods, procedures, and components have not been described in detail to avoid obscuring the embodiments described. The description is not to be considered as limited to the scope of the embodiments described herein.
[0166] In the following description, it is understood that terms such as “first,” “second,” “top,” “bottom,” “up,” “down,” “above,” “below,” and the like, are words of convenience and are not to be construed as limiting terms. Also, the terms apparatus, device, system, etc. may be used interchangeably in this text.
[0167] The many features and advantages of the disclosure are apparent from the detailed specification, and, thus, it is intended by the appended claims to cover all such features and advantages of the disclosure which fall within the scope of the disclosure. Further, since numerous modifications and variations will readily occur to those skilled in the art, it is not desired to limit the disclosure to the exact construction and operation illustrated and described, and, accordingly, all suitable modifications and equivalents may be resorted to that fall within the scope of the disclosure.
Claims
WHAT IS CLAIMED IS:1 . A guided intubation system (GIS) for guiding placement of an intubation tube during intubation of a patient, comprising: a feature extraction network configured, for one or more images of a plurality of images of a video stream captured during intubation of the patient, to detect and extract an anatomical feature of an image and compare the extracted anatomical feature to a trained artificial intelligence algorithm to determine an identified anatomical feature of the image by a feature extraction network of the guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of two or more interior anatomical structures of the patient sequentially passed during the intubation of the patient; an image classification network configured to process the determined identified anatomical feature of the image in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and a display device configured to display the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image,where the guidance information displayed on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
2. The system of claim 1 , the plurality of images of the video stream provided to the feature extraction network by a video capture device coupled to the display device.
3. The system of claim 1 , where the extracted anatomical feature is identified using a prediction feature of the trained artificial intelligence algorithm, the feature extraction network using a Fast Fourier Transform algorithm to extract the anatomical feature of the image prior to the processing of the determined identified anatomical feature of the image by the image classification network of the GIS.
4. The system of claim 3, where the extracted anatomical feature is identified using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network of the GIS renders a bounding box around the extracted anatomical feature and overlays extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted to delineate the extracted anatomical feature.
5. The system of claim 1 , the display device including a processor operable to execute instructions stored on one or more of a memory of the display device or on a memory drive of the display device to provide the guidance information viewable as guidance indicators on the display of the display device.
6. The system of claim 1 , where the graphical representation includes an output of one or more of placement indicators markings and grids displayed on the display device together with an image of a current position of the intubation tube.
7. The system of claim 6, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina, bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
8. The system of claim 1 , where the graphical representation of the determined identified anatomical feature of the image includes a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.
9. The system of claim 8, where color-coded indicators of extracted anatomical features of two or more interior anatomical structures are sequentially displayed on the display device include a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
10. The system of claim 1 , the display device including a video display device configurable to store and execute instructions by a processor of the video display device to display the image and the graphical representation of the attribute of the determined identified anatomical feature of the image, the video display device coupled to the image classification network of the GIS and external to the patient.
11. The system of claim 1 , where the instructions are instructions stored in a memory of the display device or on a memory drive of the display device.
12. The system of claim 1 , where the display device is a video display device coupled to one or more of the image classification network and the intubation tube.
13. The system of claim 12, where the intubation tube is an endotracheal intubation tube and the two or more internal anatomical structures are tracheal structures of the patient.
14. The system of claim 1 , further comprising a video capture device configured to capture the plurality of images of the video stream during intubation of the patient, the video capture device having a wired or wireless connection to the display device.
15. The system of claim 14, where the video capture device is coupled to one or more of the intubation tube and the display device, external to the two or more internal anatomical structures of the patient, configured to process video and images and perform calculations and may be a remote or a wired machine.
16. The system of claim 1 , where the display device includes one or more of a work station, a server or a server farm, a smartphone, a tablet computer, a computing device, a desktop computer, a touchscreen, a television, a video display screen, and a personal data assistant.
17. The system of claim 1 , where the image and the graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface the display device is further configured to perform one or more of: rotation of the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, access by the clinician tosaved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function to orient a view, and access to an on and off function of the trained artificial intelligence algorithm.
18. The system of claim 17, the user interface including an icon-based user interface including icons accessible by the clinician.
19. The system of claim 1 , where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN) trained to identify.
20. The system of claim 19, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of the feature extraction network and the image classification network of the GIS and deployed locally on the display device without reliance on cloud-based computing.
21. The system of claim 20, where the trained artificial intelligence algorithm is exported as a pre-trained library for integration into the display device.
22. The system of claim 21 , where the pre-trained library is a TorchScript library that is integrated into an application run by the display device using a plugin for real- time inference on the display device.
23. The system of claim 1 , the feature extraction network further configured to perform one or more of: conversion of annotation files in Pascal Visual Object Classes (VOC) ExtensibleMarkup Language (XML) formation into a You Only Look Once (YOLO) annotation format, andperform data augmentation on the plurality of images to include one or more of flipping images, jittering images, and rotating images of the plurality of images.
24. The system of claim 23, where the feature extraction network is configured for YOLO annotation format to resize one or more of the plurality of images to a resolution of 416x416 pixels for compatibility with YOLOv8.
25. The system of claim 24, the feature extraction network further configured to fine-tune a base YOLOv8 pre-trained on a Common Objects in Context (COCO) dataset using a plurality of labeled endotracheal intubation images.
26. A video display device for guiding placement during intubation of a patient, comprising: the video display device configured to be coupled to an endotracheal intubation tube and a video capture device, external to the patient, and configured to store and execute instructions by a processor of the video display device to: responsive to processing in accordance with a trained artificial intelligence algorithm and classification in accordance with a trained image classifier of one or more images of a plurality of images of a video stream captured during the intubation of the patient, provide guidance information viewable as guidance indicators on a video display of the video display device in a feedback loop to aid placement of the endotracheal intubation tube during the intubation, said guidance information providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis an internal anatomical structure of the two or more interior anatomical structures during the intubation, where the guidance information viewable as guidance indicators on the video display of the video display device is updated in the feedback loop responsive to movement of the endotracheal intubation tube during the intubation.
27. The device of claim 26, where the graphical representation includes an output of one or more of placement indicators markings and grids displayed on the display device together with an image of a current position of the intubation tube.
28. The device of claim 27, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina,bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
29. The device of claim 26, where the graphical representation of the determined identified anatomical feature of the image includes a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.
30. The device of claim 29, where color-coded indicators of the determined identified anatomical features of the two or more interior anatomical structures are sequentially displayed on the display device includes a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.31 . The device of claim 26, the video display device configured to store and execute instructions by the processor of the video display device to: display an image and a graphical representation of an attribute of a determined identified anatomical feature of the image as guidance information on a display of the display device, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image.
32. The device of claim 26, where an image and a graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface the display device is further configured to perform one or more of:rotation of the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, access by the clinician to saved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function to orient a view, and access to an on and off function of the trained artificial intelligence algorithm.
33. The device of claim 26, the user interface including an icon-based user interface including icons accessible by the clinician.
34. The device of claim 26, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
35. The device of claim 34, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of a feature extraction network and an image classification network and deployed locally on the display device without reliance on cloud-based computing.
36. A method for guiding placement of an intubation tube during intubation of a patient, comprising for one or more images of a plurality of images of a video stream captured during intubation: detecting and extracting an anatomical feature of an image and comparing the extracted anatomical feature to a trained artificial intelligence algorithm to determine an identified anatomical feature of the image by a feature extraction network of a guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of the two or more interior anatomical structures of a patient being intubated; processing by an image classification network of the GIS the determined identified anatomical feature of the image in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and displaying by a display device the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image, where the guidance information viewable on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
37. The method of claim 36, further including identifying the extracted anatomical feature using a prediction feature of the trained artificial intelligence algorithm, the extracting the anatomical feature using a Fast Fourier Transform algorithm prior to processing by the image classification network of the GIS.
38. The method of claim 37, where identifying the extracted anatomical feature includes using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network rendering a bounding box around the extracted anatomical feature in accordance with the trained image classifier and overlaying the extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted.
39. The method of claim 36, further comprising displaying by the display device one or more of placement indicators together with an image of a current position of the intubation tube, where the one or more placement indicators are one or more of markings and grids that indicate glottic, tracheal, laryngeal, carina, bronchial, or esophageal structures of the patient captured by the video capture device during the intubation.
40. The method of claim 36, further comprising displaying the graphical representation of the determined identified anatomical feature of the image as a color-coded indicator displayed on the display device, the color-coded indicator representative of one or more of the two or more interior anatomical structures.41 . The method of claim 40, further comprising sequentially displaying the color- coded indicators of determined identified anatomical features of two or more interior anatomical structures on the display device as a color-coded column indicative of asequential passing of the tube vis-a-vis the two or more interior anatomical structures.
42. The method of claim 36, further comprising a processor of the display device executing instructions stored on one or more of a memory of the display device or on a memory drive of the display device to provide the guidance information viewable as guidance indicators on the display of the display device, where executing instructions may be performed by the display device connected to a communication network or offline.
43. The method of claim 36, displaying the image and the graphical representation of the attribute of the determined identified anatomical feature of the image in a user interface displayed by the display device and where responsive to input from a clinician of the user interface execution of the instructions by the one or more processors further cause one or more of: rotating the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, accessing by the clinician saved images or videos, accessing by the clinician a picture or a video function, accessing by the clinician a compass function to orient a view, and accessing by the clinician an on and off function of the trained artificial intelligence algorithm.
44. The method of claim 36, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
45. The method of claim 44, further comprising downloading one or more of the trained artificial intelligence algorithm and the trained image classifier to one or moreof the feature extraction network and the image classification network of the GIS for local deployment on the display device without reliance on cloud-based computing.
46. One or more non-transitory computer readable media (NTCRM) comprising instructions for guiding placement of an intubation tube during intubation of a patient, where execution of the instructions by one or more processors cause a display device to: detect and extract an anatomical feature of an image and compare the extracted anatomical feature to a trained algorithm to determine an identified anatomical feature of the image by a feature extraction network of a guided intubation system (GIS), the determined identified anatomical feature of the image representative of an internal anatomical structure of the two or more interior anatomical structures of a patient being intubated; process by an image classification network of the GIS the determined identified anatomical feature of the image in accordance with a trained image classifier to generate a graphical representation of an attribute of the determined identified anatomical feature; and display by a display device the image and the graphical representation of the attribute of the determined identified anatomical feature of the image as guidance information on a display of the display device in a feedback loop to aid placement of the intubation tube during the intubation, the graphical representation of the attribute providing information on one or more of intubation placement and depth of placement of the intubation tube vis-a-vis the determined identified anatomical feature of the image, where the guidance information displayed on the display of the display device is updated in the feedback loop responsive to movement of the intubation tube during the intubation.
47. The one or more media of claim 46, where execution of the instructions by the one or more processors further cause the display device to: identify the extracted anatomical feature using a prediction feature of the trained artificial intelligence algorithm.
48. The one or more media of claim 47, where identification of the extracted anatomical feature includes using a bounding box prediction feature of the trained artificial intelligence algorithm and the image classification network renders a bounding box around the extracted anatomical feature and overlays the extracted anatomical feature with the bounding box on the image from which the extracted anatomical feature was extracted in accordance with the trained image classifier.
49. The one or more media of claim 46, where execution of the instructions by the one or more processors further cause the display device to: sequentially display the color-coded indicators of extracted anatomical features of two or more interior anatomical structures on the display device as a color-coded column indicative of a sequential passing of the tube vis-a-vis the two or more interior anatomical structures.
50. The one or more media of claim 46, where the image and the graphical representation of the attribute of the determined identified anatomical feature of the image are displayed in a user interface displayed by the display device and where responsive to input from a clinician of the user interface execution of the instructions by the one or more processors further cause one or more of: rotation of the image or the graphical representation of the attribute of the determined identified anatomical feature of the image, access by the clinician tosaved images or videos, access by the clinician to a picture or a video function, access by the clinician to a compass function, and access by the clinician to an on and off function of the trained artificial intelligence algorithm.51 . The one or more media of claim 46, the user interface including an icon-based user interface including icons accessible by the clinician.
52. The one or more media of claim 46, where one or more of the trained artificial intelligence algorithm and the trained image classifier are based on one or more of a convolutional neural network (CNN) and a generative adversarial network (GAN).
53. The one or more media of claim 52, where one or more of the trained artificial intelligence algorithm and the trained image classifier are downloadable to one or more of the feature extraction network and the image classification network of the GIS and deployed locally on the display device without reliance on cloud-based computing.
Citation Information
Patent Citations
Facilitating tracheal intubation using an articulating airway management apparatus
US20160206189A1
Image-Guided Surgery System
US20220104884A1
Medical Visualization and Intubation Systems
US20230248232A1