Image display system, training data set, image display program, and diagnosis support system
The image display system uses a machine learning model to differentiate between the airway and other anatomical structures during tracheal intubation, enhancing accuracy and speed through auxiliary images, addressing the challenge of accidental intubation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NAKAMUARA HIROKI
- Filing Date
- 2025-08-06
- Publication Date
- 2026-06-04
AI Technical Summary
Existing image display systems during tracheal intubation, particularly in infants and children, struggle to accurately distinguish between the airway and anatomical structures like the esophagus, leading to a risk of accidental intubation, and there is a need for systems that can perform tracheal intubation quickly and accurately.
An image display system utilizing a trained model generated by machine learning to discriminate between the airway and other anatomical structures, accompanied by auxiliary images to guide accurate intubation, and an image display program to facilitate rapid and precise tracheal intubation.
The system enables accurate and rapid tracheal intubation by distinguishing between the airway and other anatomical structures, reducing the risk of esophageal intubation and improving procedural efficiency.
Smart Images

Figure JP2025027860_04062026_PF_FP_ABST
Abstract
Description
Image display system, training dataset, image display program, and diagnostic support system
[0001] The present invention relates to an image display system for displaying an image of the larynx in which the airway is formed on a display screen, a training dataset used for generating a trained model used therefor, an image display program for displaying an image of the larynx on a display screen, and a diagnostic support system for supporting the diagnosis of the larynx.
[0002] It has been proposed to perform observation support by taking and displaying an image of a target site during an examination such as an endoscopic examination (see, for example, Patent Document 1 below). When performing tracheal intubation, an image of the larynx is taken using a laryngoscope. At that time, it is common practice to perform tracheal intubation while displaying the taken image of the larynx on a display screen in real time.
[0003] Japanese Patent Application Laid-Open No. 2024-31468
[0004] When the patient is a child, especially an infant under 1 year old, the time taken for tracheal intubation is short, the glottis is undeveloped, and only the bone structure may be visible, so tracheal intubation is particularly difficult. Therefore, when images of the airway and esophagus are displayed on the display screen, there is a risk of accidentally inserting the tube into the esophagus, and an image display system that can perform tracheal intubation accurately and quickly is desired.
[0005] In addition, the problems during tracheal intubation are not limited to infants under 1 year old, but can also occur when the patient is a child over 1 year old or an adult. For example, if the esophagus can be identified based on the epiglottis, the laryngoscope can be guided from the epiglottis to the esophagus, and tracheal intubation can be performed accurately and quickly.
[0006] The present invention has been made in view of the above circumstances, and an object thereof is to provide an image display system, a training dataset, an image display program, and a diagnostic support system that can perform tracheal intubation accurately and quickly.
[0007] (1) The image display system according to the present invention is an image display system for displaying an image of the larynx in which the airway is formed on a display screen, and comprises a captured image display processing unit, a discrimination processing unit, and an auxiliary image display processing unit. The captured image display processing unit displays the captured image on a display screen. When images of the airway and anatomical structures other than the airway are displayed on the display screen, the discrimination processing unit discriminates between the airway and anatomical structures other than the airway using a trained model generated by machine learning based on training data using multiple captured images, including captured images of the airway and captured images of anatomical structures other than the airway. Based on the discrimination result by the discrimination processing unit, the auxiliary image display processing unit displays an auxiliary image on the display screen for distinguishing between the airway and anatomical structures other than the airway.
[0008] With this configuration, when performing tracheal intubation, if images of the airway and other anatomical structures are displayed on the screen, a pre-trained model generated by machine learning can be used to accurately distinguish between the airway and other anatomical structures. Based on the discrimination result, auxiliary images to distinguish between the airway and other anatomical structures are displayed on the screen, allowing for accurate and rapid tracheal intubation while referring to the auxiliary images.
[0009] (2) The discrimination processing unit may, when images of the airway and esophagus are displayed on the display screen, use the trained model to distinguish between the airway and the esophagus. In this case, the auxiliary image display processing unit may display an auxiliary image on the display screen to distinguish between the airway and the esophagus based on the discrimination result by the discrimination processing unit.
[0010] With this configuration, the trachea and esophagus can be distinguished and auxiliary images displayed on the screen. Therefore, the risk of accidentally intubating the esophagus is reduced, and tracheal intubation can be performed more accurately and quickly.
[0011] (3) The discrimination processing unit may, when images of the airway and epiglottis are displayed on the display screen, use the trained model to distinguish between the airway and the epiglottis. In this case, the auxiliary image display processing unit may display an auxiliary image on the display screen to distinguish between the airway and the epiglottis based on the discrimination result by the discrimination processing unit.
[0012] With this configuration, the trachea and epiglottis can be identified and auxiliary images can be displayed on the screen. Therefore, the laryngoscope can be guided from the epiglottis to the esophagus, allowing for more accurate and rapid tracheal intubation.
[0013] (4) The auxiliary image may include an airway symbol image that is displayed in correspondence with the airway image displayed on the display screen.
[0014] With this configuration, the airways can be clearly distinguished using symbolic images, allowing for more accurate and rapid tracheal intubation.
[0015] (5) The auxiliary image may include a symbol image for the vocal cords that is displayed in correspondence with the image of the vocal cords displayed on the display screen.
[0016] With this configuration, the vocal cords can be clearly distinguished by the symbolic image of the vocal cords, allowing for more accurate and rapid tracheal intubation.
[0017] (6) The training dataset according to the present invention is a training dataset used when generating a trained model used in the image display system, and includes a plurality of captured images, including captured images of the airway and captured images of anatomical structures other than the airway.
[0018] With this configuration, a trained model can be generated using machine learning with a training dataset, and this trained model can be used to accurately distinguish between the airway and other anatomical structures.
[0019] (7) The image display program according to the present invention is an image display program for displaying an image of the larynx that forms the airway on a display screen, and causes a computer to function as an image capture display processing unit, a discrimination processing unit, and an auxiliary image display processing unit. The image capture display processing unit displays the captured image on a display screen. When images of the airway and anatomical structures other than the airway are displayed on the display screen, the discrimination processing unit discriminates between the airway and anatomical structures other than the airway using a trained model generated by machine learning based on training data using multiple captured images, including captured images of the airway and captured images of anatomical structures other than the airway. Based on the discrimination result by the discrimination processing unit, the auxiliary image display processing unit displays an auxiliary image on the display screen for distinguishing between the airway and anatomical structures other than the airway.
[0020] With this configuration, an image display program is used to display auxiliary images on the screen to distinguish between the airway and other anatomical structures, allowing for accurate and rapid tracheal intubation while referring to the auxiliary images.
[0021] (8) The diagnostic support system according to the present invention is a diagnostic support system for supporting the diagnosis of the larynx, comprising an imaging device for capturing an image of the larynx, and an image display system for displaying the image of the larynx captured by the imaging device on a display screen.
[0022] With this configuration, when diagnosing the larynx, the image of the larynx captured by the imaging device is displayed on a screen, and by using the auxiliary image as a reference, tracheal intubation can be performed accurately and quickly, thereby supporting the diagnosis.
[0023] According to the present invention, tracheal intubation can be performed accurately and quickly while referring to auxiliary images.
[0024] This is a schematic diagram showing the overall configuration of a diagnostic support system according to one embodiment of the present invention. This is a schematic diagram showing an example of a larynx image displayed on the display screen. This is a block diagram for explaining the specific configuration of the image display system. This is a flowchart showing an example of processing when analyzing images taken with a laryngoscope. This is a diagram showing the calculation results of the ROC curve and AUC. This is a diagram showing an embodiment in which the airway, vocal cords, and esophagus are identified using a trained model and auxiliary images are displayed on the display screen.
[0025] 1. Diagram 1 of the overall configuration of the diagnostic support system is a schematic diagram showing the overall configuration of a diagnostic support system according to one embodiment of the present invention. This diagnostic support system is a system for supporting the diagnosis of the larynx, and is particularly suitable for use when performing tracheal intubation.
[0026] This diagnostic support system comprises a laryngoscope 1 and a computer 2 that receives data from the laryngoscope 1. In this example, the laryngoscope 1 is connected to the computer 2 via a cable, but the transmission and reception of data between the laryngoscope 1 and the computer 2 is not limited to wired connections and may be performed wirelessly.
[0027] The laryngoscope 1 is an example of an imaging device for taking images of the larynx. The laryngoscope 1 comprises a main body 11 and an insertion part 12 that are connected to each other. The operator grasps the main body 11 of the laryngoscope 1 and inserts the insertion part 12 through the mouth of the patient H. This brings the tip of the insertion part 12 of the laryngoscope 1 close to the vicinity of the larynx inside the mouth of the patient H, enabling the laryngoscope 1 to take images of the larynx.
[0028] The main unit 11 is a so-called handle and contains a circuit board and memory (neither of which are shown). Communication between the laryngoscope 1 and the computer 2 is established via a communication interface mounted on the circuit board of the main unit 11. The main unit 11 may also be provided with a display screen for displaying images taken by the laryngoscope 1 in real time.
[0029] The insertion section 12 is a so-called blade, and either a curved shape such as a Macintosh type or a straight shape such as a mirror type is selectively used. A small camera (not shown) for taking images of the larynx is attached to the tip of the insertion section 12.
[0030] Computer 2 constitutes an image display system for displaying images of the larynx captured by the laryngoscope 1 on the display screen 21. Based on the data input from the laryngoscope 1, Computer 2 displays the images captured by the laryngoscope 1 on the display screen 21. The operator can insert tube T into the trachea of patient H while confirming the images of the larynx captured by the laryngoscope 1 on the display screen 21 of Computer 2.
[0031] 2. Image of the larynx Figure 2 is a schematic diagram showing an example of an image of the larynx 210 displayed on the display screen 21. The larynx 210 has an airway 211. The airway 211 is formed by the vocal cords 212, epiglottis 213, and arytenoid cartilage 214, and when breathing, the vocal cords 212 are open as shown in Figure 2.
[0032] Furthermore, the esophagus 215 is formed near the airway 211, adjacent to the arytenoid cartilage 214. Normally, when the esophagus 215 is open, the vocal cords 212 close, preventing food from entering the airway 211. Thus, the airway 211 and the esophagus 215 are in close proximity, and images of both the airway 211 and the esophagus 215 may be displayed simultaneously on the display screen 21.
[0033] 3. Diagram 3 of the specific configuration of the image display system is a block diagram illustrating the specific configuration of the image display system. In addition to the computer 2 mentioned above, this image display system is equipped with a storage unit 3 and the like.
[0034] Computer 2 is, for example, a personal computer, and a desktop or notebook personal computer may be used. However, a tablet or smartphone may also be used as Computer 2. In other words, Computer 2 can be any terminal equipped with a display screen 21.
[0035] Computer 2 is equipped with a control unit 22 and a display unit 23. The control unit 22 is equipped with a processor such as a CPU (Central Processing Unit), and the operation of computer 2 can be controlled by the processor executing a program. The display unit 23 is composed of, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and is equipped with a display screen 21 for displaying images. Note that the display unit 23 is not limited to being equipped with computer 2, but may be provided separately from computer 2.
[0036] The control unit 22 is connected to the display unit 23, as well as the laryngoscope 1 and the storage unit 3. These connections may be wired or wireless. The control unit 22 functions as an image display processing unit 221, a discrimination processing unit 222, and an auxiliary image display processing unit 223, etc., when the processor executes a program (image display program).
[0037] The storage unit 3 may be an external memory connected to the computer 2 by wired or wireless connection, or a cloud server connected to the computer 2 via the Internet. Alternatively, the storage unit 3 may consist of an HDD (Hard Disk Drive) or SSD (Solid State Drive), or other memory built into the computer 2.
[0038] The image capture and display processing unit 221 performs processing to display the image captured by the laryngoscope 1 on the display screen 21 of the display unit 23. In other words, the image capture and display processing unit 221 displays the image data input from the laryngoscope 1 on the display screen 21 in real time, enabling the operator to perform the work while checking the image displayed on the display screen 21. The image displayed on the display screen 21 is preferably a video displayed in real time, but still images captured at regular intervals may be switched and displayed as time progresses.
[0039] The discrimination processing unit 222 performs processing to distinguish between the airway 211 and the esophagus 215 when images of the airway 211 and esophagus 215 are displayed on the display screen 21 as shown in Figure 2. At this time, the discrimination processing unit 222 distinguishes between the airway 211 and the esophagus 215 using a trained model 31 generated by machine learning. This trained model 31 is generated by machine learning based on training data using multiple images, including images of the airway 211 and images of the esophagus 215.
[0040] The auxiliary image display processing unit 223 performs processing to display an auxiliary image on the display screen 21 to distinguish between the airway 211 and the esophagus 215, based on the discrimination result from the discrimination processing unit 222. This allows the operator to perform tracheal intubation accurately and quickly while referring to the auxiliary image. In particular, because the trachea 211 and the esophagus 215 can be distinguished and the auxiliary image displayed on the display screen, the risk of mistakenly intubating into the esophagus 215 is reduced, and tracheal intubation can be performed more accurately and quickly.
[0041] The pre-trained model 31 used for discrimination by the discrimination processing unit 222 is stored in the memory unit 3 beforehand. This pre-trained model 31 is generated by a general-purpose learning algorithm 4, such as a learning algorithm for object detection such as YOLO. The training dataset 5 used when generating the pre-trained model includes multiple images, including images of the airway 211 (airway image 51) and images of the esophagus 215 (esophageal image 52).
[0042] 4. Image Analysis Processing Figure 4 is a flowchart showing an example of the processing when analyzing images taken by the laryngoscope 1. When diagnosing the larynx using the laryngoscope 1, the images (videos) taken by the laryngoscope 1 are continuously analyzed (step S101) until the diagnosis is completed (until Yes is reached in step S108), and the processing in steps S102 to S107 is performed according to the results of the analysis.
[0043] In this example, when images of the airway 211 and the esophagus 215 are displayed on the display screen 21, not only the airway 211 and the esophagus 215 but also the vocal cords 212 that form the entrance of the airway 211 are discriminated. That is, the learned model 31 used in the discrimination by the discrimination processing unit 222 is generated by machine learning based on teacher data using a plurality of captured images including the captured image of the airway 211, the captured image of the esophagus 215, and the captured image of the vocal cords 212.
[0044] By image analysis using such a learned model 31, the airway 211, the vocal cords 212, and the esophagus 215 included in the captured image are discriminated. For the image of the airway 211 displayed on the display screen 21 (Yes in step S102), the airway symbol image is displayed in association as an auxiliary image (step S103). For the image of the vocal cords 212 (Yes in step S104), the vocal cord symbol image is displayed in association as an auxiliary image (step S105). For the image of the esophagus 215 (Yes in step S106), the esophagus symbol image is displayed in association as an auxiliary image (step S107).
[0045] However, the configuration may be such that only the airway symbol image and the esophagus symbol image are displayed without displaying the vocal cord symbol image, or the configuration may be such that only the airway symbol image is displayed.
[0046] 5. Example Hereinafter, an example of discriminating the airway 211, the vocal cords 212, and the esophagus 215 will be described. To generate a learned model, 1,179 still images were cut out from 653 cases for infants under general anesthesia, and images were prepared in 9 variations. Specifically, it is a learning data set of 1,179 images including 653 images of the airway 211 immediately before intubation, 335 images of the arytenoid cartilage 214, 139 images of the esophagus 215, and 52 images of other anatomical structures.
[0047] As the learning algorithm 4, YOLOv8 was used, and the teacher data constituting the learning dataset was prepared by dividing it into three groups of train data, validation data, and test data at a ratio of 6.4:1.6:2. The number of divisions of the learning dataset is as shown in Table 1 below, and the learning conditions of YOLOv8 are as shown in Table 2 below. Note that A1 to A5 are images of the airway 211 immediately before intubation (total 653 images), B1 to B2 are images of the arytenoid cartilage 214 (total 335 images), C is an image of the esophagus 215 (139 images), and D is an image of other anatomical structures (52 images).
[0048]
[0049] (Regarding "RandAugmentation", refer to Proc of CVPR Workshops 2020. pp. 702-703)
[0050] As an evaluation method, the classification results were evaluated by calculating recall (Re), precision (Pr), F-score, Mean Average Precision (mAP), mAP50, etc. Also, as an evaluation of the detection results, accuracy (Ac), sensitivity (Sn), and specificity (Sp) were calculated. At this time, the confidence score was based on 0.5, and the determination results of TP (True-Positive), FP (False-Positive), and FN (False-Negative) were calculated according to Table 3 below according to the combination of Ground truth, Prediction, and IoU (Intersection over Union).
[0051]
[0052] Furthermore, as an evaluation of the detection results, an ROC (Receiver Operating Characteristic) curve was drawn and the AUC (Area Under the Curve) was calculated. Figure 5 is a diagram showing the calculation results of the ROC curve and AUC.
[0053] The evaluation of the classification results is shown in Table 4 below. The evaluation of the detection results is shown in Table 5 below. In Tables 4 and 5, vc represents the vocal cords, aw represents the airway, es represents the esophagus, and All represents the overall average value.
[0054]
[0055]
[0056] First, according to the evaluation results in Table 4, the vocal cords (VC) and airway (AW) could be classified with an accuracy of over 90%, and were higher than the esophagus (ES) in all evaluation indices.
[0057] According to the evaluation results for precision (Ac), sensitivity (Sn), and specificity (Sp) in Table 5, the vocal cords (vc) and airway (aw) also showed high rectangular detection accuracy, and their detection accuracy was higher than that of the esophagus (es).
[0058] According to the AUC in Table 5 and the evaluation results in Figure 5, the vocal cords (vc) and airway (aw) had higher AUCs and superior ROC curves than the esophagus (es).
[0059] As shown in the evaluation results above, in this example, the detection accuracy of the vocal cords (vc) and airway (aw) was high, and it was possible to identify not only the airway (aw) but also the vocal cords (vc), suggesting the possibility of creating an AI that can also detect arytenoid cartilage. Furthermore, in this example, the detection accuracy of the esophagus (es) was lower compared to the vocal cords (vc) and airway (aw), but this is thought to be due to the difference in the number of images included in the training dataset. While it is possible to distinguish between the vocal cords (vc), airway (aw), and esophagus (es), it was also found that there is potential to further improve the detection accuracy of the esophagus (es).
[0060] Figure 6 shows an example in which the airway 211, vocal cords 212, and esophagus 215 are identified using a trained model, and auxiliary images 230 are displayed on the display screen 21. In this example, the airway symbol image 231, the vocal cord symbol image 232, and the esophagus symbol image 235 are displayed on the display screen 21 as auxiliary images 230 through image analysis as illustrated in Figure 4.
[0061] In other words, when images of the airway 211 and esophagus 215 are displayed on the display screen 21 as shown in Figure 6, the airway symbol image 231 is displayed in correspondence with the image of the airway 211, the vocal cord symbol image 232 is displayed in correspondence with the image of the vocal cords 212, and the esophagus symbol image 235 is displayed in correspondence with the image of the esophagus 215.
[0062] In this example, the airway symbol image 231 is displayed so as to surround an area that includes not only the airway 211 but also the epiglottis 213 and the arytenoid cartilage 214. The vocal cord symbol image 232 is displayed so as to surround the vocal cords 212 within the area surrounded by the airway symbol image 231. The esophagus symbol image 235 is displayed so as to surround the esophagus 215. Thus, in this invention, "airway" may include the airway 211 and the epiglottis 213 and arytenoid cartilage 214 that constitute the airway 211, or it may mean the airway 211 itself or the vocal cords 212 that form the entrance to the airway 211.
[0063] However, each auxiliary image 230 is not limited to a symbol image displayed so as to surround the target anatomical structure. In other words, each auxiliary image 230 can be any symbol image that clearly indicates the location of the target anatomical structure to the operator, and is not limited to a rectangular symbol image; any symbol image such as a circle, ellipse, or cross can be used as an auxiliary image 230.
[0064] 6. Modified Examples In the above embodiment, when images of the airway 211 and esophagus 215 are displayed on the display screen 21, the airway 211 and esophagus 215 are distinguished, and based on the discrimination result, an auxiliary image 230 for distinguishing the airway 211 and esophagus 215 is displayed on the display screen 21. However, the configuration is not limited to this, and any configuration that distinguishes the airway 211 and anatomical structures other than the airway 211, and based on the discrimination result, an auxiliary image 230 for distinguishing the airway 211 and anatomical structures other than the airway 211 is displayed on the display screen 21 is acceptable.
[0065] In this case, the trained model used to distinguish between the airway 211 and other anatomical structures may be generated by machine learning based on training data using multiple images, including images of the airway 211 and images of other anatomical structures. For example, by using a trained model generated by machine learning based on training data using multiple images, including images of the airway 211 and images of the epiglottis 213, it is possible to distinguish between the airway 211 and the epiglottis 213.
[0066] In other words, when images of the airway 211 and epiglottis 213 are displayed on the display screen 21, the discrimination processing unit 222 may use the trained model generated as described above to distinguish between the airway 211 and the epiglottis 213. In this case, the auxiliary image display processing unit 223 may display an auxiliary image 230 on the display screen 21 to distinguish between the airway 211 and the epiglottis 213 based on the discrimination result by the discrimination processing unit 222.
[0067] In this case, the display screen 21 may be configured to show instructions on how to move the laryngoscope 1 from the epiglottis 213 to the airway 211 (or the vocal cords 212 which form the entrance to the airway 211).
[0068] 1 Laryngoscope 2 Computer 3 Memory Unit 4 Learning Algorithm 5 Training Dataset 11 Main Unit 12 Insertion Unit 21 Display Screen 22 Control Unit 23 Display Unit 51 Airway Image 52 Esophageal Image 210 Larynx 211 Airway 212 Vocal Cords 213 Epiglottis 214 Arytenoid Cartilage 215 Esophagus 221 Image Display Processing Unit 222 Discrimination Processing Unit 223 Auxiliary Image Display Processing Unit 230 Auxiliary Image 231 Airway Symbol Image 232 Vocal Cord Symbol Image 235 Esophageal Symbol Image
Claims
1. An image display system for displaying an image of the larynx in which the airway is formed on a display screen, comprising: an image display processing unit that displays a captured image on a display screen; a discrimination processing unit that, when images of the airway and other anatomical structures are displayed on the display screen, distinguishes between the airway and other anatomical structures using a trained model generated by machine learning based on training data using multiple captured images, including images of the airway and images of other anatomical structures; and an auxiliary image display processing unit that displays an auxiliary image on the display screen to distinguish between the airway and other anatomical structures based on the discrimination result by the discrimination processing unit.
2. The image display system according to claim 1, wherein the discrimination processing unit, when images of the airway and esophagus are displayed on the display screen, uses the trained model to distinguish between the airway and the esophagus, and the auxiliary image display processing unit displays an auxiliary image for distinguishing between the airway and the esophagus on the display screen based on the discrimination result by the discrimination processing unit.
3. The image display system according to claim 1, wherein the discrimination processing unit, when images of the airway and epiglottis are displayed on the display screen, uses the trained model to distinguish between the airway and the epiglottis, and the auxiliary image display processing unit displays an auxiliary image for distinguishing between the airway and the epiglottis on the display screen based on the discrimination result by the discrimination processing unit.
4. The image display system according to claim 2 or 3, wherein the auxiliary image includes an airway symbol image that is displayed in correspondence with the airway image displayed on the display screen.
5. The image display system according to claim 4, wherein the auxiliary image includes a symbol image for vocal cords that is displayed in correspondence with the image of vocal cords displayed on the display screen.
6. A training dataset used to generate a trained model for use in the image display system described in claim 1, comprising a plurality of captured images including images of the airway and images of anatomical structures other than the airway.
7. An image display program for displaying an image of the larynx that forms the airway on a display screen, comprising: an image display processing unit that displays a captured image on a display screen; a discrimination processing unit that, when images of the airway and other anatomical structures are displayed on the display screen, distinguishes between the airway and other anatomical structures using a trained model generated by machine learning based on training data including multiple captured images of the airway and other anatomical structures; and an auxiliary image display processing unit that displays an auxiliary image on the display screen to distinguish between the airway and other anatomical structures based on the discrimination result by the discrimination processing unit, thereby causing the computer to function as an auxiliary image display processing unit.
8. A diagnostic support system for assisting in the diagnosis of the larynx, comprising: an imaging device for capturing an image of the larynx; and an image display system according to claim 1 for displaying the image of the larynx captured by the imaging device on a display screen.