Trachea cannula auxiliary guiding method based on deep learning and related device
Through the deep learning-based tracheal intubation assisted guidance method, the key features in the glottic structure are identified and positioned, and the problem that non-medical professionals find it difficult to quickly judge the airway structure in first aid tracheal intubation is solved, and the accurate auxiliary guidance of tracheal intubation is achieved, which improves the success rate and safety of first aid tracheal intubation.
Patent Information
- Application Number
- CN202510235598.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
During the first aid tracheal intubation process, it is difficult for non-medical professionals to form the ability to judge whether the airway structure is suitable for intubation in a short period of time, resulting in increased difficulty in intubation of first aid tracheal intubation.
Using deep learning-based tracheal intubation assisted guidance method, data augmentation and model training are performed by acquiring and labeling glottis images, glottis recognition models are generated, which are used to identify and classify tracheal intubation images, and to locate key features such as epiglottic cartilage, arytenoid cartilage and glottis fissures to assist in guiding tracheal intubation operation.
It effectively reduces the difficulty of on-site rescue personnel to complete first aid tracheal intubation, and provides accurate auxiliary guidance by correctly identifying and positioning key features in the glottic structure, improving the success rate and safety of first aid tracheal intubation.
Smart Images

Figure CN120164024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to a tracheal intubation assistance guidance method and related device based on deep learning. Background Art
[0002] The application of video laryngoscopes has significantly reduced the difficulty for first aid professionals to perform tracheal intubation, but it does not mean that it is equally easy for other first aid professionals to perform tracheal intubation. Emergency tracheal intubation also requires timely judgment on whether the airway structure is normal and suitable for intubation, and rapid completion of key clinical decisions. This decision-making ability usually requires long-term clinical practice and is usually difficult to obtain in ordinary simulated human tracheal intubation training. Non-medical on-site rescuers such as police, firefighters, or security guards are even more difficult to develop this judgment ability in a short time. In the current first aid system in China, those who arrive at the first aid scene fastest are often not medical professionals. Therefore, the introduction of appropriate AI-assisted decision-making tools can help on-site early rescuers make timely decisions and perform rapid rescues. Summary of the Invention
[0003] Aiming at the deficiencies in the prior art, the present invention provides a tracheal intubation guidance method and related device based on deep learning, which can correctly identify, detect, and classify tracheal intubation images and videos, classify and label the regions of the epiglottic cartilage, arytenoid cartilage, and glottis fissure structures in the images, and annotate the probability belonging to this region. The present invention can also correctly locate the key position of the glottis based on the recognition results of the glottis fissure and arytenoid cartilage, and sequentially guide the tracheal intubation operation, further reducing the difficulty for on-site rescuers to complete emergency tracheal intubation.
[0004] A tracheal intubation assistance guidance method based on deep learning includes:
[0005] Obtaining an original glottis image set;
[0006] Annotating the original glottis image set to obtain an annotated data set;
[0007] Performing data augmentation processing on the glottis images in the annotated data set through random transformation processing to obtain a sample data set;
[0008] Training an initial recognition model using the normalized sample data set to generate a glottis recognition model;
[0009] Using the glottis recognition model to perform recognition and positioning on the obtained image to be detected to obtain a recognition and positioning result.
[0010] Optionally, the annotated data set includes a glottis image set and annotation information, and the annotation information includes a classification label and a bounding box, and the classification label includes an epiglottic cartilage, an arytenoid cartilage, and a glottis fissure.
[0011] Optionally, the random transformation process includes image hue shift, image saturation shift, image brightness shift, image rotation, image translation, image scaling, image shearing, and image flipping.
[0012] Optionally, the initial recognition model is trained using the normalized sample data set to generate a glottis recognition model, including:
[0013] Build a YOLOv5m model based on the PyTorch framework, pre-train the YOLOv5m model to obtain an initial recognition model;
[0014] Iteratively train the initial recognition model with the normalized sample data set to generate a glottis recognition model.
[0015] Optionally, it also includes using five-fold cross-validation to evaluate the performance of the glottis recognition model.
[0016] Optionally, the performance evaluation metrics of the glottis recognition model include recall rate, accuracy rate, F1 value, mAP_50, and AP_50:90 value.
[0017] A glottis recognition device based on deep learning, including an acquisition unit, a feature extraction unit, a generation unit, and a recognition unit, where:
[0018] The acquisition unit is used to acquire and randomly transform the original glottis image set, normalize the image color and size, and then transfer the data set to the feature extraction unit;
[0019] The feature extraction unit is used to extract target features from the input glottis image set or video, obtain the recognition target features through a convolutional neural network, and then transfer them to the generation unit;
[0020] The generation unit is used to predict and generate the category, position, and range of all target bounding boxes, calculate the probability that all bounding boxes belong to the target structure, and transfer them to the recognition unit;
[0021] The recognition unit is used to integrate the recognition prediction results, eliminate overlapping and redundant bounding boxes according to the probability that the bounding boxes belong to the target structure, display the best target range bounding box and prompt the probability of belonging to the target, and output the final recognition result after integration.
[0022] A terminal includes a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute any one of the above-mentioned deep learning-based glottis recognition methods.
[0023] A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and the computer program includes program instructions which, when executed by a processor, cause the processor to execute any one of the above-mentioned glottis recognition methods based on deep learning.
[0024] Advantages of the present invention: The present invention can correctly identify and classify the image to be detected, and display the probability that the image to be detected belongs to the glottis. Moreover, the glottis recognition model can correctly locate and classify the key feature positions in the image to be detected, such as the epiglottic cartilage, arytenoid cartilage, and glottis fissure positions in the glottis structure, and comprehensively weight the positions of the epiglottic cartilage, arytenoid cartilage, and glottis fissure to display the recognition result of the glottis direction on the display page. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the specific embodiments of the present invention, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. In all the drawings, the elements or parts are not necessarily drawn to actual scale.
[0026] Figure 1 It is a flowchart of a tracheal intubation assistance and guidance method based on deep learning provided by an embodiment of the present invention;
[0027] Figure 2 It is part of the sample data after data augmentation processing of the glottis images in the labeled dataset provided by an embodiment of the present invention;
[0028] Figure 3 is a comparison of the training loss curves of the initial recognition model without pre-training and the initial recognition model with pre-training provided by an embodiment of the present invention;
[0029] Figure 4 It is a curve of the test accuracy and operation speed of YOLOv5 models of different sizes pre-trained on the COCO dataset provided by an embodiment of the present invention;
[0030] Figure 5 It is a curve of various indicators for visualizing the training of the initial recognition model using tensorboard provided by an embodiment of the present invention;
[0031] Figure 6 It is a schematic diagram for testing and evaluating the glottis recognition model using five-fold cross-validation provided by an embodiment of the present invention;
[0032] Figure 7 It is a schematic structural diagram of a glottis recognition device based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The embodiments of the technical solution of the present invention will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present invention more clearly, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0034] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.
[0035] In one of the embodiments, as Figure 1 shown, a tracheal intubation assistance and guidance method based on deep learning is provided, including:
[0036] 1. Obtain the original glottis image set;
[0037] The original glottis image set is obtained through an endoscope and other visualization devices such as a light source device and image processing software.
[0038] 2. Annotate the original glottis image set to obtain an annotated data set;
[0039] Preferably, in one of the embodiments, a professional emergency doctor uses an internally built CVAT platform to annotate the original glottis image set to obtain an annotated data set.
[0040] CVAT (Computer Vision Annotation Tool) is an open-source, general-purpose, web-based computer vision annotation tool used for annotating tasks such as object detection, image classification, target tracking, and semantic segmentation on image, video, and point cloud data. It also supports multiple people to collaborate on annotating the same data set, and the annotation process can be collaborated through the network.
[0041] Preferably, in one of the embodiments, the annotated data set includes a glottis image set and annotation information. The annotation information includes classification labels and bounding boxes. The classification labels include the rima glottidis, arytenoid cartilage, and epiglottic cartilage. The key positions of the rima glottidis, arytenoid cartilage, and epiglottic cartilage in the original glottis image set are selected using the bounding boxes.
[0042] 0 is used to represent the rima glottidis in the classification label, 1 is used to represent the arytenoid cartilage in the classification label, 2 is used to represent the epiglottic cartilage in the classification label, and the bounding box is represented by the numerical values of the 4 vertex coordinates.
[0043] Store the annotation information (including classification labels and bounding boxes) belonging to the same glottis image together. Here, 0 represents the glottis in the classification label, 1 represents the arytenoid cartilage in the classification label, 2 represents the epiglottis cartilage in the classification label, and the bounding box is represented by the numerical values of the four vertex coordinates, forming a file in a general format such as a txt or jpg file, like Image 01.jpg or Label 01.txt. And the content inside the picture is as follows:
[0044] --0 0.756354 0.420104 0.456822 0.762135
[0045] Among them, 0 indicates that the classification label of this glottis image is the glottis fissure, and 0.756354 0.420104 0.456822 0.762135 represents the numerical values of the four vertex coordinates of the bounding box of this glottis image.
[0046] --1 0.723117 0.611851 0.543755 0.423945
[0047] Among them, 1 indicates that the classification label of this glottis image is the arytenoid cartilage, and 0.723117 0.611851 0.543755 0.423945 represents the numerical values of the four vertex coordinates of the bounding box of this glottis image.
[0048] --2 0.886752 0.332726 0.158526 0.383494
[0049] Among them, 2 indicates that the classification label of this glottis image is the epiglottis cartilage, and 0.886752 0.332726
[0050] 0.158526 0.383494 represents the numerical values of the four vertex coordinates of the bounding box of this glottis image.
[0051] 3. Perform data augmentation processing on the glottis images in the annotation dataset through random transformation processing to obtain a sample dataset;
[0052] By performing a series of random transformations and augmentation processing on the glottis images in the annotation dataset, the data volume is increased, the generalization ability of the glottis recognition model is improved, overfitting is prevented, and the glottis recognition model is made more robust.
[0053] The random transformation processing of the present invention includes image hue shift, image saturation shift, image brightness shift, image rotation, image translation, image scaling, image shearing, and image flipping.
[0054] Partial parameter settings of the random transformation processing are as follows:
[0055] hsv_h: 0.015
[0056] hsv_s: 0.7
[0057] hsv_v: 0.4
[0058] degrees: 0.1
[0059] translate: 0.1
[0060] scale: 0.5
[0061] shear: 0.0
[0062] flipud: 0.0
[0063] fliplr: 0.5
[0064] Among them, hsv_h, hsv_s, and hsv_v represent the hue offset, saturation offset, and value offset of the glottis images in the labeled dataset respectively; degrees is the rotation angle range of the glottis images, translate is the translation range of the glottis images, scale is the scaling range of the glottis images, and shear is the shear range of the glottis images; flipud and fliplr represent the probabilities of flipping the glottis images vertically and horizontally respectively.
[0065] Part of the sample data after data augmentation of the glottis images in the labeled dataset through random transformation processing is as Figure 2 shown.
[0066] 4. Use the normalized sample dataset to train the initial recognition model to generate a glottis recognition model;
[0067] (1) Preferably, in one embodiment, provide a YOLOv5m model built based on the PyTorch framework, and pre-train the YOLOv5m model to obtain an initial recognition model;
[0068] YOLOv5 is an object detection model based on deep learning. Compared with the previous YOLOv4, YOLOv5 adopts a more lightweight network structure and a more efficient training strategy, making it have a faster inference speed and a smaller model size while maintaining a high accuracy rate, and has good versatility and practicality.
[0069] In one embodiment, use YOLOv5m as the pre-training model, and pre-train YOLOv5m with a large-scale dataset, so that the initial recognition model can learn rich feature representations and prior knowledge, which can improve the accuracy and generalization ability of the initial recognition model and reduce the training time.
[0070] Figures 3(a) and 3(b) show the comparison of the training loss curves of the initial recognition model without pre-training - Figure 3(a) and the initial recognition model with pre-training - Figure 3(b) under the same sample dataset. It can be seen that the training speed of the initial recognition model without pre-training is slower and it is prone to overfitting problems.
[0071] Figure 4 They are the curves of the test accuracy and operation speed of YOLOv5 models with different sizes under pre-training on the COCO dataset (the COCO dataset is a large-scale dataset). Comparatively speaking, the YOLOv5m pre-trained model can better balance accuracy and performance.
[0072] Some parameter settings in the YOLOv5m pre-trained model are as follows:
[0073] epochs: 300
[0074] batch_size: 16
[0075] imgsz: 640
[0076] optimizer: SGD
[0077] workers: 8
[0078] Among them, epochs is the number of training rounds (i.e., the number of times to traverse the entire COCO dataset). Appropriate epochs can better iterate the learning of the YOLOv5m model and avoid overfitting; batch_size is the number of images in each batch. The larger it is, the faster the calculation speed and the higher the GPU occupancy rate; imgsz is the size of the input image; optimizer is the type of optimizer; workers is the number of threads used by the data loader.
[0079] (2) Iteratively train the initial recognition model with the normalized sample dataset to generate a glottis recognition model.
[0080] Preferably, in one embodiment, 298 sample data are divided into a training set and a validation set in a ratio of 8:2, and the initial recognition model is iteratively trained with the normalized sample dataset to adjust the parameters of the initial recognition model in real time to generate a glottis recognition model.
[0081] Preferably, in one embodiment, as Figure 5 shown, it also includes using tensorboard to visualize various index curves of the initial recognition model training, so that developers can intuitively see how various indexes change during the training process of the initial recognition model, and then optimize the initial recognition model.
[0082] The development tool for the glottis recognition model of the present invention is vscode, the development environment is Linux, the glottis recognition model is based on the Pytorch framework, and the GPU NVIDIA Geforce RTX 3090 is used for hardware acceleration.
[0083] The specific hyperparameter settings of the initial recognition model are as follows:
[0084] lr0 to lrf: 0.01 to 0.01
[0085] momentum: 0.937
[0086] weight_decay: 0.0005
[0087] warmup_epochs: 3.0
[0088] warmup_momentum: 0.8
[0089] warmup_bias_lr: 0.1
[0090] box: 0.05
[0091] cls: 0.5
[0092] cls_pw: 1.0
[0093] obj: 1.0
[0094] obj_pw: 1.0
[0095] iou_t: 0.2
[0096] anchor_t: 4.0
[0097] fl_gamma: 0.0
[0098] Among them, lr0 to lrf is the learning rate from the initial to the final; momentum is the momentum parameter, which is used to accelerate the process of gradient descent; weight_decay is the coefficient of L2 regularization, which is used to control the complexity of the initial recognition model; warmup_epochs, warmup_momentum, and warmup_bias_lr are the number of epochs, momentum parameter, and learning rate during warm-up respectively; box, cls, obj, box_pw, cls_pw, and obj_pw are the weight coefficients of different parts in the loss function; iou_t is the IoU threshold, which is used to determine the threshold of positive and negative samples; anchor_t is the threshold for calculating the coordinate offset of the target bounding box; fl_gamma is the parameter for adjusting the Focal Loss.
[0099] Meanwhile, for the video detection task of the present invention, the consideration of the relationship between the frames before and after in time sequence is increased. That is, if the input is the video image of the glottis, then the prediction of the current frame video image will combine the recognition results of the previous several frame video images with a certain weight to obtain the final relatively smooth recognition result. The advantage of this is that it can reduce the flickering of the detection box under the condition of unstable video image or interference, and improve the stability of recognition.
[0100] Preferably, in one embodiment, it further includes using five-fold cross-validation to test and evaluate the glottis recognition model.
[0101] Five-fold cross-validation is a commonly used machine learning model evaluation method for evaluating the performance and generalization ability of the model. As Figure 6 shown, the basic idea of its five-fold cross-validation is to divide the normalized sample data set into five parts. In turn, four of them are used as the training set, and the remaining one is used as the validation set. Then, the initial recognition model is trained and the performance of the glottis recognition model is evaluated on the validation set.
[0102] This process is repeated five times, each time selecting a different part of the sample data set as the validation set, and the other four parts as the training set. Finally, the average value of the performance evaluation indexes of the glottis recognition model obtained five times is used as the final performance evaluation index of the glottis recognition model.
[0103] The advantage of five-fold cross-validation is that it can make full use of the sample data set to evaluate the glottis recognition model. Especially in the case of a small data set, it can more accurately evaluate the performance and generalization ability of the model. At the same time, because the process of cross-validation is random, it can avoid introducing bias due to unreasonable division of the data set.
[0104] As shown in Table 1, the recall rate, precision rate, F1 value, mAP_50, and AP_50:90 value are calculated as the performance evaluation indexes of the glottis recognition model:
[0105] Table 1
[0106] recall precision F1 mAP_50 AP_50:90 0.95 0.84 0.90 0.90 0.48
[0107] The recall rate, precision rate, F1, mAP_50, and AP_50:90 values of the glottis recognition model all reach a relatively high accuracy, and the recall rate and precision rate are well balanced. The F1 value is the harmonic mean of the precision rate and the recall rate, comprehensively reflecting the accurate degree of recognition of the glottis recognition model. The mAP_50 and AP_50:90 values represent the overlap degree of the detection boxes under different confidence levels. Generally speaking, the overall recognition ability of the glottis recognition model is relatively high.
[0108] 5. Analyze and identify the acquired glottis image to be detected using the glottis recognition model to obtain the recognition result.
[0109] The present invention can correctly identify and classify the image to be detected, and display the probability that the image to be detected belongs to the glottis. Moreover, the glottis recognition model can correctly locate and classify the key feature positions in the image to be detected, such as the epiglottic cartilage, arytenoid cartilage, and glottis fissure positions in the glottis structure, and display the recognition result of the glottis direction on the display page by comprehensively weighting the positions of the epiglottic cartilage, arytenoid cartilage, and glottis fissure.
[0110] In one embodiment, a glottis recognition device based on deep learning is further provided, including an acquisition unit, a feature extraction unit, a generation unit, and an identification unit, where:
[0111] The acquisition unit is configured to acquire and randomly transform the original glottis image set, normalize the image color and size, and then transfer the data set to the feature extraction unit;
[0112] The feature extraction unit is configured to extract target features from the input glottis image set or video, obtain the recognition target features through a convolutional neural network, and then transfer them to the generation unit;
[0113] The generation unit is configured to predict and generate the categories, positions, and ranges of all target bounding boxes, calculate the probability that all bounding boxes belong to the target structure, and transfer them to the identification unit;
[0114] The identification unit is configured to integrate the recognition prediction results, eliminate the overlapping and redundant bounding boxes according to the probability that the bounding boxes belong to the target structure, display the best target range bounding box and prompt the probability of belonging to the target, and output the final recognition result after integration.
[0115] In one embodiment, a terminal is further provided, including a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute some or all of the steps of any one of the tracheal intubation assistance guidance methods based on deep learning described in the above method embodiments.
[0116] In one embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute some or all of the steps of any one of the tracheal intubation assistance guidance methods based on deep learning described in the above method embodiments.
[0117] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the specification of the present invention.
Claims
1. A deep learning-based endotracheal intubation auxiliary guidance method, characterized in that: include: Obtaining a set of original glottal images; Annotating the original glottis image set to obtain an annotated data set; Performing data enhancement processing on the glottis images in the labeled data set by random transformation processing to obtain a sample data set; The normalized sample data set is used to train the initial recognition model and generate a glottal recognition model; The glottis recognition model is used to identify and locate the acquired image to be detected to obtain an identification and positioning result.
2. The deep learning-based endotracheal intubation auxiliary guidance method according to claim 1, characterized in that: The annotated data set includes a glottis image set and annotation information, the annotation information includes a classification label and a bounding box, and the classification label includes glottis fissure, arytenoid cartilage, and epiglottic cartilage.
3. The deep learning-based endotracheal intubation auxiliary guidance method according to claim 1, characterized in that: The random transformation processing includes image hue shift, image saturation shift, image brightness shift, image rotation, image translation, image scaling, image shearing and image flipping.
4. The deep learning-based endotracheal intubation auxiliary guidance method according to claim 1, characterized in that: The method of using the normalized sample data set to train the initial recognition model to generate a glottis recognition model includes: A YOLOv5m model is built based on the PyTorch framework, and the YOLOv5m model is pre-trained to obtain an initial recognition model; The initial recognition model is iteratively trained using the normalized sample data set to generate a glottal recognition model.
5. The deep learning-based endotracheal intubation auxiliary guidance method according to claim 4, characterized in that: It also includes using five-fold cross validation to evaluate the performance of the glottis recognition model.
6. The deep learning-based endotracheal intubation auxiliary guidance method according to claim 5, characterized in that: The performance evaluation indicators of the glottis recognition model include recall rate, precision rate, F1 value, mAP_50 and AP_50:90 value.
7. A glottis recognition device based on deep learning, comprising an acquisition unit, a feature extraction unit, a generation unit and a recognition unit, wherein: An acquisition unit is used to acquire and randomly transform the original glottis image set, normalize the image color and size, and then pass the data set to the feature extraction unit; A feature extraction unit is used to extract target features from an input glottis image set or video, obtain the identified target features through a convolutional neural network, and then pass them to the generation unit; The generation unit is used to predict the category, location and range of all target bounding boxes, calculate the probability that all bounding boxes belong to the target structure, and pass it to the recognition unit; The recognition unit is used to integrate the recognition prediction results, eliminate overlapping and redundant bounding boxes according to the probability that the bounding box belongs to the target structure, display the best target range bounding box and prompt the probability of belonging to the target, and output the final recognition result after integration.
8. A terminal, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 6.