Method for identifying anatomical structure of visual laryngoscope image

By constructing the MPE-UNet model, using deep learning algorithms and multi-scale feature extraction modules and other technical means, the problem of difficulty in identifying laryngeal anatomical structures in visual laryngoscopy is solved, and the success rate of intubation and medical efficiency are improved.

CN119942293APending Publication Date: 2025-05-06THE NAVAL MEDICAL UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411696018.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify the laryngeal anatomy in visual laryngoscopy, resulting in a high rate of intubation failure and affecting patient safety and medical efficiency.

Method used

A deep learning algorithm is used to combine multi-scale feature extraction module, feature fusion module and enhanced channel attention mechanism to build an MPE-UNet model to identify anatomical structures in visual laryngoscopy images.

Benefits of technology

It significantly improves the accuracy and efficiency of laryngoscopy image segmentation, reduces the recognition time of low-age anesthesiologists, liberates the energy of high-age anesthesiologists, and improves the success rate of intubation and the quality of medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942293A_ABST
    Figure CN119942293A_ABST
Patent Text Reader

Abstract

The invention provides a method for identifying an anatomical structure of a visual laryngoscope image. The method comprises the following steps: acquiring a video, intercepting a picture, labeling the anatomical structure, generating a JSON file, and performing data enhancement; constructing a neural network model, adding a multi-scale feature extraction module, replacing traditional jump connection with PFAM, and designing an enhanced channel attention mechanism; dividing the data into a training set and a verification set, training a model by using a Pytorch framework, and performing evaluation by using an Adam optimizer and a cross entropy loss function; the trained model is used for segmentation, and the performance improvement and advantages of the model are verified through experiments. The medical image segmentation method not only has excellent performance in clinical application, but also has good expansibility, and can be applied to other medical image segmentation tasks. The innovative design of the model and a high-quality data set provide a solid foundation for subsequent research, and the development of a medical image processing technology is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image processing, and in particular to a method for identifying the anatomical structure of a visual laryngoscope image. Background Art

[0002] With the development of artificial intelligence technology, deep learning algorithms have far outperformed traditional methods in the fields of image classification, detection, and segmentation. Existing technology In video laryngoscopy, inexperienced doctors may encounter challenges in tracheal intubation due to unfamiliarity with the structure of the throat. In emergency medical situations, a successful first attempt at intubation is crucial because it provides patients with a valuable opportunity for treatment. Video laryngoscopy is becoming more and more common in operating rooms and intensive care units (ICUs), mainly used for tracheal intubation of patients undergoing general anesthesia and critically ill patients. Studies have shown that the use of video laryngoscopes results in a higher success rate for intubation on the first attempt than using conventional direct laryngoscopy. However, the effective use of video laryngoscopes depends on the professional level of the operator, and accurate identification of laryngeal structures is crucial to ensuring patient safety and improving the success rate of intubation.

[0003] The problem with existing technologies is that due to the complexity and individual differences of the laryngeal anatomy, it often takes a long time for novices to accurately identify various structures, and even experienced anesthesiologists may encounter challenges in identification. In this case, junior anesthesiologists may fail to intubate due to their inability to accurately identify laryngeal structures, increasing patient risks. In addition, senior anesthesiologists may be distracted by repetitive tasks when dealing with complex cases, affecting overall medical efficiency.

[0004] The consequences of these problems include: increased intubation failure rates, patients unable to receive timely airway management in emergency situations, delayed treatment, leading to serious health risks or even life-threatening conditions; senior anesthesiologists unable to fully concentrate on handling complex cases, and medical resources not being used effectively, affecting the quality of medical services.

[0005] With the rapid development of deep learning technology, the application of artificial intelligence in the medical field, especially in medical image analysis, is becoming more and more extensive, which helps to improve the efficiency and accuracy of medical care. Based on the visual laryngoscope, combined with deep learning algorithms, it can not only help junior anesthesiologists identify the anatomical structure of the larynx, but also free up some of the energy of senior anesthesiologists, allowing them to allocate their attention to more complex and important information in airway management. At the same time, it can also be used as a teaching aid for medical students, which has important clinical application prospects. Summary of the invention

[0006] To overcome the shortcomings of the existing technology, this paper proposes a method for identifying the anatomical structure of visual laryngoscope images, which not only performs well in clinical applications, but also has good scalability and can be applied to other medical image segmentation tasks. The innovative design of the model and the high-quality data set provide a solid foundation for subsequent research and help promote the development of medical image processing technology.

[0007] To achieve the above object, the present invention provides a method for identifying the anatomical structure of a visual laryngoscope image, comprising:

[0008] Step S1: Collect videos and capture pictures from the hospital, annotate the anatomical structures, generate JSON files, and perform data enhancement.

[0009] Step S2: Build a neural network model, add a multi-scale feature extraction module MSFE, use PFAM to replace the traditional skip connection, and design an enhanced channel attention mechanism EnhancedCCA.

[0010] Step S3: Divide the data into training set and validation set, use Pytorch framework to train the model, adopt Adam optimizer and cross entropy loss function for evaluation.

[0011] Step S4: Use the trained model for segmentation and verify the improvement and advantages of the model performance through experiments. a. Construction of the dataset:

[0012] Further, step S1 is specifically as follows:

[0013] We collected 210 video laryngoscopy videos from an authoritative comprehensive tertiary-level hospital in Shanghai, cut the videos into pictures frame by frame, and selected 1416 clear and identifiable pictures;

[0014] The labeled targets include: tongue, jaw, uvula, pharyngeal wall, epiglottis, supraglottic area, glottic fissure, vocal cords, and duct;

[0015] Professional anesthesiologists use Labelme software to annotate images, generate JSON format files, and finally convert them into PascalVOC dataset format;

[0016] The original image is rotated 180 degrees for data augmentation to improve the robustness of the model;

[0017] Step S2 is specifically as follows:

[0018] A multi-scale feature extraction module MSFE is added to the encoder part to enhance the feature extraction capability through the attention mechanism of multiple spatial scales;

[0019] Use PFAM to replace the skip connection of traditional U-Net to enhance the integration of feature expression and context information;

[0020] Design an enhanced channel attention mechanism EnhancedCCA to integrate local and global information in feature maps and provide rich feature representation for feature weighting;

[0021] Step S3 is as follows:

[0022] The data is divided into training set and validation set in an 8:2 ratio;

[0023] We developed the model using the Pytorch framework and conducted experiments on a system equipped with a 256G NVIDIA RTX A6000 graphics card.

[0024] Set the batch size to 8, use the Adam optimizer, the initial learning rate to 1e-4, the loss function to cross entropy loss, and the total training cycle to 150 cycles;

[0025] Use IoU, DSC, precision, recall and accuracy as evaluation indicators and take the average value for evaluation;

[0026] Step S4 is specifically as follows:

[0027] Use the trained model to segment the test set and mark each structure with color;

[0028] The model performance was verified through ablation experiments and multi-model comparison experiments, confirming the improvements and advantages of MPE-UNet in various indicators.

[0029] Furthermore, the visual laryngoscope data in step S1 is obtained from the clinical anesthesia events of a tertiary hospital in Shanghai. The laryngoscope images of 210 patients are collected and sorted to construct a laryngeal structure dataset, which contains about 4,000 laryngeal structure photos taken by the visual laryngoscope during tracheal intubation. First, these original data images need to be renumbered to protect the privacy of patients. Then, clear and annotated laryngoscope images need to be selected, totaling 1,416 images.

[0030] Furthermore, the images were annotated by three anesthesiologists with many years of intubation experience according to the annotation standards for laryngeal structures, using labelme as the annotation tool.

[0031] Furthermore, the network model is as follows:

[0032] Add a multi-scale feature extraction module MSFE to the encoder part to enhance important features while suppressing minor information;

[0033] Using PFAM to replace the skip connection part of the traditional U-Net is designed to enhance feature expression and promote the integration of contextual information;

[0034] The enhanced channel attention mechanism EnhancedCCA module is designed to fuse the features passed down from the PFAM module and the features upsampled from the encoder. It can integrate the local and global information in the feature map and provide richer feature representation for subsequent feature weighting.

[0035] Furthermore, the segmentation and recognition module verifies the performance of the proposed MPE-UNet segmentation method on a custom dataset and analyzes the experimental results. The specific experimental process includes the following steps:

[0036] Use the improved MPE-UNet algorithm to train the custom dataset and evaluate the performance of the model.

[0037] The experimental results are compared and analyzed to confirm the improvements and advantages of the model in various indicators.

[0038] Furthermore, the enhanced channel attention mechanism (EnhancedCCA) includes a double pooling strategy to capture a wider range of feature statistics through maximum pooling and average pooling, which is then processed using two independent MLPs, and finally a sigmoid function is applied to obtain the normalized attention weights.

[0039] Furthermore, the Multi-Scale Feature Extraction module (MSFE) combines ReLU activation functions by applying point convolution, standard convolution, and dilated convolution layers at multiple spatial scales, and integrates features extracted across different scales and contexts through the “VoteConv” block.

[0040] Furthermore, the feature fusion module (PFAM) includes a downsampling block, a down-fusion block, an upsampling block and an up-fusion block, which fuses multi-level features through batch normalization and ECA mechanism to improve the accuracy of feature expression.

[0041] Furthermore, the model training uses the Adam optimizer, the initial learning rate is 1e-4, the loss function is the cross entropy loss, the total training cycle is 150 cycles, and hyperparameter tuning is used during the training process to improve the model performance.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] 1. The present invention provides a method for identifying the anatomical structure of a visual laryngoscope image. The MPE-UNet model has made many structural innovations, which significantly improves the accuracy and efficiency of laryngoscope image segmentation:

[0044] Multi-scale Feature Extraction Module (MSFE): Enhances feature extraction capability through a multi-scale attention mechanism and effectively captures contextual information in complex images.

[0045] Feature Fusion Module (PFAM): Replaces the skip connection of traditional U-Net, enhances feature expression and promotes the integration of contextual information.

[0046] Enhanced Channel Attention Mechanism (EnhancedCCA): Integrates local and global information to provide richer feature representation and improve segmentation performance.

[0047] 2. The present invention provides a method for identifying the anatomical structure of visual laryngoscope images, using the Pytorch framework to train the model on a high-performance computing system, using the Adam optimizer and the cross-entropy loss function to ensure the efficiency and stability of the model training. The reliability and practicality of the model are ensured through a rigorous training and evaluation process, including 8:2 data splitting, 150 training cycles, and the use of multiple evaluation indicators such as IoU, DSC, precision, recall rate, and accuracy.

[0048] 3. The present invention provides a method for identifying the anatomical structure of a visual laryngoscope image. The trained model is applied to the test set for segmentation, and each structure is marked by color to intuitively display the segmentation results. Through ablation experiments and multi-model comparison experiments, the significant improvement of the MPE-UNet model in various indicators is verified, proving its superior performance in segmenting the anatomical structure of laryngoscope images.

[0049] 4. The present invention provides a method for identifying the anatomical structure of visual laryngoscope images, which not only performs well in clinical applications, but also has good scalability and can be applied to other medical image segmentation tasks. The innovative design of the model and the high-quality data set provide a solid foundation for subsequent research and help promote the development of medical image processing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 It is a schematic diagram of the process of the present invention

[0052] Figure 2 It is a schematic diagram of the steps of the present invention

[0053] Figure 3 This is the MSFE module structure diagram of the present invention

[0054] Figure 4 This is the structural diagram of the PFAM module of the present invention

[0055] Figure 5 This is the structure diagram of the ECA attention mechanism of the present invention

[0056] Figure 6 This is the structure diagram of the EnhancedCCA module of the present invention

[0057] Figure 7 This is a visualization diagram of the ablation experiment results of the present invention.

[0058] Figure 8 Visualization of model comparison results DETAILED DESCRIPTION

[0059] The technical solution of the present invention will be more clearly and completely explained below through description of preferred embodiments of the present invention in combination with the accompanying drawings.

[0060] Terminology explanation:

[0061] MPE-UNet: An improved U-Net model for medical image segmentation. MPE-UNet combines multi-scale feature extraction, feature fusion, and enhanced channel attention mechanism to improve segmentation accuracy and efficiency.

[0062] U-Net: A convolutional neural network architecture commonly used for biomedical image segmentation, with a symmetric encoder and decoder structure.

[0063] Labelme: An open source image annotation tool that is often used to manually annotate target areas in images and generate annotation data.

[0064] JSON (JavaScript Object Notation): A lightweight data exchange format that is easy for humans to read and write, and easy for machines to parse and generate.

[0065] Pascal VOC (PASCAL Visual Object Classes): A commonly used image dataset format that contains annotation information for tasks such as image classification, object detection, and object segmentation.

[0066] MSFE (Multi-Scale Feature Extraction): Multi-scale feature extraction module, which enhances feature extraction capabilities by applying attention mechanisms at multiple spatial scales.

[0067] PFAM (Pyramid Fusion Attention Module): Pyramid fusion attention module replaces the jump connection part of traditional U-Net, enhances feature expression and promotes the integration of contextual information.

[0068] EnhancedCCA (Enhanced Channel-wise Contextual Attention): Enhanced channel attention mechanism, which combines global and local information through maximum pooling and average pooling to improve the accuracy of feature weighting.

[0069] IoU (Intersection over Union): An indicator for evaluating the overlap between the model prediction result and the actual target area. The calculation formula is the ratio of the intersection area of ​​the predicted area and the true area to the union area.

[0070] DSC (Dice Similarity Coefficient): Dice similarity coefficient, an indicator for evaluating the similarity between the model's predicted segmentation area and the actual segmentation area. The calculation formula is twice the intersection area divided by the sum of the predicted area and the actual area.

[0071] Pytorch: An open source deep learning framework widely used in fields such as computer vision and natural language processing.

[0072] Adam (Adaptive Moment Estimation): An optimization algorithm that combines the advantages of momentum and RMSProp and is often used to train deep learning models.

[0073] Cross-Entropy Loss: A commonly used loss function that measures the difference between the predicted distribution and the true distribution and is widely used in classification problems.

[0074] Batch Size: In deep learning, the number of samples used to train the model in one iteration.

[0075] Clinical anesthesia events: all operations and phenomena occurring during the anesthesia process, including intubation, maintenance of anesthesia, extubation, etc.

[0076] Deep Learning: A branch of machine learning that uses neural network models to learn features and patterns from large amounts of data. It is widely used in image recognition, natural language processing and other fields.

[0077] Convolutional Neural Network (CNN): A deep learning model that is particularly suitable for processing image data. It extracts features through convolutional layers and is widely used in tasks such as image classification, detection, and segmentation.

[0078] Ablation Study: An experimental method that removes or replaces some components of a model to evaluate its impact on the overall performance.

[0079] Feature weighting: In deep learning, the process of weighting the features in the feature map to highlight important features and suppress minor features.

[0080] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments:

[0081] like Figure 1 As shown:

[0082] Step S1: Collect videos and capture pictures, annotate anatomical structures, generate JSON files, and perform data enhancement;

[0083] Step S2: Build a neural network model, add a multi-scale feature extraction module, use PFAM to replace the traditional skip connection, and design an enhanced channel attention mechanism EnhancedCCA;

[0084] Step S3: Divide the data into training set and validation set, use Pytorch framework to train the model, and use Adam optimizer and cross entropy loss function for evaluation;

[0085] Step S4: Use the trained model to perform segmentation and verify the improvement and advantages of the model performance through experiments.

[0086] 1. Dataset construction: Collect data, mark each category of the anatomical structure of the larynx by drawing contour lines, and finally construct the dataset through data enhancement.

[0087] As a specific implementation method, data collection method: Due to the lack of public data sets related to visual laryngoscopes, this patent collects a large amount of clinical visual laryngoscope data. The laryngoscope images used in this study are from an authoritative comprehensive tertiary-level Class A hospital in Shanghai. 210 visual laryngoscope videos were collected and sorted by senior clinical physicians of the hospital. We cut these videos into pictures by frame and selected 1416 pictures with clear image quality and identifiable. After research and discussion, it was determined that the following targets to be segmented with clinical value will be labeled in the future: tongue, jaw, uvula, pharyngeal wall, epiglottic cartilage, supraglottic area, glottic fissure, vocal cords, and catheter. Professional anesthesiologists will also label the areas to be segmented. This dataset contains the laryngeal structure information displayed by the visual laryngoscope for tracheal intubation.

[0088] Labeling process: After obtaining the laryngoscope image data, labelme software is used to label the image. In labelme software, the anesthesiologist will use the polygon tool to click the structure boundary points on the image according to the structure of the larynx. After the labeling is completed, the label name can be entered. The generated format is json file format, and finally converted into Pascal voc data set format and input into the program.

[0089] Data enhancement method: Since collecting data to obtain high-quality data is a tedious and difficult task, data enhancement of the collected data helps to improve the robustness of the trained model in the future. This patent uses a rotation method to enhance the data, and the main operation is to rotate the original image 180 degrees.

[0090] 2. Construction of neural network model: We innovatively changed the original U-Net model structure to improve the accuracy and efficiency of laryngeal image segmentation, and named it MPE-UNet. The model structure is shown in the figure Figure 2 As shown in the figure, the specific improvement measures are as follows: (1) Add a multi-scale feature extraction module MSFE to the encoder part to strengthen important features while suppressing secondary information. (2) Use the PFAM module to replace the traditional U-Net jump connection part to enhance feature expression and promote the integration of contextual information. (3) Design an enhanced channel attention mechanism EnhancedCCA module, which can integrate local and global information in the feature map and provide richer feature representation for subsequent feature weighting.

[0091] As a specific implementation method

[0092] (1) Multi-scale feature extraction module:

[0093] The complexity of laryngeal structure images lies in the fact that a single image contains information of multiple laryngeal structures. To solve this problem, we propose an innovative module called Multiscale Feature Extraction (MSFE), which aims to enhance feature extraction capabilities through effective feature fusion and spatial attention mechanisms. The U-Net framework has a symmetrical structure of encoder and decoder, and the encoder consists of multiple downsampling models, usually stacked convolution and pooling layers. However, standard convolution operations often have difficulty in effectively capturing spatial context and multi-scale information. Therefore, we introduced the MSFE module (see Appendix). Figure 3), especially applied to the deep part of the encoder after each downsampling step. The MSFE module enhances the ability to capture contextual information by applying attention mechanisms at multiple spatial scales, including point-level convolutions, standard convolutions, and dilated convolution layers with dilations of 2, 4, and 8. The outputs of these branches are then concatenated along the channel dimension and activated by a ReLU activation function to integrate features extracted across different scales and contexts. In addition, these combined features at different scales pass through a "VoteConv" block and play the role of spy attention. The VoteConv block includes a 1x1 convolution layer, batch normalization, and a sigmoid function. The function of this block is to fuse the outputs from the first five convolutional layers and compress them back to the original number of channels. We used a 1x1 convolution kernel to promote information integration across channels without changing the spatial dimension. After passing the "VoteConv" convolution, we multiply the resulting feature map with the original feature map for attention allocation. Finally, the output features are combined with the input feature map through residual connection, which alleviates the problem of gradient disappearance, ensures the continuity of features, and improves the stability of the model, which can be expressed as:

[0094] W=x+x*voteConv[relu(cat(x1,x2,x3,x4,x5))]

[0095] Among them, x1-x5 represent different convolution operations, cat represents feature concatenation, and x represents the feature map of the original input.

[0096] (2) Feature Fusion Module

[0097] The traditional U-Net architecture combines high-resolution and low-resolution feature maps simply through skip connections to preserve detail information. However, we find that this approach is limited in fully utilizing features at different levels and processing multi-scale contextual information. Based on this, we design the Pyramid Fusion Attention Module (PFAM), as shown in the attached figure. Figure 4 As shown. The model mainly consists of four parts: downsampling block, down-fusion block, upsampling block and up-fusion block. Taking into account the multi-level characteristics and complementarity of features. In the PFAM module, the input feature maps from different levels of the encoder are first batch normalized to stabilize the learning process and improve efficiency. Subsequently, the feature maps enter the downsampling and down-fusion stages, where low-level detail features are gradually integrated into high-level features to capture a wider range of contexts. Specifically, first, the downsampling block and the down-fusion block are introduced. Initially, fp i=1 Processed through a downsampling block, denoted as fp' i=1 . Then through a downward fusion block with fpi=2 Fusion, get fp d=2 . Then, fp d=2 After a downsampling block, denoted as fp' i=2 , and from fp' i=2 The features of fp are fused together using a downward fusion block i=3 Fusion, producing fp d=3 Next, fp d=3 Processed through a downsampling block, labeled fp' i=3 , fp' i=3 The resulting features are combined with fp through a downward fusion block i=4 Fusion, get fp d=4 . This concludes the downsampling and downfusion stages. The downsampling block contains a sequence of operations where the feature maps first pass through a custom downsampling process called “downsamp”. This process reduces the spatial resolution (height and width) of the 2D feature maps by rearranging the spatial positions of the input tensor into the channel dimension, thereby retaining more spatial information at the channel level compared to traditional downsampling methods such as max pooling or average pooling. After “downsamp”, the feature maps proceed through convolution, batch normalization, and ReLU activation functions. Subsequently, these processed feature maps are fused with the feature maps of the next layer through a fusion module, which first concatenates them, then performs a 3x3 convolution, and then a residual connection. After the residual connection, an ECA (Effective Channel Attention) mechanism (as shown in the attached) is inserted. Figure 5 As shown). The ECA attention mechanism consists of global average pooling (GAP), 3x3 convolution, and sigmoid function. Initially, ECA performs global average pooling on the input feature map along the spatial dimension to reduce spatial complexity and pays attention to the weight of each channel. Subsequently, ECA uses one-dimensional convolution (k=3) to process channel information. Finally, the output of the convolution is processed using the sigmoid function. The attention weights generated by sigmoid are multiplied by the channels of the original input feature map to adjust the weights of each channel. It can be expressed as follows:

[0098] fp d =ECA[C(0.75*fp i +0.25*fp' i-1 )+fp i ]

[0099] After passing through the downsampling fusion module, the feature map is sent to the upsampling and up-fusion stage. d=4 is directly output, and for naming consistency, it is designed to be fp u=4 In addition, fp d=4 The feature after passing through an upsampling block is annotated as fp'd=4 , then with fp d=3 Fused in an upper fusion block, we get fp u=3 . fp upsamples another block after fp u=3 The characteristic is fp' d=3 , and fp d=2 After fusion, fp u=2 Finally, fp u=2 After an upsampling block, labeled fp' d=2 , fp' d=2 The characteristics of fp d=1 Fused in an upper fusion block, producing fp u=1 . This is the structural description of our PFAM. The structure of the upsampling block is similar to that of the downsampling block. The feature map first passes through a custom upsampling process called "upsamp", which achieves upsampling by rearranging the channels of the input tensor instead of using traditional interpolation methods such as bilinear or nearest neighbor interpolation. After output from upsamp, the feature undergoes convolution, batch normalization, and ReLU activation function. Subsequently, these processed features are fused with the feature map of the previous layer through an up-fusion module. In the up-fusion module, the two layers are first element-wise summed, followed by a 3x3 convolution, and then a residual connection is performed. Similarly, in the up-fusion block, we also incorporate an ECA mechanism. In the up-fusion stage, the ECA module is applied to the feature map after the residual output to dynamically adjust the channel response. This mechanism enables the model to focus on the key features related to the current task, facilitating precise and targeted adjustments, thereby improving the performance of specific tasks by refining the feature representation:

[0100] fp u =ECA[C(0.75*fp d-1 +0.25*fp' d )+fp d-1 ]

[0101] (3) Enhanced channel attention module

[0102] In the traditional U-Net design, the feature maps generated by upsampling in each layer of the decoder are directly connected to the corresponding feature maps of the encoder. This direct connection often fails to fully integrate all features, thus affecting the segmentation performance of the model. To address this problem, this study developed an innovative attention mechanism module called Enhanced Channel-wise Contextual Attention (ECCA). Figure 6As shown in the figure, the features (x) fused by PFAM and the features (g) upsampled by the decoder are both passed through maximum pooling and average pooling. This dual pooling method allows the module to capture a wider range of feature statistics, where the maximum pooling highlights the prominent features in the feature map, while the average pooling provides a global overview of the information. Through this pooling process, the ECCA module integrates local information and global information in the feature map, providing a richer feature representation for subsequent feature weighting. The ECCA module then processes these aggregated feature statistics through two independent multi-layer perceptions (MLPs), which are MLPx and MLPg. These two MLP modules achieve feature transformation and nonlinear enhancement through a combination of fully connected layers and ReLU activation functions. The ECCA module is able to accurately adjust the uniqueness of each feature set in order to effectively emphasize or suppress specific channels, because each MLP learns and modifies the weights of its related feature set separately, which is obtained from maximum pooling and average pooling. Finally, the outputs of the two MLP modules, channel_x and channel_g, are used to calculate the channel attention weights through the sigmoid activation function. These weights are used to scale the channels of the original feature maps, allowing the network to effectively learn and emphasize the most critical features when processing multi-channel feature data.

[0103] 3. Model training and evaluation

[0104] 3.1 Model Training

[0105] We divided the collected data into training and validation sets in an 8:2 ratio. MPE-UNet was developed using the Pytorch framework. All experiments were conducted on a system equipped with a 256G NVIDIA RTX A6000 graphics card running a 64-bit Ubuntu 11.2.0 operating system. Our experiments did not use any pre-trained models. The batch size was set to 8, and the Adam optimizer was used with an initial learning rate of 1e-4. The loss function was calculated using the cross entropy loss, defined as:

[0106]

[0107] Yi is a binary indicator of whether the sample belongs to class i (1 for yes, 0 for no), and Pi is the probability that the model predicts that the sample belongs to class i.

[0108] The model was trained for a total of 150 epochs. This was based on the observation that the loss function starts to converge around 150 epochs, so the training epochs were normalized to 150 epochs.

[0109] 3.2 Evaluation indicators

[0110] The evaluation metrics of the model include intersection over union (IoU), Dice Similarity Coefficient (DSC), precision, recall, and accuracy. In this patent, we take the average of these metrics, that is, the total metric of each class divided by the sum of the number of samples contained in the class.

[0111] ①IoU is a key metric to measure the overlap between the model's predicted area and the actual target area, defined as the difference between the predicted area and the true area.

[0112] The ratio of product union is given by:

[0113]

[0114] ②DSC measures the similarity between the segmented regions predicted by the model and the actual segmented regions, and its formula is:

[0115]

[0116] ③Precision is defined as the ratio of true positive samples to the total number of samples identified as positive by the model, and its formula is:

[0117]

[0118] ④Recall rate, also known as sensitivity, is the ratio of positive samples identified by the model to the total number of actual positive samples. The formula is:

[0119]

[0120] ⑤Accuracy measures the overall correctness of the model in the predicted samples, which is defined as the ratio of correctly predicted pixels to the total number of pixels. The formula is:

[0121]

[0122] In these formulas, True Positives (TP) refers to the number of samples correctly predicted as positive and positive, False Positives (FP) refers to the number of samples incorrectly predicted as positive, False Negatives (FN) refers to the number of samples incorrectly predicted as negative that are actually positive. True Negatives (TN) refers to the number of samples correctly predicted as negative.

[0123] Through hyperparameter tuning, we find the optimal hyperparameter combination. We also collect some data as a test set and perform a final evaluation on the test set to verify the generalization ability of the model.

[0124] 4. Segmentation and recognition of laryngoscope images

[0125] 4.1 Segmentation step: We use the trained model to test on the test set, and the segmentation results are as follows:

[0126] We marked the colors as: tongue (red), hard palate (blue), uvula (yellow), pharyngeal wall (green), epiglottic cartilage (purple), glottis (grey), glottic fissure (cyan), vocal cords (brown), duct (orange);

[0127] 4.2 Result Verification

[0128] We conducted ablation experiments and multi-model comparison experiments on the model. The experimental results are shown in Tables 1 and 2. As can be seen from Table 1, the three modules we added all showed good results in the segmentation of laryngeal structures. As can be seen from Table 2, compared with the baseline model (U-Net), the model has improved by 10.09%, 11.45%, 7.33%, 2.93% and 10.26% in miou, mDSC, mprecision, mrecall and maccuracy, respectively. The mIoU of MPE-UNet is 0.7426, which is 1.35% higher than the second-ranked U2net, and its mDSC is 1.26% higher than U2net. Its precision is 1.78% higher than U2net, and its accuracy is 2.43% higher than U2net. These improvements clearly prove that our model has superior segmentation capabilities. Figure 7 and Figure 8 The visualization results of the ablation experiment and the comparison experiment are shown respectively. Compared with other networks, MPE-UNet has better segmentation effects on the tongue, cartilaginous epiglottis, glottis, and duct, with a high degree of similarity to the ground truth images. In addition, all networks show relatively poor segmentation results for the uvula and vocal cords. This may be due to the similar shapes of the uvula and hard palate without a clear dividing line, the variable shapes of the vocal cords, and the small target area, which increases the difficulty of recognition.

[0129] Table 1 Ablation experiment results

[0130]

[0131] Table 2 Model comparison results

[0132]

[0133] The experimental results show that the MPE-UNet designed by this patent effectively improves the segmentation accuracy and precision of laryngeal images. In terms of model accuracy, mIoU is increased from the original 64.17% to 74.26%; mDSC is increased from the original 77.23% to 84.56%; mprecision is increased from the original 73.11% to 84.56%; mrecall is increased from the original 82.14% to 85.07%. Maccuracy is increased from the original 78.39% to 88.65%. Compared with the current mainstream segmentation network, the segmentation effect of the model is also better, which has high clinical practical significance.

[0134] The present invention proposes an algorithm for laryngeal image segmentation and recognition, which accurately assists the operation during tracheal intubation. At the same time, it introduces artificial intelligence models into the field of anesthesia and explores the solution to clinical anesthesia problems, which is of great significance.

[0135] The above specific implementations are only descriptions of the preferred implementations of the present invention, and do not limit the protection scope of the present invention. Without departing from the design concept and spirit of the present invention, various modifications, substitutions and improvements made by ordinary technicians in this field to the technical solution of the present invention based on the text description and drawings provided by the present invention should all fall within the protection scope of the present invention. The protection scope of the present invention is determined by the claims.

Claims

1. A method for identifying anatomical structures in a visual laryngoscope image, characterized in that: include: Step S1: Collect videos and capture pictures, annotate anatomical structures, generate JSON files, and perform data enhancement; Step S2: Build a neural network model, add a multi-scale feature extraction module, use PFAM to replace the traditional skip connection, and design an enhanced channel attention mechanism EnhancedCCA; Step S3: Divide the data into training set and validation set, use Pytorch framework to train the model, and use Adam optimizer and cross entropy loss function for evaluation; Step S4: Use the trained model to perform segmentation and verify the improvement and advantages of the model performance through experiments.

2. A method for identifying anatomical structures in visual laryngoscope images according to claim 1, characterized in that: Step S1 is specifically as follows: 210 video laryngoscopy videos were collected, the videos were frame-by-frame captured into images, and 1416 clear and identifiable images were selected; The labeled targets include: tongue, jaw, uvula, pharyngeal wall, epiglottis, supraglottic area, glottic fissure, vocal cords, and duct; Use Labelme software to annotate images, generate JSON format files, and finally convert them into PascalVOC dataset format; The original image is rotated 180 degrees for data augmentation to improve the robustness of the model; Step S2 is specifically as follows: A multi-scale feature extraction module MSFE is added to the encoder part to enhance the feature extraction capability through the attention mechanism of multiple spatial scales; Use PFAM to replace the skip connection of traditional U-Net to enhance the integration of feature expression and context information; Design an enhanced channel attention mechanism EnhancedCCA to integrate local and global information in feature maps and provide rich feature representation for feature weighting; Step S3 is as follows: The data is divided into training set and validation set in an 8:2 ratio; Use the Pytorch framework to develop models and conduct experiments on the system; Set the batch size to 8, use the Adam optimizer, the initial learning rate to 1e-4, the loss function to cross entropy loss, and the total training cycle to 150 cycles; Use IoU, DSC, precision, recall and accuracy as evaluation indicators and take the average value for evaluation; Step S4 is specifically as follows: Use the trained model to segment the test set and mark each structure with color; The model performance was verified through ablation experiments and multi-model comparison experiments, confirming the improvements and advantages of MPE-UNet in various indicators.

3. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: In step S1, laryngoscopic images of 210 patients were collected and organized to construct a laryngeal structure dataset, which contains about 4,000 photos of laryngeal structure taken by visual laryngoscope during tracheal intubation. First, these original data images need to be renumbered to protect the patient's privacy information; then, clear and annotated laryngoscope images are selected, totaling 1,416 images.

4. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: In step S1, the annotated images are annotated by three anesthesiologists according to the annotation standard of laryngeal structure, and the annotation tool is labelme.

5. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The neural network model is constructed in step S2 as follows: Add a multi-scale feature extraction module MSFE to the encoder part to enhance important features while suppressing minor information; Using PFAM to replace the skip connection part of the traditional U-Net is designed to enhance feature expression and promote the integration of contextual information; The enhanced channel attention mechanism EnhancedCCA module is designed to fuse the features passed down from the PFAM module and the features upsampled from the encoder, integrating the local and global information in the feature map to provide richer feature representation for subsequent feature weighting.

6. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The specific experimental process of segmentation and recognition in step S4 includes the following steps: Use the improved MPE-UNet algorithm to train the custom dataset and evaluate the performance of the model; The experimental results are compared and analyzed to confirm the improvements and advantages of the model in various indicators.

7. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The enhanced channel attention mechanism (EnhancedCCA) includes a dual pooling strategy to capture a wider range of feature statistics through maximum pooling and average pooling, which is then processed using two independent MLPs, and finally a sigmoid function is applied to obtain normalized attention weights.

8. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The Multi-Scale Feature Extraction module (MSFE) works by applying point convolution, standard convolution, and dilated convolution layers at multiple spatial scales, combined with ReLU activation functions, and integrating features extracted across different scales and contexts through the "VoteConv" block.

9. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The feature fusion module includes downsampling block, down-fusion block, upsampling block and up-fusion block. It fuses multi-level features through batch normalization and ECA mechanism to improve the accuracy of feature expression.

10. The method for identifying anatomical structures in a visual laryngoscope image according to claim 1, characterized in that: The model training uses the Adam optimizer, the initial learning rate is 1e-4, the loss function is the cross entropy loss, the total training cycle is 150 cycles, and hyperparameter tuning is used during the training process to improve the model performance.