System and method of difficult-airways assessment via an artificial intelligence scanner
The 3D facial scanner with AI-based multi-stage assessment addresses the limitations of existing airway prediction methods, providing accurate and reliable difficult airway prediction through telemedicine, enhancing surgical safety and resource efficiency.
Patent Information
- Application Number
- PCT/SG2025/050194
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-25
AI Technical Summary
Existing airway assessment methods lack the ability to accurately predict difficult airways, fail to classify mask ventilation, supraglottic insertion, and intubation, and do not utilize intermediate clinical features, leading to potential ventilation failures and anesthesia-related injuries.
A system utilizing a 3D facial scanner and artificial intelligence to generate multi-stage assessments, employing a Global-to-local Cross Pseudo Supervision method and explainable AI engine for predicting difficult airway severity, incorporating telemedicine for remote assessment.
Enables accurate and reliable prediction of difficult airways, optimizing healthcare resource utilization and improving surgical outcomes by reducing the need for physical presence during pre-operative assessments.
Smart Images

Figure SG2025050194_25092025_PF_FP_ABST
Abstract
Description
System And Method Of Difficult-Airways Assessment Via AnArtificial Intelligence ScannerRelated Applications
[0001] The present invention claims priority to Singapore patent application no. 10202400774X filed on 19 March 2024, the disclosure of which is incorporated in its entirety.Field of In ention
[0002] This invention is in the field of airway assessment, and discloses a system and an architectural structure supporting the system by combining artificial intelligence and 3D facial scanning. Tn particular, the invention relates to an intelligent Medical Cyber-Physical System (MCPS) framework for implementing an Artificial Intelligence Enabled 3D Facial Scanner for Airway Assessment system (AINFAS); tire MCPS framework is suitable for telemedicine.Background
[0003] Airway assessment is routinely performed during the pre-operative phase to identify potential difficult airways (DA). DA can lead to difficult intubation, resulting in ventilation failures and anesthesia-related injuries and deaths. Furthermore, DA can appear in up to 20% of patients during surgical procedure under general anesthesia. Among them, 93% are unanticipated and 46% occur in elective cases. Delayed identification of potential DA can lead to poor planning of airway management and potentially dangerous situations.
[0004] There have been prior art methods related to predicting potential difficult intubation of a subject, using a facial structure analysis system and facial structure data of the subject. However, there exists a need for another method and system for airway assessment that can provide reliable predictions of difficult airways (DA) and can overcome at least the following short-comings of the prior art methods:• having limited airway classes;• no classes for mask ventilation, supraglottic insertion and intubation;• no classes for overall airway difficulty;• does not predict intermediate clinical features from images / videos; and• inputs are intermediate clinical features.
[0005] The AINFAS invention thus holds the potential to revolutionize airway assessment and management, potentially saving lives and improving surgical outcomes, while optimizing the utilization of healthcare resources.Summary
[0006] The following presents a simplified summary to provide a basic understanding of the present invention. This summary is not an extensive overview of the present invention, and is not intended to identify key features of the invention. Rather, it is to present some of the inventive concepts of this invention in a generalised form as a prelude to the detailed description that is to follow.
[0007] Research has been conducted to validate the hypothesis that an artificial intelligence- based face recognition model will accurately identify patients with a potentially difficult airway. The invention aims to identify the parameters indicative of a difficult airway (DA) and develop technologies for predicting such cases. The AINFAS invention holds the potential to revolutionize airway assessment and management, potentially saving lives and improving surgical outcomes, while optimizing the utilization of healthcare resources.
[0008] Another use of the invention is in the closely related domain of sleep apnea, where it could be used for automated assessment of sleep apnea patients and automated or custom mask sizing for patients.
[0009] As will be described, the AINFAS encompasses a suite of intelligence components, cyber components, cyber-physical interfaces, and physical components, supporting the application of the AINFAS invention during prc-opcration, intra-opcration and post operation phases. Cyber and physical components are brought together to create the distribution of health-related sendees and information via electronic information and telecommunication technologies (i.e. telemedicine) to support the AINFAS invention; an explainable artificial intelligence (XAI) engine makes up the core intelligence component to augment the clinician's medical decision making. The AINFAS invention is enabled remotely for pre-operative airway assessment without the physical presence of the patient in the clinic.
[0010] Tn one embodiment, the present invention relates to a difficult-airways (DA) assessment method, compnses: generating three-dimensional scans from multiple images of a face and neck of a patient; analysing the 3D scans using a deep-learning algorithm to extract a plurality of landmarks; constructing a difficult-airway (DA) classifier for multi-stage assessment, and using an artificial intelligence model to predict the difficult-airways (DA) severity based on the multi-stage assessment.
[0011] Preferably, the multi-stage assessment comprises an intermediate clinical outcome and a final clinical outcome .
[0012] Preferably, the artificial intelligence model employs an explainable artificial intelligence engine using key clinical criteria.
[0013] Tire present invention also relates to a system for assessing difficult-airways comprising: a three-dimensional facial scanner to capture multiple images of the face and neck of a patient; a deep learning processor for extracting a plurality of landmarks in the images of the face and neck; a classifier for sorting the landmarks for multi-stage assessment; and an artificial intelligence processor to predict the difficulty of difficult-airways severitybased on the multi-stage assessment.
[0014] Preferably, the classifier comprises an intermediate clinical feature classifier and a clinical outcome classifier.
[0015] In another embodiment, the AINFAS uses a Global-to-local Cross Pseudo Supervision (G2LCPS) which is a novel end-to-end semi-supervised method for facial landmark prediction. The method employs the exponential moving average (EMA) teacher network which is common in semi-supervised learning. The novelty of G2LCPS lies in introducing an additional classification head that classifies good unlabeled images at a global level before being threshold at a local level to generate high-quality pseudo heatmaps. The use of Stacked Hourglass as the backbone reduces spatial information losses during interaction between the teacher and student networks.Brief Description of the Drawings
[0016] This invention will be described by way of non-limiting embodiments of the present invention, with reference to the accompanying drawings, in which:
[0017] FIG. 1 shows an Artificial Intelligence Enabled 3D Facial Scanner for Airway Assessment system (A1NFAS) according to an embodiment of the present invention.
[0018] FIG. 2 illustrates an intelligent Medical Cyber-Physical System (MCPS) for implementing the above AINFAS.
[0019] FIG. 3 illustrates various screenshots of the AINFAS app to guide a user in taking photographs and videos when using the MCPS.
[0020] FIGs. 4A-4B illustrate some examples of labeled images from an Annotated Facial Landmarks in the Wild (ALFW)-DA dataset that distinguish no-neck images or images with neck landmarks.
[0021] FIGs. 5A-5B depict graphs to compare the G2LCPS method according to an embodiment of the invention with that of prior art methods.
[0022] FIGs. 6A-6B illustrate the qualitative results of the G2LCPS method showing comparative upper body images, ground truth heat map and landmarks, and prediction heat map and landmarks produced by tire G2LCPS method.
[0023] FIG. 7 is a table showing normalized mean error (NME) results of an ablation study for the effectiveness of EMA, Global filter, and Local filter.
[0024] FIGs. 8A-8C illustrate facial landmarks predicted by a Practical Facial Landmark Detector (PFLD) model according to an embodiment of the present invention.
[0025] FIGs. 9A-9J illustrate facial landmarks predicted by a convolution network according to an embodiment of the present invention.Detailed Description
[0026] One or more specific and alternative embodiments of the present invention will now be described with reference to the attached drawings. It shall be apparent to one skilled in the art, however, that this invention may be practised without such specific details. Some of the details may not be described at length so as not to obscure the invention.
[0027] FIG. 1 shows an Artificial Intelligence Enabled 3D-Facial Scanner for Airway Assessment system (AINFAS) 100 according to an embodiment of the invention. The AINFAS 100 has four modules, namely, an upper body feature predictor 101, a 3D-model reconstructor 102, an intermediate clinical feature classifier 103, and a clinical outcome classifier 104.
[0028] Using data from open-source datasets (such as public dataset and clinical patient data) as input, the upper body feature predictor 101 trams the data and outputs upper body landmarks / segments. The 3D model reconstructor 102 feeds the data from open-source datasets as input into the upper body feature predictor 101. The intermediate clinical feature classifier 103 then output as intermedial clinical features that are further trained by the open source datasets as final clinical outcome.
[0029] As shown in FIG 1, there is a manual module comprising the clinician’s evaluation 105 where a clinician receives input from open-source datasets and reviews the intermediate clinical features, and the intermediate clinical features are trained by the intermediate clinical feature classifier 103 to output the evaluation as final clinical outcome; alternatively, the clinician directly evaluates input from the patient data and provides a final clinical outcome.
[0030] FIG. 2 shows an intelligent Medical Cyber-Physical System (MCPS) 200 for implementing the Artificial Intelligence Enabled 3D-Facial Scanner for Airway Assessment system (AINFAS) 100. As shown in FIG. 2, the MPCS 200 includes telemedicine devices that enable remote assessments: (i) a Patient device 201 which can be an Android tablet or phone; (ii) a Clinician device 202, such as a device used by an anesthesiologist, which can be an Android tablet or phone; and (iii) a Server 203. Naturally this invention is not limited to these three telemedicine devices 201,202,203. The software interface could be any Appthat allows patients to capture photos of themselves and allows clinicians to review the AT- generated results.
[0031] As shown in FIG. 2, a patient uploads his / her images and videos to the Patient device 201 , where an ATNFAS 100 app operating in the patient device 201 extracts relevant features using Al. After the patient captures his / her photos, the images arc converted into 3D models, followed by feature extractions in the AINFAS app operating in the patient device 201. The features extracted include landmarks of the thyroid notch, tragus, eyes, chin, sternal notch, upper lip, lower lip, iris, and nose. The features extracted also include the segmented region inside the mouth region; the features extracted constitute each patient data.
[0032] As shown in FIG. 2, the patient data is then sent to the clinician for review on the Clinician device 202. Upon receiving the results and the patient data, the clinician makes the necessary decision and returns the results to the patient. Clinicians may conduct further assessments on high-risk patients clinically if deemed necessary. Remote assessment will be sufficient for most patients, thereby reducing the wait times and costs for the patients involved. Simultaneously, these features extracted arc transferred to the secure server 203 for assessments in two stages by an explainable artificial intelligence (XAI) engine 204 which utilizes these features extracted as input. Explainability is crucial in medical decisionmaking. Unlike black boxes, explainable predictions could augment medical practitioners in making informed decisions.
[0033] The XAI engine 204 estimates the thyromental distance (TMD), neck length, neck flexion and extension angles, presence of protruding chin, Mallampati score, and interincisor distance (IID) in an intermediate stage prediction. Thus, the XAI engine 204 predicts difficult-airways (DA) through multiple stages of assessments based on well-established clinical criteria and outputs results from intermediate assessments as well as the final prediction (i.e., in Cormack and Lehane intubation grade). Due to the usage of well- established clinical criteria, clinicians can comprehend the outcomes from the intermediate assessments and possible reasoning behind the final prediction. In addition, the XAI engine 204 predicts the difficulty of mask ventilation, difficulty of supraglottic insertion and difficulty of intubation in the final stage prediction.
[0034] Whilst not shown in FIG. 1, the AINFAS 100 incorporates four components, namely: (a) an authentication system for privacy and user identification, (b) a File upload for imageand video input, (c) an Image processing for feature extraction, and (d) Results analysis for displaying of the classification results. According to an embodiment of this invention, the AINFAS 100 app is developed in Flutter, which uses a custom rendering approach that gives performances close to native applications; naturally, other development platforms can be used.
[0035] In the File upload component, images and videos can be either captured using the patient device camera or uploaded from the file storage as depicted in FIG. 3 where 301 is a screenshot showing a guideline is overlaid on the patient device 201 camera screen; and 302 shows images and videos are dragged and dropped into the respective views to ease the image processing process. Then, the files and the features extracted data are transferred to the server 203. Finally, when the classification results are returned, a summary 303 is displayed as illustrated in FIG. 3. Detailed results can be downloaded, for example, in a pdf format.
[0036] Patient dataset collection: During the course of the invention, the inventors have collected 720 patient data and have anonymized 194 of them. Among the 194 patients, 156 patients have completed their surgery, where their airway difficulty has been identified. Each subject has a total of 9 images and 3 videos taken using a mobile device, specifically to mimic the environment in atypical use case. The details of the data are listed below: The images taken are:1. Sitting, front view (mouth closed, forward facing)2. Sitting, front view (mouth opened, forward facing)3. Sitting, front view, zoomed-in view with mouth only (mouth opened)4. Sitting, profile view - nght (mouth closed, forward facing)5. Sitting, profile view - right (mouth closed, full neck extension)6. Sitting, profile view - nght (mouth closed, full neck flexion)7. Sitting, profile view - left (mouth closed, forward facing)8. Sitting, profile view - left (mouth closed, full neck extension)9. Sitting, profile view - left (mouth closed, full neck flexion)The videos taken are:1. Sitting, profile view - right (dynamic neck flexion and extension view)2. Sitting, profile view - left (dynamic neck flexion and extension view)3. 360° video of patient for facial reconstruction
[0037] Labelling of the images and videos will be described further below. Additionally, the inventors collected the following intermediate clinical features from the subjects' electronic health records (EHR), which are then anonymized for use in Al training:• Age• Gender• BMI• IRIS diameter (mm)• Medical History• Surgical History7• Anaesthetic history• Mallampati Score• Temporomandibular distance• Neck extension• Mouth opening• teeth / prosthesis• mask ventilation• Supraglottic airway insertion• Rapid sequence induction• Laryngoscope type• Intubation grade• Adjuncts• External laryngeal pressure• Any adverse outcome• Difficult ai rway
[0038] Thyromental distance feature extraction: The inventors manually labeled the eyes, chin, thyroid notch, and sternal notch on a publicly available ALFW (Annotated Facial Landmarks in the Wild) dataset (or multiview facial dataset) which is then named ALFW- DA; these landmarks are for estimating the Thyromental Distance (TMD) by calculating the pixel distance between the chin and the thyroid notch. Since tire distance between the eyesis fixed, this distance acts as a ratio to scale the pixel distance between the chin and the thyroid notch.
[0039] The invention uses a Global-to-local Cross Pseudo Supervision (G2LCPS) method; this G2LCPS method is a semi-supervised method that predicts landmarks related to DAs in two levels. The global level refers to classifying whether the image taken contains the neck region; and the local level refers to generating high-quality pseudo heatmaps for cross supervision between networks. Due to the semi-supervised nature, fewer labeled images are required, thus reducing the manual labeling process.
[0040] Results show that the G2LCPS method can classify images with neck and without neck successfully, as shown in FTGs. 4A-4B which illustrate some examples of labeled images from the ALFW-DA dataset. FIG. 4A shows “no neck” (NG) images; the first four images are NG because they do not contain visible neck landmarks; tire two following images are also NG since their quality is low, and they lack details to accurately determine the positions of neck landmarks. FIG. 4B represents acceptable examples, where all facial and neck landmarks are visible and can be easily located.
[0041] In FIGs. 5A-5B, the inventors compared normalized mean error (NME) of the facial landmarks predicted with errors occurring in three other prior art semi-supervised methods. The normalization factor is an expanded bounding box of each face provided by the ALFW- DA dataset 501 and the predicted bounding box of each face in the SidcFacc-DA 502. FIGs. 5A-5B show that the G2LCPS has achieved the lowest NME when predicting DA related landmarks.
[0042] These predicted landmarks in FIGs. 5A-5B are compared qualitatively with the ground truth in FIGs. 6A-6B obtained according to an embodiment of the invention. The top row 601 represents input images. The second row 602 and third row 603 indicate the true heatmaps and landmark ground truths. The last two rows 604, 605 are the predicted heatmaps and predicted landmarks produced by the G2LCPS method.
[0043] Tn addition, the inventors conducted an ablation study to compare the impact of an exponential moving average (EMA) filter, a global classifier filter, and a local filter on the NME. Tire results in FIG. 7 show that the impact of tire local filter on tire landmarkpredictions is the highest, whereas the impact of the EMA filter on the landmark predictions is the lowest.
[0044] The inventors also adapted a PFLD (practical facial landmark detector) model to predict landmarks for thyromental distance (TMD) estimation. The PFLD model is a supervised learning model that has a relatively small structure and provides real-time speed prediction by using a Mobilenet model: it has a backbone network that predicts the landmark coordination and an auxiliary network that estimates geometric information of the landmarks. FIGs. 8A-8C show the facial and thyroid landmarks predicted by the invention’s PFLD model using an optimizer for deep learning, such as, an AdaBelief optimizer. It can be seen from FIGs. 8A-8C that the eyes, chin, thyroid notch, and sternal notch are predicted.
[0045] Neck flcx / cxtcnsion angles feature extraction: the inventors manually labeled the tragus, forehead, nose bridge, and chin on head images with a side profile from CPLFW (cross-pose facial image), CFP (celebrities front-profile facial image), and 67 CAS1A-3D FaceVl (3D facial image) datasets. These facial landmarks are for estimating a neck flexion angle (NFA) and a neck extension angle (NEA). Prior art models for facial landmarks detection are commonly trained on frontal profiles, instead of side profiles. Hence, predicting facial landmarks on a side profile is under-explored. The inventors tuned the hyperparameters, trained, and validated a convoluted network (such as ResNet50) for this task.
[0046] FIGs. 9A-9J show examples of landmark predictions using an Adam optimizer Neck flexion angle (NFA) is estimated by comparing predictions of the image frame with an upright head pose and the image frame with the head bent forward. Conversely, neck extension angle (NEA) is estimated by comparing predictions of the image frame with an upright head pose and the image frame with the head tilted backward. It is shown that the forehead, nose bridge, chin, and tragus are predicted.
[0047] From the above description, the AINFAS 100 invention is implemented using a patient mobile camera to capture and store images and videos in the comfort of their homes. This invention is a multi-modal fusion where video and images of a patient from multiple views are used for better assessment of facial features. All data are then subsequently processed simultaneously with an ensemble of models to generate an airway assessment.
[0048] The invention is advantageous because it incorporates the intelligent medical cyberphysical system (MCPS) approach into the AINFAS 100, is unique from the prior art in terms of the specific problem domains being addressed, the algorithms used, and the valuable outcomes achieved. For example, the current clinical airway assessments are performed cither in-pcrson through standard checklists or involve medical imaging. In the present invention, mobile camera-based predictors enable the assessment to be performed remotely. Besides that, the usage of mobile cameras is affordable, accessible, and safe from radiation.
[0049] While there are many datasets and prior art relating to facial features, there is a severe lack of available labeled data for neck landmarks and difficulty airway. In addition, there are numerous challenges specific to the neck region. Due to the minimal features present on the neck, identifying crucial airway landmarks becomes challenging. These landmarks are more prominent in some individuals, specifically in thin males. When the landmarks are prominent, they may be estimated based on the shadow cast on the neck, which depends on the lighting condition. The dataset collected and used in this invention was captured in a typical room, using a mobile device. As a result, the dataset is uniquely suited to the environment the invention system is used in, contributing to the novelty of this invention.
[0050] When validated against conventional methods, the AINFAS 100 invention has the potential to be a new gold standard for airway assessment, with better specificity and sensitivity' in identifying difficult-airways (DAs). This system can also be used in surgical clinics or polyclinics, where prc-opcrativc assessment of uncomplicated patients can be done, saving them at least one clinic visit. This optimizes the utilization of healthcare resources, allowing anesthesiologists to focus on complex prc-opcrativc patients.
[0051] Black-box predictors in the prior art struggle with a lack of trust, confidence, and transparency among medical practitioners. The invention uses explainable Al (XAI engine 204) and has multiple stages with key clinical criteria as intermediate features. Explainable predictions could augment medical practitioners in making informed decisions, especially when the explanations are based on well-established clinical criteria. Multi-staged Al models enable clinical inspection at every intermediate step, exposing the rationale for each intermediate classification or measurement and facilitating clinician trust in the final assessment.
[0052] The invention helps anesthesiologists assess their patients virtually in a timely manner, prepare appropriate resources for operation without overburdening hospital resources and save more lives. The invention also serves as a teaching or assessment tool for junior and other medical staff managing airway assessments. For those with limited experience in this medical area, the invention will offer feedback therefore shortening the learning curve and enhancing their learning speed.
[0053] While specific embodiments have been described and illustrated, it is understood that many changes, modifications, variations and combinations of variations disclosed in the text description and drawings thereof could be made to the present invention without departing from the scope of the present invention. For example, this invention adapted using the above principles for automated assessment of sleep apnea patients and automated or custom mask sizing for patients.
Claims
CLAIMS:
1. A difficult-airways assessment method, comprises: generating three-dimensional scans from multiple images of a face and neck of a patient taken by the patient; analysing the 3D scans using a deep-leaming algorithm to extract a plurality of landmarks; constructing a difficult-airway classifier for multi-stage assessment; and using an artificial intelligence model to predict difficult-airways severity based on the multi-stage assessment.
2. The difficult-airways assessment according to claim 1, wherein the multi-stage assessment comprises an intermediate clinical outcome and a final clinical outcome.
3. The difficult-airways assessment according to claim 1 or 2, wherein the artificial intelligence model employs an explainable artificial intelligence (XA1) engine which uses key clinical criteria.
4. A system for assessing difficult-airway comprising: a three-dimensional facial scanner to capture multiple images of face and neck of a patient; a deep learning processor for extracting a plurality of landmarks in the multiple images of the face and neck; a classifier for sorting the landmark images for multi-stage assessment; and an artificial intelligence processor to predict difficult-airway severity based on the multi-stage assessment.
5. The system for assessing difficult-airway according to claim 4, wherein the classifier comprises an intermediate clinical feature classifier and a final clinical outcome classifier.
6. The system for assessing difficult-airway according to claim 4 or 5, wherein the deep learning processor employs a Global-to-local Cross Pseudo Supervision (G2LCPS) method, wherein the method comprises a global level and a local level.