Intelligent identification method and system for difficult airway glottis features
By embedding a micro microphone in the laryngeal mask to collect respiratory sound signals, feature extraction and separation of interference sounds, combined with multi-scale glottic feature fusion and timing modeling, the accuracy and timing of difficult airway evaluation are solved, efficient and accurate difficult airway prediction is achieved, and anesthesia safety and intubation success rate are improved.
Patent Information
- Application Number
- CN202510942328.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, difficult airway assessment has problems such as strong subjectivity, low accuracy and improper evaluation timing. It is impossible to accurately, objectively and non-invasively predict difficult airways before anesthesia, resulting in the impact of anesthesia safety and patient prognosis.
The patient's respiratory sound signal was collected through a micro microphone embedded in the laryngeal mask, and time-frequency domain features were adaptively extracted after preprocessing. The ring-shaped cartilage compression sound and secretion interference sound were separated. Multi-scale glottic feature pyramid fusion and timing modeling were used, and the Cormack-Lehane grading results were determined based on the dual-branch decision network and Bayesian decision reasoning.
The prediction accuracy of difficult airways was improved to 91.7%, the incidence of anesthesia-related complications was reduced by 15% to 20%, the success rate of first intubation was improved by 10% to 15%, the time for repeated attempts and anesthesia induction was reduced, and the allocation of medical resources was optimized.
Smart Images

Figure CN120449018A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and in particular to a method and system for intelligently identifying glottal features of difficult airways, which is suitable for predicting, evaluating, and treating difficult airways in clinical work in anesthesiology. Background Art
[0002] In clinical anesthesia, difficult airway is a common and dangerous condition that can result in serious patient injury or even death. Currently, the assessment of difficult airway relies primarily on the anesthesiologist's experience and simple predictive indicators, such as the Mallampati classification and thyromental distance. These assessment methods are subject to high subjectivity and limited accuracy. Although direct laryngoscopy can accurately assess the Cormack-Lehane (CL) classification, it can only be performed after anesthesia induction. If a difficult airway is detected at this time, the patient may face the dangerous situation of being unable to ventilate or intubate.
[0003] Some existing studies have attempted to predict difficult airway conditions using imaging methods such as ultrasound and CT, but these methods are costly and complex, hindering widespread adoption. In recent years, acoustic signal-based analysis methods have gained attention, but existing acoustic analysis technologies face challenges such as inaccurate signal acquisition, incomplete feature extraction, and significant noise interference, making them unable to accurately identify glottal features associated with difficult airway conditions.
[0004] Therefore, there is an urgent need for a method and system that can accurately, objectively and non-invasively predict difficult airway before anesthesia to improve anesthesia safety and patient prognosis. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for intelligently identifying glottal features of a difficult airway, aiming to solve the problems of high subjectivity, low accuracy and inappropriate timing of difficult airway assessment in the prior art.
[0006] The present invention proposes an intelligent recognition method for difficult airway glottal features, comprising:
[0007] Acquiring the patient's respiratory sound signal, which is collected by a micro-microphone embedded in the laryngeal mask;
[0008] Preprocessing the respiratory sound signal to obtain a time-frequency graph;
[0009] Adaptively extracting time-frequency domain features based on the time-frequency graph to obtain glottal features;
[0010] performing separation and enhancement processing on the glottis feature by performing cricoid cartilage compression sound and secretion interference sound separation and enhancement processing to obtain an enhanced glottis feature;
[0011] Based on the enhanced glottal features, a multi-scale fusion feature is obtained by fusing the multi-scale glottal feature pyramid;
[0012] Based on the multi-scale fusion features, the glottal state is analyzed by time series modeling to obtain a glottal state sequence and a time series feature vector;
[0013] Determining the Cormack-Lehane grading result through decision reasoning based on the time series feature vector and glottal state statistical characteristics;
[0014] The recommended laryngoscope model was determined based on the Cormack-Lehane classification results.
[0015] Preferably, the miniature microphone is arranged at the second curve position of the laryngeal mask, and the sensitivity of the miniature microphone is -42dBV / Pa, the frequency response is 20Hz-20kHz, and the signal-to-noise ratio is not less than 65dB.
[0016] Preferably, the pre-processing includes:
[0017] performing noise elimination processing on the respiratory sound signal;
[0018] Performing frame processing on the respiratory sound signal after noise elimination processing;
[0019] The breath sound signal after frame processing is band-pass filtered to extract the glottal characteristic frequency band of 80-300 Hz;
[0020] Performing wavelet transform processing on the extracted glottal characteristic frequency band to obtain the time-frequency graph;
[0021] The framing process includes:
[0022] Adopting a framing strategy with a 25ms frame length and a 10ms frame shift;
[0023] Multiply each frame by a Hamming window to reduce spectral leakage.
[0024] Preferably, the wavelet transform process includes:
[0025] The 9-level discrete wavelet decomposition is adopted, and the base wavelet is Daubechies-4 wavelet;
[0026] Threshold denoising is performed on level 1-4 coefficients, and the threshold is adaptively calculated based on the local energy distribution of the signal;
[0027] The reconstructed time-frequency diagram has a resolution of 10ms×5Hz.
[0028] As a priority, the time-frequency domain feature adaptive extraction includes:
[0029] Features are extracted through time domain channel, frequency domain channel and time-frequency joint channel respectively;
[0030] Based on the signal quality evaluation index, the features extracted from the time domain channel, the frequency domain channel and the time-frequency joint channel are weightedly fused to obtain the glottal feature;
[0031] Among them, the time domain channel includes three sets of parallel one-dimensional convolutional networks, and the convolution kernel sizes are 3, 5, and 7 time points respectively; the frequency domain channel includes a deformable convolutional network, and the receptive field can be adaptively adjusted; the time-frequency joint channel includes a four-layer two-dimensional convolutional network, and the convolution kernel size is 3×3.
[0032] As a priority, the enhanced separation process of the cricoid cartilage compression sound and the secretion interference sound includes:
[0033] The cricoid cartilage compression sound features and secretion interference sound features are extracted respectively through a dual-path noise feature learning network.
[0034] Enhance target glottal features and suppress noise features through spectral subtraction residual connection;
[0035] An adversarial learning strategy is used to optimize the feature extractor so that it generates glottal features that are difficult for the discriminator to distinguish.
[0036] As a priority, the multi-scale glottal feature pyramid fusion includes:
[0037] inputting the enhanced glottis features into the microstructure layer, the transient dynamic layer, the periodic variation layer and the overall pattern layer for processing respectively;
[0038] The features of each layer are weighted by a feature fusion gating unit, and inter-layer fusion is performed to obtain the multi-scale fusion feature;
[0039] Among them, the receptive field of the microstructure layer corresponds to the physical scale of 1-5ms, focusing on the tiny vibration characteristics of the opening and closing of the glottis; the receptive field of the transient dynamic layer corresponds to the physical scale of 5-20ms, focusing on the transient process characteristics of the opening and closing of the glottis; the receptive field of the periodic change layer corresponds to the physical scale of 50-200ms, focusing on the periodic change characteristics of the opening and closing of the glottis; the receptive field of the overall pattern layer corresponds to the physical scale of 500-2000ms, focusing on the overall pattern characteristics of the opening and closing of the glottis; information interaction is achieved between each layer through a bidirectional information flow path, including a top-down path and a bottom-up path.
[0040] As a priority, the timing modeling includes:
[0041] Extracting temporal dependency features through a bidirectional gated recurrent unit network;
[0042] Aggregate key temporal features through weighted temporal attention mechanism;
[0043] A glottal state transition graph is established based on weighted aggregated key temporal features;
[0044] Performing duration modeling on the glottal state sequence to obtain the time series feature vector;
[0045] The glottal state transition diagram includes four glottal states: fully open state, partially open state, partially closed state and fully closed state; the state sequence is smoothed by the conditional random field layer, and the Viterbi algorithm is used to find the global optimal state sequence.
[0046] As a priority, the decision reasoning includes:
[0047] Integrating the time series feature vector and glottal state statistical features into a decision feature vector;
[0048] The Cormack-Lehane grade probability distribution and difficult / non-difficult airway discrimination results were predicted respectively by a two-branch decision network.
[0049] Integrating the dual-branch prediction results based on Bayesian decision reasoning to obtain the Cormack-Lehane classification result;
[0050] Among them, the Bayesian decision reasoning includes: using the Cormack-Lehane distribution probability of each level based on the hospital's historical data statistics as prior knowledge; adjusting the main branch output probability based on the conditional probability of the Cormack-Lehane distribution of different population characteristics; calibrating the graded prediction using the difficult / non-difficult airway discrimination results; calculating the final posterior probability distribution, and selecting the Cormack-Lehane level with the highest posterior probability as the final prediction; when the prediction confidence is lower than the threshold, triggering the multi-model integration decision mechanism, including the integration of this deep learning model, XGBoost model and LightGBM model.
[0051] The intelligent recognition system for difficult airway glottis features includes:
[0052] A signal acquisition device for acquiring a patient's breathing sound signal, the signal acquisition device comprising a miniature microphone embedded in the laryngeal mask;
[0053] A signal preprocessing module, configured to preprocess the respiratory sound signal to obtain a time-frequency graph;
[0054] a feature extraction module for adaptively extracting time-frequency domain features based on the time-frequency graph to obtain glottal features, and performing separation and enhancement processing on the glottal features to separate cricoid cartilage compression sounds and secretion interference sounds to obtain enhanced glottal features;
[0055] A feature representation module is used to obtain a multi-scale fusion feature by fusing a multi-scale glottal feature pyramid based on the enhanced glottal feature;
[0056] A time series modeling module is used to analyze the glottal state through time series modeling based on the multi-scale fusion features to obtain a glottal state sequence and a time series feature vector;
[0057] A decision-making and reasoning module, configured to determine a Cormack-Lehane grading result through decision-making and reasoning based on the time series feature vector and the statistical features of the glottal state;
[0058] An output module is used to determine a laryngoscope model recommendation based on the Cormack-Lehane classification result;
[0059] Wherein, the feature extraction module includes:
[0060] The time-frequency domain feature adaptive extraction unit includes a time domain channel, a frequency domain channel, and a time-frequency joint channel, which are used to extract features separately and perform weighted fusion based on signal quality evaluation indicators;
[0061] The unit for separating and enhancing cricoid cartilage compression sounds from secretion interference sounds includes a dual-path noise feature learning network and a spectrum subtraction residual connection;
[0062] The feature representation module includes:
[0063] Microstructure layer, transient dynamic layer, periodic variation layer and overall pattern layer are used to extract features from different time scales;
[0064] Feature fusion gating unit, used to weight the features of each layer and perform inter-layer fusion;
[0065] The decision reasoning module includes:
[0066] A two-branch decision network is used to predict the Cormack-Lehane grade probability distribution and difficult / non-difficult airway discrimination results respectively;
[0067] The Bayesian decision inference layer is used to integrate the two-branch prediction results to obtain the final Cormack-Lehane classification result.
[0068] The beneficial effects of the present invention include:
[0069] 1. Improved the accuracy of difficult airway prediction, with the CL grading accuracy reaching 91.7%, 25% to 30% higher than traditional methods, thereby reducing the incidence of anesthesia-related complications by 15% to 20%;
[0070] 2. The micro-microphone embedded in the laryngeal mask enables non-invasive and convenient respiratory sound collection, without requiring additional operation or increasing patient discomfort;
[0071] 3. Through innovative glottal feature extraction and analysis technology, interference signals such as cricoid cartilage compression sound and secretion interference sound are effectively filtered out, improving the accuracy and robustness of feature recognition;
[0072] 4. Adopting multi-scale feature pyramid fusion technology, it comprehensively captures glottal features at different time scales from microscopic to macroscopic, thus enhancing the adaptability of the system.
[0073] 5. Based on the glottal state time series modeling, accurate analysis of the dynamic opening and closing characteristics of the glottis is achieved, providing a more comprehensive basis for difficult airway assessment;
[0074] 6. Through a self-calibrated decision-making and reasoning mechanism, the system can dynamically adjust its prediction strategy based on different patient characteristics, improving the system's generalization ability across different populations;
[0075] 7. It increased the first-time intubation success rate by 10% to 15%, reduced repeated attempts, shortened the anesthesia induction time by 3 to 5 minutes, optimized the allocation of medical resources, and improved the quality of medical care. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 Schematic diagram of the overall architecture of the intelligent recognition system for difficult airway glottis features of the present invention;
[0077] Figure 2 This is a flow chart of the method for intelligently identifying glottal features of difficult airways according to the present invention;
[0078] Figure 3 It is a structural diagram of a signal acquisition device;
[0079] Figure 4 It is a processing flow chart of the signal preprocessing module;
[0080] Figure 5 It is a structural diagram of the time-frequency domain feature adaptive extraction unit;
[0081] Figure 6 This is a schematic diagram of the structure of the unit for separating and enhancing the cricoid cartilage compression sound and secretion interference sound;
[0082] Figure 7 This is a schematic diagram of the structure of the multi-scale glottal feature pyramid fusion module;
[0083] Figure 8 The processing flow chart of the timing modeling module;
[0084] Figure 9 This is a structural diagram of the decision-making reasoning module;
[0085] Figure 10 Schematic diagram of the system's application scenario. DETAILED DESCRIPTION
[0086] Please refer to the attached Figure 1-10 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0087] like Figure 1 As shown, the intelligent recognition system for difficult airway glottis features provided by the present invention includes a signal acquisition device 1, a signal preprocessing module 2, a feature extraction module 3, a feature representation module 4, a time series modeling module 5, a decision reasoning module 6 and an application output module 7. The feature extraction module 3 includes a time-frequency domain feature adaptive extraction unit 31 and a cricoid cartilage compression sound and secretion interference sound separation and enhancement unit 32. The feature representation module 4 includes a microstructure layer 41, a transient dynamic layer 42, a periodic change layer 43, an overall pattern layer 44 and a feature fusion gating unit 45. The decision reasoning module 6 includes a dual-branch decision network 61 and a Bayesian decision reasoning layer 62.
[0088] The workflow of the system is as follows: the patient's respiratory sound signal is collected by the signal acquisition device 1, and the signal preprocessing module 2 preprocesses it to obtain a time-frequency diagram. The feature extraction module 3 extracts the glottal features based on the time-frequency diagram and performs interference sound separation and enhancement. The feature representation module 4 performs multi-scale fusion on the enhanced glottal features. The time series modeling module 5 analyzes the glottal state and extracts time series features. The decision reasoning module 6 determines the Cormack-Lehane grading result based on the time series features. Finally, the application output module 7 determines the laryngoscope model recommendation.
[0089] like Figure 2 As shown, the method for intelligently identifying glottal features of difficult airways provided by the present invention comprises the following steps:
[0090] In one embodiment of the present invention, the respiratory sound signal is collected by a micro-microphone embedded in the laryngeal mask. Figure 3 As shown, the miniature microphone is set at the second curve position of the laryngeal mask to ensure that it can be aimed at the glottis area and improve the accuracy of collection.
[0091] Preferably, the miniature microphone has a sensitivity of -42dBV / Pa, a frequency response of 20Hz-20kHz, and a signal-to-noise ratio of no less than 65dB. These parameter settings ensure that the microphone is sufficiently sensitive to low-intensity breath sounds, has low self-noise, and can cover the frequency range of the entire glottal characteristic.
[0092] In addition, the acquisition system uses an 8kHz sampling rate and 16-bit quantization accuracy, providing a 96dB dynamic range, sufficient to capture subtle changes in breathing sounds. The microphone housing uses a medical-grade metal mesh to prevent direct contact with secretions while maintaining acoustic transparency.
[0093] During signal acquisition, the raw acoustic signal is transmitted to the processing unit in real time via Bluetooth 5.0 low-power transmission technology. The system automatically begins signal acquisition when the laryngeal mask is placed using a pressure sensor. The acquisition duration is set to 30 seconds, which is sufficient to capture multiple complete respiratory cycles.
[0094] like Figure 4 As shown in FIG, the preprocessing stage includes four steps: noise elimination, signal framing, bandpass filtering and wavelet transform processing.
[0095] First, adaptive noise cancellation technology is used to pre-process the original respiratory sound signal. Specifically, the ambient noise is collected through the reference channel, and the filter parameters are adaptively adjusted based on the minimum mean square error criterion to eliminate non-physiological background noise.
[0096] Next, the noise-reduced signal is framed. In a preferred embodiment of the present invention, a framing strategy with a 25ms frame length and a 10ms frame shift is adopted, and each frame is multiplied by a Hamming window to reduce spectrum leakage. The expression of the Hamming window is:
[0097] ,
[0098] in: For the The Hamming window coefficient of each sampling point is in the range of [0.08, 1]; The number of sampling points corresponding to the frame length. At an 8kHz sampling rate, a 25ms frame length corresponds to sampling points; The sampling point index ranges from 0 to ; is a cosine function. For anesthesia cases in practical applications, this window function can effectively reduce spectral leakage, thereby more accurately capturing the short-term spectral characteristics generated by the opening and closing of the glottis.
[0099] After framing, the system designs a bandpass filter with a passband of 80–300 Hz and a stopband attenuation of ≥40 dB to extract the characteristic frequency band of glottal opening and closing. The 80–300 Hz frequency band was chosen because the primary acoustic characteristics of glottal opening and closing are concentrated in this frequency band, while most ambient noise and other physiological noise, such as heart sounds, are distributed elsewhere. In actual clinical settings, ambient noise in anesthesia rooms is typically concentrated in the low-frequency (<50 Hz) and high-frequency (>1 kHz) regions, so selecting this frequency band effectively improves the signal-to-noise ratio.
[0100] Finally, the filtered signal is subjected to wavelet transform processing. Preferably, a 9-level discrete wavelet decomposition is used, and the base wavelet is the Daubechies-4 wavelet. The Daubechies-4 wavelet has good time-frequency localization characteristics and high regularity, and is suitable for analyzing non-stationary signals such as respiratory sounds. Threshold denoising is performed on the coefficients of levels 1-4, and the threshold is adaptively calculated based on the local energy distribution of the signal. The formula is:
[0101] ,
[0102] in: For the The threshold of the level wavelet coefficient, the unit is the relative unit of the signal amplitude; For the The standard deviation of the wavelet coefficients of the level reflects the degree of fluctuation of the coefficients of the level; For the The length of the first wavelet coefficient decreases as the decomposition level increases. The length of the stage is approximately ; is the natural logarithm. This formula, based on statistical theory and assuming a Gaussian distribution, maximizes the preservation of signal characteristics while suppressing noise. During anesthesia, when the patient's airway is obstructed, breath sounds exhibit irregular variations. Using this threshold for denoising can better preserve these characteristic variations.
[0103] After reconstruction, a time-frequency diagram is obtained with a resolution of 10ms in the time dimension and 5Hz in the frequency dimension. Preferably, the system also calculates the signal-to-noise ratio, harmonic-to-noise ratio, and short-time energy stability index for adaptive adjustment of parameters for subsequent feature extraction. These indicators are crucial for evaluating glottal activity. For example, in patients with difficult airway, the common signal-to-noise ratio is less than 10dB, the harmonic-to-noise ratio is less than 0.5, and the short-time energy variation coefficient is greater than 0.4. These are all acoustic characteristic indicators of difficult airway.
[0104] like Figure 5 As shown, the time-frequency domain feature adaptive extraction unit 31 adopts a three-channel parallel architecture, including a time domain channel, a frequency domain channel and a time-frequency joint channel.
[0105] The time domain pipeline consists of three parallel one-dimensional convolutional networks with kernel sizes of 3, 5, and 7 time points, respectively, to capture glottal features at different time scales. Each convolutional network consists of four layers, downsampled layer by layer, with each layer containing 64 feature maps. The first three layers use the ReLU activation function, and the last layer uses the PReLU (parameterized ReLU) activation function, expressed as:
[0106] ,
[0107] in: is the output of the activation function; is the input value; is a learnable parameter, with an initial value of 0.25, which is adjusted as the network is trained; When , the output is equal to the input; when When , the output is the input multiplied by the parameter Compared to standard ReLU, PReLU can handle negative inputs, which is particularly important for capturing features during glottal closure (usually manifested as energy drop). In actual difficult airway recognition, incomplete glottal closure generates a series of negative features, which can be better preserved using PReLU.
[0108] The frequency domain channel uses a deformable convolutional network, and the receptive field can be adaptively adjusted. The initial convolution kernel size is 3 frequency points, the deformation parameters are learned by the network, and the initial offset is set to 0. The expression of the deformable convolution is:
[0109] ,
[0110] in: The output feature map is at position The value of For the convolution weights; is the input feature map; is the sampling position offset of the standard convolution; is the learned deformation offset; is the number of points contained in the convolution kernel; Denotes the weighted sum of all convolution kernel points. Deformed convolution allows the receptive field to be dynamically adjusted based on the input features, which is very effective for addressing individual differences in glottal features among patients. For example, glottal features in obese patients are typically in lower frequency bands, while those in thin patients are in higher frequency bands. Deformed convolution can adaptively adjust the region of interest.
[0111] In addition, the frequency channel also includes a frequency attention mechanism to learn the frequency importance weight, with special attention to the 80-300Hz range. The calculation formula of the frequency attention mechanism is:
[0112] ,
[0113] in: is the frequency attention weight vector, the dimension is the same as the frequency dimension; and is a learnable weight matrix that maps the feature dimension to the hidden dimension and back to the feature dimension; is the frequency feature vector; GELU is the Gaussian error linear unit activation function; softmax is the normalization function to ensure that the sum of all weights is 1. The expression of the GELU function is:
[0114] ,
[0115] in: is the hyperbolic tangent function; =π is the circumference constant, approximately equal to 3.14159. In difficult airway identification, different frequencies contribute differently to determining the degree of difficulty. For example, the semi-closed glottis state typically has distinct characteristics in the 120-200 Hz range. The frequency attention mechanism can automatically learn to emphasize these key frequency bands.
[0116] The joint time-frequency channel uses a four-layer two-dimensional convolutional network with a 3×3 kernel size and 32 feature maps. The pooling strategy is: max pooling for the first and third layers, and average pooling for the second and fourth layers. This alternating pooling strategy preserves both salient and smooth features, which is beneficial for processing both spikes and continuous features during glottal opening and closing.
[0117] The features of the three channels are fused through the channel attention gating module. This module first extracts the global features of each channel through global average pooling, then passes it through a two-layer fully connected network (64-32-64) with a GELU activation function in the middle, and finally generates the fusion weights of the three channels through softmax normalization. The formula is:
[0118] ,
[0119] in: For the The fusion weight of each channel is in the range of [0,1]; is the score of the corresponding channel; is the natural exponential function; Indicates the sum of three channels; The value range is 1 to 3, corresponding to the time domain channel, frequency domain channel, and combined time-frequency channel, respectively. In practical applications, the importance of the three channels varies for different types of difficult airways. For example, for difficult airways with retroglottal displacement, frequency domain features are generally more important; for difficult airways with glottal edema, time domain features may be more discriminative; and for difficult airways with a small glottal inlet, combined time-frequency features are often more effective.
[0120] Fusion weights are dynamically adjusted based on signal quality assessment metrics. For example, when the signal-to-noise ratio (SNR) falls below 5dB, the system reduces the time-domain channel weight to 0.2 and increases the frequency-domain channel weight to 0.5. When energy is unstable (coefficient of variation greater than 0.4), the time-domain channel weight is increased to 0.6. This adaptive adjustment mechanism significantly improves the robustness of feature extraction, enabling the system to adapt to changes in signal quality under varying anesthesia environments and patient conditions.
[0121] like Figure 6 As shown, the cricoid cartilage compression sound and secretion interference sound separation and enhancement unit 32 includes a dual-path noise feature learning network and a spectrum subtraction residual connection.
[0122] The dual-path noise feature learning network includes a cricoid cartilage compression sound feature extraction path and a secretion interference sound feature extraction path. The cricoid cartilage compression sound feature extraction path uses a progressive frequency selective attention structure with five layers of progressive processing. Each layer includes 16 3×3 convolution kernels, batch normalization, and GELU activation functions. The attention mechanism is implemented through the frequency domain attention matrix, and the size and feature Figure 1 The attention weight is calculated based on the frequency energy distribution characteristics. The attention weight of each layer is updated by the weighted sum of the previous layer result and the current layer calculation result. The update formula is:
[0123] ,
[0124] in: For the The final attention weight matrix of the layer; is the attention weight matrix of the previous layer; The original attention weight matrix calculated for the current layer; is the weight factor, the value range is [0,1], and 0.7 is taken in the experiment; is the layer index, ranging from 1 to 5. This progressive attention structure enables layer-by-layer refinement of the cricoid pressure sound feature extraction. Cricoid pressure is a common technique in clinical anesthesia to prevent reflux of gastric contents, but it also produces interfering sounds that can affect the recognition of glottal features. This pathway specifically learns these interfering features for subsequent separation.
[0125] The feature extraction path for secretion interference sounds is based on the temporal dependency modeling of the Transformer structure. It includes four self-attention heads, each focusing on a different temporal pattern, and a hidden dimension of 128. The position encoding uses sine-cosine encoding, and the expression is:
[0126] ,
[0127] ,
[0128] in: For position pos, dimension The positional encoding value of For position pos, dimension Position encoding value; pos is the time position index; i is the dimension index, ranging from 0 to ; The model dimension is 128; sin and cos represent the sine and cosine functions, respectively; and 10,000 is the scaling factor. Positional encoding enables the model to perceive temporal information, which is crucial for identifying irregular interference patterns caused by secretions. In patients with difficult airways, especially those with high secretions (such as those with respiratory infections or prolonged bed rest), the acoustic interference caused by secretions often exhibits specific temporal patterns, and this encoding helps the model capture these patterns.
[0129] Transformer consists of 2 layers, each of which contains a self-attention module and a feedforward neural network. The self-attention calculation uses the scaled dot product attention mechanism with a temperature parameter set to 0.1. The formula is:
[0130] ,
[0131] in: 、 、 They are query, key, and value matrices, respectively, all of which are obtained by linear transformation of input features; It represents the transpose multiplication of the query matrix and the key matrix to obtain the similarity matrix; The dimension of the key is 32; is the temperature parameter, which is set to 0.1; is a scaling factor; softmax is a normalization function that converts similarity into attention weights; the entire expression represents the weighted summation of the value matrix using the attention weights to obtain the attention output. A lower temperature parameter (0.1) makes the attention distribution sharper, helping to accurately locate the moment of secretion interference.
[0132] The spectral subtraction residual connection performs a subtraction operation in the feature space to enhance the target glottal features and suppress the noise features. The implementation steps include:
[0133] (1) Align the target feature map and the noise feature map: adjust them to the same dimension through bilinear interpolation;
[0134] (2) Feature channel mapping: The noise feature map is mapped to the target feature space through the conversion network, and the conversion network is a 1×1 convolutional layer;
[0135] (3) Residual calculation: The target feature is subtracted from the mapped noise feature to obtain the enhanced feature;
[0136] (4) Gated fusion: Introduce feature gating units to control the intensity of subtraction operations and prevent excessive suppression.
[0137] The calculation formula of the gate control unit is:
[0138] ,
[0139] in: is the gating coefficient matrix, and each element ranges from [0,1]; It is a sigmoid function that compresses the output to the O-1 range; is the learnable weight matrix; and are the target feature and the noise feature, respectively; [;] denotes a feature concatenation operation, which connects the two features along the channel dimension. The gating mechanism allows the system to dynamically adjust the suppression strength based on feature similarity, avoiding over-suppression of true glottal features.
[0140] The calculation formula of the final enhanced feature is:
[0141] ,
[0142] in: For the enhanced features; is the target feature; is the gating coefficient; represents element-wise multiplication, is the noise feature after mapping. In clinical practice, this enhancement mechanism can effectively reduce interference caused by cricoid cartilage compression and secretions, improving the clarity of glottal features. For example, when anesthesiologists apply cricoid cartilage compression before intubation, this mechanism can reduce the interference caused by compression and preserve the true glottal state information.
[0143] This unit also employs an adversarial learning strategy to optimize the feature extractor. The noise discriminator is constructed as a three-layer fully connected network (128-64-2). The training objective is for the discriminator to attempt to distinguish glottal features from noise features, while the feature extractor attempts to generate glottal features that are difficult for the discriminator to distinguish. This is achieved through adversarial training using a gradient reversal layer. The adversarial loss weight is 0.2 in the total loss function to prevent excessive adversarial forces from impacting the main task.
[0144] This separation enhancement mechanism can effectively handle common clinical interference sounds such as cricoid cartilage compression and secretion noise, significantly improving the accuracy and robustness of feature extraction. This mechanism is particularly effective for obese patients, those with gastroesophageal reflux, and those with upper respiratory tract infections, where these interference factors are more significant.
[0145] like Figure 7As shown, the feature representation module 4 adopts a multi-scale glottal feature pyramid fusion structure, including a microstructure layer 41, a transient dynamic layer 42, a periodic change layer 43, an overall pattern layer 44 and a feature fusion gating unit 45.
[0146] The resolution of the microstructure layer 41 is the original resolution, contains 2 residual blocks, has 64 features, and a receptive field size corresponding to a physical scale of 1-5ms. It focuses on the tiny vibration characteristics of the glottis opening and closing. The resolution of the transient dynamic layer 42 is 1 / 2 of the original resolution, contains 4 residual blocks, has 128 features, and a receptive field size corresponding to a physical scale of 5-20ms. It focuses on the transient process characteristics of the glottis opening and closing. The resolution of the periodic change layer 43 is 1 / 4 of the original resolution, contains 6 residual blocks, has 256 features, and a receptive field size corresponding to a physical scale of 50-200ms. It focuses on the periodic change characteristics of the glottis opening and closing. The resolution of the overall pattern layer 44 is 1 / 8 of the original resolution, contains 8 residual blocks, has 512 features, and a receptive field size corresponding to a physical scale of 500-2000ms. It focuses on the overall pattern characteristics of the glottis opening and closing.
[0147] The residual block structure is designed as follows: input → batch normalization → ReLU activation → 3×3 convolution → batch normalization → ReLU activation → 3×3 convolution → sum with input → output. When the number of input and output channels does not match, a 1×1 convolution is added to achieve channel transformation. Each residual block adds a lightweight attention mechanism, and the weights are learnable parameters. The calculation formula of the residual block is:
[0148] ,
[0149] in: is the output feature of the residual block; is the input feature of the residual block; Represents the residual map, which is implemented by a two-layer convolutional network; is the set of network layer parameters, including convolution kernel weights and batch normalization parameters; + represents element-wise addition. Residual connections allow information to be transferred directly, avoiding the vanishing gradient problem while allowing the network to learn feature differences, which is crucial for capturing subtle changes in glottal features.
[0150] The multi-scale feature pyramid contains bidirectional information flow paths, including top-down and bottom-up paths. The top-down path transfers high-level features to lower layers through upsampling. The upsampling method is nearest neighbor interpolation followed by 3×3 convolution smoothing. The number of channels is adjusted through 1×1 convolution with each downsampling layer. The bottom-up path transfers low-level features to higher layers through downsampling using 3×3 convolution with a stride of 2. The number of channels is adjusted through 1×1 convolution with each upsampling layer. Same-level feature fusion combines features from the upper, lower, and current layers.
[0151] The feature fusion gating unit 45 is based on the Squeeze-and-Excitation structure. The Squeeze operation compresses the feature map into a channel descriptor through global average pooling. The Excitation operation passes through a two-layer fully connected network with an intermediate dimension reduction ratio of 16:1 and GELU activation. It outputs channel-level importance weights to adjust the contribution of each channel feature. The calculation formula is:
[0152] ,
[0153] in: is the channel attention weight vector; and Represents Squeeze and Excitation operations respectively; is the result of global average pooling, that is, the vector composed of the average value of each channel; and is the weight matrix of the fully connected layer, Reduce the number of channels to 1 / 16 of the original, Restore it to the original number of channels; GELU is the Gaussian error linear unit activation function; The sigmoid activation function compresses the output to the range of 0-1.
[0154] The feature fusion gating unit weights the features at each level separately, then performs inter-layer fusion, ultimately outputting a multi-scale fused feature tensor with 512× the number of downsampled frames. In difficult airway identification, features at different time scales contribute differently to determining the degree of difficulty. For example, for difficult airways caused by glottal stenosis, the vibration characteristics of the microstructure layer change more significantly; for difficult airways caused by retroglottal displacement, the opening and closing characteristics of the transient dynamic layer are more discriminative; and for difficult airways caused by anatomical abnormalities, the long-term characteristics of the overall pattern layer are more important. Multi-scale fusion can integrate information from all levels to provide a comprehensive assessment of glottal status.
[0155] This multi-scale fusion architecture comprehensively captures glottal features at various timescales, from microscopic to macroscopic, effectively improving the comprehensiveness and robustness of feature representation. This architecture is particularly effective for capturing complex and difficult airway conditions, such as glottal abnormalities, vocal cord paralysis, or tumors, whose features often manifest at multiple timescales.
[0156] like Figure 8 As shown, the timing modeling module 5 first performs timing dependency modeling through the BiGRU (bidirectional gated recurrent unit) network, then weightedly aggregates key timing features through the temporal attention mechanism, then establishes a glottal state transition graph, and finally models the duration of the glottal state sequence.
[0157] The BiGRU network has a hidden layer dimension of 256, with 128 dimensions for the forward and reverse directions, a total of two layers, and residual connections between layers. The input is a multi-scale fused feature sequence from the feature representation layer, and the output is a context-enhanced feature representation for each time point. The calculation formula of the BiGRU is as follows:
[0158] ,
[0159] ,
[0160] ,
[0161] ,
[0162] in: To update the gate vector, control the degree of retention of the hidden state at the previous moment; To reset the gate vector, control the influence of the previous hidden state on the current candidate hidden state; is the candidate hidden state vector; is the hidden state vector at the current moment; is the hidden state vector at the previous moment; is the input feature vector at the current moment; 、 and is the weight matrix; is the sigmoid function; is the hyperbolic tangent function; represents element-wise multiplication; Represents the concatenation operation on the feature dimension; The BiGRU's bidirectional processing enables the model to consider both past and future information simultaneously, which is crucial for understanding the complete sequence of glottal states. In difficult airway assessment, the transition pattern of glottal states (such as opening and closing frequency and transition speed) is an important diagnostic indicator, and the BiGRU can effectively capture these time-dependent features.
[0163] The temporal attention mechanism first maps features to query, key, and value spaces through a linear projection layer. It then calculates the query-key similarity, divides it by the square root of the scaling factor (8), normalizes it through Softmax to obtain the attention weight, and finally performs a weighted summation on the value vector. Eight attention heads are used, each with a dimension of 16, for a total of 128 dimensions. Residual connections add the original features to the attention output to maintain information fluidity. The calculation formula for multi-head attention is:
[0164] ,
[0165] in: The output of multi-head attention For the The output of an attention head; 、 、 For the The linear projection matrix of the head; is the output projection matrix; Represents the concatenation operation on the feature dimension; The multi-head attention mechanism allows the model to focus on features at multiple time positions simultaneously, which is particularly effective for capturing key moments in glottal activity (such as opening and closing transition points).
[0166] The glottal state transition diagram defines four glottal states: State 1 (fully open), State 2 (partially open), State 3 (partially closed), and State 4 (fully closed). The state recognition method designs a state classifier based on the BiGRU output features. The classifier structure is a two-layer fully connected network (256-128-4), with GELU activation in the hidden layer and Softmax activation in the output layer. It outputs the probability distribution of each state at each time point. In difficult airway assessment, the distribution and transition pattern of glottal states are directly related to the Cormack-Lehane classification. For example, patients in grades C-LIII and IV typically show a significant increase in the proportion of states 3 and 4, and state 1 almost never appears.
[0167] State sequence optimization is smoothed using a conditional random field layer. The state transition matrix is a 4×4 matrix representing the transition probabilities between states and is initialized to the statistical prior probabilities. The emission probability matrix is provided by the classifier output probabilities. The optimal path inference uses the Viterbi algorithm, whose recursive formula is:
[0168] ,
[0169] in: Represents the probability value of the maximum probability path with state j at time t; is the initial state probability, which indicates the probability that the sequence starts at state i; is the state transition probability, which represents the probability of transitioning from state i to state j; Is the emission probability, indicating that state j generates an observation value probability; is the number of states, take 4; is the sequence length; represents a maximum operation; i and j are state indices, ranging from 1 to 4; and t is the time step index. The Viterbi algorithm finds the globally optimal state sequence and avoids locally incoherent state assignments, which is crucial for achieving a smooth and coherent glottal state transition.
[0170] State duration is modeled using a mixed lognormal distribution, with the mean and variance of each state predicted from the feature vector via a neural network. The network architecture consists of a shared feature extraction layer and an independent output layer for each state. The feature extraction layer is a two-layer fully connected network (256-128-64), and the output layer consists of a dual-output layer for each state, predicting the mean and variance, respectively. Abnormally short states (duration less than 20% of the reference value) or states with excessively long duration (exceeding 200% of the reference value) are flagged as risk markers. Abnormal state duration is an important risk indicator in difficult airway assessment. For example, a prolonged duration (>2 seconds) of the fully closed state (state 4) often indicates limited glottal movement and may be a sign of difficult airway.
[0171] Ultimately, the time series modeling module outputs a glottal state sequence and a 320-dimensional time series feature vector, providing a basis for subsequent decision-making and reasoning. These time series features contain complete dynamic information about glottal activity and are key indicators for assessing airway difficulty.
[0172] like Figure 9 As shown, the decision reasoning module 6 first expands the time series feature vector (320 dimensions) into a decision feature vector (512 dimensions) by mapping it through a fully connected layer and concatenating it with the glottal state statistical features. The glottal state statistical features include the proportion of each state, average duration, state transition frequency, etc.
[0173] The decision reasoning module utilizes a dual-branch decision network 61, consisting of a main branch (CL classification) and a secondary branch (difficulty / non-difficulty discrimination). The main branch is a four-layer fully connected network (512-256-128-4) with residual connections added between the first and second layers and between the second and third layers. The intermediate layers use GELU activation, and the output layer uses Softmax activation. It outputs probability distributions for the four CL levels. The secondary branch is a three-layer fully connected network (512-128-32-2) with GELU activation in the intermediate layers and Sigmoid activation in the output layer. It outputs the probability of difficulty / non-difficulty. Difficulty is defined as CL classification of III or IV, and non-difficulty is defined as CL classification of I or II. This dual-branch design enables the model to simultaneously learn the fine-grained classification task and the coarse-grained binary classification task, mutually enhancing performance. In clinical practice, this design can provide more reliable predictions, especially for borderline cases (such as those between Class II and III).
[0174] The Bayesian decision inference layer 62 integrates the main branch CL grade probability distribution and the auxiliary branch difficult / non-difficult probabilities, and combines prior knowledge to make decisions. Prior knowledge includes the distribution probability of each CL grade based on historical hospital data and the conditional probability of CL distribution based on different population characteristics (age, gender, BMI, etc.). For example, in the general population, the typical distribution ratios of C-LI, II, III, and IV are approximately 60%, 30%, 8%, and 2%; in obese patients (BMI>30), this distribution is skewed towards higher grades, approximately 40%, 35%, 20%, and 5%.
[0175] The computational steps of Bayesian decision inference are:
[0176] (1) Adjust the output probability of the main branch based on the prior distribution:
[0177] ,
[0178] in: After adjustment Level CL classification probability; The first branch prediction Level probability; For the Prior probability of level; represents the sum of the four CL levels; the denominator is a normalization factor that ensures that the adjusted probabilities sum to 1. This step combines the network output with the statistical prior, reducing the possibility of extreme predictions and increasing the robustness of the predictions.
[0179] (2) Use the auxiliary branch output to calibrate the main branch prediction:
[0180] ,
[0181] in: After calibration Level probability; The difference between the difficulty / non-difficulty probability predicted for the auxiliary branch and 0.5 indicates the confidence of the prediction; is the calibration coefficient, which is taken as 0.2; The CL index ranges from 1 to 4. The first case applies to low-grade (I or II) airways predicted to be non-difficult airways; the second case applies to high-grade (III or IV) airways predicted to be difficult airways; all other cases retain the original probability. This calibration mechanism can enhance confidence when the predictions of the main and auxiliary branches are consistent. For example, if the main branch is predicted to be CL III and the auxiliary branch is highly confident to be a difficult airway, the probability of Class III will be further increased.
[0182] (3) Calculate the final posterior probability distribution and select the CL level with the highest posterior probability as the final prediction.
[0183] The decision reasoning module also includes a decision confidence assessment mechanism. Confidence is primarily based on the difference between the highest and second-highest probability values, while also taking into account the consistency of the predictions between the primary and secondary branches and the stability assessment of the timing model. The default confidence threshold is 0.85, below which the multi-model ensemble decision mechanism is triggered. Confidence assessment is particularly important in clinical applications because it helps anesthesiologists determine the reliability of predictions. For low-confidence predictions, physicians can be more cautious and be prepared for difficult airway situations.
[0184] The multi-model ensemble decision-making mechanism includes our deep learning model (attention-enhanced residual network), XGBoost, and LightGBM models, all built based on engineered features. The ensemble method employs a stacking strategy, with logistic regression as the meta-learner. The input features are the predicted probability distributions of the three base models, and the output is the final CL classification prediction and the overall confidence score. Multi-model ensembles combine the strengths of different models to improve prediction stability and accuracy, particularly for complex or atypical cases.
[0185] Application Output Module 7 determines the laryngoscope model recommendation based on the CL grading results and confidence level from the decision-making reasoning module. CL grading results are expressed using standard medical grading (I-IV), with the probability of each level appended, accurate to two decimal places. The confidence level is visualized, and a warning indicator is displayed when it falls below the threshold. A red warning appears on the interface when the prediction is level III or IV.
[0186] The laryngoscope blade model recommendation system constructs a table based on the predicted CL grade and confidence: Grade I recommends standard curved laryngoscope blades; Grade II recommends standard curved or slightly curved laryngoscope blades; Grade III recommends curved laryngoscope blades or video laryngoscopes; and Grade IV recommends video laryngoscopes or fiberoptic bronchoscopes. These recommendations are based on the optimal device selection for different airway management difficulties in clinical practice. For example, for C-LIV (only the soft palate can be seen, not the epiglottis and glottis), traditional direct laryngoscopes generally do not provide an adequate field of view, while video laryngoscopes or fiberoptic bronchoscopes can provide better glottic visualization through their curved design and camera function.
[0187] The recommendation strategy is as follows: a single, clear recommendation is given when the prediction is high-confidence (≥0.90); a primary and alternative recommendation is given when the prediction is moderate-confidence (0.85-0.90); and multiple possible options are given with the recommendation for careful evaluation when the prediction is low-confidence (<0.85). For example, when the system predicts a grade C-LIII with a confidence of 0.92, a curved laryngoscope is clearly recommended; however, when the prediction is a grade C-LIII with a confidence of only 0.82, a curved laryngoscope, video laryngoscope, and fiberoptic bronchoscope are recommended, and the anesthesiologist is advised to carefully select the appropriate option based on the patient's specific circumstances.
[0188] In addition, the application output module provides clinical decision support information, including the risk level of difficult airway (low, medium, high), key glottal feature abnormalities, recommended airway management strategy references, and warning messages (such as when high risk is predicted). This information can help anesthesiologists develop more comprehensive airway management plans, especially for patients predicted to have a difficult airway. The system will prompt the possibility of preparing alternative airway equipment, additional personnel support, or special intubation techniques.
[0189] Data flow between modules includes the following aspects:
[0190] The signal acquisition device 1 transmits the original respiratory sound data stream (16 bit / 8 kHz) in real time to the signal preprocessing module 2 via Bluetooth low energy. The trigger mechanism is to automatically start the transmission when the laryngeal mask is placed.
[0191] Signal preprocessing module 2 transmits the wavelet-transformed time-frequency graph to feature extraction module 3. The data dimensions are 3000 × 45 (30 seconds of audio, 80–300 Hz frequency band, 5 Hz resolution). Additional information includes signal quality assessment metrics (signal-to-noise ratio, harmonic-to-noise ratio, and short-term energy stability). The processing mode is segmented, with 3 seconds of data processed at a time and a 1-second sliding window.
[0192] Inside the feature extraction module 3, the time-frequency domain feature adaptive extraction unit 31 transmits a feature map of 64×number of frames×feature dimensions to the feature fusion unit; the cricoid cartilage compression sound and secretion interference sound separation and enhancement unit 32 transmits noise features and suppression suggestions to the feature fusion unit; the feature fusion unit integrates the above inputs and outputs an enhanced glottal feature representation.
[0193] Feature extraction module 3 transmits the fused feature map (64 × number of downsampled frames × feature dimension) to feature representation module 4. The image is processed in batches of 50 frames, with an overlap of 10 frames between batches. Key frames (e.g., state transitions) are prioritized. Areas with significant noise are marked for reference in subsequent processing.
[0194] Within the feature representation module 4, the microstructure layer 41, transient dynamic layer 42, periodic change layer 43, and overall pattern layer 44 interact through bidirectional feature flows. Each layer is cascaded with residual blocks to enhance features layer by layer. The feature fusion gating unit 45 controls the contribution weight of each layer’s features. Finally, a multi-scale fused feature tensor (512 × downsampled frame number) is output.
[0195] Feature representation module 4 transfers the multi-scale fused feature tensor to time series modeling module 5. Sequences are segmented based on semantic integrity to avoid truncation at state transitions. Supplementary information includes activation maps of key features at each scale layer to assist in time series analysis. The transfer mode is full batch processing to ensure time series integrity.
[0196] Inside the timing modeling module 5, the BiGRU network transmits context-enhanced timing features to the temporal attention mechanism; the temporal attention mechanism transmits weighted key timing features to the glottal state transition graph modeling; the glottal state transition graph modeling transmits state sequence and transition features to the state duration modeling; the state duration modeling transmits the distribution parameters of each state duration to the output integration; and the output integration generates a timing modeling feature vector (320 dimensions).
[0197] The time series modeling module 5 transmits the time series modeling feature vector (320 dimensions) and glottal state statistical features to the decision reasoning module 6. This is then expanded to form a decision feature vector (512 dimensions). Quality control includes feature vector normalization and anomaly detection. Confidence estimates are preliminarily assessed based on feature stability.
[0198] Inside the decision reasoning module 6, the decision feature vector is transmitted to the dual-branch decision network 61 for parallel processing; the main branch and auxiliary branch transmit the CL classification probability and difficult / non-difficult judgment to the Bayesian decision reasoning layer 62; the Bayesian decision reasoning layer 62 transmits the posterior probability distribution to the decision confidence assessment; the decision confidence assessment triggers multi-model integration in the case of low confidence; and finally outputs the CL classification result, confidence level and model integration result (if applicable).
[0199] The decision-making and reasoning module 6 transmits the CL classification results, the probability of each level, the confidence level, and the recommendation basis to the application output module 7 in a JSON format. This transmission occurs after the complete analysis is completed. Intermediate status feedback (progress percentage) is provided during the processing.
[0200] The system also includes feedback mechanisms and dynamic optimization. For example, the application output module 7 provides feedback to the signal acquisition device 1 to adjust subsequent acquisition parameters based on the analysis results; the decision reasoning module 6 provides feedback to the feature extraction module 3 to adjust the feature extraction strategy based on confidence feedback; and the time series modeling module 5 provides feedback to the feature representation module 4 to optimize the feature representation weights based on the state recognition results.
[0201] Each module also contains adaptive mechanisms, such as the signal preprocessing module 2 adaptively adjusting the wavelet threshold based on the signal-to-noise ratio; the feature extraction module 3 dynamically adjusting the attention weight based on the cricoid cartilage compression sound detection results; the feature representation module 4 adjusting the multi-scale fusion weight based on feature importance feedback; the time series modeling module 5 dynamically adjusting the CRF parameters based on state transition stability; and the decision reasoning module 6 adaptively updating the Bayesian prior probability based on the prediction history.
[0202] This system adopts a parallel computing architecture, including three-channel parallel processing of the feature extraction layer, parallel computing of the main and auxiliary branch networks, and the use of GPU to accelerate convolution and matrix operations.
[0203] In terms of model quantization and acceleration, the system quantizes model weights from FP32 to INT8, keeping accuracy loss within 0.5%. Convolutional layer optimization uses depthwise separable convolution instead of standard convolution, reducing computational complexity by 80%. Attention mechanism optimization uses sparse attention to reduce computational complexity.
[0204] Memory optimization includes feature map reuse (intermediate feature maps are released in a timely manner), gradient checkpoints (gradient checkpoints are used to reduce memory usage when processing long sequences), and static memory allocation (pre-allocating a fixed-size memory pool to avoid dynamic allocation overhead).
[0205] The system's real-time performance indicators include: end-to-end latency <2 seconds (from completion of 30-second signal acquisition to result output), processing throughput >10 frames / second (on a standard medical tablet), and peak memory usage <2GB.
[0206] The system has the ability to adapt to signal quality. Under low signal-to-noise ratio conditions (SNR<5dB), it increases the weight of time domain features and weakens frequency domain features. When secretion interference is severe (harmonic noise ratio<0.3), it enhances the connection strength of spectrum subtraction residuals. When cricoid cartilage compression is obvious (feature similarity>0.8), it adjusts the state transition probability matrix.
[0207] In terms of adapting to individual patient differences, the system adjusts the glottal frequency focus interval based on gender differences (80-250Hz for men and 100-300Hz for women); enhances the weight of low-frequency features for elderly patients (>65 years old); and adjusts the Bayesian prior probability based on BMI (when BMI>30, the prior probability of grade III / IV increases by 15%).
[0208] The system can also handle abnormal situations, such as completing predictions based on historical fragments when the signal is interrupted; triggering re-acquisition suggestions in the event of extreme noise; and downgrading to basic feature analysis mode when pattern recognition fails.
[0209] Hardware platform requirements include: processor is ARM Cortex-A76 or equivalent performance processor, memory ≥ 4GB RAM, storage ≥ 32GB flash memory, and support for low-power Bluetooth 5.0.
[0210] The software environment includes: operating system Android 10+ or iOS 14+, deep learning framework TensorFlow Lite or ONNX Runtime, and database SQLite (for storing model parameters and reference data).
[0211] The system supports hospital information system integration, including support for HL7 standard interface, compatibility with DICOM standard, provision of REST API interface and support for electronic medical record system data exchange.
[0212] The key performance indicators of the present invention include: recognition accuracy (CL classification accuracy reaches 91.7%), timeliness index (total time from acquisition to result output is <30 seconds), robustness index (accuracy reduction is <5% under common clinical interference), adaptability index (accuracy fluctuation among different populations is <3%) and confidence accuracy (correlation coefficient between confidence and actual accuracy is >0.9).
[0213] In terms of clinical application value, this invention improves safety (the success rate of preoperative prediction of difficult airway is 25% to 30% higher than that of traditional methods, and the incidence of related complications is potentially reduced by 15% to 20%), improves clinical efficiency (increases the success rate of first-time intubation by 10% to 15%, reduces repeated attempts, and shortens anesthesia induction time by 3 to 5 minutes), optimizes medical resources (reduces unnecessary use of advanced equipment, and saves related costs by 10% to 15%), and improves medical quality (standardizes airway assessment procedures to reduce the impact of individual differences among doctors).
[0214] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. An intelligent method for identifying glottal features of difficult airways, characterized by: include: Acquiring the patient's respiratory sound signal, which is collected by a micro-microphone embedded in the laryngeal mask; Preprocessing the respiratory sound signal to obtain a time-frequency graph; Adaptively extracting time-frequency domain features based on the time-frequency graph to obtain glottal features; performing separation and enhancement processing on the glottis feature by performing cricoid cartilage compression sound and secretion interference sound separation and enhancement processing to obtain an enhanced glottis feature; Based on the enhanced glottal features, a multi-scale fusion feature is obtained by fusing the multi-scale glottal feature pyramid; Based on the multi-scale fusion features, the glottal state is analyzed by time series modeling to obtain a glottal state sequence and a time series feature vector; Determining the Cormack-Lehane grading result through decision reasoning based on the time series feature vector and glottal state statistical characteristics; The recommended laryngoscope model was determined based on the Cormack-Lehane classification results.
2. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The miniature microphone is arranged at the second curve position of the laryngeal mask, and the sensitivity of the miniature microphone is -42dBV / Pa, the frequency response is 20Hz-20kHz, and the signal-to-noise ratio is not less than 65dB.
3. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The pretreatment includes: performing noise elimination processing on the respiratory sound signal; Performing frame processing on the respiratory sound signal after noise elimination processing; The breath sound signal after frame processing is band-pass filtered to extract the glottal characteristic frequency band of 80-300 Hz; Performing wavelet transform processing on the extracted glottal characteristic frequency band to obtain the time-frequency graph; The framing process includes: Adopting a framing strategy with a 25ms frame length and a 10ms frame shift; Multiply each frame by a Hamming window to reduce spectral leakage.
4. The method for intelligently identifying glottal features of difficult airways according to claim 3, characterized in that: The wavelet transform process includes: The 9-level discrete wavelet decomposition is adopted, and the base wavelet is Daubechies-4 wavelet; Threshold denoising is performed on level 1-4 coefficients, and the threshold is adaptively calculated based on the local energy distribution of the signal; The reconstructed time-frequency diagram has a resolution of 10ms×5Hz.
5. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The time-frequency domain feature adaptive extraction includes: Features are extracted through time domain channel, frequency domain channel and time-frequency joint channel respectively; Based on the signal quality evaluation index, the features extracted from the time domain channel, the frequency domain channel and the time-frequency joint channel are weightedly fused to obtain the glottal feature; Among them, the time domain channel includes three sets of parallel one-dimensional convolutional networks, and the convolution kernel sizes are 3, 5, and 7 time points respectively; the frequency domain channel includes a deformable convolutional network, and the receptive field can be adaptively adjusted; the time-frequency joint channel includes a four-layer two-dimensional convolutional network, and the convolution kernel size is 3×3.
6. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The separation and enhancement process of the cricoid cartilage compression sound and the secretion interference sound includes: The cricoid cartilage compression sound features and secretion interference sound features are extracted respectively through a dual-path noise feature learning network. Enhance target glottal features and suppress noise features through spectral subtraction residual connection; An adversarial learning strategy is used to optimize the feature extractor so that it generates glottal features that are difficult for the discriminator to distinguish.
7. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The multi-scale glottal feature pyramid fusion includes: inputting the enhanced glottis features into the microstructure layer, the transient dynamic layer, the periodic variation layer and the overall pattern layer for processing respectively; The features of each layer are weighted by a feature fusion gating unit, and inter-layer fusion is performed to obtain the multi-scale fusion feature; Among them, the receptive field of the microstructure layer corresponds to the physical scale of 1-5ms, focusing on the tiny vibration characteristics of the opening and closing of the glottis; the receptive field of the transient dynamic layer corresponds to the physical scale of 5-20ms, focusing on the transient process characteristics of the opening and closing of the glottis; the receptive field of the periodic change layer corresponds to the physical scale of 50-200ms, focusing on the periodic change characteristics of the opening and closing of the glottis; the receptive field of the overall pattern layer corresponds to the physical scale of 500-2000ms, focusing on the overall pattern characteristics of the opening and closing of the glottis; information interaction is achieved between each layer through a bidirectional information flow path, including a top-down path and a bottom-up path.
8. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The timing modeling includes: Extracting temporal dependency features through a bidirectional gated recurrent unit network; Aggregate key temporal features through weighted temporal attention mechanism; A glottal state transition graph is established based on weighted aggregated key temporal features; Performing duration modeling on the glottal state sequence to obtain the time series feature vector; The glottal state transition diagram includes four glottal states: fully open state, partially open state, partially closed state and fully closed state; the state sequence is smoothed by the conditional random field layer, and the Viterbi algorithm is used to find the global optimal state sequence.
9. The method for intelligently identifying glottal features of difficult airways according to claim 1, characterized in that: The decision reasoning includes: Integrating the time series feature vector and glottal state statistical features into a decision feature vector; The Cormack-Lehane grade probability distribution and difficult / non-difficult airway discrimination results were predicted respectively by a two-branch decision network. Integrating the dual-branch prediction results based on Bayesian decision reasoning to obtain the Cormack-Lehane classification result; Among them, the Bayesian decision reasoning includes: using the Cormack-Lehane distribution probability of each level based on the hospital's historical data statistics as prior knowledge; adjusting the main branch output probability based on the conditional probability of the Cormack-Lehane distribution of different population characteristics; calibrating the graded prediction using the difficult / non-difficult airway discrimination results; calculating the final posterior probability distribution, and selecting the Cormack-Lehane level with the highest posterior probability as the final prediction; when the prediction confidence is lower than the threshold, triggering the multi-model integration decision mechanism, including the integration of this deep learning model, XGBoost model and LightGBM model.
10. Difficult airway glottis feature intelligent recognition system, characterized by: include: A signal acquisition device for acquiring a patient's breathing sound signal, the signal acquisition device comprising a miniature microphone embedded in the laryngeal mask; A signal preprocessing module, configured to preprocess the respiratory sound signal to obtain a time-frequency graph; a feature extraction module for adaptively extracting time-frequency domain features based on the time-frequency graph to obtain glottal features, and performing separation and enhancement processing on the glottal features to separate cricoid cartilage compression sounds and secretion interference sounds to obtain enhanced glottal features; A feature representation module is used to obtain a multi-scale fusion feature by fusing a multi-scale glottal feature pyramid based on the enhanced glottal feature; A time series modeling module is used to analyze the glottal state through time series modeling based on the multi-scale fusion features to obtain a glottal state sequence and a time series feature vector; A decision-making and reasoning module, configured to determine a Cormack-Lehane grading result through decision-making and reasoning based on the time series feature vector and the statistical features of the glottal state; An output module is used to determine a laryngoscope model recommendation based on the Cormack-Lehane classification result; Wherein, the feature extraction module includes: The time-frequency domain feature adaptive extraction unit includes a time domain channel, a frequency domain channel, and a time-frequency joint channel, which are used to extract features separately and perform weighted fusion based on signal quality evaluation indicators; The unit for separating and enhancing cricoid cartilage compression sounds from secretion interference sounds includes a dual-path noise feature learning network and a spectrum subtraction residual connection; The feature representation module includes: Microstructure layer, transient dynamic layer, periodic variation layer and overall pattern layer are used to extract features from different time scales; Feature fusion gating unit, used to weight the features of each layer and perform inter-layer fusion; The decision reasoning module includes: A two-branch decision network is used to predict the Cormack-Lehane grade probability distribution and difficult / non-difficult airway discrimination results respectively; The Bayesian decision inference layer is used to integrate the two-branch prediction results to obtain the final Cormack-Lehane classification result.
Citation Information
Cited By
Intelligent analysis and management method for throat postoperative recovery information
CN121416095A