Method and system for predicting difficult airway

By constructing a difficult airway prediction model based on the VGGish model and using the patient's voice data for prediction, the cumbersome, subjectivity and device dependence of existing evaluation methods are solved, and higher prediction accuracy and convenience are achieved.

CN120036764APending Publication Date: 2025-05-27SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358223.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing difficult airway assessment methods are cumbersome, error-prone, strong subjective, limited accuracy, and rely on expensive medical equipment, which limits the popularization and application of technology.

Method used

By collecting the patient's voice data, detecting, screening, preprocessing and feature extraction, a difficult airway prediction model including the VGGish model and a linear layer is constructed, and the Mel spectrum characteristics are predicted to predict the patient's difficult airway condition.

Benefits of technology

It significantly improves the objectivity and accuracy of difficult airway predictions, reduces misjudgment and misjudgment, avoids additional burden on patients and doctors, and reduces dependence on complex equipment and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036764A_ABST
    Figure CN120036764A_ABST
Patent Text Reader

Abstract

The invention discloses a difficult airway prediction method and system. The method comprises the following steps: acquiring voice data of a patient; detecting and screening the collected voice data to obtain clear voice data; preprocessing the clear voice data to obtain preprocessed voice data; performing feature extraction on the preprocessed voice data to obtain Mel spectrum features; constructing a difficult airway prediction model and training the difficult airway prediction model to obtain a trained difficult airway prediction model; the difficult airway prediction model comprises a VGGish model, and a linear layer, a ReLU activation function, a Dropout layer and a linear output layer which are connected in sequence, and the linear layer is connected with the output end of the VGGish model; and predicting the Mel spectrum features by using the trained difficult airway prediction model to obtain a prediction result of the difficult airway of the patient. According to the method, the burden of the patient and the doctor can be reduced, the objectivity and accuracy of predicting the difficult airway of the patient are ensured, and misjudgment and missed judgment are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anesthesiology, and particularly to a method and system for predicting difficult airways. Background Art

[0002] Before anesthetic surgery, it is necessary to evaluate the difficult airway of the patient to ensure the safety of the surgery; among them, the difficult airway refers to the type of airway that encounters difficulties during mask ventilation or tracheal intubation.

[0003] Existing methods for evaluating difficult airways require recording and measuring multiple physiological signs (such as range of neck movement, mouth opening degree, etc.), which is a cumbersome process and prone to errors, increasing the burden on patients and doctors; at the same time, existing methods for evaluating difficult airways rely heavily on the personal experience of doctors, with strong subjectivity and limited accuracy. Especially in complex environments, it is easy to increase the risk of misdiagnosis or missed diagnosis, increasing the surgical risk. In addition, existing methods for evaluating difficult airways also rely on expensive and complex medical equipment, and the use cost of these equipment is high and the environmental requirements are strict, restricting the popularization and application of the technology. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for predicting difficult airways, which can reduce the burden on patients and doctors, ensure the objectivity and accuracy of predicting the difficult airway of patients, and reduce misjudgment and missed judgment.

[0005] To achieve the above purpose, the present invention is realized through the following technical solutions:

[0006] A method for predicting difficult airways, comprising:

[0007] Collecting voice data of the patient;

[0008] Detecting and screening the collected voice data to obtain clear voice data;

[0009] Preprocessing the clear voice data to obtain preprocessed voice data; the preprocessed voice data is in a preset frequency band and has a consistent volume;

[0010] Extracting features from the preprocessed voice data to obtain Mel spectrum features;

[0011] Constructing a difficult airway prediction model and training it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model;

[0012] Use the trained difficult airway prediction model to predict the Mel spectrogram features to obtain the prediction result of the difficult airway of the patient.

[0013] Optionally, the steps of detecting and screening the collected voice data include:

[0014] Use a voice activity detection algorithm to detect the collected voice data to obtain multiple valid voice segments;

[0015] Calculate the signal-to-noise ratio of each of the valid voice segments;

[0016] Select the valid voice segments with a signal-to-noise ratio greater than or equal to a preset threshold from the multiple valid voice segments, and use the selected valid voice segments as the clear voice data.

[0017] Optionally, the range of the preset threshold is 12dB - 18dB.

[0018] Optionally, the steps of preprocessing the clear voice data include:

[0019] Perform filtering on the clear voice data to obtain voice data in the preset frequency band; and the preset frequency band is 300Hz - 3400Hz;

[0020] Perform volume normalization on the voice data in the preset frequency band so that the volume of the voice data in the preset frequency band is consistent to obtain the preprocessed voice data.

[0021] Optionally, before performing the step of using the trained difficult airway prediction model to predict the Mel spectrogram features, it further includes: pruning and quantization processing on the trained difficult airway prediction model.

[0022] On the other hand, the present invention also provides a difficult airway prediction system, including:

[0023] A voice input module, configured to collect voice data of a patient;

[0024] A quality control module, connected to the voice input module, configured to detect and screen the collected voice data to obtain clear voice data;

[0025] A preprocessing module, connected to the quality control module, configured to preprocess the clear voice data to obtain preprocessed voice data; the preprocessed voice data is in a preset frequency band and has a consistent volume;

[0026] A feature extraction module, connected to the preprocessing module, configured to extract features from the preprocessed voice data to obtain Mel spectrogram features;

[0027] A prediction module, connected to the feature extraction module, is configured to construct a difficult airway prediction model and train it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model; the prediction module also uses the trained difficult airway prediction model to predict the Mel spectrogram features to obtain the prediction result of the patient's difficult airway.

[0028] Optionally, the prediction module is further configured to perform pruning processing and quantization processing on the trained difficult airway prediction model; and the quality control module, the preprocessing module, the feature extraction module, and the prediction module are all arranged on an edge computing device.

[0029] Optionally, the quality control module includes:

[0030] A detection unit, connected to the voice input module; the detection unit uses a voice activity detection algorithm to detect the collected voice data to obtain a plurality of effective voice segments;

[0031] A screening unit, connected to the detection unit, is configured to calculate the signal-to-noise ratio of each of the effective voice segments, and screen out the effective voice segments with a signal-to-noise ratio greater than or equal to a preset threshold from the plurality of effective voice segments, and use the screened effective voice segments as the clear voice data.

[0032] Optionally, the preprocessing module includes:

[0033] A filtering unit, connected to the screening unit, is configured to perform filtering processing on the clear voice data to obtain voice data in the preset frequency band; and the preset frequency band is 300Hz - 3400Hz;

[0034] A volume processing unit, connected to the filtering unit and the feature extraction module, is configured to perform volume normalization processing on the voice data in the preset frequency band to make the volumes of the voice data in the preset frequency band consistent, so as to obtain the preprocessed voice data.

[0035] Optionally, the difficult airway prediction system further includes: a display module, connected to the prediction module, for displaying the prediction result of the patient's difficult airway.

[0036] Compared with the prior art, the present invention has at least one of the following advantages:

[0037] The present invention takes the collected voice data of patients as the core data source, analyzes the Mel-spectrum features extracted from the collected voice data through a difficult airway prediction model, significantly improves the objectivity and accuracy of the prediction results, and reduces misjudgment and missed diagnosis. The non-invasive and easy accessibility of the voice data of patients can avoid additional burdens on patients and doctors, enhance the practicability and convenience of the present invention, and make it widely applicable in diverse clinical scenarios.

[0038] In the present invention, the VGGish model is selected as the basic model for the difficult airway prediction model, and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer are additionally added at the output end of the VGGish model, enabling the difficult airway prediction model to more accurately identify voice samples containing noise or large variations, enhancing the discrimination ability for confusing boundaries; at the same time, it also enables the training process of the difficult airway prediction model to achieve better results in the balance of the number of parameters and the amount of calculation, so as to still maintain reliable prediction performance under limited computing resources and effectively reduce the risk of missed diagnosis and misdiagnosis in clinical applications.

[0039] In the present invention, pruning and quantization techniques are combined to further optimize the trained difficult airway prediction model, so as to reduce the size and computational complexity of the trained difficult airway prediction model, thereby being able to efficiently analyze voice features (i.e., Mel-spectrum features) and predict the risk level and probability of the difficult airway of patients under limited computing resources, meeting the requirements for real-time prediction or inference in the clinical environment.

[0040] In the present invention, the quality control module, the preprocessing module, the feature extraction module, and the prediction module in the difficult airway prediction system are set on an edge computing device, and data local processing is realized through edge computing technology, reducing the dependence on the network environment, outputting prediction results in real time and efficiently, and meeting the requirements for real-time and efficiency in the clinical environment; at the same time, the operation efficiency and user experience are improved. Further, the edge computing device has the characteristics of small volume, low cost, and strong performance, is applicable to a variety of clinical scenarios, reduces the dependence on complex devices and environments, and enhances the portability and economy of the difficult airway risk prediction system. Description of the Drawings

[0041] Figure 1 is a flowchart of a difficult airway prediction method provided by an embodiment of the present invention;

[0042] Figure 2 is a result schematic diagram of a difficult airway prediction system provided by an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the graphical user interface of a display module in a difficult airway prediction provided by an embodiment of the present invention. Detailed Implementation Modes

[0044] The following further elaborates on a difficult airway prediction method and system proposed by the present invention in conjunction with the accompanying drawings and specific implementation modes. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the implementation modes of the present invention. In order to make the purpose, features, and advantages of the present invention more obvious and understandable, please refer to the accompanying drawings. It should be noted that the structures, scales, sizes, etc. shown in the drawings of this specification are only used to match the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have any technical substance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed by the present invention.

[0045] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the said element.

[0046] Combined with the attached Figure 1As shown in the figure, this embodiment provides a difficult airway prediction method, including: Step S1, collecting the voice data of the patient. Step S2, detecting and screening the collected voice data to obtain clear voice data. Step S3, preprocessing the clear voice data to obtain preprocessed voice data; the preprocessed voice data is in a preset frequency band and has a consistent volume. Step S4, extracting features from the preprocessed voice data to obtain Mel spectrogram features. Step S5, constructing a difficult airway prediction model and training it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model. Step S6, using the trained difficult airway prediction model to predict the Mel spectrogram features to obtain the prediction result of the difficult airway of the patient. Optionally, the prediction result of the difficult airway of the patient includes a risk level and a prediction probability.

[0047] Specifically, in step S1, before anesthetic surgery, a high-sensitivity recording device can be used to collect the voice data of the patient. Optionally, the recording device is a microphone; the sampling rate of the microphone can be uniformly set to 16000Hz to ensure the quality and consistency of the collected voice data, and at the same time support multi-language and multi-dialect recording. Optionally, the patient can be guided to read a specified sentence or have a free conversation and record it to complete the collection of the voice data of the patient; preferably, ensure that the recording process is clear and interference-free to ensure the quality of the collected voice data.

[0048] Specifically, step S2 includes: Step S21, using a voice activity detection (VAD) algorithm to detect the collected voice data to obtain a plurality of effective voice segments. Step S22, calculating the signal-to-noise ratio (SNR) of each effective voice segment. Step S23, screening out the effective voice segments with a signal-to-noise ratio greater than or equal to a preset threshold from the plurality of effective voice segments, and using the screened effective voice segments as the clear voice data.

[0049] In this embodiment, for the collected voice data, first, the voice activity detection algorithm is used to detect the effective voice segment (i.e., the audio that truly contains voice content), so as to eliminate the silent and obvious background noise segments, ensuring the data quality and reliability of subsequent processing. Then, the signal-to-noise ratio of the detected effective voice segment is calculated to facilitate the quality screening of the effective voice segment. When the signal-to-noise ratio of the effective voice segment is lower than the preset threshold, it is determined that the clarity of the effective voice segment is insufficient and it is eliminated; when the signal-to-noise ratio of the effective voice segment is higher than or equal to the preset threshold, the effective voice segment is used as the clear voice data and subsequent preprocessing is performed. Optionally, the range of the preset threshold is 12dB to 18dB; preferably, the preset threshold is 15dB. As can be seen from the above, in this embodiment, clear voice data can be obtained by detecting and screening the collected voice data, which requires a lower requirement for the acquisition environment of the patient's voice data, so that it is still feasible to collect the patient's voice data in a noisy environment, and further makes the difficult airway prediction method provided in this embodiment not rely on special environmental conditions and can adapt to various clinical scenarios.

[0050] Optionally, the signal-to-noise ratio of the effective voice segment is calculated using the following formula:

[0051] SNR = 10×log 10 (P_signal / P_noise) (1)

[0052] where SNR represents the signal-to-noise ratio of the effective voice segment; P_signal represents the signal power of the effective voice segment; P_noise represents the noise power of the effective voice segment.

[0053] In addition, for different languages, dialects or special voice features, multiple voice analysis algorithms can be selected to obtain the clear voice data to adapt to various medical scenarios and meet the needs of different patients, but the present invention is not limited thereto.

[0054] Specifically, the step S3 includes: Step S31, filtering the clear voice data to obtain the voice data in the preset frequency band; and the preset frequency band is 300Hz to 3400Hz. Step S32, performing volume normalization processing on the voice data in the preset frequency band to make the volume of the voice data in the preset frequency band consistent, so as to obtain the preprocessed voice data.

[0055] In this embodiment, in step S31, a band-pass filter can be used to filter the clear speech data to simultaneously remove high-frequency noise and low-frequency noise in the clear speech data, so as to retain the main speech frequency band in the clear speech data, that is, obtain speech data in the preset frequency band, thereby highlighting speech information and reducing environmental interference. In step S32, by performing volume normalization on the speech data in the preset frequency band, the sound intensity of different speech data can be uniformly adjusted at the amplitude level, so that the volumes of different speech data are consistent, thereby avoiding volume differences caused by different recording devices or different vocal intensities of patients.

[0056] It can be understood that the operation of performing volume normalization on the speech data in the preset frequency band can not only reduce the amplitude difference during subsequent feature extraction, so that features such as the extracted Mel spectrum (Mel spectrum) or MFCC (Mel-scale Frequency Cepstral Coefficients) can be calculated and compared on the same scale; but also improve the consistency and reliability of features such as the Mel spectrum or MFCC in the subsequent training of the difficult airway prediction model, thereby reducing unnecessary deviations caused by inconsistent volumes.

[0057] In addition, in some embodiments, the preprocessed speech data can also be moderately enhanced, such as adding background noise, pitch changes, and speech rate changes, etc., to ensure data diversity during the subsequent training of the difficult airway prediction model, and further improve the robustness and generalization ability of the trained difficult airway prediction model to complex clinical environments, but the present invention is not limited thereto.

[0058] Specifically, in step S4, the mel-spectrogram features can be extracted from the preprocessed speech data by using the melspectrogram function in the audio processing library (Librosa library). The specific process of feature extraction includes: first, frame the preprocessed speech data, that is, set the window length (such as 25 ms) and the step size (such as 10 ms), and calculate the power spectrum; then use the melspectrogram function to map the linear spectrum (i.e., the power spectrum) to the Mel scale to obtain the mel-band energy distribution corresponding to each frame; use the power_to_db function in the Librosa library to convert the mel-band energy distribution corresponding to each frame to a logarithmic scale to enhance the simulation of the human ear's perception characteristics, thereby obtaining the mel-spectrogram features. It can be understood that the mel-spectrogram features are essentially a two-dimensional time-frequency matrix, with the horizontal axis being the time frame and the vertical axis being the Mel band. The mel-spectrogram features extracted from the preprocessed speech data not only retain the key information of the speech data in terms of time and frequency but also lay a good foundation for the trained difficult airway prediction model to capture the feature patterns related to the difficult airway risk, providing effective input, thereby improving the accuracy and robustness of the prediction results. Optionally, in the subsequent application of the difficult airway prediction model, the extracted mel-spectrogram features can be directly used as the input of the model, or aggregated or statistically analyzed in the frame or frequency band dimension as needed to extract more compact one-dimensional or two-dimensional feature vectors and used as the input of the model, but the present invention is not limited thereto.

[0059] Specifically, in step S5, the difficult airway prediction model selects the VGGish model as the basis for feature analysis, that is, as the basic model, and an additional linear layer, ReLU activation function, Dropout layer, and linear output layer are added at the output end of the VGGish model to achieve binary classification prediction.

[0060] It can be understood that, compared with directly using the VGGish model, the added linear layer in the difficult airway prediction model can perform dimensionality reduction or dimensionality increase mapping on the extracted general audio features (i.e., Mel spectrogram features), which is more suitable for the specific data distribution of difficult airway prediction. The ReLU activation function can provide stronger non-linear expression ability to help the model capture more complex speech feature patterns. The Dropout layer effectively reduces the risk of overfitting by randomly setting the outputs of some neurons to zero during training, so that the model has better stability and generalization ability in the real clinical environment. The linear output layer then completes the final binary classification mapping, and outputs the prediction result of the model as the risk level of the difficult airway and the probabilities of directly determining the corresponding level of difficult airway or non-difficult airway. Through this series of additional structures on the VGGish model, the difficult airway prediction model can more accurately identify speech samples containing noise or large variations, enhance the discrimination ability for confusing boundaries; at the same time, it also enables the training process to obtain better results in the balance of the number of parameters and the amount of calculation, so as to still maintain reliable prediction performance in the case of limited computing resources and effectively reduce the risks of missed diagnosis and misdiagnosis in clinical applications.

[0061] More specifically, the process of training the difficult airway prediction model is as follows: First, obtain the historical speech data of the patient and execute steps S2 to S4 to obtain the historical Mel spectrogram features based on the historical speech data of the patient; use the obtained historical Mel spectrogram features as the data set, and divide the data set into a training set, a validation set and a test set according to the quantity ratio of 8:1:1 and input them into the difficult airway prediction model to train the difficult airway prediction model. According to the class distribution of the training set, calculate the class weights to balance the imbalance between positive and negative samples, and use the binary cross-entropy loss function combined with the class weights to handle the class imbalance problem, and use the Adam optimizer to balance the convergence speed and stability. After training, use Youden's Index to select the best prediction threshold to balance the sensitivity and specificity of the model, and comprehensively evaluate the accuracy, precision, recall rate, F1 score and AUC value on the test set to comprehensively measure the performance of the model. In addition, automatic mixed precision training is also adopted to improve the training efficiency, but the present invention is not limited thereto.

[0062] Specifically, before executing the step S6, it also includes: pruning and quantifying the trained difficult airway prediction model to reduce the size and computational complexity of the trained difficult airway prediction model, so as to be able to meet the requirements of real-time prediction or inference in the clinical environment under the condition of limited computing resources.

[0063] More specifically, the trained difficult airway prediction model is further optimized by combining pruning and quantization techniques. The importance of weights is evaluated through pruning, and neurons or channels that have less impact on the prediction results are removed, thereby streamlining the model structure and reducing the number of parameters. Quantization, on the other hand, converts floating-point weights or activation values into a low-precision format, effectively reducing the computational and storage overheads, and significantly improving the prediction or inference efficiency of the model while ensuring the prediction accuracy, so as to meet the requirements of computational resources in terms of real-time performance and resource constraints, and thus efficiently analyze the speech features (i.e., Mel spectrogram features) and predict the risk level and probability of the patient having a difficult airway.

[0064] Specifically, in step S6, the Mel spectrogram features are input into the trained difficult airway prediction model, and the prediction results of the patient's difficult airway can be obtained, thereby quickly and clearly assisting medical staff in making decisions, greatly improving the clinical efficiency, and reducing human intervention at the same time. Optionally, the risk levels in the prediction results include high, medium, and low, but the present invention is not limited thereto.

[0065] Combined with the attached Figure 2 and Figure 3 As shown, based on the same inventive concept, this embodiment also provides a difficult airway prediction system, including: a voice input module, a quality control module, a preprocessing module, a feature extraction module, and a prediction module.

[0066] The voice input module is used to collect the voice data of the patient; optionally, the voice input module is a recording device; preferably, the recording device is a microphone, but the present invention is not limited thereto.

[0067] The quality control module is connected to the voice input module and is used to detect and screen the collected voice data to obtain clear voice data. The preprocessing module is connected to the quality control module and is used to preprocess the clear voice data to obtain preprocessed voice data; the preprocessed voice data is within a preset frequency band and has a consistent volume. The feature extraction module is connected to the preprocessing module and is used to extract features from the preprocessed voice data to obtain Mel spectrogram features. The prediction module is connected to the feature extraction module and is used to construct a difficult airway prediction model and train it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model; the prediction module also uses the trained difficult airway prediction model to predict the Mel spectrogram features to obtain the prediction results of the patient's difficult airway. Optionally, the prediction results of the patient's difficult airway include a risk level and a prediction probability.

[0068] Specifically, the quality control module includes: a detection unit and a screening unit; the detection unit is connected to the voice input module; the detection unit uses a voice activity detection algorithm to detect the collected voice data to obtain a plurality of valid voice segments. The screening unit is connected to the detection unit and is used to calculate the signal-to-noise ratio of each of the valid voice segments, and screen out the valid voice segments with a signal-to-noise ratio greater than or equal to a preset threshold from the plurality of valid voice segments, and use the screened valid voice segments as the clear voice data. Optionally, the range of the preset threshold is 12 dB to 18 dB; preferably, the preset threshold is 15 dB, but the present invention is not limited thereto.

[0069] Specifically, the preprocessing module includes: a filtering unit and a volume processing unit; the filtering unit is connected to the screening unit and is used to perform filtering processing on the clear voice data to obtain voice data in the preset frequency band; and the preset frequency band is 300 Hz to 3400 Hz; the volume processing unit is connected to the filtering unit and the feature extraction module and is used to perform volume normalization processing on the voice data in the preset frequency band to make the volume of the voice data in the preset frequency band consistent, so as to obtain the preprocessed voice data. Optionally, the filtering unit is a band-pass filter, but the present invention is not limited thereto.

[0070] In addition, in some embodiments, the preprocessing module further includes an enhancement unit, which is used to perform appropriate enhancement processing on the preprocessed voice data, such as adding background noise, pitch change, and speech rate change, etc., to ensure the data diversity during the subsequent training of the difficult airway prediction model, and further improve the robustness and generalization ability of the trained difficult airway prediction model to complex clinical environments, but the present invention is not limited thereto.

[0071] Specifically, the feature extraction module can extract Mel spectrogram features from the preprocessed voice data by using the melspectrogram function in the audio processing library (Librosa library). Optionally, in the subsequent application of the difficult airway prediction model, either the Mel spectrogram features extracted by the feature extraction module can be directly used as the input of the model, or the feature processing module can perform aggregation or statistics in the frame or frequency band dimension according to needs to extract a more compact one-dimensional or two-dimensional feature vector and use it as the input of the model; optionally, the feature processing module is connected to the feature extraction module and the prediction module, but the present invention is not limited thereto.

[0072] Specifically, the VGGish model is selected in the prediction module as the basis for feature analysis of the difficult airway prediction model, i.e., as the basic model, and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer are additionally added to the output end of the VGGish model to achieve binary classification prediction. Through this series of additional structures on the VGGish model, the difficult airway prediction model can more accurately identify speech samples containing noise or large variations, enhancing the discriminative ability for confusing boundaries; at the same time, it also enables the training process to obtain better results in the balance of the number of parameters and the amount of computation, so as to maintain reliable prediction performance even under limited computing resources and effectively reduce the risks of missed diagnosis and misdiagnosis in clinical applications. Optionally, the difficult airway prediction model can also adopt convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformers, etc., to adapt to different medical scenarios and meet the needs of different patients, but the present invention is not limited thereto.

[0073] Specifically, the quality control module, the preprocessing module, the feature extraction module, and the prediction module are all set on an edge computing device (such as Jetson Orin Nano, Raspberry Pi, NVIDIA Jetson Xavier, Google Coral) to utilize the GPU acceleration of the edge computing device to perform real-time local speech data processing and difficult airway prediction, ensuring that each speech data processing and difficult airway prediction is completed within a few hundred milliseconds to meet the real-time requirements of clinical practice. Further, the edge computing device has the characteristics of small volume, low cost, and strong performance, is applicable to a variety of clinical scenarios, reduces the dependence on complex devices and environments, and improves the portability and economy of the difficult risk prediction system. In addition, the energy consumption of the edge computing device can be optimized to ensure that the difficult risk prediction system can operate stably for a long time; the stability and reliability of the difficult risk prediction system can also be tested in different clinical environments to ensure its stable and efficient operation under various conditions.

[0074] Optionally, in one embodiment, the prediction module is further configured to perform pruning and quantization processing on the trained difficult airway prediction model to reduce the size and computational complexity of the trained difficult airway prediction model, so as to adapt to the real-time prediction or inference requirements of the edge computing device, and enable the trained difficult airway prediction model to operate efficiently on the edge computing device. Specifically, the trained difficult airway prediction model is further optimized by combining pruning and quantization techniques. The pruning process evaluates the importance of weights and removes neurons or channels that have less impact on the prediction results, thereby streamlining the model structure and reducing the number of parameters. The quantization process converts floating-point weights or activation values into low-precision formats, effectively reducing the computational and storage overhead, and significantly improving the prediction or inference efficiency of the model while ensuring the prediction accuracy, so as to meet the requirements of the edge computing device in terms of real-time performance and resource constraints, and thus efficiently analyze voice features (i.e., Mel spectrogram features) and predict the risk level and probability of the patient's difficult airway. Optionally, the trained difficult airway prediction model can be saved or exported in TorchScript format to ensure its compatibility and portability. Optionally, install the dependent libraries and tools required for the operation of the difficult airway prediction model on the edge computing device to ensure the stability and compatibility of the device. For example, install necessary software such as PyTorch, Librosa, and CUDA drivers to ensure that the model can operate efficiently.

[0075] Specifically, the difficult airway prediction system further includes: a display module, connected to the prediction module, for displaying the prediction result of the patient's difficult airway.

[0076] More specifically, the display module provides operation guidance and result presentation in a simple and intuitive manner through a graphical user interface. Figure 3 FIG. [ID] is a schematic diagram of the graphical user interface of the display module, which consists of five parts: a voice input interface, a processing status display interface, a prediction result display interface, a history record and analysis interface, and a setting and management interface. The layout of the graphical user interface is simple, the functional modules are clearly distinguished, and the operation is simple, which is convenient for patients to operate and medical staff to quickly understand and use the prediction results.

[0077] From Figure 3It can be seen that the voice input interface is provided with a recording button and a recording status indicator. The patient can press the recording button to perform voice input to collect the voice data of the patient. Moreover, the voice input interface is simple and clear, and the operation is easy. The processing status display interface will display the quality control and preprocessing status of the voice data in real time, and intuitively display the quality control progress, preprocessing progress, etc. through methods such as progress bars and status icons, so as to timely feedback the data processing status and ensure that medical staff can intuitively and quickly understand the operation of the difficult airway prediction system. The prediction result display interface can enhance the visual effect by using color coding (e.g., red indicates high risk, yellow indicates medium risk, green indicates low risk) and charts (e.g., bar charts, pie charts), and present the prediction results of the difficult airway of the patient, such as clear difficult airway risk levels (high, medium, low) and prediction probabilities, to quickly and clearly assist medical staff in making decisions. The historical record and analysis interface records the historical prediction results of the patient, provides trend analysis and comparison functions, facilitates medical staff to view and analyze the changes in the airway risk of the patient, and provides long-term monitoring and evaluation support; it supports search and screening functions. The setting and management interface allows authorized medical staff to configure system parameters (such as sampling rate, model type, alarm threshold, etc.) through secure access control, supports personalized configuration and adapts to different clinical needs; it performs data encryption to ensure the privacy and data security of the patient.

[0078] In this embodiment, the voice input module is connected to the edge computing device to ensure that the recorded voice data of the patient can be transmitted to the edge computing device in real time. The quality control module, the preprocessing module, and the feature extraction module on the edge computing device automatically perform quality control, preprocessing, and feature extraction, and automatically input the extracted voice features (i.e., Mel spectrum features) into the trained difficult airway prediction model in the prediction module for difficult airway prediction. Finally, the prediction results, that is, the risk level and prediction probability of the difficult airway, are displayed to medical staff in real time through the simple and intuitive graphical user interface in the display module.

[0079] In some embodiments, in order to ensure the security and privacy protection of patient data, the difficult airway prediction system can introduce blockchain technology to ensure the immutability and traceability of voice data and prediction results, and enhance the transparency and trust of the difficult airway prediction system. At the same time, the difficult airway prediction system adopts multi-level data encryption (including transmission encryption and storage encryption) and access control mechanisms to ensure the security of patient voice data during storage and transmission, meet strict privacy protection requirements, prevent unauthorized access, and comply with the requirements of relevant laws and regulations (such as GDPR, HIPAA, etc.).

[0080] In addition, in some other embodiments, according to the requirements of different medical institutions, the difficult airway prediction system can flexibly add multiple functional modules, such as a monitoring module, an alarm module, a data analysis report generation module, and a patient history record management module, etc. The expansion of these functions can improve the adaptability and comprehensive service ability of the difficult airway prediction system and support personalized configuration. The difficult airway prediction system can also support migrating the processing of voice data and the difficult airway prediction model to a cloud server, and using the powerful computing power of cloud computing for continuous optimization and update, which is applicable to multi-site medical institutions. At the same time, the difficult airway prediction system can also be developed into a mobile application version applicable to smartphones and tablets, enabling medical staff to perform airway risk prediction anytime and anywhere, enhancing the convenience of use, but the present invention is not limited thereto.

[0081] Furthermore, with the accumulation of clinical data, an online learning and continuous learning mechanism can also be introduced into the difficult airway prediction system, which can continuously optimize the difficult airway prediction model according to new data, improve the prediction performance and accuracy, and maintain the long-term effectiveness of the difficult airway prediction system. In large medical institutions or operating rooms, the difficult airway prediction system can support multiple users and devices to work collaboratively, allowing different medical staff to use the difficult airway prediction system simultaneously and view their respective prediction results; this multi-user management function of the difficult airway prediction system improves the efficiency of collaborative work, but the present invention is not limited thereto.

[0082] In summary, a difficult airway prediction method and system provided by this embodiment innovatively uses the collected voice data of patients as the core data source, analyzes the Mel spectrum features extracted from the collected voice data through the difficult airway prediction model, significantly improves the objectivity and accuracy of the prediction results, and reduces misjudgment and missed judgment. The non-invasive and easy accessibility of the voice data of patients can avoid additional burdens on patients and doctors, enhance the practicability and convenience of this embodiment, and make it widely applicable in diverse clinical scenarios. In addition, the difficult airway prediction system integrates functional modules such as voice input, quality control, preprocessing, feature extraction, prediction, and result display, realizes local data processing through edge computing technology, reduces dependence on the network environment, and outputs prediction results in real time and efficiently, overcoming the deficiencies of traditional cloud computing in terms of data transmission delay and privacy protection; at the same time, it improves operation efficiency and user experience. Further, the real-time processing of voice data and difficult airway prediction are realized by using the GPU acceleration of edge computing devices, meeting the requirements for real-time and efficiency in the clinical environment. Compared with expensive traditional hardware devices, the edge computing devices in this embodiment have lower costs and high system integration, reduce the complexity of device coordination, and have high economic benefits and promotion value. The difficult airway prediction system has a high degree of modularity and flexibility, can be functionally extended according to different clinical needs, such as supporting alarm and remote monitoring, and has wide applicability and good scalability. Generally speaking, this embodiment not only solves the problems of accuracy, subjectivity, and environmental dependence in traditional methods, but also has higher clinical application value and broad promotion prospects.

[0083] Although the content of the present invention has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art have read the above content, various modifications and substitutions to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.

Claims

1. A difficult airway prediction method, characterized in that: include: Collect patient voice data; Detect and screen the collected voice data to obtain clear voice data; Preprocessing the clear speech data to obtain preprocessed speech data; The preprocessed voice data is in a preset frequency band and has a consistent volume; Performing feature extraction on the preprocessed speech data to obtain Mel spectrum features; Constructing a difficult airway prediction model and training it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer, and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model; The trained difficult airway prediction model is used to predict the Mel spectrum feature to obtain a prediction result of the patient's difficult airway.

2. The difficult airway prediction method according to claim 1, characterized in that: The steps of detecting and screening the collected voice data include: Using a voice activity detection algorithm to detect the collected voice data to obtain multiple valid voice segments; Calculating the signal-to-noise ratio of each valid speech segment; The valid speech segments having a signal-to-noise ratio greater than or equal to a preset threshold are screened out from the multiple valid speech segments, and the screened out valid speech segments are used as the clear speech data.

3. The difficult airway prediction method according to claim 2, characterized in that: The preset threshold value ranges from 12dB to 18dB.

4. The difficult airway prediction method according to claim 1, characterized in that: The step of preprocessing the clear speech data comprises: Performing filtering processing on the clear voice data to obtain voice data in the preset frequency band; and the preset frequency band is 300 Hz to 3400 Hz; Volume normalization processing is performed on the voice data in the preset frequency band so that the volume of the voice data in the preset frequency band is consistent, so as to obtain the preprocessed voice data.

5. The difficult airway prediction method according to claim 1, characterized in that: Before executing the step of predicting the Mel spectrum feature using the trained difficult airway prediction model, the method further includes: performing pruning and quantization processing on the trained difficult airway prediction model.

6. A difficult airway prediction system, characterized in that: include: Voice input module, used to collect the patient's voice data; A quality control module, connected to the voice input module, for detecting and screening the collected voice data to obtain clear voice data; A preprocessing module, connected to the quality control module, for preprocessing the clear speech data to obtain preprocessed speech data; The preprocessed voice data is in a preset frequency band and has a consistent volume; A feature extraction module, connected to the preprocessing module, for extracting features from the preprocessed speech data to obtain Mel spectrum features; A prediction module is connected to the feature extraction module, and is used to construct a difficult airway prediction model and train it to obtain the trained difficult airway prediction model; the difficult airway prediction model includes a VGGish model and a linear layer, a ReLU activation function, a Dropout layer and a linear output layer connected in sequence, and the linear layer is connected to the output end of the VGGish model; the prediction module also uses the trained difficult airway prediction model to predict the Mel spectrum feature to obtain a prediction result of the patient's difficult airway.

7. The difficult airway prediction system according to claim 6, characterized in that: The prediction module is also used to perform pruning and quantization processing on the trained difficult airway prediction model; and the quality control module, the preprocessing module, the feature extraction module and the prediction module are all arranged on an edge computing device.

8. The difficult airway prediction system according to claim 6, characterized in that: The quality control module comprises: A detection unit connected to the voice input module; the detection unit detects the collected voice data using a voice activity detection algorithm to obtain a plurality of valid voice segments; A screening unit is connected to the detection unit and is used to calculate the signal-to-noise ratio of each of the valid voice segments, and to screen out the valid voice segments whose signal-to-noise ratio is greater than or equal to a preset threshold from the multiple valid voice segments, and use the screened valid voice segments as the clear voice data.

9. The difficult airway prediction system according to claim 8, characterized in that: The preprocessing module comprises: A filtering unit connected to the screening unit, configured to filter the clear voice data to obtain voice data in the preset frequency band; and the preset frequency band is 300 Hz to 3400 Hz; The volume processing unit is connected to the filtering unit and the feature extraction module, and is used to perform volume normalization processing on the voice data in the preset frequency band so that the volume of the voice data in the preset frequency band is consistent to obtain the preprocessed voice data.

10. The difficult airway prediction system according to claim 6, characterized in that: Also includes: A display module is connected to the prediction module and is used to display the prediction result of the patient's difficult airway.