Infant suffocation risk autonomous monitoring method and system based on laryngeal breathing sound

The infant asphyxiation risk monitoring system, which combines a MEMS microphone and a miniature accelerometer, utilizes a machine learning model to achieve real-time analysis of infants' throat breathing sounds. This solves the problem of autonomous monitoring of infant asphyxiation risk in the home environment and improves the accuracy of identification and early warning.

CN120899225AActive Publication Date: 2025-11-07XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

Patent Information

Application Number
CN202511436101.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing technologies lack continuous, non-invasive, and low-cost home monitoring methods, making it difficult to autonomously identify and warn of infant suffocation risks in a home environment, especially due to insufficient real-time analysis algorithms for laryngeal breathing sounds.

Method used

By using a high-sensitivity MEMS microphone to collect airflow and respiratory sound signals from an infant's throat, and combining this with a miniature accelerometer to monitor chest movement data, a suffocation anomaly classification model and a dynamic baseline model are constructed through machine learning models to achieve real-time monitoring and early warning of infant suffocation risk.

Benefits of technology

It achieves high-precision identification and early warning of infant suffocation risk in home environments, reduces environmental noise interference, improves the accuracy of early suffocation identification, and is suitable for home monitoring devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120899225A_ABST
    Figure CN120899225A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of health monitoring, in particular to an infant suffocation risk autonomous monitoring method and system based on laryngeal breathing sound, and the method comprises the steps: collecting airflow and breathing sound signals of the throat of an infant through a high-sensitivity MEMS microphone, and synchronously monitoring thoracic motion data of the infant through a miniature accelerometer; performing noise reduction processing on the airflow and the breathing sound signal to obtain a denoised breathing sound signal; performing signal separation on the denoised breathing sound signals according to a breathing event, and enhancing and combining the denoised breathing sound signals into a breathing sound data set; extracting breath sound features from the breath sound data set, constructing a suffocation anomaly classification model based on the breath sound features through a machine learning model, and establishing a dynamic baseline model adapted to individual breath differences according to the breath sound features; hardware devices of a high-sensitivity MEMS microphone and a micro accelerometer are combined with an artificial intelligence model, and infant asphyxia risk identification and early warning are autonomously completed in a home environment by using laryngeal respiration sound signals.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of health monitoring, in particular to a self-monitoring method and system for infant asphyxia risk based on laryngeal respiratory sound. BACKGROUND

[0002] Infant asphyxia is the main inducement of sudden infant death syndrome (SIDS) and acute accidental events, especially during sleep or feeding, the risk of asphyxia caused by choking, respiratory obstruction or improper body position is extremely high. For these high-risk events of asphyxia, there is currently a lack of continuous, non-invasive and low-cost home monitoring methods. Although hospital-level polysomnography is accurate, it is not suitable for routine use in home settings due to its complex equipment and numerous electrodes; wearable heart rate and oxygen saturation solutions are susceptible to motion and skin temperature interference, resulting in high false positive rates. In recent years, intelligent cameras based on crying or body movement have emerged, but they are not sensitive to silent airway obstruction and have significant privacy concerns. Laryngeal respiratory sound directly reflects the patency of the upper airway, and its spectral characteristics appear wheezing, weakening or disappearing in the early stage of asphyxia, but existing stethoscopes and acoustic patches lack real-time analysis algorithms for infants, making it difficult to achieve closed-loop interaction with care terminals. Therefore, there is an urgent need for a miniature system and method that can autonomously identify and warn of infant asphyxia risk using laryngeal respiratory sound signals in a home environment. SUMMARY

[0003] The present application aims to provide a self-monitoring method for infant asphyxia risk based on laryngeal respiratory sound, to solve the technical problem that existing technologies lack real-time analysis algorithms for infants, making it difficult to autonomously identify and warn of infant asphyxia risk in a home environment.

[0004] To solve the above technical problems, the present application specifically provides the following technical solutions: A self-monitoring method for infant asphyxia risk based on laryngeal respiratory sound, comprising the following steps: Collecting airflow and respiratory sound signals from the infant's larynx using a high-sensitivity MEMS microphone, and synchronously monitoring thoracic motion data of the infant using a miniature accelerometer; Integrating the airflow and respiratory sound signals with the thoracic motion data to perform noise reduction processing on the airflow and respiratory sound signals, obtaining denoised respiratory sound signals; Signal separating the denoised respiratory sound signals according to respiratory events, and enhancing and combining them into a respiratory sound data set; Extracting respiratory sound features from the respiratory sound data set, constructing an asphyxia anomaly classification model that distinguishes between normal, low-risk and high-risk conditions based on the respiratory sound features using a machine learning model, and establishing a dynamic baseline model that adapts to individual respiratory differences based on the respiratory sound features; The dynamic baseline model is used for monitoring and early warning of breathing abnormalities of the infant, and the asphyxia abnormality classification model is used for classifying the asphyxia risk of the breathing abnormalities fed back by the dynamic baseline model.

[0005] As a preferred scheme of the present application, the method for integrating the airflow and breath sound signals and the thoracic movement data and performing noise reduction processing on the airflow and breath sound signals comprises: The airflow and breath sound signals are decomposed into components, and the airflow and breath sound signals are: wherein, is the airflow and breath sound signal, is the target breath sound signal, is is the noise signal generated by body movement in the airflow and breath sound signal, is is the environmental noise signal in the airflow and breath sound signal; The thoracic movement data are decomposed into components, and the thoracic movement data are: wherein, is the thoracic movement data, is is the noise signal generated by body movement in the thoracic movement data, is is the environmental noise signal in the thoracic movement data; The iterative operation of the least mean square algorithm LMS makes the adaptive filter output noise estimation , and the target breath sound signal is calculated according to the noise estimation, wherein, is the noise estimation value, is the set of thoracic movement data, , is the thoracic movement data of the previous moment of is the thoracic movement data of the previous L-1 moments of is the filter weight, is the target breath sound signal; The target breath sound signal is taken as the de-noised breath sound signal. As a preferred scheme of the present application, the method for separating the de-noised breath sound signal according to the breathing event comprises: The de-noised breath sound signal is framed into a plurality of audio samples according to a preset time length, and the audio samples that are less than the preset time length are length-completed;

[0006] The short-time energy STE and the short-time zero-crossing rate ZCR of each audio sample are calculated, and the audio samples without sound are removed based on the short-time energy STE and the short-time zero-crossing rate ZCR through a double-threshold method; ​​In the remaining audio samples, the audio samples corresponding to the breathing events of inhalation, exhalation, crying, coughing and abnormal breath sounds are reserved as valuable samples based on the short-time energy (STE) and the short-time zero-crossing rate (ZCR).

[0007] As a preferred scheme of the present application, the method for combining the enhanced valuable samples into a breath sound dataset comprises: Time domain displacement: shift the valuable samples backward by 1s on the time axis, and splice the 7.0s to 8.0s part of the valuable samples to the front end of the valuable samples; Gain adjustment: adjust the amplitude of each valuable sample to be lower or higher, wherein the oscillation parameter of the amplitude changed by the adjustment is a random value between 0.1 and 5.0; Noise introduction: add a random noise signal to the valuable samples, and adjust the signal-to-noise ratio to a preset value; Mix the valuable samples and the enhanced valuable samples as the breath sound dataset.

[0008] As a preferred scheme of the present application, the method for extracting breath sound features comprises: The samples in the breath sound dataset are sequentially subjected to pre-emphasis, framing, windowing, fast Fourier transform, mel filter bank, and logarithmic energy processing to obtain a mel spectrum graph, and the mel spectrum graph is subjected to DCT transform processing to obtain MFCC features; The samples in the breath sound dataset are estimated for fundamental frequency features by the cepstrum method, and the harmonic energy features are calculated based on the fundamental frequency features; The samples in the breath sound dataset are sequentially calculated for exhalation duration and amplitude, and inhalation duration and amplitude; The MFCC features, the fundamental frequency features, the harmonic energy features, the exhalation duration and amplitude, and the inhalation duration and amplitude are taken as breath sound features.

[0009] As a preferred scheme of the present application, the method for constructing a suffocation anomaly classification model comprises: The samples in the breath sound dataset are type-labeled according to normal, low risk and high risk, and the labeled breath sound dataset is randomly divided into a training set and a test set; In the training set, the breath sound features are taken as input, and the normal, low risk and high risk are taken as output to train a CNN algorithm architecture to obtain the suffocation anomaly classification model; The CNN6 algorithm architecture comprises 4 convolution blocks (ConvBlock), 1 feature aggregation layer and 1 fully connected classification layer, and each convolution block adopts a serial combination of a convolution layer (Conv2D) and a max pooling layer (MaxPooling2D); The 4 convolution blocks are used to gradually extract deep features in the breath sound features; The feature aggregation layer is used to integrate deep features to form a global semantic representation; The fully connected classification layer is used to output classification results of normal, low risk and high risk.

[0010] As a preferred scheme of the present application, the method for constructing the dynamic baseline model comprises: selecting a normal breath sound signal of a specified duration from the denoised breath sound signal as a learning period signal; obtaining the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration duration and amplitude, the inspiration duration and amplitude of the learning period signal, and calculating the mean and standard deviation of each feature within the specified duration; establishing a baseline model for each feature according to the Gaussian distribution of the feature within the specified duration , and combining as the dynamic baseline model , wherein, is the baseline model of feature z, is the mean of feature z, is the standard deviation of feature z.

[0011] As a preferred scheme of the present application, the method for monitoring and warning the respiratory abnormalities of infants by using the dynamic baseline model and classifying the respiratory abnormalities fed back by the dynamic baseline model by using the asphyxia abnormality classification model comprises: monitoring the deviation degree of the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration duration and amplitude, the inspiration duration and amplitude of the real-time denoised breath sound signal of the infant by using the dynamic baseline model, wherein: when the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration duration and amplitude, the inspiration duration and amplitude of the real-time denoised breath sound signal of the infant deviate from the dynamic baseline model by more than a preset threshold, it is determined that the infant is in a respiratory abnormal state; when the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration duration and amplitude, the inspiration duration and amplitude of the real-time denoised breath sound signal of the infant deviate from the dynamic baseline model by less than a preset threshold, it is determined that the infant is not in a respiratory abnormal state; inputting the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration duration and amplitude, the inspiration duration and amplitude of the denoised breath sound signal corresponding to the infant in the respiratory abnormal state into the asphyxia abnormality classification model to determine the asphyxia risk category of the infant.

[0012] As a preferred scheme of the present application, the present application provides an infant asphyxia risk self-monitoring system based on throat breath sound, which is applied to an infant asphyxia risk self-monitoring method based on throat breath sound, comprising: The data acquisition unit comprises a high-sensitivity MEMS microphone and a micro accelerometer, wherein the high-sensitivity MEMS microphone is used to collect airflow and breathing sound signals of the throat of the infant, and the micro accelerometer is used to synchronously monitor thoracic movement data of the infant. The signal noise reduction unit is used to reduce noise of the airflow and breathing sound signals by integrating the airflow and breathing sound signals and the thoracic movement data, so as to obtain denoised breathing sound signals. The machine learning unit is used to separate normal breathing sound signals and abnormal breathing sound signals in the denoised breathing sound signals, and combine the normal breathing sound signals and the abnormal breathing sound signals into a breathing sound data set; breathing sound features are extracted from the breathing sound data set, an abnormal suffocation classification model for distinguishing normal, low-risk and high-risk is constructed based on the breathing sound features through a machine learning model, and a dynamic baseline model adapted to individual breathing differences is established according to the breathing sound features of the normal breathing sound signals. The autonomous monitoring unit is used to monitor and warn the breathing abnormalities of the infant by using the dynamic baseline model, and classify the breathing abnormalities fed back by the dynamic baseline model through the abnormal suffocation classification model.

[0013] As a preferred scheme of the present application, the present application provides a computer readable storage medium, wherein computer execution instructions are stored in the computer readable storage medium, and when a processor executes the computer execution instructions, a method for autonomously monitoring suffocation risk of an infant based on throat breathing sound is realized.

[0014] Compared with the prior art, the present application has the following beneficial effects: The present application directly collects throat breathing sound by using a high-sensitivity MEMS microphone, and reduces noise interference and improves suffocation early identification accuracy through noise reduction processing by a micro accelerometer. The embedded AI algorithm formed by the machine learning model and the dynamic baseline model can effectively distinguish normal breathing variation precursors and suffocation categories, and realize real-time analysis of the infant. The present application combines hardware devices such as a high-sensitivity MEMS microphone and a micro accelerometer with an artificial intelligence model, and realizes autonomous suffocation risk identification and warning of an infant in a home environment by using throat breathing sound signals. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.

[0016] Figure 1 A flow chart of a method for autonomous monitoring of suffocation risk of an infant based on laryngeal breath sounds is provided for an embodiment of the present application. Figure 2 A block diagram of a system for autonomous monitoring of suffocation risk of an infant based on laryngeal breath sounds is provided for an embodiment of the present application. Figure 3 A general architecture diagram of a suffocation risk classification model is provided for an embodiment of the present application. Figure 4 A loss value change graph and an accuracy rate change curve graph are provided for an embodiment of the present application. Figure 5 A performance index graph of a suffocation risk classification model is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0018] As shown in Figure 1 The present application provides a method for autonomous monitoring of suffocation risk of an infant based on laryngeal breath sounds, comprising the following steps: A high-sensitivity MEMS microphone is used to collect airflow and breath sound signals of the infant's larynx, and a miniature accelerometer is used to synchronously monitor thoracic movement data of the infant; The airflow and breath sound signals and the thoracic movement data are integrated to perform noise reduction processing on the airflow and breath sound signals, to obtain denoised breath sound signals; The denoised breath sound signals are separated according to respiratory events, and are enhanced and combined into a breath sound data set; Respiratory sound features are extracted from the breath sound data set, a suffocation anomaly classification model for distinguishing normal, low-risk and high-risk is constructed based on the respiratory sound features through a machine learning model, and a dynamic baseline model adapted to individual respiratory differences is established according to the respiratory sound features; The respiratory anomaly of the infant is monitored and warned by using the dynamic baseline model, and the respiratory anomaly fed back by the dynamic baseline model is classified for suffocation risk by using the suffocation anomaly classification model.

[0019] This invention acquires airflow and respiratory sound signals from an infant's throat using a high-sensitivity MEMS microphone. Since the microphone receives noise from body movement and ambient noise in addition to respiratory sound data, this invention also incorporates a miniature accelerometer. The miniature accelerometer receives only noise from body movement (which originates from the same physical motion as the noise received by the microphone and is highly correlated) and ambient noise, barely perceiving respiratory sounds. An algorithm integrates the data received by the miniature accelerometer with the data received by the high-sensitivity MEMS microphone, retaining only respiratory sound data to achieve noise reduction. The specific details are as follows: Methods for noise reduction of airflow and respiratory sound signals by integrating airflow and respiratory sound signals with chest movement data include: The airflow and breath sound signals are decomposed into their constituent parts, and the airflow and breath sound signals are as follows: ,in, These are airflow and breath sound signals. For the target breath sound signal, for Noise signals generated by body movement, for Environmental noise signals in; The thoracic motion data is decomposed into components, and the thoracic motion data is as follows: ,in, For chest movement data, for Noise signals generated by body movement, for Environmental noise signals in; The adaptive filter output noise is estimated by iterative computation using the Least Mean Square (LMS) algorithm. The target breath sound signal was calculated based on noise estimation. ,in, This is a noise estimate. A collection of chest cavity movement data. , for Chest movement data from the previous moment, for Thoracic motion data for the first L-1 time points, For filter weights, The target breath sound signal is T, which is the transpose operator. target breath sound signal As a noise-reduced breath sound signal.

[0020] The mathematical objective of the LSM algorithm is to make the mean square value E[e] of the output error signal e(k) equal to the mean square value of the output error signal e(k). 2(k)] Minimization. By continuously adjusting the filter weights w(k), y(k) becomes the best estimate of the noise component n(k) in d(k). When y(k) ≈ n(k), e(k) ≈ s(k), the denoising purpose is achieved.

[0021] First, initialize the adaptive filter: Set the filter length (order) L (e.g., L = 32 or 64), which determines the complexity of the noise channel that the algorithm can model.

[0022] Initialize the filter weight vector w, usually set to an all-zero vector: w(0) = [0, 0,..., 0] T .

[0023] Choose a step size parameter μ: Where μ is too large: the algorithm converges quickly, but is unstable, and the error will oscillate or even diverge.

[0024] μ is too small: the algorithm is stable, but converges slowly and may not be able to track changing noise.

[0025] Therefore, the present invention sets the range of μ as: .

[0026] Then, for each time k, perform the following iterative calculation: Step 1: Construct the delay input vector: x_vector(k) = [x(k), x(k-1),..., x(k-L+1)] T This vector is a collection of the current and past L-1 accelerometer signals.

[0027] Step 2: Calculate the output of the adaptive filter, y(k) = w(k) T *x_vector(k) = w(k) This y(k) is the noise estimate predicted by the algorithm based on the current weight w(k) and the accelerometer signal x(k).

[0028] Step 3: Calculate the error signal (i.e., the denoised output) e(k) = d(k) - y(k), which is the final purified target respiratory sound signal we want to obtain. Moreover, if y(k) perfectly estimates the noise n(k), then e(k) = s(k).

[0029] Step 4: Update the filter weights (LMS core), w(k+1) = w(k) + μ * e(k) * x_vector(k).

[0030] The above is the learning process of the LMS algorithm, which actually means: e(k): error measure, if e(k) is large, it means that the current estimate y(k) is poor, and the weight needs to be adjusted with great force; x_vector(k): update direction, the weight will be adjusted along the direction of the current input signal; mu: learning step, control the intensity of adjustment.

[0031] The method for separating the de-noised respiratory sound signal according to the respiratory event comprises.

[0032] The de-noised respiratory sound signal is framed into a plurality of audio samples according to a preset time length, and the audio sample less than the preset time length is time length filled; The short-time energy (STE) and the short-time zero-crossing rate (ZCR) of each audio sample are calculated, and the audio sample without sound is removed through a double-threshold method based on the short-time energy (STE) and the short-time zero-crossing rate (ZCR). In the remaining audio samples, the audio sample corresponding to the respiratory event of inhalation, exhalation, crying, coughing and abnormal respiratory sound is identified based on the short-time energy (STE) and the short-time zero-crossing rate (ZCR) and reserved as a valuable sample.

[0033] After obtaining the de-noised respiratory sound signal (i.e. the de-noised respiratory sound signal), the present application further separates it to obtain meaningful or valuable signals in the respiratory sound signal, such as valuable audio signals of exhalation, inhalation, baby voice, coughing, crying, abnormal respiratory sound and the like, and removes audio signals without sound and without meaning.

[0034] Among them, the short-time energy (Short-Time Energy, STE): measure the amplitude of a signal in a frame. The energy of events such as exhalation and crying is high, followed by inhalation, and the energy of the silent section is low.

[0035] Short-time zero-crossing rate (Short-Time Zero-Crossing Rate, ZCR): measure the number of times a signal crosses zero in a frame. High-frequency sound (such as some crying and friction sound) has a high zero-crossing rate, and low-frequency sound (such as exhalation) has a low zero-crossing rate.

[0036] By setting a low energy threshold and a high energy threshold.

[0037] If the short-time energy (STE) is greater than the high energy threshold, it is determined that there is sound frequency; If the short-time energy (STE) is less than the low energy threshold, it is determined that there is no sound frequency; If the short-time energy (STE) is between the high energy threshold and the low energy threshold, the ZCR is combined to determine: if the ZCR is lower than the preset ZCR threshold (indicating low-frequency sound), it is determined that there is a sound frame, otherwise it is determined that there is no sound frame.

[0038] The determination method of the valuable audio signal is as follows: Cough signal: very short duration, very high STE, explosive waveform (fast rise, slow decay), wide frequency spectrum (MFCC features different from breathing).

[0039] Crying / speech signal: long duration, stable fundamental frequency and harmonic structure, moderate zero-crossing rate.

[0040] Exhalation signal: moderate duration (e.g. 1-2s), high and stable energy, spectrum concentrated in low frequency, low zero-crossing rate, no stable fundamental frequency.

[0041] Inhalation: short duration (e.g. 0.5-1s), lower energy than exhalation, often accompanied by high-frequency friction sound (leading to an increase in zero-crossing rate).

[0042] The basis of model training is data. Since the data is collected from real environment, the research object is infant, and the audio of suffocation condition needs to be collected, the final sample size is bound to be small. Insufficient sample size will cause the suffocation risk classification model to overfit, making it difficult to learn effective features. In order to make the most of existing data resources, the data length should be reasonably cut first. In order to make each cut sample contain a complete breathing cycle as much as possible, the length of each cut segment is set to 8.0s. The audio samples less than 8s are aligned to 8s using zero padding.

[0043] The present application further obtains a more abundant data set through data enhancement technology. The model trained based on the data set expanded by the technology can reduce the risk of overfitting in feature learning, and finally will show more stable classification accuracy and stronger environmental adaptability. Data enhancement is to construct semantically equivalent derivative samples by adding controllable perturbations without destroying the characteristics of the original data, so as to expand the size of the data set. This method enhances the generalization ability of the model by improving the diversity of the data, while keeping the basic features of the data unchanged, as follows: The method for combining value type samples into a breathing sound data set comprises: Time domain displacement: shift the value type sample backward by 1s on the time axis, and splice the 7.0s to 8.0s part of the value type sample to the front end of the value type sample; Gain adjustment: lower or raise the amplitude of each value type sample, wherein the amplitude change of the oscillation parameter is a random value between 0.1 and 5.0; Introducing noise: add a random noise signal to the value type sample, and adjust the signal-to-noise ratio to a preset value; Mixing the value type sample and the enhanced value type sample as a breathing sound data set.

[0044] The respiratory sound dataset is randomly divided in a ratio of 8:2, 80% of which is used as a training set for the training of the airway sputum aggregation degree classification model, and the remaining 20% is used as an independent test set to evaluate the performance of the model.

[0045] Although data augmentation effectively improves the diversity of training samples, it also brings a significant computational burden. In order to balance the model performance and computational efficiency, feature extraction needs to be performed on the augmented audio data.

[0046] The respiratory sound feature extraction method comprises: The samples in the respiratory sound dataset are sequentially pre-emphasized, framed, windowed, fast Fourier transformed, Mel filter banked, and log energy processed to obtain a Mel spectrogram, and the Mel spectrogram is DCT transformed to obtain MFCC features; The samples in the respiratory sound dataset are estimated by the cepstrum method to obtain fundamental frequency features, and the harmonic energy features are calculated based on the fundamental frequency features; The samples in the respiratory sound dataset are sequentially calculated for exhalation duration and amplitude, inhalation duration and amplitude; The MFCC features, fundamental frequency features, harmonic energy features, exhalation duration and amplitude, and inhalation duration and amplitude are used as respiratory sound features.

[0047] The feature extraction technology extracts the key features related to the original signal to realize the signal representation by analyzing the original signal, and plays a key role in the medical classification and recognition task. The method can effectively remove signal redundancy and convert high-dimensional data into low-dimensional representation through feature dimension reduction, which simplifies the calculation and improves the recognition performance. The present application adopts Mel-scale Frequency Cepstral Coefficients (MFCC) as one of the extracted features. MFCC is a widely used feature extraction method in speech and audio processing, which does not need to analyze the characteristics of the original signal itself. Compared with traditional features such as short-time energy, time-domain envelope, autocorrelation coefficient, zero-crossing rate, and frequency-domain features such as spectral envelope, spectral flux, and spectral bandwidth, it can better distinguish the subtle differences between different audios, can resist the influence of background noise to a certain extent, and has better robustness. Before extracting MFCC, we first generate a Mel spectrogram, which more intuitively shows the time-frequency structure of the audio signal and is more helpful in capturing acoustic features related to human hearing. The process of extracting MFCC features and Mel spectrogram in the present application is as follows: The audio samples in the expanded training set are sequentially pre-emphasized, framed, windowed, fast Fourier transformed, Mel filter banked, and log energy processed to obtain a Mel spectrogram; The Mel spectrogram is DCT transformed to obtain MFCC features.

[0048] Pre-emphasis: The first step in the MFCC feature extraction process, its core role is to optimize the spectral characteristics of the speech signal through high-frequency enhancement. Framing: The spectral characteristics of the audio signal change over time, in order to accurately capture the frequency characteristics of different time periods, the pre-emphasized signal needs to be framed to ensure that the spectral distribution within each frame remains relatively stable. Windowing: Using the Hamming window to suppress the sidelobes to improve the continuity of adjacent frames. Short-time Fourier transform: Convert time-domain signals into time-frequency joint representation, and reveal the law of local frequency domain characteristics of signals evolving over time through the energy distribution of time-varying spectrum. Mel filter bank: Its essence is a nonlinear frequency band analysis system that simulates the hearing characteristics of the human ear, composed of a set of overlapping triangular bandpass filters. The filter bank is non-uniformly distributed under the Mel scale, with a dense low-frequency region and a sparse high-frequency region, which conforms to the nonlinear perception characteristics of human ear to pitch. Its role includes: ①Smooth the spectrum and suppress harmonics; ②Enhance the formant characteristics; ③Improve the low-frequency resolution. This processing makes the spectral features more consistent with human auditory perception, laying the perceptual foundation for speech feature extraction such as MFCC. After the log energy is mapped to the Mel scale to form the Mel spectrum, the DCT transform is performed to realize decorrelation, reduce the redundancy between features, avoid the performance degradation of the classification model due to data correlation, and make the features more compact and effective.

[0049] In addition to extracting MFCC features, the invention also extracts the fundamental frequency and harmonic features in the respiratory sound signal through time-frequency analysis, such as using autocorrelation method or cepstrum method to estimate the fundamental frequency F0, and locating the integer multiple positions (2F0, 3F0,...) of the fundamental frequency F0 on the power spectrum, and calculating the energy sum within the bandwidth near these harmonic frequencies to obtain the harmonic energy.

[0050] Based on the above high-dimensional feature extraction, the invention also extracts low-dimensional features such as expiration duration and amplitude, inspiration duration and amplitude as respiratory sound features.

[0051] The invention obtains a multi-modal respiratory sound feature (combination of high and low dimensional features), and constructs a suffocation anomaly classification model on the multi-modal respiratory sound feature, thereby providing more rich feature information for the model, improving the perception dimension of the model, and ensuring the classification performance of the classification model. The use of fusion features for classification judgment significantly enhances the robustness and anti-interference ability of the model.

[0052] The construction method of the suffocation anomaly classification model includes: The samples in the respiratory sound data set are labeled by type according to normal, low risk and high risk, mainly by judging respiratory rhythm, respiratory continuity, laryngeal stridor degree, cough, and whether there are abnormal respiratory sounds, wherein: Normal: The following conditions must be met simultaneously: the respiratory cycle interval is uniform, and the difference between adjacent cycle lengths is less than or equal to 0.5 seconds; there is no respiratory airflow interruption exceeding 2 seconds in the entire sample; no high-pitched throat humming sound is heard during inhalation or exhalation; there is no wheezing, snoring, sputum humming sound, or water bubble sound; the number of coughs within 30 seconds is less than or equal to 1, and the cough is single-voiced and non-stimulating; the average inhalation time accounts for less than 50% of the entire respiratory cycle.

[0053] Low risk: In the entire sample, clear inspiratory throat humming sound can be heard for more than 3 times, but the sound intensity is not high; or 1-3 times of barking or irritating dry cough occurs within 30 seconds. Respiratory rhythm changes, the difference between adjacent respiratory cycle lengths is greater than 1.5 seconds, and the respiratory cycle interval is less than 3 seconds; the respiratory sound amplitude is not significantly reduced.

[0054] High risk: In the entire sample, there is more than 1 time of respiratory airflow interruption with a duration of more than 5 seconds; the breathing is extremely irregular, and the difference between the longest and shortest cycle lengths in 5 consecutive respiratory cycles is more than 2 seconds; the inspiratory throat humming sound is extremely loud and irritating, and it lasts or appears as a biphasic throat humming sound; after a bout of coughing, there is more than 2 seconds of airflow interruption; compared with the baseline, the respiratory sound amplitude is reduced by less than 70% or disappears.

[0055] The labeled respiratory sound dataset is randomly divided into a training set and a test set; In the training set, the respiratory sound features are taken as input, and normal, low risk and high risk are taken as output to train the CNN algorithm architecture to obtain an asphyxia abnormality classification model; As shown in Figure 3 The CNN6 algorithm architecture includes 4 convolution blocks ConvBlock, 1 feature aggregation layer and 1 fully connected classification layer, and each convolution block adopts a series combination of a convolution layer Conv2D and a maximum pooling layer MaxPooling2D; The 4 convolution blocks are used to gradually extract deep features in the respiratory sound features; The feature aggregation layer is used to integrate the deep features to form a global semantic representation; The fully connected classification layer is used to output the classification results of normal, low risk and high risk.

[0056] The present application adopts a CNN6 algorithm architecture specially built for audio signal recognition in the AudioSet dataset to construct a classification model. The network is based on the optimization of a classic convolutional neural network architecture, and through hierarchical feature extraction and global information fusion, it reduces the computational complexity while retaining high classification accuracy, and is suitable for deployment on resource-constrained medical edge devices. The model structure includes 4 convolution blocks (ConvBlock), 1 feature aggregation layer and 1 fully connected classification layer, and the overall architecture is as shown in Figure 3As shown. Each convolutional block adopts a serial combination of convolutional layers (Conv2D) and max-pooling layers (MaxPooling2D) to gradually extract deep features in the audio signal. The size of the convolutional kernel is fixed at 3x3, which helps to capture local features while reducing computational complexity. As the convolutional block deepens, the number of channels increases exponentially (from 64 to 128, 256, and finally to 512), allowing more diverse feature information to be extracted. At the same time, through the pooling operation with a step of 2, the spatial dimension of the feature map is gradually compressed (e.g., from 128x800 to 8x50), which helps to reduce data redundancy and improve the robustness of the features. In the shallow convolutional block, the model mainly focuses on the local time-frequency texture features of the audio signal, which are crucial for distinguishing different categories of audio signals. As the convolutional block deepens, the model gradually shifts to capturing energy distribution patterns within a wide time window, which helps to further improve the accuracy of classification. The high-dimensional features output by the convolutional layer often contain redundant information and need to be compressed through pooling. This operation takes the mean or maximum value of the local region, preserving key features while reducing dimensionality. It not only effectively reduces the computational resource consumption of the model, but also suppresses the risk of overfitting caused by data redundancy. The fully connected layer, as the core prediction module of the model, integrates the local features extracted by convolution and pooling to form a global semantic representation and finally outputs the classification result. To obtain interpretable class confidence, the Softmax function needs to be applied to the output layer for probability normalization, converting the original score into a probability distribution for each class. During prediction, the class corresponding to the maximum probability value is taken as the final classification result.

[0057] To systematically evaluate the performance of the suffocation risk classification model, the following classification performance evaluation indicators are selected for comprehensive analysis: accuracy (Accuracy, Acc), recall (Recall, Rec), precision (Precision, Pre), and F1 score. The four core indicators in the confusion matrix, true positive (True Positive, TP), true negative (True Negative, TN), false positive (False Positive, FP), and false negative (False Negative, FN), are used for model evaluation: true positive (True Positive, TP): the number of true positive samples correctly predicted as positive by the model; true negative (True Negative, TN): the number of true negative samples correctly predicted as negative by the model; false positive (False Positive, FP): the number of true negative samples incorrectly predicted as positive by the model; false negative (False Negative, FN): the number of true positive samples incorrectly predicted as negative by the model; The specific calculation methods and clinical significance of each indicator are as follows: Accuracy (Acc): evaluate the overall correct rate of positive and negative predictions of the model for all samples.

[0058] Recall (Rec): evaluate the identification ability of the model for true positive samples.

[0059] Precision (Pre): reflect the accuracy of the model for predicting positive samples.

[0060] F1 Score: a balanced indicator that measures the accuracy of the model in predicting positive samples, combining the performance of precision and recall.

[0061] In addition, Receiver Operating Characteristic (ROC) and its Area Under the Curve (AUC) and Precision-Recall Curve (PR) are also used to evaluate the performance of the model.

[0062] The present application randomly divides the data set into training set and test set according to the ratio of 8:2, constructs the suffocation risk classification model, and obtains the performance evaluation results of the model on the training set and the test set, as shown in Figure 4 and 5 .

[0063] As shown in Figure 4 a, the training loss and test loss of the model gradually decrease with the increase of training rounds, and finally tend to be stable, finally converging to a lower level, with an average loss of 0.2174, which indicates that the model effectively learns the data features in the training process and gradually reduces the prediction error. As shown in Figure 4 b, the training accuracy and test accuracy of the model gradually increase with the increase of training rounds, and finally reach an overall accuracy of 96.42% on the test set. This indicates that the model shows good accuracy and stability in the classification task.

[0064] From the classification performance, as shown in Figure 5As shown in Figure b, the model performed best in identifying cases without suffocation risk, with a precision of 97.79%, a recall of 99.50%, and an F1 score of 98.65%. This indicates that the model effectively avoids misclassifying healthy children as at-risk, thereby reducing unnecessary medical examinations and parental anxiety. In clinical practice, this helps optimize resource allocation, focusing attention on children who are truly at risk. For data with low suffocation risk, the recall was 97.75%, the precision was 91.78%, and the F1 score was 94.67%. This shows that the model's ability to identify children with low suffocation risk remains strong. Although 8 low-risk cases were misclassified as high-risk, from a safety perspective, such misclassification is more acceptable than missed diagnosis, as it ensures that the risk is not underestimated and is more conducive to the early detection of children with potential risks. For data with high suffocation risk, the precision was 97.61%, the recall was 89.76%, and the F1 score was 93.52%. This indicates that the model has a low false positive rate when identifying high-risk suffocation states; almost all data judged as high-risk suffocation states are high-risk audio, with very few misjudgments. However, from... Figure 5 The confusion matrix of 'a' shows that it misclassified 34 audio segments with high suffocation risk as low suffocation risk, which means that a false negative occurred.

[0065] from Figure 5 The ROC curve of c and Figure 5 As shown by the PR curves in Figure d, the model performs best in recognizing normal breathing, with an AUC value of 0.99, almost completely distinguishing between normal and abnormal breathing states. Simultaneously, the model exhibits high precision and recall, accurately identifying most normal breathing samples while maintaining a low false positive rate, further validating its good performance. When identifying low asphyxiation risk, the model has an AUC value of 0.97, indicating high accuracy. However, its precision decreases with high recall, indicating a certain false positive rate, which may lead to misjudgment of asphyxiation risk types. When identifying high asphyxiation risk, the model has an AUC value of 0.94, which, although lower than the above results, remains at a high level, but its ability to identify high asphyxiation risk decreases. These results demonstrate that the asphyxiation risk classification model constructed in this invention exhibits good overall performance and can be used to monitor infant asphyxiation risk. However, in asphyxiation risk prediction, sensitivity should take precedence over specificity, as the cost of missed diagnoses far outweighs that of false diagnoses. Therefore, further optimization of the model can improve its sensitivity to high asphyxiation risk states.

[0066] Methods for constructing dynamic baseline models include: Select a normal breath sound signal that meets the specified duration (e.g., 24 hours) from the denoised breath sound signal signal as the learning period signal; MFCC features, fundamental frequency features, harmonic energy features, expiration length and amplitude, inspiration length and amplitude of the learning period signal are acquired, and the mean and standard deviation of each feature in a specified time length are calculated; According to the Gaussian distribution of the features in the specified time length, a baseline model of each feature is established , and is combined as a dynamic baseline model , wherein is the baseline model of feature z, is the mean of feature z, is the standard deviation of feature z.

[0067] After the dynamic baseline model enters the monitoring state, it will be continuously (or at certain time intervals) updated with new breathing data that is judged to be normal by the suffocation risk classification model, such as using recursive estimation methods such as exponential smoothing to allow the baseline model to slowly adapt to the growth and natural changes of the baby. The exponential smoothing updates the mean: , wherein is the updated mean, is the current value of feature z, is the mean before updating, is a very small learning rate, such as 0.001, to ensure that the baseline does not change dramatically due to short-term fluctuations. The standard deviation can also be updated in a similar way. The updated dynamic baseline model is .

[0068] The dynamic baseline model established by the present application is an individual digital portrait of the normal breathing pattern that is based on the individualization of the baby and continuously self-adapting to update over time. It is not a fixed threshold, but a statistical model that describes the normal fluctuation range of each breathing sound feature of the baby in a healthy state.

[0069] For example, the normal breathing patterns of different babies differ greatly. The breathing sound of a full-term healthy baby may be loud and powerful, while the breathing sound of a premature baby or a weaker baby may be naturally very weak. If a uniform and fixed threshold is used (for example, an alarm is given if the audio energy is below a certain value), the system will continuously produce false positives for babies with weak breathing sounds, and may miss reports for babies with loud breathing sounds.

[0070] After the dynamic baseline model is constructed, the unique individualized normal parameter range of the current monitored baby is learned and recorded. For example, the normal expiration length baseline of a baby is 2s±0.3s, and subsequent monitoring is compared with this baby's own baseline rather than a universal standard, and a warning is given when the baseline level is exceeded to a certain extent, thereby achieving accurate respiratory abnormality warning that adapts to individual characteristics.

[0071] The dynamic baseline model provides tailored monitoring for each infant, acknowledges and adapts to physiological diversity, can grow with the infant, always maintains monitoring accuracy without manual resetting or calibration, guarantees effectiveness for long-term use, and the dynamic baseline itself is a record of healthy trends. Doctors or parents can view the baseline curve over time to understand the respiratory development or recovery process of the infant, rather than just receiving isolated alarm events.

[0072] After the respiratory abnormality is monitored by the present application, the abnormal respiratory sound signal is transmitted to the suffocation risk classification model to distinguish whether the respiratory abnormality belongs to obstructive suffocation or central suffocation, and feedback to the guardian terminal for risk warning.

[0073] The method for monitoring and warning the respiratory abnormality of the infant by using the dynamic baseline model and classifying the suffocation risk of the respiratory abnormality fed back by the dynamic baseline model by using the suffocation abnormality classification model comprises: The dynamic baseline model is used to monitor the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the deviation degree of the expiration time length and amplitude, and the inspiration time length and amplitude of the real-time denoising respiratory sound signal of the infant, wherein: When the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration time length and amplitude, and the inspiration time length and amplitude of the real-time denoising respiratory sound signal of the infant deviate from the dynamic baseline model by more than a preset threshold value, it is determined that the infant is in a respiratory abnormality state; When the MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration time length and amplitude, and the inspiration time length and amplitude of the real-time denoising respiratory sound signal of the infant deviate from the dynamic baseline model by less than a preset threshold value, it is determined that the infant is not in a respiratory abnormality state; The MFCC feature, the fundamental frequency feature, the harmonic energy feature, the expiration time length and amplitude, and the inspiration time length and amplitude of the denoising respiratory sound signal corresponding to the determination that the infant is in a respiratory abnormality state are input into the suffocation abnormality classification model to determine the suffocation risk category of the infant.

[0074] As shown in Figure 2 The present application provides an infant suffocation risk autonomous monitoring system based on throat respiratory sound, which is applied to an infant suffocation risk autonomous monitoring method based on throat respiratory sound, comprising: A data acquisition unit comprising a high-sensitivity MEMS microphone and a micro accelerometer, the high-sensitivity MEMS microphone is used to acquire airflow and respiratory sound signals of the infant's throat, and the micro accelerometer is used to synchronously monitor the thoracic movement data of the infant; A signal denoising unit is used to integrate the airflow and respiratory sound signals and the thoracic movement data to perform denoising processing on the airflow and respiratory sound signals to obtain denoising respiratory sound signals; The machine learning unit is used for separating the normal breath sound signal and the abnormal breath sound signal in the de-noised breath sound signal, and combining the normal breath sound signal and the abnormal breath sound signal into a breath sound data set; the breath sound feature is extracted from the breath sound data set, the choking abnormality classification model distinguishing the normal, low risk and high risk is constructed based on the breath sound feature through the machine learning model, and the dynamic baseline model adapting to the individual breath difference is established according to the breath sound feature of the normal breath sound signal; The autonomous monitoring unit is used for monitoring and warning the breath abnormality of the infant by using the dynamic baseline model, and classifying the breath abnormality fed back by the dynamic baseline model into the choking risk by using the choking abnormality classification model.

[0075] The application provides a computer readable storage medium, and the computer readable storage medium stores computer execution instructions.

[0076] The application directly collects the throat breath sound by using the high-sensitivity MEMS microphone, and reduces the environmental noise interference and improves the early choking identification precision by using the micro accelerometer for de-noising processing. The embedded AI algorithm formed by the machine learning model and the dynamic baseline model can effectively distinguish the normal breath variation precursor and the choking category, and realizes the real-time analysis for the infant. The application combines the hardware device of the high-sensitivity MEMS microphone and the micro accelerometer with the artificial intelligence model, realizes the self-completion of the infant choking risk identification and early warning in the home environment by using the throat breath sound signal.

[0077] The above examples are only exemplary embodiments of the application, and are not used to limit the application, and the protection scope of the application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the application within the spirit and protection scope of the application, and the modification or equivalent replacement is also regarded as falling within the protection scope of the application.

Claims

1. A method of autonomous monitoring of risk of asphyxia in an infant based on laryngeal breath sounds, characterized in that, The method comprises the following steps: Collecting airflow and respiratory sound signals of the infant's throat by using a high-sensitivity MEMS microphone, and synchronously monitoring thoracic movement data of the infant by using a miniature accelerometer; Integrating the airflow and respiratory sound signals and the thoracic movement data, and performing noise reduction processing on the airflow and respiratory sound signals to obtain a denoised respiratory sound signal; Signal separating the denoised respiratory sound signal according to a respiratory event, and enhancing and combining to obtain a respiratory sound dataset; Extracting respiratory sound features from the respiratory sound dataset, constructing a suffocation anomaly classification model for distinguishing normal, low-risk and high-risk suffocation anomalies based on the respiratory sound features by using a machine learning model, and establishing a dynamic baseline model that adapts to individual respiratory differences based on the respiratory sound features; Monitoring and warning the respiratory anomaly of the infant by using the dynamic baseline model, and classifying the respiratory anomaly fed back by the dynamic baseline model into suffocation risk by using the suffocation anomaly classification model.

2. The method of claim 1, wherein the method is further characterized by: The method for integrating the airflow and respiratory sound signals and the thoracic movement data, and performing noise reduction processing on the airflow and respiratory sound signals comprises: constituent decomposition of an airflow and breath sound signal, wherein, is the airflow and breath sound signal, is a target breath sound signal, is is a noise signal resulting from body motion in is is an ambient noise signal in performing a constitutive decomposition on the thoracic motion data, the thoracic motion data being: wherein, is thoracic motion data, is is a noise signal resulting from body motion in the thoracic motion data, is is an ambient noise signal in the thoracic motion data; The iterative operation of the least mean square algorithm LMS causes the output noise estimate of the adaptive filter and the target breath sound signal is calculated from the noise estimate wherein is the noise estimate, is a set of thoracic movement data, is the thoracic movement data of the previous time instant of is the thoracic movement data of the previous L-1 time instants of is the filter weight, is the target breath sound signal;​ target respiratory sound signal as the denoised respiratory sound signal.

3. The method of claim 2, wherein the method is further characterized by: The method for signal separating the denoised respiratory sound signal according to a respiratory event comprises: Frame the denoised respiratory sound signal into a plurality of audio samples according to a preset time length, and fill in the audio samples that are less than the preset time length in time length; Calculate the short-time energy (STE) and the short-time zero-crossing rate (ZCR) of each audio sample, and remove the audio samples without sound based on the short-time energy (STE) and the short-time zero-crossing rate (ZCR) by using a double-threshold method; In the remaining audio samples, identify the audio samples corresponding to the respiratory events of inhalation, exhalation, crying, coughing and abnormal respiratory sound based on the short-time energy (STE) and the short-time zero-crossing rate (ZCR), and reserve the audio samples as value-type samples.

4. The method of claim 3, wherein the method is further characterized by: The method for enhancing and combining the value-type samples into a respiratory sound dataset comprises: Time domain displacement: shift the value-type samples backward by 1s on the time axis, and splice the 7.0s to 8.0s part of the value-type samples to the front end of the value-type samples; Gain adjustment: lower or raise the amplitude of each value-type sample, wherein the oscillation parameter of the amplitude changed by the lowering or raising is a random value between 0.1 and 5.0; Introducing noise: add a random noise signal to the value-type samples, and adjust the signal-to-noise ratio to a preset value; Mix the value-type samples and the enhanced value-type samples as the respiratory sound dataset.

5. The method of claim 4, wherein the method is further characterized by: The method for extracting the respiratory sound features comprises: Process the samples in the respiratory sound dataset in sequence by pre-emphasis, framing, windowing, fast Fourier transform, mel filter bank and logarithmic energy processing to obtain a mel spectrum diagram, and perform DCT transformation processing on the mel spectrum diagram to obtain MFCC features; Estimate the fundamental frequency features of the samples in the respiratory sound dataset by using a cepstrum method, and calculate the harmonic energy features based on the fundamental frequency features; Calculate the exhalation time and amplitude, and the inhalation time and amplitude of the samples in the respiratory sound dataset in sequence; Take the MFCC features, the fundamental frequency features, the harmonic energy features, the exhalation time and amplitude, and the inhalation time and amplitude as the respiratory sound features.

6. The method of claim 5, wherein the method is further characterized by: The method for constructing the suffocation anomaly classification model comprises: The respiratory sound data set samples are type-labeled according to normal, low risk and high risk, and the labeled respiratory sound data set is randomly divided into a training set and a test set; In the training set, the respiratory sound features are taken as input, and the normal, low risk and high risk are taken as output to train the CNN algorithm architecture, and the asphyxia abnormality classification model is obtained; The CNN6 algorithm architecture includes 4 convolution blocks ConvBlock, 1 feature aggregation layer and 1 fully connected classification layer, and each convolution block adopts a series combination of a convolution layer Conv2D and a maximum pooling layer MaxPooling2D; The 4 convolution blocks are used to gradually extract deep features in the respiratory sound features; The feature aggregation layer is used to integrate the deep features to form a global semantic representation; The fully connected classification layer is used to output the classification results of normal, low risk and high risk.

7. A method of self-monitoring of risk of asphyxia in an infant based on laryngeal breathing sounds according to claim 6, characterized in that: The method for constructing the dynamic baseline model comprises: A normal respiratory sound signal of a specified duration is selected from the denoised respiratory sound signal as a learning period signal; The MFCC features, fundamental frequency features, harmonic energy features, expiration duration and amplitude, and inspiration duration and amplitude of the learning period signal are obtained, and the mean and standard deviation of each feature within the specified duration are calculated; A baseline model of each feature is established according to the characteristic that the feature obeys a Gaussian distribution within a specified time length and combined as the dynamic baseline model wherein, is a baseline model of feature z, is a mean of feature z, is a standard deviation of feature z.

8. The method of claim 7, wherein the method is further characterized by: The method for monitoring and warning the respiratory abnormality of the infant by using the dynamic baseline model and classifying the respiratory abnormality fed back by the dynamic baseline model by using the asphyxia abnormality classification model comprises: The deviation degree of the MFCC features, fundamental frequency features, harmonic energy features, expiration duration and amplitude, and inspiration duration and amplitude of the real-time denoised respiratory sound signal of the infant is monitored by using the dynamic baseline model, wherein: When the deviation degree of the MFCC features, fundamental frequency features, harmonic energy features, expiration duration and amplitude, and inspiration duration and amplitude of the real-time denoised respiratory sound signal of the infant deviates from the dynamic baseline model by more than a preset threshold, it is determined that the infant is in a respiratory abnormality state; When the deviation degree of the MFCC features, fundamental frequency features, harmonic energy features, expiration duration and amplitude, and inspiration duration and amplitude of the real-time denoised respiratory sound signal of the infant deviates from the dynamic baseline model by less than a preset threshold, it is determined that the infant is not in a respiratory abnormality state; The MFCC features, fundamental frequency features, harmonic energy features, expiration duration and amplitude, and inspiration duration and amplitude of the denoised respiratory sound signal corresponding to the infant in the respiratory abnormality state are input into the asphyxia abnormality classification model to determine the asphyxia risk category of the infant.

9. An autonomous monitoring system of risk of asphyxia in infants based on laryngeal breath sounds, characterized in that it comprises: The method for autonomously monitoring the asphyxia risk of an infant based on throat respiratory sound according to any one of claims 1-8 comprises: A data acquisition unit comprising a high-sensitivity MEMS microphone and a micro accelerometer, wherein the high-sensitivity MEMS microphone is used to acquire airflow and respiratory sound signals of the throat of the infant, and the micro accelerometer is used to synchronously monitor the thoracic movement data of the infant; A signal denoising unit used to denoise the airflow and respiratory sound signals by integrating the airflow and respiratory sound signals and the thoracic movement data, and obtain denoised respiratory sound signals; The machine learning unit is configured to separate normal breath sound signals and abnormal breath sound signals in the de-noised breath sound signals, and combine the normal breath sound signals and the abnormal breath sound signals into a breath sound dataset; extract breath sound features from the breath sound dataset, construct a suffocation anomaly classification model for distinguishing normal, low-risk and high-risk suffocation anomalies based on the breath sound features through a machine learning model, and establish a dynamic baseline model adapted to individual respiratory differences according to the breath sound features of the normal breath sound signals; The autonomous monitoring unit is configured to monitor and warn the respiratory anomaly of the infant by using the dynamic baseline model, and classify the respiratory anomaly fed back by the dynamic baseline model into suffocation risk categories by using the suffocation anomaly classification model.

10. The infant asphyxia risk self-monitoring system based on laryngeal breath sounds according to claim 9, characterized in that: The input of the suffocation anomaly classification model is the breath sound features, and the output is the suffocation anomaly categories of normal, low-risk and high-risk.

Citation Information

Patent Citations

  • Artificial-intelligence-based-based early warning system and method of infant asphyxia

    CN108615333A

  • Breathing sound classification method based on deep learning

    CN111640439A

  • Sleep disorder monitoring and system thereof

    CN119907639A

  • Device and method for predicting risk of Pierre Robin syndrome by using audio

    CN120727045A

  • Portable device with multiple integrated sensors for vital signs scanning

    US20190298183A1

Cited By

  • Newborn image record detection method, system and equipment based on artificial intelligence and medium

    CN122224397A