Multi-mode child snore monitoring system and monitoring method based on mobile terminal

Through the integrated audio acquisition and multimodal data processing on the mobile terminal, the children's snoring monitoring system is used to integrate feature extraction such as the Mel frequency cepspectral coefficient and LSTM-SNN network recognition, the problem of children's OSA not being diagnosed in time is solved, and high-precision and low-latency household monitoring is achieved.

CN120388588APending Publication Date: 2025-07-29SHANGHAI STOMATOLOGICAL HOSPITAL FUDAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802879.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing technology lacks convenient and cost-effective family-based obstructive sleep apnea (OSA) monitoring methods, resulting in the failure of timely diagnosis of OSA in children, delaying treatment, and affecting healthy development.

Method used

The multimodal children's snoring monitoring system based on mobile terminals acquires data through the audio acquisition module, uses the Mel frequency cepspectral coefficient, filter group coefficient, linear prediction coefficient and short-time average energy for pre-processing, and is identified through the double-layer LSTM-SNN network to output the snoring monitoring results.

Benefits of technology

It realizes high-precision and low-latency children's OSA monitoring on mobile terminals, significantly improving classification robustness and accuracy, adapting to complex family environments, and suitable for family-based monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388588A_ABST
    Figure CN120388588A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode child snore monitoring system and detection system based on a mobile terminal, the mobile terminal comprises an audio acquisition module, a data storage module and a data processing module, the data storage module stores audio data acquired by the audio acquisition module, and the data processing module processes the audio data. The data processing module preprocesses the audio data, extracts a Mel-frequency cepstrum coefficient, a filter bank coefficient, a linear prediction coefficient and short-time average energy of each frame of audio data and stores the Mel-frequency cepstrum coefficient, the filter bank coefficient, the linear prediction coefficient and the short-time average energy in the data storage module, and an identification model is further configured in the data storage module; and the data processor module is controlled by a preset control signal, and calls the recognition model for calculation based on the Mel-frequency cepstrum coefficient, the filter bank coefficient, the linear prediction coefficient and the short-time average energy so as to output a child snore monitoring result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of children's snoring monitoring, and particularly relates to a multi-modal children's snoring monitoring system and monitoring method based on a mobile terminal. Background Art

[0002] OSA (obstructive sleep apnea) is a sleep breathing disorder in which the upper airway partially or completely collapses repeatedly during sleep, causing chronic intermittent hypoxia and sleep fragmentation. It can cause damage to multiple organs and systems such as cardiovascular and cerebrovascular diseases, cognitive impairment, and diabetes, and even sudden death, imposing a heavy health and economic burden on society. Currently, polysomnography (PSG) is considered the "gold standard" for diagnosing OSA diseases. However, due to the high price and complex operation of PSG devices, patients need to stay in the hospital for at least one night for monitoring, and the monitoring cost is relatively high. As a result, a large number of patients fail to be monitored in time, delaying the best treatment time and posing a great threat to the patients themselves, their families, and society. Therefore, convenient, home-based, and digital diagnostic technologies have received great attention.

[0003] The common causes of OSA in children include anatomical structure abnormalities, such as tonsil and adenoid hypertrophy. The children's airway is relatively narrow, and the enlarged tissues directly block the nasopharynx or oropharynx. Or maxillofacial developmental abnormalities, such as micrognathia, mandibular retrognathia, high-arched palate, etc., leading to airway stenosis. If children's OSA is not intervened, the following consequences usually occur: Difficulty in breathing: Difficulty in breathing is both the most intuitive manifestation and impact of children's sleep apnea. If this situation is not treated as soon as possible, it is very easy to induce more serious symptoms such as respiratory failure, and even threaten the life safety of children. Growth retardation: Apnea will cause the body to lack oxygen, and insufficient oxygen supply will lead to growth retardation of the body. Although growth retardation does not directly damage physical health, it will have a great negative impact on the future development of children. Sleep disorder: Over time, it may even lead to negative impacts such as inability to concentrate and memory decline in children.

[0004] Therefore, it is very necessary to effectively monitor and identify children's OSA. At present, there is no relatively convenient and cost-effective product that can provide home monitoring services for children's OSA. Summary of the Invention

[0005] To solve the above problems, the purpose of the present invention is to provide a multi-modal children's snoring monitoring system and detection method based on a mobile terminal, which can extract multi-modal information based on the original audio data, and then identify the monitoring results based on the multi-modal information, and the detection results are more accurate.

[0006] One technical solution provided by the present invention is: a multi-modal children's snoring monitoring system based on a mobile terminal. The mobile terminal includes an audio acquisition module, a data storage module, and a data processing module. The data storage module stores the audio data obtained through the audio acquisition module. The data processing module preprocesses the audio data, extracts the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data respectively, and stores them in the data storage module. An identification model is also configured in the data storage module; Controlled by a preset control signal, the data processor module calls the identification model based on the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the children's snoring monitoring result; Among them, after the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are respectively subjected to pulse coding processing through threshold coding, they are respectively input into four independent double-layer LSTM-SNN networks. The four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

[0007] Preferably, the preprocessing of the audio data by the data processing module includes: Filter the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000Hz to 15000Hz; Frame the audio data after high-pass filtering according to a preset duration.

[0008] Preferably, the preprocessing of the audio data by the data processing module further includes: Based on the fact that the high-order resonance peak bandwidth of the snoring of OSA children is 30%-50% wider than that of non-OSA children and needs to enclose a larger frequency range, resample the audio data into a signal with a sampling rate of 30000Hz to ensure signal integrity and reduce the data processing burden of the mobile terminal and speed up the processing speed.

[0009] Preferably, the preprocessing of the audio data by the data processing module further includes: Use a Wiener filter to perform noise reduction processing on the audio data, and use spectral subtraction + variational Bayesian reconstruction Wiener gain; Introduce a correction term , where is a morphological distortion penalty term used to adjust the signal gain to prevent over-suppression of certain features of children's snoring, is the observed noise spectrum, is the prior spectrum based on children's snoring, which is the statistical feature extracted from the children's snoring database, represents the Kullback-Leibler divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural characteristics of children's snoring sounds; After introducing the spectral morphology prior, construct G(k,j) of variational Bayesian estimation and rewrite the gain as: where Control the morphological distortion penalty through the variational Bayesian method to prevent the filter from overly suppressing the structural information in children's snoring sounds.

[0010] Preferably, the mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data are further subjected to pulse coding processing through threshold coding, including: Generate a truncated Gaussian distribution threshold adjustment threshold matrix with a mean of 0 and a variance of 0.8 to reduce the pulse trigger threshold of children's weak snoring sounds; Detect local maxima in the time domain window (6 frames) and frequency domain window (5 filters) in the spectrogram to strengthen key point detection and enhance the recognition of sudden snoring sounds.

[0011] Preferably, extracting the mel-frequency cepstral coefficients of each frame of the audio data includes: The number of mel-scale triangular filters is 40, covering the range of 50 - 8000 Hz, with a focus on enhancing the high-frequency harmonic concentration area of children's snoring sounds in the frequency band of 1000 - 5500 Hz; The cepstral coefficient calculation uses an improved discrete cosine transform, retaining the first 13-dimensional coefficients to suppress high-frequency noise interference.

[0012] Preferably, extracting the filter bank coefficients of each frame of the audio data includes: Use the Burg algorithm to calculate the linear prediction coefficients, with an order of 14, and extract the resonance peak positions of children's upper respiratory tracts through pole analysis: F1: 316, F2: 796, F3: 1425, F4: 2605, F5: 4510.

[0013] Preferably, the input layer of the recognition model is: Four branches are separately set to process the mel-frequency cepstral coefficients (39×50 frames), filter bank coefficients (40×50 frames), linear prediction coefficients (14×50 frames), and short-time average energy (1×50 frames); The hidden layer of the recognition model is: The first layer has 128 LSTM-SNN units, with an activation threshold σ1 = 0.3; the second layer has 64 units, σ2 = 0.1, to adapt to the sparsity of children's signals; The output layer of the recognition model is as follows: The fully connected layer fuses multiple features, and Softmax outputs the AHI grading probability (binary classification: AHI ≤ 5 vs. AHI > 5; multi-classification: mild / moderate / severe).

[0014] Based on the same concept, the present invention also provides a multi-modal children's snoring monitoring method based on a mobile terminal, which is applied to the monitoring system described in any one of the above, and includes the following steps: Store the audio data obtained by the audio acquisition module; Preprocess the audio data, and extract the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data and store them in the data storage module; Controlled by a preset control signal, call the recognition model based on the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the children's snoring monitoring result; Among them, the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are respectively pulse-coded through threshold encoding and then input into four independent double-layer LSTM-SNN networks. The four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

[0015] Preferably, the preprocessing of the audio data by the data processing module includes: Filter the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000Hz to 15000Hz; Frame the audio data after high-pass filtering according to a preset duration; Based on the fact that the high-order resonance peak bandwidth of the snoring of OSA children is 30%-50% wider than that of non-OSA children, and a larger frequency range needs to be enveloped, resample the audio data into a signal with a sampling rate of 30000Hz to ensure signal integrity and reduce the data processing burden of the mobile terminal and speed up the processing speed; Use a Wiener filter to perform noise reduction processing on the audio data, and use spectral subtraction + variational Bayesian reconstruction Wiener gain; introduce a correction term , where is a morphological distortion penalty term, which is used to adjust the signal gain to prevent over-suppression of some features of children's snoring, , is the observed noise spectrum, is the prior spectrum based on children's snoring, and the statistical features are extracted from the children's snoring database, represents the Kullback-Leibler divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural characteristics of children's snoring; After introducing the spectral morphology prior, construct the variational Bayesian estimation of G(k,j), and rewrite the gain as: where Control the morphological distortion penalty through the variational Bayesian method to prevent the filter from over-suppressing the structural information in children's snoring.

[0016] Due to the adoption of the above technical solutions, the present invention has the following advantages and positive effects compared with the prior art: The technical solution of this embodiment can comprehensively capture the respiratory acoustic characteristics during children's sleep by simultaneously extracting Mel Frequency Cepstral Coefficients (MFCC), Filter Bank Coefficients (FBANK), Linear Prediction Coefficients (LPC), and Short-Time Average Energy (STE). The four features are respectively input into independent double-layer LSTM-SNN networks to avoid interference between different features and retain their respective time dynamic information at the same time; after being combined through the fully connected layer, the softmax classifier comprehensively outputs the classification results of multi-modalities, significantly improving the classification robustness and accuracy. Children's snoring is mostly caused by adenoid hypertrophy, tonsil hyperplasia, and inferior turbinate hypertrophy, while adults are mainly caused by soft palate vibration. The narrow nasopharyngeal cavity of children attenuates high-frequency sound waves weakly, resulting in the retention of high-frequency energy; the longer pharyngeal cavity of adults enhances low-frequency resonance. In children with OSA, due to the increase in upper airway resistance, when the airway resumes patency after apnea, the air flow velocity surges, causing the tissue vibration to intensify, manifested as a sudden increase in loudness and an expansion of the formant bandwidth. Therefore, there are obvious acoustic differences between children's OSA and adults' OSA. In this embodiment, MFCC simulates the auditory characteristics of the human ear, FBANK provides frequency domain distribution information, LPC reflects the characteristics of the vocal tract, and STE is related to the respiratory intensity. The multi-modal features complement each other and enhance the recognition ability of children's OSA-related abnormalities (such as apnea and hypopnea). BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The following further details the specific embodiments of the present invention with reference to the drawings, where: Figure 1 is a schematic diagram of the composition of the multi-modal children's snoring monitoring system based on a mobile terminal of the present invention. SPECIFIC EMBODIMENTS

[0018] The following further details the present invention with reference to the drawings and specific embodiments. The advantages and features of the present invention will be clearer according to the following description and the claims. It should be noted that the drawings are all in a very simplified form and use non-precise ratios, only for the purpose of conveniently and clearly assisting in explaining the embodiments of the present invention.

[0019] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0020] First Embodiment This embodiment provides a multi-modal children's snoring monitoring system based on a mobile terminal. The mobile terminal includes an audio acquisition module, a data storage module, and a data processing module. The data storage module stores the audio data obtained through the audio acquisition module. The data processing module preprocesses the audio data, extracts the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data respectively, and stores them in the data storage module. An identification model is also configured in the data storage module; Controlled by a preset control signal, the data processor module calls the identification model based on the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the children's snoring monitoring result; Among them, after the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are respectively subjected to pulse coding processing through threshold coding, they are respectively input into four independent double-layer LSTM-SNN networks. The four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

[0021] In the technical solution of this embodiment, by simultaneously extracting Mel Frequency Cepstral Coefficients (MFCC), Filter Bank Coefficients (FBANK), Linear Prediction Coefficients (LPC), and Short-Time Average Energy (STE), the system can comprehensively capture the respiratory acoustic characteristics during children's sleep. The four features are respectively input into independent double-layer LSTM-SNN networks to avoid interference between different features while retaining their respective temporal dynamic information. After being combined through a fully connected layer, the softmax classifier synthesizes the multimodal output classification results, significantly improving the classification robustness and accuracy. Children's snoring is mostly caused by adenoid hypertrophy, tonsil hyperplasia, and inferior turbinate hypertrophy, while in adults, it is mainly caused by soft palate vibration. The narrow nasopharyngeal cavity in children attenuates high-frequency sound waves weakly, resulting in the retention of high-frequency energy; the longer pharyngeal cavity in adults enhances low-frequency resonance. In children with OSA, due to the increased upper airway resistance, when the airway resumes patency after apnea, the airflow velocity surges suddenly, causing the tissue vibration to intensify, manifested as a sudden increase in loudness and an expansion of the formant bandwidth. Therefore, there are obvious acoustic differences between children's OSA and adults' OSA. In this embodiment, MFCC simulates the auditory characteristics of the human ear, FBANK provides frequency-domain distribution information, LPC reflects the characteristics of the vocal tract, and STE is related to the respiratory intensity. The multimodal feature complementarity enhances the ability to identify abnormalities related to children's OSA (such as apnea and hypopnea).

[0022] Children's snoring has the following characteristics: I. Spectral characteristics 1. High-frequency energy distribution Fundamental frequency and harmonics: The fundamental frequency of children's snoring is significantly higher than that of adults. Research data shows that the average fundamental frequency of children is about 1270 Hz (range: 1200 - 1400 Hz), while the fundamental frequency of adults is only 130 Hz (range: 100 - 150 Hz). The harmonic energy distribution of children is wider, commonly found in high-frequency bands such as 1640 Hz, 3900 Hz, 5400 Hz, etc., while the harmonics of adults are mostly concentrated in low-frequency bands such as 226 Hz, 570 Hz, 1100 Hz, etc.

[0023] Spectral attenuation characteristics: Observable energy still exists in children's snoring above 6 kHz (such as near 15 kHz), while the high-frequency energy of adults' snoring decays rapidly and is mainly concentrated below 6 kHz.

[0024] 2. Differences in spectral stability The spectral envelope of children with simple snoring is relatively smooth and the energy distribution is concentrated; while the spectrum of children with OSA fluctuates significantly, and the bandwidth of the higher-order formants increases by 30% - 50%, indicating the instability of the acoustic characteristics caused by the collapse of the upper respiratory tract.

[0025] II. Formant structure 1. Association between frequency and anatomy The average first formant (F1) in children is 316 - 320 Hz, significantly higher than 150 - 152 Hz in adults, reflecting the anatomical characteristics of a shorter pharyngeal cavity and a narrower nasal cavity in children. The higher formants (such as the fifth formant F5) can reach 4490 - 4510 Hz in children, while only 3200 - 3312 Hz in adults, and the difference is due to the different positions of the larynx and the lengths of the vocal tracts.

[0026] 2. Impact of OSA The formant bandwidths of children with OSA increase significantly. For example, the bandwidth of the fifth formant increases from 88 Hz in simple snoring to 156 Hz, which may be related to the reduced airway muscle tension after apnea.

[0027] III. Loudness Characteristics 1. Association between Intensity Distribution and Pathology The average loudness of children with simple snoring is 40.18 ± 9.48 dBA, 44.92 ± 10.16 dBA for children with OSA, and up to 59.56 ± 7.24 dBA for adult OSA patients. The increase in loudness in children with OSA (about 12%) is lower than that in adults (23%), possibly because the airways in children are more prone to collapse and the airflow speed is limited during respiratory recovery.

[0028] 2. Dynamic Range The maximum instantaneous loudness of children's snoring can reach 60 dBA (such as explosive snoring), but the overall dynamic range (about 20 - 60 dBA) is narrower than that of adults (30 - 70 dBA).

[0029] IV. Time Domain and Signal Characteristics 1. Characteristics of Snoring Segments The average duration of children's snoring segments ranges from 1.28 ± 0.25 seconds (simple snoring) to 1.39 ± 0.28 seconds (OSA), which is similar to that of adults, but the boundary detection error is lower (about 0.074 seconds). The periodicity of children's snoring sequences is weak, and short-time autocorrelation analysis needs to be combined to enhance detection.

[0030] 2. Impact of Signal-to-Noise Ratio (SNR) The average SNR of children's snoring is 16.3 ± 2.1 dB, and the boundary detection error of low SNR segments (<9 dB) increases significantly, and the noise reduction algorithm needs to be optimized specifically.

[0031] Preferably, the preprocessing of the audio data by the data processing module includes: Filtering the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000 Hz - 15000 Hz; Framing the audio data after high-pass filtering according to a preset duration.

[0032] In this embodiment, considering the different recording conditions of mobile terminals (such as mobile phones), the robustness of the signal preprocessing algorithm is fully considered, especially the pre-filtering algorithm and noise reduction. For the preprocessing algorithm of children's snoring sounds, according to the characteristics of children's snoring sounds and taking into account the low-frequency environmental noise, this embodiment specifically adjusts some contents in the pre-filtering algorithm, including adjusting the range of the band-pass filter to 1000 Hz - 15000 Hz. Raising the lower limit to 1000 Hz is different from the 60 Hz lower limit applied to adults, which can better filter out environmental noise (<1000 Hz), fully considering the differences in different recording environments and recording devices, and providing technical support for the portability, miniaturization, and household use of the device. At the same time, due to the relatively high fundamental frequency (1270 Hz) of children's snoring sounds, the energy of children's snoring signals can be maximally retained without distorting the snoring signals, which is beneficial for subsequent data processing. And different from the characteristic that the observable energy frequency of adult snoring sounds does not exceed 6 kHz, there is still observable energy in children's snoring sounds above 6 kHz (such as near 15 kHz), so the upper limit of the filter is adjusted to 15000 Hz.

[0033] Preferably, the preprocessing of the audio data by the data processing module further includes: Based on the fact that the high-order resonance peak bandwidth of the snoring sounds of children with OSA is 30% - 50% wider than that of children without OSA and needs to envelope a larger frequency range, the audio data is resampled into a signal with a sampling rate of 30000 Hz to ensure signal integrity and reduce the data processing burden of the mobile terminal, thereby accelerating the processing speed.

[0034] In one embodiment, the imported snoring audio has a sampling rate of 44100 Hz. However, considering the relatively low spectral stability of children with OSA and the 30% - 50% wider high-order resonance peak bandwidth compared to children without OSA, which requires enclosing a larger frequency range, different from some applications, the audio is resampled into a signal with a sampling rate of 30000 Hz here. This is because according to the Nyquist rate, the minimum sampling frequency must be at least twice the highest frequency of a continuous-time signal to ensure signal integrity. Since a large amount of data needs to be processed for subsequent training of the large model, this resampling step can significantly reduce the computer's processing burden on the snoring signals and accelerate the processing speed.

[0035] Classical Wiener gain The expression: Where is the signal-to-noise ratio (SNR), that is, at frequency and time frame the ratio of the power of the signal to the power of the noise.

[0036] Wherein: is the signal power at time frame j and frequency k.

[0037] is the noise power at time frame j and frequency k.

[0038] Preferably, in order to adapt to the non-Gaussian characteristics of children's snoring, the data processing module further includes preprocessing the audio data: Performing noise reduction processing on the audio data using a Wiener filter, and using spectral subtraction + variational Bayesian reconstruction Wiener gain; Introduce a correction term , where is the morphological distortion penalty term, which is used to adjust the signal gain to prevent over-suppression of certain characteristics of children's snoring, , is the observed noise spectrum, is the statistical feature extracted from the children's snoring database based on the prior spectrum of children's snoring, represents the KL divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural characteristics of children's snoring; After introducing the spectral morphology prior, construct the variational Bayesian estimation of G(k,j), and rewrite the gain as: where Controlling the morphological distortion penalty by variational Bayesian method to prevent the filter from over-suppressing the structural information in children's snoring.

[0039] Children's snoring sounds are often of low intensity, non-stationary, and skewed distribution. Through spectral subtraction and variational Bayesian reconstruction Wiener gain, the system can adaptively estimate the noise distribution and more accurately separate snoring signals in complex backgrounds (such as environmental sounds and mattress friction sounds), reducing the "musical noise" problem of traditional spectral subtraction. By introducing a morphological distortion penalty term and using the KL divergence between the observed spectrum and the prior spectrum of children's snoring as a constraint, it prevents the excessive weakening of the structural peaks of snoring during the noise reduction process (such as the low-frequency formants of obstructive breathing events). This avoids the loss of OSA features caused by over-smoothing in traditional methods and improves the monitoring sensitivity. Utilizing the prior spectrum morphology of children's snoring (such as energy distribution and harmonic structure), the system preferentially retains the frequency bands that conform to the pathological characteristics of OSA (such as respiratory noise in the low-frequency region and intermittent airflow interruption) during noise reduction, reducing the risk of misjudgment. The variational Bayesian method jointly optimizes the noise and signal distributions through probability modeling, and is more adaptable to the signal-to-noise ratio fluctuations of children's sleep audio (such as acoustic differences caused by different body positions at night) compared to traditional Wiener filtering with fixed thresholds or empirical parameters. The combination of pulse coding and Wiener gain not only retains the temporal continuity between audio frames (required for the input of the LSTM network) but also reduces the interference of redundant information through sparsification processing, improving the classification efficiency of subsequent deep learning models.

[0040] Preferably, the further pulse coding process of the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy through threshold coding respectively includes: Generating a truncated Gaussian distribution threshold adjustment threshold matrix with a mean of 0 and a variance of 0.8 to lower the pulse trigger threshold for children's weak snoring; Detecting local maxima in the time domain window (6 frames) and frequency domain window (5 filters) in the spectrogram to enhance key point detection and improve the recognition of sudden snoring.

[0041] By generating a truncated Gaussian distribution threshold matrix with a mean of 0 and a variance of 0.8, the system dynamically assigns different threshold values to different audio features (such as MFCC, FBANK, etc.). This adaptive strategy reduces the triggering threshold of weak snoring, especially applicable to snoring with low volume but pathological significance during children's sleep (such as mild airway obstruction), reducing the risk of missed detection. The threshold design of the truncated Gaussian distribution takes into account both noise suppression and signal fidelity. It can filter background noise (such as bed sheet friction noise) and avoid the loss of weak features caused by a fixed high threshold, enhancing the robustness in complex environments. Covering an audio segment of about 1 - 2 seconds, it effectively captures short-term events within the respiratory cycle (such as sudden onset of snoring, airflow interruption), avoiding the weakening of sudden signals due to frame-level averaging. Focusing on the OSA-related frequency bands (such as the snoring formant in the low-frequency region and the airflow turbulence in the high-frequency region), through local frequency domain peak detection, it strengthens the recognition of obstructive breathing events. Threshold encoding converts continuous audio features into sparse pulse signals, reducing data redundancy; local maximum detection further extracts spatio-temporal key points, forming a "global-local" collaborative feature representation. This combination not only retains the overall pattern of snoring (such as periodicity) but also highlights the transient features of pathological events (such as sudden peaks). The generation of the threshold matrix and local detection are both lightweight operations, without the need for complex model inference, adapting to the computing power of mobile terminals; combined with a double-layer LSTM-SNN network, it can achieve low-latency real-time monitoring (such as continuous analysis at night).

[0042] Preferably, extracting the Mel-frequency cepstral coefficients of each frame of the audio data includes: The number of Mel-scale triangular filters is 40, covering the range of 50 - 8000 Hz, with a focus on enhancing the high-frequency harmonic concentration area of children's snoring in the frequency band of 1000 - 5500 Hz; The cepstral coefficient calculation uses an improved discrete cosine transform, retaining the first 13-dimensional coefficients to suppress high-frequency noise interference.

[0043] In this embodiment, through precise frequency band coverage (enhancing pathological high-frequency harmonics), noise adaptive suppression (13-dimensional cepstrum compression), and physiological adaptation for children (prioritizing high-frequency features), high-precision and low-latency snoring monitoring for children is achieved in a complex home environment, providing key technical support for mobile terminal deployment, and having both clinical diagnostic value and user experience practicality. Forty Mel triangular filters cover 50 - 8000 Hz, finely matching the frequency spectrum of children's snoring sounds, especially enhancing the 1000 - 5500 Hz frequency band (the concentrated area of high-frequency harmonics in children's snoring). Children's OSA is often accompanied by upper airway stenosis or abnormal vibration, resulting in more significant high-frequency harmonics in snoring (such as turbulent noise, laryngeal resonance). This design can highlight pathological features and avoid interference from low-frequency background noise (such as bed sheet friction noise). By retaining the first 13 cepstral coefficients and discarding high-frequency cepstral components (corresponding to rapidly changing noise in the time domain), interference from environmental noise (such as ambient speech, electronic device interference) and physiological noise (such as swallowing sounds, coughing) is effectively suppressed. Children have short vocal cords and small laryngeal cavities, with a relatively high fundamental frequency of snoring (usually > 200 Hz), and OSA is mostly caused by structural problems such as adenoid hypertrophy and tonsil enlargement, with richer high-frequency harmonics. This solution precisely adapts to children's acoustic characteristics through frequency band and filter design.

[0044] Preferably, the filter bank coefficients of each frame of the audio data are extracted including: The linear prediction coefficients are calculated using the Burg algorithm with an order of 14, and the resonance peak positions of the children's upper respiratory tract are extracted through pole analysis: F1: 316, F2: 796, F3: 1425, F4: 2605, F5: 4510.

[0045] The Burg algorithm is used to calculate the linear prediction coefficients (LPCs). By means of recursion, it can stably estimate the vocal tract model in short-time frames, effectively suppressing noise interference, especially suitable for the non-stationary noises (such as turning over and coughing) commonly found in children's sleep audio. The 14th-order LPC balances the model complexity and the feature expression ability. It can accurately model the acoustic characteristics of children's short vocal tracts (where high-frequency formants are concentrated), avoid overfitting, and retain the key formant information. Focusing on F1 (500 - 1000 Hz) and F2 (1500 - 2500 Hz), it can accurately cover the high-frequency formant shifts caused by upper respiratory tract abnormalities in children (such as adenoid hypertrophy and tonsil obstruction), directly reflecting the pathological features of OSA (such as the harmonic energy changes caused by airway stenosis). By extracting formants through LPC poles instead of direct spectrum analysis, it can weaken the influence of environmental noises (such as bed sheet friction and ambient sound) on feature extraction and enhance the sensitivity to OSA-related acoustic markers (such as intermittent airflow turbulence). The F2 frequency band (1500 - 2500 Hz) covers the concentrated area of high-order harmonics of children's snoring sounds, effectively capturing the transient features of obstructive breathing events (such as sudden changes in laryngeal vibration), making up for the deficiencies of traditional low-frequency-dominated features (such as MFCC). The recursive nature of the Burg algorithm and the moderate complexity of the 14th-order LPC model take into account both computational efficiency and feature accuracy, meeting the requirements of low power consumption and real-time processing on mobile terminals (such as continuous monitoring at night).

[0046] Preferably, the input layer of the recognition model is as follows: Four branches are separately set to process the Mel-frequency cepstral coefficients (39×50 frames), filter bank coefficients (40×50 frames), linear prediction coefficients (14×50 frames), and short-time average energy (1×50 frames); The hidden layer of the recognition model is as follows: The first layer has 128 LSTM-SNN units with an activation threshold σ1 = 0.3; the second layer has 64 units with σ2 = 0.1 to adapt to the sparsity of children's signals; The output layer of the recognition model is as follows: The fully connected layer fuses multiple features, and Softmax outputs the probability of AHI grading (binary classification: AHI ≤ 5 vs. AHI > 5; multi-classification: mild / moderate / severe).

[0047] The four-dimensional features are independently extracted. Mel Frequency Cepstral Coefficients (39×50 frames): Capture spectral details and highlight the harmonic structure and periodicity of snoring. Filter bank coefficients (40×50 frames): Strengthen the energy distribution representation of key frequency bands (such as high-frequency turbulent noise). Linear Prediction Coefficients (14×50 frames): Model the resonance characteristics of the respiratory tract and reflect the harmonic shift caused by airway stenosis. Short-time average energy (1×50 frames): Correlate with the fluctuation of respiratory intensity and assist in identifying airflow interruption events. Avoid interference between features, retain the independence and complementarity of each modality, and enhance the robustness in complex scenarios (such as noisy environments and changes in sleeping postures). LSTM captures the long-term dependencies of respiratory events (such as snoring-intermittent periods), and the sparse activation characteristics of SNN suppress background noise interference; σ1 = 0.3 balances the response to weak signals of children and noise suppression. σ2 = 0.1 further reduces the activation threshold and strengthens the capture of sudden events of children's OSA (such as short-term apnea) to avoid missed detections. The hierarchical transition from 128 to 64 units gradually extracts high-order temporal features, reduces redundant calculations, and adapts to the computing power of mobile terminals. The fully connected layer integrates the four-way features, synthesizes multi-dimensional information such as spectrum, energy, and resonance, and improves the classification confidence. The input layer of the recognition model is branched and processed in parallel, and the LSTM-SNN units in the hidden layer are streamlined, with controllable computational complexity. Pulse coding (threshold matrix) and SNN sparse activation further reduce the runtime power consumption. The model inference process is short (audio framing → feature extraction → LSTM temporal modeling → fusion classification), meeting the immediate feedback requirements for continuous nighttime monitoring (such as real-time AHI warning). Through independent extraction of multi-modal features, hierarchical LSTM-SNN temporal modeling, and dynamic threshold adaptation to the sparsity of children's signals, high-precision and low-latency monitoring of children's snoring is achieved on mobile terminals.

[0048] Second Embodiment Based on the same concept, the present invention also provides a multi-modal children's snoring monitoring method based on a mobile terminal, which is applied to the monitoring system described in any one of the above, and includes the following steps: Store the audio data obtained by the audio acquisition module; Preprocess the audio data, and separately extract the Mel Frequency Cepstral Coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data and store them in the data storage module; Controlled by a preset control signal, call the recognition model based on the Mel Frequency Cepstral Coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the monitoring result of children's snoring; Among them, the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are each input into four independent double-layer LSTM-SNN networks after being pulse-coded through threshold encoding, and the four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

[0049] The technical solution of this embodiment can comprehensively capture the respiratory acoustic characteristics during children's sleep by simultaneously extracting Mel-frequency cepstral coefficients (MFCC), filter bank coefficients (FBANK), linear prediction coefficients (LPC), and short-time average energy (STE). The four features are respectively input into independent double-layer LSTM-SNN networks to avoid interference between different features and retain their respective time dynamic information at the same time. After being combined through a fully connected layer, the softmax classifier comprehensively outputs the classification result of multi-modal data, significantly improving the classification robustness and accuracy. Snoring in children is mostly caused by adenoid hypertrophy, tonsil hyperplasia, and inferior turbinate hypertrophy, while in adults, it is mainly caused by soft palate vibration. The narrow nasopharyngeal cavity in children attenuates high-frequency sound waves weakly, resulting in the retention of high-frequency energy; the longer pharyngeal cavity in adults enhances low-frequency resonance. In children with OSA, due to the increased upper airway resistance, when the airway resumes patency after apnea, the airflow velocity surges suddenly, causing the tissue vibration to intensify, manifested as a sudden increase in loudness and an expansion of the formant bandwidth. Therefore, there are obvious acoustic differences between pediatric OSA and adult OSA. In this embodiment, MFCC simulates the auditory characteristics of the human ear, FBANK provides frequency-domain distribution information, LPC reflects the characteristics of the vocal tract, and STE is related to the respiratory intensity. The complementary multi-modal features enhance the ability to identify pediatric OSA-related abnormalities (such as apnea and hypopnea).

[0050] Preferably, the preprocessing of the audio data by the data processing module includes: Filtering the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000Hz~15000Hz; Framing the audio data after high-pass filtering according to a preset duration; Based on the fact that the high-order formant bandwidth of the snoring sound of children with OSA is 30%-50% wider than that of children without OSA and needs to envelope a larger frequency range, the audio data is resampled into a signal with a sampling rate of 30000Hz to ensure signal integrity and reduce the data processing burden of the mobile terminal and speed up the processing speed; Using a Wiener filter to perform noise reduction processing on the audio data, using spectral subtraction + variational Bayesian reconstruction Wiener gain; introducing a correction term , where is a morphological distortion penalty term used to adjust the signal gain to prevent over-suppression of some characteristics of children's snoring sounds, , is the observed noise spectrum, is the statistical feature extracted from the children's snoring database based on the prior spectrum of children's snoring, represents the KL divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural features of children's snoring; After introducing the spectral morphology prior, construct G(k,j) of variational Bayesian estimation, and rewrite the gain as: where controls the morphological distortion penalty through the variational Bayesian method to prevent the filter from over-suppressing the structural information in children's snoring.

[0051] Based on the same concept, the present invention also provides a readable storage medium, on which a processing program is stored. When the processing program is executed by a processor, it implements the multi-modal children's snoring monitoring system based on a mobile terminal described in any one of the above.

[0052] If the execution logic of the multi-modal children's snoring monitoring system based on a mobile terminal is implemented in the form of program instructions and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of software. This computer software is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0053] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific recognition content executed by the above-described system and device can refer to the corresponding process in the foregoing method embodiments.

[0054] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, provided that these changes fall within the scope of the claims of the present invention and its equivalent technologies, they still fall within the protection scope of the present invention.

Claims

1. A multi-modal children's snoring monitoring system based on a mobile terminal, characterized in that, The mobile terminal includes an audio acquisition module, a data storage module, and a data processing module. The data storage module stores the audio data obtained through the audio acquisition module. The data processing module preprocesses the audio data, extracts the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data respectively, and stores them in the data storage module. An identification model is also configured in the data storage module; Controlled by a preset control signal, the data processor module calls the identification model based on the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the monitoring result of children's snoring; Among them, after the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are respectively subjected to pulse coding processing through threshold coding, they are respectively input into four independent double-layer LSTM-SNN networks. The four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

2. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 1, characterized in that, The preprocessing of the audio data by the data processing module includes: Filtering the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000 Hz to 15000 Hz; Framing the audio data after high-pass filtering according to a preset duration.

3. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 2, characterized in that, The preprocessing of the audio data by the data processing module further includes: Based on the fact that the high-order resonance peak bandwidth of the snoring of OSA children is increased by 30%-50% compared with that of non-OSA children, and a larger frequency range needs to be enveloped, the audio data is resampled into a signal with a sampling rate of 30000 Hz to ensure signal integrity and reduce the data processing burden of the mobile terminal and speed up the processing speed.

4. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 1, characterized in that, The preprocessing of the audio data by the data processing module further includes: Using a Wiener filter to perform noise reduction processing on the audio data, and using spectral subtraction + variational Bayesian reconstruction Wiener gain; Introduce a correction term , where is a morphological distortion penalty term used to adjust the signal gain to prevent over-suppression of certain features of children's snoring, , is the observed noise spectrum, is the statistical feature extracted from the children's snoring database based on the prior spectrum of children's snoring, represents the KL divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural features of children's snoring; After introducing the spectral shape prior, constructing the variational Bayesian estimation of G(k,j), and rewriting the gain as: Among them The morphological distortion penalty is controlled by the variational Bayesian method to prevent the filter from over-suppressing the structural information in children's snoring sounds.

5. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 1, characterized in that The pulse coding processing of the Mel-frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy through threshold coding respectively further includes: Generating a truncated Gaussian distribution threshold adjustment threshold matrix with a mean of 0 and a variance of 0.8 to reduce the pulse trigger threshold of children's weak snoring; Detecting the local maximum values in the time domain window and frequency domain window in the spectrogram to strengthen the key point detection and enhance the recognition of sudden snoring.

6. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 1, wherein, Extracting the Mel-frequency cepstral coefficients of each frame of the audio data includes: The number of Mel-scale triangular filters is 40, covering the range of 50 - 8000 Hz, and focusing on enhancing the high-frequency harmonic concentration area of children's snoring in the frequency band of 1000 - 5500 Hz; The cepstral coefficient calculation uses an improved discrete cosine transform, and the first 13 dimensions of coefficients are retained to suppress high-frequency noise interference.

7. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 1, characterized in that, Extracting the filter bank coefficients of each frame of the audio data includes: Calculate the linear prediction coefficients using the Burg algorithm with an order of 14, and extract the formant positions of the upper respiratory tract of children through pole analysis: F1: 316, F2: 796, F3: 1425, F4: 2605, F5: 4510.

8. The multi-modal children's snoring monitoring system based on a mobile terminal according to claim 2, characterized in that, The input layer of the recognition model is as follows: Four branches are separately set to process Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy. The hidden layer of the recognition model is as follows: The first layer has 128 LSTM-SNN units with an activation threshold σ1 = 0.3; the second layer has 64 units with σ2 = 0.1 to adapt to the sparsity of children's signals. The output layer of the recognition model is as follows: The fully connected layer fuses multiple features, and Softmax outputs the probability of AHI grading.

9. A multi-modal children's snoring monitoring method based on a mobile terminal, applied to the monitoring system described in any one of claims 1 to 8, characterized in that, It includes the following steps: Store the audio data obtained by the audio acquisition module. Preprocess the audio data, separately extract the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy of each frame of the audio data, and store them in the data storage module. Controlled by a preset control signal, call the recognition model based on the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy for calculation to output the monitoring result of children's snoring. Among them, the Mel frequency cepstral coefficients, filter bank coefficients, linear prediction coefficients, and short-time average energy are respectively pulse-coded through threshold encoding and then input into four independent double-layer LSTM-SNN networks. The four independent double-layer LSTM-SNN networks are combined through a fully connected layer and then classified and judged through softmax.

10. The multi-modal children's snoring monitoring method based on a mobile terminal according to claim 9, wherein The preprocessing of the audio data by the data processing module includes: Filter the audio data through a band-pass filter, and the band-pass filtering range of the band-pass filter is 1000Hz - 15000Hz. Frame the high-pass filtered audio data according to a preset duration. Based on the fact that the high-order formant bandwidth of the snoring of OSA children is 30% - 50% wider than that of non-OSA children and needs to enclose a larger frequency range, resample the audio data into a signal with a sampling rate of 30000Hz to ensure signal integrity, reduce the data processing burden of the mobile terminal, and speed up the processing speed. The Wiener filter is used to perform noise reduction processing on the audio data, and the spectral subtraction + variational Bayesian reconstruction Wiener gain is used; a correction term is introduced , where is a morphological distortion penalty term used to adjust the signal gain to prevent over-suppression of certain features of children's snoring , is the observed noise spectrum is the statistical feature extracted from the children's snoring database based on the prior spectrum of children's snoring represents the KL divergence, which measures the difference between these two distributions. Using it to adjust the gain can effectively avoid over-filtering and retain the structural features of children's snoring After introducing the spectral shape prior, construct G(k,j) of variational Bayesian estimation and rewrite the gain as: Among them The morphological distortion penalty is controlled by the variational Bayesian method to prevent the filter from over-suppressing the structural information in children's snoring sounds.