Old people weakness real-time monitoring system based on multi-dimensional acoustic biomarkers
Through a multi-dimensional acoustic biomarker monitoring system, combined with a condenser microphone and accelerometer to collect acoustic parameters, logistic regression and random forest algorithms are used to achieve high-precision, low-cost real-time monitoring of frailty phenotypes, solving the problems of high invasiveness, high cost and insufficient single-dimensional assessment of existing technologies, and supporting real-time monitoring at home and early identification of frailty types.
Patent Information
- Application Number
- CN202510953472.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-05
AI Technical Summary
Existing frailty assessment technologies are highly invasive, costly, insufficient in single-dimensional assessment, and difficult to apply in real time in home/community settings.
A multi-dimensional acoustic biomarker monitoring system is used, including a sound collection unit, a background noise collection unit, a head posture detection unit and a dynamic prediction model. Acoustic parameters are collected through a condenser microphone, accelerometer and ambient noise sensor, and real-time frailty phenotype classification is performed in combination with logistic regression and random forest algorithms.
It achieves high-precision frailty phenotype classification, supports real-time monitoring at home/community, reduces equipment costs, identifies frailty types early and provides precise intervention recommendations.
Smart Images

Figure CN120585318A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical health monitoring, and in particular to a real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers. Background Art
[0002] There are two main approaches to frailty assessment in the existing frailty assessment technology system. The first relies on clinical scales for assessment. Commonly used clinical scales include the CHS, Fried phenotype, and FI index. These scales rely on specific scoring criteria and indicator systems to provide a basis for determining the degree of frailty. The second approach is to assess frailty through invasive biomarker testing. For example, testing for inflammatory factors, muscle mass, and other indicators to obtain biological information reflecting an individual's frailty status. In recent years, with the continuous advancement of science and technology and the continued deepening of research, the emerging concept of digital biomarkers has been introduced into the field of frailty assessment. Among them, data obtained through wearable sensors monitoring gait, grip strength, etc. fall into the category of digital biomarkers, which provide new perspectives and methods for frailty assessment.
[0003] However, although digital biomarkers bring new possibilities, current frailty assessment technology as a whole still has some urgent problems that need to be solved: ① Invasiveness and cost: blood tests and imaging analysis require professional equipment and are expensive; ② Single dimension: Existing wearable devices focus more on physical activity and cannot distinguish frail phenotypes; ③ Traditional acoustic analysis usually requires a laboratory environment and is difficult to apply in real time in home / community scenarios.
[0004] EBF: Energy-related Frailty Phenotype (related to energy, activity, fatigue, depression).
[0005] SBF: Sarcopenia-related Frailty Phenotype (associated with loss of muscle mass / strength, falls, and fractures).
[0006] HBF-E: Hybrid Frailty Phenotype-Energy Dominant.
[0007] HBF-S: Hybrid Frailty Phenotype-Sarcopenia Dominant. Summary of the Invention
[0008] The technical problem to be solved by the present invention is: to address the shortcomings of the existing technology and provide a real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers.
[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0010] A real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers, comprising a sound collection unit, a background noise collection unit, an acoustic parameter collection unit, a head posture detection unit, a dynamic prediction model, and a cloud platform; the background noise collection unit is used to collect ambient noise information around the user for noise suppression;
[0011] The acoustic parameter acquisition unit is used to perform signal preprocessing on the user's voice information and obtain the user's acoustic parameters based on the preprocessed user's voice information. The acoustic parameters include the average number of zero crossings, local peak-to-valley changes, formant frequency changes, and spectral energy ratio;
[0012] The head posture detection unit is used to detect the head posture angle, which is input into the dynamic prediction model together with the output of the acoustic parameter acquisition unit;
[0013] The dynamic prediction model receives the A1-A4 parameters output by the acoustic parameter acquisition unit and the Δθ output by the head posture detection unit, and outputs the frailty phenotype classification result and the corresponding probability;
[0014] The cloud platform communicates bidirectionally with the dynamic prediction model to record historical acoustic parameters, posture data and frailty phenotype probabilities and optimize model parameters.
[0015] As a further improvement, the sound collection unit is a condenser microphone, and the frequency response range of the condenser microphone is 50Hz-16kHz.
[0016] As a further improvement, the background noise collection unit is an environmental noise sensor.
[0017] For further improvement, the steps of signal preprocessing are as follows:
[0018] ① Endpoint detection: Use the double threshold method to detect the endpoint of the user's voice information to obtain valid voice;
[0019] ② Collect background noise through the environmental noise sensor and calculate the average spectrum of its silent section as the noise spectrum |N(ω)|;
[0020] ③ Perform short-time Fourier transform on the effective speech to obtain the amplitude spectrum of the noisy speech |Y(ω)|.
[0021] Subtract the noise template from |Y(ω)| to obtain the preprocessed speech Y_enhanced(ω):
[0022] |Y_enhanced(ω)|=max(|Y(ω)|-α*|N(ω)|,β*|N(ω)|) where |Y(ω)| is the amplitude spectrum of the noisy speech and |N(ω)| is the estimated noise spectrum;
[0023] α is the over-subtraction factor, which is a coefficient greater than 1 and is used to over-subtract the noise spectrum in spectral subtraction; β is the spectrum lower limit constraint, which is a coefficient between 0 and 1 and is used to set the minimum amplitude lower limit of the processed spectrum to prevent excessive suppression from causing speech distortion and generating "musical noise"; max means taking the maximum value;
[0024] The over-subtraction factor α and the spectrum lower limit coefficient β are dynamically adjusted according to the real-time signal-to-noise ratio (SNR): when the SNR is higher than a first preset threshold, a smaller α value and a smaller β value are used; when the SNR is lower than a second preset threshold, a larger α value and a larger β value are used. For example:
[0025] ① When the signal-to-noise ratio (SNR) is higher than the preset threshold of 20dB: the α value range is 1.5-2, and the β value range is 0.05-0.1 to avoid excessive distortion;
[0026] ② When the signal-to-noise ratio (SNR) is lower than the preset threshold of 10dB: the α value range is 3-5, and the β value range is 0.2-0.4, ensuring that the noise is effectively suppressed while controlling the residual noise level;
[0027] SNR≈10*log10((speech signal energy) / (noise signal energy)).
[0028] As a further improvement, the head posture detection unit is an accelerometer, which is used to monitor the head posture angle and construct a posture offset matrix:
[0029]
[0030] When the head posture detection unit determines that Δθ>15°, the system discards the acoustic parameters of the current analysis window and triggers a re-collection instruction;
[0031] Where Δθ represents the total offset of the head posture angle, α, β, and γ are Euler angles, where α represents the pitch angle, β represents the yaw angle, and γ represents the roll angle. Δα, Δβ, and Δγ are the angular deviations of the current posture angle (α, β, and γ) relative to a reference posture (the initial posture of the user when standing upright).
[0032] As a further improvement, the method for obtaining the acoustic parameters is as follows:
[0033] After the user pronounces the vowel " / a / " and undergoes signal preprocessing, the dynamic time warping method is used to align the speech segments, extract stable pronunciation segments with a duration of ≥ 1 second, and then intercept 500ms of the stable pronunciation segments as the analysis window: Average zero crossings A1: The number of zero crossings in each 10ms window within the analysis window is calculated and the average is taken as the average zero crossings A1:
[0034]
[0035] Among them, x i [n] is the i-th frame signal, L is the frame length, and N is the total number of frames; sgn() is the sign function, and n is the signal sample point index, ranging from 1 to L in each frame;
[0036] Local peak-to-valley variation A2: Calculate the peak-to-valley difference of the signal within every 20ms window of the analysis window, and take the variance as the stability indicator:
[0037] A2=Var(max(x k )-min(x k )),k=1,2,...,M
[0038] Where Var represents the variance function, max() represents the maximum value, min() represents the minimum value of the signal in the specified window, M represents the total number of 20ms windows divided in the analysis window, k represents the index of the 20ms window, k = 1, 2, ..., M; Var() is the variance function;
[0039] Formant frequency variation A3: Linear predictive coding is used to extract the first three formants in the analysis window and calculate their standard deviation:
[0040]
[0041] Among them F t represents the formant frequency vector extracted at the tth frame, T represents the total number of valid frames used to calculate the formant standard deviation within the analysis window, Represents the frequency mean vector of the three resonance peaks F1, F2, and F3 within the analysis window; Spectral energy ratio A4: Divide the spectrum of the analysis window into low-frequency 50-500Hz and high-frequency 2-4kHz, and calculate the energy ratio:
[0042]
[0043] X(f) represents (). This is the spectrum representation of the signal, usually obtained by converting the time domain signal to the frequency domain through Fourier transform. X(f) represents the spectral component of the signal at frequency f.
[0044] Represents the sum of the energy of all spectral components from 2-4kHz.
[0045] It represents the sum of the energy of all spectral components from 50Hz to 500Hz.
[0046] For further improvement, the dynamic prediction model first normalizes the input vector [A1, A2, A3, A4, Δθ] to the interval [0, 1]; then the probability P of frailty phenotype i is obtained through the Logistic regression analysis model and the random forest RF model. final (y=i|x), take the frailty phenotype corresponding to the maximum probability value (if there are multiple identical maximum values, output the phenotype with the highest risk level) as the output result:
[0047] P final (y=i|x)=w·P RF (y=i|x)+(1-w)·P LR (y=i|x)
[0048] P final (y=i|x): The final predicted probability of the input feature belonging to frailty phenotype i by the Logistic regression analysis model, with a value range of [0,1]
[0049] P RF (y = i | x) probability of frailty phenotype i predicted by the random forest RF model;
[0050] w is the adjustment coefficient, obtained through pre-training.
[0051] As a further improvement, the frailty phenotype includes energy-based frailty EBF, sarcopenia-based frailty SBF, energy-based mixed frailty phenotype HBF-E, and sarcopenia-based mixed frailty phenotype HBF-S.
[0052] A further improvement also includes a mobile app for receiving and displaying the output results of the dynamic prediction model.
[0053] A wearable device comprising a neck-mounted housing with an integrated condenser microphone, a three-axis accelerometer, and an ambient noise sensor, connected to a mobile app via Bluetooth Low Energy. The wearable device is used to run the aforementioned real-time elderly frailty monitoring system based on multi-dimensional acoustic biomarkers.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. High-precision classification: The A1-A4 parameter combination improves accuracy by 30% compared to a single biomarker
[0056] 2. Real-time and portability: supports home / community scenarios and does not require professional operation;
[0057] 3. Early identification of EBF (associated depression) and SBF (associated bone fracture) to facilitate precise intervention;
[0058] 4. Cost advantage: The unit price of the equipment is 80% lower than that of traditional detection methods, making it suitable for large-scale promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic diagram of the algorithm flow of the present invention. DETAILED DESCRIPTION
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] (1) Hardware module:
[0062] Sound collection unit: directional condenser microphone (frequency response range 50Hz-16kHz), synchronously recording pronunciation signals;
[0063] Auxiliary sensors: accelerometer (monitoring head posture stability), ambient noise sensor (dynamic noise reduction);
[0064] Data processing unit: low-power embedded chip (embedded microprocessor) that supports real-time signal processing.
[0065] (2) Software algorithm:
[0066] Signal preprocessing: endpoint detection, noise suppression, and standardized vowel “ / a / ” extraction;
[0067] Feature extraction: Calculate A1-A4 parameters and correct pronunciation deviations based on motion data;
[0068] Classification model: Based on logistic regression and random forest algorithm, it outputs the probability of frailty phenotype.
[0069] (3) User interface:
[0070] Mobile App: Displays real-time monitoring results, risk warnings, and health recommendations;
[0071] Cloud platform: stores historical data and supports remote review and analysis by doctors.
[0072] (1) Hardware module:
[0073] Sound collection unit: directional condenser microphone (frequency response range 50Hz-16kHz), synchronously recording pronunciation signals;
[0074] Auxiliary sensors: accelerometer (monitoring head posture stability), environmental noise sensor (dynamic noise reduction); data processing unit: low-power embedded chip (embedded microprocessor) that supports real-time signal processing.
[0075] (2) Software algorithm:
[0076] Signal preprocessing: endpoint detection, noise suppression, and standardized vowel “ / a / ” extraction;
[0077] Feature extraction: Calculate A1-A4 parameters and correct pronunciation deviations based on motion data;
[0078] Classification model: Based on logistic regression and random forest algorithm, it outputs the probability of frailty phenotype.
[0079] (3) User interface:
[0080] Mobile App: Displays real-time monitoring results, risk warnings, and health recommendations;
[0081] Cloud platform: stores historical data and supports remote review and analysis by doctors.
[0082] Acoustic parameter combination analysis algorithm: collects and analyzes the following four acoustic biomarker parameters in real time:
[0083] A1 (average number of zero crossings): reflects airflow and respiratory function and is used to identify energy-based failure (EBF);
[0084] A2 (local peak-to-trough variation): assesses glottal airflow stability for identification of sarcopenic weakness (SBF);
[0085] A3 (formant frequency change) and A4 (spectral energy ratio): Jointly analyze complex weakness phenotypes (HBF-E / HBF-S). Embedded multimodal sensor fusion: Integrates a high-precision condenser microphone, motion sensor (to monitor pronunciation stability), and ambient noise suppression module to ensure the accuracy of acoustic data.
[0086] Dynamic prediction model: Based on logistic regression and dynamic prediction model, frailty phenotype (EBF / SBF / HBF-E / HBF-S) is classified in real time and a personalized risk score is generated.
[0087] Non-invasive wearable design: The device is implemented in the form of a neck-worn or ear-worn style, supports long-term wearing and remote data transmission, and is adapted to a smartphone application for real-time feedback.
[0088] 1. Signal preprocessing
[0089] ① Endpoint detection: The double-threshold method is adopted, combined with short-time energy (STE) and short-time zero-crossing rate (ZCR):
[0090] Short-time energy: Calculate the sum of squares of the signal frames, and set an energy threshold to distinguish speech segments from silent segments.
[0091] Short-time zero-crossing rate: Count the number of times the signal crosses zero, and assist in identifying the清音 / 浊音 boundary.
[0092] Dynamic threshold adjustment: Dynamically adjust the double-threshold values according to the ambient noise sensor data to enhance robustness. It is necessary to use both the "user's voice information" (speech signal) and the ambient noise information collected by the "ambient noise acquisition unit" simultaneously.
[0093] Calculation method:
[0094] ① Calculate the ambient noise level (N): Usually take the short-time energy (STE_noise) or RMS value of the signal collected by the ambient noise sensor in the silent segment (when there is no user speech).
[0095] ② Set the reference threshold: In a quiet environment (N < N0, N0 is the preset quiet threshold), set the initial energy thresholds (STE_th_high_base, STE_th_low_base) and the zero-crossing rate threshold (ZCR_th_base).
[0096] ③ Dynamic adjustment: Adjust the threshold according to the real-time noise level N:
[0097] STE_th_high = STE_th_high_base + k_ste * NSTE_th_low = STE_th_low_base + k_ste * N (k_ste is the energy adjustment coefficient, usually < 1)
[0098] ZCR_th = ZCR_th_base + k_zcr * N (k_zcr is the zero-crossing rate adjustment coefficient)
[0099] For example: k_ste can be taken as 0.5, k_zcr can be taken as 0.2, and N0 can be taken as -40dB
[0100] ② Noise suppression: Adopt an improved spectral subtraction method:
[0101] 1. Collect the background noise spectrum through the ambient noise sensor and construct a noise template.
[0102] 2. Perform a short-time Fourier transform (STFT) on the speech signal and subtract the noise template from the amplitude spectrum.
[0103] 3. Introduce over-subtraction factor and spectrum lower limit constraint to avoid residual music noise.
[0104] Oversubtraction factor (α): A coefficient greater than 1 (e.g., α = 2-4). Used to oversubtract the noise spectrum during spectral subtraction to ensure adequate noise suppression. Spectral floor constraint (β): A coefficient between 0 and 1 (e.g., β = 0.1-0.3). Used to set the minimum amplitude limit of the processed spectrum (relative to the noise spectrum) to prevent oversuppression from causing speech distortion and "musical noise."
[0105] |Y_enhanced(ω)|=max(|Y(ω)|-α*|N(ω)|,β*|N(ω)|), where |Y(ω)| is the amplitude spectrum of the noisy speech and |N(ω)| is the estimated noise spectrum.
[0106] Avoid residual music noise and dynamically adjust according to the ambient noise sensor data;
[0107] The over-subtraction factor α and the spectrum lower limit coefficient β can be dynamically adjusted according to the real-time signal-to-noise ratio (SNR):
[0108] ① High SNR (quiet): Use a smaller α (such as 1.5-2) and a smaller β (such as 0.05-0.1) to avoid excessive distortion. ② Low SNR (noisy): Use a larger α (such as 3-5) and a larger β (such as 0.2-0.4) to ensure that noise is effectively suppressed while controlling the residual noise level.
[0109] SNR≈10*log10((speech signal energy) / (noise signal energy))
[0110] ③ Standardized vowel extraction
[0111] Based on the pronunciation stability of the vowel " / a / ", dynamic time warping (DTW) was used to align speech segments. Stable pronunciation segments lasting ≥ 1 second were extracted, and the middle 500ms was used as the analysis window.
[0112] 2. Feature Extraction
[0113] ①A1 (average number of zero crossings): Calculate the number of zero crossings in each 10ms window within the standardized vowel segment and take the average as A1:
[0114]
[0115] Among them, xi[n] is the i-th frame signal, L is the frame length, and N is the total number of frames.
[0116] ②A2 (local peak-to-valley variation): Calculate the peak-to-valley difference of the signal within a 20ms window and use the variance as a stability indicator:
[0117] A2=Var(max(x k )-min(x k )),k=1,2,...,M
[0118] ③A3 (formant frequency variation): Linear predictive coding (LPC) is used to extract the first three formants (F1-F3) and calculate their standard deviation:
[0119]
[0120] ④A4 (spectral energy ratio): Divide the spectrum into low frequency (50-500Hz) and high frequency (2-4kHz) and calculate the energy ratio:
[0121]
[0122] 3. Sensor data fusion: pronunciation stability correction
[0123] The head attitude angle (pitch, yaw) is monitored by the accelerometer and the attitude offset matrix is constructed:
[0124] Δθ=α·Acc x +β·Acc y
[0125] When the head posture detection unit determines that Δθ>15°, the system discards the acoustic parameters of the current analysis window and triggers a re-collection instruction.
[0126] 4. Classification Model
[0127] ① Feature input
[0128] Input vector: [A1, A2, A3, A4, Δθ], normalized to the interval [0, 1].
[0129] ②Model architecture
[0130] Using a mixed model:
[0131] 1) Logistic regression.
[0132] 2) Random forest (100 trees): handles nonlinear relationships and outputs the probability of frailty phenotype (EBF / SBF / HBFE / HBFS).
[0133] 3) Dynamic weight allocation: Adjust the output weights of the two models based on the real-time data confidence.
[0134] ③ Training optimization
[0135] Dataset: Contains 500 acoustic samples of elderly people (labeled with frailty phenotypes).
[0136] The dataset contains acoustic samples of 500 elderly people aged 65 and above, collected from the geriatric department of a tertiary hospital. Clinicians annotated the frailty types (EBF / SBF / HBF-E / HBF-S) according to the Fried frailty phenotype criteria. Sample distribution: 120 cases of EBF, 130 cases of SBF, 110 cases of HBF-E, and 140 cases of HBF-S.
[0137] Cross-validation: 5-fold cross-validation, accuracy ≥ 92%.
[0138] Based on the test results, EBF may be more dependent on A1 (respiratory airflow), SBF may be more dependent on A2 (glottal stability), and HBF-E / HBF-S may be more dependent on A3 / A4 (spectral characteristics). These results help us understand which features are more important for distinguishing all frailty phenotypes overall and can be used to guide computational resource allocation during model simplification and embedded optimization.
[0139] 5. Embedded Optimization
[0140] ① Lightweight computing: Fixed-point arithmetic is deployed on the STM32 chip to reduce floating-point overhead. A sliding window mechanism is used to update classification results every 200ms. The LPC coefficient table is pre-calculated to reduce real-time computation.
[0141] 6. Clinical Validation
[0142] Tested in 100 community-dwelling elderly people
[0143] The overall accuracy of this system was 92.3% (95% CI: 90.1–94.5%);
[0144] Compared with single biomarker methods (using only A1 or A2), the accuracy increased by 28.5% to 32.7%;
[0145] The unit price of the device is US$80, while the average unit price of traditional testing methods (such as handgrip dynamometer + gait analyzer) is US$400, which reduces the cost by 80%.
[0146] 7. Algorithm process as follows Figure 1 As shown:
[0147] Signal input → endpoint detection → noise suppression → vowel extraction → feature calculation → sensor fusion → classification model → risk score output.
[0148] The process implements closed-loop operation on the embedded chip, ensuring low power consumption (<50mW) and high real-time performance.
[0149] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers, characterized by: It includes a sound collection unit, a background noise collection unit, an acoustic parameter collection unit, a head posture detection unit, a dynamic prediction model and a cloud platform; the sound collection unit is used to collect the user's sound information; The background noise collection unit is used to collect environmental noise information around the user for noise suppression; The acoustic parameter acquisition unit is used to perform signal preprocessing on the user's voice information and obtain the user's acoustic parameters based on the preprocessed user's voice information. The acoustic parameters include the average number of zero crossings, local peak-to-valley changes, formant frequency changes, and spectral energy ratio; The head posture detection unit is used to detect the head posture angle, which is input into the dynamic prediction model together with the output of the acoustic parameter acquisition unit; The dynamic prediction model receives the acoustic parameters output by the acoustic parameter acquisition unit and the Δθ output by the head posture detection unit, and outputs the frailty phenotype classification result and the corresponding probability; The cloud platform communicates bidirectionally with the dynamic prediction model to record historical acoustic parameters, posture data and frailty phenotype probabilities and optimize the parameters of the dynamic prediction model.
2. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The sound collection unit is a condenser microphone, and the frequency response range of the condenser microphone is 50Hz-16kHz.
3. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The background noise collection unit is an environmental noise sensor.
4. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The steps for signal preprocessing are as follows: ① Endpoint detection: Use double threshold method to detect the endpoint of the user's voice information to obtain valid voice; ② Collect background noise through the environmental noise sensor and calculate the average spectrum of its silent section as the noise spectrum |N(ω)|; ③ Perform short-time Fourier transform on the effective speech to obtain the noisy speech amplitude spectrum |Y(ω)|. Subtract the noise template from the noisy speech amplitude spectrum |Y(ω)| to obtain the preprocessed speech Y_enhanced(ω): |Y_enhanced(ω)|=max(|Y(ω)|-α*|N(ω)|,β*|N(ω)|) where |Y(ω)| is the amplitude spectrum of the noisy speech and |N(ω)| is the estimated noise spectrum; α is the over-subtraction factor, a coefficient greater than 1, used to over-subtract the noise spectrum in spectral subtraction; β is the spectrum lower limit constraint, a coefficient between 0 and 1, used to set the minimum amplitude lower limit of the processed spectrum to prevent excessive suppression from causing speech distortion and generating "musical noise"; max means taking the maximum value; The over-subtraction factor α and the spectrum lower limit coefficient β are dynamically adjusted according to the real-time signal-to-noise ratio SNR: ① When the signal-to-noise ratio (SNR) is higher than the preset threshold of 20dB: the α value range is 1.5-2, and the β value range is 0.05-0.1 to avoid excessive distortion; ② When the signal-to-noise ratio (SNR) is lower than the preset threshold of 10dB: the α value range is 3-5, and the β value range is 0.2-0.4, ensuring that the noise is effectively suppressed while controlling the residual noise level; SNR≈10*log10((speech signal energy) / (noise signal energy)).
5. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The head posture detection unit is an accelerometer, which is used to monitor the head posture angle and construct a posture offset matrix: If Δθ>15°, the pronunciation is determined to be unstable, the acoustic parameters of the current analysis window are discarded, and a re-collection instruction is triggered; Where Δθ represents the total offset of the head posture angle, α, β, and γ are Euler angles, where α represents the pitch angle, β represents the yaw angle, and γ represents the roll angle. Δα, Δβ, and Δγ are the angular deviations of the current posture angle (α, β, and γ) relative to a reference posture, where the reference posture is the initial posture of the user when standing upright.
6. The real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers according to claim 5, characterized in that: The acoustic parameter acquisition method is as follows: All parameters are calculated in the noise-reduced speech segment. After the user pronounces the vowel " / a / " and performs signal preprocessing, the speech segments are aligned using the dynamic time warping method, and stable pronunciation segments with a duration of ≥ 1 second are extracted. 500ms of the stable pronunciation segment are then intercepted as the analysis window: Average number of zero crossings A1: The number of zero crossings per 10ms window in the analysis window is calculated, and the average is taken as the average number of zero crossings A1: Among them, x i [n] is the i-th frame signal, L is the frame length, and N is the total number of frames; sgn() is the sign function, and n is the signal sample point index, ranging from 1 to L in each frame; Local peak-to-valley variation A2: Calculate the peak-to-valley difference of the signal within every 20ms window of the analysis window, and take the variance as the stability indicator: A2=Var(max(x k )-min(x k )), k = 1, 2, ..., M, where Var represents the variance function, max() represents the maximum value, min() represents the minimum value of the signal in the specified window, M represents the total number of 20ms windows divided in the analysis window, k represents the index of the 20ms window, k = 1, 2, ..., M; Var() is the variance function; Formant frequency variation A3: Linear predictive coding is used to extract the first three formants in the analysis window and calculate their standard deviation: Among them F t represents the formant frequency vector extracted at the tth frame, T represents the total number of valid frames used to calculate the formant standard deviation within the analysis window, Indicates the frequency mean vector of the three resonance peaks F1, F2, and F3 within the analysis window; Spectral energy ratio A4: Divide the spectrum of the analysis window into low-frequency 50-500Hz and high-frequency 2-4kHz, and calculate the energy ratio: X(f) represents the spectral representation of the signal (usually obtained by Fourier transform) converted from the time domain signal to the frequency domain; X(f) represents the spectral component of the signal at frequency f; Represents the sum of the energy of all spectral components from 2-4kHz. It represents the sum of the energy of all spectral components from 50Hz to 500Hz.
7. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The dynamic prediction model first normalizes the input vector [A1, A2, A3, A4, Δθ] to the interval [0, 1]; then the probability P of frailty phenotype i is obtained through the Logistic regression analysis model and the random forest RF model. final (y=i|x), take the frailty phenotype corresponding to the maximum probability value (if there are multiple identical maximum values, output the phenotype with the highest risk level) as the output result: P final (y=i|x)=w·P RF (y=i|x)+(1-w)·P LR (y=i|x) P final (y=i|x): The final predicted probability of the input feature belonging to frailty phenotype i by the Logistic regression analysis model, with a value range of [0,1] P RF (y = i | x) probability of frailty phenotype i predicted by the random forest RF model; w is the adjustment coefficient, obtained through pre-training.
8. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1 is characterized in that: The frailty phenotypes include energy-based frailty EBF, sarcopenia-based frailty SBF, mixed frailty phenotype HBF-E with energy as the main factor, and mixed frailty phenotype HBF-S with sarcopenia as the main factor.
9. The real-time monitoring system for elderly frailty based on multi-dimensional acoustic biomarkers according to claim 1, characterized in that: It also includes a mobile app for receiving and displaying the output results of the dynamic prediction model.
10. A device, characterized in that The device is a wearable device, which includes a neck-hanging shell with an integrated condenser microphone, a three-axis accelerometer and an ambient noise sensor, and is connected to a mobile app via low-power Bluetooth; the wearable device is used to run the real-time monitoring system for frailty in the elderly based on multi-dimensional acoustic biomarkers as described in any one of claims 1-9.