Rapid voice screening system and method for acute stroke
By calculating the energy distribution of phonemes and dynamic noise reduction in the frequency band, analyzing the pause and respiratory flow cycle characteristics of the speech signal, and building an individual baseline model, solving the problem of insufficient recognition ability in noise environments in the existing technology, and achieving rapid and accurate speech screening for acute stroke.
Patent Information
- Application Number
- CN202510431874.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing acute stroke speech screening technology has limited recognition capabilities in a noisy environment, and its feature extraction is not comprehensive enough, making it difficult to adapt to individual differences, resulting in limited accuracy and efficiency of screening results, which cannot meet the needs of rapid diagnosis.
By calculating the probability of phoneme energy distribution in the frequency band, dynamically adjusting the noise reduction weight, analyzing the pause and respiratory flow cycle characteristics of the speech signal, extracting the formant frequency and the energy distribution between phonemes, constructing an individual baseline model and comparing it with the stroke speech database, and optimizing the feature modeling and matching process.
It improves the clarity and usability of speech data, enhances the accuracy of speech pause feature recognition, improves the efficiency and accuracy of stroke screening, and achieves fast and accurate early diagnosis support.
Smart Images

Figure CN120260545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical speech recognition, and particularly to a rapid speech screening system and method for acute stroke. Background Art
[0002] The technical field of medical speech recognition involves using speech signal processing and analysis technologies to identify, classify, and extract medical-related information. By processing speech data with a computer, it assists in medical diagnosis, condition monitoring, rehabilitation evaluation, and patient communication. Medical speech recognition covers multiple links such as speech signal acquisition, feature extraction, pattern matching, speech transcription, and medical information analysis, and is applied to various scenarios such as disease screening, medical record generation, and speech-assisted medical treatment. The development of this technology mainly relies on speech feature analysis, machine learning modeling, medical knowledge graphs, and speech processing methods based on neural networks to efficiently identify the speech features of patients, extract key information, and assist in medical decision-making.
[0003] Among them, the rapid speech screening system for acute stroke refers to collecting, extracting features, and performing pattern matching on the speech signals of patients to identify acute stroke symptoms. The system mainly targets the impaired language ability of stroke patients and covers multiple technical links such as real-time speech signal collection, audio signal preprocessing, extraction of abnormal speech features, and comparative analysis of speech patterns. The specific methods include using audio input devices such as microphones to collect speech signals, performing signal preprocessing through time-frequency transformation methods to remove noise and irrelevant frequency components, and extracting abnormal speech features through acoustic feature parameters such as pronunciation pauses, speech rate changes, and abnormal pitch. Pattern matching technology is used to compare the speech features of patients with stroke-related abnormal speech patterns to judge the stroke risk of patients. The system relies on the constructed speech data model and stroke pathological feature library, and analyzes in combination with the real-time speech data of patients to complete the rapid screening task.
[0004] Traditional acute stroke speech screening technologies have certain limitations in the process of speech signal processing and analysis. They rely on time-frequency transformation methods for signal preprocessing, and have limited ability to distinguish noise in complex environments, resulting in the loss of key speech information. In the feature extraction link, single acoustic feature parameters are mostly used, lacking comprehensive analysis of the stability of formant frequency, bandwidth, and energy distribution, resulting in insufficient comprehensiveness of feature extraction and affecting the accuracy of subsequent pattern matching. In the pattern matching process, it relies on pre-constructed speech data models and stroke pathological feature libraries for comparative analysis, lacking dynamic modeling and real-time evaluation of individual speech features, and it is difficult to adapt to the individual differences of different patients, resulting in weak generalization ability of screening results. The screening accuracy and efficiency in complex environments are limited, and it is difficult to meet the needs of clinical rapid diagnosis. Summary of the Invention
[0005] The object of the present invention is to solve the disadvantages existing in the prior art, and to propose a rapid voice screening system and method for acute stroke.
[0006] To achieve the above object, the present invention adopts the following technical solution: A rapid voice screening system for acute stroke includes:
[0007] The signal preprocessing module acquires the voice signal of the screening object, performs time-frequency transformation, obtains the energy distribution data of phonemes in multiple frequency bands, calculates the occurrence probability of phonemes in each frequency band, screens the abnormal phoneme intervals in the voice signal, divides the key voice frequency band, secondary voice frequency band and environmental noise frequency band, and sets a noise reduction weight for each frequency band. Combining the dynamic changes of the signals in each frequency band, the frequency band classification boundary is updated to obtain the noise-reduced voice data;
[0008] The pause analysis module calls the noise-reduced voice data, detects and marks the pause timestamps of the voice signal, obtains the respiratory airflow cycle data of the stroke screening object, extracts the respiratory peak and valley values, calculates the offset value of the pause point relative to the respiratory cycle, obtains the respiratory phase angle of the voice pause point, establishes the alignment degree parameter between the pause point and the respiratory rhythm, identifies the distribution pattern of the pause point, and obtains the phase offset index;
[0009] The feature modeling module calls the phase offset index, uses the voice signal to obtain the voice baseline feature sample of the stroke screening object, extracts the voice spectrum features through short-time Fourier transform, calculates the formant frequency, bandwidth, and energy distribution between phonemes of the voice signal, obtains the stability of the spectrum features, calls the stability parameter of the spectrum features, records the time drift range of the formant, calls the drift range, screens the spectrum data, and obtains the individual baseline model;
[0010] The real-time evaluation module calls the individual baseline model to obtain the time features of the model, including syllable duration, pause interval, and speech rate change rate, calculates the time deviation degree of the current voice sample, compares it with the time features of the individual baseline model to obtain the time deviation calculation value, obtains the frequency features of the individual baseline model, including fundamental frequency distribution, formant trajectory, and spectrum energy ratio, obtains the voice detection sample of the stroke screening object, calculates the frequency deviation degree of the current voice sample, compares it with the frequency features of the individual baseline model to obtain the frequency deviation calculation value, calls the time deviation calculation value and the frequency deviation calculation value, and obtains the comprehensive feature deviation rate;
[0011] The feature matching module calls the comprehensive feature deviation rate, obtains the formant time series data of the stroke screening object, calculates the trajectory changes of formant frequency, bandwidth, and amplitude, calls the trajectory change data, calculates the offset rate of the formant at multiple time points, analyzes the smoothness of the formant trajectory between multiple syllables, calls the trajectory smoothness data, calculates the formant time series trajectory mapping data, calls the trajectory mapping data, and calculates the feature matching degree by comparing and analyzing with the voice features of known stroke patients, screens and records stroke patients, and generates a stroke screening record.
[0012] As a further solution of the present invention, the noise-reduced voice data includes frequency band energy distribution adjustment values, voice segment recognition results, and noise reduction weight coefficients. The phase offset degree index includes pause timestamps, peak and valley values of the respiratory airflow cycle, and pause point phase distribution patterns. The individual baseline model includes the energy stability between phonemes, the formant time drift range, and the screened spectral data. The comprehensive feature deviation rate includes speech rate fluctuation deviation, fundamental frequency distribution deviation, and formant trajectory deviation. The stroke screening record includes formant trajectory parameters, trajectory offset rate, and stroke voice matching degree.
[0013] As a further solution of the present invention, the signal preprocessing module includes:
[0014] The frequency band energy distribution calculation sub-module obtains the voice signal of the screening object, performs time-frequency transformation, extracts the energy distribution data of phonemes in multiple frequency bands, calculates the occurrence probability of phonemes in multiple frequency bands, detects the abnormal phoneme distribution area, calls the abnormal phoneme distribution area data, screens the frequency band range of the voice signal, and obtains the frequency band energy distribution parameters;
[0015] The abnormal phoneme interval screening sub-module calls the frequency band energy distribution parameters, calculates the energy change amplitude of phonemes in multiple frequency bands, obtains the energy fluctuation curve of phonemes on the time axis, detects the intervals with prominent energy changes in the voice signal, and obtains the abnormal phoneme interval data;
[0016] The dynamic noise reduction weight adjustment sub-module calls the abnormal phoneme interval data, screens the frequency bands corresponding to abnormal phonemes, analyzes the energy change rate of the abnormal phoneme interval, calculates the relative proportion of the abnormal phoneme interval in the voice signal, updates the frequency band classification boundary, and uses the formula:
[0017]
[0018] Calculates the adjusted noise reduction weight parameter, adjusts the noise reduction weight of the abnormal phoneme interval, sets the weight ratio of the key voice frequency band, the secondary voice frequency band, and the environmental noise frequency band, calculates the dynamic changes of the signals in multiple frequency bands, performs noise reduction processing on the voice signal, and obtains the noise-reduced voice data;
[0019] Among them, W adj represents the adjusted noise reduction weight, and W base represents the basic noise reduction weight, E var represents the energy change value in the abnormal phoneme interval, and E norm represents the average energy in the normal phoneme interval, and ∈ is a small value to prevent the denominator from being zero, and E dev represents the energy deviation in the abnormal phoneme interval, and E sum represents the total energy value.
[0020] As a further solution of the present invention, the pause analysis module includes:
[0021] The pause timestamp detection sub-module calls the noise-reduced speech data, identifies the user's pauses in the speech signal by analyzing the time-axis characteristics of the speech signal, calculates the pause duration, combines the spectral changes of the speech signal, identifies the distribution state of the pause timestamps in the speech, and obtains the speech pause timestamp data;
[0022] The respiratory airflow period calculation sub-module calls the speech pause timestamp data, obtains the respiratory airflow data of the stroke screening object, identifies the respiratory cycle by extracting the peaks of the respiratory signal, matches the respiratory phase curve, and obtains the respiratory cycle data;
[0023] The phase offset degree calculation sub-module calls the respiratory cycle data, identifies the position of the pause point in the respiratory cycle, calculates the time offset of the pause point relative to the respiratory peak, identifies the distribution pattern of the pause point, and uses the formula:
[0024]
[0025] calculates the phase offset angle of the pause point, establishes the alignment degree parameter between the pause point and the respiratory rhythm, and obtains the phase offset degree index;
[0026] Among them, θ offset represents the phase offset angle of the pause point, t pause represents the time point when the pause occurs, t peak represents the time point of the nearest respiratory peak, T cycle represents the duration of a complete respiratory cycle.
[0027] As a further solution of the present invention, the feature modeling module includes:
[0028] The baseline feature extraction sub-module calls the phase offset index, uses the voice samples of the stroke screening object, decomposes the spectral features of the time-domain signal through short-time Fourier transform, calculates the energy distribution of phonemes, extracts the fundamental frequency, formant frequency, and bandwidth in multiple frequency ranges, screens the syllable paragraphs with stable pronunciation in the voice samples, records the energy distribution range of multiple syllables, calculates the spectral transition rate between phonemes, analyzes the rate of energy change in multiple frequency bands, and calculates the baseline feature parameters;
[0029] The spectral stability calculation sub-module calls the baseline feature parameters and uses the formula:
[0030]
[0031] Calculate the spectral stability parameter to obtain the spectral stability data;
[0032] where S represents the spectral stability, M represents the number of time windows of the voice sample, f c represents the formant frequency of the c-th time window, f c+1 represents the formant frequency of the (c + 1)-th time window, b c represents the formant bandwidth of the c-th time window, b c+1 represents the formant bandwidth of the (c + 1)-th time window, e c represents the energy distribution of the c-th time window, e c+1 represents the energy distribution of the (c + 1)-th time window, c represents the index of the time window, w1 represents the weight factor of bandwidth change, and w2 represents the weight factor of energy distribution change;
[0033] The individual baseline screening sub-module calls the spectral stability data, calculates the formant time drift range of the voice signal, extracts the formant frequency of each time window, compares the frequency changes of adjacent time windows, calculates the change rate of the formant, sets the drift threshold and screens the voice spectrum, obtains the voice feature data of the target screening object, and establishes an individual baseline model.
[0034] As a further solution of the present invention, the real-time evaluation module includes:
[0035] The time feature comparison sub-module calls the individual baseline model, extracts the time feature parameters of the current voice sample, including syllable duration, pause interval, and speech rate change rate, calculates the time deviation degree of the current voice sample by comparing with the time features of the baseline model, and obtains the time deviation calculation value;
[0036] The frequency feature comparison sub-module combines the calculated time deviation value to extract the frequency features of the current voice sample, including fundamental frequency distribution, formant trajectory, and spectral energy ratio. By comparing with the frequency features of the individual baseline model, it calculates the frequency deviation degree of the current voice sample to obtain the frequency deviation calculation value;
[0037] The comprehensive feature calculation sub-module calls the calculated frequency deviation value and combines it with the time deviation value, using the formula:
[0038]
[0039] to calculate the comprehensive feature deviation rate;
[0040] where CFD is the comprehensive feature deviation rate, TPD is the time deviation calculation value, FPD is the frequency deviation calculation value, α is the weight parameter of the time deviation, β is the weight parameter of the frequency deviation, γ is the weight parameter of the time-frequency deviation interaction term, and δ is the smoothing factor to prevent the denominator from being zero.
[0041] As a further solution of the present invention, the feature matching module includes:
[0042] The formant trajectory calculation sub-module calls the comprehensive feature deviation rate, extracts the formant time series data, analyzes the frequency, bandwidth, and amplitude of the formants in the voice signal, calculates the offset rate of the formants at multiple time points, analyzes the smoothness of the formant trajectory between multiple syllables, and obtains the formant trajectory smoothness parameter by calculating the frequency change rate of the formants at multiple time points;
[0043] The trajectory mapping calculation sub-module calls the formant trajectory smoothness parameter, analyzes the change trend of the formant trajectory in the voice sample, calculates the time series trajectory mapping data through the smoothness of the trajectory at multiple time points, performs interpolation processing on the formant trajectory data at multiple time points to form a continuous trajectory mapping curve, constructs the feature mapping model of the formant trajectory, and obtains the formant time series trajectory mapping data;
[0044] The feature comparison sub-module calls the formant time series trajectory mapping data and combines it with the formant features of the voice samples of known stroke patients, using the formula:
[0045]
[0046] performs operations to obtain the feature matching degree, calls the matching degree data, screens and records stroke patients, and generates stroke screening records;
[0047] where M s is the feature matching degree, X k is the k-th feature value of the current voice sample, and Y kis the corresponding eigenvalue of the baseline sample of the stroke patient, K is the total number of features, and k is the index number of the currently calculated feature.
[0048] An acute stroke rapid voice screening method, which is performed based on the above-mentioned acute stroke rapid voice screening system, and includes the following steps:
[0049] S1: Obtain the voice signal of the screening object, perform time-frequency analysis on the voice signal, calculate the phoneme energy distribution probability of each frequency band, detect the abnormal interval in the voice signal, perform noise reduction on the voice data, and output the noise-reduced voice data;
[0050] S2: Based on the noise-reduced voice data, locate the voice pause timestamp, analyze the pause duration and distribution law, extract the periodic peak-valley characteristics of the voice breathing airflow, analyze the corresponding relationship between the airflow change and the pause point, and form a phase offset index;
[0051] S3: Based on the phase offset index, extract the formant frequency, bandwidth and inter-phoneme energy distribution of the voice data, obtain the stability information of the energy distribution, record the time drift range of the formant, screen the spectrum data, and form an individual baseline model;
[0052] S4: Based on the individual baseline model, compare the time-frequency characteristics of the real-time voice and the baseline model, calculate the time deviation degree of the syllable duration and speech rate fluctuation, calculate the frequency deviation degree of the fundamental frequency distribution and formant trajectory, and generate a comprehensive feature deviation rate;
[0053] S5: Based on the comprehensive feature deviation rate, extract the formant trajectory parameters, calculate the offset rate and smoothness, construct a trajectory mapping model, and compare it with the stroke voice database to generate a stroke screening record.
[0054] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0055] In the present invention, by calculating the phoneme energy distribution probability of the frequency band and dynamically adjusting the noise reduction weight coefficient, the clarity and usability of the voice data are improved, the voice signal processing ability in a noisy environment is optimized, the analysis of the periodic peak-valley characteristics of the breathing airflow and the pause point distribution pattern enhances the recognition accuracy of the voice pause characteristics, by extracting the formant frequency, bandwidth and inter-phoneme energy distribution, the pertinence and adaptability of feature modeling are improved, constructing a trajectory mapping model and comparing it with the stroke voice database improve the efficiency and accuracy of stroke screening, realize rapid screening and accurate recognition, and provide more reliable technical support for the early diagnosis of acute stroke. Description of the Drawings
[0056] Figure 1 is the system flow chart of the present invention;
[0057] Figure 2 It is the flowchart of the signal preprocessing module of the present invention;
[0058] Figure 3 It is the flowchart of the pause analysis module of the present invention;
[0059] Figure 4 It is the flowchart of the feature modeling module of the present invention;
[0060] Figure 5 It is the flowchart of the real-time evaluation module of the present invention;
[0061] Figure 6 It is the flowchart of the feature matching module of the present invention. Specific embodiments
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0064] Please refer to Figure 1 , the present invention provides a technical solution: an acute stroke rapid voice screening system includes:
[0065] The signal preprocessing module acquires the voice signal of the screening object, performs time-frequency transformation, obtains the energy distribution data of phonemes in multiple frequency bands, calculates the occurrence probability of phonemes in each frequency band, screens the abnormal phoneme intervals in the voice signal, divides the key voice frequency band, the secondary voice frequency band and the environmental noise frequency band, sets a noise reduction weight for each frequency band, combines the dynamic changes of the signals in each frequency band, updates the frequency band classification boundary, and obtains the noise-reduced voice data;
[0066] The pause analysis module calls the noise-reduced speech data, detects and marks the pause timestamps of the speech signal, obtains the respiratory airflow cycle data of the stroke screening subject, extracts the respiratory peak and valley values, calculates the offset value of the pause point relative to the respiratory cycle, obtains the respiratory phase angle of the speech pause point, establishes the alignment parameter between the pause point and the respiratory rhythm, identifies the distribution pattern of the pause point, and obtains the phase offset index;
[0067] The feature modeling module calls the phase offset index, uses the speech signal to obtain the speech baseline feature samples of the stroke screening subject, extracts the speech spectrum features through short-time Fourier transform, calculates the formant frequency, bandwidth, and energy distribution between phonemes of the speech signal, obtains the stability of the spectrum features, calls the stability parameter of the spectrum features, records the formant time drift range, calls the drift range, filters the spectrum data, and obtains the individual baseline model;
[0068] The real-time evaluation module calls the individual baseline model to obtain the time features of the model, including syllable duration, pause interval, and speech rate change rate, calculates the time deviation degree of the current speech sample, compares it with the time features of the individual baseline model to obtain the time deviation calculation value, obtains the frequency features of the individual baseline model, including fundamental frequency distribution, formant trajectory, and spectrum energy ratio, obtains the speech detection sample of the stroke screening subject, calculates the frequency deviation degree of the current speech sample, compares it with the frequency features of the individual baseline model to obtain the frequency deviation calculation value, and calls the time deviation calculation value and the frequency deviation calculation value to obtain the comprehensive feature deviation rate;
[0069] The feature matching module calls the comprehensive feature deviation rate, obtains the formant time series data of the stroke screening subject, calculates the trajectory changes of the formant frequency, bandwidth, and amplitude, calls the trajectory change data, calculates the offset rate of the formant at multiple time points, analyzes the smoothness of the formant trajectory between multiple syllables, calls the trajectory smoothness data, calculates the formant time series trajectory mapping data, calls the trajectory mapping data, and through comparative analysis with the speech features of known stroke patients, calculates the feature matching degree, screens and records the stroke patients, and generates a stroke screening record.
[0070] The noise-reduced speech data includes the frequency band energy distribution adjustment value, the speech segment recognition result, and the noise reduction weight coefficient. The phase offset index includes the pause timestamp, the respiratory airflow cycle peak and valley values, and the pause point phase distribution pattern. The individual baseline model includes the energy stability between phonemes, the formant time drift range, and the filtered spectrum data. The comprehensive feature deviation rate includes the speech rate fluctuation deviation degree, the fundamental frequency distribution deviation degree, and the formant trajectory deviation degree. The stroke screening record includes the formant trajectory parameters, the trajectory offset rate, and the stroke speech matching degree.
[0071] Please refer to Figure 2 , the signal preprocessing module includes:
[0072] The frequency band energy distribution calculation sub-module obtains the voice signal of the screening object, performs time-frequency transformation, extracts the energy distribution data of phonemes in multiple frequency bands, calculates the occurrence probability of phonemes in multiple frequency bands, detects the abnormal phoneme distribution area, calls the data of the abnormal phoneme distribution area, screens the frequency band range of the voice signal, and obtains the frequency band energy distribution parameters;
[0073] The frequency band energy distribution calculation sub-module obtains the voice signal of the stroke screening object. First, it uses a high-sensitivity microphone for audio acquisition, records a 10-second voice sample at a sampling rate of 44.1 kHz, converts the sample data into a frequency-domain signal, performs a short-time Fourier transform, decomposes it to obtain the spectral data of different time segments, extracts the energy distribution of phonemes in multiple frequency bands, calculates the proportion of the total energy in each frequency band, and uses the formula:
[0074]
[0075] where E freq represents the total energy of the frequency band, S t (f) is the power spectral density at frequency f at time t (unit: dB), Δt is the window step size, and N = 1000 is taken. Calculate the total energy of different phonemes in the frequency range of 0 - 4 kHz, and obtain the energy ratio after normalization. Assume that the power spectral density of a certain frequency band is S = [15, 18, 20, 16, 14, 17, 19, 20, 18, 16] dB in 10 time segments, and the duration of each time segment is Δt = 10 ms, that is, 0.01 s. Substitute it into the formula to calculate the total energy:
[0076] E freq = (15 + 18 + 20 + 16 + 14 + 17 + 19 + 20 + 18 + 16) × 0.01;
[0077] E freq = 173 × 0.01 = 1.73 dB;
[0078] This result shows that the energy value of this frequency band is 1.73 dB. According to the calculation results of different frequency bands, the energy distribution of the frequency band can be determined, and the frequency band energy distribution parameters can be obtained.
[0079] The abnormal phoneme interval screening sub-module calls the frequency band energy distribution parameters, calculates the energy change amplitude of phonemes in multiple frequency bands, obtains the energy fluctuation curve of phonemes on the time axis, detects the intervals with prominent energy changes in the voice signal, and obtains the abnormal phoneme interval data;
[0080] The abnormal phoneme interval screening sub-module calls the frequency band energy distribution parameters, extracts the main speech frequency interval (300 Hz - 3 kHz) and the background noise frequency interval (below 100 Hz or above 4 kHz) according to the calculated energy ratio of each frequency band, compares the instantaneous energy and the long-term average energy of each frequency band in the speech signal, calculates the energy change rate, and uses the formula:
[0081]
[0082] Among them, R var represents the energy change rate, E inst is the energy value of the current time segment (unit: dB), E avg is the long-term average energy (unit: dB), and ∈ is set to 0.01 to prevent the denominator from approaching zero. Assume that the instantaneous energy E inst of a certain time segment = 12.5 dB, and the long-term average energy E avg = 10.0 dB. Substitute into the formula to calculate the energy change rate:
[0083]
[0084] The result shows that the energy change rate is 0.25. According to the calculated rate value, determine whether there is an abnormal phoneme interval in this frequency band, and obtain the abnormal phoneme interval data.
[0085] The dynamic noise reduction weight adjustment sub-module calls the abnormal phoneme interval data, screens the frequency bands corresponding to the abnormal phonemes, analyzes the energy change rate of the abnormal phoneme interval, calculates the relative proportion of the abnormal phoneme interval in the speech signal, updates the frequency band classification boundary, and uses the formula:
[0086]
[0087] Calculate the adjusted noise reduction weight parameter, adjust the noise reduction weight of the abnormal phoneme interval, set the weight ratio of the key speech frequency band, the secondary speech frequency band and the environmental noise frequency band, calculate the dynamic changes of the signals in multiple frequency bands, perform noise reduction processing on the speech signal, and obtain the noise-reduced speech data;
[0088] Among them, W adj represents the adjusted noise reduction weight, W base represents the basic noise reduction weight, E var represents the energy change value of the abnormal phoneme interval, E norm represents the average energy of the normal phoneme interval, ∈ is a small value to prevent the denominator from being zero, E dev represents the energy deviation of the abnormal phoneme interval, E sum represents the total energy value;
[0089] The dynamic noise reduction weight adjustment sub-module calls the data of abnormal phoneme intervals, filters the frequency bands corresponding to abnormal phonemes, analyzes the energy change rate of the abnormal phoneme intervals, calculates the relative proportion of the abnormal phoneme intervals in the speech signal. If the proportion exceeds 10%, the noise reduction weight is appropriately reduced to avoid losing important speech features. Calculate the dynamic change values of each frequency band, call the dynamic change values, and set the noise reduction weights for the key speech frequency band (300 Hz - 3 kHz), the secondary speech frequency band (100 Hz - 300 Hz, 3 kHz - 4 kHz), and the ambient noise frequency band (below 100 Hz or above 4 kHz). The weight calculation uses the formula:
[0090]
[0091] Among them, W adj is the adjusted noise reduction weight, W base is the basic noise reduction weight (the initial value is set to 0.7), E var represents the energy change value of the abnormal phoneme interval (unit: dB), E norm represents the average energy of the normal phoneme interval (unit: dB), ∈ is set to 0.01 to prevent the denominator from approaching zero, E dev represents the energy deviation of the abnormal phoneme interval, E sum represents the total energy value. Set the energy change value E var of the abnormal phoneme interval = 12.5 dB, the average energy E norm of the normal phoneme interval = 10.0 dB, the energy deviation E dev of the abnormal phoneme interval = 5.2 dB, the total energy value E sum = 50.0 dB, substitute into the formula for calculation:
[0092]
[0093] W adj = 0.7×0.75 + 0.104 = 0.525 + 0.104 = 0.629;
[0094] The result shows that in the current speech signal, the noise reduction weight for the abnormal phoneme interval should be set to 0.629, which is slightly lower than the basic weight of 0.7. Call the adjusted noise reduction weight parameter and apply it to the noise reduction processing of the corresponding frequency band to obtain the noise-reduced speech data.
[0095] Table 1: Example of energy distribution parameters for frequency bands
[0096] Frequency band (Hz) Energy percentage (%) Phoneme anomaly coefficient Noise reduction weight 100-300 12.5 0.15 0.75 300-1000 35.0 0.10 0.70 1000-3000 40.0 0.05 0.65 3000-4000 7.5 0.20 0.80 >4000 5.0 0.25 0.85
[0097] As shown in Table 1, the energy proportion in different frequency bands and the phoneme anomaly coefficient affect the adjustment of the noise reduction weight. For example, the energy proportion in the 100 - 300 Hz frequency band is 12.5%, the phoneme anomaly coefficient is 0.15, and the corresponding noise reduction weight is 0.75. While in the 3000 - 4000 Hz frequency band, the phoneme anomaly coefficient is relatively high, and the noise reduction weight is relatively increased to 0.80 to enhance the noise reduction ability.
[0098] Please refer to Figure 3 , the pause analysis module includes:
[0099] The pause timestamp detection sub-module calls the noise-reduced speech data. By analyzing the time-axis characteristics of the speech signal, it identifies the pauses of the user in the speech signal, calculates the pause duration, combines with the spectral changes of the speech signal, identifies the distribution state of the pause timestamps in the speech, and obtains the speech pause timestamp data;
[0100] The pause timestamp detection sub-module calls the noise-reduced speech data, parses the time axis of the speech signal. First, it detects the instantaneous energy of the speech signal and analyzes the power changes of the speech signal at different moments, extracting the time periods with energy lower than the set threshold. For each identified pause point, it records its occurrence timestamp and calculates the time interval between adjacent pause points to determine whether the pause meets the set physiological criteria. Subsequently, using time series data, it identifies the distribution pattern of the pause points and checks whether the pause points are concentrated in specific speech paragraphs. It calls the extracted pause timestamp data, compares the pause time interval with the set threshold, filters out the pause points that meet the stroke screening characteristics, and obtains the speech pause timestamp data.
[0101] Calculate the pause point interval time using the formula:
[0102] T pause =t pause,a -t pause,a-1 ;
[0103] Among them, T pause represents the pause point interval time (seconds), t pause,a represents the timestamp of the current pause point (seconds), t pause,a-1 represents the timestamp of the previous pause point (seconds). Assuming the timestamps of two consecutive pause points are t pause,1 =1.2 seconds and t pause,2 =2.5 seconds respectively, substituting into the formula for calculation:
[0104] T pause =2.5 - 1.2 = 1.3 seconds;
[0105] This result indicates that the pause point interval time is 1.3 seconds. By calling this interval time and combining with the speech signal characteristics, it further analyzes the physiological rationality of the pause points, and finally filters out the pause points that meet the screening criteria.
[0106] The respiratory airflow cycle calculation sub-module calls the voice pause timestamp data, obtains the respiratory airflow data of the stroke screening object, identifies the respiratory cycle by extracting the peak value of the respiratory signal, matches the respiratory phase curve, and obtains the respiratory cycle data;
[0107] The respiratory airflow cycle calculation sub-module calls the voice pause timestamp data, synchronously obtains the respiratory airflow data of the stroke screening object, analyzes the fluctuation pattern of the respiratory signal, and smooths the airflow waveform to reduce instantaneous noise interference. Then, using time series data, it identifies the local extreme points of the respiratory signal, extracts the peak and valley values of each respiratory cycle, and calculates the time interval between two consecutive respiratory peaks. When calculating the respiratory cycle, it is necessary to check the time stability between adjacent respiratory peaks. If the time interval fluctuates too much, the cycle data is excluded to ensure data accuracy. Based on the time-axis distribution of the respiratory airflow signal, it further calculates the respiratory rhythm change trend and identifies the rhythm characteristics of the respiratory fluctuation. Finally, it calls the calculated respiratory cycle data, matches the respiratory phase curve, and obtains the respiratory airflow cycle data.
[0108] To calculate the respiratory cycle, the formula is used:
[0109] T cycle =t peak,b -t peak,b-1 ;
[0110] Where, T cycle represents the respiratory cycle time (seconds), t peak,b represents the timestamp of the current respiratory peak (seconds), t peak,b-1 represents the timestamp of the previous respiratory peak (seconds). Assuming that the time points of two consecutive respiratory peaks are, t peak,1 =3.0 seconds, t peak,2 =5.5 seconds, substituting into the formula for calculation:
[0111] T cycle =5.5 - 3.0 = 2.5 seconds;
[0112] This result indicates that the respiratory cycle is 2.5 seconds. Call the calculated respiratory cycle data, check the cycle stability, and input the valid cycle data into the subsequent respiratory phase calculation step.
[0113] The phase offset degree calculation sub-module calls the respiratory cycle data, identifies the position of the pause point in the respiratory cycle, calculates the time offset of the pause point relative to the respiratory peak, identifies the distribution pattern of the pause point, and uses the formula:
[0114]
[0115] Calculate the phase offset angle of the pause point, establish the alignment parameter between the pause point and the breathing rhythm, and obtain the phase offset index;
[0116] Among them, θ offset represents the phase offset angle of the pause point, t pause represents the time point when the pause occurs, t peak represents the time point of the nearest breathing peak, T cycle represents the duration of a complete breathing cycle;
[0117] The phase offset calculation sub-module calls the breathing airflow cycle data, obtains the position of the pause point in the breathing cycle, and calculates the time offset of the pause point relative to the breathing peak. First, identify the breathing state when the pause point occurs and determine the complete breathing cycle in which the pause point is located. Then, based on the breathing phase curve, calculate the phase offset angle of the pause point. If the pause point occurs in the late exhalation or early inhalation, the phase offset angle is relatively large; if the pause point is close to the breathing peak, the phase offset angle is relatively small. The formula for calculating the phase offset angle is as follows:
[0118]
[0119] Among them, θ offset represents the phase offset angle (degrees) of the pause point, t pause represents the time point (seconds) when the pause occurs, t peak represents the time point (seconds) of the nearest breathing peak, T cycle represents the duration (seconds) of a complete breathing cycle. Assume that the pause time point t pause = 4.0 seconds, the time point of the nearest breathing peak t peak = 3.0 seconds, and the breathing cycle duration T cycle = 2.5 seconds. Substitute into the formula for calculation:
[0120]
[0121] The result shows that the phase offset angle of the pause point is 144 degrees. Call the phase offset angle of the pause point, establish the alignment parameter between the pause point and the breathing rhythm, and obtain the phase offset index.
[0122] Table 2: Example of Breathing Cycle and Pause Point Characteristics
[0123]
[0124]
[0125] As shown in Table 2, different respiratory cycles and pause point characteristics affect the calculation of phase shift. For example, when the pause point occurs at 4.0 seconds and the nearest respiratory peak is at 3.0 seconds, the calculated phase shift is 144 degrees. When the pause point is at 9.5 seconds and the nearest respiratory peak time is 8.0 seconds, the phase shift reaches 270 degrees.
[0126] Please refer to Figure 4 , the feature modeling module includes:
[0127] The baseline feature extraction sub-module calls the phase shift index, uses the speech samples of stroke screening subjects, decomposes the spectral features of the time-domain signal through short-time Fourier transform, calculates the energy distribution of phonemes, extracts the fundamental frequency, formant frequency, and bandwidth within multiple frequency ranges, screens the syllable paragraphs with stable pronunciation in the speech samples, records the energy distribution range of multiple syllables, calculates the spectral transition rate between phonemes, analyzes the rate of energy change in multiple frequency bands, and calculates the baseline feature parameters;
[0128] The baseline feature extraction sub-module calls the phase shift index, obtains the speech baseline feature samples of stroke screening subjects, performs short-time Fourier transform on the speech signal, converts the original speech signal from the time domain to the frequency domain, analyzes the fundamental frequency, formant frequency, and amplitude characteristics of the speech signal, extracts the time-frequency distribution characteristics of syllables, calculates the average duration of syllables, and measures the time interval between adjacent syllables to identify abnormal pause intervals. Using the sliding window method, calculate the energy mean of the speech signal within different time segments, obtain the dispersion degree of the energy distribution, and judge the energy fluctuation trend by calculating the energy distribution variance, and eliminate the abnormal intervals where the energy fluctuation exceeds the set range to ensure the extraction of stable speech samples.
[0129] The standard deviation of the syllable time interval is calculated using the formula:
[0130]
[0131] where σ YT is the standard deviation of the syllable time interval (seconds), YM is the number of syllables, YT d is the duration of the d-th syllable (seconds), is the average duration of syllables (seconds).
[0132] Suppose there are five syllables in a certain section of speech, and their durations are YT1 = 0.4s, YT2 = 0.5s, YT3 = 0.3s, YT4 = 0.6s, YT5 = 0.5s respectively. Then calculate the average duration of syllables:
[0133]
[0134] Calculate the standard deviation of the syllable time interval:
[0135]
[0136] This result indicates that the syllable interval stability of this segment of speech sample is 0.104 s. If it is higher than the set threshold, there may be abnormal pauses.
[0137] The spectral stability calculation sub-module calls the baseline feature parameters and uses the formula:
[0138]
[0139] Calculate the spectral stability parameter and obtain the spectral stability data;
[0140] Among them, S represents the spectral stability, M represents the number of time windows of the speech sample, f c represents the formant frequency of the c-th time window, f c+1 represents the formant frequency of the (c + 1)-th time window, b c represents the formant bandwidth of the c-th time window, b c+1 represents the formant bandwidth of the (c + 1)-th time window, e c represents the energy distribution of the c-th time window, e c+1 represents the energy distribution of the (c + 1)-th time window, c represents the index of the time window, w1 represents the weight factor of the bandwidth change, and w2 represents the weight factor of the energy distribution change;
[0141] The spectral stability calculation sub-module calls the baseline feature parameters, calculates the spectral stability of the speech signal in different time windows, extracts the formant information in different time segments of the speech sample through the sliding window technology, obtains the formant frequency distribution in multiple time windows, calculates the formant change rate between time windows, measures the amplitude of the frequency offset, compares with the set stability threshold, screens the frequency intervals lower than the threshold to remove the frequency bands greatly affected by environmental noise or speech dynamic changes, and only retains the speech samples that meet the stable standard.
[0142] Calculate the spectral stability using the formula:
[0143]
[0144] Among them, S is the spectral stability (Hz), M is the number of time windows, f b is the formant frequency (Hz) of the b-th time window, b b is the formant bandwidth (Hz) of the b-th time window, e b is the energy distribution (dB) of the b-th time window, and w2 and w3 are the weight factors of the bandwidth and energy changes respectively.
[0145] Assume that there are three time windows in the speech sample, and their formant frequencies are f1 = 2500 Hz, f2 = 2520 Hz, and f3 = 2495 Hz, respectively. The corresponding bandwidths are b1 = 130 Hz, b2 = 128 Hz, and b3 = 132 Hz, respectively. The energy distributions are e1 = 60 dB, e2 = 62 dB, and e3 = 61 dB, respectively. Assume that the weight factors w2 = 0.3 and w3 = 0.4, and calculate the spectrum stability:
[0146]
[0147] The result shows that the spectrum stability of the sample is 24 Hz. If it exceeds the set threshold, there may be abnormal spectrum drift.
[0148] The individual baseline screening submodule calls the spectrum stability data, calculates the time drift range of the formant of the speech signal, extracts the formant frequency of each time window, compares the frequency changes of adjacent time windows, calculates the change rate of the formant, sets the drift threshold and screens the speech spectrum, obtains the speech feature data of the target screening object, and establishes an individual baseline model;
[0149] The individual baseline screening submodule calls the spectrum stability data, analyzes the time drift range of the formant of the speech signal, records the distribution of the formants in the stability frequency band, compares the change trend of the formant frequency on the time axis, calculates the rate at which the formant frequency drifts along the time axis, screens the drift interval that exceeds the normal range, and determines whether the drift trend meets the standard of the stable speech mode. For the frequency bands that drift beyond the set range, analyze the duration and frequency of abnormal drift, remove the parts with more drastic changes, and retain only the spectrum data that meets the stability threshold.
[0150] The drift interval duration is calculated using the formula:
[0151]
[0152] Where PD is the total duration of the drift interval, PM is the number of drift events, and PT is g is the starting time of the g-th drift, I g It is an indicator variable of the drift persistence state, 0 or 1, indicating whether it is in the drift state.
[0153] Assume that there are three formant drifts in a speech sample, and their starting times are PT1 = 0.5s, PT2 = 1.2s, and PT3 = 2.0s. Assume that all drift events are abnormal, that is, I1 = I2 = I3 = 1, then calculate the duration of the drift interval:
[0154] PD=(1.2-0.5)·1+(2.0-1.2)·1=0.7+0.8=1.5;
[0155] The result shows that the total duration of the abnormal drift interval of the sample is 1.5 s. If it exceeds the set threshold, there may be a significant abnormal resonance peak drift.
[0156] Please refer to Figure 5 , the real-time evaluation module includes:
[0157] The time feature comparison sub-module calls the individual baseline model to extract the time feature parameters of the current speech sample, including syllable duration, pause interval, and speech rate change rate. By comparing with the time features of the baseline model, the time deviation degree of the current speech sample is calculated to obtain the time deviation calculation value;
[0158] The time feature comparison sub-module calls the individual baseline model to extract the time feature parameters of the current speech sample, including syllable duration, pause interval, and speech rate change rate. It segments each syllable paragraph from the speech signal, calculates the start and end times of each syllable, analyzes the position where the pause appears and its interval length, and measures the change trend of the speech rate within different time windows to obtain the statistical distribution of each time feature. By calculating the difference between each feature value and the corresponding feature value of the individual baseline model, the deviation degree of the time feature is determined.
[0159] The time deviation is calculated using the following formula:
[0160]
[0161] where D t is the time deviation calculation value, T z is the z-th time feature value of the current speech sample, T bz is the corresponding time feature value of the individual baseline model, and TM is the total number of time features.
[0162] Suppose the time feature data of the current speech sample is as follows: T = {0.35, 0.40, 0.30, 0.38, 0.42},
[0163] The time features of the individual baseline model are: T b = {0.32, 0.39, 0.28, 0.36, 0.40},
[0164] Then the time deviation is calculated as follows:
[0165]
[0166] This result indicates that the average deviation of the time features of the current speech sample compared with the individual baseline model is 0.02 seconds.
[0167] The frequency feature comparison sub-module combines the time deviation calculation value to extract the frequency features of the current speech sample, including fundamental frequency distribution, formant trajectory, and spectral energy ratio. By comparing with the frequency features of the individual baseline model, it calculates the frequency deviation degree of the current speech sample and obtains the frequency deviation calculation value.
[0168] The frequency feature comparison sub-module extracts the frequency features of the current speech sample according to the time deviation calculation value, including fundamental frequency distribution, formant trajectory, and spectral energy ratio. It calls the short-time Fourier transform to obtain the energy distribution of the speech signal in different frequency intervals, identifies the main fundamental frequency components, analyzes the frequency trajectory of the formants and their drift on the time axis, and calculates the energy proportion of the speech signal in different frequency bands to form an overall distribution map of the spectral features. For these features, it calculates the mean square error with the feature data stored in the individual baseline model to determine the frequency offset degree of the current speech sample.
[0169] The frequency deviation is calculated using the following formula:
[0170]
[0171] where D f is the frequency deviation calculation value, F h is the h-th frequency feature value of the current speech sample, and F bh is the corresponding frequency feature value of the individual baseline model, and PN is the total number of frequency features.
[0172] Suppose the frequency feature data of the current speech sample is as follows: F = {210, 215, 220, 218, 212},
[0173] The frequency features of the individual baseline model are: F b = {208, 213, 218, 215, 210},
[0174] Then the frequency deviation is calculated as follows:
[0175]
[0176] This result indicates that the average deviation of the frequency features of the current speech sample compared to the individual baseline model is 2.24 Hz.
[0177] The comprehensive feature calculation sub-module calls the frequency deviation calculation value and combines it with the time deviation value using the formula:
[0178]
[0179] to calculate the comprehensive feature deviation rate;
[0180] Among them, CFD is the comprehensive feature deviation rate, TPD is the calculated value of time deviation, FPD is the calculated value of frequency deviation, α is the weight parameter of time deviation, β is the weight parameter of frequency deviation, γ is the weight parameter of the time-frequency deviation interaction term, and δ is the smoothing factor to prevent the denominator from being zero;
[0181] Call the calculated value of time deviation and the calculated value of frequency deviation. First, extract the time feature and frequency feature data of the target individual. The time feature data includes syllable duration, pause interval, and speech rate change rate. The frequency feature data includes fundamental frequency distribution, formant trajectory, and spectral energy ratio. By performing time series analysis on these data, identify their change trends over time, and use the sliding window method for smoothing processing to eliminate the influence of short-term fluctuations. On this basis, calculate the time deviation degree of the current speech sample, compare the time feature data with the time feature data in the individual baseline model to obtain the calculated value of time deviation, and the calculated value of frequency deviation is obtained by calculating the deviation degree of frequency parameters such as the fundamental frequency and formant trajectory of the current speech sample. Based on the above calculation results, combine the calculated value of time deviation and the calculated value of frequency deviation, and use the weight parameters to normalize the two to make them comparable on the same scale, and finally calculate the comprehensive feature deviation rate.
[0182] Use the formula:
[0183]
[0184] Assume TPD = 0.15, FPD = 0.10, α = 0.6, β = 0.4, γ = 0.2, δ = 0.05, substitute the set values for calculation:
[0185] 0.6×0.15 + 0.4×0.10 = 0.09 + 0.04 = 0.13;
[0186] 0.2×(0.15×0.10) = 0.2×0.015 = 0.003;
[0187] 0.05 + |0.15 - 0.10| = 0.05 + 0.05 = 0.10;
[0188]
[0189] The calculation results show that compared with the baseline model, the comprehensive feature deviation rate of the current speech sample of the target individual is 1.33, indicating that there is a certain degree of change between the time deviation and the frequency deviation. This value can be used to evaluate its abnormal degree in the follow-up.
[0190] Please refer to Figure 6 , the feature matching module includes:
[0191] The formant trajectory calculation sub-module calls the comprehensive feature deviation rate, extracts the formant time series data, analyzes the frequency, bandwidth, and amplitude of the formants in the speech signal, calculates the deviation rate of the formants at multiple time points, analyzes the smoothness of the formant trajectory between multiple syllables, and obtains the formant trajectory smoothness parameter by calculating the rate of change of the formant frequency at multiple time points;
[0192] The formant trajectory calculation sub-module calls the comprehensive feature deviation rate, extracts the formant time series data, analyzes the speech signal, obtains the frequency, bandwidth, and amplitude of the formants, measures the changes of these parameters at multiple time points, and forms a set of time series feature data. In the analysis process, first, the collected speech data needs to be framed, and each frame contains a fixed number of sampling points to ensure the continuity and stability of the time series data. Next, the frequency characteristics of the formants need to be calculated for each frame of data. A common method is to extract the formant information based on the short-time Fourier transform (STFT) or the Mel spectrum analysis method.
[0193] When calculating the formant deviation rate, select the formant frequencies f′ i , f′ i+1 at adjacent time points t′ i , f′ i+1 for calculation:
[0194]
[0195] where V′ f is the formant frequency change rate, f′ i+1 , f′ i are the formant frequencies at time points t′ i+1 , t′ i respectively, and t′ i+1 - t′ i is the time interval.
[0196] To illustrate this calculation process specifically, assume that the formant frequencies of a certain speech signal at t′1 = 0.5 s and t′2 = 1.0 s are f′1 = 250 Hz and f′2 = 280 Hz respectively, then calculate:
[0197]
[0198] This rate reflects the change of frequency over time. If the change rate exceeds a certain threshold, it indicates that the frequency change is large and there may be an anomaly. Then, calculate the mean value of the change rates at all time points to obtain the smoothness of the formant trajectory, and calculate the formant trajectory smoothness parameter for subsequent speech feature analysis.
[0199] The trajectory mapping calculation sub-module calls the formant trajectory smoothness parameter, analyzes the change trend of the formant trajectory in the speech sample, calculates the time series trajectory mapping data through the trajectory smoothness at multiple time points, interpolates the formant trajectory data at multiple time points to form a continuous trajectory mapping curve, constructs a feature mapping model of the formant trajectory, and obtains the formant time series trajectory mapping data;
[0200] The trajectory mapping calculation sub-module calls the formant trajectory smoothness parameter, analyzes the change trend of the formant trajectory in the speech sample. First, it statistically analyzes the trajectory changes at each time point to construct the time series trajectory mapping data. Generally, in order to reduce the discreteness and sudden change of the signal, it is necessary to interpolate the trajectory data. The interpolation methods include linear interpolation, spline interpolation, and mean interpolation, etc.
[0201] To ensure the smoothness of the trajectory, adjust the time points with large trajectory changes. Assume that at time point t i-1 ,t i ,t i+1 the formant trajectory data changes violently, and the mean interpolation method can be used to adjust the formant frequency at t i :
[0202]
[0203] where, pf i is the smoothed frequency value at time point t i , pf i-1 , pf i+1 are the frequency values at adjacent time points respectively.
[0204] Assume that the frequencies at t1, t2, and t3 are pf1 = 240Hz and pf3 = 260Hz respectively, then the interpolation calculation at t2 is:
[0205]
[0206] Through this method, a continuous trajectory mapping curve is formed, and a feature mapping model of the formant trajectory is established. This mapping model is not only used to detect the stability of individual speech signals, but also can be used for the comparative analysis of specific speech features to obtain the formant time series trajectory mapping data.
[0207] In practical applications, the change of the formant trajectory is affected by multiple factors, such as background noise, the volume change of the speaker, and the sensitivity of the voice recording device, etc. Therefore, when establishing the trajectory mapping data, it is necessary to normalize these factors. For example, through normalization, all frequency values are mapped between [0,1] to eliminate the benchmark differences between different speech samples, making the trajectory data more suitable for feature matching analysis.
[0208] The feature comparison sub-module calls the formant time series trajectory mapping data, combines the formant features of the voice samples of known stroke patients, and uses the formula:
[0209]
[0210] Performs calculations to obtain the feature matching degree, calls the matching degree data, screens and records stroke patients, and generates stroke screening records;
[0211] where, M s is the feature matching degree, X k is the k-th feature value of the current voice sample, Y k is the corresponding feature value of the baseline sample of the stroke patient, K is the total number of features, and k is the index number of the currently calculated feature;
[0212] The feature comparison sub-module calls the formant time series trajectory mapping data, performs feature matching analysis on the formant trajectory of the current voice sample and the trajectory data of the voice samples of known stroke patients, and calculates the feature matching degree. When calculating the feature matching degree, the formula is used:
[0213]
[0214] where, M s is the feature matching degree, X k is the k-th feature value of the current voice sample, Y k is the corresponding feature value of the baseline sample of the stroke patient, K is the total number of features, and k is the index number of the currently calculated feature.
[0215] In practical applications, different voice features (such as formant frequency, syllable duration, energy distribution, etc.) are selected for matching, and different weights are assigned to different features to highlight the influence of key features on the overall similarity calculation. Assume X = [0.8, 0.6, 0.7], Y = [0.75, 0.65, 0.68], then:
[0216] ∑X k Y k = (0.8 × 0.75) + (0.6 × 0.65) + (0.7 × 0.68) = 0.998;
[0217]
[0218] This result indicates that the similarity between the target voice sample and the voice samples of known stroke patients is relatively high. According to the set matching threshold, further screen and record whether the target sample meets the characteristics of stroke patients to obtain the feature matching degree.
[0219] A rapid voice screening method for acute stroke, the rapid voice screening method for acute stroke is executed based on the above-mentioned rapid voice screening system for acute stroke, and includes the following steps:
[0220] S1: Obtain the voice signal of the screening object, perform time-frequency analysis on the voice signal, calculate the phoneme energy distribution probability of each frequency band, detect the abnormal interval in the voice signal, perform noise reduction on the voice data, and output the noise-reduced voice data;
[0221] S2: Based on the noise-reduced voice data, locate the voice pause timestamps, analyze the pause duration and distribution law, extract the periodic peak-valley characteristics of the voice breathing airflow, analyze the corresponding relationship between the airflow change and the pause point, and form a phase deviation index;
[0222] S3: Based on the phase deviation index, extract the formant frequency, bandwidth and energy distribution between phonemes of the voice data, obtain the stability information of the energy distribution, record the time drift range of the formant, screen the spectral data, and form an individual baseline model;
[0223] S4: Based on the individual baseline model, compare the time-frequency characteristics of the real-time voice and the baseline model, calculate the time deviation degree of the syllable duration and speech rate fluctuation, calculate the frequency deviation degree of the fundamental frequency distribution and formant trajectory, and generate a comprehensive feature deviation rate;
[0224] S5: Based on the comprehensive feature deviation rate, extract the formant trajectory parameters, calculate the deviation rate and smoothness, construct a trajectory mapping model, and compare it with the stroke voice database to generate a stroke screening record.
[0225] The above is only the preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. An acute stroke rapid voice screening system, characterized in that, The system includes: The signal preprocessing module acquires the voice signal of the screening object, calculates the phoneme energy distribution probability of each frequency band, identifies the abnormal interval, divides the key voice segments, secondary voice segments and noise segments, adjusts the noise reduction weight coefficient, and outputs the noise-reduced voice data; The pause analysis module, based on the noise-reduced voice data, locates the voice pause timestamps, extracts the peak-valley characteristics of the respiratory airflow cycle, identifies the distribution pattern of the pause points, and generates the phase offset index; The feature modeling module calls the phase offset index, extracts the formant frequency, bandwidth, and inter-phoneme energy distribution of the voice data, obtains the energy distribution stability, records the formant time drift range, and filters the spectral data to form an individual baseline model; The real-time evaluation module uses the individual baseline model to compare the time-frequency characteristics of the real-time voice and the baseline model, calculates the time deviation degree of the syllable duration and speech rate fluctuation, calculates the frequency deviation degree of the fundamental frequency distribution and formant trajectory, and generates the comprehensive feature deviation rate; The feature matching module calls the comprehensive feature deviation rate, extracts the formant trajectory parameters, calculates the offset rate and smoothness, constructs a trajectory mapping model and compares it with the stroke voice database to generate a stroke screening record.
2. The acute stroke rapid voice screening system according to claim 1, characterized in that, The noise-reduced voice data includes the frequency band energy distribution adjustment value, the voice segment recognition result, and the noise reduction weight coefficient. The phase offset index includes the pause timestamp, the peak-valley value of the respiratory airflow cycle, and the pause point phase distribution pattern. The individual baseline model includes the inter-phoneme energy stability, the formant time drift range, and the filtered spectral data. The comprehensive feature deviation rate includes the speech rate fluctuation deviation degree, the fundamental frequency distribution deviation degree, and the formant trajectory deviation degree. The stroke screening record includes the formant trajectory parameters, the trajectory offset rate, and the stroke voice matching degree.
3. The acute stroke rapid voice screening system according to claim 2, wherein The signal preprocessing module includes: The frequency band energy distribution calculation sub-module acquires the voice signal of the screening object, performs time-frequency transformation, extracts the energy distribution data of phonemes in multiple frequency bands, calculates the occurrence probability of phonemes in multiple frequency bands, detects the abnormal area of phoneme distribution, calls the abnormal area data of phoneme distribution, filters the frequency band range of the voice signal, and obtains the frequency band energy distribution parameters; The abnormal phoneme interval screening sub-module calls the frequency band energy distribution parameters, calculates the energy change amplitude of phonemes in multiple frequency bands, obtains the energy fluctuation curve of phonemes on the time axis, detects the interval with prominent energy change in the voice signal, and obtains the abnormal phoneme interval data; The dynamic noise reduction weight adjustment sub-module calls the abnormal phoneme interval data, filters the frequency bands corresponding to the abnormal phonemes, analyzes the energy change rate of the abnormal phoneme interval, calculates the relative proportion of the abnormal phoneme interval in the voice signal, updates the frequency band classification boundary, and uses the formula: Calculate the adjusted noise reduction weight parameter, adjust the noise reduction weight of the abnormal phoneme interval, set the weight ratio of the key voice frequency band, secondary voice frequency band, and environmental noise frequency band, calculate the dynamic change of the signal in multiple frequency bands, perform noise reduction processing on the voice signal, and obtain the noise-reduced voice data; Among them, W adj represents the adjusted noise reduction weight, and W base represents the basic noise reduction weight. E var represents the energy change value in the abnormal phoneme interval, and E norm represents the average energy in the normal phoneme interval. ∈ is a small value to prevent the denominator from being zero, and E dev represents the energy deviation in the abnormal phoneme interval, and E sum represents the total energy value.
4. The acute stroke rapid voice screening system according to claim 3, wherein The pause analysis module includes: The pause timestamp detection sub-module calls the denoised speech data, identifies the pauses of the user in the speech signal by analyzing the time-axis characteristics of the speech signal, calculates the pause duration, combines the spectral changes of the speech signal, identifies the distribution state of the pause timestamps in the speech, and obtains the speech pause timestamp data; The respiratory airflow cycle calculation sub-module calls the speech pause timestamp data, obtains the respiratory airflow data of the stroke screening object, identifies the respiratory cycle by extracting the peaks of the respiratory signal, matches the respiratory phase curve, and obtains the respiratory cycle data; The phase offset degree calculation sub-module calls the respiratory cycle data, identifies the position of the pause point in the respiratory cycle, calculates the time offset of the pause point relative to the respiratory peak, identifies the distribution pattern of the pause point, and uses the formula: Calculate the phase offset angle of the pause point, establish the alignment degree parameter between the pause point and the respiratory rhythm, and obtain the phase offset degree index; Among them, θ offset represents the phase offset angle of the pause point, t pause represents the time point when the pause occurs, t peak represents the time point of the most recent respiratory peak, T cycle represents the duration of a complete respiratory cycle.
5. The acute stroke rapid voice screening system according to claim 4, wherein The feature modeling module includes: The baseline feature extraction sub-module calls the phase offset degree index, uses the speech samples of the stroke screening object, decomposes the spectral features of the time-domain signal through short-time Fourier transform, calculates the energy distribution of phonemes, extracts the fundamental frequency, formant frequency, and bandwidth in multiple frequency ranges, screens the syllable paragraphs with stable pronunciation in the speech samples, records the energy distribution range of multiple syllables, calculates the spectral transition rate between phonemes, analyzes the rate of energy change in multiple frequency bands, and calculates the baseline feature parameters; The spectral stability calculation sub-module calls the baseline feature parameters and uses the formula: Calculate the spectral stability parameter and obtain the spectral stability data; Among them, S represents the spectral stability, M represents the number of time windows of the speech sample, f c represents the formant frequency of the c-th time window, f c+1 represents the formant frequency of the (c + 1)-th time window, b c represents the formant bandwidth of the c-th time window, b c+1 represents the formant bandwidth of the (c + 1)-th time window, e c represents the energy distribution of the c-th time window, e c+1 represents the energy distribution of the (c + 1)-th time window, c represents the index of the time window, w1 represents the weight factor of the bandwidth change, and w2 represents the weight factor of the energy distribution change; The individual baseline screening sub-module calls the spectral stability data, calculates the formant time drift range of the speech signal, extracts the formant frequency of each time window, compares the frequency changes of adjacent time windows, calculates the change rate of the formant, sets the drift threshold and screens the speech spectrum, obtains the speech feature data of the target screening object, and establishes an individual baseline model.
6. The acute stroke rapid voice screening system according to claim 5, characterized in that, The real-time evaluation module includes: The time feature comparison sub-module calls the individual baseline model, extracts the time feature parameters of the current speech sample, including syllable duration, pause interval, and speech rate change rate, calculates the time deviation degree of the current speech sample by comparing with the time features of the baseline model, and obtains the time deviation calculation value; The frequency feature comparison sub-module combines the time deviation calculation value, extracts the frequency features of the current speech sample, including fundamental frequency distribution, formant trajectory, and spectral energy ratio, calculates the frequency deviation degree of the current speech sample by comparing with the frequency features of the individual baseline model, and obtains the frequency deviation calculation value; The comprehensive feature calculation sub-module calls the frequency deviation calculation value, combines the time deviation value, and uses the formula: Calculate the comprehensive feature deviation rate; Among them, CFD is the comprehensive feature deviation rate, TPD is the time deviation calculation value, FPD is the frequency deviation calculation value, α is the weight parameter of the time deviation, β is the weight parameter of the frequency deviation, γ is the weight parameter of the time-frequency deviation interaction term, and δ is the smoothing factor to prevent the denominator from being zero.
7. The acute stroke rapid voice screening system according to claim 6, characterized in that The feature matching module includes: The formant trajectory calculation sub-module calls the comprehensive feature deviation rate, extracts the formant time series data, analyzes the frequency, bandwidth, and amplitude of the formants in the speech signal, calculates the deviation rate of the formants at multiple time points, analyzes the smoothness of the formant trajectory among multiple syllables, and obtains the formant trajectory smoothness parameter by calculating the formant frequency change rate at multiple time points; The trajectory mapping calculation sub-module calls the formant trajectory smoothness parameter, analyzes the change trend of the formant trajectory in the speech sample, calculates the time series trajectory mapping data through the trajectory smoothness at multiple time points, performs interpolation processing on the formant trajectory data at multiple time points, forms a continuous trajectory mapping curve, constructs a feature mapping model of the formant trajectory, and obtains the formant time series trajectory mapping data; The feature comparison sub-module calls the formant time series trajectory mapping data, combines the formant characteristics of the speech samples of known stroke patients, and uses the formula: Performs operations to obtain the feature matching degree, calls the matching degree data, screens and records stroke patients, and generates a stroke screening record; Among them, M s is the feature matching degree, X k is the k-th eigenvalue of the current voice sample, Y k is the corresponding eigenvalue of the baseline sample of stroke patients, K is the total number of features, and k is the index number of the currently calculated feature.
8. A rapid voice screening method for acute stroke, characterized in that, Executed according to the acute stroke rapid speech screening system described in any one of claims 1-7, including the following steps: S1: Obtain the speech signal of the screening object, perform time-frequency analysis on the speech signal, calculate the phoneme energy distribution probability of each frequency band, detect the abnormal interval in the speech signal, perform noise reduction on the speech data, and output the noise-reduced speech data; S2: Based on the noise-reduced speech data, locate the speech pause time stamps, analyze the pause duration and distribution law, extract the periodic peak-valley characteristics of the speech breathing airflow, analyze the corresponding relationship between the airflow change and the pause point, and form a phase deviation index; S3: Based on the phase deviation index, extract the formant frequency, bandwidth, and energy distribution between phonemes of the speech data, obtain the stability information of the energy distribution, record the time drift range of the formants, screen the spectral data, and form an individual baseline model; S4: Based on the individual baseline model, compare the time-frequency characteristics of the real-time speech and the baseline model, calculate the time deviation degree of the syllable duration and speech rate fluctuation, calculate the frequency deviation degree of the fundamental frequency distribution and the formant trajectory, and generate a comprehensive feature deviation rate; S5: Based on the comprehensive feature deviation rate, extract the formant trajectory parameters, calculate the deviation rate and smoothness, construct a trajectory mapping model, and compare it with the stroke speech database to generate a stroke screening record.