Voice-based learner cognitive state assessment method, system, device and medium
By extracting features of microvibration, vocal rhythm control, and respiratory-voice coordination efficiency from speech data, the accuracy and interpretability of learner cognitive state assessment in distance learning are addressed, resulting in more stable and reliable assessment outcomes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU OPEN UNIVERSITY (THE CITY VOCATIONAL COLLEGE OF JIANGSU)
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies cannot accurately assess learners’ cognitive state in remote learning. Traditional methods are highly subjective and cannot deeply explore biomarkers closely related to cognitive state, resulting in low accuracy and poor interpretability of assessment results.
By collecting speech data, this study extracts features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency using LPC, ICA, Hilbert transform, and nonlinear dynamics methods. These features are then mapped to attention, fatigue, psychological load, and stress level. A weighted fusion approach is used for comprehensive decision-making, providing a speech-based method for assessing learners' cognitive state.
It improves the accuracy and interpretability of learners' cognitive status assessment in distance learning, has an objective biological basis, enhances the robustness of the assessment, and facilitates precise teaching intervention.
Smart Images

Figure CN122020096B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of remote learning cognitive assessment, specifically relating to a speech-based method, system, device, and medium for assessing learners' cognitive state. Background Technology
[0002] With the increasing prevalence of distance learning and online education, timely and accurate assessment of students' cognitive status has become a key challenge. Traditional assessment methods have several limitations, relying primarily on teacher observation, self-report questionnaires, or periodic exams to assess student learning. These methods are highly subjective and unsuitable for the educational context of distance learning.
[0003] Existing voice or video-based analytics technologies typically focus on a single dimension, such as voice tone or facial expression, failing to delve into the subtle biomarkers generated by the neuromuscular and central nervous system control systems that are closely related to cognitive states. This results in low accuracy of assessment results, making it impossible to achieve precise early warning and intervention. Furthermore, the use of a purely data black box approach for assessment leads to poor interpretability of the results. Summary of the Invention
[0004] This invention addresses the shortcomings of existing technologies by providing a speech-based method, system, device, and medium for assessing learners' cognitive states. It can deeply mine biomarkers that influence cognition from collected speech data, thereby improving the accuracy and interpretability of learners' cognitive state assessment in distance learning.
[0005] This invention provides the following technical solution:
[0006] Firstly, a speech-based method for assessing learners' cognitive states is provided, comprising the following steps: Collect and preprocess speech data under different vocalization tasks. The speech data includes reference speech, reading speech, and answer speech. The micro-vibration features reflecting speech tremor were separated from the reference speech using LPC and ICA analysis methods. The vocal rhythm control features were reconstructed from the reading speech using Hilbert transform and nonlinear dynamics methods. The respiratory speech coordination efficiency features of the answer speech were extracted using energy envelope analysis and phase coordination calculation. The features of speech microvibration, vocal rhythm control, and respiratory-voice coordination efficiency are mapped to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. Weighted fusion is then used to make a comprehensive decision, resulting in the learner's final scores for the four cognitive dimensions.
[0007] Optionally, the method further includes the following steps: performing real-time validity analysis on the different collected voice data respectively; if the preset validity conditions are met, the current voice data is retained; if the preset validity conditions are not met, the current voice data is deleted and the voice data is re-collected.
[0008] Optionally, the step of separating speech micro-vibration features reflecting speech tremor from the reference speech using LPC and ICA analysis methods specifically involves: LPC analysis is performed on each frame of reference speech to obtain a single-channel residual signal; A multi-channel observation signal is constructed by embedding the single-channel residual signal through a time delay; ICA analysis was performed on the multi-channel observation signals to obtain several independent components; Based on the physiological characteristics of speech microvibrations, the component with the most concentrated energy was selected from all independent components as the microvibration signal. Extract speech microvibration features from microvibration signals, including vibration energy, frequency band energy ratio, and sample entropy.
[0009] Optionally, the reconstructing of vocal rhythm control features from spoken speech using Hilbert transform and nonlinear dynamics methods specifically involves: The speech envelope signal is extracted from the read-out speech using Hilbert transform; The scaling exponent of the speech envelope signal is extracted using a detrended fluctuation analysis method. ; The MSE curve of the speech envelope signal was plotted using multi-scale entropy analysis, and the area under the curve was extracted. and complexity index ; The recursive graph of the speech envelope signal is calculated using the phase space reconstruction method, and deterministic indices of the recursive graph are extracted using recursive quantitative analysis. and laminar flow index ; The width of the multifractal spectrum is extracted after calculating the multifractal spectrum of the speech envelope signal using multifractal analysis. ; Scaling index Width of multifractal spectrum The area under the MSE curve and complexity index and deterministic indicators of recursion graphs and laminar flow index This is integrated into vocal rhythm control characteristics.
[0010] Optionally, the respiratory-voice coordination efficiency features of the answer speech are extracted through energy envelope analysis and phase coordination calculation, specifically: Respiratory cycles were detected from the voice responses using energy envelope analysis, and the inspiratory duration, expiratory duration, respiratory cycle length, inspiratory-expiratory ratio, and peak expiratory energy of each respiratory cycle were recorded. Within each expiratory segment, the number of syllables was detected, and the coefficients of variation for syllable rate, speech production ratio, and inspiratory-expiratory ratio were calculated. Map the timeline of each respiratory cycle to On the phase ring, the phase consistency index of all breathing cycles is obtained by using the phase value of the start time of each syllable on the phase ring; The energy envelope of the speech during all exhalation phases is compared with the preset ideal envelope, and the ratio between the two is calculated to obtain the energy output ratio. The ideal envelope is the rectangle formed by maintaining the peak energy of exhalation during all exhalation phases. The overall stress index of the current learner is calculated based on the phase consistency index, energy output ratio, and the current learner's baseline respiratory rate. The inspiratory-to-expiratory ratio, syllable rate, speech production ratio, energy output ratio, coefficient of variation of the inspiratory-to-expiratory ratio, phase consistency index, and comprehensive stress index are integrated into the respiratory-speech coordination efficiency characteristics.
[0011] Optionally, the process of mapping the speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features to initial scores for four cognitive dimensions—attention, fatigue, psychological load, and stress level—is performed, and a weighted fusion is used for comprehensive decision-making to obtain the learner's final scores for the four cognitive dimensions, specifically: The extracted speech microvibration features, vocal rhythm control features, and respiratory speech coordination efficiency features were preprocessed respectively. Three independent random forest models were used to predict the initial scores of each feature in the four cognitive dimensions of attention, fatigue, psychological load, and stress level. The final score for each cognitive dimension is calculated and output according to the following formula: ; in, For learners The final score for each cognitive dimension, , and These are the characteristics of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency, respectively, at the [number]th [year]. The weighting coefficients of each cognitive dimension, , and The features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model are respectively the features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model in the 1st... Initial scores for each cognitive dimension.
[0012] Optionally, the method also includes the following steps: determining the learner's cognitive state level based on the learner's final scores across the four cognitive dimensions; The determination of the learner's cognitive state level specifically includes: comparing the final score of each cognitive dimension with the corresponding preset cognitive dimension abnormality interval to obtain the abnormality level of each cognitive dimension, and outputting the learner's cognitive state level based on the abnormality level of each cognitive dimension and preset rules.
[0013] Secondly, a speech-based learner cognitive state assessment system is provided, including: The voice acquisition module is used to acquire and preprocess voice data under different vocal tasks. The voice data includes reference voice, reading voice, and answering voice. The feature extraction module is used to separate speech micro-vibration features reflecting speech tremor from the reference speech through LPC and ICA analysis methods, reconstruct vocal rhythm control features from the reading speech through Hilbert transform and nonlinear dynamics methods, and extract the respiratory speech coordination efficiency features of the answer speech through energy envelope analysis and phase coordination calculation. The assessment module maps speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. It then performs a weighted fusion to make a comprehensive decision and obtain the learner's final scores for the four cognitive dimensions.
[0014] Thirdly, a computer device is provided, including a processor and a memory; wherein the processor executes a computer program stored in the memory to implement the steps of the speech-based learner cognitive state assessment method according to any one of the first aspects.
[0015] Fourthly, a computer-readable storage medium is provided for storing a computer program; when executed by a processor, the computer program implements the steps of the speech-based learner cognitive state assessment method as described in any one of the first aspects.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features extracted in this invention are all biomarkers that deeply influence cognition. Speech microvibration features map fatigue and stress levels; the stability and complexity of vocal rhythm control features directly reflect attention and psychological load; and the dysregulation of respiratory-voice coordination efficiency features is a sensitive indicator of cognitive resource exhaustion and stress. Using these three features for cognitive assessment in distance learning eliminates reliance on subjective scales and has an objective biological basis, thereby improving the accuracy and interpretability of learner cognitive state assessment in distance learning and facilitating precise teaching interventions for learners. In addition, after mapping the cognitive dimensions of the above three features separately, a comprehensive decision is made through weighted fusion. Compared with a purely data-driven black box model, this method is more stable and reliable under small sample conditions. Even if the quality of a certain type of signal deteriorates, this method can still perform robust assessment, enhancing the robustness of this method in practical applications. Attached Figure Description
[0017] Figure 1 This is a flowchart of the speech-based learner cognitive state assessment method of the present invention; Figure 2 This is a step sequence diagram of the speech-based learner cognitive state assessment method of the present invention. Detailed Implementation
[0018] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variations thereof in the specification, claims and the above-mentioned drawings of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or devices.
[0019] Example 1: like Figure 1 and Figure 2 As shown, the speech-based learner cognitive state assessment method includes the following steps: S1: Collect and preprocess speech data under different vocalization tasks. The speech data includes baseline speech, reading speech, and answer speech. S2: The speech micro-vibration features reflecting speech tremor are separated from the reference speech by LPC and ICA analysis methods. The vocal rhythm control features are reconstructed from the reading speech by Hilbert transform and nonlinear dynamics methods. The respiratory speech coordination efficiency features of the answer speech are extracted by energy envelope analysis and phase coordination calculation. S3: The features of speech microvibration, vocal rhythm control, and respiratory-voice coordination efficiency are mapped to the initial scores of four cognitive dimensions: attention, fatigue, psychological load, and stress level, respectively. The final scores of the learner's four cognitive dimensions are obtained by weighted fusion for comprehensive decision-making.
[0020] In this embodiment, the reference speech is the learner's continuous vowel pronunciation, such as continuously pronouncing the vowel / a: / for 10 seconds. The reading speech is the learner's pronunciation when reading a pre-given standard text of 200-300 words; the answering speech is the learner's pronunciation when answering preset questions, such as answering speech that lasts for 1-2 minutes.
[0021] When collecting different types of speech data, independent segmented acquisition or continuous acquisition can be used. For continuously acquired speech, segmentation is required based on learner metadata tags. Existing technologies can be used for segmentation. Learner metadata includes: student ID, timestamp, device information, task type tag, and sequence record.
[0022] Step S1 involves preprocessing the acquired speech data, including: filtering out low-frequency noise below 50Hz and high-frequency components above 20kHz, while retaining the core speech frequency band. A first-order high-pass filter is used to boost the high-frequency components to compensate for natural attenuation during speech propagation. The amplitude of the speech data is scaled proportionally to the peak value to a standard range to eliminate the influence of volume differences. Further preprocessing steps, such as pre-emphasis, framing, and windowing, are also required for the acquired speech; specific details can be found in existing technologies.
[0023] In this embodiment, a learner cognitive state assessment method based on speech signals further includes: performing real-time validity analysis on different collected speech data; if the preset validity conditions are met, the current speech data is retained; if the preset validity conditions are not met, the current speech data is deleted and the re-collection of speech data is initiated.
[0024] The preset validity conditions for real-time validity analysis include: signal-to-noise ratio. , the maximum amplitude Continuous silence lasts less than 10% of the total duration.
[0025] Step S2 specifically includes: S2.1: Acquisition of speech microvibration features.
[0026] Speech micro-vibration features reflecting speech tremor were separated from the reference speech using LPC and ICA analysis methods. Speech micro-vibration features are quantitative biomarkers that convert the reference speech into electromyographic signal states.
[0027] Step S2.1 specifically includes the following steps: S2.1.1: Perform LPC analysis on each frame of reference speech to obtain a single-channel residual signal.
[0028] The specific steps of Linear Predictive Coding (LPC) analysis are as follows: calculate the autocorrelation function of each frame of the reference speech, and use the autocorrelation method or covariance method to solve for the LPC coefficients. Use the obtained LPC coefficients to form an inverse filter to perform inverse filtering on the reference speech to obtain the residual of each frame. Finally, use the overlapping addition method to concatenate the parameters of all frames to obtain a single-channel residual signal.
[0029] S2.1.2: A multi-channel observation signal is constructed by embedding the single-channel residual signal through time delay.
[0030] Multi-channel observation signals for: , where superscript This indicates the transpose operation. For the first Single-channel residual signal at each sampling point The set time delay parameter, This represents the sampling point number of the single-channel residual signal. The dimension embedded for time delay.
[0031] S2.1.3: Perform ICA analysis on the multi-channel observation signal to obtain several independent components.
[0032] Independent Component Analysis (ICA) specifically involves: Assuming the multi-channel observed signal is a linear mixture of several statistically independent source signals, the goal of ICA analysis is to find a separation matrix. Specifically, the FastICA (Fast Independent Component Analysis) algorithm is used to find the separation matrix that maximizes the negative entropy of the output independent components. Once the separation matrix is obtained, the multi-channel observed signal can be decomposed into several independent components. The formula for maximizing negative entropy is: ; in, It is an approximation of negative entropy. For the expected value, As the independent components after separation, It is not a quadratic function. Let them be Gaussian variables. It indicates that they are directly proportional.
[0033] S2.1.4: Based on the physiological characteristics of speech microvibrations, select the component with the most concentrated energy from all independent components as the microvibration signal.
[0034] Power spectral density is calculated for each independent component, and then the component with the most concentrated energy in the 8-30Hz range is selected as the final microvibration signal based on the physiological characteristics of speech microvibration (the frequency of speech microvibration is usually in the range of 8-30Hz).
[0035] S2.1.5: Extract speech microvibration features from the microvibration signal, including vibration energy, frequency band energy ratio, and sample entropy.
[0036] Tremor energy is used to measure the overall intensity of microtremor signals; the specific formula is as follows: ; in, The tremor energy is the microtremor signal. An increase in the tremor energy value relative to the daily value usually indicates increased tension. microtremor signal Length, For the first The final micro-tremor signal at each sampling point This indicates taking the square modulo 1.
[0037] The band energy ratio is used to reveal the energy frequency distribution pattern of microtremor signals. The changes in the energy proportion of different bands are related to different physiological and psychological states. The band energy ratio of this application includes the low-frequency band energy ratio and the high-frequency band energy ratio. The specific calculation formula is as follows: ; ; in, For low-frequency energy ratio, For high-frequency band energy ratio, The micro-tremor signal estimated using the Welch method (a classic algorithm for power spectral density estimation) The power spectral density, For frequency. 8-12Hz: usually corresponds to the core frequency band of physiological tremor; 12-20Hz: usually reflects higher frequency neural oscillation components or myogenic tremor; 8-30 Hz: preset total energy reference broadband range. Reduce at the same time An increase indicates a shift in the tremor spectrum to higher frequencies, suggesting that the learner may be experiencing fatigue. An elevated level indicates that the tremor energy is concentrated in the physiological tremor frequency range, which may indicate that the learner is experiencing tension.
[0038] Sample entropy is used to quantify the complexity and unpredictability of microtremor signals, and the formula is: ; in, The sample entropy is the microtremor signal. A decrease in sample entropy may indicate fatigue or excessive cognitive load. The embedding dimension used to construct the vector length. This is the tolerance threshold. In order to be within the tolerance Inside, with the indivual The number of templates that are similar to the template vectors in the dimension.
[0039] Voice microvibration characteristics for: ,in, This indicates the transpose operation.
[0040] S2.2: Acquisition of vocal rhythm control characteristics.
[0041] By analyzing the fluctuation patterns of rhythms using Hilbert transform, the vocal control signals that drive speech in the brain are reconstructed, and their complexity is quantified using nonlinear dynamics methods to extract the vocal rhythm control features that reflect the brain's cognitive control function.
[0042] Step S2.2 specifically includes the following steps: S2.2.1: Extract the speech envelope signal from the read-out speech using Hilbert transform.
[0043] The rhythmic variation of speech amplitude is directly controlled by vocalization signals emitted by the brain. Therefore, analyzing the rhythmic variation of speech amplitude can extract the speech envelope signal from the read-aloud speech using the Hilbert transform. The formula for extracting the speech envelope signal using the Hilbert transform is: ; in, The extracted speech envelope signal, For preprocessing The audio recording of the moment. Indicates to Perform the Hilbert transform; the Hilbert transform will A 90-degree phase shift is used to obtain the imaginary part of the analytic signal.
[0044] After extracting the speech envelope signal, it is usually necessary to preprocess the speech envelope signal, which includes resampling and normalization operations.
[0045] S2.2.2: Extract the scaling exponent of the speech envelope signal using a detrended fluctuation analysis method. .
[0046] Detrended Fluctuation Analysis (DFA) is primarily used to analyze whether long-range correlations exist in vocal rhythm control, i.e., whether the current instruction is intrinsically linked to past instructions. Step S2.2.2 specifically involves: The speech envelope signal is converted into a random walk process by integration. The formula for integration is: ; in, The integrated signal The speech envelope signal after preprocessing The envelope value of each sampling point The number of sampling points for the speech envelope signal. This represents the mean of the speech envelope signal.
[0047] The integrated signal is divided into several non-overlapping intervals of equal length, with interval length being... After performing local detrending analysis on each interval, the average fluctuation of all intervals is calculated, and the scaling exponent is obtained by fitting the average fluctuation onto a double logarithmic coordinate system.
[0048] In local detrending analysis, a straight line is fitted using the least squares method within each interval, the detrending variance within that interval is calculated, and the average fluctuation is calculated for all intervals based on each detrending variance. Finally, the slope is obtained by fitting the line onto a log-log coordinate system, which is the scaling exponent. The specific formula is as follows: ; in, Interval length The average fluctuation below is the intercept constant. When This indicates that the rhythm of the reading aloud is close to random noise, thus suggesting a cognitive state such as inattention or fatigue. To set a threshold, it is usually less than or equal to 0.1; when This indicates that the rhythm of the reading aloud is close to the ideal cognitive state. To set a threshold, it is usually less than or equal to 0.1; when Some believe that the rhythm of the reading aloud is too regular and rigid, indicating deliberate control or tension.
[0049] S2.2.3: Using multi-scale entropy analysis, plot the MSE curve of the speech envelope signal and extract the area under the curve. and complexity index .
[0050] Multiscale entropy (MSE) analysis is used to assess the brain's information processing capacity and cognitive resilience across different time scales. Specifically, it involves analyzing the speech envelope signal at different scales. The sample is smoothed down to coarsely divide it into several coarse-grained sequences. The sample entropy is calculated for each coarse-grained sequence. The calculation formula can refer to the calculation formula of the reference speech sample entropy.
[0051] Plotting MSE curves: by scale Plot the MSE curve with the x-axis representing the sample entropy corresponding to the scale and the y-axis representing the scale, and calculate the area under the curve. and complexity index .
[0052] Calculate the area under the curve and complexity index The formula is: ; ; in, For the largest scale, The scale corresponding to the coarse-grained sequence The sample entropy below; A higher value indicates that the brain maintains a high degree of complexity and adaptability across all time scales, meaning better cognitive resilience. An elevated value indicates that the brain needs to rely more on long-term information integration, which means that the learner's processing speed is slower or the task is more difficult.
[0053] S2.2.4: The recursive graph of the speech envelope signal is calculated using the phase space reconstruction method, and the deterministic indices of the recursive graph are extracted using the recursive quantitative analysis method. and laminar flow index .
[0054] Deterministic indicators for extracting recursion graphs and laminar flow index Specifically, the one-dimensional speech envelope signal is first reconstructed into a multi-dimensional phase space trajectory. Then, for any two points in the reconstructed multi-dimensional phase space trajectory, the distance between them is calculated (usually Euclidean distance is chosen). If the distance is less than a set value, a point is marked at the corresponding position in the recursive graph, thus generating the recursive graph. Finally, deterministic indices are extracted from the recursive graph. and laminar flow index Certainty indicators The proportion of recursive points in the diagonal structure of the recursive graph, a laminar flow index. This represents the proportion of recursive points in the vertical structure of the recursive graph.
[0055] S2.2.5: Using multifractal analysis, the width of the multifractal spectrum is extracted after calculating the multifractal spectrum of the speech envelope signal. .
[0056] Multifractal Detrended Fluctuation Analysis (MF-DFA) is used to detect whether a speech envelope signal exhibits multifractal properties. Specifically, it transforms the speech envelope signal into a random walk process through integration, divides the integrated sequence into several equal-length non-overlapping intervals, and then applies a set of predefined parameters. Order value, calculation The first-order wave function, and for each The generalized Hurst exponent is obtained by fitting the first-order wave function to a double logarithmic coordinate system. Based on the generalized Hurst exponent After calculating the singularity index of each fractal spectrum, the width of the multifractal spectrum is output. .
[0057] The formula for calculating the singularity index is: ; ; in, The singularity index, For order, for The derivative of the multifractal spectrum is the fractal dimension corresponding to the singular index. and These are the maximum and minimum values of the singularity index of the multifractal spectrum.
[0058] S2.2.6: Scale index Width of multifractal spectrum The area under the MSE curve and complexity index and deterministic indicators of recursion graphs and laminar flow index Integrated into vocal rhythm control characteristics .
[0059] ; in, This indicates the transpose operation.
[0060] S2.3: Acquisition of respiratory speech coordination efficiency characteristics.
[0061] By analyzing the energy fluctuation patterns in the answering voice through energy envelope analysis and phase coordination calculation, the breathing cycle is automatically identified, and breathing parameters, voice output efficiency, and coordination between breathing and voice are quantified. Features reflecting the breathing-voice coordination efficiency that reflect physical and mental state and stress level are extracted.
[0062] Step S2.3 specifically includes the following steps: S2.3.1: Detect respiratory cycles from the voice responses using energy envelope analysis, and record the inspiratory duration, expiratory duration, respiratory cycle length, inspiratory-expiratory ratio, and peak expiratory energy for each respiratory cycle.
[0063] Specifically, the speech energy envelope of the answer speech is extracted, and a short-time energy sequence of the answer speech is obtained. Its rapid fluctuations are captured, and a low-pass filter (cutoff frequency 2-4Hz) is applied to the short-time energy sequence to extract a smooth energy envelope. A peak detection algorithm is used to detect local minima on the smooth energy envelope, and each local minima is marked as the start of inspiration. Similarly, a peak detection algorithm is used to detect local maxima on the smooth energy envelope, and these local maxima are marked as expiratory peaks. Two adjacent local minima constitute one respiratory cycle. In this embodiment, the length of the expiratory segment (i.e., the interval between the start of the rise in speech energy in the smooth energy envelope and the next local minima) is verified to be within a physiologically reasonable range (0.5-5 seconds) to eliminate false detections. After the above processing, several respiratory cycles are obtained, and the inspiratory duration, expiratory duration, respiratory cycle length, inspiratory-expiratory ratio, and expiratory peak energy of each respiratory cycle are recorded. The formula is: ; Inhalation duration, This refers to the duration of exhalation.
[0064] S2.3.2: Within each expiratory segment, detect the number of syllables and calculate the coefficients of variation for syllable rate, speech production ratio, and inhalation-expiration ratio.
[0065] The method for detecting the number of syllables can refer to existing technologies, such as peak detection based on the speech energy envelope. The coefficient of variation of the inspiratory-to-expiratory ratio can be calculated using the standard deviation and the mean; the formulas for syllable rate and speech production ratio are: ; ; in, Syllable rate, The number of syllables detected during the expiratory phase; an abnormally high syllable rate indicates rapid breathing, while a low rate indicates fatigue or hesitation. For speech production ratio, The time period with voice activity (non-silent). This is the respiratory cycle.
[0066] S2.3.3: Map the timeline of each respiratory cycle to... On the phase ring, the phase consistency index of all breathing cycles is obtained by using the phase value of the start time of each syllable on the phase ring.
[0067] ; ; in, The phase consistency index has a value of [value missing]. , for The complex vector form, The imaginary unit, For the first The phase value of the syllable start time on the phase ring For the first The start time of each syllable. The closer the value is to 1, the more consistent the starting phase of all breathing cycles is, and the better the coordination between breathing and speech. The closer the value is to 0, the more it indicates that breathing cannot effectively support the process, suggesting tension or dysfunction.
[0068] S2.3.4: Compare the vocal energy envelope of all exhalation phases with the preset ideal envelope, calculate the ratio between the two, and obtain the energy output ratio. The ideal envelope is the rectangle formed by maintaining peak expiratory energy throughout all exhalation phases.
[0069] ; in, Indicates the energy output ratio. A value of approximately 0.5 indicates that the energy distribution is symmetrical, and breathing support is economical and stable. A value less than 0.5 indicates an asymmetrical energy distribution, shallow or rapid breathing, or insufficient support, making it difficult to speak. The exhalation phase The exhaled energy during the voice recording of answering questions. Peak expiratory energy.
[0070] S2.3.5: Calculate the learner's overall stress index based on the phase consistency index, energy output ratio, and the learner's current respiratory rate baseline.
[0071] The overall stress index is a quantitative score of the learner's overall stress level. The calculation formula is: ; in, Respiratory rate, The baseline value of the current learner's respiratory rate, as measured in advance. The standard deviation of baseline respiratory rate, , and The weights are for breathing, coordination, and energy efficiency, respectively.
[0072] S2.3.6: Integrate the inspiratory-to-expiratory ratio, syllable rate, speech production ratio, energy output ratio, coefficient of variation of the inspiratory-to-expiratory ratio, phase consistency index, and comprehensive stress index into a respiratory-speech coordination efficiency feature.
[0073] ; in, The coefficient of variation of the inspiratory-to-expiratory ratio. This refers to the efficiency characteristics of respiratory-voice coordination.
[0074] Step S3 specifically includes the following steps: S3.1: Preprocess the extracted speech microvibration features, vocal rhythm control features, and respiratory speech coordination efficiency features respectively.
[0075] To each , and The three types of features are standardized separately, for example, using the Z-score method, and then PCA is used to reduce the dimensionality of each type of feature.
[0076] S3.2: Three independent random forest models are used to predict the initial scores of each feature in the four cognitive dimensions of attention, fatigue, psychological load, and stress level.
[0077] All three random forest models were trained before use. Specifically, features (vocal microvibration features, vocal rhythm control features, or respiratory-vocal coordination efficiency features) and labels (initial scores across four cognitive dimensions) were used to train the random forest models. The structure of the random forest models referenced existing techniques.
[0078] S3.3: Calculate and output the final scores for each cognitive dimension according to the following formula: ; in, For learners The final score for each cognitive dimension, , and These are the characteristics of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency, respectively, at the [number]th [year]. The weighting coefficients of each cognitive dimension, , and The features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model are respectively the features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model in the 1st... Initial scores for each cognitive dimension.
[0079] In this embodiment, step S2.4 is also included: determining the learner's cognitive state level based on the learner's final scores in the four cognitive dimensions; Determining the learner's cognitive state level specifically includes: comparing the final score of each cognitive dimension with the corresponding preset abnormal range of the cognitive dimension to obtain the abnormal level of each cognitive dimension, and outputting the learner's cognitive state level based on the abnormal level of each cognitive dimension and preset rules.
[0080] Cognitive abnormality levels are typically categorized into normal, Level I abnormality, and Level II abnormality based on a defined abnormality range. Level II abnormality deviates more significantly from the normal level compared to Level I abnormality. A specific classification rule for cognitive state levels is provided in Table 1: Table 1 Cognitive State Classification Rules
[0081] Example 2: A speech-based learner cognitive state assessment system, comprising: The voice acquisition module is used to acquire and preprocess voice data under different vocal tasks. The voice data includes reference voice, reading voice, and answering voice. The feature extraction module is used to separate speech micro-vibration features reflecting speech tremor from the reference speech through LPC and ICA analysis methods, reconstruct vocal rhythm control features from the reading speech through Hilbert transform and nonlinear dynamics methods, and extract the respiratory speech coordination efficiency features of the answer speech through energy envelope analysis and phase coordination calculation. The assessment module maps speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. It then performs a weighted fusion to make a comprehensive decision and obtain the learner's final scores for the four cognitive dimensions.
[0082] For more specific details about the above system, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0083] Example 3: The present invention provides a computer device, including a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the steps of the above-described speech-based learner cognitive state assessment method.
[0084] For more detailed information on the above methods, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0085] Example 4: The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, it implements the steps of the above-described speech-based learner cognitive state assessment method.
[0086] For more detailed information on the above methods, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0087] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The systems, devices, and storage media disclosed in the embodiments are described simply because they correspond to the methods disclosed in the embodiments; relevant details can be found in the method section.
[0088] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0089] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A speech-based method for assessing learners' cognitive states, characterized in that, Includes the following steps: Collect and preprocess speech data under different vocalization tasks. The speech data includes reference speech, reading speech, and answer speech. The micro-vibration features reflecting speech tremor were separated from the reference speech using LPC and ICA analysis methods. The vocal rhythm control features were reconstructed from the reading speech using Hilbert transform and nonlinear dynamics methods. The respiratory speech coordination efficiency features of the answer speech were extracted using energy envelope analysis and phase coordination calculation. The features of speech microvibration, vocal rhythm control, and respiratory-voice coordination efficiency are mapped to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. Weighted fusion is then used to make a comprehensive decision to obtain the learner's final scores for the four cognitive dimensions. The process of separating speech micro-vibration features reflecting speech tremor from the reference speech using LPC and ICA analysis methods specifically involves: LPC analysis is performed on each frame of reference speech to obtain a single-channel residual signal; A multi-channel observation signal is constructed by embedding the single-channel residual signal through a time delay; ICA analysis was performed on the multi-channel observation signal to obtain several independent components; Based on the physiological characteristics of speech microvibrations, the component with the most concentrated energy was selected from all independent components as the microvibration signal. Extract speech microvibration features from microvibration signals, including vibration energy, frequency band energy ratio, and sample entropy; The method of reconstructing the vocal rhythm control features from spoken text using Hilbert transform and nonlinear dynamics specifically involves: Extracting the speech envelope signal from the read-out speech using Hilbert transform; The scaling exponent of the speech envelope signal is extracted using a detrended fluctuation analysis method. ; The MSE curve of the speech envelope signal was plotted using multi-scale entropy analysis, and the area under the curve was extracted. and complexity index ; The recursive graph of the speech envelope signal is calculated using the phase space reconstruction method, and deterministic indices of the recursive graph are extracted using recursive quantitative analysis. and laminar flow index ; The width of the multifractal spectrum is extracted after calculating the multifractal spectrum of the speech envelope signal using multifractal analysis. ; Scaling index Width of multifractal spectrum The area under the MSE curve and complexity index and deterministic indicators of recursion graphs and laminar flow index This is integrated into vocal rhythm control characteristics; The respiratory-voice coordination efficiency features of the answer speech were extracted through energy envelope analysis and phase coordination calculation, specifically: Respiratory cycles were detected from the voice responses using energy envelope analysis, and the inspiratory duration, expiratory duration, respiratory cycle length, inspiratory-expiratory ratio, and peak expiratory energy of each respiratory cycle were recorded. Within each expiratory segment, the number of syllables was detected, and the coefficients of variation for syllable rate, speech production ratio, and inspiratory-expiratory ratio were calculated. Map the timeline of each respiratory cycle to On the phase ring, the phase consistency index of all breathing cycles is obtained by using the phase value of the start time of each syllable on the phase ring; The energy envelope of the speech during all exhalation phases is compared with the preset ideal envelope, and the ratio between the two is calculated to obtain the energy output ratio. The ideal envelope is the rectangle formed by maintaining the peak energy of exhalation during all exhalation phases. The overall stress index of the current learner is calculated based on the phase consistency index, energy output ratio, and the current learner's baseline respiratory rate. The inspiratory-to-expiratory ratio, syllable rate, speech production ratio, energy output ratio, coefficient of variation of the inspiratory-to-expiratory ratio, phase consistency index, and comprehensive stress index are integrated into the respiratory-speech coordination efficiency characteristics.
2. The speech-based learner cognitive state assessment method according to claim 1, characterized in that, It also includes the following steps: The system performs real-time validity analysis on the different voice data collected. If the preset validity conditions are met, the current voice data is retained. If the preset validity conditions are not met, the current voice data is deleted and the voice data is re-collected.
3. The speech-based learner cognitive state assessment method according to claim 1, characterized in that, The process involves mapping speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. A weighted fusion method is then used for comprehensive decision-making to obtain the learner's final scores across these four cognitive dimensions. Specifically: The extracted speech microvibration features, vocal rhythm control features, and respiratory speech coordination efficiency features were preprocessed respectively. Three independent random forest models were used to predict the initial scores of each feature in the four cognitive dimensions of attention, fatigue, psychological load, and stress level. The final score for each cognitive dimension is calculated and output according to the following formula: ; in, For learners The final score for each cognitive dimension, , and These are the characteristics of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency, respectively, at the [number]th [year]. The weighting coefficients of each cognitive dimension, , and The features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model are respectively the features of speech microvibration, vocal rhythm control, and respiratory-speech coordination efficiency predicted by the random forest model in the 1st... Initial scores for each cognitive dimension.
4. The speech-based learner cognitive state assessment method according to claim 1, characterized in that, It also includes the following steps: The learner's cognitive state level is determined based on the final scores of the learner's four cognitive dimensions; The determination of the learner's cognitive state level specifically includes: comparing the final score of each cognitive dimension with the corresponding preset cognitive dimension abnormality interval to obtain the abnormality level of each cognitive dimension, and outputting the learner's cognitive state level based on the abnormality level of each cognitive dimension and preset rules.
5. A speech-based learner cognitive state assessment system, comprising the steps of the speech-based learner cognitive state assessment method according to any one of claims 1-4, characterized in that, include: The voice acquisition module is used to acquire and preprocess voice data under different vocal tasks. The voice data includes reference voice, reading voice, and answering voice. The feature extraction module is used to separate speech micro-vibration features reflecting speech tremor from the reference speech through LPC and ICA analysis methods, reconstruct vocal rhythm control features from the reading speech through Hilbert transform and nonlinear dynamics methods, and extract the respiratory speech coordination efficiency features of the answer speech through energy envelope analysis and phase coordination calculation. The assessment module maps speech microvibration features, vocal rhythm control features, and respiratory-voice coordination efficiency features to initial scores for four cognitive dimensions: attention, fatigue, psychological load, and stress level. It then performs a weighted fusion to make a comprehensive decision and obtain the learner's final scores for the four cognitive dimensions.
6. A computer device, characterized in that, It includes a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the steps of the speech-based learner cognitive state assessment method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, Used to store computer programs; when executed by a processor, the computer programs implement the steps of the speech-based learner cognitive state assessment method according to any one of claims 1-4.