English oral pronunciation quality evaluation method based on multi-modal speech feature analysis
By employing multimodal speech feature analysis methods, a joint feature map and a multidimensional representation model are constructed, which solves the limitations of existing technologies in assessing the quality of spoken English pronunciation. This enables accurate assessment of pronunciation details and personalized teaching support, thereby improving the adaptability and scientific rigor of the assessment system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-03-31
AI Technical Summary
Existing English pronunciation quality assessment technologies lack comprehensive utilization of the rich time-frequency domain information and phase features in speech signals. This results in limitations in the assessment system when dealing with subtle differences in pronunciation and variations in individual characteristics. Consequently, they cannot meet the personalized needs of different native language backgrounds, different learning stages, and different learning goals, thus affecting the scientific and personalized teaching effectiveness of language learning.
A multimodal speech feature analysis method is adopted to construct a joint feature map by acquiring the amplitude spectrum, phase spectrum, group delay features and cepstral domain features of the speech signal. Combined with the phase-group delay coupling response value, the articulation detail feature sequence is extracted, and a multidimensional representation model is constructed. Prosodic rhythm parameters are integrated for comprehensive evaluation and a visual diagnostic report is generated.
It significantly enhances the ability to capture transient changes and micro-dynamic processes in pronunciation, accurately describes the multi-dimensional distribution characteristics of pronunciation problems, supports fair assessment and personalized teaching strategies for learners of different levels, improves the adaptability and scientific rigor of the assessment system, and provides a precise pronunciation quality assessment tool for modern personalized language education.
Smart Images

Figure CN121528247B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech analysis technology, and more specifically, to a method for assessing the quality of spoken English pronunciation based on multimodal speech feature analysis. Background Technology
[0002] In an era of accelerated globalization and increasingly frequent international exchanges, English, as a global language, plays an irreplaceable bridging role in education, business, science and technology, and spoken English proficiency has become a core indicator for measuring comprehensive language literacy. With the widespread application of artificial intelligence in education and the deepening popularity of personalized learning concepts, intelligent pronunciation assessment systems based on speech signal processing are playing an increasingly important role in modern language teaching. However, the phonetic mechanisms involved in spoken English are extremely complex, encompassing the interaction of multiple levels, including speech production, acoustic propagation, and perceptual cognition. In particular, the interference from the native language's phonetic system, individual differences in speech organs, and cognitive habit transfer exhibited by non-native speakers during pronunciation pose serious challenges to the accurate assessment and detailed analysis of pronunciation quality, directly affecting the scientific level of language learning and the effectiveness of personalized teaching.
[0003] Current pronunciation quality assessment technologies typically rely on basic acoustic parameter extraction and simplified pattern recognition algorithms, employing single-dimensional feature analysis and fixed threshold judgment strategies. This lack of comprehensive utilization of the rich time-frequency domain information and phase features inherent in speech signals leads to significant limitations in handling subtle pronunciation differences, complex speech phenomena, and individual characteristic variations. With the rapid expansion of online education and the increasing diversity of learners' backgrounds, traditional "one-size-fits-all" assessment models are increasingly unable to meet the personalized needs of learners with different native language backgrounds, learning stages, and learning goals. They fail to accurately identify specific pronunciation problems and lack targeted improvement guidance mechanisms. This technological bottleneck not only restricts the teaching effectiveness of intelligent language learning systems but also severely hinders the development of speech assessment technology towards higher accuracy, greater adaptability, and a better user experience.
[0004] In view of this, the present invention proposes an English spoken pronunciation quality assessment method based on multimodal speech feature analysis to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution:
[0006] Methods for assessing the quality of spoken English pronunciation based on multimodal speech feature analysis include:
[0007] The speech signal of the target user's spoken English is acquired, and the speech signal is decomposed into multiple domains to obtain multimodal speech features, which include amplitude spectrum features, phase spectrum features, group delay features, and cepstral domain features.
[0008] Based on the time-frequency correspondence between the phase spectrum features and the group delay features, a joint feature map is constructed; based on the temporal evolution pattern of the energy accumulation region in the joint feature map and combined with the transient change features of the amplitude spectrum features, feature extraction is performed to obtain the articulation detail feature sequence;
[0009] The sequence of detailed pronunciation features is matched with a preset standard pronunciation template at multiple scales, and a multi-dimensional representation model is constructed based on the multi-scale matching results. The fine-grained quality score of the phoneme unit in each multi-dimensional representation model is calculated by combining the prosodic rhythm parameters in the cepstral domain features. The prosodic coherence score and the overall fluency score are then integrated to generate a comprehensive pronunciation quality assessment result and a visualized diagnostic report of pronunciation deviation.
[0010] Furthermore, the process of constructing the joint feature map includes:
[0011] The phase spectrum features are subjected to phase expansion processing to obtain a continuous phase spectrum. The continuous phase spectrum is subjected to frequency dimension difference operation to obtain the phase derivative spectrum. The phase derivative spectrum is then subjected to median filtering processing to obtain the group delay estimate.
[0012] The deviation between the estimated group delay and the actual measured value of the group delay characteristics is calculated, and a group delay reliability weight map is constructed based on the deviation calculation results.
[0013] Based on the reliable weighted map of the group delay, the weighted correlation coefficient between the phase derivative spectrum and the group delay feature within a preset time-frequency analysis window is calculated and denoted as the phase-group delay coupled response value; the phase-group delay joint feature map is constructed based on the obtained phase-group delay coupled response value.
[0014] Furthermore, the process of obtaining the continuous phase spectrum includes:
[0015] The phase spectrum features are scanned frequency by frequency to identify and mark the frequency points with phase abrupt changes, which are then marked as potential jump points; and local frequency estimates corresponding to multiple frequency points adjacent to the potential jump points are obtained.
[0016] Based on the local frequency estimate, predict the expected phase value at the corresponding potential transition point, and calculate the residual between the expected phase value and the actual phase value to obtain the phase residual.
[0017] Based on the phase residual, determine whether the corresponding potential transition point is a real phase entanglement point. If it is not a phase entanglement point, no other operation is performed; if it is a phase entanglement point, perform phase compensation on the phase entanglement point to obtain a continuous phase spectrum.
[0018] Furthermore, the process of obtaining the pronunciation detail feature sequence includes:
[0019] Energy accumulation region detection is performed on the joint feature map to identify connected regions whose phase-group delay coupling response values exceed a preset response threshold, which are then used as feature response regions.
[0020] The temporal boundary change rate within the characteristic response region is obtained, and the phoneme boundary position is determined by combining the transient energy of the amplitude spectrum feature. The consonant phoneme interval and vowel phoneme interval are located based on the phoneme boundary position.
[0021] Within the located consonant phoneme range, the phase jump mode and group delay feature impulse response mode of the joint feature map in the high-frequency band are extracted and recorded as the consonant initiation transient features. The consonant initiation transient features include the phase jump amplitude of plosives, the group delay fluctuation period of fricatives, and the mixed transient mode of affricates.
[0022] Within the located vowel phoneme range, extract the formant frequency trajectory from the amplitude spectrum features, and calculate the phase stability index of the formant frequency trajectory at the corresponding position in the joint feature map, which is denoted as the vowel formant migration trajectory.
[0023] The transition segments between adjacent phonemes are obtained based on the phoneme boundary positions, and the continuity measure of the phase spectrum feature and the smoothness measure of the group delay feature within the corresponding transition segment are calculated; these are used as phase continuity features; the continuity measure is obtained by calculating the energy of the second derivative of the phase spectrum feature, and the smoothness measure is obtained by calculating the local variance of the group delay feature.
[0024] The consonant initiation transient features, vowel formant migration trajectories, and phase continuity features are arranged in chronological order to construct the articulation detail feature sequence.
[0025] Furthermore, the method for determining the phoneme boundary position is that the time difference between the peak moment of the boundary change rate and the peak moment of the energy rise slope is less than a preset time tolerance threshold.
[0026] Furthermore, the process of obtaining the phoneme boundary position includes:
[0027] The temporal boundary of the feature response region is calculated by first-order difference to obtain the time series of the boundary change rate; and the time series of the boundary change rate is subjected to adaptive threshold peak detection to obtain the peak candidate time set of the boundary change rate.
[0028] Energy integration is performed on the amplitude spectrum feature over a wide frequency range to obtain the transient energy envelope curve. Differentiation is then performed on the transient energy envelope curve to obtain the energy rise slope at different times. The energy rise slopes are arranged in chronological order to form a time series of energy rise slopes. A fixed threshold peak detection is performed on the time series of energy rise slopes to obtain a set of candidate times for energy rise slope peaks.
[0029] Calculate the time matching degree between the set of candidate peak times of the boundary change rate and the set of candidate peak times of the energy rise slope, confirm the phoneme boundary position based on the matching degree, and make fine adjustments to the confirmed phoneme boundary position.
[0030] Furthermore, the construction process of the multidimensional representation model includes:
[0031] Select a standard pronunciation template corresponding to the acquired speech signal from a preset standard pronunciation template library. The standard pronunciation template contains a standard pronunciation detail feature sequence.
[0032] The articulation detail feature sequence is dynamically time-warped and aligned with the articulation detail feature sequence of the standard articulation template to obtain a time mapping function, and the time-domain deviation component of each phoneme unit in the articulation detail feature sequence is calculated based on the time mapping function.
[0033] After time normalization and alignment are completed, the deviation components of the consonant initial transient features and vowel formant migration trajectories are calculated with respect to the standard articulation template to obtain frequency domain deviation components and phase domain deviation components. The frequency domain deviation components include the plosive frequency center offset, the fricative energy distribution deviation, and the consonant bandwidth difference value. The phase domain deviation components include the phase stability deviation of the formant position, the phase continuity deviation of the vowel transition segment, and the phase dispersion deviation of the vowel core segment.
[0034] The multidimensional representation model is constructed based on the time-domain deviation component, the frequency-domain deviation component, and the phase-domain deviation component.
[0035] Furthermore, the process of generating a comprehensive pronunciation quality assessment result includes:
[0036] Based on the aforementioned multidimensional representation model, and combined with the prosodic rhythm parameters in the cepstral domain features, an adaptive pronunciation quality assessment model is established.
[0037] The deviation components of each phoneme unit in the multidimensional representation model are normalized according to the adaptive pronunciation quality assessment model to obtain the normalized deviation value, and the basic quality score corresponding to each phoneme unit is calculated based on it.
[0038] The prosodic coherence score between adjacent phoneme units is calculated based on the phase continuity feature; and the pause distribution, speech rate fluctuation, and stress rhythm of the speech signal are extracted based on the prosodic rhythm parameters in the cepstral domain feature; and the overall fluency score is calculated.
[0039] The basic quality score, prosodic coherence score, and overall fluency score are weighted and fused at multiple levels to obtain a comprehensive pronunciation quality assessment result.
[0040] The obtained comprehensive pronunciation quality assessment results are output and a visual diagnostic report is generated.
[0041] Furthermore, the process of establishing the adaptive pronunciation quality assessment model includes:
[0042] Based on the statistical distribution of the deviation components of all phoneme units in the multidimensional representation model, and combined with a predefined level template, the overall pronunciation level estimate of the user is obtained. The overall pronunciation level estimate is obtained by performing weighted cluster analysis on the deviation components.
[0043] Select a dynamic assessment benchmark from a standard pronunciation template library that matches the level template of the overall pronunciation level estimate, wherein the dynamic assessment benchmark includes differentiated error tolerance thresholds for learners of different levels;
[0044] Based on the prosodic rhythm parameters in the cepstral domain features, the native language migration interference pattern of the speech signal is extracted. The native language migration interference pattern includes syllable duration distribution preference, stress position shift pattern and intonation curve morphology features.
[0045] Based on the aforementioned mother tongue transfer interference pattern, identify the systematic and random deviations caused by mother tongue habits in the pronunciation deviations;
[0046] Based on the dynamic evaluation benchmark, different evaluation weights are set for the systematic deviation and the random deviation to construct an adaptive pronunciation quality evaluation model.
[0047] Furthermore, the process of extracting mother tongue transfer interference patterns includes:
[0048] The fundamental frequency contour is extracted from the cepstral domain features to obtain the pitch curve of the speech signal. The pitch curve is then subjected to piecewise linear fitting to obtain the pitch curve morphological features, which include the slope distribution of the curve segments and the position distribution of the curve inflection points.
[0049] The speech signal is segmented into syllables, the duration of each syllable is calculated, and a statistical distribution histogram of syllable duration is constructed, which is denoted as the syllable duration distribution preference.
[0050] By performing cluster analysis on the syllable duration distribution preferences, learners' preferred syllable duration patterns are identified, and compared with the standard syllable duration distribution of the target language to calculate the distribution difference.
[0051] Based on the energy envelope curve of the amplitude spectrum feature, the position of stressed syllables in the speech signal is identified, and the position distribution pattern of stressed syllables in the sentence is calculated and denoted as the stress position offset pattern.
[0052] The intonation curve morphology features, syllable duration distribution preferences, and stress position shift patterns are vectorized and matched with a preset native language prosodic feature library to identify native language transfer interference patterns; the native language prosodic feature library contains typical prosodic feature templates of learners from various native language backgrounds.
[0053] The technical effects and advantages of the English spoken pronunciation quality assessment method based on multimodal speech feature analysis of this invention are as follows:
[0054] This invention utilizes short-time Fourier transform multi-domain decomposition technology to simultaneously extract four modal features of speech signals: amplitude spectrum, phase spectrum, group delay, and cepstral domain. This provides a richer and more accurate acoustic data foundation for comprehensive analysis of pronunciation quality compared to traditional single-feature methods. It innovatively introduces a phase expansion algorithm and a reliability weighting mechanism to construct a joint phase-group delay feature map, effectively mining phase information that is traditionally difficult to utilize, significantly improving the ability to capture transient changes and micro-dynamic processes in pronunciation. A three-dimensional tensor representation model incorporating time-domain, frequency-domain, and phase-domain deviations is constructed, accurately describing the multi-dimensional distribution characteristics and statistical laws of pronunciation problems, solving the technical challenge of comprehensively quantifying complex pronunciation deviations using traditional methods. Furthermore, by combining the physiological mechanisms of speech production, a mechanism for extracting detailed features of different pronunciation types is established, deeply revealing the phase abrupt changes of plosives, the group delay fluctuations of fricatives, and the stability of vowel formants. This study investigates the time-frequency evolution of key pronunciation elements, such as qualitative and phoneme transition continuity; employs dynamic time warping alignment and multi-scale matching techniques to achieve multi-timescale pronunciation quality prediction capabilities, ranging from microsecond-level phoneme boundary detection to second-level sentence fluency assessment; establishes a dynamic assessment benchmark and differentiated weighting mechanism based on overall level estimation, effectively supporting fair assessment and personalized teaching strategy development for learners at different levels (beginner, intermediate, and advanced); significantly improves the system's adaptability to learners with multilingual backgrounds and the scientific rigor of assessment results through mother tongue transfer interference pattern recognition and systematic bias differentiation techniques combined with big data-driven parameter optimization methods; and ultimately forms a complete technical solution integrating multimodal feature analysis, adaptive assessment modeling, and intelligent diagnostic feedback, providing precise pronunciation quality assessment tools and scientific learning improvement decision support for modern personalized language education. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the English spoken pronunciation quality assessment method based on multimodal speech feature analysis according to the present invention. Detailed Implementation
[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] Example 1
[0058] Please see Figure 1 As shown in this embodiment, the English spoken pronunciation quality assessment method based on multimodal speech feature analysis includes:
[0059] The speech signal of the target user's spoken English is acquired and decomposed into multi-domain features to obtain multimodal speech features, including key indicators such as amplitude spectrum features, phase spectrum features, group delay features, and cepstral domain features. The audio signal is acquired in real time using an audio acquisition device, with the acquisition frequency typically set to 16kHz to ensure the capture of the complete frequency range of spoken pronunciation. The multi-domain decomposition process employs short-time Fourier transform technology, ensuring both time and frequency resolution. Amplitude spectrum features reflect the energy distribution of the speech signal, phase spectrum features contain temporal information of pronunciation, group delay features describe the time delay characteristics of frequency components, and cepstral domain features extract the resonance characteristics of the vocal tract. These multimodal features characterize the detailed features of pronunciation from different perspectives, providing a comprehensive data foundation for subsequent refined analysis.
[0060] Based on the time-frequency correspondence between phase spectrum features and group delay features, a joint feature map is constructed. Based on the temporal evolution pattern of energy accumulation regions in the joint feature map and combined with the transient change features of amplitude spectrum features, feature extraction is performed to obtain a sequence of articulation detail features. The joint feature map uses time-frequency as the coordinate axis and the coupled response value of phase gradient and group delay as the pixel intensity, integrating the advantages of two complementary features. Phase spectrum features are sensitive to transient changes in articulation and can capture the onset features of consonants; group delay features are sensitive to formant positions and can accurately locate the frequency structure of vowels. By calculating the weighted correlation between the phase derivative spectrum and group delay features, the constructed joint feature map can clearly present the microscopic dynamic process of articulation on the time-frequency plane, solving the problem that traditional single features cannot simultaneously... The technical challenge lies in capturing transient and steady-state information. Simultaneously, the sequence of articulation details includes consonant initiation transient features, vowel formant migration trajectories, and phase continuity features of phoneme transition segments. Next, energy-concentrated characteristic response regions are identified in the joint feature map, corresponding to key moments and frequency ranges of articulation. Then, the energy rise slope of the amplitude spectrum features is combined to accurately detect phoneme boundary positions. Finally, specific detailed features are extracted for different phoneme types. Consonant initiation transient features capture the instantaneous dynamics of consonants such as plosives and fricatives; vowel formant migration trajectories describe the frequency evolution during vowel articulation; and phase continuity features of phoneme transition segments reflect the fluency of continuous articulation. This multi-level feature extraction lays the foundation for accurate assessment of articulation quality.
[0061] The system performs multi-scale matching of pronunciation detail feature sequences with preset standard pronunciation templates, and constructs a multi-dimensional representation model based on the multi-scale matching results. It then calculates the fine-grained quality score of phoneme units in each multi-dimensional representation model by combining the prosodic rhythm parameters in the cepstral domain features. Finally, it integrates the prosodic coherence score and the overall fluency score to generate a comprehensive pronunciation quality assessment result and a visualized diagnostic report of pronunciation deviations. First, the user's pronunciation is time-aligned with the standard pronunciation using dynamic time warping technology to eliminate the influence of speech rate differences. Then, the deviation components in the time domain, frequency domain, and phase domain are calculated on the aligned time axis. The temporal deviation component reflects the accuracy of pronunciation rhythm and duration, the frequency deviation component characterizes deviations in timbre and frequency characteristics, and the phase deviation component reveals issues with pronunciation coherence and stability. The multidimensional representation model organizes data in tensor form: the first dimension indexes the phoneme sequence, the second dimension indexes the deviation type, and the third dimension stores the statistical characteristic parameters of the deviation, achieving a structured and comprehensive representation of pronunciation deviations. Next, a weighted cluster analysis of all phoneme deviation components is used to estimate the user's overall proficiency level. Then, a dynamic evaluation benchmark matching this level is selected from a standard pronunciation template library to provide a basis for further evaluation. Differentiated error tolerance thresholds are set for users of the same skill level. Furthermore, by analyzing the prosodic parameters of the cepstral domain, systematic and random biases caused by mother tongue transfer are identified, and different evaluation weights are assigned to these two types of biases. This adaptive mechanism ensures the fairness and relevance of the evaluation; novice users are not overly penalized for systematic mother tongue habits, while advanced users are subject to stricter detail requirements. Secondly, a basic quality score is calculated at the phoneme level, using a nonlinear mapping function to convert normalized bias values into 0-100 scores. The design of the mapping function ensures a gradual decrease in scores for small biases and a rapid decrease for large biases. The system first calculates fluency at the speed level, aligning with human perception. Then, it calculates prosodic coherence at the syllable level, focusing on the phase continuity of phoneme transitions. Finally, it calculates overall fluency at the sentence level, considering pause distribution, speech rate fluctuations, and stress rhythm. These three levels of scores are fused using adaptive weighting, with weights dynamically determined based on the user's overall proficiency estimate, ultimately generating a comprehensive pronunciation quality assessment. Finally, the comprehensive pronunciation quality assessment is transformed into an intuitive visual report, including an overall score, detailed phoneme-level scores, highlighted problematic phonemes, formant trajectory comparison charts, and targeted improvement suggestions. This visual diagnostic report helps users clearly understand their pronunciation problems and provides scientific guidance for targeted practice.
[0062] In embodiments of the present invention, the detailed implementation process of constructing the joint feature map includes:
[0063] Phase unrolling is performed on the phase spectrum features to eliminate 2π jumps and obtain a continuous phase spectrum. Phase unrolling is a crucial preprocessing step in constructing the joint feature map. The original phase spectrum contains periodic 2π jumps due to principal value constraints; these jumps are not true phase changes and severely interfere with subsequent phase gradient calculations. The process first scans the phase spectrum frequency-by-frequency to detect potential jump points where the phase difference between adjacent frequencies exceeds π. Then, at each potential jump point, local frequency estimates for its multiple adjacent frequencies are calculated. The slope is obtained by linear regression of the phase differences between adjacent frequencies; this slope represents the local frequency. The expected phase value at the jump point is predicted based on the local frequency estimates and compared with the actual measured value to calculate the residual. If the absolute value of the residual is close to an integer multiple of 2π (the judgment threshold is usually 1.5π), it is confirmed as a true phase entanglement point, and phase compensation of an integer multiple of 2π is performed. The compensation direction is determined by the sign of the local frequency. Compared to the traditional simple differential unrolling method, this adaptive algorithm can effectively handle noise interference and rapid frequency changes, obtaining a more stable continuous phase spectrum.
[0064] A frequency-dimensional difference operation is performed on the continuous phase spectrum to obtain the phase derivative spectrum. Then, median filtering is applied to the phase derivative spectrum in the time dimension to obtain a smoothed group delay estimate. The process first performs a first-order difference operation on the continuous phase spectrum in the frequency dimension, calculating the ratio of the phase difference to the frequency difference between adjacent frequency points to obtain a preliminary phase derivative spectrum. Then, median filtering is applied in the time dimension, with a filtering window length set to 3-5 frames. This effectively suppresses impulse noise and outliers while preserving the true dynamic trend. Median filtering has better edge-preserving properties, making it suitable for processing transient components in the speech signal. The smoothed group delay estimate provides a stable input for subsequent coupled response calculations.
[0065] Based on the deviation between the original measured value and the estimated value of the group delay feature, a group delay reliability weight map is constructed. This reliability weight map is a key mechanism for distinguishing effective information from noise interference, addressing the instability of group delay features within certain frequency ranges. The construction process first calculates the point-by-point deviation between the original group delay measurement (directly calculated via Fourier transform phase spectrum) and the estimated value based on the phase derivative. The absolute value of the deviation reflects the uncertainty of the group delay measurement at that time-frequency point. Then, through statistical distribution analysis of the deviation, a reliability threshold is determined. Regions with deviations less than the threshold are assigned high weights, while regions with deviations greater than the threshold are assigned low weights. Linear interpolation is used in the intermediate regions. Regions with high weights typically correspond to voiced segments and frequency ranges near formants, where group delay feature measurements are stable and phonologically significant. Regions with low weights typically correspond to high-frequency noise segments and spectral dips in voiceless consonants, where group delay features are susceptible to noise. The reliability weight map provides an adaptive confidence modulation mechanism for subsequent weighted analysis. The calculation formula for the weight assignment process is as follows: In the formula, Representing time and frequency points Reliability weight at the location; This represents the original measurement value corresponding to the group delay feature; This represents the group delay estimate. The adjustment parameter, which controls the rate of weight decrease, is set by those skilled in the art based on experience; regions with high weights correspond to frequency ranges where group delay measurements are stable, and these regions will receive greater attention in subsequent analyses.
[0066] Within each time-frequency analysis window, the weighted correlation coefficient between the phase derivative spectrum and the group delay feature is calculated, denoted as the phase-group delay coupling response value. The coupling response value is a core indicator for quantifying the consistency between phase and group delay features, reflecting the cooperative pattern of the two features in the local time-frequency region. The calculation process is performed within each time frame of the short-time Fourier transform. First, the phase derivative spectrum vector and group delay feature vector corresponding to that time frame are extracted; the dimensions of these two vectors are equal to the number of frequency analysis points. Then, the weight vector corresponding to the time-frequency window in the group delay reliability weight map is introduced to weight the two feature vectors, highlighting the contribution of reliable frequency points and suppressing the influence of unreliable frequency points. Finally, the correlation coefficient of the weighted vector is calculated and used as the phase-group delay coupling response value. High correlation indicates that the two features consistently indicate a certain articulation event in that time-frequency region, usually corresponding to consonant plosives, fricatives, or vowel formants; low correlation may correspond to transition segments or silent segments. The calculation of the coupling response value fully utilizes the complementarity of the two features: the phase derivative spectrum is sensitive to time transients, and the group delay is sensitive to frequency structure; their coupling can more comprehensively capture the dynamic characteristics of articulation.
[0067] Based on the distribution of the coupled response values in the time-frequency plane, a phase-group delay joint feature map is constructed. The joint feature map is a visual representation of multimodal feature fusion, providing a unified time-frequency feature basis for subsequent pronunciation analysis. The construction process uses the coupled response value of each time-frequency point of the short-time Fourier transform as the pixel intensity of the spectrum. The time axis corresponds to the analysis frame sequence, and the frequency axis corresponds to the frequency analysis points, forming a two-dimensional matrix structure. The temporal resolution of the spectrum is set according to the frame shift of the analysis window to ensure that it can capture the rapid transient changes of consonants. The frequency resolution is set according to the frequency sampling interval of the Fourier transform to distinguish adjacent formants. To enhance the readability of the spectrum, the coupled response values are dynamically compressed, and details in weak response areas are highlighted through logarithmic or power transforms. At the same time, pseudo-color mapping is applied to map the response values to colors, with high response values displayed as warm colors (red, yellow) and low response values displayed as cool colors (blue, green). The final generated joint feature spectrum intuitively shows the time-frequency evolution pattern of phase-group delay features during the pronunciation process. Brightly colored areas with concentrated energy correspond to key pronunciation events, providing a clear time-frequency localization basis for phoneme boundary detection and detailed feature extraction.
[0068] In embodiments of the present invention, the detailed implementation steps for obtaining the pronunciation detail feature sequence include:
[0069] Energy accumulation regions are detected in the joint feature map to identify connected regions whose coupling response values exceed a preset response threshold, and these regions are denoted as feature response regions. Energy accumulation regions are used to detect local high-value regions in the joint feature map that are manifested as coupling response values. The detection process first sets a dynamic threshold based on the statistical distribution of the coupled response values, typically using the mean of all response values plus 1.5 times the standard deviation as the criterion. Then, binarization is applied to convert the joint feature map into a binary image, marking points with coupled response values greater than the threshold as 1 and those less than the threshold as 0. Next, morphological connectivity analysis is used to identify connected regions in the binary image based on the eight-neighborhood connectivity criterion. Each connected region represents a set of temporally and frequency-adjacent feature responses. To filter out small-scale responses caused by noise, morphological processing is performed on each connected region to remove excessively small fragment noise, retaining meaningful feature response areas. A minimum area threshold is set, and connected regions with areas smaller than this threshold are removed. The retained feature response areas correspond to important articulation events in speech, and their position, size, and shape contain rich articulation information. Position corresponds to the time and frequency of articulation, size reflects the duration and frequency range of articulation, and shape is related to the dynamic process of articulation. These feature response areas provide clear target regions for subsequent phoneme boundary detection and detail feature extraction.
[0070] Phoneme boundary positions are detected by combining the temporal boundary change rate of the feature response region with the transient energy rise slope of the amplitude spectrum features. Phoneme boundary detection is a key step in the extraction of articulation detail features, and accurate boundary detection ensures the relevance and precision of feature extraction. The detection process employs a multi-feature fusion strategy. First, a first-order difference calculation is performed on the temporal boundary of the feature response region to obtain the time series of the boundary change rate. The peak of the boundary change rate usually corresponds to the start or end moment of the phoneme. Simultaneously, energy integration is performed on the amplitude spectrum features over a wide frequency range (usually 100Hz-8kHz). By calculating the total energy value of each time frame, a transient energy envelope curve that changes with time is formed. Differentiation is performed on the transient energy envelope curve, and the energy rise slope at each moment is obtained by calculating the ratio of the energy difference between adjacent time frames to the time interval. Finally, the slope values at all moments are arranged in chronological order to form a time series of energy rise slopes. The peak of the time series of energy rise slopes corresponds to the transient change moment of rapid energy rise in speech, which is used to assist in the detection of phoneme boundary positions such as the start of consonants.
[0071] Peak detection was then performed on the two time series separately. The boundary change rate was detected using an adaptive threshold, with the threshold determined by adding 1.5 times the local standard deviation to the local mean. The energy rise slope was detected using a fixed threshold. The corresponding peak candidate time sets for each time series were obtained. Finally, the phoneme boundary positions were determined based on the two sets of peak candidate time sets. The determination condition was that there were paired times in the two candidate sets with a time difference less than a preset time tolerance threshold. This dual verification mechanism ensured that the boundary position had significant changes in both time-frequency features and significant energy transitions, greatly reducing the false detection rate. The confirmed phoneme boundary positions were finely adjusted by searching for extreme points of the coupled response values in the joint feature map within a small range (±5 milliseconds) near the boundary. The boundary was corrected to the time corresponding to the extreme point, further improving the boundary positioning accuracy to the single-frame level. This dual feature verification mechanism significantly improved the accuracy and robustness of boundary detection, effectively avoiding the false detection and missed detection problems of single feature methods.
[0072] Within the detected consonant phoneme intervals, the phase transition patterns and impulse response patterns of the group delay features in the high-frequency band of the joint feature map are extracted and denoted as consonant onset transient features. Consonant onset transient features are a key indicator for evaluating the quality of consonant pronunciation; different types of consonants have different transient characteristics. The extraction process first locates the consonant phoneme intervals based on phoneme boundary positions and phoneme annotation information. Then, phase transition patterns are extracted in the high-frequency band of the joint feature map (typically 2kHz-8kHz, representing the frequency range where consonant energy is concentrated). Phase transitions are characterized by calculating the local variance and peak amplitude of the phase derivative spectrum. Simultaneously, the impulse response patterns of the group delay features are extracted, and the impulse response is characterized by detecting the peaks and durations of the group delay curve. Specific feature descriptions were defined for different consonant types: plosives (e.g., / p / , / t / , / k / ) are characterized by short, strong phase abrupt changes and sharp group delay impulses, quantified by the amplitude of the phase abrupt change (the maximum phase difference between adjacent time frames in the phase derivative spectrum) and the half-width at half-maximum (WHM) of the impulse response; fricatives (e.g., / s / , / f / , / θ / ) are characterized by sustained high-frequency phase fluctuations and relatively wide group delay fluctuation periods, quantified by the standard deviation of the phase derivative spectrum and the main period of the group delay fluctuations; affricates (e.g., / tʃ / , / dʒ / ) exhibit a mixed transient pattern, including a phase abrupt change in the initial stage (corresponding to the plosive component) and subsequent group delay fluctuations (corresponding to the fricative component), capturing the complex characteristics of affricates through temporal segmentation analysis. These transient features accurately characterize the subtle differences in consonant articulation, providing richer information for consonant quality assessment than traditional energy features.
[0073] Within the detected vowel phoneme intervals, formant frequency trajectories are extracted from the amplitude spectrum features, and the phase stability index of the corresponding positions of these formant frequency trajectories in the joint feature map is calculated, denoted as the vowel formant migration trajectory. Vowel features are the core content for evaluating pronunciation quality, and the position and stability of formants directly reflect the accuracy of the articulatory organ configuration. The extraction process first uses the Linear Predictive Coding (LPC) method to estimate the formant frequencies from the amplitude spectrum. By solving the polynomial roots corresponding to the LPC coefficients, the frequency positions of the first three formants F1, F2, and F3 are identified. Formant tracking is performed for each frame during the vowel duration, forming a trajectory curve of the formant frequency changing over time. Then, in the joint feature map, the time-frequency positions corresponding to the formant frequency trajectories are located, and the phase derivative spectra at these positions are extracted to calculate the phase stability index. Phase stability is characterized by the variance of the phase derivative spectrum in the time dimension, and the formula is: In the formula, This represents the phase stability index at the i-th resonance peak. Indicates at time and frequency points The phase value at (t is time 1) (This indicates the phase at that time-frequency position). The frequency of the i-th formant at time t is represented; T represents the duration of the vowel phoneme interval. The variance represents the average time over the entire vowel duration T (the overline indicates averaging over time t). A smaller variance indicates a more stable phase characteristic of the formant position, reflecting a strong ability to maintain steady-state vowel pronunciation; a larger variance indicates drastic phase characteristic fluctuations, potentially indicating inaccurate formant positions or unstable vowel pronunciation. The vowel formant migration trajectory integrates the frequency information (obtained from the amplitude spectrum) and stability information (obtained from the phase characteristics) of the formants, providing a multi-dimensional characteristic description for assessing vowel quality.
[0074] In the transition segment between adjacent phonemes, the continuity measure of the phase spectrum feature and the smoothness measure of the group delay feature are calculated, denoted as the phase continuity feature of the phoneme transition segment. The transition segment feature reflects the coordination and fluency of the articulation and is an important clue to distinguish natural articulation from mechanical splicing. The calculation process first locates the time range of the phoneme transition segment, usually the middle region between the boundaries of two phonemes. Then, within this interval, the continuity measure of the phase spectrum feature is calculated by calculating the energy of the second derivative of the phase spectrum (i.e., phase acceleration). The second derivative reflects the smoothness of the phase change; low energy indicates a smooth and continuous phase change, while high energy indicates a sharp abrupt change in the phase change. At the same time, the smoothness measure of the group delay feature is calculated by calculating the average of the local variance of the group delay within the transition segment (sliding window size of 3-5 frames). Small variance indicates a smooth group delay change, while large variance indicates a drastic fluctuation in the group delay. Natural phoneme transitions are characterized by low phase second derivative energy and low group delay variance, reflecting the coordinated movement of the articulatory organs; while unnatural transitions (such as phoneme breaks caused by native language habits) are characterized by high-energy phase jumps and high-variance group delay fluctuations; these transition characteristics provide quantitative indicators for assessing the coherence of pronunciation.
[0075] The speech signal is segmented into a sequence of phoneme units based on the transient features of consonant initiation, the migration trajectory of vowel formants, and the phase continuity features of phoneme transition segments, arranged chronologically to construct a sequence of articulation detail features. This feature sequence construction is the final step in integrating multiple feature types, organizing the scattered phoneme-level features into a structured temporal representation. The construction process, based on phoneme boundary detection results, divides the speech signal into a sequence of phoneme units. Each unit includes an initiation time, an end time, a phoneme type label, and corresponding detail features. For consonant units, the transient features of consonant initiation are associated (phase jump amplitude, group delay fluctuation period, or mixing mode parameters). For vowel units, the migration trajectory of vowel formants is associated (frequency trajectories of F1, F2, and F3 and phase stability indices). For adjacent phonemes, the phase continuity features of the phoneme transition segments are associated (second derivative energy and group delay variance). The resulting sequence of articulation detail features is a time-ordered set of feature vectors, each corresponding to a phoneme unit or transition segment. The vector dimension is dynamically determined according to the phoneme type, comprehensively recording the fine-grained features of the articulation process and providing a data foundation for subsequent comparative analysis with standard templates.
[0076] In embodiments of the present invention, the detailed implementation steps for constructing the multidimensional representation model include:
[0077] Standard pronunciation templates corresponding to the speech signal are selected from a pre-built standard pronunciation template library. These templates contain standard pronunciation detail feature sequences. The selection of standard templates serves as the reference benchmark for deviation analysis, directly impacting the accuracy and relevance of the evaluation. The selection process begins by retrieving standard pronunciation samples of the same text from the standard pronunciation templates based on the text content of the speech signal (obtained through speech recognition or pre-annotation). The standard pronunciation templates contain high-quality pronunciation samples recorded by native speakers, which undergo the same feature extraction process as the user samples to generate standard pronunciation detail feature sequences. To accommodate different pronunciation standard requirements, the template library includes various accent variants (e.g., American English, British English) and various speech speed styles (slow, normal, fast), allowing for the selection of appropriate template types based on the application scenario. For a single text, multiple standard samples may exist. By calculating the coarse feature similarity between the user sample and each candidate template (e.g., global features such as phoneme count, total duration, and average fundamental frequency), the best-matching template is selected as the reference. The feature sequence structure of the standard template is consistent with the user sample, containing phoneme-level detail features, providing a benchmark for subsequent phoneme-by-phoneme comparisons.
[0078] Dynamic time warping (DTW) is used to align the user's feature sequence with the feature sequence of the standard pronunciation template, obtaining a time mapping function. DTW is a classic algorithm for aligning temporal data, capable of handling differences in speech rate and local time scaling. The alignment process treats the user's feature sequence and the standard template's feature sequence as two time series, using a dynamic programming algorithm to find the time mapping path that minimizes the cumulative distance between the two sequences. Euclidean or cosine distance is used as the distance metric to calculate the difference in feature vectors at corresponding time points. To constrain the rationality of the mapping path, global path constraints such as Sakoe-Chiba bands or Itakura parallelograms are applied to limit the magnitude of time distortion and avoid unreasonable jump alignment. After alignment, the time mapping function is obtained, which describes the time scaling relationship between the user's pronunciation and the standard pronunciation, revealing whether the user pronounces each phoneme too fast or too slow. The advantage of the DTW algorithm is its ability to handle uneven speech rates; even if the user's overall speech rate differs significantly from the standard template, it can still find accurate phoneme-level correspondences.
[0079] Based on the time mapping function, the temporal deviation component of each phoneme unit in the articulation detail feature sequence is calculated. Temporal deviation reflects the deviation of pronunciation in the time dimension and is the most intuitive indicator of pronunciation problems. The calculation process first determines the correspondence between the user phoneme and the standard template phoneme for each phoneme unit according to the time mapping function. Then, three types of temporal deviations are calculated: phoneme duration deviation (representing the duration of the user phoneme minus the duration of the corresponding phoneme in the standard template), phoneme start time deviation (quantifying the advance or delay of phoneme appearance by comparing the start time of the user phoneme with the expected start time of the corresponding phoneme in the standard template (calculated based on the cumulative duration of preceding phonemes), and the degree of time distortion within the phoneme, characterized by calculating the standard deviation of the derivative of the time mapping function within that phoneme interval. The formula for representing the degree of time distortion is: In the formula, Indicates the degree of time distortion. The derivative of the time mapping function M; The mean of the derivative within the corresponding phoneme interval is represented by N, which represents the number of sampling points within the phoneme interval. A large standard deviation of the derivative indicates that the pronunciation speed within the phoneme is uneven, with unnatural acceleration or deceleration. A small standard deviation of the derivative indicates that the time flow within the phoneme is uniform and the pronunciation is stable. These three types of time-domain deviation components quantify the pronunciation problems in the time dimension from different perspectives, providing multi-dimensional indicators for time-domain scoring.
[0080] After dynamic time warping alignment is completed, the frequency domain deviation components between the corresponding features of the consonant's initial transient characteristics and the standard pronunciation template are calculated on the aligned time axis. The calculation process is performed on the unified time axis after alignment by the DTW algorithm to ensure that the corresponding consonant phonemes are being compared. For different types of consonants, different frequency domain deviations are calculated, including the frequency center offset of plosives (calculated by comparing the high-frequency energy distribution centers of the user and the standard template, the offset reflects the deviation of the articulation point), the energy distribution deviation of fricatives (calculated by calculating the KL divergence of the spectral energy distribution of the user and the standard template, a large divergence value indicates a significant difference in the spectral morphology of the fricative), and the consonant bandwidth difference (calculated by comparing the bandwidth of the main energy concentration frequency bands of the user and the standard template, the bandwidth difference reflects the accuracy and control of pronunciation). These frequency domain deviation components accurately quantify the problems of consonant pronunciation in the frequency dimension, providing an acoustic basis for consonant quality scoring.
[0081] On the aligned timeline, the phase domain deviation component between the vowel formant migration trajectory and the corresponding trajectory of the standard pronunciation template is calculated. The phase domain deviation is used to assess subtle differences in pronunciation by utilizing phase information. The calculation process is also performed on the aligned timeline. For each vowel phoneme, the phase characteristics of the user and the standard template are compared, and three types of phase domain deviations are calculated, including the phase stability deviation of the formant position (by comparing the phase stability indices of the user and the standard template at the formant frequency and calculating the difference between the two; a positive difference indicates that the user's formant phase stability is lower than the standard, which may indicate that the vowel pronunciation is unstable), the phase continuity deviation of the vowel transition segment (by comparing the phase second derivative energy of the vowel's beginning and end transition segments and calculating the difference between the user and the standard template; a large difference indicates that the transition is not as smooth as the standard), and the phase dispersion deviation of the vowel core segment (steady-state part). By comparing the dispersion of the phase characteristics of the middle steady-state segment of the vowel in the frequency dimension, the dispersion difference between the user and the standard template is calculated; high dispersion may indicate that the formant structure is unclear. Phase domain deviation provides microscopic information that traditional frequency domain features cannot capture, and it has a high degree of discriminative power for subtle differences in pronunciation among advanced users.
[0082] A multidimensional representation model is constructed based on the temporal, frequency, and phase domain deviation components. This model is represented in tensor form, a highly efficient structure for organizing multidimensional deviation data, facilitating subsequent multidimensional analysis and machine learning processing. The construction process first determines the three dimensions of the tensor representation: the first dimension corresponds to the phoneme sequence index, with a size equal to the number of phoneme units in the speech; the second dimension corresponds to the deviation type, including temporal, frequency, and phase domain deviations; and the third dimension corresponds to the statistical characteristic parameters of the deviation, including mean, standard deviation, maximum, and minimum values. The resulting multidimensional representation model uses a three-dimensional tensor representation, where each element quantifies the value of a specific phoneme, deviation type, and statistical characteristic. This structured representation preserves phoneme-level details while facilitating global statistical analysis, such as extracting the distribution of specific deviation types through tensor slicing and calculating the overall deviation level through tensor aggregation, providing rich input features for the adaptive evaluation model.
[0083] In embodiments of the present invention, the detailed steps for establishing an adaptive pronunciation quality assessment model include:
[0084] Based on the statistical distribution of deviation components of all phoneme units in the multidimensional representation model, and combined with a predefined level template, the overall pronunciation level estimate of the user is obtained. The calculation process first extracts the values of all deviation components from the multidimensional representation model, forming a high-dimensional feature vector (dimension: number of phonemes × number of deviation types × number of statistical parameters). Then, the feature vector is normalized to eliminate dimensional differences between different deviation types. Next, weighted clustering is applied to match the user's deviation distribution with the predefined level template. The level template is typically divided into three levels: beginner, intermediate, and advanced. Each level corresponds to a set of typical deviation distribution characteristics: beginner users exhibit large and scattered deviations in multiple types; intermediate users show significant improvement in some deviations but still have systematic problems; advanced users show relatively small overall deviations but still have deviations in specific details. The clustering process calculates the distance (such as Mahalanobis distance or cosine distance) between the user's deviation distribution and the center of each level, classifying the user to the nearest level center. This level is the overall pronunciation level estimate; this estimate provides the basis for subsequent dynamic benchmark selection.
[0085] A dynamic assessment benchmark is selected from the standard pronunciation template library to match the level template of the corresponding overall pronunciation level estimate. In addition to containing perfect standard pronunciation, the standard pronunciation template library also contains a library of tolerance thresholds for users of different levels. These thresholds are obtained through statistical analysis of a large amount of user data and reflect the typical performance range of users at each level. The selection process retrieves the threshold parameter set of the corresponding level based on the overall pronunciation level estimate, including the acceptable range of each deviation type and the parameters of the scoring mapping function. The selection of the dynamic assessment benchmark ensures the fairness of the assessment and avoids unreasonable situations caused by using a single standard to judge all users.
[0086] Based on the prosodic rhythm parameters in the cepstral domain features, the native language migration interference pattern of the speech signal is extracted. Native language migration interference is a common phenomenon among non-native speakers, who unconsciously bring their native language pronunciation habits into the target language. The extraction process first extracts the fundamental frequency profile from the cepstral domain features. By performing piecewise linear fitting on the fundamental frequency profile, the morphological features of the intonation curve are identified and extracted, including the slope distribution of the curve segments and the location distribution of the curve inflection points. Then, by segmenting the speech signal into syllables, the duration of each syllable is calculated, and a statistical distribution histogram of syllable duration is constructed. This histogram can be used to reflect the user's syllable duration preference pattern. Next, based on the energy of the amplitude spectrum features, the positions of stressed syllables are identified and marked to obtain the corresponding syllables in the speech signal. The text describes the distribution patterns of syllable positions within sentences. It then describes the vectorization of the acquired intonation curve morphology features, syllable duration distribution, and stressed syllable position distribution patterns, and performs similarity matching with a pre-defined native language prosodic feature database. This database contains typical prosodic patterns from users with various native language backgrounds (such as Chinese, Japanese, Korean, and Spanish). Cosine similarity or Mahalanobis distance is used to measure similarity and identify users' native language backgrounds and typical transfer patterns. The identification of native language transfer interference patterns provides a basis for distinguishing between systematic and random biases.
[0087] Based on the mother tongue transfer interference pattern, this study identifies systematic and random deviations in pronunciation caused by mother tongue habits. Distinguishing between systematic and random deviations is crucial for fair assessment. Systematic deviations are deep-seated habits that users find difficult to overcome at the current stage, while random deviations are caused by negligence or instability. The identification process is based on cross-phoneme consistency analysis of deviations. First, based on the mother tongue transfer interference pattern, it predicts the probability and direction of phoneme types exhibited by users in systematic deviations. For example, Chinese users may systematically shorten the duration of unstressed syllables, weaken or omit consonant endings, etc. Then, in a multidimensional representation model, it checks whether these predicted phoneme types and deviation directions appear consistently in multiple related phoneme units. If the proportion of a certain phoneme type's deviation direction appearing consistently in all related phonemes is greater than a preset overlap threshold, it is determined to be a systematic deviation. If it shows an independent distribution in different phoneme units without obvious patterns, it is determined to be a random deviation. This pattern recognition-based distinction method is more scientific and reliable than simple threshold judgment.
[0088] Based on a dynamic evaluation benchmark, different evaluation weights are assigned to systematic and random deviations to construct an adaptive pronunciation quality evaluation model. The core idea of the user-adaptive pronunciation quality evaluation model is to dynamically adjust the evaluation criteria according to the user's level and the nature of the deviation. The construction process first defines the evaluation weight function. For systematic deviations, the weight function increases with the overall pronunciation level, with lower weights at beginner levels and higher weights at advanced levels, reflecting the gradual requirements for overcoming native language transfer during the learning process. The corresponding mathematical formula for the weight function is:
[0089] In the formula, and These represent the corresponding weighting coefficients for systematic bias and random bias, respectively. This represents the pre-set benchmark weight for systematic deviations; Indicates the adjustment factor (in this application) L represents the user's current skill level (L=1, 2, 3, corresponding to beginner, intermediate and advanced levels respectively). This represents the intermediate level value (i.e., L=2). This linear adjustment mechanism enables a smooth transition of the evaluation criteria, avoids abrupt changes in scores during level transitions, and ensures the continuity and fairness of the evaluation. The final adaptive evaluation model includes dynamic benchmark thresholds and differentiated weight configurations, providing a personalized evaluation framework for subsequent comprehensive score calculations.
[0090] In embodiments of the present invention, the detailed steps for generating a comprehensive pronunciation quality assessment result include:
[0091] The deviation components of each phoneme unit in the multidimensional representation model are normalized according to the adaptive pronunciation quality assessment model to obtain the normalized deviation value. The processing involves normalizing each deviation type (sub-types in the time domain, frequency domain, and phase domain) based on the tolerance threshold for that level in the adaptive assessment model. The normalization formula uses relative deviation calculation, dividing the actual deviation value by the corresponding tolerance threshold to obtain the normalized deviation value. Specifically, when the actual deviation exceeds the tolerance threshold, the normalized deviation value is greater than 1 and is recorded as 1; if it is within the tolerance range, the normalized value is less than 1. The advantage of this normalization method is that it automatically adapts to the user's level; the same absolute deviation will yield different normalized values at different levels, reflecting a dynamic standard. Simultaneously, the same normalization logic is applied to all deviation types, ensuring the comparability of deviations across different dimensions. The normalized deviation value provides a unified input range for subsequent nonlinear mapping, simplifying the design of the scoring function.
[0092] Based on the normalized deviation value, a basic quality score is calculated for each phoneme unit; the basic quality score is used to convert the normalized deviation value into an intuitive score. The calculation process uses a nonlinear mapping function to convert the normalized deviation value into a score. The nonlinear mapping function results in a gradual decrease in score when the deviation is small, and a rapid decrease when the deviation is large. In this application, an exponential decay function is used to characterize this, and the corresponding formula is:
[0093] In the formula, Indicates the basic quality score. To control the rate of score decrease using the attenuation coefficient, The degree of nonlinearity is controlled by shape parameters; wherein, in this application ; ; This represents the weighted average of the normalized bias values;
[0094] For each phoneme unit, base scores are calculated separately for consonants and vowels. The consonant score focuses on transient characteristic deviations in the frequency and phase domains, while the vowel score focuses on formant deviations in the frequency domain and stability deviations in the phase domain. These base quality scores provide fine-grained phoneme-level evaluation results for subsequent multi-level fusion.
[0095] Based on the phase continuity characteristics of phoneme transition segments, a prosodic coherence score is calculated between adjacent phoneme units. The prosodic coherence score assesses the fluency and naturalness of pronunciation, focusing on the quality of connection between phonemes. The calculation process is based on the previously extracted phase continuity characteristics of phoneme transition segments, including a continuity measure of the phase spectrum and a smoothness measure of group delay. First, the deviation between the phase continuity characteristics and the corresponding characteristics of the standard pronunciation template is calculated. Then, combined with the continuity evaluation of the time mapping function, normalized deviation values corresponding to phase domain continuity and time domain continuity are obtained respectively. Finally, a weighted average is used to obtain the coherence score corresponding to the phoneme transition segment. Furthermore, the coherence scores of the phoneme transition segments contained within each syllable are averaged to obtain the prosodic coherence score for the corresponding syllable.
[0096] Based on the prosodic rhythm parameters in the cepstral domain features, the pause distribution, speech rate fluctuations, and stress rhythms of the speech signal are extracted to calculate the overall fluency score. The calculation process first analyzes the pause distribution, detects silent segments in the speech, and evaluates the rationality of pause positions (whether pauses occur at grammatical boundaries) and the appropriateness of pause duration (pauses that are too long or too short affect fluency). Then, speech rate fluctuations are analyzed, calculating the average pronunciation rate of syllables or phonemes, as well as the standard deviation of the rate; a stable speech rate reflects good control. Further analysis of stress rhythms is performed, evaluating the intensity contrast, duration contrast, and positional accuracy of stressed syllables. English is a stress-rhythmic language, and correct stress patterns are an important component of fluency. The results of the three dimensions are weighted and fused to obtain the overall fluency score. The overall fluency score provides a macro-level quality assessment, complementing the micro-level phoneme score.
[0097] The basic quality score, prosodic coherence score, and overall fluency score are weighted and fused at multiple levels to obtain a comprehensive pronunciation quality assessment result. The comprehensive pronunciation quality assessment result is the final output overall score, which integrates assessment information at the phoneme, syllable, and sentence levels. The fusion process employs a layered weighting strategy. First, at the phoneme level, the basic quality scores of all phoneme units are weighted and averaged according to duration to obtain a phoneme-level comprehensive score; duration weighting ensures that longer phonemes contribute more to the total score. Then, at the syllable level, the phoneme-level comprehensive score is weighted and fused with the prosodic coherence score to obtain a syllable-level comprehensive score, where the weight of the phoneme-level comprehensive score in the weighted fusion process must be greater than that of the prosodic coherence score. Finally, at the sentence level, the syllable-level comprehensive score is further weighted and fused with the overall fluency score to obtain the final comprehensive score. The weight of the syllable-level comprehensive score must be greater than that of the overall fluency score; the fusion weight of each level is adaptively adjusted according to the estimated value of the overall pronunciation level. Beginner users pay more attention to the accuracy of basic phonemes, so the weight of the phoneme level is higher; advanced users pay more attention to prosody and fluency, so the weight of the syllable-level and sentence-level scores is increased accordingly. Multi-level weighted fusion realizes a comprehensive evaluation from micro to macro, ensuring that the comprehensive score reflects both the accuracy of details and the overall performance, providing users with a scientific and comprehensive evaluation of pronunciation quality.
[0098] Next, the comprehensive assessment results are output and a visual diagnostic report is generated to help users intuitively understand their personal pronunciation problems and obtain targeted improvement suggestions.
[0099] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0100] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0101] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for evaluating the pronunciation quality of oral English based on multi-modal speech feature analysis, characterized in that, The method comprises the following steps: obtaining a speech signal of a target user speaking English, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features, including amplitude spectrum features, phase spectrum features, group delay features, and cepstrum domain features; constructing a joint feature map according to the time-frequency correspondence relationship between the phase spectrum features and the group delay features; extracting features according to the time sequence evolution mode of the energy accumulation area in the joint feature map, and combining the transient change characteristics of the amplitude spectrum features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, constructing a multi-dimensional representation model based on the multi-scale matching result, calculating the fine-grained quality score of each phoneme unit in the multi-dimensional representation model combined with the rhythm rhythm parameters in the cepstrum domain features, and fusing the rhythm coherence score and the overall fluency score to generate a comprehensive pronunciation quality evaluation result and a visual diagnostic report of pronunciation deviation; The process of constructing a joint feature map comprises: performing phase unwrapping processing on the phase spectrum features to obtain a continuous phase spectrum, performing difference operation on the continuous phase spectrum in the frequency dimension to obtain a phase derivative spectrum, and performing median filtering processing on the phase derivative spectrum to obtain a group delay estimation value; performing deviation calculation on the group delay estimation value and the actual measurement value of the group delay features, and constructing a group delay reliability weight map based on the deviation calculation result; calculating the weighted correlation coefficient of the phase derivative spectrum and the group delay features in the preset time-frequency analysis window based on the group delay reliability weight map, denoted as the phase-group delay coupling response value; constructing a phase-group delay joint feature map according to the obtained phase-group delay coupling response value; The process of obtaining a continuous phase spectrum comprises: performing frequency point by frequency point scanning on the phase spectrum features, identifying and marking frequency points with phase mutations as potential jump points; and obtaining local frequency estimation values corresponding to multiple frequency points adjacent to the potential jump points; predicting the expected phase value at the corresponding potential jump point according to the local frequency estimation value, and performing residual calculation on the expected phase value and the actual phase value to obtain a phase residual; determining whether the corresponding potential jump point is a real phase wrap point based on the phase residual, if it is not a phase wrap point, no other operation is performed; if it is a phase wrap point, phase compensation is performed on the phase wrap point to obtain a continuous phase spectrum.
2. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 1, characterized in that, The process of obtaining a pronunciation detail feature sequence comprises: detecting the energy accumulation area of the joint feature map to identify the connected area where the phase-group delay coupling response value exceeds the preset response threshold as a feature response area; obtaining the time sequence boundary change rate in the feature response area, and combining the transient energy of the amplitude spectrum features to determine the phoneme boundary position, and positioning the consonant phoneme interval and the vowel phoneme interval according to the phoneme boundary position; in the positioned consonant phoneme interval, extract the phase jump mode of the joint feature map in the high frequency band and the impulse response mode of the group delay features, and denote them as consonant initial transient features, the consonant initial transient features include the phase mutation amplitude of plosive, the group delay fluctuation period of fricative, and the mixed transient mode of affricate. Within the located vowel phoneme interval, a formant frequency trajectory in the amplitude spectrum feature is extracted, and a phase stability index of a corresponding position in a joint feature map is calculated, denoted as a vowel formant migration trajectory Based on the phoneme boundary position, a transition section between adjacent phonemes is obtained, and a continuity measure of the phase spectrum feature and a smoothness measure of the group delay feature in the corresponding transition section are calculated; they are taken as phase continuity features; the continuity measure is obtained by calculating the energy of the second derivative of the phase spectrum feature, and the smoothness measure is obtained by calculating the local variance of the group delay feature; The consonant onset transient feature, the vowel formant migration trajectory and the phase continuity feature are arranged in time sequence to construct the pronunciation detail feature sequence.
3. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 2, characterized in that, The determination mode of the phoneme boundary position is that the time difference between the peak time of the boundary change rate and the peak time of the energy rising slope is less than a preset time tolerance threshold.
4. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 2, characterized in that, The phoneme boundary position acquisition process includes: First-order difference calculation is performed on the time sequence boundary of the feature response section to obtain a time sequence of boundary change rates; and adaptive threshold peak value detection is performed on the time sequence of the boundary change rates to obtain a peak candidate time set of the boundary change rates; The amplitude spectrum feature is energy integrated in a wide frequency range to obtain a transient energy envelope curve, and differential operation is performed on the transient energy envelope curve to obtain energy rising slopes at different times, which are arranged in time sequence to form a time sequence of energy rising slopes; fixed threshold peak value detection is performed on the time sequence of the energy rising slopes to obtain a peak candidate time set of the energy rising slopes; The time matching degree of the peak candidate time set of the boundary change rates and the peak candidate time set of the energy rising slopes is calculated, the phoneme boundary position is confirmed based on the matching degree, and the confirmed phoneme boundary position is finely adjusted.
5. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 2, characterized in that, The construction process of the multi-dimensional representation model includes: A standard pronunciation template corresponding to the obtained speech signal is selected from a preset standard pronunciation template library, and the standard pronunciation template includes a standard pronunciation detail feature sequence; Dynamic time warping alignment is performed on the pronunciation detail feature sequence and the pronunciation detail feature sequence of the standard pronunciation template to obtain a time mapping function, and a time domain deviation component of each phoneme unit in the pronunciation detail feature sequence is calculated based on the time mapping function; After time warping alignment is completed, deviation component calculation is performed on the consonant onset transient feature and the vowel formant migration trajectory and the standard pronunciation template respectively to obtain frequency domain deviation components and phase domain deviation components; the frequency domain deviation components include burst frequency center offset, friction sound energy distribution deviation and consonant bandwidth difference value; the phase domain deviation components include phase stability deviation of formant position, phase continuity deviation of vowel transition section and phase dispersion deviation of vowel core section; The multi-dimensional representation model is constructed according to the time domain deviation component, the frequency domain deviation component and the phase domain deviation component.
6. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 5, characterized in that, The process of generating a comprehensive pronunciation quality evaluation result includes: An adaptive pronunciation quality evaluation model is established based on the multi-dimensional representation model and in combination with the prosodic rhythm parameters in the cepstrum domain feature; According to the adaptive pronunciation quality evaluation model, the deviation components of each phoneme unit in the multi-dimensional representation model are normalized to obtain normalized deviation values, and the basis quality scores corresponding to each phoneme unit are calculated based on the normalized deviation values; According to the phase continuity feature, the prosodic coherence scores between adjacent phoneme units are calculated, and the pause distribution, speech rate fluctuation, and stress rhythm of the speech signal are extracted according to the prosodic rhythm parameters in the cepstrum domain feature; and the overall fluency score is calculated; The basis quality scores, prosodic coherence scores, and overall fluency scores are multi-level weighted and fused to obtain a comprehensive pronunciation quality evaluation result; The obtained comprehensive pronunciation quality evaluation result is output, and a visual diagnostic report is generated.
7. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 6, characterized in that, The process of establishing the adaptive pronunciation quality evaluation model includes: According to the statistical distribution of the deviation components of all phoneme units in the multi-dimensional representation model, and in combination with the pre-defined level grade template, an overall pronunciation level estimation value of the user is obtained, which is obtained by weighted clustering analysis on the deviation components; A dynamic evaluation benchmark that matches the level grade template to which the overall pronunciation level estimation value belongs is selected from a standard pronunciation template library, and the dynamic evaluation benchmark includes differentiated fault tolerance thresholds for different level learners; According to the prosodic rhythm parameters in the cepstrum domain feature, the native language transfer interference mode of the speech signal is extracted, which includes syllable duration distribution preference, stress position shift rule, and intonation curve shape feature; According to the native language transfer interference mode, the systematic deviation and random deviation caused by native language habits in the pronunciation deviation are identified; Based on the dynamic evaluation benchmark, different evaluation weights are set for the systematic deviation and random deviation to construct an adaptive pronunciation quality evaluation model.
8. The method for assessing the quality of oral English pronunciation based on multi-modal speech feature analysis according to claim 7, characterized in that, The process of extracting the native language transfer interference mode includes: The cepstrum domain feature is subjected to fundamental frequency contour extraction to obtain a pitch curve of the speech signal, and the pitch curve is subjected to piecewise linear fitting to obtain an intonation curve shape feature, which includes the slope distribution of the curve segment and the position distribution of the curve inflection point; The speech signal is subjected to syllable segmentation, the duration of each syllable is calculated, and a statistical distribution histogram of the syllable duration is constructed, denoted as the syllable duration distribution preference; The preferred syllable duration mode of the learner is identified by clustering analysis on the syllable duration distribution preference, and compared with the standard syllable duration distribution of the target language to calculate the distribution difference degree; According to the energy envelope curve of the amplitude spectrum feature, the stress syllable position in the speech signal is identified, and the position distribution mode of the stress syllable in the sentence is calculated, denoted as the stress position shift rule; The prosody curve shape feature, the syllable duration distribution preference, and the accent position offset rule are vectorized as feature vectors, and similarity matching is performed with a preset native language rhythm feature library to identify a native language migration interference mode; the native language rhythm feature library contains typical rhythm feature templates of learners with various native language backgrounds.
Citation Information
Patent Citations
Mobile English learning system based on interaction
CN120278862A
Power transformer fault diagnosis method and system based on voiceprint recognition
CN120766712A