Mandarin pronunciation comparison training method based on real-time voice recognition

Through real-time speech recognition and tongue position tracking technology, pronunciation feature vectors are constructed and matched with standard templates, which solves the problem of inaccurate tongue position control in Mandarin pronunciation training, realizes real-time pronunciation deviation detection and adjustment prompts, and improves the accuracy and efficiency of Mandarin learning.

CN120748445APending Publication Date: 2025-10-03HANGZHOU HUANCHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511098703.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing Mandarin pronunciation training systems are deficient in tongue position tracking and tongue position adjustment guidance, and are unable to effectively distinguish between retroflex and flat tongue sounds. This makes it difficult for trainees to accurately grasp the pronunciation differences between the initial consonants "zh", "ch", "sh" and "z", "c", "s", and lacks real-time and accurate tongue position comparison information.

Method used

Through real-time speech recognition technology, the voice data of the trainee is obtained, and the changes in tongue position are monitored using acoustic feature extraction algorithms and tongue position trajectory tracking algorithms. The pronunciation feature vector is constructed and matched with the standard feature template, and the tongue position comparison results and adjustment prompts are presented in real time.

Benefits of technology

It improves the accuracy and real-time performance of pronunciation deviation detection, helping trainees to adjust tongue position movements in a timely manner at key pronunciation nodes, and improving the accuracy of Mandarin pronunciation and learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748445A_ABST
    Figure CN120748445A_ABST
Patent Text Reader

Abstract

The invention provides a mandarin pronunciation comparison training method based on real-time speech recognition, which relates to the technical field of speech recognition, and generates a pronunciation deviation value and a tongue position feedback control signal in real time by dynamically monitoring tongue position characteristic parameters at different time points in a target training time period and combining acoustic characteristic matching degree calculation. Thus, the accuracy and the real-time performance of pronunciation deviation detection can be improved, a trainer can visually know the dynamic change of the tongue position in the pronunciation process, the tongue position action can be adjusted in time at the key pronunciation node, and the accuracy and the learning efficiency of mandarin pronunciation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and more particularly to a Mandarin pronunciation comparison training method based on real-time speech recognition. Background Art

[0002] As China's official language, accurate pronunciation of Mandarin is a key component in its learning and promotion. With the advancement of phonetics, linguistics, and artificial intelligence technologies, pronunciation training methods based on speech recognition have gradually become an important technical tool in Mandarin learning. Traditional Mandarin pronunciation training relies primarily on teacher guidance and learner self-imitation, a method characterized by high subjectivity and low training efficiency. In recent years, the rapid development of speech recognition technology has enabled the development of automated pronunciation training systems. These systems collect user voice data and, in combination with acoustic feature analysis and speech comparison algorithms, can quantitatively assess learners' pronunciation and provide feedback. However, existing speech training systems often focus on assessing the overall matching of acoustic features, ignoring the crucial role of tongue position in language pronunciation. Tongue position is a crucial factor in determining Mandarin pronunciation accuracy, and is particularly important for distinguishing retroflex and flat consonants. The shortcomings of existing technologies in tongue position tracking and tongue position adjustment guidance have become a major bottleneck restricting the further development of Mandarin pronunciation training systems.

[0003] Existing Mandarin pronunciation training technology has shortcomings in the following areas: First, although some systems can identify pronunciation deviations in users through acoustic features, they lack dynamic tracking and analysis of tongue position details, making it difficult for trainees to intuitively understand the specific sources of pronunciation deviations. Second, existing technologies often use a single acoustic parameter as the evaluation basis, failing to combine tongue position trajectory with acoustic features to construct a multi-dimensional pronunciation feature vector. This single-dimensional analysis method reduces the accuracy of pronunciation deviation analysis. Furthermore, when generating feedback signals, current pronunciation training systems typically only provide simple voice prompts or scoring results, lacking specific guidance on key tongue position change points, making it difficult to meet trainees' needs for detailed pronunciation adjustments. In particular, in the pronunciation comparison training of retroflex and flat consonants, parameters such as the height of the tongue tip and the degree of tongue curl are crucial to pronunciation. However, existing systems fail to provide real-time, accurate tongue position comparison information, and are unable to effectively address the specific problems learners encounter during pronunciation training. Summary of the Invention

[0004] To address the above-mentioned technical problems, the present invention provides a Mandarin pronunciation comparison training method based on real-time speech recognition. This method can, to a certain extent, address the difficulty learners face in accurately distinguishing the pronunciation of the initial consonants "zh," "ch," and "sh" from those of "z," "c," and "s" during Mandarin pronunciation training, particularly the problem of pronunciation confusion caused by inaccurate tongue position control during continuous pronunciation.

[0005] According to one aspect of the present invention, a method for comparative training of Mandarin pronunciation based on real-time speech recognition is provided, which comprises:

[0006] Acquiring real-time speech data of a trainee, and converting the real-time speech data into a first acoustic feature sequence through an acoustic feature extraction algorithm;

[0007] Based on the first acoustic feature sequence, a tongue position trajectory tracking algorithm is used to dynamically monitor the tongue position feature parameters to obtain a tongue tip lift height curve and a tongue surface curl degree curve, and key points of tongue position changes during pronunciation are identified based on the curves;

[0008] According to the tongue position change key points, combined with the first acoustic feature sequence, a pronunciation feature vector is constructed, and the pronunciation feature vector is matched with a preset standard feature template of retroflex and flat tongue sounds to obtain a real-time pronunciation deviation value;

[0009] When the deviation value exceeds the preset range, the comparison result between the ideal tongue position curve and the actual tongue position curve is presented in real time through a visual display, and targeted tongue position adjustment prompts are output at key pronunciation nodes.

[0010] Furthermore, the first acoustic feature sequence includes tongue position feature parameters, airflow turbulence parameters and friction noise spectrum parameters;

[0011] The tongue position feature parameters are extracted by a linear prediction analysis method, including the first N order linear prediction coefficients;

[0012] The airflow turbulence parameters are extracted using zero-crossing rate and short-time energy analysis methods;

[0013] The friction noise spectrum parameters are calculated by fast Fourier transform.

[0014] Furthermore, obtaining the tongue tip lift height curve and the tongue surface curling degree curve includes:

[0015] Performing mapping transformation on the tongue position feature parameters in the first acoustic feature sequence to determine the vocal tract cross-sectional area distribution;

[0016] Mapping the vocal tract cross-sectional area distribution into a three-dimensional space coordinate system and calculating the tongue position space coordinates;

[0017] The tongue position spatial coordinates are projected and features are extracted to obtain the tongue tip lift height curve and the tongue surface curling degree curve.

[0018] Furthermore, the tongue tip lifting height curve and the tongue surface curling degree curve are shown as follows:

[0019]

[0020]

[0021] in,

[0022]

[0023]

[0024] in, is the height of the tongue tip at time t, is the y component of the tongue position coordinate after smoothing, is the time-varying weight function of tongue tip motion, is the degree of tongue curling at time t, is the number of sampling points on the tongue contour, is the vector from the mth sampling point to the m+1th sampling point, is the angle between adjacent vectors, is the time-varying weight function of tongue curl, is the rate of change of the height of the tongue tip, is the rate of change of tongue curl, is the time constant of the tongue tip movement, is the time constant of tongue curling.

[0025] Furthermore, when the key change point is identified, the features of the tongue tip lift height curve and the tongue surface curling degree curve are fused;

[0026] When the local change of the fused feature function exceeds the preset threshold and the second-order derivative is zero, the moment is marked as the tongue position change key point.

[0027] Furthermore, the pronunciation feature vector includes tongue position features, airflow turbulence features and friction noise spectrum features. The tongue position features are the height of the tongue tip and the degree of tongue curling at the certain moment.

[0028] Furthermore, the airflow turbulence feature is centered on the key point of tongue position change to construct a local analysis window, and the short-time zero-crossing rate and short-time energy value of the speech signal are calculated within the window to form an airflow turbulence feature sequence;

[0029] The friction noise spectrum feature is the energy distribution feature of each sub-band of the speech signal at the key point.

[0030] Furthermore, the calculation of the airflow turbulence characteristic is shown in the following formula:

[0031]

[0032] in, For the moment The characteristic value of airflow turbulence calculated at is the length of the local time window, which is used to determine the time range for calculating turbulence characteristics. is the displacement in the time window, ranging from -L / 2 to L / 2, is the zero-crossing rate function, which indicates the number of times the speech signal passes through the zero point in unit time. is the short-time energy function.

[0033] Furthermore, the matching degree calculation is performed by respectively calculating the tongue tip height, airflow turbulence and deviation of frequency band energy distribution, and dynamically adjusting the weight coefficients to obtain an initial matching degree score through weighted summation;

[0034] The initial matching score is subjected to time constraints to obtain the real-time pronunciation deviation value.

[0035] Furthermore, the real-time presentation of the comparison results of the ideal tongue position curve and the actual tongue position curve includes superimposing and displaying the ideal tongue position trajectory curve of the standard pronunciation, and the ideal tongue position trajectory curve is marked in translucent yellow; at the same time, superimposing and displaying the actual tongue position trajectory curve of the trainee, and the actual tongue position trajectory curve is marked in translucent blue, which is the tongue tip lifting height curve and the tongue surface curling degree curve; when there is overlap between the two trajectory curves, the overlapping area is displayed in green.

[0036] Compared to existing technologies, the Mandarin pronunciation comparison training method based on real-time speech recognition provided by the present invention dynamically monitors tongue position feature parameters at different time points within the target training period and, combined with acoustic feature matching calculations, generates pronunciation deviation values ​​and tongue position feedback control signals in real time. This improves the accuracy and real-time nature of pronunciation deviation detection, helping trainees intuitively understand the dynamic changes in tongue position during pronunciation, allowing them to adjust tongue position movements at key pronunciation nodes in a timely manner, improving the accuracy and learning efficiency of Mandarin pronunciation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0038] Figure 1 4 is a flowchart of a method for comparative Mandarin pronunciation training based on real-time speech recognition according to an embodiment of the present invention.

[0039] Figure 2 Schematic diagram of a comparison result between an ideal tongue position curve and an actual tongue position curve presented in real time based on real-time speech recognition according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein.

[0041] Figure 1 Flowchart of the Mandarin pronunciation comparison training method based on real-time speech recognition according to an embodiment of the present invention. Figure 1 As shown, in the Mandarin pronunciation comparison training method based on real-time speech recognition, the following steps are included:

[0042] S1: Acquire real-time speech data of a trainee, and convert the real-time speech data into a first acoustic feature sequence through an acoustic feature extraction algorithm, wherein the first acoustic feature sequence includes tongue position feature parameters, airflow turbulence parameters, and friction noise spectrum parameters;

[0043] Acquiring real-time speech data of a trainee and converting the real-time speech data into a first acoustic feature sequence through an acoustic feature extraction algorithm, specifically including:

[0044] The real-time speech data generated by the trainee when reading the training text is collected by a microphone device, wherein the real-time speech data is an audio signal with a sampling rate of 16kHz and a quantization accuracy of 16 bits; the real-time speech data is preprocessed, including pre-emphasis, framing and windowing, wherein the pre-emphasis coefficient is set to 0.97, the frame length is 25ms, the frame shift is 10ms, and the windowing is performed using a Hamming window;

[0045] The pre-processed real-time speech data is subjected to an acoustic feature extraction algorithm to extract acoustic features, including: extracting tongue position feature parameters using a linear prediction analysis method, wherein the tongue position feature parameters include the first N-order linear prediction coefficients, where N is 12; extracting airflow turbulence parameters using a zero-crossing rate and short-time energy analysis method, wherein the airflow turbulence parameters are used to characterize airflow characteristics during pronunciation; and calculating friction noise spectrum parameters using a fast Fourier transform, wherein the friction noise spectrum parameters are used to characterize the spectrum characteristics of consonant pronunciation.

[0046] The extracted tongue position feature parameters, airflow turbulence parameters and friction noise spectrum parameters are organized into a first acoustic feature sequence in chronological order. Each frame of the first acoustic feature sequence contains the above three types of parameters for subsequent tongue position trajectory tracking and pronunciation feature analysis; the first acoustic feature sequence is stored in matrix form, each row of the matrix corresponds to an analysis frame, and each column corresponds to a feature parameter.

[0047] S2: Based on the first acoustic feature sequence, dynamically monitor the tongue position feature parameters using a tongue position trajectory tracking algorithm to obtain a tongue tip lift height curve and a tongue surface curling degree curve, and identify key points of tongue position changes during pronunciation based on the curves;

[0048] First, the tongue position feature parameters in the first acoustic feature sequence are mapped and transformed. After obtaining the 12th-order linear prediction coefficient of each frame, the vocal tract area function is obtained by calculating the reflection coefficient, wherein the linear prediction coefficient of each frame is converted into the reflection coefficient by recursion. The intermediate coefficients in the recursive process need to maintain numerical stability. When the absolute value of the reflection coefficient is close to 1, normalization is used to avoid numerical overflow. If the reflection coefficients of adjacent vocal tract sections are determined, the area ratio of the adjacent vocal tract sections is calculated according to the acoustic transmission line theory to obtain a complete vocal tract area function, as shown in the following formula:

[0049]

[0050] in, is the area of ​​the m+1th channel cross section at the i-th frame, is the area of ​​the mth channel cross section at the i-th frame, is the reflection coefficient of the mth position at the i-th frame, and its value range is [-1,1].

[0051] Based on the obtained vocal tract area function, a vocal tract shape reconstruction model is constructed. After the distribution of the vocal tract cross-sectional area is determined, it is mapped to a three-dimensional spatial coordinate system, where the x-axis represents the front-to-back direction, the y-axis represents the up-down direction, and the z-axis represents the left-right direction. For each vocal tract cross section, the azimuth and longitudinal position parameters of the cross section are determined according to its relative position in the vocal tract. That is, the spatial coordinates of the tongue position at the i-th frame are set to P_i(x, y, z), and the calculation formula is:

[0052]

[0053] in, is the three-dimensional tongue position spatial coordinate at the i-th frame, is the coordinate transformation matrix, which is determined by three rotation angle parameters. is the rotation angle around the x-axis, is the rotation angle around the y-axis, is the rotation angle around the z-axis, is the azimuth angle of the mth sound channel section, is the longitudinal position of the mth sound channel section.

[0054] If the tongue position coordinates between the current frame and the previous frame differ significantly, measurement noise may be present, and the coordinates need to be adaptively smoothed. During this adaptive smoothing process, the smoothing factor is dynamically adjusted based on the severity of the coordinate changes between adjacent frames. The smoothing strength increases when the changes are severe, and decreases when the changes are gentle.

[0055] Furthermore, the reconstructed tongue position spatial coordinates are projected and feature extracted. After the three-dimensional tongue position coordinates are determined, they are projected onto the sagittal plane to extract two key parameters: the tongue tip lifting height and the tongue surface curling angle. The tongue tip lifting height is obtained by calculating the displacement of the tongue tip area in the y-axis direction, and a time-varying weight function is introduced to modulate it. When the tongue tip movement speed is large, the weight is increased to highlight the characteristics of the rapidly changing interval. The tongue surface curling angle is obtained by calculating the average value of the angles between adjacent tongue sites. A time-varying weight function is also introduced to increase the weight when the tongue surface shape changes significantly, as shown in the following formula:

[0056]

[0057]

[0058] in,

[0059]

[0060]

[0061] in, is the height of the tongue tip at time t, is the y component of the tongue position coordinate after smoothing, is the time-varying weight function of tongue tip motion, is the degree of tongue curling at time t, is the number of sampling points on the tongue contour, is the vector from the mth sampling point to the m+1th sampling point, is the angle between adjacent vectors, is the time-varying weight function of tongue curl, is the rate of change of the height of the tongue tip, is the rate of change of tongue curl, is the time constant of the tongue tip movement, is the time constant of tongue curling.

[0062] Subsequently, the obtained curve of tongue tip lifting height and tongue surface curling degree are subjected to feature analysis. When it is necessary to identify key change points, the features of the two curves are fused to construct a comprehensive feature function. The comprehensive feature function takes into account both position features and dynamic features, where the position features reflect the absolute position information of the tongue position, and the dynamic features reflect the change trend information of the tongue position. When the local change of the comprehensive feature function exceeds the preset threshold and its second-order derivative is zero, the moment is marked as a key point of tongue position change. The comprehensive feature function As shown in the following formula:

[0063]

[0064] When F(t) meets the following conditions, it is considered as a key point of tongue position change:

[0065]

[0066] in, is the weight coefficient of the height of the tongue tip, is the weight coefficient of the tongue curling degree, is the weight coefficient of the dynamic feature, is the second-order time derivative, representing the acceleration characteristics, is the time window length, is the detection threshold, is the second-order derivative of the comprehensive characteristic function.

[0067] Finally, the sequence of identified tongue position change key points is post-processed. Once the key point sequence is obtained, it needs to be time-domain aligned with the pre-calibrated pronunciation expert tongue position change template. The time-domain alignment process uses a dynamic programming strategy to iteratively calculate the cumulative distance matrix to find the optimal time correspondence. If the deviation between a key point and the template exceeds the tolerance range, it is marked as a potential pronunciation error point.

[0068] It is worth noting that although the above description outlines the general process of identifying the key points of tongue position changes during pronunciation, the specific implementation details may vary depending on the trainee. For example, when the trainee begins to pronounce the sound "weave", the system first captures the process of the tongue tip rapidly lifting from a calm position. At this time, the curve of the height of the tongue tip lifting shows a clear upward trend, and the curve of the degree of tongue curling also increases. Through the calculation of the comprehensive feature function, the system identifies the first key point of the tongue position change, which represents the starting position where the tongue tip begins to approach the hard palate. When the tongue tip moves to the highest point, the system captures the second key point, at which time the tongue surface shows a clear curling state. If the trainee's tongue position control is accurate, the position and timing characteristics of these two key points will be highly consistent with the standard template.

[0069] In contrast, when the trainee pronounces the sound "zi," the system captures a noticeably different tongue tip trajectory. The tongue tip lifts lower than in the sound "zhi," and the tongue curl curve remains within a smaller range. The system also identifies key points, but the characteristic parameters of these key points differ significantly from those for the sound "zhi."

[0070] It should be noted that when the trainee pronounces "zh", "ch", and "sh", the system can capture the entire process of tongue tip lifting and tongue curling in real time. The tongue tip lifting height curve can be used to monitor whether the tongue tip is actually lifted to the hard palate; the tongue curling degree curve can be used to determine whether the tongue has formed the necessary curling shape. In contrast, when pronouncing "z", "c", and "s", the system monitors whether the tongue tip is only lifted to the upper back of the teeth, while confirming whether the tongue remains flat. This real-time monitoring allows trainees to intuitively understand the fundamental difference in tongue position control between the two groups of initials.

[0071] S3: Constructing a pronunciation feature vector based on the tongue position change key point, in combination with the airflow turbulence parameter and the friction noise spectrum parameter, and calculating a matching degree between the pronunciation feature vector and a preset standard feature template of retroflex and flat tongue sounds to obtain a real-time pronunciation deviation value;

[0072] The airflow turbulence characteristics are extracted and analyzed within the neighborhood of each key point of tongue position change. When the pronunciation process reaches a certain key moment, the system takes this moment as the center and extends 25 milliseconds forward and backward to construct a local analysis window. Within this window, the short-time zero-crossing rate of the speech signal is first calculated to characterize the rapid change characteristics of the signal; at the same time, the short-time energy value is calculated to characterize the intensity change of the airflow. When the product of the zero-crossing rate and the short-time energy is greater than the preset threshold, it indicates that there is obvious airflow turbulence at this moment. The system will continuously track the time distribution and intensity changes of these turbulence characteristics to obtain a complete airflow turbulence characteristic sequence. Let t_k be the time position of the kth key point, then the airflow turbulence characteristic is calculated as follows:

[0073]

[0074] in,

[0075]

[0076]

[0077] in, For the moment The characteristic value of airflow turbulence calculated at is the length of the local time window, which is used to determine the time range for calculating turbulence characteristics. is the displacement in the time window, ranging from -L / 2 to L / 2, is the zero-crossing rate function, which indicates the number of times the speech signal passes through the zero point in unit time. is the short-time energy function, which indicates the energy of the speech signal in a short time window. is the input speech signal sequence, is the Hamming window function, which is used to reduce the spectrum leakage caused by signal truncation. is the length of the analysis frame, that is, the number of sampling points used in each calculation, is a sign function that outputs 1 when the input is positive, -1 when it is negative, and 0 when it is zero. |·| is an absolute value operator. is the kth key time point, is the offset position relative to the key time point, The previous sampling point is offset relative to the key time point.

[0078] At the same time, the speech signal at each key point is subjected to spectral analysis. First, the entire frequency band is divided into several sub-bands, where the low-frequency band (0-1000Hz) adopts a narrower frequency band interval, with one sub-band divided every 100Hz; the mid-frequency band (1000-3000Hz) adopts a frequency band interval of 200Hz; and the high-frequency band (above 3000Hz) adopts a frequency band interval of 500Hz. After the system obtains the spectrum of the speech signal, it calculates the energy distribution characteristics in each sub-band. If the energy value of a certain frequency band is significantly higher than that of the adjacent frequency band, the frequency band is marked as a characteristic frequency band. By analyzing the distribution pattern and energy ratio of these characteristic frequency bands, the pronunciation characteristics of different consonants can be effectively distinguished. That is, at each key point, the energy distribution of M frequency bands is extracted:

[0079]

[0080] in,

[0081]

[0082] in, For the moment Place The energy distribution characteristic value of the frequency band is is the frequency variable, which represents the frequency component of the signal, is the starting frequency of the mth frequency band, is the end frequency of the mth frequency band, is the complex spectrum obtained by performing short-time Fourier transform on the speech signal at time t_k, is the energy density of the spectrum, that is, the square of the modulus of the complex spectrum, is the weighting function of the mth frequency band, which is used to adjust the importance of different frequency components. is the center frequency of the mth frequency band, is the standard deviation of the mth frequency band, controlling the width of the frequency band, is the importance weight coefficient of the mth frequency band, is the natural exponential function, It is the square of the distance between the frequency point and the center frequency of the band.

[0083] After the system obtains the tongue position features, airflow turbulence features, and spectral features, it needs to organically integrate these features. First, each feature is normalized so that its numerical range is unified between 0 and 1. Then, the system will dynamically adjust the weight coefficient of each feature according to the noise level of the current pronunciation environment. If the environmental noise is large, the weight of the spectral feature will be reduced and the weight of the tongue position feature will be increased accordingly; if the speaker's pronunciation speed is detected to be abnormal, the weight of the timing feature will be adjusted accordingly.

[0084] Furthermore, when calculating the matching degree between the pronunciation feature vector and the preset standard feature templates of retroflex and flat tongue sounds, the system first retrieves the standard feature template corresponding to the target phoneme from the preset sound library. After the retrieval is completed, the system compares and analyzes the real-time pronunciation feature vector with the standard template in multiple dimensions: first, in the tongue position feature dimension, the tongue position matching degree is obtained by calculating the Euclidean distance between the height of the tongue tip and the degree of tongue curling. If the deviation of the tongue position feature exceeds the preset threshold, the system marks the moment as a key focus area; then, in the airflow feature dimension, the system compares the real-time pronunciation and the standard template in terms of similarity in airflow turbulence parameters. When it is found that there are obvious differences in the airflow characteristics, the system will further analyze the specific deviations in airflow intensity and duration; then, in the spectrum feature dimension, the system compares the real-time pronunciation and the standard template in each frequency band. Energy distribution difference. If the energy difference in a certain frequency band is significant, the frequency band will be marked as a characteristic frequency band. After completing the comparison of the features of each dimension, the system dynamically adjusts the weight coefficients of the features of each dimension according to the current environmental noise level and the characteristics of the speaker, and obtains the initial matching score by weighted summation. Then the system analyzes the dynamic changes of matching at multiple time scales, including local feature matching of 10-20 milliseconds, trend analysis of 50-100 milliseconds and overall coordination evaluation of 200-300 milliseconds. If a sudden change in matching is detected, the system will retrospectively analyze the feature changes before and after, and eliminate the influence of instantaneous fluctuations through smoothing. Finally, the system comprehensively considers the matching results of each time scale to generate a final matching score, which reflects both the pronunciation accuracy at the current moment and the dynamic characteristics evaluation results of the pronunciation process.

[0085] The specific process of obtaining the tongue position matching degree by calculating the Euclidean distance between the tongue tip lifting height and the tongue surface curling degree is to obtain the tongue tip lifting height parameter and the tongue surface curling degree parameter during the real-time pronunciation process. After obtaining these two parameters, the system will combine them into a two-dimensional feature point; at the same time, the standard tongue position parameter of the corresponding phoneme is extracted from the standard template library, which also constitutes a two-dimensional feature point; then the system calculates the Euclidean distance between the two two-dimensional feature points. If the calculated distance value is less than the preset first threshold value, it is considered that the tongue position at the current moment is highly consistent with the standard pronunciation, and the system will set the tongue position matching degree at that moment to a high matching degree. interval value; if the distance value falls between the first threshold and the second threshold, the tongue position matching degree at that moment is set to the medium matching degree interval value; when the distance value is greater than the second threshold, further analysis is performed on whether the deviation of the tongue tip lifting height is large or the deviation of the tongue surface curling degree is significant. If the deviation of the tongue tip lifting height is large, then focus on the continuity of the tongue tip movement trajectory. If the deviation of the tongue surface curling degree is significant, then focus on analyzing the coordination of tongue surface muscle control; when a sudden change in the tongue position parameters is detected, retrospectively analyze the parameter changes before and after the moment, and eliminate the influence of instantaneous fluctuations through smoothing processing; thereby obtaining an accurate tongue position matching evaluation result.

[0086] It should be noted that the process of determining the first threshold and the second threshold specifically includes: first, through large-scale experimental data collection, the system collects tongue position parameter data of multiple standard speakers when reading retroflex consonants and flat tongue consonants in different scenarios. Each speaker needs to repeat the reading multiple times to ensure the reliability of the data; then the collected data is statistically analyzed to calculate the mean and standard deviation of the tongue position parameters of the standard speaker group. After obtaining these statistical features, the system sets 0.8 times the standard deviation as the first threshold, which reflects the allowable range of normal pronunciation fluctuations; 1.5 times the standard deviation is set as the second threshold, which represents the judgment boundary of obvious pronunciation deviation; then the system will dynamically adjust these two benchmark thresholds according to the physiological characteristics of speakers of different age groups.

[0087] On the other hand, when comparing the similarity between the real-time pronunciation and the standard template in terms of airflow turbulence parameters, the real-time sequence and the standard sequence are time-aligned through the dynamic time warping algorithm. After the sequence alignment is completed, the correlation coefficient and mean square error of the two sequences at the corresponding time are calculated for judgment, and the comparison results of multiple consecutive frames are weighted averaged to obtain a stable airflow turbulence similarity evaluation value.

[0088] When comparing the energy distribution differences between real-time pronunciation and standard templates in each frequency band, the differences are evaluated using Euclidean distance and KL divergence. If the energy difference in a frequency band exceeds 30% of the average energy, it is marked as a critical frequency band, and its energy ratio with adjacent frequency bands and whether it constitutes a characteristic resonance peak are further analyzed, with a focus on the characteristic differences in the 2000-4000Hz frequency band. In the presence of noise interference, the system adjusts the frequency band weights according to the noise characteristics, and ultimately weights the difference results of each frequency band into a comprehensive spectral difference measure.

[0089] S4: When the deviation value exceeds the preset range, the comparison result between the ideal tongue position curve and the actual tongue position curve is presented in real time through a visual display, and targeted tongue position adjustment prompts are output at key pronunciation nodes.

[0090] A tongue position feedback control signal is established based on the real-time pronunciation deviation value. When the deviation value exceeds a preset range, a multi-dimensional feedback display system is established on the visual display interface. In the main view area of ​​the display interface, a three-dimensional vocal tract model is drawn using real-time dynamic rendering technology. The model includes the precise anatomical structure of pronunciation organs such as the mouth and tongue. The ideal tongue position trajectory curve of the standard pronunciation is superimposed on the model and marked with translucent yellow; at the same time, the actual tongue position trajectory curve of the trainee is displayed and marked with translucent blue. The overlapping area of ​​the two curves is displayed in green, which intuitively reflects the accuracy of tongue position control. A parameter monitoring panel is set on the right side of the interface to display the numerical changes of key parameters such as the height of the tongue tip lifting and the degree of tongue curling in real time.

[0091] Provide precise tongue position adjustment prompts at key pronunciation nodes. The system has pre-stored a library of diagnostic rules for various common pronunciation problems by pronunciation experts, including criteria for typical problems such as tongue position being too high, tongue position being too low, tongue curling being insufficient, tongue tip being too forward, and tongue tip being too backward. When a tongue position deviation is detected, the system automatically matches the most similar problem type and retrieves corresponding adjustment suggestions from the rule library. These suggestions are displayed in the prompt area of ​​the interface in concise text, such as "Please slightly lower the tongue tip position" and "The tongue surface needs to curl back appropriately." At the same time, the system will also mark the ideal tongue position adjustment direction and amplitude with arrows on the three-dimensional model. For example, when the key pronunciation node "zh" is detected, if the height of the tongue tip is lower than the standard value, the system prompts "Please raise the tongue tip to the back of the upper gum"; when the pronunciation node is "ch", if the tongue curl is not enough, the system prompts "Please retract the tongue and curl it moderately"; when the pronunciation node is "sh", if the airflow strength is weak, the system prompts "Please increase the airflow strength and keep the tongue curled"; for the flat tongue sound "z", when the tongue tip position is too high, the system prompts "Please lower the tongue tip to the front of the upper gum"; for the "c" sound, if the airflow explosive force is insufficient, the system prompts "Please increase the airflow explosive force and keep the tongue tip close to the upper gum"; when the pronunciation node is "s", if the gap between the tongue tip and the gum is too large, the system prompts "Please reduce the distance between the tongue tip and the upper gum"; during continuous pronunciation, when the tongue position conversion between adjacent syllables is not fast enough, the system prompts "Please speed up the tongue position change speed"; if the tongue position is continuously unstable, the system prompts "Please keep the tongue position stable and avoid swinging left and right". These targeted prompts are displayed in real time on the system interface in concise and clear text, helping learners quickly identify and correct pronunciation problems.

[0092] In summary, the Mandarin pronunciation contrast training method based on real-time speech recognition according to the embodiment of the present invention is explained, which generates pronunciation deviation values ​​and tongue position feedback control signals in real time by dynamically monitoring the tongue position feature parameters at different time points within the target training time period and combining the acoustic feature matching calculation. In this way, the accuracy and real-time performance of pronunciation deviation detection can be improved, which helps trainees to intuitively understand the dynamic changes of tongue position during the pronunciation process, thereby adjusting tongue position movements in time at key pronunciation nodes, and improving the accuracy and learning efficiency of Mandarin pronunciation.

[0093] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned Mandarin pronunciation comparison training method based on real-time speech recognition have been described in the above reference. Figure 1 and Figure 2 The method has been described in detail in the description of the Mandarin pronunciation contrast training method based on real-time speech recognition, and therefore, its repeated description will be omitted.

[0094] In summary, the Mandarin pronunciation contrast training method based on real-time speech recognition according to the embodiment of the present invention is explained, by dynamically monitoring the tongue position feature parameters at different time points within the target training time period, and combining the acoustic feature matching calculation, the pronunciation deviation value and the tongue position feedback control signal are generated in real time. In this way, the accuracy and real-time performance of pronunciation deviation detection can be improved, which helps the trainee to intuitively understand the dynamic changes of the tongue position during the pronunciation process, thereby adjusting the tongue position movement in time at the key pronunciation nodes, and improving the accuracy and learning efficiency of Mandarin pronunciation.

Claims

1. A Mandarin pronunciation comparison training method based on real-time speech recognition, characterized in that: include: Acquiring real-time speech data of a trainee, and converting the real-time speech data into a first acoustic feature sequence through an acoustic feature extraction algorithm; Based on the first acoustic feature sequence, a tongue position trajectory tracking algorithm is used to dynamically monitor the tongue position feature parameters to obtain a tongue tip lift height curve and a tongue surface curl degree curve, and key points of tongue position changes during pronunciation are identified based on the curves; According to the tongue position change key points, combined with the first acoustic feature sequence, a pronunciation feature vector is constructed, and the pronunciation feature vector is matched with a preset standard feature template of retroflex and flat tongue sounds to obtain a real-time pronunciation deviation value; When the deviation value exceeds the preset range, the comparison result between the ideal tongue position curve and the actual tongue position curve is presented in real time through a visual display, and targeted tongue position adjustment prompts are output at key pronunciation nodes.

2. The method for comparative Mandarin pronunciation training based on real-time speech recognition according to claim 1, wherein: The first acoustic feature sequence includes tongue position feature parameters, airflow turbulence parameters and friction noise spectrum parameters; The tongue position feature parameters are extracted by a linear prediction analysis method, including the first N order linear prediction coefficients; The airflow turbulence parameters are extracted using zero-crossing rate and short-time energy analysis methods; The friction noise spectrum parameters are calculated by fast Fourier transform.

3. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 1, wherein: Obtaining the tongue tip lift height curve and the tongue surface curling degree curve includes: Performing mapping transformation on the tongue position feature parameters in the first acoustic feature sequence to determine the vocal tract cross-sectional area distribution; Mapping the vocal tract cross-sectional area distribution into a three-dimensional space coordinate system and calculating the tongue position space coordinates; The tongue position spatial coordinates are projected and features are extracted to obtain the tongue tip lift height curve and the tongue surface curling degree curve.

4. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 3, wherein: The tongue tip lifting height curve and tongue surface curling degree curve are shown as follows: in, in, is the height of the tongue tip at time t, is the y component of the tongue position coordinate after smoothing, is the time-varying weight function of tongue tip motion, is the degree of tongue curling at time t, is the number of sampling points on the tongue contour, is the vector from the mth sampling point to the m+1th sampling point, is the angle between adjacent vectors, is the time-varying weight function of tongue curl, is the rate of change of the height of the tongue tip, is the rate of change of tongue curl, is the time constant of tongue tip movement, is the time constant of tongue curling.

5. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 4, wherein: When the key change point is identified, the features of the tongue tip lift height curve and the tongue surface curling degree curve are merged; When the local change of the fused feature function exceeds the preset threshold and the second-order derivative is zero, the moment is marked as the tongue position change key point.

6. The method for comparative Mandarin pronunciation training based on real-time speech recognition according to claim 1, wherein: The pronunciation feature vector includes tongue position features, airflow turbulence features and friction noise spectrum features. The tongue position features are the height of the tongue tip and the degree of tongue curling at the certain moment.

7. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 6, wherein: The airflow turbulence feature is centered on the key point of tongue position change, and a local analysis window is constructed. The short-time zero-crossing rate and short-time energy value of the speech signal are calculated within the window to form an airflow turbulence feature sequence; The friction noise spectrum feature is the energy distribution feature of each sub-band of the speech signal at the key point.

8. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 8, wherein: The calculation of the airflow turbulence characteristic is shown in the following formula: in, For the moment The characteristic value of airflow turbulence calculated at is the length of the local time window, which is used to determine the time range for calculating turbulence characteristics. is the displacement in the time window, ranging from -L / 2 to L / 2, is the zero-crossing rate function, which indicates the number of times the speech signal passes through the zero point in unit time. is the short-time energy function.

9. The method for comparative training of Mandarin pronunciation based on real-time speech recognition according to claim 8, characterized in that: include: The matching degree calculation is performed by respectively calculating the tongue tip height, airflow turbulence and deviation of frequency band energy distribution, and dynamically adjusting the weight coefficient to obtain an initial matching degree score through weighted summation; The initial matching score is subjected to time constraints to obtain the real-time pronunciation deviation value.

10. The method for comparative Mandarin pronunciation training based on real-time speech recognition according to claim 1, characterized in that: The real-time presentation of the comparison results of the ideal tongue position curve and the actual tongue position curve includes superimposing and displaying the ideal tongue position trajectory curve of the standard pronunciation, and the ideal tongue position trajectory curve is marked in translucent yellow; at the same time, superimposing and displaying the actual tongue position trajectory curve of the trainee, and the actual tongue position trajectory curve is marked in translucent blue, which is the tongue tip lifting height curve and the tongue surface curling degree curve; when there is overlap between the two trajectory curves, the overlapping area is displayed in green.