Spoken Chinese pronunciation automatic evaluation and correction system based on AI speech recognition

By designing an automatic evaluation and correction system for spoken Chinese pronunciation based on AI speech recognition, the problem of low accuracy in pronunciation evaluation and correction of existing systems is solved, and comprehensive evaluation and personalized correction of user pronunciation are achieved, and the efficiency and effect of pronunciation correction are improved.

CN120199272AInactive Publication Date: 2025-06-24HUBEI UNIV OF EDUCATION
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510244648.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing speech recognition system has problems such as low accuracy in the evaluation and correction of Chinese oral pronunciation, inability to fully recognize and evaluate multiple dimensions of pronunciation, and lack of personalized correction solutions.

Method used

An automatic evaluation and correction system for Chinese spoken pronunciation based on AI speech recognition is designed, including a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module and a pronunciation correction module. The system analyzes the characteristics of syllables, tone, pitch and speech speed in real time, compares them with standard pronunciation templates, calculates syllable deviation data, and provides a personalized correction plan based on the scoring results.

Benefits of technology

It achieves a more accurate and comprehensive evaluation of user pronunciation, provides targeted correction solutions, and improves the efficiency and effectiveness of pronunciation correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199272A_ABST
    Figure CN120199272A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition and artificial intelligence, in particular to an automatic evaluation and correction system for spoken Chinese pronunciation based on AI speech recognition, which comprises a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module and a pronunciation correction module. Wherein the voice input module is used for receiving a Chinese spoken pronunciation signal of a user; the voice preprocessing module is used for preprocessing the voice data; the pronunciation analysis module is used for analyzing pronunciation of the user in real time by using an AI voice recognition technology; the pronunciation scoring module is used for calculating a score value of pronunciation of the user; and the pronunciation correction module is used for providing a targeted pronunciation correction scheme for the user. According to the invention, through the multi-dimensional pronunciation evaluation and personalized correction scheme based on the AI speech recognition technology, the pronunciation deviation of the user can be accurately recognized, real-time feedback can be provided, and the pronunciation correction accuracy and efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of speech recognition and artificial intelligence, and particularly to an automatic evaluation and correction system for Chinese spoken pronunciation based on AI speech recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, speech recognition technology has gradually become an important tool in Chinese spoken language learning; especially in the fields of foreign language learning, Chinese pronunciation correction, etc., speech recognition technology can help users monitor and evaluate the accuracy of their pronunciation in real time; however, most of the existing speech recognition systems focus on simple speech-to-text functions and lack in-depth analysis of pronunciation details and personalized correction solutions; although some systems can provide basic pronunciation scores, they often only involve single audio comparison, unable to comprehensively identify and evaluate pronunciation features in multiple dimensions such as syllables, tones, pitch, speech rate, etc. in pronunciation, and lack accurate feedback and effective correction methods for users' pronunciation deviations.

[0003] Most of the current pronunciation correction systems rely on preset templates to provide standardized scoring criteria and lack the ability to make personalized adjustments according to the actual pronunciation of users; this results in the problem of low accuracy in pronunciation evaluation for many systems, unable to accurately capture the subtle differences in users' pronunciation, thereby affecting the pronunciation improvement effect of users; at the same time, based on the scoring results, most of the existing systems lack targeted and effective correction measures, usually relying only on simple text or audio feedback, lacking intuitive visual assistance and pronunciation practice guidance. Therefore, there is an urgent need for an automatic evaluation and correction system for Chinese spoken pronunciation based on AI speech recognition to solve the above problems. Summary of the Invention

[0004] Based on the above purposes, the present invention provides an automatic evaluation and correction system for Chinese spoken pronunciation based on AI speech recognition.

[0005] An automatic evaluation and correction system for Chinese spoken pronunciation based on AI speech recognition includes a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module, and a pronunciation correction module; wherein:

[0006] The speech input module: is used to receive the Chinese spoken pronunciation signal of the user and convert it into digital speech data;

[0007] The speech preprocessing module: is used to preprocess the speech data provided by the speech input module, including noise reduction processing and echo cancellation processing, and generate optimized speech data;

[0008] Pronunciation analysis module: Based on the optimized voice data, it uses AI voice recognition technology to analyze the user's pronunciation in real time, identify the pronunciation features of syllables, tones, pitch, and speech rate, and compare them with the standard pronunciation template to calculate the syllable deviation data;

[0009] Pronunciation scoring module: Receives the syllable deviation data output by the pronunciation analysis module, and combines the standard pronunciation model and scoring rules to calculate the scoring value of the user's pronunciation. This scoring value includes syllable accuracy, tone accuracy, and speech rate moderation, forming a comprehensive scoring result;

[0010] Pronunciation correction module: According to the comprehensive scoring result generated by the pronunciation scoring module, combined with the syllable deviation data provided by the pronunciation analysis module, it provides a targeted pronunciation correction plan for the user, including voice prompts, mouth shape demonstrations, and pronunciation practice guidance to help the user effectively correct pronunciation deviations.

[0011] Optionally, the voice input module includes a signal receiving unit, a signal conversion unit, and a data transmission unit; where:

[0012] Signal receiving unit: Used to receive the user's Chinese spoken pronunciation signal through an audio acquisition device, and this signal is an analog audio signal;

[0013] Signal conversion unit: Used to convert the analog audio signal received by the signal receiving unit into digital audio data through an analog-to-digital converter;

[0014] Data transmission unit: Used to transmit the digital audio data output by the signal conversion unit to the voice preprocessing module.

[0015] Optionally, the voice preprocessing module includes a denoising unit and an echo cancellation unit; where:

[0016] Denoising unit: Used to receive the digital voice data from the voice input module and suppress the background noise in the voice data using the Wiene filtering algorithm. The formula is: , where is the denoised voice signal, is the frequency response of the Wiener filter, is the input noisy voice signal;

[0017] Echo cancellation unit: Used to perform echo cancellation processing on the denoised voice data, and uses an adaptive filtering algorithm to eliminate the echo interference in the voice signal. The formula is: , where is the error signal after echo cancellation, is the input signal containing echo, is the weight coefficient of the adaptive filter, is the reference signal, is the order of the filter.

[0018] Optionally, the pronunciation analysis module includes a syllable recognition unit, a tone analysis unit, a pitch analysis unit, a speech rate analysis unit, and a pronunciation deviation calculation unit; where:

[0019] Syllable recognition unit: used to receive the optimized speech data from the speech preprocessing module, perform syllable segmentation on the speech signal using AI speech recognition technology, recognize each syllable in the user's pronunciation, and compare the recognized syllables with the corresponding syllables in the standard pronunciation template to obtain the syllable deviation data for each syllable;

[0020] Tone analysis unit: used to analyze the tone changes in the user's pronunciation, perform tone recognition on the speech signal, recognize the tone of each syllable in the user's pronunciation, and compare it with the standard tone model to calculate the tone accuracy and obtain the tone deviation data;

[0021] Pitch analysis unit: used to detect the pitch changes in the user's pronunciation, calculate the pitch deviation value of the user's pronunciation by analyzing the frequency characteristics in the user's pronunciation and combining with the standard pitch template, and compare the pitch deviation value with the standard pitch value in the standard pitch template to obtain the pitch deviation data;

[0022] Speech rate analysis unit: used to analyze the speech rate changes in the user's pronunciation, calculate the speech rate of the user's pronunciation, including the pronunciation duration of each syllable and the speech fluency, and compare it with the standard speech rate model to obtain the speech rate deviation data;

[0023] Pronunciation deviation integration unit: used to comprehensively integrate the deviation data output by the syllable recognition unit, the tone analysis unit, the pitch analysis unit, and the speech rate analysis unit to form the syllable deviation data.

[0024] Optionally, the syllable recognition unit includes:

[0025] Speech signal preprocessing: receive the optimized speech data from the speech preprocessing module, and convert the speech signal into a spectrogram using the short-time Fourier transform;

[0026] Syllable segmentation: perform feature extraction on the spectrogram, identify the syllable boundaries in the speech signal, and automatically determine the start and end positions of the syllables by learning the time and frequency characteristics in the speech signal through multi-layer convolution operations;

[0027] Syllable recognition: analyze the time series characteristics of the speech signal to identify the specific content of each syllable;

[0028] Syllable comparison: Compare the recognized syllables with the corresponding syllables in the standard pronunciation template, and use the dynamic time warping algorithm to compare the temporal characteristics of the syllables to obtain the matching degree between the syllables;

[0029] Deviation calculation: Based on the syllable comparison results, quantify the time difference and frequency difference of the syllables, and calculate the syllable deviation data of each syllable.

[0030] Optionally, the tone analysis unit includes:

[0031] Tone feature extraction: Receive the optimized speech data from the syllable recognition unit, and use the Mel-frequency cepstral coefficients to extract the tone features in the speech signal;

[0032] Tone classification and recognition: Classify the extracted Mel-frequency cepstral coefficients through a trained deep neural network to identify the tone type of each syllable;

[0033] Tone accuracy calculation: Compare the recognized tone type with the standard tone model, and calculate the tone accuracy;

[0034] Tone deviation calculation: Calculate the tone deviation data according to the difference between the tone accuracy and the standard tone model.

[0035] Optionally, the pronunciation scoring module includes a standard pronunciation model unit, a scoring rule application unit, and a scoring calculation unit; where:

[0036] Standard pronunciation model unit: Used to store the preset standard pronunciation model, including standard syllables, tones, pitches, and speech rates, as the benchmark for scoring;

[0037] Scoring rule application unit: Based on the preset scoring rules, used to compare and analyze the received syllable deviation data with the standard pronunciation model, and allocate corresponding scoring criteria according to the weights of different pronunciation features;

[0038] Scoring calculation unit: According to the analysis results of the scoring rule application unit, calculate the scoring values of syllable accuracy, tone accuracy, and speech rate moderation respectively, and generate a comprehensive scoring result of the user's pronunciation by integrating the scoring values of each item.

[0039] Optionally, the scoring rule application unit includes:

[0040] Weight assignment: According to the preset scoring rules, assign weights to syllable accuracy, tone accuracy, and speech rate moderation respectively; let the weight of syllable accuracy be , the weight of tone accuracy be , the weight of speech rate moderation be , where ;

[0041] Determination of scoring criteria: According to the weights of each pronunciation feature, corresponding scoring criteria are assigned; assuming the total score is 100 points, the scoring criteria for syllable accuracy are as follows: ; The scoring criteria for tone accuracy are as follows: ; The scoring criteria for appropriate speech rate are as follows: ;

[0042] Application of scoring criteria: Apply the assigned scoring criteria for each pronunciation feature and to their corresponding pronunciation feature score values.

[0043] Optionally, the scoring calculation unit includes:

[0044] Calculation of syllable accuracy: Receive syllable deviation data from the pronunciation analysis module, including the number of syllables pronounced correctly and the total number of syllables pronounced, and calculate the syllable accuracy according to the following formula : , where represents the number of syllables pronounced correctly by the user, represents the total number of syllables pronounced by the user;

[0045] Calculation of tone accuracy: Receive tone deviation data from the pronunciation analysis module, including the number of tones pronounced correctly and the total number of tones pronounced, and calculate the tone accuracy according to the following formula : , where represents the number of tones pronounced correctly by the user, represents the total number of tones pronounced by the user;

[0046] Calculation of appropriate speech rate: Receive speech rate deviation data from the pronunciation analysis module, including the actual speech rate of the user and the standard speech rate , and calculate the appropriate speech rate according to the following formula : , where represents the actual speech rate of the user, represents the standard speech rate;

[0047] Generation of comprehensive score: Receive the , and calculation results from the above steps; and according to the preset weights and , use the following formula to calculate the comprehensive score : , where are the weights of syllable accuracy, tone accuracy and appropriate speech rate respectively, and .

[0048] Optionally, the pronunciation correction module includes a voice prompt unit, a mouth shape demonstration unit, and a pronunciation practice guidance unit; where:

[0049] Voice prompt unit: Used to generate and play targeted voice prompts according to the comprehensive scoring result generated by the pronunciation scoring module, combined with the syllable deviation data provided by the pronunciation analysis module, to guide the user to make real-time adjustments during the pronunciation process;

[0050] Mouth shape demonstration unit: Used to generate a standard mouth shape animation or video demonstration corresponding to the syllables that the user needs to improve according to the pronunciation deviation data, and help the user correctly imitate and practice the pronunciation actions through visual assistance;

[0051] Pronunciation practice guidance unit: Used to provide personalized pronunciation practice plans and guidance according to the comprehensive scoring result and the pronunciation deviation of each syllable, including pronunciation practice questions, repeated practice prompts, and progress tracking.

[0052] Advantages of the present invention:

[0053] In the present invention, through the comprehensive analysis of pronunciation features such as syllables, tones, pitches, and speech rates, the deviations in the user's pronunciation can be identified in detail, and a quantitative scoring result can be generated based on these deviations, providing more accurate pronunciation feedback; compared with the traditional simple audio comparison method, it can analyze all aspects of pronunciation more comprehensively, improving the accuracy and comprehensiveness of pronunciation evaluation.

[0054] In the present invention, through real-time voice prompts, standard mouth shape demonstrations, and targeted pronunciation practice guidance, it helps the user to more effectively improve pronunciation deviations, thereby achieving faster and more accurate pronunciation correction; this intelligent and personalized pronunciation correction method has higher practical application value and learning efficiency than the single scoring feedback and templated training methods in the prior art. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 Schematic diagram of the automatic Chinese spoken language pronunciation evaluation and correction system according to an embodiment of the present invention;

[0057] Figure 2 Schematic diagram of the pronunciation analysis module according to an embodiment of the present invention. Detailed Embodiments

[0058] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0059] As Figure 1 - Figure 2 shown, a Chinese spoken language pronunciation automatic evaluation and correction system based on AI speech recognition includes a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module, and a pronunciation correction module; among them:

[0060] The speech input module: is used to receive the Chinese spoken language pronunciation signal of the user and convert it into digital speech data;

[0061] The speech preprocessing module: is used to preprocess the speech data provided by the speech input module, including noise reduction processing and echo cancellation processing, and generate optimized speech data for the pronunciation analysis module to perform accurate recognition;

[0062] The pronunciation analysis module: based on the optimized speech data, uses AI speech recognition technology to perform real-time analysis on the user's pronunciation, recognizes the pronunciation features of syllables, tones, pitch, and speech rate, and compares them with the standard pronunciation template, thereby calculating the syllable deviation data for the pronunciation scoring module to further evaluate;

[0063] The pronunciation scoring module: receives the syllable deviation data output by the pronunciation analysis module, and combines the standard pronunciation model and the scoring rules to calculate the scoring value of the user's pronunciation. This scoring value includes syllable accuracy rate, tone accuracy, and speech rate moderation, forming a comprehensive scoring result for the pronunciation correction module to further process;

[0064] The pronunciation correction module: according to the comprehensive scoring result generated by the pronunciation scoring module, combines the syllable deviation data provided by the pronunciation analysis module, and provides a targeted pronunciation correction plan for the user, including voice prompts, mouth shape demonstrations, and pronunciation practice guidance, to help the user effectively correct pronunciation deviations.

[0065] The speech input module includes a signal receiving unit, a signal conversion unit, and a data transmission unit; among them:

[0066] The signal receiving unit: is used to receive the Chinese spoken language pronunciation signal of the user through an audio acquisition device, and this signal is an analog audio signal;

[0067] Signal conversion unit: It is used to convert the analog audio signal received by the signal receiving unit into digital audio data through an analog-to-digital converter (ADC). The converted audio data is a digital format sound waveform with a certain sampling frequency and bit depth, facilitating subsequent processing and analysis;

[0068] Data transmission unit: It is used to transmit the digital audio data output by the signal conversion unit to the speech preprocessing module for subsequent signal processing and analysis; Through the organic combination of the above units, it ensures the accurate reception and conversion of the user's spoken signal, providing high-quality input data for subsequent speech processing and analysis.

[0069] The speech preprocessing module includes a denoising unit and an echo cancellation unit; Among them:

[0070] Denoising unit: It is used to receive the digital speech data from the speech input module and suppress the background noise in the speech data using the Wiene filtering algorithm. The formula is: , where, is the denoised speech signal, is the frequency response of the Wiener filter, is the input noisy speech signal;

[0071] Echo cancellation unit: It is used to perform echo cancellation processing on the denoised speech data, and uses an adaptive filtering algorithm to eliminate the echo interference in the speech signal. The formula is: , where, is the error signal after echo cancellation, is the input signal containing echo, is the weight coefficient of the adaptive filter, is the reference signal, is the order of the filter; By setting the denoising unit and the echo cancellation unit in the speech preprocessing module and using the Wiener filtering algorithm and the adaptive filtering algorithm, the background noise and echo interference in the speech signal are effectively reduced, and the clarity and purity of the speech data are improved, which provides high-quality input data for the subsequent pronunciation analysis module.

[0072] The pronunciation analysis module includes a syllable recognition unit, a tone analysis unit, a pitch analysis unit, a speech rate analysis unit, and a pronunciation deviation calculation unit; Among them:

[0073] Syllable recognition unit: It is used to receive the optimized speech data from the speech preprocessing module, use AI speech recognition technology to segment the speech signal into syllables, identify each syllable in the user's pronunciation, and compare the recognized syllables with the corresponding syllables in the standard pronunciation template to obtain the syllable deviation data of each syllable;

[0074] Tone analysis unit: It is used to analyze the tone changes in the user's pronunciation, identify the tones of the speech signal, recognize the tone of each syllable in the user's pronunciation, compare it with the standard tone model, calculate the tone accuracy, and obtain the tone deviation data;

[0075] Pitch analysis unit: It is used to detect the pitch changes in the user's pronunciation. By analyzing the frequency characteristics in the user's pronunciation and combining with the standard pitch template, calculate the pitch deviation value of the user's pronunciation, and compare this pitch deviation value with the standard pitch value in the standard pitch template to obtain the pitch deviation data; The calculation formula for the above pitch deviation value is: , where, represents the frequency of the user's pronunciation of the th syllable, represents the standard pronunciation frequency of the th syllable, is the pitch deviation value;

[0076] Speech rate analysis unit: It is used to analyze the speech rate changes in the user's pronunciation, calculate the speech rate of the user's pronunciation, including the pronunciation duration of each syllable and the speech fluency, and compare it with the standard speech rate model to obtain the speech rate deviation data, and output the speech rate deviation data together with other syllable deviation data; The speech rate calculation formula is: , where, is the speech rate, is the number of syllables included in the pronunciation, is the total duration of the pronunciation (in seconds);

[0077] Pronunciation deviation integration unit: It is used to integrate the deviation data output by the syllable recognition unit, tone analysis unit, pitch analysis unit and speech rate analysis unit to form syllable deviation data, which will be further provided to the pronunciation scoring module for scoring processing; Through the independent analysis and accurate calculation of pronunciation features such as syllables, tones, pitches and speech rates, it can comprehensively and accurately evaluate the deviation of the user's Chinese spoken pronunciation, provide detailed feedback data, and provide a basis for subsequent pronunciation correction.

[0078] The syllable recognition unit includes:

[0079] Speech signal preprocessing: Receive the optimized speech data from the speech preprocessing module, and use the short-time Fourier transform (STFT) to convert the speech signal into a spectrogram to represent the time-frequency characteristics of the speech signal; The STFT calculation formula is: , where, is the signal representation in the time-frequency domain, is the time-domain signal, is the window function, is the frequency, is the time;

[0080] Syllable segmentation: Extract features from the spectrogram, identify the syllable boundaries in the speech signal, and learn the temporal and frequency features in the speech signal through multi-layer convolutional operations, so as to automatically determine the start and end positions of syllables;

[0081] Syllable recognition: Analyze the time series features of the speech signal, identify the specific content of each syllable, and the RNN model accurately identifies the pronunciation of each syllable by memorizing the information of previous time steps;

[0082] Syllable comparison: Compare the recognized syllables with the corresponding syllables in the standard pronunciation template, and use the dynamic time warping (DTW) algorithm to compare the temporal features of the syllables to obtain the matching degree between syllables. The calculation formula of the dynamic time warping algorithm is:

[0083] , where is the matching distance, and are the feature vectors of the th syllable and the th standard syllable respectively, is the distance metric between feature vectors;

[0084] Deviation calculation: Based on the syllable comparison results, quantify the time difference and frequency difference of syllables, and calculate the syllable deviation data of each syllable; Through the precise syllable segmentation and comparison process of the syllable recognition unit, it can efficiently and accurately identify each syllable in the user's pronunciation and calculate the pronunciation deviation, providing a precise basis for subsequent pronunciation evaluation and correction.

[0085] The tone analysis unit includes:

[0086] Tone feature extraction: Receive the optimized speech data from the syllable recognition unit, and use the Mel-frequency cepstral coefficients (MFCC) to extract the tone features in the speech signal; The calculation formula of the Mel-frequency cepstral coefficients is: , where represents the th Mel-frequency cepstral coefficient, represents the short-time Fourier transform (STFT) result of the speech signal, is the frequency response of the Mel filter bank, is the total number of filter banks;

[0087] Tone classification and recognition: Classify the extracted Mel-frequency cepstral coefficients through a trained deep neural network (DNN) to identify the tone type of each syllable; The calculation formula of the deep neural network is: , where is the output of the network (i.e., the tone category), is the weight matrix, is the input Mel Frequency Cepstral Coefficient (MFCC) feature vector, is the bias term, is the activation function;

[0088] Tone accuracy calculation: Compare the recognized tone types with the standard tone model to calculate the tone accuracy; the formula is: , where represents the total number of syllables, and the number of matching tones represents the number of syllables that match the corresponding tone types in the standard pronunciation model;

[0089] Tone deviation calculation: Calculate the tone deviation data based on the difference between the tone accuracy and the standard tone model, the formula is: , where represents the recognized tone of the th syllable in the user's pronunciation, is the tone type in the standard tone model, and the tone deviation represents the average deviation value of all syllables; Through the above steps, use Mel Frequency Cepstral Coefficients to extract tone features, and combine with a deep neural network for tone classification and recognition, which can accurately identify the tone types in the user's pronunciation; By comparing with the standard tone model and calculating the accuracy and deviation, it can effectively evaluate the accuracy of the user's tone.

[0090] The pronunciation scoring module includes a standard pronunciation model unit, a scoring rule application unit, and a scoring calculation unit; among them:

[0091] Standard pronunciation model unit: Used to store the preset standard pronunciation model, including standard syllables, tones, pitches, and speaking speeds, as the benchmark for scoring;

[0092] Scoring rule application unit: Based on the preset scoring rules, used to compare and analyze the received syllable deviation data with the standard pronunciation model, and allocate corresponding scoring criteria according to the weights of different pronunciation features;

[0093] Scoring calculation unit: According to the analysis results of the scoring rule application unit, calculate the scoring values of syllable accuracy rate, tone accuracy, and speaking speed moderation respectively, and generate the comprehensive scoring result of the user's pronunciation by integrating the scoring values of each item; Through the design of the above units, it can comprehensively and accurately evaluate the user's pronunciation performance, and combine the standard pronunciation model and scoring rules to provide a targeted and objective comprehensive scoring result for the user, effectively guiding pronunciation correction.

[0094] The scoring rule application unit includes:

[0095] Weight assignment: According to the preset scoring rules, assign weights to the syllable accuracy rate, tone accuracy, and speaking speed moderation respectively; Let the weight of the syllable accuracy rate be , the weight of tone accuracy is , the weight of appropriate speech rate is , where ;

[0096] Determination of scoring criteria: According to the weights of each pronunciation feature, corresponding scoring criteria are assigned; assuming the total score is 100 points, the scoring criteria for syllable accuracy rate are: ; the scoring criteria for tone accuracy are: ; the scoring criteria for appropriate speech rate are: ;

[0097] Application of scoring criteria: Apply the assigned scoring criteria for each pronunciation feature and to their corresponding pronunciation feature score values, ensuring that the score of each pronunciation feature is assigned a corresponding score range according to its weight; through the above steps, the weights of each pronunciation feature can be reasonably assigned according to the preset scoring rules, and the corresponding scoring criteria can be determined, providing a reliable scoring basis for subsequent pronunciation correction.

[0098] The scoring calculation unit includes:

[0099] Calculation of syllable accuracy rate: Receive syllable deviation data from the pronunciation analysis module, including the number of syllables with correct pronunciation and the total number of syllables pronounced, and calculate the syllable accuracy rate according to the following formula : , where represents the number of syllables with correct pronunciation by the user, represents the total number of syllables pronounced by the user;

[0100] Calculation of tone accuracy: Receive tone deviation data from the pronunciation analysis module, including the number of tones with correct pronunciation and the total number of tones pronounced, and calculate the tone accuracy according to the following formula : , where represents the number of tones with correct pronunciation by the user, represents the total number of tones pronounced by the user;

[0101] Calculation of appropriate speech rate: Receive speech rate deviation data from the pronunciation analysis module, including the actual speech rate of the user and the standard speech rate , and calculate the appropriate speech rate according to the following formula : , where represents the actual speech rate of the user, represents the standard speech rate;

[0102] Generation of comprehensive score: Receive from the above steps , and the calculation result; and according to the preset weights and , use the following formula to calculate the comprehensive score : , where are the weights of syllable accuracy, tone accuracy, and speech rate moderation respectively, and .

[0103] The pronunciation correction module includes a voice prompt unit, a mouth shape demonstration unit, and a pronunciation practice guidance unit; among them:

[0104] Voice prompt unit: Used to generate and play targeted voice prompts according to the comprehensive score result generated by the pronunciation scoring module and combined with the syllable deviation data provided by the pronunciation analysis module, guiding the user to make real-time adjustments during the pronunciation process;

[0105] Mouth shape demonstration unit: Used to generate standard mouth shape animations or video demonstrations corresponding to the syllables that the user needs to improve according to the pronunciation deviation data, and help the user correctly imitate and practice the pronunciation actions through visual assistance;

[0106] Pronunciation practice guidance unit: Used to provide personalized pronunciation practice plans and guidance according to the comprehensive score result and the pronunciation deviation of each syllable, including pronunciation practice questions, repeated practice prompts, and progress tracking, to help the user systematically improve the pronunciation quality; By setting a voice prompt unit, a mouth shape demonstration unit, and a pronunciation practice guidance unit in the pronunciation correction module, a multi-dimensional and personalized pronunciation correction solution can be provided according to the user's comprehensive score result and specific pronunciation deviation.

[0107] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without these detailed descriptions. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0108] The above description is only a preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition, characterized in that: It includes a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module and a pronunciation correction module; wherein: Voice input module: used to receive the user's Chinese spoken pronunciation signal and convert it into digital voice data; Speech preprocessing module: used to preprocess the speech data provided by the speech input module, including denoising and echo cancellation, to generate optimized speech data; Pronunciation analysis module: Based on the optimized voice data, AI speech recognition technology is used to analyze the user's pronunciation in real time, identify the pronunciation characteristics of syllables, intonations, pitch and speech speed, and compare them with the standard pronunciation template to calculate the syllable deviation data; Pronunciation scoring module: receives the syllable deviation data output by the pronunciation analysis module, and calculates the user's pronunciation score based on the standard pronunciation model and scoring rules. The score includes syllable accuracy, tone accuracy, and speech speed moderation, forming a comprehensive scoring result. Pronunciation Correction Module: Based on the comprehensive scoring results generated by the pronunciation scoring module and the syllable deviation data provided by the pronunciation analysis module, it provides users with targeted pronunciation correction solutions, including voice prompts, lip shape demonstrations and pronunciation practice guidance, to help users effectively correct pronunciation deviations.

2. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 1 is characterized in that: The voice input module includes a signal receiving unit, a signal conversion unit and a data transmission unit; wherein: Signal receiving unit: used to receive the user's Chinese spoken pronunciation signal through the audio acquisition device, and the signal is an analog audio signal; Signal conversion unit: used for converting the analog audio signal received by the signal receiving unit into digital audio data through an analog-to-digital converter; Data transmission unit: used to transmit the digitized audio data output by the signal conversion unit to the speech preprocessing module.

3. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 1 is characterized in that: The speech preprocessing module includes a denoising unit and an echo cancellation unit; wherein: De-noising unit: used to receive the digitized voice data from the voice input module and suppress the background noise in the voice data using the Wiene filtering algorithm. The formula is: ,in, is the denoised speech signal, is the frequency response of the Wiener filter, is the input noisy speech signal; Echo cancellation unit: used to perform echo cancellation processing on the denoised voice data, using an adaptive filtering algorithm to eliminate echo interference in the voice signal. The formula is: ,in, To eliminate the error signal after echo, is the input signal containing the echo, is the weight coefficient of the adaptive filter, is the reference signal, is the order of the filter.

4. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 1 is characterized in that: The pronunciation analysis module includes a syllable recognition unit, a tone analysis unit, a pitch analysis unit, a speech rate analysis unit and a pronunciation deviation calculation unit; wherein: Syllable recognition unit: used to receive the optimized voice data from the voice preprocessing module, use AI voice recognition technology to segment the voice signal into syllables, identify each syllable in the user's pronunciation, and compare the identified syllables with the corresponding syllables in the standard pronunciation template to obtain the syllable deviation data of each syllable; Tone analysis unit: used to analyze the tone changes in the user's pronunciation, perform tone recognition on the voice signal, identify the tone of each syllable in the user's pronunciation, compare it with the standard tone model, calculate the tone accuracy, and obtain the tone deviation data; Pitch analysis unit: used to detect pitch changes in user pronunciation, calculate the pitch deviation value of the user's pronunciation by analyzing the frequency characteristics of the user's pronunciation, combined with the standard pitch template, and compare the pitch deviation value with the standard pitch value in the standard pitch template to obtain pitch deviation data; Speech rate analysis unit: used to analyze the speech rate changes in the user's pronunciation, calculate the user's pronunciation rate, including the pronunciation duration of each syllable and the fluency of the speech, and compare it with the standard speech rate model to obtain speech rate deviation data; Pronunciation deviation integration unit: used to integrate the deviation data output by the syllable recognition unit, the tone analysis unit, the pitch analysis unit and the speech rate analysis unit to form syllable deviation data.

5. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 4 is characterized in that: The syllable recognition unit comprises: Speech signal preprocessing: receiving the optimized speech data from the speech preprocessing module and converting the speech signal into a spectrogram using short-time Fourier transform; Syllable segmentation: Extract features from the spectrogram, identify syllable boundaries in the speech signal, and learn the time and frequency features in the speech signal through multi-layer convolution operations to automatically determine the start and end positions of the syllable; Syllable recognition: Analyze the time series characteristics of speech signals to identify the specific content of each syllable; Syllable matching: The recognized syllables are compared with the corresponding syllables in the standard pronunciation template, and the dynamic time warping algorithm is used to compare the timing characteristics of the syllables to obtain the matching degree between the syllables; Deviation calculation: Based on the syllable comparison results, the time difference and frequency difference of the syllables are quantified, and the syllable deviation data of each syllable is calculated.

6. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 5, characterized in that: The tone analysis unit comprises: Tone feature extraction: Receives optimized speech data from the syllable recognition unit and uses Mel-frequency cepstral coefficients to extract tone features in the speech signal; Tone classification and recognition: The extracted Mel-frequency cepstral coefficients are classified through a trained deep neural network to identify the tone type of each syllable; Tone accuracy calculation: compare the recognized tone type with the standard tone model to calculate the tone accuracy; Tone deviation calculation: Tone deviation data is calculated based on the difference between the tone accuracy and the standard tone model.

7. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 1, characterized in that: The pronunciation scoring module includes a standard pronunciation model unit, a scoring rule application unit and a scoring calculation unit; wherein: Standard pronunciation model unit: used to store the preset standard pronunciation model, including standard syllables, intonation, pitch and speaking speed, as the basis for scoring; Scoring rule application unit: based on preset scoring rules, used to compare and analyze the received syllable deviation data with the standard pronunciation model, and assign corresponding scoring standards according to the weights of different pronunciation features; Scoring calculation unit: based on the analysis results of the scoring rule application unit, the scoring values ​​of syllable accuracy, tone accuracy and speech speed moderation are calculated respectively, and the comprehensive scoring result of the user's pronunciation is generated by combining the various scoring values.

8. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 7 is characterized in that: The scoring rule application unit comprises: Weight allocation: According to the preset scoring rules, weights are allocated to syllable accuracy, tone accuracy and speaking speed moderation respectively; the weight of syllable accuracy is set to , the weight of tone accuracy is , the weight of moderate speaking speed is ,in ; Scoring criteria: According to the weight of each pronunciation feature, the corresponding scoring criteria are assigned; assuming the total score is 100 points, the scoring criteria for syllable accuracy are: ; The scoring criteria for tone accuracy are: ; The scoring criteria for moderate speaking speed are: ; Application of scoring criteria: The scoring criteria for each pronunciation feature are assigned and Applied to the corresponding pronunciation feature score values.

9. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 8, characterized in that: The scoring calculation unit comprises: Syllable accuracy calculation: Receive syllable deviation data from the pronunciation analysis module, including the number of correctly pronounced syllables and the total number of pronounced syllables, and calculate the syllable accuracy according to the following formula : ,in, Indicates the number of syllables that the user pronounced correctly. Indicates the total number of syllables pronounced by the user; Tone accuracy calculation: Receive tone deviation data from the pronunciation analysis module, including the number of correctly pronounced tones and the total number of pronounced tones, and calculate the tone accuracy according to the following formula : ,in, Indicates the number of tones that the user pronounced correctly, Indicates the total number of tones pronounced by the user; Speech speed moderation calculation: Receive speech speed deviation data from the pronunciation analysis module, including the user's actual speech speed Standard speaking speed , and calculate the moderate speaking speed according to the following formula : ,in, Indicates the user's actual speaking speed. Indicates standard speaking speed; Comprehensive score generation: received from the above steps , and The calculation results; and according to the preset weights and , use the following formula to calculate the comprehensive score : ,in, are the weights of syllable accuracy, tone accuracy and speaking speed moderation, respectively, and .

10. The Chinese spoken pronunciation automatic evaluation and correction system based on AI speech recognition according to claim 1, characterized in that: The pronunciation correction module includes a voice prompt unit, a lip shape demonstration unit and a pronunciation practice guidance unit; wherein: Voice prompt unit: used to generate and play targeted voice prompts based on the comprehensive scoring results generated by the pronunciation scoring module and the syllable deviation data provided by the pronunciation analysis module, so as to guide the user to make real-time adjustments during the pronunciation process; Lip shape demonstration unit: used to generate standard lip shape animation or video demonstration corresponding to the syllables that users need to improve based on pronunciation deviation data, and help users correctly imitate and practice pronunciation movements through visual assistance; Pronunciation practice guidance unit: used to provide personalized pronunciation practice plans and guidance based on the comprehensive scoring results and pronunciation deviations of each syllable, including pronunciation practice questions, repetition practice prompts and progress tracking.

Citation Information

Cited By

  • Chinese tone error correction system and method based on cross-modal adversarial generation

    CN120412556A

  • Mandarin pronunciation evaluation system based on deep learning

    CN120431970A

  • A Mandarin pronunciation evaluation system based on deep learning

    CN120431970B

  • Spoken language pronunciation training correction system based on intelligent equipment

    CN121034351A