Real-time pronunciation correction equipment for English learner

Through deep neural network and convolutional neural network technology, errors in pronunciation in English learners are analyzed and corrected in real time, and the problems of low accuracy and lack of personalized feedback in the existing technology are solved, achieving efficient pronunciation correction and learning efficiency improvement.

CN120108425APending Publication Date: 2025-06-06NANJING COLLEGE OF CHEM TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510259049.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is less accurate when facing the accent, speech speed, pronunciation or language barriers in English learners' pronunciation, and lacks personalized and real-time feedback, making it difficult to conduct detailed analysis and in-depth adjustments.

Method used

Advanced technologies such as deep neural networks and convolutional neural networks are adopted to collect real-time pronunciation data through microphones, perform noise filtering and speech enhancement, extract MFCC features, build an English pronunciation model, compare the differences between learners' pronunciation and standard pronunciation in real time, combine acoustics and language models for error evaluation, generate corrective feedback, and adjust training strategies through machine learning.

Benefits of technology

It realizes real-time pronunciation correction with high accuracy, provides personalized feedback and guidance, and helps learners improve pronunciation errors in a targeted manner and improve learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108425A_ABST
    Figure CN120108425A_ABST
Patent Text Reader

Abstract

A real-time pronunciation correction device for English learners belongs to the field of voice processing and comprises an acquisition and preprocessing module, a voice feature extraction module, a standard pronunciation model construction module, a data customization management unit, a pronunciation correction suggestion module, a correction and progress tracking module and an adjustment and optimization module. By combining the deep neural network and the convolutional neural network deep learning technology, the pronunciation characteristics of the learner can be efficiently analyzed and processed. By comparing the difference between the pronunciation of the learner and the standard pronunciation in real time, pronunciation errors are accurately detected, so that high-accuracy pronunciation correction is realized. And in combination with machine learning and data analysis, the equipment can dynamically adjust a pronunciation training strategy according to individual requirements of the learner, and provide customized pronunciation guidance and suggestions. The personalized feedback helps the learner to improve the pronunciation error of the learner in a targeted manner, and the learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of speech processing, and more specifically, particularly relates to a real-time pronunciation correction device for English learners. Background Art

[0002] With the development of deep learning technology, speech recognition technology has made significant progress, especially in natural language processing and automatic speech recognition (ASR). Equipment can efficiently convert speech into text, but there are still challenges in the accuracy and complexity of individual pronunciation. Existing speech recognition technology has low accuracy when faced with accents, speaking speed, non-standard pronunciation or language barriers. In addition, the subtle differences between standard pronunciation and non-standard pronunciation are difficult to correct automatically through existing equipment. Traditional pronunciation correction technology mainly relies on manual evaluation and guidance, such as voice coaches or pronunciation applications. However, these technologies have a low degree of intelligence and often cannot provide personalized feedback based on learners' real-time performance. Most of the existing pronunciation correction technologies are based on rules or limited data sets, lack personalized and real-time feedback, and it is difficult to conduct detailed analysis and in-depth adjustments to learners' pronunciation.

[0003] Most current pronunciation assessment technologies rely on static audio recording and post-processing, but lack an instant feedback mechanism. Although some language learning applications provide pronunciation correction functions, most devices still rely on manual editing and simple comparisons, failing to achieve accurate real-time feedback. Real-time speech analysis faces problems such as noise interference and background sound effects. How to accurately extract and analyze learners' pronunciation without affecting the user experience and give instant correction feedback remains a technical challenge. Summary of the invention

[0004] In view of the above or existing problems of real-time pronunciation correction devices for English learners, the present invention is proposed.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] The embodiment of the present invention provides a real-time pronunciation correction device for English learners, including: a collection and preprocessing module, which is used to collect the learner's real-time pronunciation data using a microphone, perform noise filtering and speech enhancement, and use a deep neural network to process background noise;

[0007] Speech feature extraction module, used to extract learner's pronunciation features using MFCC;

[0008] The standard pronunciation model building module is used to build an English pronunciation model and use convolutional neural networks to learn and simulate standard pronunciation;

[0009] The data customization management unit is used to detect whether there are pronunciation errors by comparing the learner's pronunciation with the standard pronunciation in real time, and to conduct error assessment using a similarity measurement method by combining the acoustic model with the language model;

[0010] The pronunciation correction suggestion module is used to generate correction feedback based on real-time detection results, provide audio feedback or text prompts, and point out problems in pronunciation;

[0011] Correction and Progress Tracking module, used to track learners’ progress by recording every pronunciation error and improvement;

[0012] The adjustment and optimization module is used to adjust pronunciation training strategies and provide pronunciation guidance based on learners' progress and learning style through machine learning and data analysis.

[0013] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the method of collecting the real-time pronunciation data of the learners by using a microphone, performing noise filtering and speech enhancement, and processing background noise by using a deep neural network includes:

[0014] Divide the speech signal into frames, convert it to the frequency domain, estimate the noise spectrum of each frame, subtract the noise spectrum from the speech spectrum to obtain the purified speech spectrum, perform inverse transform on the purified spectrum, and restore the time domain signal:

[0015] S(f, t) = X(f, t) - N(f, t)

[0016] Where S(f, t) is the denoised speech signal spectrum, X(f, t) is the noisy speech spectrum, and N(f, t) is the estimated noise spectrum.

[0017] During the training process, a noisy speech signal is input and the speech signal is recovered from the noisy signal; the noisy speech signal is represented as a feature vector, multiple hidden layers are used, features are extracted layer by layer through a nonlinear activation function, and the denoised speech signal is output; during the training process, the difference between the predicted speech signal and the real speech signal is minimized by optimizing the loss function.

[0018] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the method of extracting the pronunciation features of the learners using MFCC includes:

[0019] The continuous speech signal is divided into several short time frames, and the signal of each frame is windowed by a window function:

[0020]

[0021] Where N is the length of the frame and w(n) is the windowing function;

[0022] Each frame of windowed signal is converted to the frequency domain through fast Fourier transform to obtain the spectrum of the frame. Through FFT, the frequency components of the signal are obtained:

[0023]

[0024] Among them, X(f) is the frequency domain signal, and x(n) is the windowed time domain signal;

[0025] The spectrum is processed by the Mel filter bank, the frequency range is mapped to the Mel scale, and the Mel spectrum is transformed using DCT to obtain the final MFCC features. The DCT formula is:

[0026]

[0027] Where Cn is the nth MFCC coefficient, M is the number of Mel bands, and Sm(t) is the energy value of the Mel band.

[0028] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the construction of the English pronunciation model and the use of a convolutional neural network to learn and simulate standard pronunciation include:

[0029] The input layer receives a two-dimensional matrix containing MFCC features. The extraction process of each MFCC coefficient generates a feature matrix of shape (T, F), where T is the number of time frames and F is the number of MFCC coefficients. Spatial local features are extracted from the input features. Each convolution kernel captures the local pattern in the signal through the convolution operation.

[0030] Then, the local time-frequency features are extracted from the input feature matrix through the convolution operation. The input is set to X, the convolution kernel is set to W, and the convolution operation formula is as follows:

[0031]

[0032] Among them, Y(t,f) is the feature map output by the convolution operation, W(i,j) is the convolution kernel with size k×k, t and f are the indices of time frame and frequency,

[0033] The convolution operation extracts local audio patterns from the input features by sliding the convolution kernel.

[0034] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the method of detecting whether there is a pronunciation error by comparing the difference between the learner's pronunciation and the standard pronunciation in real time includes:

[0035] By calculating the similarity between the learner's pronunciation and the standard pronunciation, pronunciation errors are detected. The cosine angle between the learner's pronunciation feature vector and the standard pronunciation feature vector is calculated to measure the similarity between the two. The formula is:

[0036]

[0037] Among them, A and B are two pronunciation feature vectors, ||A|| and ||B|| are their norms respectively;

[0038] According to the comparison result, a threshold is set, and if the similarity is lower than the threshold, a pronunciation error is detected.

[0039] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the method of combining the acoustic model with the language model and using a similarity measurement method to perform error evaluation includes:

[0040] The extracted MFCC features are input into the acoustic model, the pronunciation is decoded using a decoding algorithm, and a candidate speech unit sequence is output. The acoustic model outputs the phoneme prediction at each moment to form a time series A = (a1, a2, ..., aT), where T is the number of speech frames and aT is the phoneme prediction at the Tth moment;

[0041] The language model evaluates whether the learner's pronunciation conforms to the grammatical and semantic rules of the language, and judges its rationality by calculating the probability of the learner's pronunciation; for the input phoneme sequence A = (a1, a2, ..., aT), the language model calculates the probability P(A) of the sequence.

[0042] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the method of generating correction feedback based on the real-time detection results, providing audio feedback or text prompts, and pointing out problems in pronunciation includes:

[0043] When the learner makes a mistake in pronunciation, the device plays a short prompt tone to remind the learner to pay attention to the pronunciation problem;

[0044] For incorrectly pronounced parts, the device plays audio of standard pronunciation for learners to compare with. The device plays the correct pronunciation demonstration word by word and points out the location of the learner's pronunciation deviation.

[0045] If the learner makes mistakes in intonation or rhythm, the device gives suggestions for adjustments through audio feedback.

[0046] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, wherein: the correction feedback is generated based on the real-time detection result, audio feedback or text prompts are provided, and problems in pronunciation are pointed out, and the following is also included:

[0047] When it is detected that the learner has pronounced a certain phoneme incorrectly, a text prompt points out the wrong phoneme and gives the correct pronunciation; if the error occurs in a word or phrase, the device displays the wrong word on the screen and highlights or underlines the wrong part;

[0048] For grammatical or phonetic deviations, the device provides grammatical or language regularity hints.

[0049] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the progress tracking is performed by recording each pronunciation error and improvement of the learner, including:

[0050] Regularly evaluate learners' overall pronunciation improvement and give an overall evaluation report every week, which includes the change in the number of errors and the improvement in pronunciation accuracy;

[0051] By counting the error patterns of each learner, and providing improvement plans based on the problems.

[0052] As a preferred solution of the real-time pronunciation correction device for English learners of the present invention, the pronunciation training strategy is adjusted through machine learning and data analysis according to the learner's progress and learning style, and pronunciation guidance is provided, including:

[0053] The device adjusts the feedback method according to the learner's feedback preference. For learners who like to hear the comparison between their own pronunciation and the standard pronunciation, the device will provide multi-audio feedback; for learners who like to correct themselves through text prompts, the device provides multi-pronunciation suggestions through text;

[0054] For learners who progress quickly, the device adjusts the training frequency and reduces the time of a single exercise; for learners who progress slowly, the device extends the exercise time and increases the frequency of exercise.

[0055] The beneficial effects of the present invention are as follows: the present invention can efficiently analyze and process the pronunciation characteristics of learners by combining advanced deep learning technologies such as deep neural networks and convolutional neural networks. By comparing the differences between the learner's pronunciation and the standard pronunciation in real time, pronunciation errors can be accurately detected, thereby achieving highly accurate pronunciation correction. Combining machine learning with data analysis, the device can dynamically adjust the pronunciation training strategy according to the learner's personalized needs and provide customized pronunciation guidance and suggestions. This personalized feedback helps learners improve their pronunciation errors in a targeted manner and improve learning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0057] Figure 1 A schematic diagram of the structure of a real-time pronunciation correction device for English learners provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0059] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0060] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0061] Example

[0062] S1: Use a microphone to collect learners’ real-time pronunciation data, perform noise filtering and speech enhancement, and use a deep neural network to process background noise.

[0063] Preferably, the speech signal is divided into frames, converted to the frequency domain, the noise spectrum of each frame is estimated, the noise spectrum is subtracted from the speech spectrum to obtain a purified speech spectrum, and the purified spectrum is inversely transformed to restore the time domain signal:

[0064] S(f, t) = X(f, t) - N(f, t)

[0065] Where S(f, t) is the denoised speech signal spectrum, X(f, t) is the noisy speech spectrum, and N(f, t) is the estimated noise spectrum.

[0066] During the training process, a noisy speech signal is input and the speech signal is recovered from the noisy signal; the noisy speech signal is represented as a feature vector, multiple hidden layers are used, features are extracted layer by layer through a nonlinear activation function, and the denoised speech signal is output; during the training process, the difference between the predicted speech signal and the real speech signal is minimized by optimizing the loss function.

[0067] Furthermore, the input noisy speech signal is preprocessed, including the following steps:

[0068] The input noisy speech signal x(t) is divided into frames according to a certain frame length. Each frame contains a time window of a certain length, such as 20ms, and there is a 50% overlap between each frame.

[0069] Perform short-time Fourier transform on the signal x(t) of each frame to obtain the spectrum representation of the frame:

[0070]

[0071] Where X(f,t) is the spectrum of the noisy speech signal;

[0072] Estimate the noise spectrum N(f,t). By training a deep neural network, the device can estimate the noise spectrum of each frame based on the input noisy speech signal. This process can be achieved through the following steps:

[0073] The noisy speech signal X(f,t) is converted into a feature vector, usually using Mel-frequency cepstral coefficients or other frequency domain features as input.

[0074] We use a multi-layer deep neural network to estimate the noise spectrum N(f,t). Each hidden layer of the network extracts features layer by layer through a nonlinear activation function and outputs the noise spectrum of each frame.

[0075] Assume that the input of the network is the feature vector of the noisy speech signal, and the output of the network is the noise spectrum of each frame:

[0076] N(f, t) = DNN(X(f, t))

[0077] By training a deep neural network, the model can learn an estimate of the noise spectrum;

[0078] After obtaining the noise spectrum, the noise can be removed by spectral subtraction:

[0079] S(f, t) = X(f, t) - N(f, t)

[0080] Among them, S(f,t) is the speech spectrum after denoising;

[0081] The denoised speech spectrum S(f,t) is restored to the time domain signal through inverse short-time Fourier transform. During the training process, the device uses the noisy speech signal and the corresponding denoising target (the real clean speech signal) as input and label;

[0082] The noisy speech signal x(t) and its corresponding clean speech signal y(t) are framed and STFTed to obtain a spectral feature vector; the spectrum of the noisy speech signal is input into a deep neural network, and the difference between the predicted denoised speech signal and the true clean speech signal is minimized by optimizing the loss function.

[0083] S2: Use MFCC to extract the learner’s pronunciation features.

[0084] Preferably, the continuous speech signal is divided into several short time frames, and the signal of each frame is windowed by a window function:

[0085]

[0086] Where N is the length of the frame and w(n) is the windowing function;

[0087] Each frame of windowed signal is converted to the frequency domain through fast Fourier transform to obtain the spectrum of the frame. Through FFT, the frequency components of the signal are obtained:

[0088]

[0089] Among them, X(f) is the frequency domain signal, and x(n) is the windowed time domain signal;

[0090] The spectrum is processed by the Mel filter bank, the frequency range is mapped to the Mel scale, and the Mel spectrum is transformed using DCT to obtain the final MFCC features. The DCT formula is:

[0091]

[0092] Where Cn is the nth MFCC coefficient, M is the number of Mel bands, and Sm(t) is the energy value of the Mel band.

[0093] Furthermore, suppose there is a speech signal x(t) of length L, and the length of each frame is set to N = 256 sampling points, and the frame shift is 128 sampling points. A Hamming window is used for windowing, and FFT calculation is performed to obtain the spectrum, and 20 Mel filters are used to process the spectrum, and finally DCT transformation is performed to obtain 12 MFCC features.

[0094] The signal x(t) is divided into several frames, each frame length is N = 256, and there are 128 overlapping sampling points between each frame. Each frame signal is subjected to Hamming windowing processing, and each frame of the windowed signal is subjected to fast Fourier transform to obtain the frequency domain representation X(f) of each frame;

[0095] Use 20 Mel filters to process the spectrum X(f), map the frequency range to the Mel scale, and obtain the energy value Sm(t) of the Mel band; perform DCT transform on the energy value of the Mel band to obtain 12 MFCC coefficients C1, C2, …, C12; finally, the output 12 MFCC features can be used for subsequent speech analysis tasks, such as speech recognition, speaker recognition, etc.

[0096] S3: Build an English pronunciation model and use convolutional neural networks to learn and simulate standard pronunciation.

[0097] Preferably, the input layer receives a two-dimensional matrix containing MFCC features, and the extraction process of each MFCC coefficient generates a feature matrix of shape (T, F), where T is the number of time frames and F is the number of MFCC coefficients. Spatial local features are extracted from the input features, and each convolution kernel captures a local pattern in the signal through a convolution operation;

[0098] Then, the local time-frequency features are extracted from the input feature matrix through the convolution operation. The input is set to X, the convolution kernel is set to W, and the convolution operation formula is as follows:

[0099]

[0100] Among them, Y(t,f) is the feature map output by the convolution operation, W(i,j) is the convolution kernel with size k×k, t and f are the indices of time frame and frequency,

[0101] The convolution operation extracts local audio patterns from the input features by sliding the convolution kernel.

[0102] Further, assume that the input matrix X is an MFCC feature matrix extracted from a speech signal, with a shape of (T=15, F=13). That is, each frame of the signal has 13 MFCC coefficients and the number of time frames is 15.

[0103]

[0104] Select a 3×3 convolution kernel W to extract local time-frequency features:

[0105]

[0106] When performing a convolution operation, the convolution kernel W starts from the upper left corner of the input matrix X and slides in sequence, calculating the weighted sum of the convolution kernel and the corresponding input area each time. For example, for the position t=1, f=1, the calculation of the convolution operation is:

[0107] Y(1,1)=X(1,1)·W(0,0)+X(1,2)·W(0,1)+X(1,3)·W(0,2)+X(2,1)·W(1,0)+X(2,2)·W(1,1)+X(2,3)·W(1,2)+X(2,1)+X(3,1)·W(2,0)+X(3,2)·W(2,1)+X(3,3)·W(2,2), and so on. Through the convolution operation, the output feature map Y is generated on the entire feature matrix X.

[0108] After the convolution operation, the output feature map Y will have different time-frequency patterns. For example, for k = 3, assuming that after convolution, a feature map of size (13, 11) is obtained:

[0109]

[0110] The feature map Y contains the time-frequency local features extracted by the convolution kernel, which represents the local patterns of different time frames and frequency bands in the speech signal.

[0111] S4: By comparing the learner's pronunciation with the standard pronunciation in real time, it detects whether there are pronunciation errors. It combines the acoustic model with the language model and uses the similarity measurement method to perform error assessment.

[0112] Preferably, the pronunciation errors are detected by calculating the similarity between the learner's pronunciation and the standard pronunciation, and the cosine angle between the learner's pronunciation feature vector and the standard pronunciation feature vector is calculated to measure the similarity between the two. The formula is:

[0113]

[0114] Among them, A and B are two pronunciation feature vectors, ||A|| and ||B|| are their norms respectively;

[0115] According to the comparison result, a threshold is set, and if the similarity is lower than the threshold, a pronunciation error is detected.

[0116] Preferably, the extracted MFCC features are input into the acoustic model, the pronunciation is decoded using a decoding algorithm, and a candidate speech unit sequence is output. The acoustic model outputs the phoneme prediction at each time to form a time series A=(a1, a2, ..., aT), where T is the number of speech frames and aT is the phoneme prediction at the Tth time;

[0117] The language model evaluates whether the learner's pronunciation conforms to the grammatical and semantic rules of the language, and judges its rationality by calculating the probability of the learner's pronunciation; for the input phoneme sequence A = (a1, a2, ..., aT), the language model calculates the probability P(A) of the sequence.

[0118] Furthermore, assuming that the learner is practicing pronouncing the word "bat", the device works as follows: the learner reads out "bat", the audio signal is received by the device and converted into an MFCC feature matrix; through the acoustic model, the MFCC features are decoded into phoneme sequences, such as "b", "ae", and "t"; the language model receives the phoneme sequence ("b", "ae", "t") and calculates the generation probability of the sequence; if the probability of the sequence is high, the device considers the pronunciation correct and returns a no error prompt; otherwise, if the pronunciation does not conform to the standard phoneme sequence, a pronunciation error prompt is returned; assuming that the language model probability is low, the device feedback "Your pronunciation 'ae' is inconsistent with the standard pronunciation and should be changed to . " and prompt to continue practicing.

[0119] S5: Generate correction feedback based on real-time detection results, provide audio feedback or text prompts, and point out problems in pronunciation.

[0120] Preferably, when the learner makes a pronunciation error, the device plays a short prompt tone to remind the learner to pay attention to the pronunciation problem;

[0121] For incorrectly pronounced parts, the device plays audio of standard pronunciation for learners to compare with. The device plays the correct pronunciation demonstration word by word and points out the location of the learner's pronunciation deviation.

[0122] If the learner makes mistakes in intonation or rhythm, the device gives suggestions for adjustments through audio feedback.

[0123] Preferably, when it is detected that the learner has pronounced a certain phoneme incorrectly, a text prompt points out the incorrect phoneme and gives the correct pronunciation; if the error occurs in a certain word or phrase, the device displays the incorrect word on the screen and highlights or underlines the incorrect part;

[0124] For grammatical or phonetic deviations, the device provides grammatical or language regularity hints.

[0125] Further, assuming that the learner is practicing pronouncing the word "bat", the device works as follows:

[0126] The learner uses a microphone to record the pronunciation, "bat". The device decodes the learner's voice through the acoustic model and obtains a phoneme sequence (such as "b", "ae", "t"). The language model calculates the language rationality of the sequence and finds that the learner's "ae" phoneme pronunciation is not standard;

[0127] The device plays a short beep ("beep") to remind the learner that there is a pronunciation problem. The device plays the audio of the standard pronunciation of "bat" and plays it word by word, emphasizing the deviation between the standard phoneme "ae" and the learner's pronunciation. "The pronunciation of 'ae' in 'bat' is too low and should be adjusted to sound. ";

[0128] If the learner has problems with intonation, the device detects that the learner's intonation is too flat by analyzing the pitch and rhythm, and the device plays audio feedback: "The pronunciation of 'bat' should have an upward intonation. Please increase the intonation appropriately and slow down the speaking speed." After the learner adjusts the pronunciation based on the feedback, re-records and submits it, the device detects the pronunciation again; if the pronunciation is correct, the device plays a "correct" prompt sound (such as a "ding" sound effect) to encourage the learner to continue practicing.

[0129] Furthermore, assuming that the learner is practicing pronouncing the word "necessary", the device works as follows:

[0130] The learner uses a microphone to record the pronunciation of "necessary", and the device converts it into a feature vector through feature extraction;

[0131] Through the acoustic model, the device detects the phoneme sequence pronounced by the learner (such as "n", "ε", "s", "r", "i"); the device found that the learner's pronunciation did not follow the standard pronunciation rules (for example, the stress position was wrong and should be on the first syllable); the device compared the learner's pronunciation with the standard pronunciation and found that the learner did not put the stress on the "ne" syllable correctly when pronouncing;

[0132] The device displays a text prompt on the screen: "You pronounced the 'n' correctly, but the 'a' part is not pronounced correctly. The correct pronunciation should be At the same time, the device plays the audio of the correct pronunciation; the word "necessary" is displayed on the screen, where the "ne" part is highlighted in bold or underlined, instructing learners to pay special attention to the pronunciation of this syllable;

[0133] The device pointed out through text and voice prompts: "The stress on the first syllable of 'necessary' should be pronounced as Rather than ”;

[0134] After the learner hears the standard pronunciation, he or she tries to record his or her pronunciation again. When the device detects that the pronunciation is correct, it plays a "correct" prompt tone to encourage the learner to continue practicing.

[0135] S6: Track learners’ progress by recording each pronunciation error and improvement.

[0136] Preferably, the learner's overall pronunciation improvement is regularly evaluated and an overall evaluation report is given every week, which includes the change in the number of errors and the improvement in pronunciation accuracy;

[0137] By counting the error patterns of each learner, and providing improvement plans based on the problems.

[0138] Furthermore, assuming that the learner often makes mistakes in pronouncing the word “necessary”, how can the device conduct regular assessment and generate improvement plans:

[0139] The learner pronounced "necessary" and the device detected pronunciation errors. The main errors occurred in the pronunciation and stress position of the phonemes. The specific errors included: the learner failed to pronounce the / ε / sound accurately; the stress was not placed on the correct syllable (it should be on the first syllable "ne");

[0140] The device records the learner's pronunciation errors and counts the pronunciation errors of "necessary" in each practice. By comparing the learner's pronunciation data over the past week, the device identifies "necessary" as a high-frequency error word, with errors in vowel pronunciation and stress position.

[0141] The device calculated that there were 15 errors this week, a 10% decrease from last week. This week, the learner's overall pronunciation accuracy increased from 80% to 85%. The device showed that "necessary" was the word with the most errors this week, especially in terms of pronunciation syllables and stress. To improve the pronunciation of "necessary", the device suggested that learners focus on practicing the pronunciation of the vowel / ε / and pay special attention to the stress position in the word. The device will recommend similar words, such as the / ε / sound in "necessary", for special training.

[0142] S7: Adjust pronunciation training strategies and provide pronunciation guidance based on learners’ progress and learning style through machine learning and data analysis.

[0143] Preferably, the device adjusts the feedback method according to the learner's feedback preference. For learners who like to hear the comparison between their own pronunciation and the standard pronunciation, the device will provide multi-audio feedback; for learners who like to correct themselves through text prompts, the device will provide multi-pronunciation suggestions through text;

[0144] For learners who progress quickly, the device adjusts the training frequency and reduces the time of a single exercise; for learners who progress slowly, the device extends the exercise time and increases the frequency of exercise.

[0145] Further, the learner began to practice pronouncing the word "necessary". The device detected that the learner failed to pronounce the vowels correctly. and there is a deviation in the stress;

[0146] The device first plays the learner's pronunciation, and then plays the standard pronunciation, allowing the learner to compare the differences between the two; then, the device plays a standard word pronunciation demonstration;

[0147] The device displays an error message on the screen: "Error message: The correct pronunciation is Please place your tongue in the middle of your mouth, with the tip of your tongue slightly below your upper teeth. "; The incorrect word "necessary" is highlighted in the text, and the incorrect part is highlighted, pointing out: "The stress should be on the syllable 'ne', not on the syllable 'ces'.";

[0148] The learner's pronunciation accuracy has improved by 10% in the past week. Therefore, the device decided to reduce the single training time from 30 minutes to 20 minutes, and focus the practice on complex pronunciation tasks. For learners who are progressing slowly, the device will increase the training time to 40 minutes during the next practice, and increase the frequency of repeated practice. The device automatically generates an evaluation report every week to show the learner's pronunciation improvement. For example: "This week you have made significant progress in the pronunciation of the word 'necessary', with a 10% reduction in the error rate." For learners who are progressing slowly, the device reminds: "You still have a large deviation in the pronunciation of vowels. It is recommended to increase the pronunciation practice time."

[0149] In the description of the present invention, it should be noted that the terms “first”, “second” and “third” are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0151] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0152] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0154] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0155] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

Claims

1. A real-time pronunciation correction device for English learners, characterized in that: include: The acquisition and preprocessing module is used to collect learners' real-time pronunciation data using microphones, perform noise filtering and speech enhancement, and use deep neural networks to process background noise; Speech feature extraction module, used to extract learner's pronunciation features using MFCC; The standard pronunciation model building module is used to build an English pronunciation model and use convolutional neural networks to learn and simulate standard pronunciation; The data customization management unit is used to detect whether there are pronunciation errors by comparing the learner's pronunciation with the standard pronunciation in real time, and to conduct error assessment using a similarity measurement method by combining the acoustic model with the language model; The pronunciation correction suggestion module is used to generate correction feedback based on real-time detection results, provide audio feedback or text prompts, and point out problems in pronunciation; Correction and Progress Tracking module, used to track learners’ progress by recording every pronunciation error and improvement; The adjustment and optimization module is used to adjust pronunciation training strategies and provide pronunciation guidance based on learners' progress and learning style through machine learning and data analysis.

2. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The method uses a microphone to collect the learner's real-time pronunciation data, performs noise filtering and speech enhancement, and uses a deep neural network to process background noise, including: Divide the speech signal into frames, convert it to the frequency domain, estimate the noise spectrum of each frame, subtract the noise spectrum from the speech spectrum to obtain the purified speech spectrum, perform inverse transform on the purified spectrum, and restore the time domain signal: S(f,t)=X(f,t)-N(f,t) Where S(f, t) is the denoised speech signal spectrum, X(f, t) is the noisy speech spectrum, and N(f, t) is the estimated noise spectrum. During the training process, a noisy speech signal is input and the speech signal is recovered from the noisy signal; the noisy speech signal is represented as a feature vector, multiple hidden layers are used, features are extracted layer by layer through a nonlinear activation function, and the denoised speech signal is output; during the training process, the difference between the predicted speech signal and the real speech signal is minimized by optimizing the loss function.

3. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The method of using MFCC to extract the learner's pronunciation features includes: The continuous speech signal is divided into several short time frames, and the signal of each frame is windowed by a window function: Where N is the length of the frame and w(n) is the windowing function; Each frame of windowed signal is converted to the frequency domain through fast Fourier transform to obtain the spectrum of the frame. Through FFT, the frequency components of the signal are obtained: Among them, X(f) is the frequency domain signal, and x(n) is the windowed time domain signal; The spectrum is processed by the Mel filter bank, the frequency range is mapped to the Mel scale, and the Mel spectrum is transformed using DCT to obtain the final MFCC features. The DCT formula is: Where Cn is the nth MFCC coefficient, M is the number of Mel bands, and Sm(t) is the energy value of the Mel band.

4. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The English pronunciation model is constructed by using a convolutional neural network to learn and simulate standard pronunciation, including: The input layer receives a two-dimensional matrix containing MFCC features. The extraction process of each MFCC coefficient generates a feature matrix of shape (T, F), where T is the number of time frames and F is the number of MFCC coefficients. Spatial local features are extracted from the input features. Each convolution kernel captures the local pattern in the signal through the convolution operation. Then, the local time-frequency features are extracted from the input feature matrix through the convolution operation. The input is set to X, the convolution kernel is set to W, and the convolution operation formula is as follows: Among them, Y(t,f) is the feature map output by the convolution operation, W(i,j) is the convolution kernel with size k×k, t and f are the indices of time frame and frequency, The convolution operation extracts local audio patterns from the input features by sliding the convolution kernel.

5. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The method of detecting whether there is a pronunciation error by comparing the learner's pronunciation with the standard pronunciation in real time includes: By calculating the similarity between the learner's pronunciation and the standard pronunciation, pronunciation errors are detected. The cosine angle between the learner's pronunciation feature vector and the standard pronunciation feature vector is calculated to measure the similarity between the two. The formula is: Among them, A and B are two pronunciation feature vectors, ||A|| and ||B|| are their norms respectively; According to the comparison result, a threshold is set, and if the similarity is lower than the threshold, a pronunciation error is detected.

6. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The method of combining the acoustic model with the language model and using a similarity measurement method to perform error assessment includes: The extracted MFCC features are input into the acoustic model, the pronunciation is decoded using a decoding algorithm, and a candidate speech unit sequence is output. The acoustic model outputs the phoneme prediction at each moment to form a time series A = (a1, a2, ..., aT), where T is the number of speech frames and aT is the phoneme prediction at the Tth moment; The language model evaluates whether the learner's pronunciation conforms to the grammatical and semantic rules of the language, and judges its rationality by calculating the probability of the learner's pronunciation; for the input phoneme sequence A = (a1, a2, ..., aT), the language model calculates the probability P(A) of the sequence.

7. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: Based on the real-time detection results, correction feedback is generated, audio feedback or text prompts are provided, and problems in pronunciation are pointed out, including: When the learner makes a pronunciation error, the device plays a short prompt tone to remind the learner to pay attention to the pronunciation problem; For incorrectly pronounced parts, the device plays audio of standard pronunciation for learners to compare with. The device plays the correct pronunciation demonstration word by word and points out the location of the learner's pronunciation deviation. If the learner makes mistakes in intonation or rhythm, the device gives suggestions for adjustments through audio feedback.

8. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The method of generating correction feedback based on the real-time detection result, providing audio feedback or text prompts, and pointing out problems in pronunciation also includes: When it is detected that the learner has pronounced a certain phoneme incorrectly, a text prompt points out the wrong phoneme and gives the correct pronunciation; if the error occurs in a word or phrase, the device displays the wrong word on the screen and highlights or underlines the wrong part; For grammatical or phonetic deviations, the device provides grammatical or language regularity hints.

9. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: The learner's progress is tracked by recording each pronunciation error and improvement, including: Regularly evaluate learners' overall pronunciation improvement and give an overall evaluation report every week, which includes the change in the number of errors and the improvement in pronunciation accuracy; By counting the error patterns of each learner, and providing improvement plans based on the problems.

10. The real-time pronunciation correction device for English learners according to claim 1, characterized in that: According to the learner's progress and learning style, the pronunciation training strategy is adjusted through machine learning and data analysis to provide pronunciation guidance, including: The device adjusts the feedback method according to the learner's feedback preference. For learners who like to hear the comparison between their own pronunciation and the standard pronunciation, the device will provide multi-audio feedback; for learners who like to correct themselves through text prompts, the device provides multi-pronunciation suggestions through text; For learners who progress quickly, the device adjusts the training frequency and reduces the time of a single exercise; for learners who progress slowly, the device extends the exercise time and increases the frequency of exercise.

Citation Information

Cited By

  • Language learning assisting method and device

    CN120708650A

  • English pronunciation correction method using AI speech recognition

    CN120783751A

  • English pronunciation correction method using AI voice recognition

    CN120783751B