English pronunciation error correction system and device based on artificial intelligence

Through an English pronunciation error correction system based on artificial intelligence, combined with multi-dimensional feedback and precise pronunciation evaluation, the problem that existing devices cannot correct accent details is solved, and all-round pronunciation error correction and personalized exercises are achieved, improving error correction efficiency.

CN120279946AInactive Publication Date: 2025-07-08陈海贝
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510563314.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing English pronunciation error correction device cannot correct students' accent details, and cannot make in-depth corrections, so it can only store errors.

Method used

The English pronunciation error correction system based on artificial intelligence is adopted, including speech acquisition, speech recognition, pronunciation evaluation, error classification and error correction feedback modules, and error correction is corrected through multi-dimensional visual, auditory and text feedback, and combined with clarity score, acoustic model and phoneme alignment algorithm for accurate error correction.

Benefits of technology

It realizes all-round error correction for students' pronunciation, can timely detect the clarity of accents, help students discover the source of errors and conduct targeted training, and improve error correction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279946A_ABST
    Figure CN120279946A_ABST
Patent Text Reader

Abstract

The invention discloses an English pronunciation error correction system and device based on artificial intelligence, and relates to the field of English pronunciation error correction. The system comprises a voice acquisition module, a voice recognition module, a voice program definition judgment module, a pronunciation evaluation module, a voice classification recognition module, an error classification module, a voice processing module, a voice processing module, a voice processing module and a voice processing module, wherein the voice of a user is input and recorded through an audio acquisition module; the voice program definition judgment module judges voice program definition through a definition comprehensive scoring formula; according to the English pronunciation training system, classified input voice features are compared with standard voice features to carry out detailed error evaluation, an error correction feedback module carries out multi-dimensional integrated error correction feedback through multi-dimensional vision, hearing and characters, and a user receiving module carries out evaluation and personalized practice on English pronunciation by setting an interactive interface. According to the system, errors caused by insufficient definition and error types of pronunciation of students are prevented, and the students are helped to find error sources better and faster and correct the error sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of English pronunciation error correction, and specifically to an English pronunciation error correction system and device based on artificial intelligence. Background Art

[0002] Currently, the number of global English learners has exceeded 2 billion, and the demand for language communication has increased sharply. However, traditional classrooms are restricted by teachers and the environment, and oral training often becomes a mere formality. Pronunciation deviations not only lead to communication barriers but also weaken learning confidence, forming a vicious cycle. Traditional error correction relies on subjective feedback from teachers, which has problems of delay and inconsistent standards. Therefore, more and more English pronunciation error correction systems are used to guide students in deeper English learning and correction.

[0003] Existing English pronunciation error correction devices can only simply point out students' pronunciation errors, unable to correct the details of students' accents, and unable to conduct subsequent in-depth correction for students, only being able to store students' errors simply. In view of the above situation, the present invention provides an English pronunciation error correction system and device based on artificial intelligence to solve the above problems. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides an English pronunciation error correction system and device based on artificial intelligence, which solves the problem that existing English pronunciation error correction devices can only simply point out students' pronunciation errors, unable to correct the simple accent details of students, and unable to conduct subsequent in-depth correction for students, only being able to store students' errors.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: an English pronunciation error correction system and device based on artificial intelligence, the system includes;

[0006] A voice collection module, which records the input of the user's voice through an audio collection module;

[0007] A voice recognition module, which judges the clarity of the voice program through a clarity comprehensive scoring formula;

[0008] A pronunciation evaluation module, which further classifies and recognizes the voice through acoustic model analysis of voice features;

[0009] An error classification module, which conducts a detailed error assessment by comparing the classified input voice features with standard voice features, and determines and corrects errors in sequence according to the order of phonemes, stress, and intonation;

[0010] An error correction feedback module, which conducts multi-dimensional integrated error correction feedback through multi-dimensional vision, audition, and text;

[0011] The user reception module evaluates and conducts personalized practice on English pronunciation by setting up an interactive interface.

[0012] Preferably, the voice collection module collects English pronunciation by setting up a high-sensitivity microphone and further removes background noise from the collected sound using a noise reduction algorithm.

[0013] Preferably, the clarity comprehensive scoring formula is

[0014] α = 0.4SNR + 0.2(1 - ZCR) + 0.3STE + 0.1(1 - F0CV)

[0015] Where,

[0016] α is the clarity evaluation value;

[0017] SNR is the signal-to-noise ratio value, that is, the numerical ratio of the voice signal power to the background noise power;

[0018] ZCR is the zero-crossing rate, that is, the number of times the signal crosses zero per second;

[0019] STE is the short-time energy, that is, the energy value of each 20ms frame;

[0020] F0 - CV is the fundamental frequency variation coefficient, that is, the numerical ratio of the fundamental frequency standard deviation to the fundamental frequency mean;

[0021] Among them, the numerical values of each symbol are normalized, and each parameter needs to be linearly normalized to the interval [0, 1];

[0022] Where,

[0023] When α is greater than 0.75, it is determined to be clear and the next pronunciation evaluation is carried out;

[0024] When α is lower than 0.45, it is determined to be unclear and re-recording is required.

[0025] Preferably, the pronunciation evaluation module analyzes the voice and collects each feature through an acoustic model and a phoneme alignment algorithm for the next error classification.

[0026] Preferably, in the error classification module, the phoneme evaluation is first carried out, and the phoneme evaluation formula is:

[0027]

[0028] Where,

[0029] X is the phoneme evaluation value;

[0030] DTW is the dynamic time warping algorithm, which is used to calculate the acoustic feature distance between the user and the standard phoneme;

[0031] D_user is the sequence of acoustic feature of phonemes pronounced by the user;

[0032] D_standard is the sequence of acoustic feature of standard phonemes;

[0033] len is the sequence length, i.e., the number of frames;

[0034] Among them,

[0035] When X is greater than 0.85, it is a standard phoneme pronunciation and no correction is needed;

[0036] When X is less than 0.85, it is an inaccurate phoneme pronunciation and correction is needed.

[0037] Preferably, in the error classification module, the stress is evaluated next. The evaluation formula for the stress is:

[0038]

[0039] Among them,

[0040] Y is the stress evaluation value;

[0041] E_user is the peak position of stress energy of the user's pronunciation;

[0042] E_standard is the peak position of stress energy of the standard pronunciation;

[0043] Among them,

[0044] When Y is greater than 0.7, the stress position is correct and no correction is needed;

[0045] When Y is less than 0.7, the stress position is incorrect and correction is needed.

[0046] Preferably, in the error classification module, the intonation is evaluated last. The evaluation formula for the intonation is:

[0047] Z = P(F_user,F_standard)

[0048] Among them,

[0049] Z is the intonation evaluation value;

[0050] P is the Pearson correlation coefficient, which calculates the linear correlation of two groups of fundamental frequency sequences;

[0051] F_user is the peak position of stress energy of the user's pronunciation;

[0052] F_standard is the peak position of stress energy of the standard pronunciation;

[0053] Among them,

[0054] When Z is greater than 0.6, it is the intonation matching standard and no error correction is required;

[0055] When Z is less than 0.6, it is the intonation matching deviation and error correction is required.

[0056] Preferably, the error correction feedback module corrects errors by simultaneously setting three feedback methods: visual, auditory, and text;

[0057] Among them, the visual feedback includes displaying a waveform diagram and a spectrogram for marking the error position;

[0058] The auditory feedback is provided by playing a comparison between the standard pronunciation and the user's pronunciation;

[0059] The text feedback provides detailed pronunciation improvement suggestions.

[0060] Preferably, the error correction device includes an audio acquisition module, a calculation unit, and a feedback output device;

[0061] Among them, the audio acquisition module includes:

[0062] A high-sensitivity microphone array for directional sound collection to suppress ambient noise;

[0063] A sound card, which is an ADC converter for 24-bit analog-to-digital conversion;

[0064] A preprocessing circuit for real-time hardware noise reduction;

[0065] Among them, the calculation unit includes an edge computing device and a coprocessor;

[0066] Among them, the feedback output device includes:

[0067] A visualization screen for displaying learning content, spectrograms, and tongue position animations;

[0068] Bone conduction headphones for playing the standard sound and the user's pronunciation in real-time comparison;

[0069] A force feedback device for simulating the oral airflow during pronunciation.

[0070] The present invention discloses an English pronunciation error correction system and device based on artificial intelligence, and the beneficial effects thereof are as follows:

[0071] 1. The English pronunciation error correction system and device based on artificial intelligence pre-judge the English accent by setting a clarity evaluation before error correction, timely detect whether the accent is clear, and prevent subsequent pronunciation judgments from being incorrect due to insufficient clarity.

[0072] 2. The English pronunciation error correction system and device based on artificial intelligence helps students correct pronunciation in all aspects by classifying and detecting the input pronunciation, points out the types of pronunciation errors of students, and helps students better and faster discover the sources of errors and correct them.

[0073] 3. The English pronunciation error correction system and device based on artificial intelligence enables students to conduct targeted training according to the common mistakes made in the pronunciation error correction process through artificial intelligence-assisted personalized practice, improving the error correction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0075] Figure 1 It is a schematic diagram of the system of the present invention;

[0076] Figure 2 It is a general diagram of error correction feedback of the present invention;

[0077] Figure 3 It is a schematic diagram of logical judgment of the present invention;

[0078] Figure 4 It is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0080] By providing an English pronunciation error correction system and device based on artificial intelligence, the embodiments of the present application solve the problem that the existing English pronunciation error correction devices can only simply point out the pronunciation errors of students, cannot correct the simple accent details of students, and cannot perform subsequent in-depth correction on students, and can only store the errors of students.

[0081] To better understand the above technical solutions, the following will describe the above technical solutions in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0082] The embodiments of the present invention disclose an English pronunciation error correction system and device based on artificial intelligence.

[0083] According to the attached Figures 1-4 shown,

[0084] The system includes:

[0085] A voice collection module that records the user's voice input through a microphone. The voice collection module collects English pronunciations by setting an audio collection module and further removes background noise in the collected sound using a noise reduction algorithm. The audio collection modules used are a high-sensitivity microphone array and a sound card, which are used for directional sound collection and suppressing ambient noise; the sound card is an ADC converter for 24-bit analog-to-digital conversion; here, a preprocessing circuit is used for real-time hardware noise reduction.

[0086] A voice recognition module that judges the clarity of the voice program through a clarity comprehensive scoring formula. The clarity comprehensive scoring formula is

[0087] α = 0.4SNR + 0.2(1 - ZCR) + 0.3STE + 0.1(1 - F0CV)

[0088] where,

[0089] α is the clarity evaluation value;

[0090] SNR is the signal-to-noise ratio value, that is, the numerical ratio of the voice signal power to the background noise power;

[0091] ZCR is the zero-crossing rate, that is, the number of times the signal crosses zero per second;

[0092] STE is the short-time energy, that is, the energy value of each 20ms frame;

[0093] F0-CV is the fundamental frequency variation coefficient, that is, the numerical ratio of the fundamental frequency standard deviation to the fundamental frequency mean;

[0094] where, the numerical values of each symbol are normalized, and each parameter needs to be linearly normalized to the [0, 1] interval;

[0095] where,

[0096] When α is greater than 0.75, it is judged as clear and the next pronunciation evaluation is carried out;

[0097] When α is lower than 0.45, it is judged as unclear and re-recording is required;

[0098] For voice recognition, methods such as Google Speech-to-Text and Whisper are used for support, and deep learning and further judgment are carried out through deep learning models Transformer and RNN technologies.

[0099] The pronunciation evaluation module analyzes speech features through an acoustic model for further speech classification and recognition. The pronunciation evaluation module uses an acoustic model and a phoneme alignment algorithm to parse speech and collect various features for the next step of error classification. Among them, pronunciation features are recognized using methods such as Kaldi and DeepSpeech. By establishing acoustic models such as HMM and DNN and using phoneme alignment algorithm technology as a technical solution, a larger-scale and diverse speech dataset is used to train the model, and transfer learning is introduced. The pre-trained models Whisper and DeepSpeech are used to improve AI performance to assist students in further pronunciation evaluation.

[0100] The error classification module conducts a detailed error assessment by comparing the classified input speech features with the standard speech features, and makes sequential judgments and corrections according to the order of phonemes, stress, and intonation.

[0101] In the error classification module, the evaluation of phonemes is carried out first. The evaluation formula for phonemes is:

[0102]

[0103] Among them,

[0104] X is the phoneme evaluation value;

[0105] DTW is the dynamic time warping algorithm, which is used to calculate the acoustic feature distance between the user and the standard phoneme;

[0106] D_user is the phoneme acoustic feature sequence of the user's pronunciation;

[0107] D_standard is the acoustic feature sequence of the standard phoneme;

[0108] len is the sequence length, that is, the number of frames;

[0109] Among them,

[0110] When X is greater than 0.85, the phoneme pronunciation is standard and no correction is required;

[0111] When X is less than 0.85, the phoneme pronunciation is inaccurate and correction is required;

[0112] In the error classification module, the evaluation of stress is carried out secondly. The evaluation formula for stress is:

[0113]

[0114] Among them,

[0115] Y is the stress evaluation value;

[0116] E_user is the peak position of the stress energy of the user's pronunciation;

[0117] E_standard is the peak position of the stress energy of the standard pronunciation;

[0118] Among them,

[0119] When Y is greater than 0.7, the stress position is correct and no correction is needed;

[0120] When Y is less than 0.7, the stress position is incorrect and correction is needed.

[0121] In the error classification module, the intonation is finally evaluated. The evaluation formula for intonation is:

[0122] Z = P(F_user, F_standard)

[0123] Among them,

[0124] Z is the intonation evaluation value;

[0125] P is the Pearson correlation coefficient, which calculates the linear correlation of two groups of fundamental frequency sequences;

[0126] F_user is the peak position of the stress energy of the user's pronunciation;

[0127] F_standard is the peak position of the stress energy of the standard pronunciation;

[0128] Among them,

[0129] When Z is greater than 0.6, it is the intonation matching standard and no correction is needed;

[0130] When Z is less than 0.6, it is the intonation matching deviation and correction is needed;

[0131] Among them, the error classification uses the Montreal Forced Aligner method for classification, and is technically supported by machine learning classification algorithms such as SVM and random forest. It is sorted according to the importance of phonemes, stress, and intonation. The most important phonemes for spoken language are recognized first. If the phonemes are incorrect, correction feedback is directly generated, allowing students to correct the phonemes first. When the phonemes are correct, the next stress judgment can be made. Finally, it is the intonation. Through this meticulous correction, students who are completely wrong can master the pronunciation of phonemes first. For students who have a good grasp of phoneme pronunciation, their spoken language skills are further refined and improved. Further correction and learning are carried out in the order of stress and intonation, making their spoken language better and better.

[0132] The error correction feedback module provides multi-dimensional error correction feedback through vision, hearing, and text. The error correction feedback module corrects errors by setting three feedback methods: vision, hearing, and text simultaneously;

[0133] Among them,

[0134] Visual feedback includes displaying waveform diagrams and spectrograms for marking error positions, and through a visualization screen, presenting learning content, displaying spectrograms, and tongue position animations;

[0135] Auditory feedback is provided by playing a comparison between the standard pronunciation and the user's pronunciation, and through bone conduction headphones, the standard sound and the user's pronunciation are compared and played in real time;

[0136] Text feedback is provided by offering detailed pronunciation improvement suggestions;

[0137] In addition, a force feedback device is used to simulate the oral airflow during pronunciation;

[0138] Error correction feedback uses the Google Text-to-Speech and Amazon Polly methods for feedback, and natural language generation NLG and text-to-speech TTS are used for technical support;

[0139] Through various feedback methods, students are corrected in all aspects, enabling them to clearly know their error points.

[0140] The user reception module evaluates English pronunciation and conducts personalized practice by setting up an interactive interface, collects and analyzes students' completion status and data through artificial intelligence, and uses a reinforcement learning algorithm to optimize the recommendation strategy. Combining the user's historical data with real-time performance, a personalized learning path and personalized practice process are generated to help students better correct their spoken English.

[0141] The error correction device includes an audio acquisition module, a calculation unit, and a feedback output device, where the calculation unit includes an edge computing device and a coprocessor.

[0142] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An English pronunciation error correction system based on artificial intelligence, characterized in that, The system includes; A voice acquisition module that records the user's voice input through an audio acquisition module; A voice recognition module that judges the clarity of the voice program through a clarity comprehensive scoring formula; A pronunciation evaluation module that further classifies and recognizes voices by analyzing voice features through an acoustic model; An error classification module that carefully evaluates errors by comparing the classified input voice features with the standard voice features, and determines and corrects errors in sequence according to the order of phonemes, stress, and intonation; An error correction feedback module that provides multi-dimensional integrated error correction feedback through multi-dimensional vision, hearing, and text; A user reception module that evaluates English pronunciation and provides personalized practice by setting an interactive interface.

2. The English pronunciation correction system based on artificial intelligence according to claim 1, characterized in that: The voice acquisition module collects English pronunciations by setting a high-sensitivity microphone and further removes background noise in the collected sounds using a noise reduction algorithm.

3. The English pronunciation error correction system based on artificial intelligence according to claim 1, characterized in that: The clarity comprehensive scoring formula is α = 0.4SNR + 0.2(1 - ZCR) + 0.3STE + 0.1(1 - F0CV) Where, α is the clarity evaluation value; SNR is the signal-to-noise ratio value, that is, the numerical ratio of the voice signal power to the background noise power; ZCR is the zero-crossing rate, that is, the number of times the signal crosses zero per second; STE is the short-time energy, that is, the energy value of each 20ms frame; F0-CV is the fundamental frequency variation coefficient, that is, the numerical ratio of the fundamental frequency standard deviation to the fundamental frequency mean; Where, the numerical values of each symbol are normalized, and each parameter needs to be linearly normalized to the interval [0,1]; Where, When α is greater than 0.75, it is judged as clear, and the next pronunciation evaluation is carried out; When α is lower than 0.45, it is judged as unclear and needs to be re-recorded.

4. The English pronunciation error correction system based on artificial intelligence according to claim 1, wherein: The pronunciation evaluation module analyzes voices and collects various features through an acoustic model and a phoneme alignment algorithm for the next error classification.

5. The English pronunciation error correction system based on artificial intelligence according to claim 1, characterized in that: In the error classification module, first, the evaluation of phonemes is carried out. The evaluation formula of phonemes is: Where, X is the phoneme evaluation value; DTW is the dynamic time warping algorithm, which is used to calculate the acoustic feature distance between the user and the standard phoneme; D_user is the phoneme acoustic feature sequence of the user's pronunciation; D_standard is the acoustic feature sequence of the standard phoneme; len is the sequence length, that is, the number of frames; Where, When X is greater than 0.85, the phoneme pronunciation is standard and no correction is needed; When X is less than 0.85, the phoneme pronunciation is inaccurate and needs to be corrected.

6. The English pronunciation error correction system based on artificial intelligence according to claim 1, characterized in that: In the error classification module, secondly, the evaluation of stress is carried out. The evaluation formula of stress is: Where, Y is the stress evaluation value; E_user is the peak position of the stress energy of the user's pronunciation; E_standard is the peak position of the stress energy of the standard pronunciation; Where, When Y is greater than 0.7, the stress position is correct and no correction is needed; When Y is less than 0.7, the stress position is incorrect and needs to be corrected.

7. The English pronunciation error correction system based on artificial intelligence according to claim 1, characterized in that: In the error classification module, finally, the evaluation of intonation is carried out. The evaluation formula of intonation is: Z = P(F_user,F_s tandard) Where, Z is the intonation evaluation value; P is the Pearson correlation coefficient, which calculates the linear correlation of two groups of fundamental frequency sequences; F_user is the peak position of the stress energy of the user's pronunciation; F_standard is the peak position of the stress energy of the standard pronunciation; Among them, When Z is greater than 0.6, it is the intonation matching standard and no correction is required; When Z is less than 0.6, it is the intonation matching deviation and correction is required.

8. The English pronunciation correction system based on artificial intelligence according to claim 1, characterized in that: The error correction feedback module corrects errors by setting three feedback methods: visual, auditory, and text; Among them, Visual feedback includes displaying waveform diagrams and spectrograms for marking error positions; Auditory feedback is provided by playing a comparison between the standard pronunciation and the user's pronunciation; Text feedback provides detailed pronunciation improvement suggestions.

9. An English pronunciation error correction device based on artificial intelligence, comprising the English pronunciation error correction system based on artificial intelligence according to any one of claims 1-8, characterized in that: The error correction device includes an audio acquisition module, a calculation unit, and a feedback output device; Among them, the audio acquisition module includes: A high-sensitivity microphone array for directional sound collection and suppressing ambient noise; A sound card, which is an ADC converter for 24-bit analog-to-digital conversion; A preprocessing circuit for real-time hardware noise reduction; Among them, the calculation unit includes an edge computing device and a coprocessor; Among them, the feedback output device includes: A visualization screen for displaying learning content, spectrograms, and tongue position animations; Bone conduction headphones for playing the standard sound and the user's pronunciation in real-time comparison; A force feedback device for simulating oral airflow during pronunciation.