A sound production detection method and a sound production detection device
By providing personalized auditory and visual feedback to native Chinese speakers, and combining speech synthesis and multi-dimensional audio visualization, the problem of users having difficulty recognizing their own pronunciation errors is solved, achieving the effect of quickly locating and correcting pronunciation errors.
Patent Information
- Application Number
- CN202310432880.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-04-21
AI Technical Summary
Existing computer-aided pronunciation training systems have failed to effectively help users accurately recognize their own pronunciation errors, especially native Chinese speakers who are influenced by their mother tongue when learning English and have difficulty perceiving pronunciation errors. Furthermore, the systems lack clear teaching interventions and personalized feedback.
This invention provides a pronunciation detection method that, through mispronunciation detection and diagnosis, combined with auditory and visual feedback, performs comparative speech analysis on common error types made by native Chinese speakers when learning English, and generates personalized audiovisual feedback results, including standard pronunciation synthesis, error reconstruction audio, and multi-dimensional audio visualization, to help users understand and correct their errors.
It can quickly locate users' high-frequency errors in a short period of time, improve users' awareness of their own pronunciation problems, extend the interaction path of the computer-aided pronunciation training system, and ensure that users can accurately locate and correct pronunciation errors.
Smart Images

Figure CN116403607B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer-aided pronunciation training, and in particular to a pronunciation detection method and a pronunciation detection device. BACKGROUND
[0002] As the foundation of human intelligence and culture, language not only helps the communication between individuals, but also promotes the accumulation and dissemination of knowledge. Mastering a second language has gradually become a basic need for human beings. In the process of second language acquisition, some learners have not received systematic pronunciation training due to limited conditions, and it is difficult for them to perceive pronunciation errors due to the influence of mother tongue transfer. With the maturity of human-computer voice interaction technology based on deep learning, computer-aided pronunciation training (CAPT) systems effectively alleviate the problem of time and space resources, and have been widely applied to vocabulary learning, oral English learning and other products. However, the current system usually uses standard English phonetic symbols and general English teaching materials to explain words to users and provide correct pronunciation. The knowledge learned by users of different levels is completely consistent, and only scoring results are provided without clear guidance for users, which is not conducive to users' accurate understanding of their pronunciation errors.
[0003] Early oral English teaching completely relies on teacher manpower teaching, and is affected by the unbalanced allocation of educational resources. Many people in China have not received systematic pronunciation training. Chinese native speakers often use Chinese rules for reading in the process of learning English, resulting in obvious Chinese accent and tune in pronunciation. With the development of deep learning and other technologies, mispronunciation detection and speech synthesis technologies support the birth of computer-aided pronunciation training systems. However, the current computer-aided pronunciation training system pays more attention to the improvement of its own performance, without considering the real needs of users, which is not conducive to users' accurate understanding of their pronunciation errors.
[0004] Abbreviations and key terms definition
[0005] MDD: Mispronunciation Diagnose and Detection, mispronunciation detection and diagnosis, refers to detecting errors in speech and providing diagnostic results.
[0006] CAPT: Computer Assisted Pronunciation Training, computer-aided pronunciation training, refers to a method of using computer technology to help users improve pronunciation.
[0007] TTS: Text-to-speech, text-to-speech, also known as speech synthesis. Refers to converting text information into standard and fluent pronunciation.
[0008] VC: Voice Conversion, refers to the technology of converting the voice of one person into the voice of another person while preserving the content and emotional characteristics of the voice.
[0009] ASR: Automatic Speech Recognization, refers to the conversion of speech into corresponding text information.
[0010] Cited documents:
[0011] [1] Z. Zhang, Y. Wang and J. Yang, "Masked Acoustic Unit for Mispronunciation Detection and Correction," ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, Singapore, 2022, pp. 6832-6836, doi: 10.1109 / ICASSP43922.2022.9747414.
[0012] [2] Alif Silpachai, Ivana Rehman, Taylor Anne Barriuso, John Levis, Evgeny Chukharev-Hudilainen, Guanlong Zhao, Ricardo Gutierrez-Osuna: Effects of Voice Type and Task on L2 Learners'
[0013] Awareness of Pronunciation Errors. Interspeech 2021: 1952-1956
[0014] [3] Pronunciation training method and device, electronic equipment and storage medium, CN115273898A, Anhui
[0015] Tao Yun Technology Co., Ltd.
[0016] [4] Bu Y, Ma T, Li W, et al. PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback [C] / / Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 2021: 1-14. SUMMARY
[0017] The present application aims to solve the technical problem that the prior art is not conducive for users to accurately recognize their own pronunciation errors, and to provide a pronunciation detection method and a pronunciation detection device.
[0018] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0019] A pronunciation detection method comprises the following steps: S1: providing a text for a user and collecting user pronunciation audio of the user reading the text; S2: detecting and diagnosing pronunciation errors of the user pronunciation audio, positioning error pronunciation and performing error labeling to obtain a text with error labeling; S3: visualizing the error pronunciation to obtain a comparison feedback result and feeding back the comparison feedback result to the user, so that the user understands their own pronunciation problems and error causes, and thus corrects according to the comparison feedback result.
[0020] In some embodiments of the present application, based on pronunciation errors commonly made by Chinese native speakers in learning English, a phonological contrast analysis is performed from the aspects of vowels, consonants and suprasegmentals to form the text for detecting the correctness of English pronunciation for Chinese native speakers.
[0021] In some embodiments of the present application, the comparison feedback result comprises an auditory comparison feedback result and a visual comparison feedback result.
[0022] In some embodiments of the present application, step S3 comprises the following steps: S31: performing speech synthesis on the text in standard American pronunciation to obtain standard pronunciation audio as the first auditory contrast feedback result; S32: performing speech synthesis on the text with error annotation in standard American pronunciation, and obtaining error reconstruction audio with the user's pronunciation based on sound conversion technology as the second auditory contrast feedback result; S33: the text with error annotation comprises phoneme text, converting the phoneme text into English phonetic symbols as the first visual contrast feedback result; S34: based on digital audio processing technology, audio visualization is performed on the user pronunciation audio, the standard pronunciation audio and the error reconstruction audio to obtain phonetic symbol contrast chart, waveform chart, spectrogram, pitch contour and formant scatter chart as the second visual contrast feedback result.
[0023] In some embodiments of the present application, a dual-channel mode is used to provide contrast feedback results for the user, wherein one channel is the second auditory contrast feedback result and the other channel is the user pronunciation audio.
[0024] In some embodiments of the present application, when the number of error pronunciations is greater than or equal to two, only the contrast feedback result of one error pronunciation is fed back at a time.
[0025] The present application also provides a pronunciation detection device comprising an evaluation test module, a detection positioning module and a visualization contrast module, wherein: the evaluation test module receives user pronunciation audio as input, provides text for the user, and collects user pronunciation audio of the user reading the text; the detection positioning module is used for error pronunciation detection and diagnosis of the user pronunciation audio, error pronunciation positioning and error annotation, thereby obtaining text with error annotation; the visualization contrast module is used for visualizing the error pronunciation to obtain contrast feedback results, and feeding back the contrast feedback results to the user, so that the user understands the pronunciation problems and error reasons, and corrects according to the contrast feedback results.
[0026] In some embodiments of the present application, the contrast feedback result includes an auditory contrast feedback result and a visual contrast feedback result; the auditory contrast feedback result includes a first auditory contrast feedback result and a second auditory contrast feedback result, and the visual contrast feedback result includes a first visual contrast feedback result and a second visual contrast feedback result; the visualization contrast module synthesizes the text into a standard American pronunciation to obtain a standard pronunciation audio as the first auditory contrast feedback result; the visualization contrast module synthesizes the text with error annotations into a standard American pronunciation, and obtains an error reconstruction audio with the user's pronunciation based on a voice conversion technology as the second auditory contrast feedback result; the text with error annotations includes a phoneme text, and the visualization contrast module converts the phoneme text into an English phonetic symbol as the first visual contrast feedback result; the visualization contrast module performs audio visualization on the user pronunciation audio, the standard pronunciation audio and the error reconstruction audio based on a digital audio processing technology to obtain a waveform graph, a spectrogram, a pitch contour and a formant scatter plot as the second visual contrast feedback result.
[0027] In some embodiments of the present application, the visualization contrast module provides the contrast feedback result to the user in a dual-channel manner, wherein one channel is the second auditory contrast feedback result, and the other channel is the user pronunciation audio.
[0028] In some embodiments of the present application, when the number of error pronunciations is greater than or equal to two, the visualization contrast module only feeds back the contrast feedback result of one error pronunciation at a time.
[0029] The present application has the following beneficial effects:
[0030] The pronunciation detection method and the pronunciation detection device provided by the present application can help the user to correct the error pronunciation by providing the user with a text, collecting the user pronunciation audio of the user reading the text, positioning the error pronunciation through error pronunciation detection and diagnosis, visualizing the error pronunciation to obtain a contrast feedback result, and finally feeding back the contrast feedback result to the user.
[0031] In addition, in some embodiments of the present application, the following beneficial effects are also achieved:
[0032] By analyzing the common error types of pronunciation confusion of English pronunciation of Chinese native speakers, a reading text is designed and provided, so that the user can expose more error pronunciations in a short time, thereby quickly positioning the high-frequency error types of the user, which helps the user to correct the error pronunciation.
[0033] By providing multi-dimensional multi-modal visual and auditory correct and incorrect pronunciation comparison feedback results, the learner is explicitly informed of the location of his own errors, improving his ability to perceive his own incorrect pronunciation, and extending the interaction path between the computer-aided pronunciation training system and the user, thereby helping the user to better correct and improve.
[0034] By comparing and correcting only one incorrect pronunciation at a time, the user is allowed to focus on one pronunciation problem at a time in a controlled variable manner, thereby improving the accuracy of the user's error localization and proceeding step by step in self-correction.
[0035] Other benefits of the embodiments of the present application will be further described below. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a step flowchart of the pronunciation detection method in the embodiments of the present application;
[0037] Figure 2 is a schematic diagram of the pronunciation detection device in the embodiments of the present application;
[0038] Figure 3 is a usage flowchart of the pronunciation detection method in embodiment 1;
[0039] Figure 4 is a principle diagram of the pronunciation detection method in the embodiments of the present application;
[0040] Figure 5a is a waveform diagram of the user's reading audio in the embodiments of the present application;
[0041] Figure 5b is a waveform diagram of the standard American pronunciation in the embodiments of the present application;
[0042] Figure 5c is a waveform diagram of the incorrect reconstructed audio in the embodiments of the present application;
[0043] Figure 6a is a waveform diagram in the visual feedback comparison result in the embodiments of the present application;
[0044] Figure 6b is a spectrogram in the visual feedback comparison result in the embodiments of the present application;
[0045] Figure 6c is a pitch contour in the visual feedback comparison result in the embodiments of the present application;
[0046] Figure 6d is a resonance peak scatter plot in the visual feedback comparison result in the embodiments of the present application. DETAILED DESCRIPTION
[0047] The application will be further described below with reference to the drawings and in conjunction with the preferred embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0048] It should be noted that the terms such as left, right, up, down, top, bottom, etc. in the embodiments are only relative concepts or are referenced to the normal use state of the product, and should not be considered as limiting.
[0049] In the prior art, for example, paper [1] proposes to use an acoustic unit (AU) as an intermediate feature for pronunciation error detection and correction, rather than directly using an ASR-based method, to avoid the expensive requirement of data set annotation for an ASR-based CAPT system. Paper [2] conducts a pronunciation perception experiment using Golden Speech and Silver Speech, Golden Speech being audio in which the learner's pronunciation is converted into standard pronunciation using a voice conversion technique, and Silver Speech being standard pronunciation audio, to verify whether the voice conversion technique can improve the learner's error awareness by having the learner listen to the audio and then mark the places where he or she thinks there are pronunciation problems. Patent [3] proposes a method of calculating the similarity between the read-aloud audio and the standard audio to display the pronunciation training information of the word and intuitively understand the pronunciation of each syllable. Paper [4] proposes a method of introducing exaggerated feedback into a computer-aided pronunciation training system, which sets up four different levels of exaggerated audio-visual speech, namely zero, low, medium and high, for the user to learn independently.
[0050] Currently, most AI spoken language applications provide functions such as textbook synchronization, read-aloud scoring, and phonetic symbol learning, but there are relatively few systems that can provide targeted pronunciation training. The following shortcomings exist:
[0051] (1) The user's own pronunciation level and mother tongue background are not taken into account, and all learners at the initial level and with the same pronunciation background use the same set of test questions, so the difficulty of the provided knowledge may not be suitable for the learner; it is difficult to find all the pronunciation errors of the learner in a short period of time (papers [1], [2], [4], and patent [3]);
[0052] (2) There is no word-level focused explanation of the user's own errors, and too many errors in a sentence will make it difficult for the learner to accurately locate the errors and thus improve them (paper [2]);
[0053] (3) There is no visual feedback of the user's errors, and only the pronunciation score and error position are provided, so the user can only know his or her own level, and the user cannot realize the reasons for his or her own errors, resulting in the user's perceptual ability not being improved (papers [1], [2]).
[0054] To solve the above problems, the present application provides a pronunciation detection method and a pronunciation detection device. Figure 1 As shown in the figure, the pronunciation detection method comprises the following steps: S1: providing a text to a user and collecting user pronunciation audio of the user reading the text;
[0055] S2: detecting and diagnosing the user pronunciation audio, positioning and labeling the errors to obtain a text with error labels; S3: visualizing the errors to obtain a comparison feedback result and feeding back the comparison feedback result to the user, so that the user understands the pronunciation problems and error causes, and corrects the errors according to the comparison feedback result.
[0056] Preferably, based on the pronunciation errors commonly made by Chinese native speakers in learning English, the phonology of vowels, consonants and suprasegmentals is analyzed to form the text for detecting the correctness of English pronunciation of Chinese native speakers.
[0057] The embodiment of the present application designs a more targeted reading text for learners with Chinese native background, which covers all possible pronunciation confusion error types to expose more errors in a short time.
[0058] (1) providing a text to a user and collecting user pronunciation audio of the user reading the text.
[0059] Mother tongue transfer is a widespread influence in language learning, which is particularly evident in pronunciation learning and may exist in any mother tongue background. The pronunciation rules of Chinese are based on pinyin, while the pronunciation of English is based on phonetic symbols. Different pronunciation models are one of the reasons for mother tongue transfer. However, the common points between the two, i.e. the common points between the phonetics, are the characteristics of vowels, consonants and suprasegmentals, and the International Phonetic Alphabet can provide symbolic language support for all vowels, consonants and suprasegmentals. Therefore, the embodiment of the present application compares the differences between the vowels, consonants and suprasegmentals of Chinese and English. The negative transfer effects caused by the similarities and differences between the two are summarized as shown in Table 1. The comparison results are expanded from the phoneme level to the words containing the phonemes, and further expanded to the sentences containing the words. According to the process from local to global, the principle of minimum redundancy is followed, and the final evaluation test is presented in the form of words and sentences as shown in Table 2. The content in the table includes most of the possible errors of Chinese native speakers when speaking English, which can expose the errors in a short time.
[0060] The method for limiting reading text in the embodiment of the present application is obviously superior to the current AI spoken language products such as FluentU and Open English, which can comprehensively stimulate the user to expose the pronunciation of the user himself. Compared with the method in papers [1], [2] and [4] which only use large-scale corpus, the embodiment of the present application can quickly locate the user's incorrect pronunciation and avoid the evaluation results caused by the mismatch between the text and the user's own level.
[0061] Table 1 English pronunciation confusion existing in Chinese native speakers
[0062]
[0063] Table 2 part of the content of the evaluation test
[0064]
[0065] (2) Targeted mispronunciation detection and diagnosis
[0066] After the user finishes reading the provided text, the mispronunciation detection and diagnosis technology in deep learning is used for error recognition and diagnosis, the user's error pronunciation and its diagnosis result are obtained, the error pronunciation is located and error labeling is performed, the text with error labeling is obtained, that is, the phoneme text content corresponding to the user's actual pronunciation, the text with error labeling contains error pronunciation letters, words and specific error forms, etc., and the user's pronunciation problem is understood.
[0067] (3) Implementation of personalized audio-visual feedback
[0068] The embodiment of the present application introduces a comparative feedback strategy into a computer-aided pronunciation training system, as shown in Figure 4 It is composed of two parts of auditory comparative feedback results and visual comparative feedback results. After locating the user's error pronunciation, how to emphasize and highlight the user's awareness of his own errors is one of the capabilities that the current CAPT system lacks. The embodiment of the present application introduces the comparative feedback mechanism in the traditional teaching method, so that the user can face the comparison between correct and incorrect pronunciation, and through the comparison of multi-dimensional and multi-modal forms, the user can deeply understand the pronunciation difference between correct and incorrect.
[0069] To realize personalized audio-visual feedback, the comparative feedback result of visual error pronunciation is obtained, the embodiment of the present application synthesizes the prompt text into standard American pronunciation, obtains the standard pronunciation audio (as shown in Figure 5b ), as the first auditory comparative feedback of the embodiment of the present application; the text with error labeling is synthesized into standard American pronunciation, and is converted into error reconstruction audio with the user's tone based on sound conversion technology (as shown in Figure 5cThe phoneme text with error labels is converted into English phonetic symbols, as the first visual contrast of the embodiment of the application; the user's pronunciation audio (such as Figure 5a The standard pronunciation audio, the error reconstruction audio are audio-visualized using digital audio processing technology to obtain a phonetic symbol contrast chart, a waveform chart (as shown in Figure 6a A spectrogram (as shown in Figure 6b A pitch contour (as shown in Figure 6c A resonance peak scatter chart (as shown in Figure 6d The second visual contrast of the embodiment of the application. Since the training audio is generated according to the user's own pronunciation, the audio-visual contrast feedback for the user can be accurately generated to help the user improve the perception of pronunciation errors. Figures 5a to 5c In the above figures, the horizontal coordinates are time, and the vertical coordinates are amplitude; Figures 5a to 6d In the above figures, the word "issue" is taken as an example.
[0070] In the preferred embodiment, the user is provided with contrast feedback results using a dual-channel mode, one channel being the second auditory contrast feedback result, and the other channel being the user's pronunciation audio.
[0071] After obtaining the auditory contrast feedback results, in order to enable the user to more intuitively learn the reasons for the errors, the embodiment of the application uses the knowledge of digital audio processing to visually draw the audio signals, as shown in Figures 6a to 6d
[0072] Waveform chart: as shown in Figure 6a The change of the sound signal with time is visually represented, and the basic characteristics of the sound, such as the loudness change and the pronunciation duration, are roughly displayed. In the above figures, the horizontal coordinates are time, and the vertical coordinates are amplitude.
[0073] Spectrogram: as shown in Figure 6b The time-domain signal is converted into a frequency-domain signal, which is a two-dimensional representation of the frequency spectrum of the sound signal changing with time, and can intuitively display the sound frequency components at different times and the change process. In the above figures, the gray scale is the amplitude (the darker the color or brightness, the greater the signal strength at the time point and the frequency).
[0074] Pitch contour: as shown in Figure 6c The curve of the perceived pitch of the sound is tracked with time, indicating the change of the fundamental frequency of the speech signal with time, and carrying certain prosody information. The pitch contour is very necessary for the contrast of the tone change at the word level. Chinese native speakers are affected by the single tone of Chinese characters, and it is difficult for them to understand the need for tone change in a word, so the pitch contour is needed to show the pronunciation change in the word in an intuitive form. The horizontal coordinates are time, and the vertical coordinates are pitch (hertz).
[0075] Formant scatter plot: As shown in FIG. 6, the frequencies that can resonate in the vocal tract appear as a series of discrete high-energy regions in the spectrogram. Formants are strongly related to the content of pronunciation, and thus are crucial for distinguishing vowels, consonants, etc. The frequency and energy of formants can be estimated by Linear Predictive Coding (LPC). The horizontal axis is time, and the vertical axis is frequency. Figure 6d
[0076] Accordingly, the embodiment of the present application obtains the personalized audio-visual contrast feedback result for the pronunciation of a certain user, and provides in-depth pronunciation guidance for the user, so as to improve the user's ability to distinguish and perceive correct and incorrect pronunciation.
[0077] In a preferred embodiment, when the number of incorrect pronunciations is greater than or equal to two, only one contrast feedback result of incorrect pronunciation is fed back at a time.
[0078] The present application also provides a pronunciation detection device, comprising an evaluation test module, a detection positioning module, and a visual contrast module, wherein: the evaluation test module receives user pronunciation audio as input, and is used to provide text for the user, and collect user pronunciation audio of the user reading the text; the detection positioning module is used for incorrect pronunciation detection and diagnosis of user pronunciation audio, performs incorrect pronunciation positioning and error labeling, so as to obtain text with error labeling; the visual contrast module is used for visualizing the contrast feedback result of the incorrect pronunciation, and feeding back the contrast feedback result to the user, so that the user understands the pronunciation problem and the error reason, and thus corrects according to the contrast feedback result. The evaluation test module is also used to provide reading text for the user.
[0079] In a preferred embodiment, the visual contrast module synthesizes the text into standard American pronunciation to obtain standard pronunciation audio as the first auditory contrast feedback result; the visual contrast module synthesizes the text with error labeling into standard American pronunciation, and obtains error reconstruction audio with the user's voice based on sound conversion technology as the second auditory contrast feedback result; the visual contrast module converts the phoneme text into English phonetic symbols as the first visual contrast feedback result; and the visual contrast module performs audio visualization on the user pronunciation audio, the standard pronunciation audio, and the error reconstruction audio based on digital audio processing technology to obtain a waveform diagram, a spectrogram, a pitch contour, and a formant scatter plot as the second visual contrast feedback result.
[0080] In a preferred embodiment, the visual contrast module uses a dual-channel mode to provide the contrast feedback result for the user, wherein one channel is the second auditory contrast feedback result, and the other channel is the user pronunciation audio.
[0081] In a preferred embodiment, when the number of incorrect pronunciations is greater than or equal to two, the visualization comparison module only provides the comparison feedback result for one of the incorrect pronunciations at a time.
[0082] When there are two or more incorrect pronunciations, the visualization module will show which pronunciation is problematic in the comparison feedback results. However, clicking on the audiovisual feedback only applies to one of the pronunciations. Taking "It is a sentence" as an example, if both 'a' and 'en' are incorrect, both syllables will be marked and displayed simultaneously. However, the audiovisual comparison feedback generated for 'a' will not affect 'en' in any way; it's simply a matter of "wrong is wrong," depending on which syllable the user wants to correct. Each time the user reads aloud, comparison feedback is generated based on that reading result for further instruction, unlike other products that directly provide teaching materials. This is a dynamic teaching process.
[0083] In this embodiment of the invention, the user selects Chinese as their native language background and reads the text aloud as prompted. The assessment module receives the audio file read aloud by the user as input, and the detection and localization module processes it using mispronunciation detection and diagnosis technology to obtain the phonemes of the user's incorrect pronunciation and their diagnostic results. The visualization comparison module synthesizes audio-visual comparison feedback based on the diagnosis. The audio-visual comparison feedback results are directly presented to the user so that they can understand their specific pronunciation problems and the reasons for their errors, and correct them based on the comparison feedback results. Figure 2 As shown, taking "issue" as an example, the phonetic symbol / I / does not have a separate pronunciation in the initials and finals of Mandarin Chinese. The pronunciation of "issue" is / Some learners pronounce the short vowel / I / as the long / eI / , resulting in a pronunciation that is / This invention requires first returning a phonetic symbol comparison to the user, while also providing the user with their original recording and correct pronunciation. / Audio and Reconstruction / / audio, and the corresponding audio visualization images of the three.
[0084] Still taking issue as an example, the generated auditory contrast feedback audio (illustrated in spectrogram) is as follows. In a specific embodiment, the left ear can apply the user's original audio, and the right ear can apply the error reconstruction audio. The user is provided with contrast feedback in two channels. If the user has too many errors in a sentence, the embodiment of the present application generates contrast feedback results only for a certain error pronunciation according to the diagnostic results, and does not correct other errors temporarily, so as to realize the process of "one word at a time". In this process, the embodiment of the present application uses a TTS algorithm with a standard American pronunciation dataset to generate natural and authentic English expressions. At the same time, a VC algorithm for any speaker is used to obtain the user's acoustic characteristics, and the error pronunciation reconstruction is performed with the preservation of the tone, so that the user can "hear his own voice".
[0085] Embodiment 1
[0086] The pronunciation detection method of the present embodiment utilizes a computer-aided pronunciation training system with the introduction of audio-visual contrast feedback strategy. As shown in Figure 3 The pronunciation detection method of the present embodiment consists of three parts: (1) for Chinese native background learners, a template is provided for the learner's reading, which is an easy-to-mistake and easy-to-confuse pronunciation text for the learner to learn English, helping the learner to more obviously expose his own pronunciation errors, and the template is also used as an assessment test; the pronunciation detection device inputs the reading text to the user, and the audio after the user's reading is the input of the pronunciation detection device; (2) user error pronunciation diagnosis based on mispronunciation detection and diagnosis technology, scoring and correcting the learner's pronunciation to accurately locate the problem; (3) speech generation based on speech synthesis and voice conversion technology, audio visualization based on digital audio processing as contrast feedback results, the contrast feedback results provide contrast feedback from two angles of vision and hearing, vision provides contrast of phonetic symbols, waveform, spectrogram, pitch contour and formant scatter diagram, and hearing provides contrast between the user's pronunciation audio, standard pronunciation audio and error reconstruction audio.
[0087] The pronunciation detection method provided by the embodiment of the present application utilizes a computer-aided pronunciation training system, which mainly consists of three parts, pronunciation evaluation and error diagnosis for user background, and personalized audio-visual contrast feedback generated according to error results. The first part, the specific pronunciation diagnosis takes into account the common pronunciation errors of Chinese native speakers in learning English, and conducts a phonetic contrast analysis from the aspects of vowels, consonants and suprasegmentals, and summarizes a set of evaluation test questions suitable for learners of this background. The second part, according to the reading of the learner on the text, conducts comprehensive mispronunciation detection and diagnosis to locate the user's own high-frequency error; the third part, according to the error results, carries out standard American pronunciation and error pronunciation reconstruction, and carries out audio visualization of the synthesis results to obtain the corresponding waveform diagram, spectrogram, pitch contour and resonance peak scatter diagram, so as to provide personalized audio-visual contrast feedback. Compared with the previous method, the embodiment of the present application has the following advantages:
[0088] The learners of Chinese native background are designed to read the text more specifically, which includes all possible pronunciation confusion error types, so as to expose more errors in a short time.
[0089] Each time only one error pronunciation of the user on one word is corrected specifically, which ensures that the learner only focuses on the problem of a certain pronunciation, and can improve the user's positioning ability of the error.
[0090] Meanwhile, various forms of audio-visual contrast feedback are provided to clearly inform the learner of the problem of his own error, prolong the interaction path between the computer-aided pronunciation training system and the user, and better correct and improve.
[0091] The embodiment of the present application has important value for human-computer interaction and foreign language teaching industry, and can be used in various application scenarios such as oral learning, accent correction and pronunciation training, and can provide personalized feedback for the user's own level to improve the user's own error pronunciation awareness.
[0092] The embodiment of the present application has no special requirement for the hardware environment and can be realized on a general computer.
[0093] The characteristics of the embodiment of the present application can be summarized as:
[0094] The pronunciation evaluation reading text for Chinese native background English pronunciation is provided, which can quickly locate the high-frequency error type of the user in a short time.
[0095] It is ensured that each time only one error pronunciation is corrected to control the variable form to accurately locate the user's own error, and to proceed step by step for self-correction.
[0096] Multi-dimensional multi-modal personalized audio-visual correct and wrong contrast feedback is provided, and the ability provided by the computer-aided pronunciation training system is further extended, and the user's perception ability of the user's own wrong pronunciation is improved.
[0097] The system design realizes the process of dynamic teaching and personalized correction feedback, and has certain application significance.
[0098] The embodiment of the present application can also be expanded as follows:
[0099] (1) For the feedback mode, the embodiment of the present application uses visual and auditory contrast feedback, if it can be applied to intelligent terminal equipment, it can also use the contrast form of hearing and touch (vibration) to make the learner aware of the error through more obvious stimulation.
[0100] (2) For algorithm implementation, the embodiment of the present application uses the sound conversion method to reconstruct the pronunciation, compare the user's wrong pronunciation with the user's correct pronunciation, and also can directly use the speaker-related natural speech synthesis algorithm to compare the correct standard American pronunciation with the standard American pronunciation with the user's own error, which is also a kind of contrast feedback.
[0101] The above is a further detailed description of the present application in combination with a specific preferred embodiment, and the specific implementation of the present application cannot be limited to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of equivalent alternatives or obvious modifications can be made, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present application.
Claims
1. A method of detecting a sound production, characterized by, The method comprises the following steps: S1: providing a text for a user and collecting user pronunciation audio of the user reading the text; S2: detecting and diagnosing mispronunciation of the user pronunciation audio, positioning and marking errors to obtain a text with error marking; S3: visualizing the error pronunciation to obtain a contrast feedback result and feeding back the contrast feedback result to the user, so that the user understands the pronunciation problem and the error reason, and corrects according to the contrast feedback result; wherein step S3 comprises the following steps: S31: synthesizing the text into standard American pronunciation to obtain standard pronunciation audio as a first auditory contrast feedback result; S32: synthesizing the text with error marking into standard American pronunciation, and obtaining error reconstruction audio with user voice based on sound conversion technology as a second auditory contrast feedback result; S33: the text with error marking comprises a phoneme text, which is converted into English phonetic symbols as a first visual contrast feedback result; S34: based on digital audio processing technology, the user pronunciation audio, the standard pronunciation audio and the error reconstruction audio are visualized to obtain phonetic symbol contrast chart, waveform chart, spectrogram, pitch contour and resonance peak scatter chart as a second visual contrast feedback result.
2. The sound production detection method of claim 1, wherein, Based on the pronunciation errors commonly made by Chinese native speakers in learning English, a text for detecting the correctness of English pronunciation of Chinese native speakers is formed by comparing and analyzing the pronunciation from the aspects of vowels, consonants and suprasegmentals in phonology.
3. The sound production detection method of claim 1, wherein, The contrast feedback result comprises auditory contrast feedback result and visual contrast feedback result.
4. The sound production detection method of claim 1, wherein, The contrast feedback result is provided to the user in a dual-channel manner, one channel being the second auditory contrast feedback result and the other channel being the user pronunciation audio.
5. The sound production detection method of claim 1, wherein, When the number of error pronunciations is greater than or equal to two, only one contrast feedback result of error pronunciation is fed back at a time.
6. A sound production detection apparatus characterized by comprising: The method comprises an evaluation test module, a detection positioning module and a visual contrast module, wherein: The evaluation test module receives user pronunciation audio as input to provide a text for a user and collect user pronunciation audio of the user reading the text; The detection positioning module is used for detecting and diagnosing mispronunciation of the user pronunciation audio, positioning and marking errors to obtain a text with error marking; The visual contrast module is used for visualizing the error pronunciation to obtain a contrast feedback result and feeding back the contrast feedback result to the user, so that the user understands the pronunciation problem and the error reason, and corrects according to the contrast feedback result; the contrast feedback result comprises auditory contrast feedback result and visual contrast feedback result; the auditory contrast feedback result comprises first auditory contrast feedback result and second auditory contrast feedback result, and the visual contrast feedback result comprises first visual contrast feedback result and second visual contrast feedback result; The visual contrast module synthesizes the text into standard American pronunciation to obtain standard pronunciation audio as a first auditory contrast feedback result; The visual contrast module synthesizes the text with error annotations into standard American pronunciation speech, and obtains the error reconstruction audio with the user's voice based on sound conversion technology as the second auditory contrast feedback result; The text with error annotations includes phoneme text, and the visual contrast module converts the phoneme text into English phonetic symbols as the first visual contrast feedback result; The visual contrast module uses digital audio processing technology to visually display the user's pronunciation audio, the standard pronunciation audio, and the error reconstruction audio as waveform graphs, spectrograms, pitch contours, and formant scatter plots as the second visual contrast feedback result.
7. The sound production detection apparatus of claim 6, wherein The visual contrast module provides the contrast feedback results to the user in a dual-channel manner, where one channel is the second auditory contrast feedback result, and the other channel is the user's pronunciation audio.
8. The sound production detection apparatus of claim 6, wherein When the number of error pronunciations is greater than or equal to two, the visual contrast module only feeds back the contrast feedback result of one error pronunciation at a time.
Citation Information
Patent Citations
Pronunciation correcting method, pronunciation correcting device, pronunciation correcting equipment and computer readable storage medium
CN110085261A