Pronunciation learning program
The computer program addresses the challenge of unnoticed pronunciation defects by employing reverse playback and focused phoneme comparison, enhancing learners' intuitive awareness of pronunciation differences.
Patent Information
- Application Number
- JP2024089586
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2044-05-31
AI Technical Summary
Learners fail to notice defects in their own pronunciation when comparing model speech with their recorded speech due to the limitations of existing forward playback methods.
A computer program that utilizes reverse playback to compare model and recorded speech, highlighting pronunciation differences by reversing the audio waveform, and optionally using a reverse playback start position determination to focus on specific phonemes.
Enhances learners' ability to intuitively identify pronunciation flaws by comparing reversed model and recorded speech, making it suitable for young learners and improving pronunciation accuracy.
Abstract
Description
[Technical Field]
[0001] The present invention relates to a pronunciation learning program for the purpose of acquiring a foreign language. [Background technology]
[0002] Some general-purpose audio processing software has a reverse playback function that reverses the time axis of the audio waveform of recorded audio or audio read from an audio file and plays it back. However, no computer program for learning foreign language pronunciation has been proposed that systematically uses this reverse playback function.
[0003] On the other hand, there is a foreign language pronunciation learning method in which a learner listens to a model speech, imitates the pronunciation, records it, and then compares the model speech with the recorded speech while listening to it again, thereby becoming aware of flaws in their own pronunciation. Naturally, this playback is performed in the forward direction, keeping the original time axis of the speech waveform. In this invention, this forward playback method will be called "forward playback." Summary of the Invention [Problem to be solved by the invention]
[0004] The problem to be solved is that learners often fail to notice defects in their own pronunciation when comparing model speech with their own recorded speech in order. [Means for solving the problem]
[0005] The most important feature of this invention is that it is a computer program that helps you notice pronunciation defects that would not be noticed by comparing a model voice with your own recorded voice by playing them in reverse.
[0006] The inventors have found that, under certain conditions, even ordinary learners who are not phonetic experts can judge for themselves whether a pronunciation is appropriate by playing the two sounds backwards and comparing them. For example, the English word "egg" is pronounced in katakana as "egg." Translated into romaji, it is "eggu," so when played backwards, it should sound like "ugge," but in reality, it tends to sound like "uhhe." This is thought to be due to the fact that Japanese speakers swallow their breath when pronouncing double consonants. On the other hand, when pronounced more politely in English, it sounds like "(u)gge." Similarly, when "add" is played backwards, it tends to sound like "otta" instead of "odda," and when "up" is played backwards, it tends to sound like "uhha" instead of "uppa." If you can hear "(ug)ge" when you play the sound of "egg" in reverse, you will have come closer to an English pronunciation. In reality, the two sounds are not always clearly distinguishable, so by comparing the model sound played in reverse with your own pronunciation played in reverse and trying to get the feeling that the two sounds are similar, you can achieve a new type of pronunciation learning.
[0007] In this specification, the native language is Japanese and the foreign language is English, but the native language and foreign language are not limited to specific languages. For example, the present invention can be used as a computer program for learning Japanese as a foreign language when the native language is English.
[0008] The first solution in the present invention is to provide a program that causes a computer to function as an audio input means that inputs audio and converts it into audio data, an audio output means that outputs the audio data as audio, an audio data storage means that stores audio data, an audio model forward playback means that outputs an audio model to the audio output means in a forward direction on the time axis, an audio recording means that stores the audio data input by the audio input means in the audio data storage means, a recorded audio reverse playback means that outputs the audio data stored in the audio data storage means by the audio recording means to the audio output means in a reverse direction on the time axis, and an audio model reverse playback means that outputs the audio model to the audio output means in a reverse direction on the time axis.
[0009] The computer may be a personal computer, smartphone, mobile phone, or other type of computer, but is assumed to be equipped with a microphone as an audio input device and a speaker, headphones, earphones, or other audio output device. The audio input means is a means for inputting audio data from the audio input device. The audio output means is a means for outputting audio to the audio output device based on the audio data. The device in which the audio data storage means stores audio data is the internal memory of this computer, an external storage device connected to this computer, an external storage device connected to the network to which this computer is connected, or the internal memory or external storage device of another computer connected to the network to which this computer is connected.
[0010] It is desirable to have a recorded audio sequential playback means for playing back the audio data stored in the audio data storage means on the audio output means in the forward direction on the time axis. In this case, not only comparison in reverse playback but also comparison in forward playback can be performed in combination.
[0011] The model audio forward playback means / model audio reverse playback means may forward / reverse play model audio data pre-stored in the audio data storage means, or may use audio data generated at any time by a voice synthesis program whose parameters have been adjusted so that the difference between the two is simply that the time axis is reversed.
[0012] The second solution of the present invention is a computer program according to the first solution, characterized in that it has a reverse playback start position determination means for determining the start position on the time axis of reverse playback from audio data, the model audio reverse playback means reverse-plays the model audio from the start position determined by the reverse playback start position determination means from the audio data of the model audio, and the recorded audio reverse playback means reverse-plays the recorded audio from the start position determined by the reverse playback start position determination means from the audio data of the recorded audio.
[0013] When comparing the reversed katakana pronunciation of "egg" ("uh-he") with the reversed English pronunciation ("(uh)-g-ge"), learners may be confused about what to focus their attention on. Therefore, it is more intuitive to see only the "he" and "ge" sounds reversed, which are the sounds that should be focused on. For example, the reverse playback start position determining means calculates the correspondence between the phoneme string and the waveform of the voice data using a segmentation program, and determines the start position of reverse playback so that the phoneme to which attention should be paid, which has been determined through trial and error by the provider of the pronunciation learning program, is played back in reverse. Alternatively, if the purpose of the reverse playback start position determining means is to specifically remove the effects of the pronunciation of Japanese choked consonants, it is simple and useful to determine the reverse playback start position as the minimum value of the sound volume over time, which is the position where the corresponding stop consonant closes, without using a general-purpose complex segmentation program. [Effects of the Invention]
[0014] The pronunciation learning program of the present invention has the advantage of being able to make learners aware of differences that would not be noticed by comparing a model speech and a learner's recorded speech in forward order. The pronunciation learning program of the second solution of the present invention can make learners more intuitively aware of differences between the two speeches played back in reverse, making it suitable for use by even young learners. DETAILED DESCRIPTION OF THE INVENTION [Example]
[0015] In Example 1, a "model audio playback button," a "recording button," a "recorded audio reverse playback button," and a "model audio reverse playback button" are arranged on the computer operation screen. Each button triggers a model audio forward playback means, an audio recording means, a recorded audio reverse playback means, and a model audio reverse playback means. The learner first presses the "Model Audio Playback Button" to play the model audio forward and listen to it. Next, they press the "Record Button" to start recording and speak by imitating the audio they heard. After recording is complete, they press the "Recorded Audio Reverse Playback Button" to listen to the recorded audio in reverse, and then press the "Model Audio Reverse Playback Button" to listen to the model audio in reverse, and compare the two. If they feel that the two are not similar, they start again by pressing the "Model Audio Playback Button" to play the same model audio forward. When the two are close enough to be perceived as similar, learning of that audio is complete. [Example]
[0016] In Example 2, the model audio sequential playback means is triggered by a voice command such as "next" spoken into the smartphone. When the model audio sequential playback means has completed forward playback of the model audio, the voice recording means is triggered. When a certain time has elapsed since the start of recording, recording is completed and the model audio sequential playback means is triggered again. When the model audio sequential playback means has completed forward playback of the model audio, the model audio reverse playback means is triggered. When the model audio reverse playback means has completed reverse playback of the model audio, the recorded audio sequential playback means is triggered. When the recorded audio sequential playback means has completed forward playback of the recorded audio, the recorded audio reverse playback means is triggered. When the recorded audio reverse playback means has completed reverse playback of the recorded audio, the device enters a state of waiting for a voice command. In this way, after listening to and imitating the model audio, the learner can compare the automatically played set of forward and reverse playback of the model audio with the automatically played set of forward and reverse playback of his or her own recorded audio. If the two voices are not considered similar, the system can repeat the learning of the same voice from the beginning by speaking the voice command "again." When the two voices become close enough to be considered similar, the learning of that voice is considered complete and the system can move on to learning the next voice by speaking the voice command "next." [Example]
[0017] In the third embodiment, the purpose is specialized to remove the influence of the pronunciation of Japanese geminate consonants, and an example is given of a reverse playback start position determination means that determines the minimum value of the sound volume over time as the position of the stop of the corresponding stop consonant as the reverse playback start position. First, the absolute value of the waveform of the audio data to be reverse-played is smoothed using a Gaussian filter or the like to create a transition graph of the volume of the sound that changes smoothly on the time axis. In the case of one-syllable speech data such as egg, the position on the graph where the first maximum is found by searching forward from the beginning is determined to be the vowel position, and the first minimum found by searching forward again is determined to be the position of the stop sound, i.e., the position equivalent to the Japanese geminate consonant. Then, by starting reverse playback from this stop position, only the part that requires attention is presented in reverse playback, such as "he" and "ge," rather than "uhhe" or "(u)gge."
Claims
1. Computer, a voice input means for inputting voice and converting it into voice data; an audio output means for outputting the audio data as audio; a voice data storage means for storing voice data; a model audio sequential playback means for outputting the model audio to the audio output means in a forward direction on a time axis; a voice recording means for storing voice data input by the voice input means in the voice data storage means; a recorded voice reverse playback means for outputting the voice data stored in the voice data storage means by the voice recording means to the voice output means in a reverse direction on the time axis; a model audio reverse playback means for outputting the model audio to the audio output means in a reverse direction on the time axis; A program to function as a
2. a reverse playback start position determining means for determining a start position on a time axis of reverse playback from audio data; the model audio reverse playback means reverse-plays the model audio from the start position determined by the reverse playback start position determination means from the audio data of the model audio; the recorded voice reverse playback means plays back the recorded voice in reverse from the start position determined by the reverse playback start position determination means from the voice data of the recorded voice.
2. The computer program of claim 1.