Speech similarity calculation engine, and learning assistance program, learning assistance system, and learning assistance method using same
The voice similarity calculation engine addresses the challenges of evaluating speech in language learning by comparing learner audio with teaching material audio, using advanced analysis techniques to capture voice changes effectively and improve conversation skills at a lower cost.
Patent Information
- Application Number
- PCT/JP2024/038721
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for evaluating speech during language learning, such as shadowing, require specialized personnel and are costly due to the need for human evaluation or complex AI models that are difficult to develop and operate effectively.
A voice similarity calculation engine that acquires and compares audio data to determine the similarity between a learner's speech and teaching material audio, using techniques such as amplitude change extraction, change rate pattern comparison, and vector sequence analysis to calculate Euclidean distance or cosine similarity.
This solution allows for accurate capture of regular voice changes in language learning, reducing development and operation costs while enabling learners to quickly improve their practical conversation skills.
Smart Images

Figure JP2024038721_08052025_PF_FP_ABST
Abstract
Description
Speech similarity calculation engine, and learning support program, learning support system, and learning support method using the same
[0001] The present invention relates to a speech similarity calculation engine, which is a program for determining the similarity between first speech data and second speech data, for example, in learning pronunciation of a foreign language, and to a program, system, and method for providing learning support and teaching materials using the same.
[0002] In recent years, with the advancement and falling prices of information terminal devices such as tablet PCs and smartphones, as well as the development of communication networks such as the Internet, learning methods using personal computers, tablet PCs, smartphones, etc. via the Internet (e-learning) have become commonplace. Because e-learning allows learning to be done anytime and anywhere via the Internet, it is increasingly being used for language learning to prepare for qualifications and school entrance exams, and to improve professional skills.
[0003] In recent years, the so-called "shadowing" method, which simultaneously trains listening and speaking skills, has been gaining attention in e-learning. Shadowing is a learning method in which students listen to the speech of a native speaker or audio materials and repeatedly follow the speech like a "shadow" (see, for example, Patent Document 1).
[0004] In the production of a language, acoustic factors are as important as spelling, and training in regular sound changes in the target language, such as intonation, liaison, and Korean "patchim," is particularly important. For example, in languages such as English, linking occurs, creating a flowing connection between the boundaries of consecutive spoken words. This linking is an important element for natural pronunciation and fluent speaking. For this reason, in speaking training, not just shadowing, consciously pronouncing regular sound changes reduces pauses and stumbles between words. Understanding and becoming accustomed to these regular sound changes also improves listening ability in the target language.
[0005] In the past, in language learning such as shadowing training, in addition to the traditional method of native speakers or teachers evaluating students' pronunciation as a way to judge the success of regular sound changes such as intonation and liaison, tools using machine learning such as AI (artificial intelligence) have also been developed.
[0006] Patent No. 5911630
[0007] However, while the method of direct human evaluation described above has the advantage of being able to accurately capture how a speech actually sounds, as the human ear can detect subtle differences in pronunciation, accents, rhythm, and other elements, it has the disadvantage of requiring the hiring of specialized personnel such as language teachers and native speakers, which can lead to cost issues and complicated procedures for conducting the actual evaluation, such as coordinating schedules with such personnel.
[0008] Furthermore, in the evaluation of speech production using machines such as the above-mentioned AI, the characteristics of speech production can be captured by analyzing the resonant frequencies in the vocal tract, and this has become the mainstream of speech analysis, but currently, in order to take into account and evaluate all the elements of complex speech production in regular speech changes such as intonation and liaison, the elements related to language production are extremely diverse, making it difficult to cover everything with a single model.As a result, accurate evaluation requires the development of large amounts of training data and a wide range of advanced models, which requires specialized knowledge and skills, and the development itself is highly difficult, resulting in problems such as excessive costs for development and operation.
[0009] Therefore, the present invention has been made to solve such problems, and aims to provide a speech similarity calculation engine that can precisely capture regular speech changes in input speech when evaluating pronunciation in language learning such as shadowing, while suppressing increases in the costs of its development and operation, and to provide a learning support program, learning support system, and learning support method that allow learners to quickly improve their practical conversational ability.
[0010] In order to solve the above problem, the audio similarity calculation engine of the present invention is characterized by comprising: an audio data acquisition unit that acquires the first audio data and the second audio data; an outline extraction unit that normalizes amplitude changes in time-series sound pressure data of each of the first audio data and the second audio data and extracts an outline; an outline comparison unit that calculates a degree of similarity for each of the extracted outlines of the first audio data and the second audio data based on a pattern of change rate along time, and extracts parts with high degrees of similarity as similar parts; a vector sequence calculation unit that converts the degree of similarity and similar parts obtained by the outline comparison unit into vector sequences and calculates the Euclidean distance or cosine similarity of these vector sequences; and a similarity calculation unit that calculates the similarity between the first audio data and the second audio data based on the degree of similarity obtained by the outline comparison unit and the Euclidean distance or cosine similarity obtained by the vector sequence calculation unit.
[0011] The audio similarity determination engine of the present invention can be used in the learning support system and method of the present invention, and the learning support system and method of the present invention preferably include: a shadowing processing unit that executes, in parallel, a playback process that plays back teaching material audio that has been recorded in advance as teaching material, and a recording process that records audio data spoken by the learner as speech data; an audio input unit that uses the teaching material audio as the first speech data and inputs the speech data as the second speech data to the audio data acquisition unit; and a display data generation unit that generates display data for displaying the similarity calculated by the similarity calculation unit so that it can be compared with the speech position of text data corresponding to the teaching material audio.
[0012] In the learning support system and method of the present invention, it is preferable that a detailed calculation setting unit is further provided which sets, among single or consecutive words contained in the text data or the first audio data, a position where a regular sound change occurs as a characteristic utterance position, and that the display data generation unit generates the display data by including the characteristic utterance position in the text data.
[0013] Furthermore, in the learning support system or method using the audio similarity determination engine of the above invention, it is preferable to provide: a multilingual display unit that displays multilingual display data obtained by translating audio recorded in advance as learning material into a language different from the language of the learning material audio; an audio input unit that inputs the learning material audio as the first audio data and the vocalization data as the second audio data to the audio data acquisition unit; an utterance start determination unit that determines the length of the vocalization data up to the start of utterance; and a display data generation unit that generates display data for displaying the similarity calculated by the similarity calculation unit and the determination result regarding the start of utterance by the utterance start determination unit in a manner that allows them to be compared with the vocalization position in text data corresponding to the learning material audio.
[0014] Furthermore, the present invention provides a learning support program, system, and method for providing the learning materials in the above-mentioned learning support system or learning support method, characterized in that the above-mentioned learning support program is provided with: a reference unit that collects information related to at least one of various occupations, various hobbies, and various specialized fields as teacher data, and refers to a neural network trained by layering distribution patterns of features extracted from the collected teacher data; a specific learning material creation unit that compares, via the reference unit, the neural network with the distribution pattern of features extracted from learner information specific to the learner set by the learner, and extracts information related to the learner information from the teacher data as specific topic information based on the degree of agreement between the two features, and creates the learning material audio and learning material text related to this specific topic information as specific learning materials; and a lesson management unit that manages the progress of each learner's lesson based on a provision plan for the specific learning material created by the specific learning material creation unit.
[0015] The above-described system and method according to the present invention can be realized by executing a program of the present invention written in a predetermined language on a computer. That is, by installing the program of the present invention in an IC chip or memory device of a general-purpose computer such as a portable terminal device, smartphone, wearable terminal, mobile PC or other information processing terminal, personal computer or server computer, and executing it on the CPU, a system having the above-described functions can be constructed and the method according to the present invention can be implemented.
[0016] The program of the present invention can be distributed, for example, via a communication line, or transferred as a package application that runs on a stand-alone computer by recording it on a computer-readable recording medium. This recording medium can be recorded on a variety of recording media, including magnetic recording media such as flexible disks and cassette tapes, optical disks such as CD-ROMs and DVD-ROMs, and RAM cards. The computer-readable recording medium on which the program is recorded makes it possible to easily implement the above-described system and method using a general-purpose computer or a dedicated computer, and also makes it easy to store, transport, and install the program.
[0017] As described above, these inventions make it possible to provide a speech similarity calculation engine that can precisely capture regular phonetic changes in input speech, such as intonation, liaison (linking), and "patchim" in Korean, when evaluating pronunciation in language learning such as shadowing, thereby enabling learners to quickly improve their practical conversational ability while suppressing increases in development and operation costs.
[0018] Specifically, the speech similarity calculation engine executes a series of processes to calculate the similarity between the first and second speech data, acquiring speech data, extracting its outline, comparing the outlines, and extracting similar parts. This information is then converted into a vector sequence, and the similarity between the two speech data is specifically calculated using Euclidean distance and cosine similarity. This allows for precise capture of subtle elements of speech.
[0019] By applying the above-mentioned speech similarity calculation engine, the speech data spoken by the learner can be compared with the speech data from the teaching materials, and the similarity and accompanying information such as the word recognition status in the speech recognition model can be displayed to support shadowing learning. This allows the learner to specifically understand the differences between their own speech and the speech data from the teaching materials and to quickly make any necessary corrections.
[0020] Furthermore, according to the present invention, by identifying the positions where regular vocal changes, such as intonation changes and liaisons, occur in the audio material and presenting the judgment results for the learner's vocal data, it is possible to perform an evaluation that is specialized for training in regular vocal changes.
[0021] By applying the above-mentioned speech similarity calculation engine, it is possible to translate the learning material speech into a different language and display it, and then compare the speech data spoken by the learner with the learning material speech, thereby implementing pattern practice with a function for evaluating the similarity of the speech to the learning material speech. In this pattern practice, for example, by using a function for measuring the length of time from when the native language is displayed until the learner speaks, it is possible to train the learner to reflexively speak the foreign language after seeing the translation of the learning material speech to be spoken.
[0022] Furthermore, according to the present invention, a neural network can be used to collect information related to the characteristics of a learner based on information specific to the learner, and specific teaching materials can be created based on that information, making it possible to provide personalized teaching materials tailored to the learner's needs.
[0023] 1 is a block diagram showing the overall configuration of a learning assistance system according to an embodiment. FIG. 2 is a block diagram showing the internal configuration of a user terminal according to an embodiment, where (a) shows a case where a speech recognition processing unit is implemented on the user terminal side, and (b) shows a case where the speech recognition processing unit is implemented on the server side. FIG. 3 is a block diagram showing the internal configuration of a lesson execution unit according to an embodiment. FIG. 4 is a block diagram showing the internal configuration of a speech similarity calculation engine according to an embodiment. FIG. 5 is a block diagram showing the internal configuration of a learning assistance server according to an embodiment, where (a) shows a case where a speech recognition processing unit is implemented on the user terminal side, and (b) shows a case where the speech recognition processing unit is implemented on the server side. FIG. 6 is a flow diagram showing the operation of a learning assistance system according to an embodiment. FIG. 7 is a flow diagram showing the operation of a speech similarity calculation engine according to an embodiment. FIG. 8 is an explanatory diagram showing a display screen of a user terminal according to an embodiment. FIG. 9 is an explanatory diagram showing a processing concept related to speech waveforms and similarity calculation in a speech similarity calculation engine according to an embodiment. FIG. 10 is an explanatory diagram showing extraction of speech waveforms and one-sided data thereof in a speech similarity calculation engine according to an embodiment. FIG. 11 is an explanatory diagram showing speech waveforms and outline extraction thereof in a speech similarity calculation engine according to an embodiment. FIG. 12 is an explanatory diagram showing a method for calculating a degree of match based on a waveform change rate in a speech similarity calculation engine according to an embodiment. FIG. 13 is an explanatory diagram showing extraction of similar waveforms in a speech similarity calculation engine according to an embodiment. 1 is an explanatory diagram showing a regression process and noise removal in an audio similarity calculation engine according to an embodiment; FIG. 2 is an explanatory diagram showing vectorization of amplitude data in an audio similarity calculation engine according to an embodiment; FIG. 3 is an explanatory diagram showing a detailed calculation setting unit in an audio similarity calculation engine according to an embodiment;
[0024] The following describes in detail embodiments of the speech similarity calculation engine according to the present invention, and a learning support system, method, and program using the same, with reference to the accompanying drawings. The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above embodiments. For example, some components may be omitted from all of the components shown in the embodiments.
[0025] (Overview of Learning Support Service) In this embodiment, the present invention will be described as being applied to English vocabulary learning. The purpose of this learning support service is to use individual learning materials customized for each learner to repeatedly train them, focusing on speaking and listening, so that they can naturally acquire native English. The learners who will be using this system are expected to be diverse, including students and working adults, with a wide range of occupations, hobbies, and regions, and the system focuses on general listening and speaking training, and also incorporates training methods such as shadowing and pattern practice.
[0026] In particular, in this embodiment, a learning course is constructed that is suited to the unique characteristics, learning environment, and set goals of each learner, and is designed to allow learners to actively come into contact with English words that they encounter frequently by using teaching material content that is likely to interest each learner or that is adapted to the learner's needs.
[0027] Specifically, the learning content outline is that rather than speaking and listening in a predetermined order, students can stresslessly learn words they use frequently at work or words they like related to their hobbies. Therefore, the system provides the AI with information specific to each student, and creates learning materials based on genres and topics that allow students to learn in accordance with their specific hobbies and preferences. The entire learning process is structured so that students can prioritize learning English sentences and vocabulary that interest them most.
[0028] Learners become proficient by repeatedly performing shadowing and pattern practice, and the difficulty and format of the questions they are given can be adjusted to suit each learner's learning situation. If they can produce the words with a certain level of accuracy, their proficiency level increases, and once they reach a certain level, the questions are changed to more difficult formats as the lesson progresses. Conversely, if they answer incorrectly, their proficiency level decreases, and once they fall below a certain level, the questions are changed to less difficult formats. These methods enable adaptive learning that is optimized for each individual learner.
[0029] In particular, this embodiment provides various lesson-style modes, such as a "shadowing" mode and a "pattern practice" mode. For example, the "shadowing" mode is a learning method in which a learner listens to a recording of a model speaker of the target language, such as a native speaker, and repeatedly follows the audio like a "shadow." In this embodiment, the learner plays back the audio recorded in advance as a learning material, and while listening to the played-back audio, the learner imitates the audio and records the resulting audio as speech data. The learner's proficiency is evaluated based on the degree of similarity between the audio and the speech data, as well as on additional information such as the word recognition status of the speech recognition model.
[0030] On the other hand, the "pattern practice" mode is a learning method in which a learner is presented with a sentence in their native language, which they translate into the target language and then speak as is. In this embodiment, a speech recording of a sentence in their native language translated into English is pre-recorded as learning material audio. A text translated into a language other than the learning material audio, such as their native language, Japanese, is output as other-language display data, and the learner looks at the text and immediately speaks the translated text in the target language. This "pattern practice" mode not only improves real-time translation ability, but also allows the learner to naturally acquire the grammatical structure and phrase usage of the target language by repeating this learning method. Learning regular sound changes based on the results of speech similarity calculations also strengthens listening ability.
[0031] (Overall Configuration of Learning Support System) Fig. 1 is a conceptual diagram showing the overall configuration of a learning support system according to this embodiment. As shown in Fig. 1, the learning support system according to this embodiment is generally composed of smartphones 1 (1a, 1b) used by multiple learners U (Ua, Ub), who are learners, and a learning support server 3 installed on the Internet 2. In this embodiment, the smartphone 1 will be described as an example of an information processing terminal device, but in this embodiment, a personal computer, a tablet PC, or the like can be used as the information processing terminal device in addition to a smartphone.
[0032] In this embodiment, the learning assistance server 3 is a server that performs learning assistance progress processing, and can be realized by a single server device or a group of multiple server devices, with multiple functional modules virtually constructed on a CPU, and each functional module cooperating to perform various processes. In addition, this learning assistance server 3 can send and receive data via the Internet 2 using its communication function, and can present web pages through browser software using its web server function.
[0033] The smartphones 1 (1a and 1b) are portable information processing terminal devices that use wireless communication. The mobile phones communicate wirelessly with relay points such as wireless base stations 22, allowing users to receive communication services such as phone calls and data communications while moving. Examples of communication methods for the mobile phones include 3G (3rd Generation), LTE (Long Term Evolution), 4G, 5G, FDMA, TDMA, CDMA, W-CDMA, and PHS (Personal Handyphone System). In addition to communication functions, the smartphones 1 are equipped with various functions, such as a digital camera, an application software execution function, and a location information acquisition function using GPS (Global Positioning System) or the like.
[0034] The location information acquisition function is a function for acquiring and recording location information indicating the location of the device itself, and examples of this location information acquisition function include a method for detecting the location of the device itself using signals from satellites, such as GPS, and a method for detecting the location using the radio wave strength from a mobile phone wireless base station 22 or a Wi-Fi communication access point, as shown in Fig. 1. In this embodiment, the location information acquisition function is used, for example, to predict the address of a learner from the learner's current location and to suggest input character candidates when entering the address or the name of the elementary school the learner belongs to.
[0035] The smartphone 1 is equipped with a liquid crystal display as a display unit that displays information, as well as operation devices such as operation buttons that allow the learner to perform input operations, and the operation devices include a touch panel that is superimposed on the liquid crystal display and serves as an input unit that acquires operation signals through touch operations such as specifying coordinate positions on the liquid crystal display. Specifically, the touch panel is an input device that inputs operation signals through pressure or electrostatic detection through touch operations using the user's fingertip or pen, and is configured by superimposing a liquid crystal display that displays graphics and a touch sensor that receives operation signals corresponding to the coordinate positions of the graphics displayed on the liquid crystal display.
[0036] (Internal Structure of Each Device) Next, the internal structure of each device constituting the above-described learning support system will be described in detail. Figures 2 to 4 are block diagrams showing the internal structure of the smartphone 1 according to this embodiment, and Figure 5 is a block diagram showing the internal structure of the learning support server 3 according to this embodiment. Note that the term "module" used in the description refers to a functional unit configured by hardware such as a device or equipment, software having such functionality, or a combination of these, and for achieving a predetermined operation.
[0037] (1) Smartphone 1 Next, we will explain the internal configuration of the smartphone 1. As shown in Fig. 2(a) , the smartphone 1 includes a communication interface 11, an input interface 12, an output interface 13, an application execution unit 14, and a memory 15.
[0038] The communication interface 11 is a communication interface for performing data communication, and has the function of contactless communication such as wireless communication, and contact (wired) communication using a cable, adapter means, etc. The input interface 12 is a device for inputting user operations, such as a mouse, keyboard, operation buttons, or touch panel 12a. The output interface 13 is a device for outputting video and audio, such as a display or speaker. In particular, the output interface 13 includes a display unit 13a such as a liquid crystal display, and this display unit 13a is superimposed on the touch panel 12a, which is an input interface.
[0039] The memory 15 is a storage device that stores the OS (Operating System), firmware, programs for various applications, other data, etc., and stores a user ID that identifies the learner, learning assistance application data downloaded from the learning assistance server 3, as well as learning assistance data processed by the application execution unit 14. In particular, in this embodiment, the memory 15 stores scenario data, movie data used in the option mode, audio learning questions, recognition dictionary files, etc., acquired from the learning assistance server 3.
[0040] The application execution unit 14 is a module that executes applications such as a general OS, a learning support application, and browser software, and is usually realized by a CPU, etc. In this application execution unit 14, the voice similarity calculation engine and learning support program according to the present invention are executed, thereby virtually constructing a display data generation unit 141, a display control unit 142, a voice recognition processing unit 143, and a learning progress processing unit 144.
[0041] The voice recognition processing unit 143 is a module that recognizes the voice of learner U from the voice signal input through the input interface 12 and generates a control signal for the learning support application based on the voice recognition results, and in this embodiment has a voice similarity calculation engine 140 and a voice input unit 143a.
[0042] The audio input unit 143a is a module that provides an interface for appropriately acquiring teaching material audio recorded in advance as teaching material and audio data uttered by the learner, and inputting these audio data as first audio data and second audio data, respectively, to the audio similarity calculation engine 140. The audio input unit 143a according to this embodiment has a function of aligning the beginning portions of the teaching material audio and the audio data uttered by the learner as necessary in the shadowing mode, and inputting them to the audio data acquisition unit.
[0043] The audio similarity calculation engine 140 is an engine program for determining the similarity between first audio data and second audio data, and extracts the outline of amplitude changes in the sound pressure data of each audio data and calculates the degree of match based on the pattern of the change rate of these outlines. Furthermore, it has the function of extracting highly similar portions as similar parts and calculating Euclidean distance and cosine similarity. The audio similarity calculation engine 140 according to this embodiment is an audio similarity calculation program for determining the similarity between the first audio data and the second audio data, and by executing this program on the CPU of a smartphone or server device, a functional module such as that shown in FIG. 4 is virtually constructed.
[0044] Specifically, by executing the audio similarity calculation program, the following are constructed on the CPU of the information processing terminal device: a first audio data acquisition unit 140a and a second audio data acquisition unit 140b that acquire the first audio data and the second audio data, respectively; an utterance start determination unit 140c that executes a process for determining the utterance start position; an outline extraction unit 140d that extracts the outline of the amplitude change in the sound pressure data according to the time series of each of the first and second audio data; and an outline comparison unit 140e that calculates the degree of similarity based on the pattern of the change rate according to the time series for the outline of each of the extracted first and second audio data, and extracts parts with a high degree of similarity as similar parts.
[0045] The speech start determination unit 140c is a module that determines the speech start position (time) of speech data, particularly the presence and length of silent sections in pattern practice learning and shadowing learning, and provides accurate speech evaluation and feedback. The silent section determination process is performed, for example, by detecting sections where the signal energy falls below a specific threshold. Specifically, an integral value can be used to calculate the sum of squares of the signal amplitude (energy) over a certain time range and determine the speech start position based on this.
[0046] In this embodiment, the outline extraction unit 140d or the outline comparison unit 140e may be provided with a function for automatically aligning the start points of the speech data uttered by the learner through outline extraction and comparison processing when the start points of the learning material speech and the speech uttered by the learner are misaligned due to shadowing in the shadowing mode. Furthermore, the processing for aligning the start points of the speech data uttered by the learner and the learning material speech in this shadowing mode may be performed by a speech recognition model.
[0047] Furthermore, by executing the audio similarity calculation program, the CPU of the information processing terminal device is configured with a vector sequence calculation unit 140f that converts similar parts into vector sequences based on the extraction results by the outline comparison unit 140e and calculates the Euclidean distance or cosine similarity of these vector sequences, a similarity calculation unit 140g that calculates the similarity between the first and second audio data based on the degree of match and the Euclidean distance or cosine similarity, and an output unit 140h that outputs the audio similarity that is the calculation result. Note that the vector sequence calculation unit 140f in this embodiment converts similar parts into vector sequences based on the extraction results by the outline comparison unit 140e, but when vectorizing the audio sound pressure data, the vector sequence may be converted by including similarity calculation constants such as the DTW score and DDTW score calculated by the outline comparison unit 140e in the vector elements together with the similar parts, as necessary.
[0048] Here, the speech signal is preprocessed using noise reduction and normalization, then divided into frames of a fixed length (e.g., 10 ms to 30 ms), and a window function (e.g., a Hamming or Hann window) is applied to each frame to smooth discontinuities at the edges of the frame. The instantaneous energy of the signal in each frame is then squared and integrated (or summed). Frames with energy below a certain threshold are marked as silence. Post-processing procedures such as smoothing or a minimum duration filter may be applied to prevent short speech fragments from being mistakenly detected as silence.
[0049] Here, an example is shown in which the voice recognition processing unit is implemented on the smartphone 1 side; however, as shown in Figures 2(b) and 5(b), for example, the voice recognition processing unit may be implemented on the learning assistance server 3 or an external server device, and voice data input on the smartphone side may be sent to the external server side through synchronous processing via a communication interface, and voice recognition processing including voice similarity calculation may be executed on the server side, and the recognition processing results may be returned to the smartphone side.
[0050] The display data generation unit 141 is a module that generates display data to be displayed on the display unit 13a under the control of the display control unit 142. The display data is data that is generated by combining graphic data, image data, character data, video data, audio data, and other data. In particular, in this embodiment, the display data generation unit 141 has a function of generating display data for displaying the similarity calculated by the audio similarity calculation engine 140 in comparison with the utterance position of text data corresponding to the learning material audio during audio recognition in shadowing learning or pattern practice learning.
[0051] The display data generation unit 141 also has a function of generating display data by combining vocalization results according to various learning methods, such as the positions where regular sound changes occur in the text data and the determination of unvoiced sections in pattern practice learning and shadowing learning. The display control unit 142 includes a display mode switching unit 145a, which displays graphics according to the display mode.
[0052] The learning progress processing unit 144 is a module that progresses lessons in accordance with the "learning course," which is a teaching material provision plan created by the learning support server 3, and is equipped with a lesson execution unit 146 and a synchronization processing unit 145.
[0053] The lesson execution unit 146 is a module that, as a learning support service, progresses each lesson based on the learner's unique learning process. For example, the lesson execution unit 146 executes, as one of event processes, learning modes such as assignments for shadowing learning and assignments for pattern practice learning that are included in learning materials for evaluating the accuracy of pronunciation by learner U.
[0054] In this embodiment, the lesson execution unit 146 includes a shadowing processing unit 146a, a detailed calculation setting unit 146b, a teaching material management unit 146c, a schedule management unit 146d, a pattern practice learning control unit 146e, and a grade reflection unit 146f.
[0055] The shadowing processing unit 146a is a module that supports shadowing learning, in which a learner speaks while following the learning material audio being played. It plays learning material audio recorded in advance as learning material and records the speech data spoken by the learner as speech data. In this shadowing mode, if the start positions of the learning material audio and the speech spoken by the learner are misaligned due to shadowing, the speech input unit 143a aligns the start positions of the learning material audio and the speech data spoken by the learner as necessary and inputs them into the speech data acquisition unit. Note that the process of aligning the start positions of the speech by the speech input unit 143a can be omitted if the speech similarity calculation engine itself has a function for automatically aligning the start positions of speech using an outline extraction function or a similar part extraction function.
[0056] The detailed calculation setting unit 146b is a module that sets, as characteristic vocalization positions, positions in single or consecutive words in the text data where regular vocalization changes occur. For example, within single or consecutive words in the text data, it sets positions where regular vocalization changes occur, such as regular intonation changes, liaisons, or Korean consonants. When setting liaisons, for example, the detailed calculation setting unit 146b compares each word and phrase in the text data of the learning material audio to be used and connects words (zones) according to a certain rule (e.g., the end of the preceding word is a consonant and the first letter of the following word is a vowel), thereby labeling portions that have vocalization characteristics relative to the reference audio. In this way, using labeling data for characteristic vocalization positions such as intonation changes, liaisons, and consonants, the audio similarity calculation engine can evaluate the vocalization of those positions from the learning material audio and the comparison audio.
[0057] The learning material management unit 146c is a module that selects and manages learning materials to be used, manages learning material audio and learning material text, and provides appropriate learning materials. The schedule management unit 146d is a module that manages the learner's learning schedule and progress. The grade reflection unit 146f is a module that evaluates the learner's learning results and progress, and reflects and displays them as grades.
[0058] The pattern practice learning control unit 146e is a module that supports pattern practice learning by utilizing a speech similarity calculation engine. This pattern practice learning control unit 146e is a module that executes a learning method in which a learner looks at a sentence written in a language other than the target language, immediately translates it into the target language, and speaks it. Specifically, the pattern practice learning control unit 146e functions as a other language display unit that displays, as other language display data, teaching material audio in the language that the learner should speak as the target language, i.e., the target language (English in this case), and text in a language other than the language of the teaching material audio (Japanese in this case).
[0059] In this embodiment, learning material audio recorded in advance as learning material is used as a base, and the text data is translated to generate display data in a language other than the target language. However, the native language provided as text data may be translated into the target language (English in this case) and then vocalized to generate learning material audio. In this case, the learning material audio generated by vocalizing the translated text is input to the audio similarity calculation engine 140 as first audio data. In some cases, speech data spoken by the learner may be recognized and converted into text, and translated into a language other than the target language (Japanese in this case), without using the audio similarity calculation engine. The learner's performance may then be evaluated by comparing the text sentences in the native language provided as learning material with the transcript of the vocalized speech.
[0060] The speech input unit 143a treats the teaching material speech as the first speech data and acquires the learner's speech data as the second speech data for the pattern practice learning control unit 146e. The display data generation unit 141 also performs the function of generating display data that enables comparison of the speech positions of the text data corresponding to the teaching material speech based on the similarity calculated by the similarity calculation unit.
[0061] Furthermore, in this embodiment, the learning progress processing unit 144 generates and executes various learning tasks as event processing based on the learning process corresponding to the learning level of the learner, while synchronizing with the learning support processing unit 36 on the learning support server 3 side through the synchronization processing unit 145. That is, in this embodiment, the learning progress processing unit 144 cooperates with the learning progress processing unit 144 on the learning support server 3 side through the synchronization processing unit 145, so that part of the learning support progress processing is performed on the learning support server 3 side, and part of the graphic processing and event processing is performed by the learning progress processing unit 144 on the smartphone 1a side.
[0062] (2) Learning Support Server 3 Figure 5(a) shows the internal configuration of the learning support server 3. The learning support server 3 is a server device located on the Internet 2, and is capable of sending and receiving data to and from each smartphone 1 (1a, 1b) via the Internet 2.
[0063] The learning support server 3 includes a communication interface 31 for data communication via the Internet 2, an authentication unit 33 for authenticating the authority of users and user terminals, an access management unit 32 for managing access from each learner's information processing terminal device, a learning support processing unit 36 for executing learning support progress processing, a learning material providing unit 34 for distributing learning material data to each learner, and a group of various databases 35 (35a to 35c).
[0064] The database group 35 includes a teaching material database 35a that stores learning materials specific to each learner, a user database 35b that stores information about the learners, and a lesson management database 35c that stores various data necessary for managing lessons, such as learning history information as information about the learning progress process and each learner's learning progress process, and the schedules of registered instructors. Each of these databases may be a single database, or may be divided into multiple databases and a relational database in which each piece of data is linked together by setting relationships between them.
[0065] The learning material database 35a is a storage device that stores assignments and tests to be given during learning support. It stores audio data included in learning materials in a language other than the target language (here, English) as well as audio data included in learning materials in a language other than the target language (here, Japanese). In this embodiment, two types of audio data are used as learning materials because bidirectional translations between two languages (here, English and Japanese) are also used as questions to be recognized. However, depending on the format of the questions, the storage of one type of audio data may be omitted. Furthermore, three or more types of audio data may be implemented to support multiple languages. In particular, the learning material database 35a stores questions whose difficulty can be adjusted by changing the learning mode or question format for the same question, and the corresponding answers, in association with each other. In this embodiment, the learning materials included in the learning material database 35a are stored by classifying each question by genre.
[0066] The information stored in the user database 35b includes authentication information linked to an identifier (user ID or terminal ID) that identifies the learner or the information processing terminal device used by the learner, such as a password, as well as personal information about the user linked to the user ID and the model of the information processing terminal device. The user database 35b also functions as a learning history database, storing learning history information for each learner, by associating assessment result information indicating the question format and correct / incorrect assessment results when the question was asked with question time information indicating the date and time the question was answered. For example, the user database 35b stores, in addition to authentication history (access history) for each learner or user device, learning history for each user regarding the progress of learning support (learning level, status during the lesson, score, usage history, etc.), payment information related to learning support, etc., through its relationship with the lesson management database 35c.
[0067] The information stored in the lesson management database 35c includes lesson schedules and learning processes for the progress of learning support, information on various event processes, graphic information, and the like as lesson management data.
[0068] The authentication unit 33 is a module that establishes a communication session with each smartphone 1 via the communication interface 31 and performs authentication processing for each established communication session. This authentication processing involves obtaining authentication information from the smartphone 1 of the accessing learner, referencing the user database 35b to identify the learner, and authenticating their authority. The authentication results (user ID, authentication time, session ID, etc.) by the authentication unit 33 are sent to the learning support processing unit 36 and are also stored as authentication history in the user database 35b via the access management unit 32.
[0069] The learning support processing unit 36 is a module that generates various event processes to progress learning support in order to progress each learner's lesson in accordance with the lesson schedule, which is a plan for providing teaching materials, or the learning process.It executes a learning support program that includes certain rules, logic, or algorithms, and generates event processes for learning modes that conform to the teaching materials as the learning support progresses.
[0070] Here, the learning process is data described as an event generation process for gradually developing a lesson, including questions to be asked, the question format, question intervals, difficulty level, etc., in accordance with the progress of the learning support (the learner's learning level, the difficulty level specified according to that learning level, and the history of test scores), and is composed of a list of execution commands and arguments for the event processing to be generated on a time schedule (values specifying the questions to be asked and the question format, file names to be played, etc.).
[0071] The learning support processing unit 36 also serves as a learning history recording unit that records learning history information in the lesson management database 35c. That is, in this embodiment, the learning support processing unit 36 on the learning support server 3 side cooperates with the learning progress processing unit 144 on the smartphone 1 side through the synchronization processing unit 145, so that part of the learning support progress processing is performed on the learning support server 3 side, and part of the graphic processing and event processing is performed by the learning progress processing unit 144 on the smartphone 1 side. For example, the learning support server 3 side predicts possible event processing based on the generation and progress management of the learning process and lesson schedule, management of the learner's learning level, etc., generates the conditions for the event processing to occur on the learning support server 3 side, and transmits the conditions to the smartphone 1 side. The actual generation of the event processing and the graphic processing required for it are then performed on the smartphone 1 side based on the conditions for the event processing received from the learning support server 3.
[0072] The teaching material providing unit 34 is a module that distributes the teaching material data files generated by the specific teaching material creating unit 36b to each learner's smartphone 1 periodically or on demand via the communication interface 31 in order to synchronize these files with the smartphone 1 based on the learning level, characteristics, and expertise of each learner, under the control of the lesson management unit 36c.
[0073] Although the example shown here is one in which the speech recognition processing unit is implemented on the smartphone 1 side and all of the speech recognition processing is performed on the smartphone side, some or all of the modules related to the speech recognition processing may be implemented on the learning assistance server 3 side. For example, as shown in Figures 2(b) and 5(b), a module related to speech recognition processing may be implemented on the learning assistance server 3 side as a speech recognition processing unit 37, and speech data input on the smartphone side may be sent to the learning assistance server 3 side by synchronous processing via a communication interface, and speech recognition processing including speech similarity calculation may be performed on the speech recognition processing unit 37 side of the learning assistance server 3, and the recognition processing result may be returned to the smartphone side.
[0074] (Operation of the Learning Support System) The learning support method of the present invention can be implemented by operating the learning support system or learning support program described above. Figure 6 is a flow diagram showing the operation of the learning support system. Note that the processing procedure described below is merely an example, and each process may be modified as much as possible. Furthermore, steps in the processing procedure described below can be omitted, replaced, or added as appropriate depending on the embodiment.
[0075] 6, first, a learning assistance application is started on the smartphone 1 used by the learner, and the learning assistance server 3 is accessed, and an authentication procedure is executed (S201). If this authentication procedure is successful, the accessing learner is identified and permitted to log in (S101).
[0076] Next, as necessary, learner information, which is information specific to the learner, such as their occupation and hobbies, is sent from the smartphone 1 to the learning assistance server 3. Once the learner-specific information has been set, the neural network is referenced (S102) and learning materials specific to the learner are created (S103). When this specific learning material is created, a learning material provision plan is collated based on the learner's level, lesson progress, and test results. Note that setting the learner information and creating the specific learning material do not need to be performed every time, but can be performed as needed, such as for the first lesson or when settings need to be changed.
[0077] That is, if the learner's unique information has already been set and unique learning materials have been created, the application references the current learning history or lesson schedule, reads out the learner's current learning level (test results) and progress in the lesson schedule, references the provision plan according to the progress of learning support, and provides the appropriate learning materials (S104).The provision of this learning material does not need to be performed every time learning starts, and if the learning material has already been downloaded, the continuation of that learning material will be performed.
[0078] Next, the lesson execution unit 146 on the smartphone 1 analyzes the provided teaching materials and lesson schedule (S203), and when a dramatic event or a to-do event for the lesson is set according to the progress of the learning support, the event processing to be performed is called by referring to the provision plan, and messages and videos are displayed on the top screen, etc., and selectable lesson modes are displayed, prompting the learner to select the lesson mode they desire.
[0079] The learning method (lesson mode) to be executed is selected in response to a selection operation by the learner (S204). Note that these modes may be determined automatically based on the lesson schedule or teaching plan, in addition to the learner's selection operation. For example, explanations and reports by message display or video playback may be automatically started as optional modes depending on the progress of learning support, or an automatically selected lesson mode may be started as appropriate when a predetermined lesson branch point in the teaching plan is reached.
[0080] If shadowing learning is selected ("Shadowing" in S204), shadowing learning begins (S205), and the learning material audio is played back while the learner's voice is recorded in a manner that follows the audio. In this embodiment, the learning material audio is played back after the English sentence is displayed, but the shadowing learning method may also be such that, depending on the level, the learning material audio is played back without displaying the English sentence. On the other hand, if pattern practice learning is selected in step S204 ("Pattern Practice" in S204), pattern practice learning begins (S206), and the learning material Japanese sentence is displayed while the learner's voice is recorded as it corresponds to the Japanese sentence.
[0081] In the speech recognition process for the speech data uttered by the learner in each of these modes, the speech similarity calculation engine 140 calculates the similarity between the learning material speech and the speech data uttered by the learner (comparison speech) (S207). That is, in the speech-based learning methods (here, shadowing and pattern practice), the accuracy of the learner's speech is evaluated based on the similarity calculated by the speech similarity calculation engine 140 and accompanying information such as the word recognition status by the speech recognition model, and the correctness of the answer is determined and the score is calculated according to the evaluation.
[0082] Additionally, the locations of single or consecutive words in the text data where regular vocalization changes occur are designated as characteristic vocalization locations. Methods for detecting characteristic vocalizations include a human listening to the audio material to identify locations where characteristic vocalizations occur, or using a machine learning speech analysis model to identify locations in the audio material that differ from normal vocalizations and define those locations as characteristic vocalizations. For example, ELSA Speak (https: / / elsaspeak.com / en / product) can output the pronunciation of words recognized using an AI engine using phonetic symbols, thereby identifying characteristic vocalizations. For example, if the correct pronunciation of "apple" is "?apl," but the AI model recognizes the audio material as "?epl," it can be determined that a characteristic vocalization occurred there.
[0083] In the pattern practice according to the present embodiment, learning material audio recorded in advance as learning material is used as a base, and the textual data is translated to generate display data in a language other than the target language (Japanese in this case). However, for example, a language other than the target language provided as text data may be translated into the target language (English in this case) and then vocalized to generate learning material audio. In this case, the learning material audio generated by vocalizing the translated text is input to the audio similarity calculation engine 140 as first audio data. In some cases, speech data spoken by a learner may be recognized and converted into text, and the text may be translated into a language other than the target language (Japanese in this case), and the learner's performance may be evaluated by comparing the presented text sentence with the vocalized text sentence.
[0084] In this speech recognition, regular vocalization changes may be determined and labeled. For example, in determining vocalization changes such as liaison, the text data of the teaching material speech to be used may be compared with each word or phrase, and parts with vocal characteristics may be labeled with respect to the reference speech, such as linking words (zones) according to a certain rule (e.g., the preceding word ends with a consonant and the following word begins with a vowel), and this labeling data may be used to determine whether linking has been successful from the teaching material speech and the comparison speech, and the vocalization of that part may be evaluated.
[0085] These similarities and correct / incorrect judgments are then displayed as evaluation result information. At this time, display data is generated that also includes the characteristic utterance positions in the text data. The learning support processing unit 36 then records the results, together with the question format at the time the question was asked, for each question and for each learner as learning history information in the lesson management database 35c (S208).
[0086] Although the example shown here is one in which the speech recognition processing unit is implemented on the smartphone 1 side and all of the speech recognition processing is executed on the smartphone side, some or all of the modules related to the speech recognition processing may be implemented on the learning assistance server 3 side and executed on the learning assistance server 3 side. For example, as indicated by the dotted line in Figure 6, steps S207 and S208 related to the speech recognition processing may be executed on the learning assistance server 3 side. In this case, by synchronous processing via the communication interface, speech data input on the smartphone side is sent to the learning assistance server 3 side, and speech recognition processing including speech similarity calculation is executed on the speech recognition processing unit 37 side of the learning assistance server 3, and the recognition processing result is returned to the smartphone side.
[0087] The above-described processes S204 to S208 are repeated ("N" in S209) until the lesson is completed ("Y" in S209). After that, when the lesson is completed ("Y" in S209), the progress of the lesson is reported to the learning assistance server 3 as a termination process, and is recorded in the learning assistance server 3, completing the termination process.
[0088] (Audio Similarity Calculation Process) The above-described audio recognition process in the audio similarity calculation engine 140 is performed according to the following procedure. Fig. 7 shows the procedure for the audio recognition process in the audio similarity calculation engine 140, and Figs. 9 to 16 show the details of each process. In this embodiment, as shown in Fig. 9, the sound pressure changes of two pieces of audio data that are the subject of similarity comparison are compared to calculate the degree of match. In other words, the accuracy of the utterance is determined by comparing only the sound pressure changes, rather than analyzing the frequency components of the audio.
[0089] First, the learning material voice is input as the first voice data and the utterance data is input as the second voice data to the voice data acquisition unit of the voice similarity calculation engine 140 via the voice input unit 143a (S301). For the input first and second voice data, the voice start determination unit 140c of the voice similarity calculation engine 140 detects the position (time) of the first utterance in the voice data next to the voice data acquisition unit (S302).
[0090] Only the amplitude of one side of the sound pressure data of each sound is extracted, and the amplitude data of the sound pressure of the sound data is normalized within a certain range (S303). First, when stereo recorded sound data is basically recorded using two channels allocated to the left and right, only one of the channels needs to be used, so one of the upper and lower waveforms is extracted as shown in Figure 10. In the case of stereo recording, the upper and lower waveforms differ, while in the case of monaural recording, the upper and lower waveforms are symmetrical.
[0091] Next, the amplitude data of the sound pressure of the first and second audio data from which the amplitude of one side has been extracted is normalized within a certain range to conform to a certain standard. Note that the outline extraction unit of this embodiment uses only one channel of the audio data, but if the sound source includes multiple channels such as stereo or surround recording, differences will occur in the sound pressure data for each channel, but any channel may be extracted as long as the audio is recorded. If the sound source is a monaural recording, only one channel is recorded, so data selection processing is not required.
[0092] Each piece of audio data records sound pressure changes as amplitude, and since the reference sound pressure level varies depending on the equipment configuration and environment, comparison must be made when the sound pressure levels are consistent. To align the sound pressure levels of the two pieces of audio data, the data is normalized by dividing all the data by the maximum absolute value of the sound pressure using the following formula.
[0093] Thereafter, unnecessary noise is reduced or removed from each piece of audio data (S304). Because the recorded data contains white noise and the like due to the recording environment, equipment, etc., in this embodiment, a noise filter is used to reduce or remove sound pressure data below the mode. Note that, since outline extraction is assumed here, a filter that removes data below the mode (below the mode) is used as a noise removal method, but neural networks or other well-known means can also be used.
[0094] Next, the overall shape and structure of the sound pressure data are analyzed, and an outline is extracted (S305). Specifically, as shown in Fig. 11, a scanning window, which is a frame of a certain size and a certain amount of movement, is used to extract an outline from a sound pressure waveform that continues for a certain period of time. Here, scanning is started from a base point with a certain scanning interval, and the maximum value within the scanning window range is extracted in each scanning interval and stored as outline data. At this time, if the same location in the previous scanning interval reaches a maximum value in the next scanning interval, the value is not stored and scanning continues, and when scanning of the entire continuous waveform is completed, outline data of the continuous waveform can be extracted.
[0095] Regression data for performing regression analysis is created from the extracted outline data (S306). That is, since the outline data extracted in step S305 contains noise components that were not completely removed in step S304 and other steps, in this embodiment, Ridge regression is performed on the outline data to smooth the data and remove the noise components. This Ridge regression, also known as Tikhonov regularization, is a type of linear regression that is a normalization process that is particularly effective when there are many collinear explanatory variables or when the number of variables is greater than the number of samples, and introduces a regularization term to prevent overfitting of the model.
[0096] Specifically, the objective function (sum of squared residuals) of a normal linear regression is minimized by adding the L2 norm of the parameters (the sum of the squares of each parameter). As shown in FIG. 14, this regularization term removes or reduces minute vibrations (noise) from the outline data before the regression process, as shown in FIG. 14(a), as shown in FIG. 14(b). This prevents the absolute values of the regression coefficients from becoming too large, improving the generalization performance of the model. While the present invention illustrates a case where Ridge regression is performed on the outline data, the present invention is not limited to this. Other regression algorithms with equivalent accuracy can also be used, such as: polynomial regression; nonlinear regression and kernel regression using basis expansion; lasso regression; and support vector regression (SVR).
[0097] Next, the regression data created in step S306 is used to calculate the degree of similarity between the waveforms of the two pieces of speech data as a score using the DTW method or DDTW (Derivative Dynamic Time Warping) method, and scoring is performed in which this score is defined as a "degree of similarity 1" (S307). As a result, as shown in Figure 12, an evaluation can be obtained that comprehensively determines the difference in time axis and the difference in speech change. Note that in this embodiment, the DTW method or DDTW method is used as an example to calculate the degree of similarity, but the present invention is not limited to this. For example, a method can be used in which specific parts of speech data are extracted in order and stored as vector sequences, and when the Euclidean distance or cosine similarity exceeds a threshold, they are set as "similar parts."
[0098] The DDTW method is a technique for calculating similarity using the rate of change (derivative or differential) of time-series data, and performs comparison by focusing on the changes and movements of the time-series data rather than the shape or pattern of the original data. Specifically, the DDTW method performs calculations using the differential value (rate of change) of the time-series data, and performs matching based on the change pattern of the data rather than the absolute position or value of the data, thereby reducing the influence of noise contained in the original data.
[0099] For example, the two pieces of speech data (regression data) created in step S306 are each differentiated to calculate the rate of change, and a normal DTW (Dynamic Time Warping) algorithm is applied to this differentiated data to find the optimal matching path. This DTW algorithm is an algorithm for calculating the similarity between time series data, and can evaluate the similarity even when the two pieces of time series data are stretched or contracted over time.
[0100] Specifically, the distance between each point of two regression data is calculated to create a matrix, and dynamic programming is used to calculate the minimum cumulative distance based on the distance matrix. Based on this minimum cumulative distance, the optimal matching path between the two time series data is found. This dynamic programming decomposes the problem into small subproblems, creates a table or data structure to store the solutions to the subproblems, and combines the stored solutions to the subproblems to construct an optimal solution for the overall problem.
[0101] Similarly, as shown in Fig. 13, the DDTW method is used to extract particularly similar portions of the two pieces of speech data (S308). Here, corresponding points are found from the two matching paths found in step S308, and two similar waveforms are extracted from each point that corresponds to the original data.
[0102] The point cloud data of the similar waveforms of each voice data extracted in step S308 is stored in a vector sequence (S309). Because the number of point cloud data of the similar waveforms extracted in step S308 is the same between the two voices, the two vectors have the same dimension, and it is possible to define the similar parts of the two voices as vector positions of the same dimension in time space, as shown in Figure 15.
[0103] Thereafter, the Euclidean distance and cosine similarity between the two vector sequences stored in step S309 are calculated, and these are defined as match 2 and match 3 (S310), and the similarity is calculated based on the defined match 1 to match 3 (S311). Since the Euclidean distance and cosine similarity between the two speech data vectorized in step S309 are calculated and similar parts of the two waveforms are extracted, it is possible to calculate the similarity between the two speeches without adjusting for differences in utterance timing using a speech recognition model or the like.
[0104] If two speech vectors are A and B, the Euclidean distance and cosine similarity are expressed by the following equations.
[0105] Finally, the speech is evaluated (scored) by comprehensively evaluating the matching score 1 (DDTW score) calculated in step S307 and the matching scores 2 (Euclidean distance) and 3 (cosine similarity) calculated in step S310. When calculating this comprehensive evaluation, since the matching scores 1 to 3 have judgment characteristics, the score evaluating the regular sound changes is calculated on three axes (three-dimensional space) combining the matching scores 1 to 3. In this embodiment, for example, a final evaluation may be given on a three-level scale, such as "very good," "good," or "a little better," based on a three-dimensional evaluation score and a three-dimensional threshold. This threshold setting can be achieved by comparing the matching score between the actual evaluation score and a human-based evaluation. Here, the scoring process can be carried out by "setting a threshold (classifier setting) based on three axes (three matching points)" to "performing a three-level evaluation." However, the classifier accuracy correction process can also be implemented using machine learning based on human feedback.
[0106] (Actions and Effects) As described above, according to this embodiment, when evaluating pronunciation in language learning such as shadowing and pattern practice, a speech similarity calculation engine is provided that can precisely capture subtle speech elements such as intonation and linking while minimizing increases in development and operation costs, enabling learners to quickly improve their practical conversational ability.
[0107] In more detail, this embodiment does not require the creation of a separate, special speech recognition model developed by machine learning or the like, but instead uses an engine that measures the similarity of a learner's speech to a model speech, i.e., a learning material speech. This avoids the need to collect and select training data tailored to the purpose, and then train the data, as is the case with developing a speech recognition model, thereby significantly reducing implementation costs. Furthermore, it is possible to comprehensively determine elements that are prominent as sound pressure elements, such as intonation and linking, without relying on a model.
[0108] In this way, in this embodiment, the similarity between the learning material voice and the learner's voice is calculated, so that basically any vocal phenomenon that appears as a change in sound pressure can be judged as correct or incorrect. Therefore, it is possible to judge whether the vocalization has the same content (word sequence) as the reference voice for a variety of languages, not just English, such as Chinese, Swahili, and Thai.
[0109] Furthermore, according to this embodiment, it can be used as a shadowing tool or a pattern practice tool. To improve listening skills in English learning, it is necessary to train the ability to recognize words, and in particular, it is necessary to improve elements that should be referenced in actual usage, such as the pronunciation of linked words, such as linking.
[0110] The system of this embodiment can utilize shadowing, a learning method in which a learner listens to learning material audio and produces the same pronunciation as the audio, to improve listening ability. For example, as shown in Figure 16, by comparing each word or phrase, and linking words (zones) to determine whether the reference audio has regular pronunciation changes, such as linking, the audio data is labeled, and the relevant parts are extracted and evaluated from the reference audio and comparison audio using the DDTW method and this labeling data, making it possible to determine whether the shadowing has regular pronunciation changes.
[0111] The above-described embodiments are merely examples of the present invention. Therefore, the present invention is not limited to the above-described embodiments, and various modifications can be made depending on the design, etc., as long as they do not deviate from the technical concept of the present invention. For example, the above-described voice similarity calculation engine 140 can be applied to shadowing and pattern practice training, as well as imitation practice of everyday conversation, support for multilingual learning, and tracking of changes in singing practice and vocalization.
[0112] Furthermore, the present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining the multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments.
[0113] DESCRIPTION OF THE REFERENCE NUMERALS U...Learner 1...Smartphone 2...Internet 3...Learning support server 11...Communication interface 12...Input interface 12a...Touch panel 13...Output interface 13a...Display unit 14...Application execution unit 15...Memory 22...Wireless base station 31...Communication interface 32...Access management unit 33...Authentication unit 34...Learning material provision unit 35...Database group 35a...Learning material database 35b...User database 35c...Lesson management database 36...Learning support processing unit 36a...Neural network reference unit 36b...Unique teaching material creation unit 36c...Lesson management unit 140...Voice similarity calculation engine 141...Display data generation unit 142...Display control unit 143...Voice recognition processing unit 143a...Voice input unit 144...Learning progress processing unit 145...Synchronization processing unit 145a...Display mode switching unit 146...Lesson execution unit 146a...Shadowing processing unit 146b...Detailed calculation setting unit 146c...Learning material management unit 146d...Schedule management unit 146e...Pattern practice learning control unit 146f...Score reflection unit
Claims
1. A program for determining a similarity between first voice data and second voice data, comprising: a voice data acquisition unit that acquires the first voice data and the second voice data; an outline extraction unit that normalizes amplitude changes in time-series sound pressure data of each of the first voice data and the second voice data and extracts outlines of the data; an outline comparison unit that calculates a degree of similarity for each of the extracted outlines of the first voice data and the second voice data based on a pattern of change rate over time and extracts parts with a high degree of similarity as similar parts; a vector sequence calculation unit that converts the similar parts obtained by the outline comparison unit into vector sequences and calculates a Euclidean distance or a cosine similarity of the vector sequences; and a similarity calculation unit that calculates the similarity between the first voice data and the second voice data based on the degree of similarity obtained by the outline comparison unit and the Euclidean distance or the cosine similarity obtained by the vector sequence calculation unit.
2. A learning support program using the voice similarity determination engine described in claim 1, characterized in that the learning support program causes a computer to function as: a shadowing processing unit that executes in parallel a playback process for playing teaching material voice recorded in advance as teaching material, and a recording process for recording voice data uttered by a learner as speech data; a voice input unit that treats the teaching material voice as the first speech data and inputs the speech data as the second speech data to the speech data acquisition unit; and a display data generation unit that generates display data for displaying the similarity calculated by the similarity calculation unit in a manner that allows it to be compared with the utterance position of text data corresponding to the teaching material voice.
3. A learning support program using the voice similarity judgment engine described in claim 1, characterized in that the learning support program causes a computer to function as: a other language display unit that displays other language display data translated into a language different from the language to be studied; a voice input unit that inputs the teaching material voice in the language to be studied as the first voice data and the spoken data as the second voice data to the voice data acquisition unit; a speech start judgment unit that judges the length of time until the speech start of the spoken data; and a display data generation unit that generates display data for displaying the similarity calculated by the similarity calculation unit and the judgment result regarding the start of speech by the speech start judgment unit in a manner that allows it to be compared with the speaking position of text data corresponding to the teaching material voice.
4. The learning support program according to claim 2 or 3, further comprising causing the computer to function as a detailed calculation setting unit which sets, among single or consecutive words contained in the text data, points at which regular changes in pronunciation occur as characteristic utterance positions, and wherein the display data generation unit generates the display data by including the characteristic utterance positions in the text data.
5. A learning support program for providing said teaching materials to the learning support program according to claim 2 or 4, characterized in that the computer is made to function as: a reference unit which collects information relating to at least one of various occupations, various hobbies, and various specialized fields as teacher data, and refers to a neural network trained by stacking distribution patterns of features extracted from the collected teacher data; a unique teaching material creation unit which compares, through said reference unit, a distribution pattern of features extracted from learner information unique to the learner set by the learner with said neural network, and extracts information related to the learner information from the teacher data as unique topic information based on the degree of agreement between the two features, and creates the teaching material audio and teaching material text related to this unique topic information as unique teaching materials; and a lesson management unit which manages the progress of each learner's lesson based on a provision plan for the unique teaching materials created by said unique teaching material creation unit.
6. A learning support system using the voice similarity determination engine described in claim 1, comprising: a shadowing processing unit which executes in parallel a playback process for playing teaching material voice recorded in advance as teaching material and a recording process for recording voice data uttered by a learner as voice data; a voice input unit which treats the teaching material voice as the first voice data and inputs the voice data to the voice data acquisition unit as the second voice data; and a display data generation unit which generates display data for displaying the similarity calculated by the similarity calculation unit in comparison with the utterance position of text data corresponding to the teaching material voice.
7. A learning support system using the voice similarity judgment engine described in claim 1, comprising: a other language display unit that displays other language display data translated into a language different from the language to be learned; a voice input unit that inputs the teaching material voice in the language to be learned as the first voice data and the spoken data as the second voice data to the voice data acquisition unit; a speech start judgment unit that judges the length until the start of speaking of the spoken data; and a display data generation unit that generates display data for displaying the similarity calculated by the similarity calculation unit and the judgment result regarding the start of speaking by the speech start judgment unit in a manner that allows comparison with the speaking position of text data corresponding to the teaching material voice.
8. The learning support system according to claim 6 or 7, further comprising a detailed calculation setting unit which sets, among single or consecutive words contained in the text data, points at which regular changes in pronunciation occur as characteristic utterance positions, and the display data generation unit generates the display data by including the characteristic utterance positions in the text data.
9. A learning support system for providing said learning materials to the learning support program according to claim 6 or 8, comprising: a reference unit which collects information relating to at least one of various occupations, various hobbies, and various specialized fields as teacher data, and refers to a neural network trained by stacking distribution patterns of features extracted from the collected teacher data; a unique teaching material creation unit which compares, via said reference unit, the distribution pattern of features extracted from learner information unique to the learner set by the learner with said neural network, and extracts information related to the learner information from the teacher data as unique topic information based on the degree of agreement between the two features, and creates the teaching material audio and teaching material text related to this unique topic information as unique teaching materials; and a lesson management unit which manages the progress of each learner's lesson based on a provision plan for the unique teaching material created by said unique teaching material creation unit.
10. A learning support method using the voice similarity determination engine described in claim 1, comprising: a shadowing processing step in which a shadowing processing unit executes in parallel a playback process for playing teaching material voice recorded in advance as teaching material and a recording process for recording voice data spoken by a learner as speech data; a voice input step in which a voice input unit inputs the teaching material voice as the first speech data and the speech data as the second speech data to the voice data acquisition unit; and a display data generation step in which a display data generation unit generates display data for displaying the similarity calculated by the similarity calculation unit in a manner that allows comparison with the speech position of text data corresponding to the teaching material voice.
11. A learning support method using the voice similarity determination engine described in claim 1, comprising: a other language display step in which a other language display unit displays other language display data translated into a language different from the language to be learned; a voice input step in which a voice input unit inputs the teaching material voice of the learning language as the first voice data and the spoken data as the second voice data to the voice data acquisition unit; a speech start determination step in which a speech start determination unit determines the length until the start of speech of the spoken data; and a display data generation step in which a display data generation unit generates display data for displaying the similarity calculated by the similarity calculation unit and the determination result regarding the speech start portion in a manner that allows comparison with the speaking position of text data corresponding to the teaching material voice.
12. A learning support method as described in claim 10 or 11, further comprising a detailed calculation setting step in which a detailed calculation setting unit sets, among single or consecutive words contained in the text data, a location where regular changes in pronunciation occur as a characteristic speech position, and in the display data generation step, the display data generation unit generates the display data by including the characteristic speech positions in the text data.
13. A learning support method for providing said learning materials to the learning support program according to claim 10 or 12, comprising: a reference step in which a reference unit refers to a neural network trained by collecting information related to at least one of various occupations, various hobbies, and various specialized fields as teacher data and stacking distribution patterns of features extracted from the collected teacher data; a unique teaching material creation step in which a unique teaching material creation unit compares, via the reference unit, the distribution pattern of features extracted from learner information unique to the learner set by the learner with the neural network, and extracts information related to the learner information from the teacher data as unique topic information based on the degree of agreement between the two features, and creates the teaching material audio and teaching material text related to this unique topic information as unique teaching materials; and a lesson management step in which a lesson management unit manages the progress of each learner's lesson based on a provision plan for the unique teaching material created by the unique teaching material creation unit.
Citation Information
Patent Citations
Pronunciation grading device, and program
JP2006201491A
Speech evaluating device
JP2007163976A
Musical piece practice assisting device, dynamic time warping module, and program
JP2008040258A
Method and system for automatically generating fill-in-the-blank questions of foreign language sentence
KR102189894B1
Foreign matter selection and sterilization system for marine raw material of salted seafood
KR102465187B1