Voice conversion device, voice conversion method, and voice conversion program
The voice conversion device addresses the issue of incorrect foreign language pronunciation by replacing specific voice data portions with native-like audio, enhancing language learning accuracy.
Patent Information
- Application Number
- JP2024002877
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing language education devices struggle to accurately convey the correct pronunciation of a foreign language when an instructor's pronunciation differs from that of a native speaker, leading to potential miscommunication in language learning.
A voice conversion device and method that acquires first voice data, identifies and replaces specific language type portions with corresponding second voice data, using features of a native speaker's pronunciation, and outputs the converted voice along with character information.
Enables accurate conversion of an instructor's voice to native-like pronunciation, improving language learning by ensuring correct pronunciation is conveyed to students.
Smart Images

Figure 2025109140000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a voice conversion device, a voice conversion method, and a voice conversion program, and particularly to a voice conversion device that processes voices uttered in a language that is not a person's native language, a voice conversion method executed by the voice conversion device, and a voice conversion program that causes a computer to execute the voice conversion method.
Background Art
[0002] In recent years, in learning where an instructor teaches a foreign language to students, the instructor teaches using two languages, the native language and the foreign language. For this reason, the voice uttered by the instructor contains a mixture of the native language voice and the foreign language voice. There are cases where the pronunciation of the foreign language by the instructor is different from the pronunciation of a person whose native language is that foreign language. A technique for recognizing a voice in which a native language voice and a foreign language voice are mixed is known. For example, Japanese Patent Application Laid-Open No. 3-226785 discloses a language education device having a symbol-based conversation education processing unit that inputs a symbolized input sentence and performs response processing for language training. The language education device includes a voice recognition device that inputs a voice input sentence and outputs a pronunciation symbol string in word units and a word symbol string used when indicating a normal word, a voice recognition error correction table that stores correct words and error-prone pronunciation symbol patterns of those words, and a pronunciation correction processing unit that inputs the pronunciation symbol string and the word symbol string, collates the pronunciation symbol string with the pronunciation symbol pattern, and outputs candidates for correct words in the voice recognition error correction table based on the collation result and provides them to the symbol-based conversation education processing unit in the form of a symbolized input sentence. A language education device with a voice recognition device characterized by the above is known.
[0003] According to the language education device described in Japanese Patent Application Laid-Open No. 3-226785, voice recognition can be performed even when the pronunciation of a foreign language by an instructor is different from the pronunciation of a person whose native language is that foreign language. However, there is a problem that when the instructor cannot pronounce in the pronunciation of a person whose native language is the foreign language, the correct pronunciation of the foreign language may not be conveyed to the students.
Prior Art Documents
Patent Document
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] One object of the present invention is to provide an audio conversion device capable of converting the audio uttered by a user into audio of pronunciation corresponding to the language type.
[0006] Another object of the present invention is to provide an audio conversion method capable of converting the audio uttered by a user into audio of pronunciation corresponding to the language type.
[0007] Still another object of the present invention is to provide an audio conversion program capable of converting the audio uttered by a user into audio of pronunciation corresponding to the language type.
Means for Solving the Problems
[0008] According to an aspect of the present invention, an audio conversion device includes an audio acquisition unit that acquires first audio data corresponding to the audio uttered by a user, and a replacement unit that replaces a replacement target portion corresponding to a predetermined language type in the first audio data with second audio data different from the first audio data.
[0009] According to another aspect of the present invention, an audio conversion method causes an audio conversion device to execute an audio acquisition step of acquiring first audio data corresponding to the audio uttered by a user, and a replacement step of replacing a replacement target portion corresponding to a predetermined language type in the first audio data with second audio data different from the first audio data.
[0010] According to still another aspect of the present invention, a voice conversion program causes a computer to execute a voice acquisition step of acquiring first voice data corresponding to the voice uttered by a user, and a replacement step of replacing a portion to be replaced corresponding to a predetermined language type in the first voice data with second voice data different from the first voice data.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the following description, the same parts are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions thereof will not be repeated.
[0013] FIG. 1 is a diagram showing an example of the configuration of a foreign language learning system according to one embodiment of the present invention. Referring to FIG. 1, the foreign language learning system 1 includes a server 100 and personal computers (hereinafter referred to as "PCs") 200, 200-1 to N. However, N is a positive integer. N is 1 or more. The server 100 and each of the PCs 200-1 to N are connected to the Internet 5 and can communicate with each other.
[0014] Each of the PCs 200, 200-1 to N is equipped with a camera, a microphone for collecting sound, and a speaker for outputting sound. Each of the PCs 200, 200-1 to N is a general computer, and its main hardware configuration and functions are the same. Therefore, unless otherwise specified, the PC 200 will be described as an example.
[0015] Note that instead of the PCs 200-200, 1 to N, an information communication device such as a PDA (Personal Digital Assistance) or a smartphone may be used as long as it is equipped with a camera, a microphone, a speaker, and a communication function. Also, not limited to the Internet 5, other networks may be used as long as the server 100 and the PCs 200, 200-1 to N can communicate with each other. The network may be, for example, a local area network (LAN) or a wide area network (WAN).
[0016] In the foreign language learning system 1 according to the present embodiment, a case where a Japanese native speaker instructor holds a lecture to teach English to a plurality of students will be described as an example. In this case, the instructor operates the PC 200, and a plurality of students operate any one of the PCs 200-1 to N.
[0017] For each of PC200-200,1 to N, a program for participating in a foreign language learning lecture is installed. By each of PC200-200,1 to N communicating with server 100, a lecture is held. The program installed in each of PC200,200-1 to N may be a dedicated program for communicating with server 100, or may be a general browser program if server 100 provides a web service.
[0018] Server 100 realizes the foreign language learning system 1 by executing a voice conversion program. Server 100 communicates with PC200,200-1 to N and transmits the data received from each of PC200,200-1 to N to the other PC200,200-1 to N.
[0019] The data transmitted and received between each of PC200,200-1 to N and server 100 includes voice data indicating voice, image data indicating an image, and application data. The image includes a still image and a moving image. The data transmitted and received between each of PC200,200-1 to N and server 100 may be compressed data or uncompressed data.
[0020] Server 100 coordinates the data to be transmitted to each of PC200,200-1 to N. For example, server 100 can transmit the voice data received from each of PC200,200-1 to N to all the others of PC200,200-1 to N. Here, mainly, the voice data received from the PC200 operated by the instructor is transmitted to PC200-1 to N operated by one or more students.
[0021] Server 100 determines and transmits the voice data, image data, and application data to be transmitted in response to requests from each of the PCs 200, 200-1 to N. Therefore, the images displayed on each of the PCs 200, 200-1 to N may be the same or different, and the voices output from each of the PCs 200, 200-1 to N may be the same or different. Note that each of the PCs 200, 200-1 to N may process and output the voice data received from Server 100. In this case, the process of processing voice data by Server 100 becomes unnecessary, so the load is reduced. Also, each of the PCs 200, 200-1 to N may process and display a plurality of image data received from Server 100. In this case, the process of processing a plurality of image data by Server 100 becomes unnecessary, so the load is reduced.
[0022] Figure 2 is a block diagram showing an example of the hardware configuration of the server. Referring to Figure 2, Server 100 is a computer that performs arithmetic processing, and includes a central processing unit (CPU) 101 for controlling the entire Server 100, which is connected to a bus 113 respectively, a ROM (Read Only Memory) 102 that stores programs to be executed by the CPU 101, a RAM (Random Access Memory) 103 used as a working area of the CPU 101, an HDD 104 that stores data non-volatilely, a communication unit 105 that connects the CPU 101 to the Internet 5, a display unit 106 that displays images, an operation unit 107 that accepts input of operations, and an external storage device 110.
[0023] The communication unit 105 is an interface for connecting Server 100 to the Internet 5. Therefore, the CPU 101 can communicate with the PCs 200-1 to 100N connected to the Internet 5 via the communication unit 105.
[0024] The external storage device 110 has a CD-ROM (Compact Disk Read Only Memory) 111 attached thereto. The CPU 101 controls the external storage device 110 to read out the data stored in the CD-ROM 111.
[0025] In the present embodiment, the CPU 101 executes a program stored in the ROM 102 or the HDD 104. Further, the CPU 101 may control the external storage device 110 to read out a program for the CPU 101 to execute from the CD-ROM 111, store the read program in the RAM 103, and execute it.
[0026] Furthermore, the CPU 101 downloads a program from a computer connected to the Internet 5 and stores it in the HDD 104. Also, when a computer connected to the Internet 5 writes a program to the HDD 104, the program is stored in the HDD 104. The CPU 101 may load the program stored in the HDD 104 into the RAM 103 and execute it.
[0027] Note that the recording medium for storing the program to be executed by the CPU 101 is not limited to the CD-ROM 111, and may be a medium such as a flexible disk, a cassette tape, an optical disk (MO (Magnetic Optical Disc) / MD (Mini Disc) / DVD (Digital Versatile Disc)), an IC card, an optical card, a mask ROM, a semiconductor memory such as an EPROM (Erasable Programmable ROM), etc. The program referred to here includes not only a program directly executable by the CPU 101, but also a source program, a compressed program, an encrypted program, and the like.
[0028] FIG. 3 is a block diagram showing an example of the hardware configuration of a PC. Referring to FIG. 3, PC 200 is a computer that performs arithmetic processing, and includes a CPU 201 for controlling the entire PC 200, which is connected to a bus 212, a ROM 202 for storing a program to be executed by the CPU 201, a RAM 203 used as a work area for the CPU 201, an HDD 204 for storing data non-volatilely, a communication unit 205 for connecting the CPU 201 to the Internet 5, a display unit 206 for displaying images, an operation unit 207 for receiving input of operations by a participant who is a user, a camera 208 for imaging the participant, a speaker 209 for outputting sound, a microphone 210 for collecting the voice of an operator, and an external storage device 211.
[0029] An external storage device 211 has a CD-ROM 211A mounted thereon. The CPU 201 controls the external storage device 211 and reads out data stored in the CD-ROM 211A.
[0030] A module in which at least two of the camera 208, the speaker 209, and the microphone 210 are integrated may be connected to the PC 200. The module includes, for example, a headset in which the speaker 209 and the microphone 210 are integrated.
[0031] FIG. 4 is a diagram showing an example of functions of the CPU included in the server in the present embodiment. The functions shown in FIG. 4 are functions realized by the CPU 101 included in the server 100 by executing a voice conversion program stored in the ROM 102, the HDD 104, or the CD-ROM 111.
[0032] Referring to FIG. 4, the CPU 101 includes an audio acquisition unit 11, a replacement part determination unit 13, a replacement unit 15, an output unit 17, and a history information generation unit 19. The audio acquisition unit 11 acquires first audio data corresponding to the voice uttered by the instructor. The voice uttered by the instructor is collected by the microphone 210 of the PC 200 operated by the instructor. The PC 200 converts the voice collected by the microphone 210 into first audio data. The PC 200 controls the communication unit 205 to transmit the first audio data to the server 100. The audio acquisition unit 11 controls the communication unit 105 to acquire the first audio data received by the communication unit 105 from the PC 200. The audio acquisition unit 11 outputs the first audio data to the replacement part determination unit 13.
[0033] The replacement part determination unit 13 determines a replacement target part from the first audio data. The replacement part determination unit 13 includes an audio recognition unit 21, a language type classification unit 23, an extraction unit 25, a determination unit 27, a related information acquisition unit 29, a foreign word acquisition unit 31, and a probability output unit 33. The audio recognition unit 21 performs audio recognition on the audio data and outputs character information to the language type classification unit 23.
[0034] The related information acquisition unit 29 acquires related information related to the audio data. The related information includes information related to the instructor who uttered the voice of the first audio data. The related information includes materials distributed to the students during the instructor's guidance and characters obtained from the image of the image data obtained by photographing the situation where the instructor is uttering the voice. For example, the related information acquisition unit 29 acquires the data of the materials distributed by the instructor for teaching a foreign language from the PC 200. Also, the related information acquisition unit 29 acquires the image data obtained by imaging with the camera 208 of the PC 200 from the PC 200. The related information acquisition unit 29 extracts characters from the image of the image data by analyzing the image data and determines them as related information. For example, there may be a case where the instructor writes words or sentences in a foreign language on the blackboard. The related information acquisition unit 29 extracts the words or sentences written in a foreign language on the blackboard from the image data obtained by photographing the blackboard as related information. The related information acquisition unit 29 outputs the related information to the audio recognition unit 21.
[0035] The voice recognition unit 21 may not be able to recognize a single word from the voice. For example, there may be a plurality of different words for a single voice. For example, in Japanese, there are homophones, and in English, there are a plurality of words with similar pronunciations. For example, "read" and "lead" have similar pronunciations. The voice recognition unit 21 refers to the related information and determines one from among the plurality of word candidates. For example, the voice recognition unit 21 determines, among the plurality of candidates, the same candidate as the word included in the material distributed by the instructor to teach a foreign language, or the same candidate as the word included in the characters extracted from the image by analyzing the image data.
[0036] The language type classification unit 23 analyzes the character information and classifies each of the plurality of words included in the character information into any one of a plurality of language type groups. When the language types are predetermined, the plurality of predetermined language types may be given to the language type classification unit 23. The language type classification unit 23 divides the character information into a plurality of words, identifies each of the plurality of words, and identifies the language type for each of the plurality of words. Here, the first language type is Japanese and the second language type is English. The language type classification unit 23 classifies each of the plurality of words included in the character information into either the group of the first language type or the group of the second language type. Note that the number of language types may be three or more.
[0037] The foreign word acquisition unit 31 acquires foreign word information. The foreign word information defines a foreign word of the first language type corresponding to the word of the second language type. Here, since English is taught to students whose native language is Japanese, the foreign words are words that have become established as Japanese words incorporated from English into Japanese. The foreign word information that defines the foreign words is stored in advance in the HDD 104. The foreign word acquisition unit 31 acquires the foreign word information stored in the HDD 104. The foreign word acquisition unit 31 outputs the foreign word information to the language type classification unit 23.
[0038] The language type classification unit 23 identifies the language type for a word by referring to foreign language information. When a word corresponds to both Japanese and English and is a foreign word, the language type classification unit 23 classifies the word into the group of the Japanese language type. By not classifying foreign words into the group of the English language type, it can be reproduced in Japanese pronunciation without being reproduced in English pronunciation.
[0039] The probability output unit 33 outputs a probability indicating the likelihood of being classified into any of the groups of multiple language types by the language type classification unit 23. When classified by referring to the relevant information by the language type classification unit 23, the probability is lower compared to the case of not referring to the relevant information. Also, the probability of a word classified into the group of the Japanese language type by referring to the foreign word database is lower compared to a word classified into the group of the English language type. The probability output unit 33 controls the communication unit 105 to transmit probability information including the word, the name of the group of the language type into which the word is classified, and the probability, to the PC 200. The instructor operating the PC 200 can refer to the probability information to correct the group of the language type into which the word is classified.
[0040] The extraction unit 25 extracts one or more words classified into the group of the English language type by the language type classification unit 23. Note that when the group of the language type into which a word is classified is corrected by the instructor, the extraction unit 25 extracts one or more words classified into the corrected group of the English language type.
[0041] The determination unit 27 determines the part of the first audio data corresponding to one or more words classified into the group of the English language type extracted by the extraction unit 25 as the replacement target part. The determination unit 27 outputs a pair of the replacement target part and the word to the replacement unit 15.
[0042] The replacement unit 15 replaces the replacement target part in the first audio data with the second audio data. The second audio data is English audio data. The English audio data corresponds to the audio of a native English speaker uttering words. Since the replacement target part is replaced with the second audio data, the probability output by the probability output unit 33 indicates the probability that the replacement target part of the first audio data corresponding to the word is replaced with the second audio data for the words classified into the group of English language types.
[0043] The replacement unit 15 prepares in advance the feature data of the voice of a native English speaker, and generates the second data by synthesizing the voice of the word using the feature data. Also, without using voice synthesis, for each of a plurality of words, the second audio data corresponding to the audio of a native English speaker uttering the word may be stored in advance in 104. The replacement unit 15 acquires the second audio data by reading out the second audio data corresponding to the word from the HDD 104. The replacement unit 15 outputs the third audio data obtained by replacing the replacement target part of the first audio data with the second audio data to the output unit 17.
[0044] The output unit 17 controls the communication unit 105 to transmit the audio data and the character information to each of the PCs 200-1 to N. The output unit 17 includes an audio output unit 41, a character output unit 43, an audio switching unit 45, and a character switching unit 47. The audio output unit 41 transmits the third audio data to each of the PCs 200-1 to N by controlling the communication unit 105. In each of the PCs 200-1 to N, the third audio data is played back, and the played-back audio is output from the speaker 209. For this reason, the English audio uttered by the instructor can be heard as the audio uttered by a native English speaker. For this reason, the student can hear English with a more accurate pronunciation than the pronunciation of the audio uttered by the instructor.
[0045] Further, when the voice output unit 41 has previously acquired feature data indicating the features of the instructor's voice, the voice output unit 41 may convert the second voice data corresponding to the replacement target portion among the third voice data into the instructor's voice using the feature data. As a result, the second voice data of the replacement target portion is converted into a voice closer to the voice uttered by the instructor. For this reason, the student can listen to the third voice data being played back without a sense of discomfort.
[0046] The voice switching unit 45 switches the voice data transmitted by the voice output unit 41 for each of the PCs 200-1 to N according to an operation received from any one of the PCs 200-1 to N. For example, when the voice switching unit 45 receives an operation to switch to the first voice data from the PC 200-1, it switches so that the first voice data is output to the PC 200-1. When the voice switching unit 45 receives an operation to switch to the first voice data, the voice output unit 41 transmits the first voice data to the PC 200-1 instead of the third voice data. Also, when the voice switching unit 45 receives an operation to switch to the third voice data from the PC 200-1, it switches so that the third voice data is output to the PC 200-1. When the voice switching unit 45 receives an operation to switch to the third voice data, the voice output unit 41 transmits the third voice data to the PC 200-1 instead of the first voice data.
[0047] The character output unit 43 transmits the character information obtained by voice recognition of the first voice data to each of the PCs 200-1 to N by controlling the communication unit 105. At each of the PCs 200-1 to N, the character information is displayed on the display unit 206. For this reason, the character information corresponding to the voice uttered by the instructor can be visually recognized. For this reason, the student can simultaneously confirm the English pronunciation and the spelling of the word. Note that the character output unit 43 may transmit the character information obtained by voice recognition of the third voice data to each of the PCs 200-1 to N by controlling the communication unit 105.
[0048] The character switching unit 47 switches, for each of the PCs 200-1 to N, whether to transmit character information according to an operation received from any of the PCs 200-1 to N. For example, when the character switching unit 47 receives an operation to request character information from the PC 200-1, it switches so that the character information is output by the PC 200-1. When the character output unit 43 receives an operation to request character information by the character switching unit 47, it transmits the character information to the PC 200-1. Also, when the character switching unit 47 receives an operation that does not request character information from the PC 200-1, it switches so that the character information is not output to the PC 200-1. When the character output unit 43 receives an operation that does not output character information by the character switching unit 47, it does not transmit the character information to the PC 200-1.
[0049] The history information generation unit 19 generates history information. The history information includes the replacement target portion of the first voice data and the replaced second voice data. The history information generation unit 19 controls the communication unit 105 to transmit the history information to the PC 200. The instructor who operates the PC 200 can compare the pronunciation of his or her own words with the pronunciation of a native English speaker by playing back the history information, and can correct his or her own pronunciation.
[0050] FIG. 5 is a flowchart showing an example of the flow of the voice conversion process. The voice conversion process is a process executed by the CPU 101 of the server 100 by executing a voice conversion program stored in the ROM 102, HDD 104, or CD-ROM 111. Referring to FIG. 5, the CPU 101 acquires the first voice data (step S01) and advances the process to step S02. The CPU 101 controls the communication unit 105 to acquire the first voice data received by the communication unit 105 from the PC 200. The first voice data corresponds to the voice uttered by the instructor. The PC 200 is operated by the instructor, and the voice uttered by the instructor collected by the microphone 210 provided therein is converted into the first voice data, and the first voice data is transmitted to the server 100.
[0051] In step S02, the first language type is specified, and the process proceeds to step S03. The first language type is a predetermined language type that is the language type of the instructor's native language. Here, the first language type is set to Japanese.
[0052] In step S03, foreign language information is acquired, and the process proceeds to step S04. The foreign language information that defines the foreign words of the first language type is stored in advance in the HDD 104. The foreign language information stored in the HDD 104 is acquired. In step S04, related information is acquired, and the process proceeds to step S05. The related information related to the first audio data acquired in step S01 is acquired. The related information includes information related to the instructor who uttered the audio of the first audio data. The related information includes the materials distributed to the students for the instructor to conduct the guidance, and the characters obtained from the images of the image data obtained by photographing the situation where the instructor is uttering the audio.
[0053] In step S05, the process of performing speech recognition on the first audio data is executed, and the process proceeds to step S05. The first audio data acquired in step S01 is speech-recognized and converted into character information. When performing speech recognition, the related information acquired in step S04 is referred to. There may be a case where a plurality of words are candidates for the audio. In this case, the candidate included in the related information is selected from among the plurality of candidates.
[0054] In step S06, one of the one or more words included in the character information converted from the first audio data is selected as the processing target, and the process proceeds to step S07. In step S07, the word is classified, and the process proceeds to step S08. The word selected as the processing target in step S06 is classified into any one of a plurality of predetermined language type groups. Here, the word is classified into either the group of Japanese as the first language type or the group of English as the second language type. When performing this classification, the foreign word information acquired in step S03 is referred to. If the word to be processed exists in both English and Japanese and the word corresponds to a foreign word, it is classified into the group of the Japanese language type.
[0055] In step S08, it is determined whether there is a word that was not selected as the processing target in step S06. If there is an unselected word, the process returns to step S06; if not, the process proceeds to step S09.
[0056] In step S09, audio feature data is acquired, and the process proceeds to step S10. From the PC200, audio feature data indicating the features of the instructor's voice is acquired.
[0057] In step S10, one of the plurality of words included in the character information converted from the first audio data is selected as the processing target, and the process proceeds to step S11. In step S11, the portion of the first audio data corresponding to the word selected as the processing target is specified, and the process proceeds to step S12. In step S12, it is determined whether the word selected as the processing target is of the first language type. If the word is of the first language type, the process proceeds to step S13; if it is of the second language type, the process proceeds to step S15.
[0058] In step S13, the first audio data acquired in step S01 is output as it is, and the process proceeds to step S14. The first audio data is transmitted to each of the PC200-1 to N.
[0059] In step S14, it is determined whether there is a word that was not selected as the processing target in step S10. If there is an unselected word, the process returns to step S10; otherwise, the process ends.
[0060] In step S15, the part of the first audio data specified in step S11 is determined as the replacement target part, and the process proceeds to step S16. In step S16, the second audio data is acquired, and the process proceeds to step S17. The second audio data is audio data corresponding to the voice obtained by pronouncing the word selected as the processing target in step S10 in English, which is the second language type. For a plurality of words, the audio data of the voice uttered by a person whose native language is English is pre-stored in the HDD 104, and the audio data corresponding to the word is read out as the second audio data. Also, instead of the second audio data read from the HDD 104 itself, the voice feature data of the instructor acquired in step S09 is used to change the sound quality of the second audio data. Thereby, the second audio data can be made similar to the sound quality of the voice uttered by the instructor.
[0061] In step S17, the replacement target part of the first audio data is replaced with the second audio data and output, and the process proceeds to step S18. Instead of the first audio data acquired in step S01, the second audio data is output. Specifically, the second audio data is transmitted to each of the PCs 200-1 to N.
[0062] <Modification Example> In the foreign language learning system 1 in the modification example, the server 100 determines whether the audio is in the first language type or the second language type using the learned learning model. The CPU 101 provided in the server 100 in the modification example changes the replacement part determination unit 13 of the functions shown in FIG. 4.
[0063] FIG. 6 is a diagram showing an example of the functions of the CPU included in the server in the modified example. Referring to FIG. 6, the differences from the functions shown in FIG. 4 are that a learning model generation unit 51 is added and the replacement part determination unit 13 is changed to a replacement part determination unit 13A. Since the other functions are the same as those shown in FIG. 4, the description will not be repeated here. The learning model generation unit 51 generates a learned learning model by performing machine learning on the learning model.
[0064] The learning model generation unit 51 performs machine learning on the learning model using learning data with the explanatory variable being voice and the target variable being language type. The learning model generation unit 51 outputs the learned learning model to the replacement part determination unit 13A. The learning data includes first learning data and second learning data. The first learning data includes, as the explanatory variable, the voice of a person whose native language is Japanese uttering English words, and includes English as the target variable. The second learning data includes, as the explanatory variable, the voice of a person whose native language is Japanese uttering Japanese words, and includes Japanese as the target variable. The first learning data and the second learning data include the voices uttered by the same person. Here, the case where the first learning data and the second learning data include the voices uttered by a person whose native language is Japanese will be described as an example.
[0065] The replacement part determination unit 13A includes a language type prediction unit 23A, a probability output unit 33A, an extraction unit 25, and a determination unit 27. The language type prediction unit 23A receives the learned learning model from the learning model generation unit 51 and the first voice data from the voice acquisition unit 11. The language type prediction unit 23A gives the first voice data to the learning model and predicts the language type of the voice of the first voice data. The language type prediction unit 23A divides the voice data into a plurality of word data in word units and gives each of the plurality of word data to the learning model. The language type prediction unit 23A outputs a pair of word data and language type to the extraction unit 25 for each of the plurality of word data.
[0066] The probability output unit 33A outputs a probability indicating the reliability of the language type predicted by the learning model from the word data. The probability output unit 33A controls the communication unit 105 to transmit probability information including a word, the name of the language type to which the word is classified, and the probability to the PC 200. The instructor operating the PC 200 can refer to the probability information and modify the group of language types to which the word is classified, etc.
[0067] The extraction unit 25 extracts one or more words classified into the English language type by the language type prediction unit 23A. When the group of language types to which the word is classified is modified by the instructor, the extraction unit 25 extracts one or more words classified into the modified English language type group.
[0068] FIG. 7 is a flowchart showing an example of the flow of the voice conversion process in the modification example. Referring to FIG. 7, the differences from the voice conversion process shown in FIG. 5 are that step S07 is changed to step S07A, and steps S03 to S06, step S08, and step S11 are deleted. Since the other processes are the same as those shown in FIG. 5, the description will not be repeated here.
[0069] In step S01, the first voice data is acquired. In the next step S02, the first language type is determined. Here, Japanese is determined as the first language type. In the next step S07A, the first voice data is given to the learning model, the language type corresponding to the first voice data is predicted by the learning model, and the process proceeds to step S09. In step S09, the voice feature data of the person who uttered the voice corresponding to the first voice data, here the instructor, is acquired, and the process proceeds to step S12. In step S12, the process branches based on the language type predicted in step S07A. If the language type of the first voice data is the first language type, the process proceeds to step S13, and if it is the second language type, the process proceeds to step S15.
[0070] <Other Embodiments> In the above-described embodiment, the first language type is Japanese and the second language type is English. The first language type is not limited to Japanese, and the second language type is not limited to English. The first language type and the second language type only need to be different language types. For example, the first language type may be English and the second language type may be Japanese.
[0071] As described above, the server 100 in the foreign language learning system 1 in the present embodiment acquires first audio data corresponding to the voice uttered by an instructor whose first language type (Japanese) is the native language. Then, the server 100 replaces the part to be replaced corresponding to the second language type (English) of the first audio data with second audio data corresponding to the voice uttered by a person whose native language is English. Therefore, when the instructor pronounces an English word, the word is output with the same pronunciation as that uttered by a person whose native language is English. Therefore, the English voice uttered by the instructor can be converted into a voice with a native pronunciation.
[0072] Further, the server 100 classifies a plurality of words included in the character information obtained by voice-recognizing the first audio data into either Japanese or English, and determines second audio data corresponding to each of one or more words classified as English. Therefore, it is possible to easily extract the part to be replaced corresponding to the English word from the first audio data corresponding to the voice uttered by an instructor whose native language is Japanese.
[0073] Further, the server 100 acquires related information related to the first audio data, performs voice recognition based on the related information, and determines character information. Therefore, the accuracy of voice recognition can be improved.
[0074] Further, for an English word, a Japanese foreign word is pre-registered, and the server 100 does not determine second audio data corresponding to the foreign word. The server 100 outputs the part corresponding to the foreign word in the first audio data as a Japanese voice. Therefore, a Japanese word can be output as the voice uttered by the instructor.
[0075] In addition, the server 100 in the modification example uses a learning model obtained by performing machine learning on learning data with explanatory variables as voices and objective variables as language types, predicts the language type of the first voice data, and replaces the replacement target part predicted to be English in the first voice data with the second voice data. Therefore, the English part can be easily determined from the first voice data.
[0076] In addition, the server 100 converts the second voice data using voice feature data indicating the features of the instructor's voice. Therefore, the voice corresponding to the second voice data can be output as the instructor's voice.
[0077] In addition, the server 100 generates history information including the replacement target part of the first voice data and the replaced second voice data. Therefore, the instructor can compare the voice he / she uttered with the voice corresponding to the second voice data.
[0078] In addition, the server 100 outputs the probability that the replacement target part is replaced with the second voice data. Therefore, the instructor can easily determine whether the replacement with the second voice data is correct.
[0079] In addition, the server 100 outputs replacement post-character information obtained by performing voice recognition on either the first voice data or the third voice data obtained by replacing the replacement target part of the first voice data with the second voice data. Therefore, character information is displayed on each of the PCs 200-1 to N. Therefore, a plurality of students operating each of the PCs 200-1 to N can visually recognize the character information while listening to the voice played back from the third voice data, and can view it while comparing the voice and the character information.
[0080] In addition, the server 100 outputs either the first voice data or the third voice data. Therefore, a plurality of students operating each of the PCs 200-1 to N can select and play back either the first voice data or the third voice data.
[0081] <Summary of the Embodiment> (Item 1) A voice acquisition means for acquiring first voice data corresponding to the voice uttered by the user, and a replacement means for replacing a part to be replaced corresponding to a predetermined language type in the first voice data with second voice data different from the first voice data. A voice conversion device comprising the same.
[0082] According to this aspect, the part to be replaced corresponding to the predetermined language type in the first voice data is replaced with the second voice data. For this reason, when the user utters a word corresponding to a predetermined language type, the word is output in another voice. Therefore, it is possible to provide a voice conversion device capable of converting the voice uttered by the user into a voice of pronunciation corresponding to the language type.
[0083] (Item 2) The voice conversion device according to Item 1, wherein the second voice data is voice data of the predetermined language type.
[0084] According to this aspect, since the second voice data is voice data of a predetermined language type, if the second voice data corresponds to the voice uttered by a person whose native language is the predetermined language type, when the user utters a word corresponding to the predetermined language type, the word is output with the same pronunciation as the pronunciation uttered by a person whose native language is the predetermined language type.
[0085] (Item 3) Voice recognition means for performing voice recognition on the first voice data and outputting character information, Classification means for classifying a plurality of words included in the character information into either the first language type or the second language type, Extraction means for extracting one or more of the words classified into the second language type, The voice conversion device according to Item 1 or 2, further comprising determination means for determining the second voice data corresponding to one or more of the extracted words.
[0086] According to this aspect, a plurality of words included in the character information obtained by voice recognition of the first voice data are classified into either the first language type or the second language type, and second voice data corresponding to one or more words classified into the second language type is determined. Therefore, it is possible to extract a portion corresponding to a word corresponding to the second language type from the first voice data corresponding to the voice uttered by the first user whose native language is the first language type.
[0087] (Item 4) The voice conversion device further includes related information acquisition means for acquiring related information related to the first voice data. The voice recognition means determines the word based on the related information, and the voice conversion device according to item 3.
[0088] According to this aspect, a word is determined based on related information related to the first voice data. Therefore, the accuracy of voice recognition is improved.
[0089] (Item 5) Foreign words are registered in advance corresponding to words of a predetermined language type. The determination means does not determine the second voice data corresponding to the word corresponding to the foreign word among the plurality of words included in the character information, and the voice conversion device according to item 3 or 4.
[0090] According to this aspect, foreign words are registered in advance corresponding to words of a predetermined language type, and second voice data corresponding to a word corresponding to a foreign word among a plurality of words included in the character information is not determined. Therefore, the portion corresponding to the foreign word in the first voice data is output as it is without being replaced by the second voice data. Therefore, the foreign word can be output as the voice uttered by the first user.
[0091] (Item 6) The replacement means uses a learning model obtained by machine learning of learning data with an explanatory variable as voice and a target variable as a language type to determine, as the replacement target portion, a portion of the first voice data predicted to be the predetermined language type by the learning model. The learning data of the voice conversion device according to claim 1 includes first learning data containing voices of the same person uttering words of the predetermined language type, and second learning data containing voices of the same person uttering words of a language type different from the predetermined language type.
[0092] According to this aspect, a learning model obtained by performing machine learning on learning data including voices of the same person uttering words of a predetermined language type and voices of the same person uttering words of a language type different from the predetermined language type is used to predict the language type of the first voice data. Therefore, it is possible to easily determine the part of the first voice data that is of the predetermined language type.
[0093] (Clause 7) The voice conversion device according to any one of claims 1 to 6, further comprising voice synthesis means for converting the second voice data using voice feature data indicating the characteristics of the user's voice.
[0094] According to this aspect, since the second voice data is converted using voice feature data indicating the characteristics of the user's voice, the voice corresponding to the second voice data can be output as the user's voice.
[0095] (Clause 8) The voice conversion device according to any one of claims 1 to 7, further comprising output means for outputting a history of a pair of the replacement target and the second voice data obtained by replacing the replacement target part of the first voice data.
[0096] According to this aspect, a history of a pair of the replacement target part of the first voice data and the replaced second voice data is output. Therefore, the user can compare the voice uttered by the user with the voice corresponding to the second voice data.
[0097] (Clause 9) The voice conversion device according to any one of claims 1 to 8, further comprising probability output means for outputting the probability that the replacement target part is replaced with the second voice data.
[0098] According to this aspect, since the probability of being replaced with the second audio data is output, the user can easily determine whether the replacement with the second audio data is correct.
[0099] (Item 10) A character output means for outputting replacement character information obtained by performing speech recognition on either the first audio data or third audio data obtained by replacing the replacement target portion of the first audio data with the second audio data, the speech conversion device according to any one of Items 1 to 9 further comprising the same.
[0100] According to this aspect, since the replacement character information obtained by performing speech recognition on either the first audio data or the third audio data is output, it is possible to view while comparing the voice and character information of the third audio data.
[0101] (Item 11) A voice output means for outputting either the first audio data or third audio data obtained by replacing the replacement target portion of the first audio data with the second audio data, the speech conversion device according to any one of Items 1 to 10 further comprising the same.
[0102] Either the first audio data or the third audio data can be selectively output.
[0103] (Item 12) A voice acquisition step of acquiring first audio data corresponding to the voice uttered by the user, A replacement step of replacing a replacement target portion corresponding to a predetermined language type in the first audio data with second audio data different from the first audio data, a voice conversion method for causing a voice conversion device to execute the same.
[0104] According to this aspect, it is possible to provide a voice conversion method capable of converting the voice uttered by the user into a voice of pronunciation corresponding to the language type.
[0105] (Item 13) A voice acquisition step of acquiring first audio data corresponding to the voice uttered by the user, A voice conversion program that causes a computer to execute a replacement step of replacing a replaceable part corresponding to a predetermined language type in the first voice data with second voice data different from the first voice data.
[0106] According to this aspect, it is possible to provide a voice conversion program capable of converting the voice uttered by the user into a voice with pronunciation corresponding to the language type.
[0107] The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present invention is shown not by the above description but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims are included.
Explanation of Reference Numerals
[0108] 1 Foreign language learning system, 5 Internet, 11 Voice acquisition unit, 13, 13A Replacement part determination unit, 15 Replacement unit, 17 Output unit, 19 History information generation unit, 21 Voice recognition unit, 23 Language type classification unit, 23A Language type prediction unit, 25 Extraction unit, 27 Determination unit, 29 Related information acquisition unit, 31 Foreign word acquisition unit, 33, 33A Probability output unit, 41 Voice output unit, 43 Character output unit, 45 Voice switching unit, 47 Character switching unit, 51 Learning model generation unit, 100 Server, 101 CPU, 102 ROM, 103 RAM, 104 HDD, 105 Communication unit, 106 Display unit, 107 Operation unit, 110 External storage device, 111 CD-ROM, 201 CPU, 202 ROM, 203 RAM, 204 HDD, 205 Communication unit, 206 Display unit, 207 Operation unit, 208 Camera, 209 Speaker, 210 Microphone, 211 External storage device, 211A CD-ROM.
Claims
1. An audio acquisition means for acquiring first audio data corresponding to the audio uttered by the user, A voice conversion device comprising: a replacement means for replacing a replacement target portion corresponding to a predetermined language type in the first audio data with second audio data different from the first audio data.
2. The voice conversion device according to claim 1, wherein the second audio data is audio data of the predetermined language type.
3. An audio recognition means for performing audio recognition on the first audio data and outputting character information, A classification means for classifying a plurality of words included in the character information by language type, An extraction means for extracting one or more of the words classified into the predetermined language type, The voice conversion device according to claim 1, further comprising: a determination means for determining the second audio data corresponding to one or more of the extracted words.
4. Further comprising a related information acquisition means for acquiring related information related to the first audio data, The voice conversion device according to claim 3, wherein the audio recognition means determines the word based on the related information.
5. Foreign words are registered in advance corresponding to words of a predetermined language type, The voice conversion device according to claim 3, wherein the determination means does not determine the second audio data corresponding to the word corresponding to the foreign word among the plurality of words included in the character information.
6. The replacement means uses a learning model obtained by machine learning of learning data with explanatory variables as audio and target variables as language types, and determines, as the replacement target portion, a portion of the first audio data predicted to be the predetermined language type by the learning model, The voice conversion device according to claim 1, wherein the learning data includes first learning data including audio uttered by the same person with words of the predetermined language type, and second learning data including audio uttered by the same person with words of a language type different from the predetermined language type.
7. The voice conversion device according to any one of claims 1 to 6, further comprising a voice synthesis means for converting the second audio data using personal characteristic data indicating the characteristics of the user's voice.
8. The voice conversion device according to any one of claims 1 to 6, further comprising an output means for outputting a history of a set of the replacement target portion of the first audio data and the second audio data obtained by replacing the replacement target portion.
9. The voice conversion device according to any one of claims 1 to 6, further comprising probability output means for outputting a probability that the replacement target part is replaced with the second voice data.
10. The voice conversion device according to any one of claims 1 to 6, further comprising character output means for outputting replacement character information obtained by performing voice recognition on either one of the first voice data and third voice data obtained by replacing the replacement target part of the first voice data with the second voice data.
11. The voice conversion device according to any one of claims 1 to 6, further comprising voice output means for outputting either one of the first voice data and third voice data obtained by replacing the replacement target part of the first voice data with the second voice data.
12. A voice acquisition step of acquiring first voice data corresponding to the voice uttered by the user, A voice conversion method for causing a voice conversion device to execute a replacement step of replacing a replacement target part corresponding to a predetermined language type in the first voice data with second voice data different from the first voice data.
13. A voice acquisition step of acquiring first voice data corresponding to the voice uttered by the user, A voice conversion program for causing a computer to execute a replacement step of replacing a replacement target part corresponding to a predetermined language type in the first voice data with second voice data different from the first voice data.
Citation Information
Patent Citations
Linguistic education device with voice recognition device
JP1991226785A