Language learning system and electronic device using the language learning system
The language learning system translates and generates voice in a second language for real-time pronunciation assessment, addressing the limitations of conventional methods by enabling anytime and anywhere practice.
Patent Information
- Application Number
- JP2025001420U
- Authority / Receiving Office
- JP · JP
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2035-05-07
AI Technical Summary
Conventional language learning methods are limited by teacher resources and fixed classroom time, preventing real-time and anywhere language training.
A language learning system that translates voice content from a first language to text in a second language, generates voice in the second language based on the text, and compares the learner's voice with the generated voice to assess pronunciation accuracy.
Enables real-time and anywhere language learning by providing immediate feedback on pronunciation accuracy, allowing learners to practice anytime and anywhere.
Smart Images

Figure 0003251957000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to a learning device, and more particularly, to a language learning system and an electronic device using the language learning system.
Background Art
[0002] With the development of the economy and cultural exchanges, the requirements for an individual's language ability are increasing.
[0003] The conventional language learning method is that students and teachers directly face each other and learn in an interactive communication method. The teacher directly experiences the student's language pronunciation, and corrects it when the student's pronunciation is incorrect, and can ensure that the student's pronunciation is accurate. Such a learning method can achieve the purpose of real-time correction. However, due to limited teacher resources, not all students can be ensured to have the opportunity to practice individually one-on-one. And due to the limited classroom teaching time, language learning cannot be carried out anytime and anywhere.
[0004] Therefore, how to efficiently conduct language training has become an issue to be solved.
Summary of the Invention
Problems to be Solved by the Invention
[0005] One embodiment in the present disclosure discloses a language learning system for accessing at least one command, comprising a storage element for storing at least one command, and a processing element coupled to the storage element. The processing element includes a translation module for translating voice content of a first language into text content of a second language different from the first language, a voice synthesis module for generating a voice of the second language based on the text content of the second language, a voice comparison module for comparing a learned voice of the second language generated by a learner learning the voice of the second language with the voice of the second language, and generating a corresponding score based on a similarity value between the learned voice of the second language and the voice of the second language.
Means for Solving the Problem
[0006] In some embodiments, the voice comparison module generates the similarity value based on an edit distance between the learned voice of the second language and the voice of the second language.
[0007] In some embodiments, the voice comparison module executes a Levenshtein mathematical algorithm to calculate the edit distance between the learned voice of the second language and the voice of the second language.
[0008] In some embodiments, the voice comparison module further sets a base weight value and generates the corresponding score based on the similarity value and the base weight value.
[0009] In some embodiments, the voice synthesis module is further used to set the intonation and speaking speed of the second language and generate the voice of the second language based on the intonation and the speaking speed.
[0010] Another embodiment of the present disclosure includes a storage element for storing at least one command, a processing element coupled to the storage element, a voice acquisition element coupled to the processing element for acquiring a voice in a first language uttered by a learner, a display element coupled to the processing element, and a voice element coupled to the processing element. The processing element further includes a translation module for translating the content of the voice in the first language into text content in a second language different from the first language and displaying it on the display element, a voice synthesis module for generating a voice in the second language based on the text content in the second language and uttering the voice in the second language by the voice element, and a voice comparison module for comparing the learned voice in the second language generated by the learner learning the voice in the second language acquired by the voice acquisition element with the voice in the second language and generating a corresponding score based on a similarity value between the learned voice in the second language and the voice in the second language and displaying it on the display element. An electronic device for accessing the at least one command is disclosed.
[0011] In some embodiments, the voice comparison module generates the similarity value based on an edit distance between the learned voice in the second language and the voice in the second language.
[0012] In some embodiments, the voice comparison module executes the Levenshtein mathematical algorithm to calculate the edit distance between the learned voice in the second language and the voice in the second language.
[0013] In some embodiments, the electronic device further includes an input interface coupled to the processing element for receiving commands, and the voice synthesis module generates the voice in the second language in response to the commands.
[0014] In some embodiments, the voice synthesis module is further used to set the intonation and speaking speed of the second language and generate the voice in the second language based on the intonation and the speaking speed.
Advantages of the Invention
[0015] The language learning system and corresponding electronic device of the present application can generate text content in a second language based on voice content in a first language input by a learner, and generate a voice in the second language of this text content based on the text content in the second language. Therefore, the learner can perform real-time learning of the second language based on this voice in the second language. In addition, in the present application, the learner obtains in real time the learning voice of the second language that the learner learns based on this text content, performs voice comparison with the generated voice in the second language, determines the pronunciation accuracy score of the learner's learning voice in the second language, and thereby realizes the purpose of performing language learning anytime and anywhere.
Brief Description of the Drawings
[0016] To make the above and other objects, features, advantages and embodiments of the present disclosure clearer and easier to understand, the drawings are described as follows.
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0017] Examples will be given below and described in detail with reference to the drawings. However, the provided examples are not intended to limit the scope included in this disclosure. The description of the structural operations is not intended to limit their execution order. Structures in which elements are recombined and devices having equivalent effects produced are within the scope included in this disclosure. Also, the drawings are for illustrative purposes only and are not drawn according to the original dimensions. For ease of understanding, the same or similar elements in the following description will be denoted by the same reference numerals.
[0018] The terms used throughout the specification and in the claims for utility model registration, unless otherwise specified, are generally used in this field and have their ordinary meanings in the context of the content disclosed herein.
[0019] In addition, the terms "comprising," "including," "having," "containing," etc., used in this specification are all open-ended terms, meaning "including but not limited to." Also, the "and / or" used in this specification includes any one or more of the relatedly listed items and all combinations thereof.
[0020] In this specification, when an element is described as being "connected", "coupled", or "electrically connected" to another element, this element may be a direct connection, direct coupling, or direct electrical connection, or there may be additional elements between the two elements, and this element is indirectly connected, indirectly coupled, or indirectly electrically connected to this other element. However, when an element is described as being "directly connected", "directly coupled", or "directly electrically connected" to another element, it should be understood that there are no additional elements. Note that when an element is described as being "linked" or "communicatively connected" to another element, this element may be indirectly wired and / or wirelessly communicated to another element via other elements, or an element may be physically connected to another element without passing through other elements. Note that in this specification, terms such as "first", "second",... are used to describe different elements, but this term is only used to distinguish elements or operations described by the same technical term.
[0021] In the conventional language learning method, since the student and the teacher learn in a face-to-face and interactive communication method, such a learning method can achieve the purpose of real-time correction. However, it is limited by the limited resources of the teacher and the fixed class time in the classroom, and the purpose of conducting language learning anytime and anywhere cannot be achieved. Therefore, this application provides a language learning system and a corresponding electronic device, which can generate text content in a second language based on the voice content in the first language input by the learner, and generate a reference voice for this text content based on the text content in the second language. Therefore, the learner can perform real-time second language learning based on this reference voice. The language learning system of this application can obtain the learning voice in the second language generated by the learner based on this text content in real time, compare it with this reference voice, and judge the pronunciation accuracy score of the learning voice in the second language by the learner, thereby achieving the purpose of conducting language learning anytime and anywhere.
[0022] FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. FIG. 2 is a schematic diagram of a language learning system according to an embodiment of the present disclosure. The language learning system 100 includes a translation module 110, a voice synthesis module 120, and a voice comparison module 130. In some embodiments, the language learning system 100 may be disposed in an electronic device, which may be an electronic device such as a smartphone, a tablet computer, etc. The electronic device 10 includes a processing element 20, a storage element 30, a voice acquisition element 40, a display element 50, a voice element 60, and an input interface 70. The above elements may communicate via, for example, a bus, but are not limited thereto. The processing element 20 may be realized by a central processing unit or a computing unit. The storage element 30 may be realized by a random access memory (RAM), a read only memory (ROM), a flash memory, a hard disk, or other storage devices that can be used to store data. The voice acquisition element 40 may be realized by a microphone, a voice sensor, a force sensor. The voice element 60 may be realized by a horn, a speaker. The input interface 70 may be, for example, a keyboard, a mouse, a touch screen, etc. However, the present application is not limited to the above. In some embodiments, the storage element 30 further stores a plurality of pen commands, and the processing element 20 executes different ones of these commands to implement the translation module 110, the voice synthesis module 120, and the voice comparison module 130 of the language learning system 100 to provide language training to the learner, and at the same time, can compare and score the voice intonation of the learner's pronunciation. Please refer to FIGS. 1 and 2 simultaneously.
[0023] In some embodiments, the voice acquisition element 40 of the electronic device 10 is used to acquire the voice content of the learner's first language. The processing element 20 of the electronic device 10 is coupled to the voice acquisition element 40 and is used to receive the voice content of the first language acquired by the voice acquisition element 40. In some embodiments, the processing element 20 executes different ones of these commands in the storage element 30 to implement the translation module 110 of the language learning system 100, translate the voice content of the first language into text content of the second language, and display it on the display element 50. In some embodiments, if the learner is a Japanese person who wants to learn Chinese through Japanese, the first language to be used is Japanese and the second language is Chinese. The learner can utter the voice content in Japanese, a language they are familiar with. After the voice acquisition element 40 acquires this voice content, it transmits it to the processing element 20. The processing element 20 implements the translation module 110 of the language learning system 100, translates the Japanese voice content into Chinese character content, and then displays it on the display element 50.
[0024] Next, the learner controls the electronic device 10 to read the text content in the second language displayed on the display element 50 by inputting a reading command via the input interface 70 of the electronic device 10, and the voice element 60 can emit the voice in the second language corresponding to this text content. Therefore, the learner can directly listen to and learn the voice in the second language corresponding to this text content. In some embodiments, since the voice in the second language emitted by the voice element 60 is used as a voice criterion for subsequently judging the score, the processing element 20 stores the voice in the second language emitted by the voice element 60 in the storage element 30 and uses it as a subsequent criterion. In some embodiments, the processing element 20 of the electronic device 10 executes different ones of these commands in the storage element 30 to implement the voice synthesis module 120 of the language learning system 100, reads the text content based on the text content in the second language displayed on the display element 50, and emits this second voice by the voice element 60. In some embodiments, the learner learns Chinese through Japanese, the first language used is Japanese, and the second language is Chinese. Therefore, the voice element 60 emits the Chinese voice that the learner wants to learn for the Chinese text content displayed on the display element 50, and the Chinese voice emitted by the voice element 60 is stored in the storage element 30.
[0025] FIG. 3 is a schematic flowchart of a voice synthesis module according to an embodiment of the present disclosure. In some embodiments, the processing element 20 of the electronic device 10 executes different ones of these commands in the storage element 30 to implement the voice synthesis module 120 to read the text content aloud. First, in step 121, it monitors a reading command input by the learner. Then, in step 122, it determines whether there is text content that needs to be read. In one embodiment, after monitoring the input of the reading command from the learner, the processing element 20 determines whether there is text content in a second language that needs to be read on the display element 50 of the electronic device 10. If not, step 121 is re-executed to continue monitoring the reading command input by the learner. If so, step 123 is executed to obtain the correspondingly set language, speech rate, and intonation based on the input text content. In some embodiments, the processing element 20 obtains the correspondingly set language, speech rate, and intonation based on the text content displayed on the display element 50. Then, in step 124, it reads the text content aloud based on the obtained language, speech rate, and intonation and plays it back by the voice element 60. In some embodiments, the read voice played back by the voice element 60 is stored in the storage element 30 as a subsequent voice similarity comparison criterion.
[0026] In some embodiments, when a learner learns Chinese through Japanese, the first language to be used is Japanese and the second language is Chinese. Therefore, when the processing element 20 of the electronic device 10 implements the voice synthesis module 120 to read the text content, in step 121, it monitors the reading command input by the learner. After monitoring the input of the reading command from the learner, in step 122, it determines whether there is Chinese text content that needs to be read on the display element 50 of the electronic device 10. If this Chinese text content exists on the display element 50, in step 123, the processing element 20 obtains the language setting of Chinese and the reading speech rate and intonation set for the Chinese language. In step 124, based on this, it reads the Chinese text content on the display element 50 and plays it back through the voice element 60. In some embodiments, the reading voice played back by the voice element 60 is stored in the storage element 30 as a subsequent voice similarity comparison criterion.
[0027] After the learner listens to the voice of the second language corresponding to the text content through the voice element 60 of the electronic device 10, the learner can learn the second language based on the speech rate and intonation emitted by the voice element 60. In some embodiments, the voice acquisition element 40 of the electronic device 10 acquires and stores the learning voice of the second language uttered by the learner, for example, stores it in the storage element 30. The processing element 20 executes different ones of these commands in the storage element 30 to implement the voice comparison module 130 of the language learning system 100, compares the learning voice of the second language uttered by the learner in the storage element 30 with the voice of the second language emitted by the voice element 60, analyzes the similarity of both voices, gives a corresponding score, and can display it on the score display element 50. In some embodiments, in addition to displaying the score, a language level corresponding to this score may be displayed for the language learner to refer to. Therefore, the learner can know whether his / her second language ability and voice pronunciation are accurate and the degree to which they should be improved based on this score or level. In some embodiments, the processing element 20 implements the voice comparison module 130 to determine the similarity of both based on the edit distance between the learning voice of the second language uttered by the learner and the voice of the second language emitted by the voice element 60, and gives a corresponding similarity value. In some embodiments, the processing element 20 implements the voice comparison module 130 to execute the Levenshtein mathematical algorithm to calculate the edit distance between the learning voice of the second language uttered by the learner and the voice of the second language emitted by the voice element 60, thereby determining the similarity of both and giving a corresponding similarity value. However, this application is not limited to the above-mentioned Levenshtein mathematical algorithm. In other embodiments, other mathematical algorithms may be used to calculate the similarity value. In some embodiments, for the learner to learn Chinese through Japanese, the first language used is Japanese and the second language is Chinese. Therefore, the voice element 60 emits a Chinese voice based on the acquired Chinese language settings, speech rate, and intonation, and the learner can synchronously emit a Chinese voice based on the speech rate and intonation emitted by the voice element 60. The voice acquisition element 40 acquires the Chinese voice uttered by the learner and stores it in the storage element 30.The processing element 20 implements the voice comparison module 130 to compare the Chinese voice uttered by the learner in the storage element 30 with the Chinese voice uttered by the voice element 60, analyze the similarity of both Chinese voices, give corresponding scores, and display them on the display element 50.
[0028] Figure 4 is a schematic flowchart of a voice comparison module according to an embodiment of the present disclosure. In some embodiments, the processing element 20 of the electronic device 10 executes different ones of these commands in the storage element 30 to implement the voice comparison module 130. When comparing the similarity between the learning voice of the second language uttered by the learner in the storage element 30 and the voice of the second language uttered by the voice element 60, first, in step 131, it is determined whether there is a voice of the second language uttered by the voice element. In some embodiments, since the voice of the second language uttered by the voice element 60 is a reference voice for determination, first, it is confirmed whether there is a voice of the second language uttered by the voice element 60 in the storage element 30. If not, the comparison flow is terminated. Conversely, if there is a voice of the second language uttered by the voice element 60 in the storage element 30, step 132 is executed to set the base weight value. In some embodiments, the base weight value is used to determine the similarity score and may be set to, for example, 95. Then, step 133 of calculating the similarity value is executed. In some embodiments, the voice comparison module 130 determines the similarity between the learning voice of the second language uttered by the learner and the voice of the second language uttered by the voice element 60 based on the edit distance between them and gives the corresponding similarity value. In some embodiments, the processing element 20 implements the voice comparison module 130 to execute the Levenshtein mathematical algorithm to calculate the edit distance between the learning voice of the second language uttered by the learner and the voice of the second language uttered by the voice element 60, thereby determining the similarity values of both. In step 134, the learner's score is calculated. In some embodiments, the learner's score is calculated based on the similarity value and the weight value. In some embodiments, the learner's score is equal to the multiplication result of the similarity value and the weight value. Finally, in step 135, the score is displayed on the display element 50.
[0029] As described above, the present application provides a language learning system and a corresponding electronic device, which can generate text content in a second language based on voice content in a first language input by a learner, and generate a voice in the second language of this text content based on the text content in the second language. Therefore, the learner can perform real-time learning of the second language of this text content based on this voice in the second language. In addition, in the present application, the learner can obtain in real time the learning voice in the second language that the learner learns based on this text content, perform voice comparison with the generated voice in the second language, determine the pronunciation accuracy score of the learning voice in the second language of the learner, and thereby realize the purpose of performing language learning at any time and anywhere.
[0030] Although the present disclosure has been disclosed in the above embodiments, the above-described embodiments are not intended to limit the present invention. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present disclosure. The protection scope of the present disclosure should be based on what is defined by the scope of claims for utility model registration attached later.
Description of Reference Numerals
[0031] In order to make the above and other objects, features, advantages and embodiments of the present disclosure clearer and easier to understand, the description of the reference numerals attached is as follows. 10: Electronic device 20: Processing element 30: Memory element 40: Voice acquisition element 50: Display element 60: Audio element 70: Input interface 100: Language learning system 110: Translation module 120: Voice synthesis module 130: Voice comparison module 121, 122, 123, 124, 131, 132, 133, 134, 135: Steps
Claims
1. A storage element for storing at least one command, A processing element coupled to the storage element, Comprising, The processing element, A translation module for translating voice content in a first language into text content in a second language different from the first language, A voice synthesis module for generating a voice in a second language based on the text content in the second language, A voice comparison module that compares a learned voice in the second language generated by a learner with the voice in the second language and generates a corresponding score based on a similarity value between the learned voice in the second language and the voice in the second language, A language learning system for accessing the at least one command so as to execute.
2. The language learning system according to claim 1, wherein the voice comparison module generates the similarity value based on an edit distance between the learned voice in the second language and the voice in the second language.
3. The language learning system according to claim 2, wherein the voice comparison module executes a Levenshtein mathematical algorithm to calculate the edit distance between the learned voice in the second language and the voice in the second language.
4. The language learning system according to claim 2, wherein the voice comparison module further sets a basic weight value and generates the corresponding score based on the similarity value and the basic weight value.
5. The language learning system according to claim 1, wherein the voice synthesis module further sets the intonation and speaking speed of the second language and is used to generate the voice in the second language based on the intonation and the speaking speed.
6. A storage element for storing at least one command, A processing element coupled to the storage element, A voice acquisition element coupled to the processing element for acquiring a voice in a first language uttered by a learner, A display element coupled to the processing element, An audio element coupled to the processing element, Comprising, The processing element further, A translation module for translating the content of the voice in the first language into text content in a second language different from the first language and displaying it on the display element, A voice synthesis module for generating a voice in a second language based on the text content in the second language and uttering the voice in the second language by the audio element, An electronic device for accessing at least one command to execute a voice comparison module that compares a second language learning voice generated by the learner learning the voice of the second language obtained by the voice acquisition element with the voice of the second language, generates a corresponding score based on a similarity value between the second language learning voice and the voice of the second language, and displays the score on the display element. **Claim 7** The electronic device according to claim 6, wherein the voice comparison module generates the similarity value based on an edit distance between the second language learning voice and the voice of the second language. **Claim 8** The electronic device according to claim 7, wherein the voice comparison module executes a Levenshtein mathematical algorithm to calculate the edit distance between the second language learning voice and the voice of the second language. **Claim 9** The electronic device according to claim 6, further comprising an input interface coupled to the processing element for receiving commands, wherein the voice synthesis module generates the voice of the second language in response to the commands. **Claim 10** The electronic device according to claim 6, wherein the voice synthesis module further sets an intonation and a speaking speed of the second language, and generates the voice of the second language based on the intonation and the speaking speed.