system
A system using voice input, recognition, translation, and synthesis technologies provides real-time foreign language learning in daily life, addressing the limitations of conventional methods by enabling efficient and sustainable language acquisition with emotion recognition and personalized lessons.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Existing systems fail to provide an effective and affordable environment for children to learn foreign languages naturally in their daily lives, and conventional learning methods are costly and limited to specific times and places.
A system utilizing voice input, recognition, translation, and synthesis technologies to convert ambient conversations into a foreign language in real-time, played back through wearable devices, with emotion recognition and personalized learning features.
Enables children to learn foreign languages efficiently and naturally in daily life, with real-time translation and personalized lessons, promoting sustainable language acquisition.
Smart Images

Figure 2026047964000001_ABST
Abstract
Description
Technical Field
[0005] ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In order to enable children to acquire a foreign language at an early age, it is necessary to send them to a dedicated English conversation school or a national school, etc., but this is costly and can only be realized by families with sufficient funds. Furthermore, it is difficult to provide an environment where a foreign language can be learned naturally in daily life. Therefore, there is a need for an effective system that can be easily used by more families and allows learning while actually using a foreign language in daily life.
Means for Solving the Problems
[0005] The present invention is a system including voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, and translated voice playback means. Specifically, the system acquires surrounding Japanese conversations through a device worn by the user and transmits them to a server. The server converts the Japanese conversations into text using voice recognition technology, and then translates that text into English using a translation engine. Subsequently, the text translated into English is converted into voice data using voice synthesis technology, transmitted to the device, and played back to the user. This makes it possible to provide a low-cost environment in the home where children can be exposed to foreign languages on a daily basis.
[0006] "Voice input means" refers to a microphone or other voice capture device used to acquire ambient sounds.
[0007] "Voice data transmission means" refers to a communication module or protocol for transmitting acquired voice data to a remote server.
[0008] "Speech recognition means" refers to a speech recognition engine or software used to convert acquired speech data into text format.
[0009] "Translation means" refers to translation engines or algorithms used to translate text obtained by speech recognition means into another language.
[0010] "Speech synthesis means" refers to a speech synthesis engine or software used to convert text generated by translation means into speech.
[0011] "Synthesized speech transmission means" refers to a communication module or protocol for transmitting generated speech data to a user device.
[0012] "Translated audio playback means" refers to an audio output device such as a speaker or bone conduction earphones for playing back transmitted audio data. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system for providing an environment in which children can naturally learn a foreign language in their daily lives, and is implemented as follows.
[0035] Device setup and connection
[0036] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. Then, a dedicated application is launched and an internet connection is established.
[0037] Voice acquisition and transmission
[0038] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and sent to a server via the internet. During this process, the audio data is packetized into fixed-capacity chunks to ensure reliable transmission.
[0039] Speech recognition and translation
[0040] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. This text data is then translated into the required foreign language (e.g., English) by a translation engine. The translated text is then checked for grammatical and semantic accuracy using natural language processing (NLP) techniques.
[0041] Speech synthesis and transmission
[0042] The translated text data is converted into English audio data by a speech synthesis engine. This English audio data is compressed and sent to the terminal. To minimize translation delays, the transmitted audio data is processed in real time, packet by packet.
[0043] Audio playback
[0044] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0045] Specific example
[0046] The following are some specific situations.
[0047] The user's parent asks in Japanese, "What did you study at school today?"
[0048] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0049] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0050] The server's translation engine translates "What did you study at school today?" into English.
[0051] The server converts the English text translated by the speech synthesis engine into audio data.
[0052] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0053] This system allows users to listen to everyday conversations in a foreign language in real time, enabling them to efficiently acquire that language.
[0054] The following describes the processing flow.
[0055] Step 1:
[0056] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0057] Step 2:
[0058] The device launches a dedicated app and establishes an internet connection.
[0059] Step 3:
[0060] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0061] Step 4:
[0062] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0063] Step 5:
[0064] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0065] Step 6:
[0066] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0067] Step 7:
[0068] The server passes the translated English text to the speech synthesis engine, which converts it into English speech data.
[0069] Step 8:
[0070] The server compresses the generated English audio data, packets it, and sends it to the terminal.
[0071] Step 9:
[0072] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0073] Step 10:
[0074] The device transmits English audio data to the user's device and plays it back in real time. The user hears the surrounding Japanese conversation as translated English. Efficient data transfer and processing are performed to minimize translation delays.
[0075] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to a server, where it performs translation and speech synthesis before playing the English audio. Through this process, the user can hear the English audio "What did you study at school today?" in real time.
[0076] (Example 1)
[0077] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0078] In modern society, opportunities for children to naturally learn foreign languages in their daily lives are limited. Conventional learning systems and speech recognition systems have difficulty enabling foreign language learning through natural, real-time conversation, and it has been particularly challenging to provide instantly translated foreign languages in a home conversational environment. This invention aims to solve these problems and provide a system that enables children to efficiently learn foreign languages in their daily lives.
[0079] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0080] In this invention, the server includes speech recognition means, translation means, natural language processing means, speech synthesis means, and synthesized speech transmission means. This makes it possible to translate acquired Japanese speech into a foreign language in real time with high accuracy, compress the resulting speech, transmit it to the user's device, and play it back immediately.
[0081] "Voice input means" refers to devices or methods for acquiring ambient sounds or conversations.
[0082] "Voice data transmission means" refers to a device or method for transmitting acquired voice data to other devices or servers.
[0083] "Speech recognition means" refers to a technology or device that analyzes acquired speech data and converts it into corresponding text data.
[0084] "Translation means" refers to a technology or device that converts text data generated by speech recognition means into another language.
[0085] "Speech synthesis means" refers to a technology or device that converts text data generated by translation means into speech data.
[0086] "Synthesized speech transmission means" refers to a technology or device for transmitting generated speech data to another device.
[0087] "Translated audio playback means" refers to technology or equipment for playing back transmitted audio data.
[0088] "Voice input devices" refer to devices used to acquire sound, such as microphones, audio glasses, and bone conduction earphones.
[0089] "Pairing method" refers to a technology or method for connecting an audio input device and a terminal via wireless communication (e.g., Bluetooth or Wi-Fi).
[0090] "Real-time audio data buffering means" refers to a technology or device for temporarily storing acquired audio data and processing it in real time.
[0091] "Method for packetizing voice data" refers to a technology or device that divides voice data to be transmitted into packets of a certain size and transmits them as packets.
[0092] "Natural language processing means" refers to technologies or devices for verifying and correcting the accuracy of the grammar and semantics of translated text or audio data.
[0093] "Audio data compression means" refers to a technology or device that compresses audio data to reduce its size when transmitting it.
[0094] This invention provides a system that offers users an environment in which they can naturally learn a foreign language in their daily lives. Specific embodiments are described below.
[0095] Device setup and connection
[0096] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. This operation launches a dedicated application and establishes an internet connection.
[0097] Hardware and software to use
[0098] Hardware: Audio glasses, bone conduction earphones, smartphones, tablets
[0099] Software: Dedicated applications, speech recognition engine, translation engine, natural language processing (NLP) technology, speech synthesis engine
[0100] Voice acquisition and transmission
[0101] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and transmitted to a server via the internet with high reliability. During this process, the audio data is packetized into fixed-capacity chunks.
[0102] Speech recognition and translation
[0103] The server inputs the received audio data into a speech recognition engine, which converts the Japanese audio into text. This text data is then translated into the required foreign language (e.g., English) by a translation engine. Finally, natural language processing (NLP) techniques are used to verify the accuracy of the grammar and meaning.
[0104] Speech synthesis and transmission
[0105] The translated text data is converted into English audio data by a speech synthesis engine. This data is compressed and sent to the terminal. The audio data is processed in real time on a packet-by-packet basis to minimize transmission delay.
[0106] Audio playback
[0107] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0108] Specific example
[0109] The following are some specific situations.
[0110] The user's parent asks in Japanese, "What did you study at school today?"
[0111] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0112] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0113] The server's translation engine translates "What did you study at school today?" into English.
[0114] The server converts the English text translated by the speech synthesis engine into audio data.
[0115] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0116] Example of a prompt
[0117] The following are specific examples of prompt statements for the generative AI model related to this system.
[0118] "Please tell me how to translate what my parents say in Japanese into English in real time and listen to it using audio glasses."
[0119] "Please explain the mechanism of a system that allows children to naturally learn a foreign language through everyday conversation, including specific devices and steps."
[0120] "Please describe in detail your approach to building a learning system that combines speech recognition and translation functions, including the Japanese-to-English translation process."
[0121] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0122] Step 1:
[0123] The user wears audio glasses or bone conduction earphones that handle audio input and output. Next, the terminal pairs with these devices via Bluetooth or Wi-Fi. The user then launches a dedicated application on the terminal and establishes an internet connection. Input is the confirmation of wearing the audio device and connecting to the terminal, while output is the completion of the connection and the launch of the application. Specific actions include the user turning on the audio device, the terminal detecting the device, and selecting pairing.
[0124] Step 2:
[0125] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time, packetized into fixed-capacity chunks, and sent to a server via the internet. The input is ambient sound, and the output is packetized audio data. Specifically, the device's microphone continuously collects ambient sound, temporarily stores it internally, and then divides the data into packets at the appropriate time.
[0126] Step 3:
[0127] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. Next, the text data is translated into the required foreign language (e.g., English) by a translation engine. Furthermore, the accuracy of the grammar and meaning is checked using natural language processing (NLP) techniques. The input is packetized audio data, and the output is translated text data. Specifically, the process involves the speech recognition engine analyzing the audio data to generate Japanese text, the translation engine translating that text into English, and NLP techniques checking the grammar and meaning.
[0128] Step 4:
[0129] The translated text data is converted into audio data by a speech synthesis engine. This English audio data is compressed, repacked, and sent to the terminal. The input is the translated text data, and the output is the packetized audio data. Specifically, the process involves the speech synthesis engine converting the text data into speech, compressing that audio data, dividing it into packets, and sending them to the terminal.
[0130] Step 5:
[0131] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. The input is packetized audio data, and the output is the audio for the user to listen to. Specifically, the operation involves the device receiving packetized audio data, decompressing it, reconstructing it into audio data, and finally playing it back through the audio device.
[0132] Through the steps described above, this system enables users to listen to translated foreign language audio in real time during their daily lives.
[0133] (Application Example 1)
[0134] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0135] In today's globalized world, acquiring multiple languages from an early age is considered crucial. However, providing an effective environment for children to naturally learn foreign languages is not easy. Traditional foreign language education is often limited to specific times and places, lacking opportunities for natural language acquisition in daily life. Furthermore, systems that provide customized lessons tailored to individual learners while managing progress are still not adequately developed. This reduces the efficiency and sustainability of learning, making the language acquisition process difficult.
[0136] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0137] In this invention, the server includes voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, real-time educational translation means, learning progress management means, and dialogue simulation means. This makes it possible to provide an environment in which children can naturally learn a foreign language in their daily lives. Furthermore, by tracking the user's learning progress and providing customized lessons tailored to each individual's progress, the efficiency of learning is improved and sustained learning is promoted.
[0138] A "voice input device" is a device that has the function of acquiring sounds from the user's surroundings.
[0139] A "voice data transmission means" is a device that has the function of transmitting acquired voice data to a server.
[0140] "Speech recognition means" refers to technology that converts transmitted speech data into text data.
[0141] "Translation means" refers to technology that translates text data generated by speech recognition means into another language.
[0142] "Speech synthesis means" refers to a technology that converts text data generated by translation means into speech data.
[0143] "Synthesized speech transmission means" refers to a function for transmitting speech data generated by speech synthesis means to the user's device.
[0144] "Translated audio playback means" refers to a function that plays back audio data received on the user's device.
[0145] "Educational real-time translation technology" is a technology that translates surrounding conversations in real time so that children can naturally learn a foreign language in their daily lives.
[0146] A "learning progress management system" is a technology that tracks a user's learning progress and provides customized lessons according to their progress.
[0147] A "dialogue simulation method" is a technology that provides a simulation of a conversation based on a specific situation.
[0148] The present invention is a system that provides an environment in which children can naturally learn a foreign language in their daily lives. This system includes the following functions: voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, educational real-time translation means, learning progress management means, and dialogue simulation means.
[0149] First, the user wears audio glasses or bone conduction earphones as a means of voice input and uses a smartphone or tablet as a terminal. The terminal is paired with the device via Bluetooth or Wi-Fi and a dedicated application is launched. The terminal captures ambient sound through its microphone, and the acquired audio data is sent to a server via the internet. At this time, the audio data is packetized to ensure reliable transmission. The server inputs the received audio data into a speech recognition engine and converts the Japanese speech into text format.
[0150] The converted text data is translated into the required foreign language (e.g., English) by a translation engine. The translated text is then converted into audio data by a speech synthesis engine, compressed, and sent to the terminal. The terminal decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in a foreign language in real time.
[0151] The real-time educational translation system translates surrounding conversations in real time, providing children with an environment where they can naturally learn a foreign language. This feature makes it easier for them to develop the habit of listening to a foreign language in their daily lives. Furthermore, the learning progress management system tracks the user's learning progress and provides customized lessons according to their individual progress. This feature enables efficient and sustainable learning. The dialogue simulation system provides conversation simulations based on specific situations, supporting practical language learning.
[0152] Hardware and software to be used
[0153] Hardware: Smartphone or tablet, audio glasses or bone conduction earphones, microphone
[0154] Software: Python, speech_recognition library, googletrans library, pyttsx3 library
[0155] Data processing and data calculation
[0156] 1. The voice input device captures voice data through the microphone.
[0157] 2. The audio data transmission means sends the acquired audio data to the server.
[0158] 3. The server's speech recognition system converts the speech data into text data.
[0159] 4. The translation method translates text data into a foreign language using a translation engine.
[0160] 5. The speech synthesis means converts the translated text into speech data.
[0161] 6. The synthesized speech transmission means transmits the voice data to the terminal.
[0162] 7. The translated audio playback device decodes the audio data and plays it back to the user.
[0163] 8. Educational real-time translation tools translate everyday conversations in real time.
[0164] 9. A learning progress management system manages the user's learning status and provides customized lessons.
[0165] 10. The dialogue simulation means provides conversation simulation.
[0166] Specific example
[0167] For example, if the user's parent says in Japanese, "What did you study at school today?", the device picks up the parent's voice with its microphone and sends it to the server in real time. The server uses a speech recognition engine to convert it into Japanese text, "What did you study at school today?", and then a translation engine translates it into English as "What did you study at school today?". The translated English text is converted into audio data by a speech synthesis engine, and this converted audio data is sent to the device, which decodes it and sends it to the user's device for playback.
[0168] Example of a prompt
[0169] "I want to develop a tool that translates everyday conversations into English to help with foreign language learning. The core function is real-time voice translation. Please tell me about the specific functions of this tool and how to implement them."
[0170] As a result, by using the system of the present invention, an environment is created in which children can naturally learn foreign languages in their daily lives, promoting efficient and sustainable learning.
[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0172] Step 1:
[0173] The user wears audio glasses or bone conduction earphones and pairs them with a smartphone or tablet. The input is the pairing information between the terminal and the device, and each establishes a stable connection via Bluetooth or Wi-Fi. The output is the status of successful pairing.
[0174] Step 2:
[0175] The device captures ambient sound through its microphone. The input is ambient sound, which the device's microphone picks up. The output is audio data, which is buffered.
[0176] Step 3:
[0177] Audio data is sent to a server via the internet. The input is buffered audio data, which is reliably transmitted to the server. The output is the audio data stored on the server.
[0178] Step 4:
[0179] The server uses a speech recognition engine to convert audio data into text data. The input is audio data stored on the server, which the speech recognition engine analyzes and converts into text. The output is text data.
[0180] Step 5:
[0181] The server translates text data into a foreign language using a translation engine. The input is text data generated by speech recognition, which the translation engine then translates. The output is the translated foreign language text.
[0182] Step 6:
[0183] The server converts translated text data into speech data using a speech synthesis engine. The input is translated foreign language text, which the speech synthesis engine then converts into speech data. The output is synthesized speech data.
[0184] Step 7:
[0185] The server compresses the synthesized speech data and sends it to the terminal via the internet. The input is synthesized speech data, which is sent to the terminal after compression. The output is the audio data received by the terminal.
[0186] Step 8:
[0187] The terminal decodes the received audio data and plays it back through the user's device. The input is the received audio data, which is decoded and then played back. The output is a foreign language audio heard by the user.
[0188] Step 9:
[0189] The server manages the user's learning progress based on their current status and provides personalized lessons. The input consists of the user's learning data and progress information, which are then analyzed to generate appropriate lessons. The output is a customized lesson plan.
[0190] Step 10:
[0191] The device performs a dialogue simulation based on a specific situation. The input is pre-configured situation information, and the dialogue content experienced by the user is generated through the simulation process. The output is the result of the dialogue simulation with the user.
[0192] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0193] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[0194] Device setup and connection
[0195] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0196] Voice acquisition and transmission
[0197] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0198] Speech recognition and translation
[0199] The server receives the audio data and converts it into text format using a speech recognition engine. The converted Japanese text data is then passed to a translation engine and translated into English text data.
[0200] Emotion recognition and analysis
[0201] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. The analysis results are added as supplementary information to the text data.
[0202] Speech synthesis and emotion regulation
[0203] The translated English text and added sentiment information are passed to the speech synthesis engine. Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0204] Audio transmission and playback
[0205] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[0206] Specific example
[0207] The following are some specific situations.
[0208] The user's parent asks in Japanese, "What did you study at school today?"
[0209] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0210] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0211] The server uses a translation engine to translate "What did you study at school today?" into English.
[0212] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[0213] The English text, with added emotional information, is passed to the speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[0214] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[0215] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[0216] This system allows children to hear everyday conversations in a foreign language in real time, while simultaneously providing natural translations that help them understand the speaker's emotions.
[0217] The following describes the processing flow.
[0218] Step 1:
[0219] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0220] Step 2:
[0221] The device launches a dedicated app and establishes an internet connection.
[0222] Step 3:
[0223] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0224] Step 4:
[0225] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0226] Step 5:
[0227] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0228] Step 6:
[0229] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0230] Step 7:
[0231] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. Emotional information is added to the text data as supplementary information.
[0232] Step 8:
[0233] The server passes the translated English text and added sentiment information to the speech synthesis engine, which converts it into English speech data. The sentiment engine adjusts the tone and pitch of the speech synthesis based on the sentiment information.
[0234] Step 9:
[0235] The server compresses the generated audio data, packets it, and sends it to the terminal.
[0236] Step 10:
[0237] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0238] Step 11:
[0239] The device transmits English audio data to the user's device and plays it back in real time. The user can hear the surrounding Japanese conversation as natural English that reflects emotions.
[0240] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to the server. After the server performs translation and speech synthesis, it uses emotion recognition to analyze the user's parent's emotions (for example, "joy of asking"), and sends audio data reflecting the results back to the device. The device then plays the English audio, allowing the user to hear the English phrase "What did you study at school today?" in real time, conveying the emotion behind it.
[0241] (Example 2)
[0242] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0243] Conventional speech translation systems simply translate speech without considering emotional information, resulting in unnatural-sounding translations that fail to convey the speaker's emotions. Furthermore, there has been no system that provides natural-sounding translations that reflect the speaker's emotions, which is crucial for children learning foreign languages in their daily lives. This can lead to low user satisfaction and a decrease in the efficiency of foreign language learning.
[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0245] In this invention, the server includes speech recognition means, translation means, emotion recognition means, emotion information addition means, and speech synthesis means. This makes it possible to consider emotion information when translating speech and generate natural speech that reflects the speaker's emotions.
[0246] "Voice input means" refers to a device or function for acquiring ambient sounds.
[0247] "Voice data transmission means" refers to a means for transmitting acquired voice data to a specific server or terminal.
[0248] "Speech recognition means" refers to a technology or device for converting speech data into text format.
[0249] "Translation means" refers to a technology or device for translating converted text data into another language.
[0250] "Emotion recognition means" refers to a technology or device that analyzes a speaker's emotions from audio data.
[0251] A "means for adding emotional information" refers to a means for adding analyzed emotional information to text data.
[0252] "Speech synthesis means" refers to a technology or device for converting text data into speech data.
[0253] A "synthesized speech transmission means" is a means for transmitting generated speech data to a specific terminal or device.
[0254] A "translation audio playback means" is a function for playing back audio data received by a terminal or device.
[0255] This invention provides a system that allows children to naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system consists of the following elements:
[0256] Device setup and connection
[0257] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones). Next, a terminal (e.g., a smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Subsequently, the terminal launches a dedicated application and establishes an internet connection.
[0258] Specific example:
[0259] The user puts on the bone conduction earphones, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[0260] Voice acquisition and transmission
[0261] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[0262] Specific example:
[0263] When the user's parent asks, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to a server via the internet.
[0264] Speech recognition and translation
[0265] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[0266] Specific example:
[0267] The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[0268] Emotion recognition and analysis
[0269] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[0270] Specific example:
[0271] The server inputs the audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[0272] Speech synthesis and emotion regulation
[0273] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[0274] Specific example:
[0275] The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[0276] Audio transmission and playback
[0277] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[0278] Specific example:
[0279] The translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is then sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[0280] Example of a prompt
[0281] "When a parent asks, 'What did you study at school today?', please translate this into English and play an audio recording that reflects emotional information."
[0282] The flow of the specific process in Example 2 will be described using FIG. 13.
[0283] Step 1: Device Setup and Connection
[0284] The user wears an audio device (e.g., audio glasses or bone conduction earphones) for voice input and output, and the terminal (e.g., smartphone or tablet) is paired with the audio device via Bluetooth or Wi-Fi. Next, the terminal launches a dedicated application and establishes an Internet connection.
[0285] Specific operation: The user wears bone conduction earphones on the ears, opens the Bluetooth settings of the smartphone to pair with the earphones. Then launches the dedicated application and connects to the Internet. The application displays a connection confirmation message.
[0286] Input: The user wears the audio device and pairs it with the terminal.
[0287] Output: The dedicated application of the terminal is launched and an Internet connection is established.
[0288] Step 2: Voice Acquisition and Transmission
[0289] The terminal uses the built-in microphone to acquire ambient conversation voices in real time. The acquired voice data is buffered and transmitted to the server via the Internet.
[0290] Specific operation: When the user's parent says, "What did you study at school today?", the microphone of the terminal acquires the voice. The voice data is buffered and transmitted to the server via the Internet.
[0291] Input: The conversation voice of the user's parent.
[0292] Output: The voice data transmitted to the server.
[0293] Step 3: Speech Recognition and Translation
[0294] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[0295] Specific operation: The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[0296] Input: Acquired audio data.
[0297] Output: Translated English text data.
[0298] Step 4: Emotion Recognition and Analysis
[0299] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[0300] Specific operation: The server inputs audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[0301] Input: Audio data and translated English text data.
[0302] Output: English text data with added sentiment information.
[0303] Step 5: Speech synthesis and emotion regulation
[0304] The translated English text and the emotion information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The emotion engine adjusts the tone and pitch of the voice based on the emotion information to generate natural voice data.
[0305] Specific operation: The English text "What did you study at school today?" and the emotion information are passed to the speech synthesis engine. Tone and pitch adjustments based on the emotion are made, and natural voice data is generated.
[0306] Input: English text data with emotion information added.
[0307] Output: Voice data adjusted based on the emotion.
[0308] Step 6: Transmission and playback of the voice
[0309] The generated voice data is sent by the server to the terminal. The terminal sends the decoded voice data to the user's device and plays it back in real time.
[0310] Specific operation: The translated voice data is sent from the server to the terminal and is immediately decoded. The decoded voice is sent to the bone conduction earphone, and the user can listen to the voice "What did you study at school today?" in real time.
[0311] Input: The generated voice data.
[0312] Output: The voice played back from the user's device.
[0313] (Application Example 2)
[0314] Next, Application Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0315] Modern children do not have sufficient opportunities to naturally learn foreign languages in their daily lives. Furthermore, there is a lack of systems that provide natural-sounding translations that take emotional states into account in real time. This creates challenges in improving learning efficiency and the depth of foreign language comprehension.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0317] In this invention, the server includes a voice input means, an emotion recognition means, and a speech synthesis means. This allows children to listen to everyday conversations in real time as a foreign language while also understanding the emotions of the speakers.
[0318] "Voice input means" refers to hardware and software for acquiring ambient sounds.
[0319] "Voice data transmission means" refers to a means for transmitting acquired voice data to a server.
[0320] A "speech recognition means" is an engine for converting speech data into text format.
[0321] A "translation tool" is an engine used to translate text data into another language.
[0322] An "emotion recognition tool" is an algorithm used to analyze the emotional state of the speaker.
[0323] A "speech synthesis means" is an engine for generating speech from text data.
[0324] A "synthesized voice transmission means" is a means for transmitting generated voice data to a terminal.
[0325] A "translation audio playback means" is a means for playing back audio transmitted to a terminal.
[0326] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[0327] Device setup and connection
[0328] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0329] Voice acquisition and transmission
[0330] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0331] Speech recognition and translation
[0332] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[0333] Emotion recognition and analysis
[0334] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[0335] Speech synthesis and emotion regulation
[0336] The translated English text and added sentiment information are passed to a speech synthesis engine (e.g., pyttsx3). Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0337] Audio transmission and playback
[0338] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[0339] Specific example
[0340] The following are some specific situations.
[0341] The user's parent asks in Japanese, "What did you study at school today?"
[0342] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0343] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0344] The server uses a translation engine to translate "What did you study at school today?" into English.
[0345] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[0346] English text with added emotional information is passed to a speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[0347] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[0348] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[0349] As a concrete example of using a generative AI model to perform emotion recognition and reflecting the results in speech synthesis, the following prompt sentence is used:
[0350] Analyze the speaker's emotions based on the question, "What did you study at school today?"
[0351] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0352] Step 1:
[0353] Device setup and connection
[0354] The user wears audio glasses or bone conduction earphones for voice input and output. Next, a device (smartphone or tablet) is paired with these devices via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0355] Input: Audio glasses or bone conduction earphones, device (smartphone or tablet)
[0356] Output: Paired device, launched dedicated application
[0357] Step 2:
[0358] Voice acquisition and transmission
[0359] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0360] Input: Surrounding conversation audio
[0361] Output: Audio data sent to the server (after buffering)
[0362] Step 3:
[0363] Speech recognition and translation
[0364] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[0365] Input: Audio data sent to the server
[0366] Output: Translated English text data
[0367] Step 4:
[0368] Emotion recognition and analysis
[0369] The server uses voice data acquired via voice input and translated text data to analyze the user's speech emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[0370] Input: Translated English text data, acquired audio data
[0371] Output: Text data with emotional information added.
[0372] Step 5:
[0373] Speech synthesis and emotion regulation
[0374] The English text data, which includes emotional information, is passed to a speech synthesis engine (e.g., PyttsX3). Based on the emotional information, the emotion engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0375] Input: English text data with emotional information attached
[0376] Output: Adjusted natural audio data
[0377] Step 6:
[0378] Audio transmission and playback
[0379] The generated audio data is sent from the server to the terminal. The decoded audio data is sent to the user's device and played back in real time. This allows the user to hear surrounding Japanese conversations as natural-sounding English audio with corresponding emotions.
[0380] Input: Adjusted natural voice data
[0381] Output: Audio played on the user's device
[0382] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0383] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0384] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0385] [Second Embodiment]
[0386] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0387] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0388] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0389] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0390] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0392] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0393] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0394] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0395] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0396] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0397] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0398] This invention is a system for providing an environment in which children can naturally learn a foreign language in their daily lives, and is implemented as follows.
[0399] Device setup and connection
[0400] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. Then, a dedicated application is launched and an internet connection is established.
[0401] Voice acquisition and transmission
[0402] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and sent to a server via the internet. During this process, the audio data is packetized into fixed-capacity chunks to ensure reliable transmission.
[0403] Speech recognition and translation
[0404] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. This text data is then translated into the required foreign language (e.g., English) by a translation engine. The translated text is then checked for grammatical and semantic accuracy using natural language processing (NLP) techniques.
[0405] Speech synthesis and transmission
[0406] The translated text data is converted into English audio data by a speech synthesis engine. This English audio data is compressed and sent to the terminal. To minimize translation delays, the transmitted audio data is processed in real time, packet by packet.
[0407] Audio playback
[0408] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0409] Specific example
[0410] The following are some specific situations.
[0411] The user's parent asks in Japanese, "What did you study at school today?"
[0412] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0413] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0414] The server's translation engine translates "What did you study at school today?" into English.
[0415] The server converts the English text translated by the speech synthesis engine into audio data.
[0416] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0417] This system allows users to listen to everyday conversations in a foreign language in real time, enabling them to efficiently acquire that language.
[0418] The following describes the processing flow.
[0419] Step 1:
[0420] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0421] Step 2:
[0422] The device launches a dedicated app and establishes an internet connection.
[0423] Step 3:
[0424] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0425] Step 4:
[0426] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0427] Step 5:
[0428] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0429] Step 6:
[0430] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0431] Step 7:
[0432] The server passes the translated English text to the speech synthesis engine, which converts it into English speech data.
[0433] Step 8:
[0434] The server compresses the generated English audio data, packets it, and sends it to the terminal.
[0435] Step 9:
[0436] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0437] Step 10:
[0438] The device transmits English audio data to the user's device and plays it back in real time. The user hears the surrounding Japanese conversation as translated English. Efficient data transfer and processing are performed to minimize translation delays.
[0439] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to a server, where it performs translation and speech synthesis before playing the English audio. Through this process, the user can hear the English audio "What did you study at school today?" in real time.
[0440] (Example 1)
[0441] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0442] In modern society, opportunities for children to naturally learn foreign languages in their daily lives are limited. Conventional learning systems and speech recognition systems have difficulty enabling foreign language learning through natural, real-time conversation, and it has been particularly challenging to provide instantly translated foreign languages in a home conversational environment. This invention aims to solve these problems and provide a system that enables children to efficiently learn foreign languages in their daily lives.
[0443] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0444] In this invention, the server includes speech recognition means, translation means, natural language processing means, speech synthesis means, and synthesized speech transmission means. This makes it possible to translate acquired Japanese speech into a foreign language in real time with high accuracy, compress the resulting speech, transmit it to the user's device, and play it back immediately.
[0445] "Voice input means" refers to devices or methods for acquiring ambient sounds or conversations.
[0446] "Voice data transmission means" refers to a device or method for transmitting acquired voice data to other devices or servers.
[0447] "Speech recognition means" refers to a technology or device that analyzes acquired speech data and converts it into corresponding text data.
[0448] "Translation means" refers to a technology or device that converts text data generated by speech recognition means into another language.
[0449] "Speech synthesis means" refers to a technology or device that converts text data generated by translation means into speech data.
[0450] "Synthesized speech transmission means" refers to a technology or device for transmitting generated speech data to another device.
[0451] "Translated audio playback means" refers to technology or equipment for playing back transmitted audio data.
[0452] "Voice input devices" refer to devices used to acquire sound, such as microphones, audio glasses, and bone conduction earphones.
[0453] "Pairing method" refers to a technology or method for connecting an audio input device and a terminal via wireless communication (e.g., Bluetooth or Wi-Fi).
[0454] "Real-time audio data buffering means" refers to a technology or device for temporarily storing acquired audio data and processing it in real time.
[0455] "Method for packetizing voice data" refers to a technology or device that divides voice data to be transmitted into packets of a certain size and transmits them as packets.
[0456] "Natural language processing means" refers to technologies or devices for verifying and correcting the accuracy of the grammar and semantics of translated text or audio data.
[0457] "Audio data compression means" refers to a technology or device that compresses audio data to reduce its size when transmitting it.
[0458] This invention provides a system that offers users an environment in which they can naturally learn a foreign language in their daily lives. Specific embodiments are described below.
[0459] Device setup and connection
[0460] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. This operation launches a dedicated application and establishes an internet connection.
[0461] Hardware and software to use
[0462] Hardware: Audio glasses, bone conduction earphones, smartphones, tablets
[0463] Software: Dedicated applications, speech recognition engine, translation engine, natural language processing (NLP) technology, speech synthesis engine
[0464] Voice acquisition and transmission
[0465] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and transmitted to a server via the internet with high reliability. During this process, the audio data is packetized into fixed-capacity chunks.
[0466] Speech recognition and translation
[0467] The server inputs the received audio data into a speech recognition engine, which converts the Japanese audio into text. This text data is then translated into the required foreign language (e.g., English) by a translation engine. Finally, natural language processing (NLP) techniques are used to verify the accuracy of the grammar and meaning.
[0468] Speech synthesis and transmission
[0469] The translated text data is converted into English audio data by a speech synthesis engine. This data is compressed and sent to the terminal. The audio data is processed in real time on a packet-by-packet basis to minimize transmission delay.
[0470] Audio playback
[0471] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0472] Specific example
[0473] The following are some specific situations.
[0474] The user's parent asks in Japanese, "What did you study at school today?"
[0475] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0476] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0477] The server's translation engine translates "What did you study at school today?" into English.
[0478] The server converts the English text translated by the speech synthesis engine into audio data.
[0479] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0480] Example of a prompt
[0481] The following are specific examples of prompt statements for the generative AI model related to this system.
[0482] "Please tell me how to translate what my parents say in Japanese into English in real time and listen to it using audio glasses."
[0483] "Please explain the mechanism of a system that allows children to naturally learn a foreign language through everyday conversation, including specific devices and steps."
[0484] "Please describe in detail your approach to building a learning system that combines speech recognition and translation functions, including the Japanese-to-English translation process."
[0485] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0486] Step 1:
[0487] The user wears audio glasses or bone conduction earphones that handle audio input and output. Next, the terminal pairs with these devices via Bluetooth or Wi-Fi. The user then launches a dedicated application on the terminal and establishes an internet connection. Input is the confirmation of wearing the audio device and connecting to the terminal, while output is the completion of the connection and the launch of the application. Specific actions include the user turning on the audio device, the terminal detecting the device, and selecting pairing.
[0488] Step 2:
[0489] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time, packetized into fixed-capacity chunks, and sent to a server via the internet. The input is ambient sound, and the output is packetized audio data. Specifically, the device's microphone continuously collects ambient sound, temporarily stores it internally, and then divides the data into packets at the appropriate time.
[0490] Step 3:
[0491] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. Next, the text data is translated into the required foreign language (e.g., English) by a translation engine. Furthermore, the accuracy of the grammar and meaning is checked using natural language processing (NLP) techniques. The input is packetized audio data, and the output is translated text data. Specifically, the process involves the speech recognition engine analyzing the audio data to generate Japanese text, the translation engine translating that text into English, and NLP techniques checking the grammar and meaning.
[0492] Step 4:
[0493] The translated text data is converted into audio data by a speech synthesis engine. This English audio data is compressed, repacked, and sent to the terminal. The input is the translated text data, and the output is the packetized audio data. Specifically, the process involves the speech synthesis engine converting the text data into speech, compressing that audio data, dividing it into packets, and sending them to the terminal.
[0494] Step 5:
[0495] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. The input is packetized audio data, and the output is the audio for the user to listen to. Specifically, the operation involves the device receiving packetized audio data, decompressing it, reconstructing it into audio data, and finally playing it back through the audio device.
[0496] Through the steps described above, this system enables users to listen to translated foreign language audio in real time during their daily lives.
[0497] (Application Example 1)
[0498] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0499] In today's globalized world, acquiring multiple languages from an early age is considered crucial. However, providing an effective environment for children to naturally learn foreign languages is not easy. Traditional foreign language education is often limited to specific times and places, lacking opportunities for natural language acquisition in daily life. Furthermore, systems that provide customized lessons tailored to individual learners while managing progress are still not adequately developed. This reduces the efficiency and sustainability of learning, making the language acquisition process difficult.
[0500] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0501] In this invention, the server includes voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, real-time educational translation means, learning progress management means, and dialogue simulation means. This makes it possible to provide an environment in which children can naturally learn a foreign language in their daily lives. Furthermore, by tracking the user's learning progress and providing customized lessons tailored to each individual's progress, the efficiency of learning is improved and sustained learning is promoted.
[0502] A "voice input device" is a device that has the function of acquiring sounds from the user's surroundings.
[0503] A "voice data transmission means" is a device that has the function of transmitting acquired voice data to a server.
[0504] "Speech recognition means" refers to technology that converts transmitted speech data into text data.
[0505] "Translation means" refers to technology that translates text data generated by speech recognition means into another language.
[0506] "Speech synthesis means" refers to a technology that converts text data generated by translation means into speech data.
[0507] "Synthesized speech transmission means" refers to a function for transmitting speech data generated by speech synthesis means to the user's device.
[0508] "Translated audio playback means" refers to a function that plays back audio data received on the user's device.
[0509] "Educational real-time translation technology" is a technology that translates surrounding conversations in real time so that children can naturally learn a foreign language in their daily lives.
[0510] A "learning progress management system" is a technology that tracks a user's learning progress and provides customized lessons according to their progress.
[0511] A "dialogue simulation method" is a technology that provides a simulation of a conversation based on a specific situation.
[0512] The present invention is a system that provides an environment in which children can naturally learn a foreign language in their daily lives. This system includes the following functions: voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, educational real-time translation means, learning progress management means, and dialogue simulation means.
[0513] First, the user wears audio glasses or bone conduction earphones as a means of voice input and uses a smartphone or tablet as a terminal. The terminal is paired with the device via Bluetooth or Wi-Fi and a dedicated application is launched. The terminal captures ambient sound through its microphone, and the acquired audio data is sent to a server via the internet. At this time, the audio data is packetized to ensure reliable transmission. The server inputs the received audio data into a speech recognition engine and converts the Japanese speech into text format.
[0514] The converted text data is translated into the required foreign language (e.g., English) by a translation engine. The translated text is then converted into audio data by a speech synthesis engine, compressed, and sent to the terminal. The terminal decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in a foreign language in real time.
[0515] The real-time educational translation system translates surrounding conversations in real time, providing children with an environment where they can naturally learn a foreign language. This feature makes it easier for them to develop the habit of listening to a foreign language in their daily lives. Furthermore, the learning progress management system tracks the user's learning progress and provides customized lessons according to their individual progress. This feature enables efficient and sustainable learning. The dialogue simulation system provides conversation simulations based on specific situations, supporting practical language learning.
[0516] Hardware and software to be used
[0517] Hardware: Smartphone or tablet, audio glasses or bone conduction earphones, microphone
[0518] Software: Python, speech_recognition library, googletrans library, pyttsx3 library
[0519] Data processing and data calculation
[0520] 1. The voice input device captures voice data through the microphone.
[0521] 2. The audio data transmission means sends the acquired audio data to the server.
[0522] 3. The server's speech recognition system converts the speech data into text data.
[0523] 4. The translation method translates text data into a foreign language using a translation engine.
[0524] 5. The speech synthesis means converts the translated text into speech data.
[0525] 6. The synthesized speech transmission means transmits the voice data to the terminal.
[0526] 7. The translated audio playback device decodes the audio data and plays it back to the user.
[0527] 8. Educational real-time translation tools translate everyday conversations in real time.
[0528] 9. A learning progress management system manages the user's learning status and provides customized lessons.
[0529] 10. The dialogue simulation means provides conversation simulation.
[0530] Specific example
[0531] For example, if the user's parent says in Japanese, "What did you study at school today?", the device picks up the parent's voice with its microphone and sends it to the server in real time. The server uses a speech recognition engine to convert it into Japanese text, "What did you study at school today?", and then a translation engine translates it into English as "What did you study at school today?". The translated English text is converted into audio data by a speech synthesis engine, and this converted audio data is sent to the device, which decodes it and sends it to the user's device for playback.
[0532] Example of a prompt
[0533] "I want to develop a tool that translates everyday conversations into English to help with foreign language learning. The core function is real-time voice translation. Please tell me about the specific functions of this tool and how to implement them."
[0534] As a result, by using the system of the present invention, an environment is created in which children can naturally learn foreign languages in their daily lives, promoting efficient and sustainable learning.
[0535] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0536] Step 1:
[0537] The user wears audio glasses or bone conduction earphones and pairs them with a smartphone or tablet. The input is the pairing information between the terminal and the device, and each establishes a stable connection via Bluetooth or Wi-Fi. The output is the status of successful pairing.
[0538] Step 2:
[0539] The device captures ambient sound through its microphone. The input is ambient sound, which the device's microphone picks up. The output is audio data, which is buffered.
[0540] Step 3:
[0541] Audio data is sent to a server via the internet. The input is buffered audio data, which is reliably transmitted to the server. The output is the audio data stored on the server.
[0542] Step 4:
[0543] The server uses a speech recognition engine to convert audio data into text data. The input is audio data stored on the server, which the speech recognition engine analyzes and converts into text. The output is text data.
[0544] Step 5:
[0545] The server translates text data into a foreign language using a translation engine. The input is text data generated by speech recognition, which the translation engine then translates. The output is the translated foreign language text.
[0546] Step 6:
[0547] The server converts translated text data into speech data using a speech synthesis engine. The input is translated foreign language text, which the speech synthesis engine then converts into speech data. The output is synthesized speech data.
[0548] Step 7:
[0549] The server compresses the synthesized speech data and sends it to the terminal via the internet. The input is synthesized speech data, which is sent to the terminal after compression. The output is the audio data received by the terminal.
[0550] Step 8:
[0551] The terminal decodes the received audio data and plays it back through the user's device. The input is the received audio data, which is decoded and then played back. The output is a foreign language audio heard by the user.
[0552] Step 9:
[0553] The server manages the user's learning progress based on their current status and provides personalized lessons. The input consists of the user's learning data and progress information, which are then analyzed to generate appropriate lessons. The output is a customized lesson plan.
[0554] Step 10:
[0555] The device performs a dialogue simulation based on a specific situation. The input is pre-configured situation information, and the dialogue content experienced by the user is generated through the simulation process. The output is the result of the dialogue simulation with the user.
[0556] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0557] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[0558] Device setup and connection
[0559] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0560] Voice acquisition and transmission
[0561] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0562] Speech recognition and translation
[0563] The server receives the audio data and converts it into text format using a speech recognition engine. The converted Japanese text data is then passed to a translation engine and translated into English text data.
[0564] Emotion recognition and analysis
[0565] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. The analysis results are added as supplementary information to the text data.
[0566] Speech synthesis and emotion regulation
[0567] The translated English text and added sentiment information are passed to the speech synthesis engine. Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0568] Audio transmission and playback
[0569] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[0570] Specific example
[0571] The following are some specific situations.
[0572] The user's parent asks in Japanese, "What did you study at school today?"
[0573] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0574] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0575] The server uses a translation engine to translate "What did you study at school today?" into English.
[0576] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[0577] The English text, with added emotional information, is passed to the speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[0578] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[0579] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[0580] This system allows children to hear everyday conversations in a foreign language in real time, while simultaneously providing natural translations that help them understand the speaker's emotions.
[0581] The following describes the processing flow.
[0582] Step 1:
[0583] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0584] Step 2:
[0585] The device launches a dedicated app and establishes an internet connection.
[0586] Step 3:
[0587] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0588] Step 4:
[0589] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0590] Step 5:
[0591] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0592] Step 6:
[0593] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0594] Step 7:
[0595] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. Emotional information is added to the text data as supplementary information.
[0596] Step 8:
[0597] The server passes the translated English text and added sentiment information to the speech synthesis engine, which converts it into English speech data. The sentiment engine adjusts the tone and pitch of the speech synthesis based on the sentiment information.
[0598] Step 9:
[0599] The server compresses the generated audio data, packets it, and sends it to the terminal.
[0600] Step 10:
[0601] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0602] Step 11:
[0603] The device transmits English audio data to the user's device and plays it back in real time. The user can hear the surrounding Japanese conversation as natural English that reflects emotions.
[0604] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to the server. After the server performs translation and speech synthesis, it uses emotion recognition to analyze the user's parent's emotions (for example, "joy of asking"), and sends audio data reflecting the results back to the device. The device then plays the English audio, allowing the user to hear the English phrase "What did you study at school today?" in real time, conveying the emotion behind it.
[0605] (Example 2)
[0606] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0607] Conventional speech translation systems simply translate speech without considering emotional information, resulting in unnatural-sounding translations that fail to convey the speaker's emotions. Furthermore, there has been no system that provides natural-sounding translations that reflect the speaker's emotions, which is crucial for children learning foreign languages in their daily lives. This can lead to low user satisfaction and a decrease in the efficiency of foreign language learning.
[0608] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0609] In this invention, the server includes speech recognition means, translation means, emotion recognition means, emotion information addition means, and speech synthesis means. This makes it possible to consider emotion information when translating speech and generate natural speech that reflects the speaker's emotions.
[0610] "Voice input means" refers to a device or function for acquiring ambient sounds.
[0611] "Voice data transmission means" refers to a means for transmitting acquired voice data to a specific server or terminal.
[0612] "Speech recognition means" refers to a technology or device for converting speech data into text format.
[0613] "Translation means" refers to a technology or device for translating converted text data into another language.
[0614] "Emotion recognition means" refers to a technology or device that analyzes a speaker's emotions from audio data.
[0615] A "means for adding emotional information" refers to a means for adding analyzed emotional information to text data.
[0616] "Speech synthesis means" refers to a technology or device for converting text data into speech data.
[0617] A "synthesized speech transmission means" is a means for transmitting generated speech data to a specific terminal or device.
[0618] A "translation audio playback means" is a function for playing back audio data received by a terminal or device.
[0619] This invention provides a system that allows children to naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system consists of the following elements:
[0620] Device setup and connection
[0621] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones). Next, a terminal (e.g., a smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Subsequently, the terminal launches a dedicated application and establishes an internet connection.
[0622] Specific example:
[0623] The user puts on the bone conduction earphones, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[0624] Voice acquisition and transmission
[0625] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[0626] Specific example:
[0627] When the user's parent asks, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to a server via the internet.
[0628] Speech recognition and translation
[0629] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[0630] Specific example:
[0631] The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[0632] Emotion recognition and analysis
[0633] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[0634] Specific example:
[0635] The server inputs the audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[0636] Speech synthesis and emotion regulation
[0637] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[0638] Specific example:
[0639] The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[0640] Audio transmission and playback
[0641] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[0642] Specific example:
[0643] The translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is then sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[0644] Example of a prompt
[0645] "When a parent asks, 'What did you study at school today?', please translate this into English and play an audio recording that reflects emotional information."
[0646] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0647] Step 1: Device setup and connection
[0648] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones), and the terminal (e.g., smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Next, the terminal launches a dedicated application and establishes an internet connection.
[0649] Specific steps: The user places the bone conduction earphones in their ears, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[0650] Input: The user wears an audio device and pairs it with the terminal.
[0651] Output: The terminal's dedicated application launches and an internet connection is established.
[0652] Step 2: Acquiring and transmitting audio
[0653] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[0654] Specific operation: When the user's parent says, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to the server via the internet.
[0655] Input: Audio recording of the user's parent's conversation.
[0656] Output: Audio data sent to the server.
[0657] Step 3: Speech Recognition and Translation
[0658] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[0659] Specific operation: The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[0660] Input: Acquired audio data.
[0661] Output: Translated English text data.
[0662] Step 4: Emotion Recognition and Analysis
[0663] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[0664] Specific operation: The server inputs audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[0665] Input: Audio data and translated English text data.
[0666] Output: English text data with added sentiment information.
[0667] Step 5: Speech synthesis and emotion regulation
[0668] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[0669] Specific process: The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[0670] Input: English text data with added sentiment information.
[0671] Output: Voice data adjusted based on emotion.
[0672] Step 6: Sending and Playing Audio
[0673] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[0674] Specific operation: Translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[0675] Input: Generated audio data.
[0676] Output: Audio played from the user's device.
[0677] (Application Example 2)
[0678] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0679] Modern children do not have sufficient opportunities to naturally learn foreign languages in their daily lives. Furthermore, there is a lack of systems that provide natural-sounding translations that take emotional states into account in real time. This creates challenges in improving learning efficiency and the depth of foreign language comprehension.
[0680] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0681] In this invention, the server includes a voice input means, an emotion recognition means, and a speech synthesis means. This allows children to listen to everyday conversations in real time as a foreign language while also understanding the emotions of the speakers.
[0682] "Voice input means" refers to hardware and software for acquiring ambient sounds.
[0683] "Voice data transmission means" refers to a means for transmitting acquired voice data to a server.
[0684] A "speech recognition means" is an engine for converting speech data into text format.
[0685] A "translation tool" is an engine used to translate text data into another language.
[0686] An "emotion recognition tool" is an algorithm used to analyze the emotional state of the speaker.
[0687] A "speech synthesis means" is an engine for generating speech from text data.
[0688] A "synthesized voice transmission means" is a means for transmitting generated voice data to a terminal.
[0689] A "translation audio playback means" is a means for playing back audio transmitted to a terminal.
[0690] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[0691] Device setup and connection
[0692] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0693] Voice acquisition and transmission
[0694] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0695] Speech recognition and translation
[0696] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[0697] Emotion recognition and analysis
[0698] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[0699] Speech synthesis and emotion regulation
[0700] The translated English text and added sentiment information are passed to a speech synthesis engine (e.g., pyttsx3). Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0701] Audio transmission and playback
[0702] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[0703] Specific example
[0704] The following are some specific situations.
[0705] The user's parent asks in Japanese, "What did you study at school today?"
[0706] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0707] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0708] The server uses a translation engine to translate "What did you study at school today?" into English.
[0709] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[0710] English text with added emotional information is passed to a speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[0711] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[0712] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[0713] As a concrete example of using a generative AI model to perform emotion recognition and reflecting the results in speech synthesis, the following prompt sentence is used:
[0714] Analyze the speaker's emotions based on the question, "What did you study at school today?"
[0715] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0716] Step 1:
[0717] Device setup and connection
[0718] The user wears audio glasses or bone conduction earphones for voice input and output. Next, a device (smartphone or tablet) is paired with these devices via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0719] Input: Audio glasses or bone conduction earphones, device (smartphone or tablet)
[0720] Output: Paired device, launched dedicated application
[0721] Step 2:
[0722] Voice acquisition and transmission
[0723] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0724] Input: Surrounding conversation audio
[0725] Output: Audio data sent to the server (after buffering)
[0726] Step 3:
[0727] Speech recognition and translation
[0728] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[0729] Input: Audio data sent to the server
[0730] Output: Translated English text data
[0731] Step 4:
[0732] Emotion recognition and analysis
[0733] The server uses voice data acquired via voice input and translated text data to analyze the user's speech emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[0734] Input: Translated English text data, acquired audio data
[0735] Output: Text data with emotional information added.
[0736] Step 5:
[0737] Speech synthesis and emotion regulation
[0738] The English text data, which includes emotional information, is passed to a speech synthesis engine (e.g., PyttsX3). Based on the emotional information, the emotion engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0739] Input: English text data with emotional information attached
[0740] Output: Adjusted natural audio data
[0741] Step 6:
[0742] Audio transmission and playback
[0743] The generated audio data is sent from the server to the terminal. The decoded audio data is sent to the user's device and played back in real time. This allows the user to hear surrounding Japanese conversations as natural-sounding English audio with corresponding emotions.
[0744] Input: Adjusted natural voice data
[0745] Output: Audio played on the user's device
[0746] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0747] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0748] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0749] [Third Embodiment]
[0750] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0751] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0752] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0753] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0754] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0755] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0756] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0757] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0758] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0759] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0760] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0761] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0762] This invention is a system for providing an environment in which children can naturally learn a foreign language in their daily lives, and is implemented as follows.
[0763] Device setup and connection
[0764] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. Then, a dedicated application is launched and an internet connection is established.
[0765] Voice acquisition and transmission
[0766] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and sent to a server via the internet. During this process, the audio data is packetized into fixed-capacity chunks to ensure reliable transmission.
[0767] Speech recognition and translation
[0768] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. This text data is then translated into the required foreign language (e.g., English) by a translation engine. The translated text is then checked for grammatical and semantic accuracy using natural language processing (NLP) techniques.
[0769] Speech synthesis and transmission
[0770] The translated text data is converted into English audio data by a speech synthesis engine. This English audio data is compressed and sent to the terminal. To minimize translation delays, the transmitted audio data is processed in real time, packet by packet.
[0771] Audio playback
[0772] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0773] Specific example
[0774] The following are some specific situations.
[0775] The user's parent asks in Japanese, "What did you study at school today?"
[0776] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0777] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0778] The server's translation engine translates "What did you study at school today?" into English.
[0779] The server converts the English text translated by the speech synthesis engine into audio data.
[0780] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0781] This system allows users to listen to everyday conversations in a foreign language in real time, enabling them to efficiently acquire that language.
[0782] The following describes the processing flow.
[0783] Step 1:
[0784] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0785] Step 2:
[0786] The device launches a dedicated app and establishes an internet connection.
[0787] Step 3:
[0788] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0789] Step 4:
[0790] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0791] Step 5:
[0792] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0793] Step 6:
[0794] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0795] Step 7:
[0796] The server passes the translated English text to the speech synthesis engine, which converts it into English speech data.
[0797] Step 8:
[0798] The server compresses the generated English audio data, packets it, and sends it to the terminal.
[0799] Step 9:
[0800] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0801] Step 10:
[0802] The device transmits English audio data to the user's device and plays it back in real time. The user hears the surrounding Japanese conversation as translated English. Efficient data transfer and processing are performed to minimize translation delays.
[0803] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to a server, where it performs translation and speech synthesis before playing the English audio. Through this process, the user can hear the English audio "What did you study at school today?" in real time.
[0804] (Example 1)
[0805] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0806] In modern society, opportunities for children to naturally learn foreign languages in their daily lives are limited. Conventional learning systems and speech recognition systems have difficulty enabling foreign language learning through natural, real-time conversation, and it has been particularly challenging to provide instantly translated foreign languages in a home conversational environment. This invention aims to solve these problems and provide a system that enables children to efficiently learn foreign languages in their daily lives.
[0807] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0808] In this invention, the server includes speech recognition means, translation means, natural language processing means, speech synthesis means, and synthesized speech transmission means. This makes it possible to translate acquired Japanese speech into a foreign language in real time with high accuracy, compress the resulting speech, transmit it to the user's device, and play it back immediately.
[0809] "Voice input means" refers to devices or methods for acquiring ambient sounds or conversations.
[0810] "Voice data transmission means" refers to a device or method for transmitting acquired voice data to other devices or servers.
[0811] "Speech recognition means" refers to a technology or device that analyzes acquired speech data and converts it into corresponding text data.
[0812] "Translation means" refers to a technology or device that converts text data generated by speech recognition means into another language.
[0813] "Speech synthesis means" refers to a technology or device that converts text data generated by translation means into speech data.
[0814] "Synthesized speech transmission means" refers to a technology or device for transmitting generated speech data to another device.
[0815] "Translated audio playback means" refers to technology or equipment for playing back transmitted audio data.
[0816] "Voice input devices" refer to devices used to acquire sound, such as microphones, audio glasses, and bone conduction earphones.
[0817] "Pairing method" refers to a technology or method for connecting an audio input device and a terminal via wireless communication (e.g., Bluetooth or Wi-Fi).
[0818] "Real-time audio data buffering means" refers to a technology or device for temporarily storing acquired audio data and processing it in real time.
[0819] "Method for packetizing voice data" refers to a technology or device that divides voice data to be transmitted into packets of a certain size and transmits them as packets.
[0820] "Natural language processing means" refers to technologies or devices for verifying and correcting the accuracy of the grammar and semantics of translated text or audio data.
[0821] "Audio data compression means" refers to a technology or device that compresses audio data to reduce its size when transmitting it.
[0822] This invention provides a system that offers users an environment in which they can naturally learn a foreign language in their daily lives. Specific embodiments are described below.
[0823] Device setup and connection
[0824] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. This operation launches a dedicated application and establishes an internet connection.
[0825] Hardware and software to use
[0826] Hardware: Audio glasses, bone conduction earphones, smartphones, tablets
[0827] Software: Dedicated applications, speech recognition engine, translation engine, natural language processing (NLP) technology, speech synthesis engine
[0828] Voice acquisition and transmission
[0829] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and transmitted to a server via the internet with high reliability. During this process, the audio data is packetized into fixed-capacity chunks.
[0830] Speech recognition and translation
[0831] The server inputs the received audio data into a speech recognition engine, which converts the Japanese audio into text. This text data is then translated into the required foreign language (e.g., English) by a translation engine. Finally, natural language processing (NLP) techniques are used to verify the accuracy of the grammar and meaning.
[0832] Speech synthesis and transmission
[0833] The translated text data is converted into English audio data by a speech synthesis engine. This data is compressed and sent to the terminal. The audio data is processed in real time on a packet-by-packet basis to minimize transmission delay.
[0834] Audio playback
[0835] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[0836] Specific example
[0837] The following are some specific situations.
[0838] The user's parent asks in Japanese, "What did you study at school today?"
[0839] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0840] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0841] The server's translation engine translates "What did you study at school today?" into English.
[0842] The server converts the English text translated by the speech synthesis engine into audio data.
[0843] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[0844] Example of a prompt
[0845] The following are specific examples of prompt statements for the generative AI model related to this system.
[0846] "Please tell me how to translate what my parents say in Japanese into English in real time and listen to it using audio glasses."
[0847] "Please explain the mechanism of a system that allows children to naturally learn a foreign language through everyday conversation, including specific devices and steps."
[0848] "Please describe in detail your approach to building a learning system that combines speech recognition and translation functions, including the Japanese-to-English translation process."
[0849] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0850] Step 1:
[0851] The user wears audio glasses or bone conduction earphones that handle audio input and output. Next, the terminal pairs with these devices via Bluetooth or Wi-Fi. The user then launches a dedicated application on the terminal and establishes an internet connection. Input is the confirmation of wearing the audio device and connecting to the terminal, while output is the completion of the connection and the launch of the application. Specific actions include the user turning on the audio device, the terminal detecting the device, and selecting pairing.
[0852] Step 2:
[0853] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time, packetized into fixed-capacity chunks, and sent to a server via the internet. The input is ambient sound, and the output is packetized audio data. Specifically, the device's microphone continuously collects ambient sound, temporarily stores it internally, and then divides the data into packets at the appropriate time.
[0854] Step 3:
[0855] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. Next, the text data is translated into the required foreign language (e.g., English) by a translation engine. Furthermore, the accuracy of the grammar and meaning is checked using natural language processing (NLP) techniques. The input is packetized audio data, and the output is translated text data. Specifically, the process involves the speech recognition engine analyzing the audio data to generate Japanese text, the translation engine translating that text into English, and NLP techniques checking the grammar and meaning.
[0856] Step 4:
[0857] The translated text data is converted into audio data by a speech synthesis engine. This English audio data is compressed, repacked, and sent to the terminal. The input is the translated text data, and the output is the packetized audio data. Specifically, the process involves the speech synthesis engine converting the text data into speech, compressing that audio data, dividing it into packets, and sending them to the terminal.
[0858] Step 5:
[0859] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. The input is packetized audio data, and the output is the audio for the user to listen to. Specifically, the operation involves the device receiving packetized audio data, decompressing it, reconstructing it into audio data, and finally playing it back through the audio device.
[0860] Through the steps described above, this system enables users to listen to translated foreign language audio in real time during their daily lives.
[0861] (Application Example 1)
[0862] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0863] In today's globalized world, acquiring multiple languages from an early age is considered crucial. However, providing an effective environment for children to naturally learn foreign languages is not easy. Traditional foreign language education is often limited to specific times and places, lacking opportunities for natural language acquisition in daily life. Furthermore, systems that provide customized lessons tailored to individual learners while managing progress are still not adequately developed. This reduces the efficiency and sustainability of learning, making the language acquisition process difficult.
[0864] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0865] In this invention, the server includes voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, real-time educational translation means, learning progress management means, and dialogue simulation means. This makes it possible to provide an environment in which children can naturally learn a foreign language in their daily lives. Furthermore, by tracking the user's learning progress and providing customized lessons tailored to each individual's progress, the efficiency of learning is improved and sustained learning is promoted.
[0866] A "voice input device" is a device that has the function of acquiring sounds from the user's surroundings.
[0867] A "voice data transmission means" is a device that has the function of transmitting acquired voice data to a server.
[0868] "Speech recognition means" refers to technology that converts transmitted speech data into text data.
[0869] "Translation means" refers to technology that translates text data generated by speech recognition means into another language.
[0870] "Speech synthesis means" refers to a technology that converts text data generated by translation means into speech data.
[0871] "Synthesized speech transmission means" refers to a function for transmitting speech data generated by speech synthesis means to the user's device.
[0872] "Translated audio playback means" refers to a function that plays back audio data received on the user's device.
[0873] "Educational real-time translation technology" is a technology that translates surrounding conversations in real time so that children can naturally learn a foreign language in their daily lives.
[0874] A "learning progress management system" is a technology that tracks a user's learning progress and provides customized lessons according to their progress.
[0875] A "dialogue simulation method" is a technology that provides a simulation of a conversation based on a specific situation.
[0876] The present invention is a system that provides an environment in which children can naturally learn a foreign language in their daily lives. This system includes the following functions: voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, educational real-time translation means, learning progress management means, and dialogue simulation means.
[0877] First, the user wears audio glasses or bone conduction earphones as a means of voice input and uses a smartphone or tablet as a terminal. The terminal is paired with the device via Bluetooth or Wi-Fi and a dedicated application is launched. The terminal captures ambient sound through its microphone, and the acquired audio data is sent to a server via the internet. At this time, the audio data is packetized to ensure reliable transmission. The server inputs the received audio data into a speech recognition engine and converts the Japanese speech into text format.
[0878] The converted text data is translated into the required foreign language (e.g., English) by a translation engine. The translated text is then converted into audio data by a speech synthesis engine, compressed, and sent to the terminal. The terminal decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in a foreign language in real time.
[0879] The real-time educational translation system translates surrounding conversations in real time, providing children with an environment where they can naturally learn a foreign language. This feature makes it easier for them to develop the habit of listening to a foreign language in their daily lives. Furthermore, the learning progress management system tracks the user's learning progress and provides customized lessons according to their individual progress. This feature enables efficient and sustainable learning. The dialogue simulation system provides conversation simulations based on specific situations, supporting practical language learning.
[0880] Hardware and software to be used
[0881] Hardware: Smartphone or tablet, audio glasses or bone conduction earphones, microphone
[0882] Software: Python, speech_recognition library, googletrans library, pyttsx3 library
[0883] Data processing and data calculation
[0884] 1. The voice input device captures voice data through the microphone.
[0885] 2. The audio data transmission means sends the acquired audio data to the server.
[0886] 3. The server's speech recognition system converts the speech data into text data.
[0887] 4. The translation method translates text data into a foreign language using a translation engine.
[0888] 5. The speech synthesis means converts the translated text into speech data.
[0889] 6. The synthesized speech transmission means transmits the voice data to the terminal.
[0890] 7. The translated audio playback device decodes the audio data and plays it back to the user.
[0891] 8. Educational real-time translation tools translate everyday conversations in real time.
[0892] 9. A learning progress management system manages the user's learning status and provides customized lessons.
[0893] 10. The dialogue simulation means provides conversation simulation.
[0894] Specific example
[0895] For example, if the user's parent says in Japanese, "What did you study at school today?", the device picks up the parent's voice with its microphone and sends it to the server in real time. The server uses a speech recognition engine to convert it into Japanese text, "What did you study at school today?", and then a translation engine translates it into English as "What did you study at school today?". The translated English text is converted into audio data by a speech synthesis engine, and this converted audio data is sent to the device, which decodes it and sends it to the user's device for playback.
[0896] Example of a prompt
[0897] "I want to develop a tool that translates everyday conversations into English to help with foreign language learning. The core function is real-time voice translation. Please tell me about the specific functions of this tool and how to implement them."
[0898] As a result, by using the system of the present invention, an environment is created in which children can naturally learn foreign languages in their daily lives, promoting efficient and sustainable learning.
[0899] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0900] Step 1:
[0901] The user wears audio glasses or bone conduction earphones and pairs them with a smartphone or tablet. The input is the pairing information between the terminal and the device, and each establishes a stable connection via Bluetooth or Wi-Fi. The output is the status of successful pairing.
[0902] Step 2:
[0903] The device captures ambient sound through its microphone. The input is ambient sound, which the device's microphone picks up. The output is audio data, which is buffered.
[0904] Step 3:
[0905] Audio data is sent to a server via the internet. The input is buffered audio data, which is reliably transmitted to the server. The output is the audio data stored on the server.
[0906] Step 4:
[0907] The server uses a speech recognition engine to convert audio data into text data. The input is audio data stored on the server, which the speech recognition engine analyzes and converts into text. The output is text data.
[0908] Step 5:
[0909] The server translates text data into a foreign language using a translation engine. The input is text data generated by speech recognition, which the translation engine then translates. The output is the translated foreign language text.
[0910] Step 6:
[0911] The server converts translated text data into speech data using a speech synthesis engine. The input is translated foreign language text, which the speech synthesis engine then converts into speech data. The output is synthesized speech data.
[0912] Step 7:
[0913] The server compresses the synthesized speech data and sends it to the terminal via the internet. The input is synthesized speech data, which is sent to the terminal after compression. The output is the audio data received by the terminal.
[0914] Step 8:
[0915] The terminal decodes the received audio data and plays it back through the user's device. The input is the received audio data, which is decoded and then played back. The output is a foreign language audio heard by the user.
[0916] Step 9:
[0917] The server manages the user's learning progress based on their current status and provides personalized lessons. The input consists of the user's learning data and progress information, which are then analyzed to generate appropriate lessons. The output is a customized lesson plan.
[0918] Step 10:
[0919] The device performs a dialogue simulation based on a specific situation. The input is pre-configured situation information, and the dialogue content experienced by the user is generated through the simulation process. The output is the result of the dialogue simulation with the user.
[0920] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0921] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[0922] Device setup and connection
[0923] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[0924] Voice acquisition and transmission
[0925] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[0926] Speech recognition and translation
[0927] The server receives the audio data and converts it into text format using a speech recognition engine. The converted Japanese text data is then passed to a translation engine and translated into English text data.
[0928] Emotion recognition and analysis
[0929] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. The analysis results are added as supplementary information to the text data.
[0930] Speech synthesis and emotion regulation
[0931] The translated English text and added sentiment information are passed to the speech synthesis engine. Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[0932] Audio transmission and playback
[0933] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[0934] Specific example
[0935] The following are some specific situations.
[0936] The user's parent asks in Japanese, "What did you study at school today?"
[0937] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[0938] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[0939] The server uses a translation engine to translate "What did you study at school today?" into English.
[0940] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[0941] The English text, with added emotional information, is passed to the speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[0942] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[0943] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[0944] This system allows children to hear everyday conversations in a foreign language in real time, while simultaneously providing natural translations that help them understand the speaker's emotions.
[0945] The following describes the processing flow.
[0946] Step 1:
[0947] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[0948] Step 2:
[0949] The device launches a dedicated app and establishes an internet connection.
[0950] Step 3:
[0951] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[0952] Step 4:
[0953] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[0954] Step 5:
[0955] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[0956] Step 6:
[0957] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[0958] Step 7:
[0959] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. Emotional information is added to the text data as supplementary information.
[0960] Step 8:
[0961] The server passes the translated English text and added sentiment information to the speech synthesis engine, which converts it into English speech data. The sentiment engine adjusts the tone and pitch of the speech synthesis based on the sentiment information.
[0962] Step 9:
[0963] The server compresses the generated audio data, packets it, and sends it to the terminal.
[0964] Step 10:
[0965] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[0966] Step 11:
[0967] The device transmits English audio data to the user's device and plays it back in real time. The user can hear the surrounding Japanese conversation as natural English that reflects emotions.
[0968] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to the server. After the server performs translation and speech synthesis, it uses emotion recognition to analyze the user's parent's emotions (for example, "joy of asking"), and sends audio data reflecting the results back to the device. The device then plays the English audio, allowing the user to hear the English phrase "What did you study at school today?" in real time, conveying the emotion behind it.
[0969] (Example 2)
[0970] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0971] Conventional speech translation systems simply translate speech without considering emotional information, resulting in unnatural-sounding translations that fail to convey the speaker's emotions. Furthermore, there has been no system that provides natural-sounding translations that reflect the speaker's emotions, which is crucial for children learning foreign languages in their daily lives. This can lead to low user satisfaction and a decrease in the efficiency of foreign language learning.
[0972] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0973] In this invention, the server includes speech recognition means, translation means, emotion recognition means, emotion information addition means, and speech synthesis means. This makes it possible to consider emotion information when translating speech and generate natural speech that reflects the speaker's emotions.
[0974] "Voice input means" refers to a device or function for acquiring ambient sounds.
[0975] "Voice data transmission means" refers to a means for transmitting acquired voice data to a specific server or terminal.
[0976] "Speech recognition means" refers to a technology or device for converting speech data into text format.
[0977] "Translation means" refers to a technology or device for translating converted text data into another language.
[0978] "Emotion recognition means" refers to a technology or device that analyzes a speaker's emotions from audio data.
[0979] A "means for adding emotional information" refers to a means for adding analyzed emotional information to text data.
[0980] "Speech synthesis means" refers to a technology or device for converting text data into speech data.
[0981] A "synthesized speech transmission means" is a means for transmitting generated speech data to a specific terminal or device.
[0982] A "translation audio playback means" is a function for playing back audio data received by a terminal or device.
[0983] This invention provides a system that allows children to naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system consists of the following elements:
[0984] Device setup and connection
[0985] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones). Next, a terminal (e.g., a smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Subsequently, the terminal launches a dedicated application and establishes an internet connection.
[0986] Specific example:
[0987] The user puts on the bone conduction earphones, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[0988] Voice acquisition and transmission
[0989] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[0990] Specific example:
[0991] When the user's parent asks, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to a server via the internet.
[0992] Speech recognition and translation
[0993] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[0994] Specific example:
[0995] The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[0996] Emotion recognition and analysis
[0997] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[0998] Specific example:
[0999] The server inputs the audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[1000] Speech synthesis and emotion regulation
[1001] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[1002] Specific example:
[1003] The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[1004] Audio transmission and playback
[1005] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[1006] Specific example:
[1007] The translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is then sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[1008] Example of a prompt
[1009] "When a parent asks, 'What did you study at school today?', please translate this into English and play an audio recording that reflects emotional information."
[1010] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1011] Step 1: Device setup and connection
[1012] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones), and the terminal (e.g., smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Next, the terminal launches a dedicated application and establishes an internet connection.
[1013] Specific steps: The user places the bone conduction earphones in their ears, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[1014] Input: The user wears an audio device and pairs it with the terminal.
[1015] Output: The terminal's dedicated application launches and an internet connection is established.
[1016] Step 2: Acquiring and transmitting audio
[1017] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[1018] Specific operation: When the user's parent says, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to the server via the internet.
[1019] Input: Audio recording of the user's parent's conversation.
[1020] Output: Audio data sent to the server.
[1021] Step 3: Speech Recognition and Translation
[1022] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[1023] Specific operation: The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[1024] Input: Acquired audio data.
[1025] Output: Translated English text data.
[1026] Step 4: Emotion Recognition and Analysis
[1027] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[1028] Specific operation: The server inputs audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[1029] Input: Audio data and translated English text data.
[1030] Output: English text data with added sentiment information.
[1031] Step 5: Speech synthesis and emotion regulation
[1032] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[1033] Specific process: The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[1034] Input: English text data with added sentiment information.
[1035] Output: Voice data adjusted based on emotion.
[1036] Step 6: Sending and Playing Audio
[1037] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[1038] Specific operation: Translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[1039] Input: Generated audio data.
[1040] Output: Audio played from the user's device.
[1041] (Application Example 2)
[1042] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1043] Modern children do not have sufficient opportunities to naturally learn foreign languages in their daily lives. Furthermore, there is a lack of systems that provide natural-sounding translations that take emotional states into account in real time. This creates challenges in improving learning efficiency and the depth of foreign language comprehension.
[1044] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1045] In this invention, the server includes a voice input means, an emotion recognition means, and a speech synthesis means. This allows children to listen to everyday conversations in real time as a foreign language while also understanding the emotions of the speakers.
[1046] "Voice input means" refers to hardware and software for acquiring ambient sounds.
[1047] "Voice data transmission means" refers to a means for transmitting acquired voice data to a server.
[1048] A "speech recognition means" is an engine for converting speech data into text format.
[1049] A "translation tool" is an engine used to translate text data into another language.
[1050] An "emotion recognition tool" is an algorithm used to analyze the emotional state of the speaker.
[1051] A "speech synthesis means" is an engine for generating speech from text data.
[1052] A "synthesized voice transmission means" is a means for transmitting generated voice data to a terminal.
[1053] A "translation audio playback means" is a means for playing back audio transmitted to a terminal.
[1054] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[1055] Device setup and connection
[1056] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[1057] Voice acquisition and transmission
[1058] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[1059] Speech recognition and translation
[1060] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[1061] Emotion recognition and analysis
[1062] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[1063] Speech synthesis and emotion regulation
[1064] The translated English text and added sentiment information are passed to a speech synthesis engine (e.g., pyttsx3). Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[1065] Audio transmission and playback
[1066] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[1067] Specific example
[1068] The following are some specific situations.
[1069] The user's parent asks in Japanese, "What did you study at school today?"
[1070] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[1071] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[1072] The server uses a translation engine to translate "What did you study at school today?" into English.
[1073] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[1074] English text with added emotional information is passed to a speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[1075] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[1076] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[1077] As a concrete example of using a generative AI model to perform emotion recognition and reflecting the results in speech synthesis, the following prompt sentence is used:
[1078] Analyze the speaker's emotions based on the question, "What did you study at school today?"
[1079] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1080] Step 1:
[1081] Device setup and connection
[1082] The user wears audio glasses or bone conduction earphones for voice input and output. Next, a device (smartphone or tablet) is paired with these devices via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[1083] Input: Audio glasses or bone conduction earphones, device (smartphone or tablet)
[1084] Output: Paired device, launched dedicated application
[1085] Step 2:
[1086] Voice acquisition and transmission
[1087] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[1088] Input: Surrounding conversation audio
[1089] Output: Audio data sent to the server (after buffering)
[1090] Step 3:
[1091] Speech recognition and translation
[1092] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[1093] Input: Audio data sent to the server
[1094] Output: Translated English text data
[1095] Step 4:
[1096] Emotion recognition and analysis
[1097] The server uses voice data acquired via voice input and translated text data to analyze the user's speech emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[1098] Input: Translated English text data, acquired audio data
[1099] Output: Text data with emotional information added.
[1100] Step 5:
[1101] Speech synthesis and emotion regulation
[1102] The English text data, which includes emotional information, is passed to a speech synthesis engine (e.g., PyttsX3). Based on the emotional information, the emotion engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[1103] Input: English text data with emotional information attached
[1104] Output: Adjusted natural audio data
[1105] Step 6:
[1106] Audio transmission and playback
[1107] The generated audio data is sent from the server to the terminal. The decoded audio data is sent to the user's device and played back in real time. This allows the user to hear surrounding Japanese conversations as natural-sounding English audio with corresponding emotions.
[1108] Input: Adjusted natural voice data
[1109] Output: Audio played on the user's device
[1110] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1111] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1112] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1113] [Fourth Embodiment]
[1114] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1115] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1116] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1117] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1118] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1119] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1120] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1121] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1122] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1123] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1124] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1125] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1126] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1127] This invention is a system for providing an environment in which children can naturally learn a foreign language in their daily lives, and is implemented as follows.
[1128] Device setup and connection
[1129] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. Then, a dedicated application is launched and an internet connection is established.
[1130] Voice acquisition and transmission
[1131] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and sent to a server via the internet. During this process, the audio data is packetized into fixed-capacity chunks to ensure reliable transmission.
[1132] Speech recognition and translation
[1133] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. This text data is then translated into the required foreign language (e.g., English) by a translation engine. The translated text is then checked for grammatical and semantic accuracy using natural language processing (NLP) techniques.
[1134] Speech synthesis and transmission
[1135] The translated text data is converted into English audio data by a speech synthesis engine. This English audio data is compressed and sent to the terminal. To minimize translation delays, the transmitted audio data is processed in real time, packet by packet.
[1136] Audio playback
[1137] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[1138] Specific example
[1139] The following are some specific situations.
[1140] The user's parent asks in Japanese, "What did you study at school today?"
[1141] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[1142] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[1143] The server's translation engine translates "What did you study at school today?" into English.
[1144] The server converts the English text translated by the speech synthesis engine into audio data.
[1145] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[1146] This system allows users to listen to everyday conversations in a foreign language in real time, enabling them to efficiently acquire that language.
[1147] The following describes the processing flow.
[1148] Step 1:
[1149] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[1150] Step 2:
[1151] The device launches a dedicated app and establishes an internet connection.
[1152] Step 3:
[1153] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[1154] Step 4:
[1155] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[1156] Step 5:
[1157] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[1158] Step 6:
[1159] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[1160] Step 7:
[1161] The server passes the translated English text to the speech synthesis engine, which converts it into English speech data.
[1162] Step 8:
[1163] The server compresses the generated English audio data, packets it, and sends it to the terminal.
[1164] Step 9:
[1165] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[1166] Step 10:
[1167] The device transmits English audio data to the user's device and plays it back in real time. The user hears the surrounding Japanese conversation as translated English. Efficient data transfer and processing are performed to minimize translation delays.
[1168] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to a server, where it performs translation and speech synthesis before playing the English audio. Through this process, the user can hear the English audio "What did you study at school today?" in real time.
[1169] (Example 1)
[1170] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1171] In modern society, opportunities for children to naturally learn foreign languages in their daily lives are limited. Conventional learning systems and speech recognition systems have difficulty enabling foreign language learning through natural, real-time conversation, and it has been particularly challenging to provide instantly translated foreign languages in a home conversational environment. This invention aims to solve these problems and provide a system that enables children to efficiently learn foreign languages in their daily lives.
[1172] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1173] In this invention, the server includes speech recognition means, translation means, natural language processing means, speech synthesis means, and synthesized speech transmission means. This makes it possible to translate acquired Japanese speech into a foreign language in real time with high accuracy, compress the resulting speech, transmit it to the user's device, and play it back immediately.
[1174] "Voice input means" refers to devices or methods for acquiring ambient sounds or conversations.
[1175] "Voice data transmission means" refers to a device or method for transmitting acquired voice data to other devices or servers.
[1176] "Speech recognition means" refers to a technology or device that analyzes acquired speech data and converts it into corresponding text data.
[1177] "Translation means" refers to a technology or device that converts text data generated by speech recognition means into another language.
[1178] "Speech synthesis means" refers to a technology or device that converts text data generated by translation means into speech data.
[1179] "Synthesized speech transmission means" refers to a technology or device for transmitting generated speech data to another device.
[1180] "Translated audio playback means" refers to technology or equipment for playing back transmitted audio data.
[1181] "Voice input devices" refer to devices used to acquire sound, such as microphones, audio glasses, and bone conduction earphones.
[1182] "Pairing method" refers to a technology or method for connecting an audio input device and a terminal via wireless communication (e.g., Bluetooth or Wi-Fi).
[1183] "Real-time audio data buffering means" refers to a technology or device for temporarily storing acquired audio data and processing it in real time.
[1184] "Method for packetizing voice data" refers to a technology or device that divides voice data to be transmitted into packets of a certain size and transmits them as packets.
[1185] "Natural language processing means" refers to technologies or devices for verifying and correcting the accuracy of the grammar and semantics of translated text or audio data.
[1186] "Audio data compression means" refers to a technology or device that compresses audio data to reduce its size when transmitting it.
[1187] This invention provides a system that offers users an environment in which they can naturally learn a foreign language in their daily lives. Specific embodiments are described below.
[1188] Device setup and connection
[1189] The user wears audio glasses or bone conduction earphones that handle voice input and output. Next, a device (e.g., a smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. This operation launches a dedicated application and establishes an internet connection.
[1190] Hardware and software to use
[1191] Hardware: Audio glasses, bone conduction earphones, smartphones, tablets
[1192] Software: Dedicated applications, speech recognition engine, translation engine, natural language processing (NLP) technology, speech synthesis engine
[1193] Voice acquisition and transmission
[1194] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time and transmitted to a server via the internet with high reliability. During this process, the audio data is packetized into fixed-capacity chunks.
[1195] Speech recognition and translation
[1196] The server inputs the received audio data into a speech recognition engine, which converts the Japanese audio into text. This text data is then translated into the required foreign language (e.g., English) by a translation engine. Finally, natural language processing (NLP) techniques are used to verify the accuracy of the grammar and meaning.
[1197] Speech synthesis and transmission
[1198] The translated text data is converted into English audio data by a speech synthesis engine. This data is compressed and sent to the terminal. The audio data is processed in real time on a packet-by-packet basis to minimize transmission delay.
[1199] Audio playback
[1200] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in real time as a foreign language.
[1201] Specific example
[1202] The following are some specific situations.
[1203] The user's parent asks in Japanese, "What did you study at school today?"
[1204] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[1205] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[1206] The server's translation engine translates "What did you study at school today?" into English.
[1207] The server converts the English text translated by the speech synthesis engine into audio data.
[1208] The converted audio data is sent to the terminal, which decodes it and sends it to the user's device for playback.
[1209] Example of a prompt
[1210] The following are specific examples of prompt statements for the generative AI model related to this system.
[1211] "Please tell me how to translate what my parents say in Japanese into English in real time and listen to it using audio glasses."
[1212] "Please explain the mechanism of a system that allows children to naturally learn a foreign language through everyday conversation, including specific devices and steps."
[1213] "Please describe in detail your approach to building a learning system that combines speech recognition and translation functions, including the Japanese-to-English translation process."
[1214] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1215] Step 1:
[1216] The user wears audio glasses or bone conduction earphones that handle audio input and output. Next, the terminal pairs with these devices via Bluetooth or Wi-Fi. The user then launches a dedicated application on the terminal and establishes an internet connection. Input is the confirmation of wearing the audio device and connecting to the terminal, while output is the completion of the connection and the launch of the application. Specific actions include the user turning on the audio device, the terminal detecting the device, and selecting pairing.
[1217] Step 2:
[1218] The device captures ambient sound through its microphone. The acquired audio data is buffered in real time, packetized into fixed-capacity chunks, and sent to a server via the internet. The input is ambient sound, and the output is packetized audio data. Specifically, the device's microphone continuously collects ambient sound, temporarily stores it internally, and then divides the data into packets at the appropriate time.
[1219] Step 3:
[1220] The server inputs the received audio data into a speech recognition engine, converting the Japanese speech into text format. Next, the text data is translated into the required foreign language (e.g., English) by a translation engine. Furthermore, the accuracy of the grammar and meaning is checked using natural language processing (NLP) techniques. The input is packetized audio data, and the output is translated text data. Specifically, the process involves the speech recognition engine analyzing the audio data to generate Japanese text, the translation engine translating that text into English, and NLP techniques checking the grammar and meaning.
[1221] Step 4:
[1222] The translated text data is converted into audio data by a speech synthesis engine. This English audio data is compressed, repacked, and sent to the terminal. The input is the translated text data, and the output is the packetized audio data. Specifically, the process involves the speech synthesis engine converting the text data into speech, compressing that audio data, dividing it into packets, and sending them to the terminal.
[1223] Step 5:
[1224] The device decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. The input is packetized audio data, and the output is the audio for the user to listen to. Specifically, the operation involves the device receiving packetized audio data, decompressing it, reconstructing it into audio data, and finally playing it back through the audio device.
[1225] Through the steps described above, this system enables users to listen to translated foreign language audio in real time during their daily lives.
[1226] (Application Example 1)
[1227] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1228] In today's globalized world, acquiring multiple languages from an early age is considered crucial. However, providing an effective environment for children to naturally learn foreign languages is not easy. Traditional foreign language education is often limited to specific times and places, lacking opportunities for natural language acquisition in daily life. Furthermore, systems that provide customized lessons tailored to individual learners while managing progress are still not adequately developed. This reduces the efficiency and sustainability of learning, making the language acquisition process difficult.
[1229] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1230] In this invention, the server includes voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, real-time educational translation means, learning progress management means, and dialogue simulation means. This makes it possible to provide an environment in which children can naturally learn a foreign language in their daily lives. Furthermore, by tracking the user's learning progress and providing customized lessons tailored to each individual's progress, the efficiency of learning is improved and sustained learning is promoted.
[1231] A "voice input device" is a device that has the function of acquiring sounds from the user's surroundings.
[1232] A "voice data transmission means" is a device that has the function of transmitting acquired voice data to a server.
[1233] "Speech recognition means" refers to technology that converts transmitted speech data into text data.
[1234] "Translation means" refers to technology that translates text data generated by speech recognition means into another language.
[1235] "Speech synthesis means" refers to a technology that converts text data generated by translation means into speech data.
[1236] "Synthesized speech transmission means" refers to a function for transmitting speech data generated by speech synthesis means to the user's device.
[1237] "Translated audio playback means" refers to a function that plays back audio data received on the user's device.
[1238] "Educational real-time translation technology" is a technology that translates surrounding conversations in real time so that children can naturally learn a foreign language in their daily lives.
[1239] A "learning progress management system" is a technology that tracks a user's learning progress and provides customized lessons according to their progress.
[1240] A "dialogue simulation method" is a technology that provides a simulation of a conversation based on a specific situation.
[1241] The present invention is a system that provides an environment in which children can naturally learn a foreign language in their daily lives. This system includes the following functions: voice input means, voice data transmission means, voice recognition means, translation means, voice synthesis means, synthesized voice transmission means, translated voice playback means, educational real-time translation means, learning progress management means, and dialogue simulation means.
[1242] First, the user wears audio glasses or bone conduction earphones as a means of voice input and uses a smartphone or tablet as a terminal. The terminal is paired with the device via Bluetooth or Wi-Fi and a dedicated application is launched. The terminal captures ambient sound through its microphone, and the acquired audio data is sent to a server via the internet. At this time, the audio data is packetized to ensure reliable transmission. The server inputs the received audio data into a speech recognition engine and converts the Japanese speech into text format.
[1243] The converted text data is translated into the required foreign language (e.g., English) by a translation engine. The translated text is then converted into audio data by a speech synthesis engine, compressed, and sent to the terminal. The terminal decodes the received audio data and plays it back to the user through audio glasses or bone conduction earphones. This allows the user to hear surrounding Japanese conversations in a foreign language in real time.
[1244] The real-time educational translation system translates surrounding conversations in real time, providing children with an environment where they can naturally learn a foreign language. This feature makes it easier for them to develop the habit of listening to a foreign language in their daily lives. Furthermore, the learning progress management system tracks the user's learning progress and provides customized lessons according to their individual progress. This feature enables efficient and sustainable learning. The dialogue simulation system provides conversation simulations based on specific situations, supporting practical language learning.
[1245] Hardware and software to be used
[1246] Hardware: Smartphone or tablet, audio glasses or bone conduction earphones, microphone
[1247] Software: Python, speech_recognition library, googletrans library, pyttsx3 library
[1248] Data processing and data calculation
[1249] 1. The voice input device captures voice data through the microphone.
[1250] 2. The audio data transmission means sends the acquired audio data to the server.
[1251] 3. The server's speech recognition system converts the speech data into text data.
[1252] 4. The translation method translates text data into a foreign language using a translation engine.
[1253] 5. The speech synthesis means converts the translated text into speech data.
[1254] 6. The synthesized speech transmission means transmits the voice data to the terminal.
[1255] 7. The translated audio playback device decodes the audio data and plays it back to the user.
[1256] 8. Educational real-time translation tools translate everyday conversations in real time.
[1257] 9. A learning progress management system manages the user's learning status and provides customized lessons.
[1258] 10. The dialogue simulation means provides conversation simulation.
[1259] Specific example
[1260] For example, if the user's parent says in Japanese, "What did you study at school today?", the device picks up the parent's voice with its microphone and sends it to the server in real time. The server uses a speech recognition engine to convert it into Japanese text, "What did you study at school today?", and then a translation engine translates it into English as "What did you study at school today?". The translated English text is converted into audio data by a speech synthesis engine, and this converted audio data is sent to the device, which decodes it and sends it to the user's device for playback.
[1261] Example of a prompt
[1262] "I want to develop a tool that translates everyday conversations into English to help with foreign language learning. The core function is real-time voice translation. Please tell me about the specific functions of this tool and how to implement them."
[1263] As a result, by using the system of the present invention, an environment is created in which children can naturally learn foreign languages in their daily lives, promoting efficient and sustainable learning.
[1264] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1265] Step 1:
[1266] The user wears audio glasses or bone conduction earphones and pairs them with a smartphone or tablet. The input is the pairing information between the terminal and the device, and each establishes a stable connection via Bluetooth or Wi-Fi. The output is the status of successful pairing.
[1267] Step 2:
[1268] The device captures ambient sound through its microphone. The input is ambient sound, which the device's microphone picks up. The output is audio data, which is buffered.
[1269] Step 3:
[1270] Audio data is sent to a server via the internet. The input is buffered audio data, which is reliably transmitted to the server. The output is the audio data stored on the server.
[1271] Step 4:
[1272] The server uses a speech recognition engine to convert audio data into text data. The input is audio data stored on the server, which the speech recognition engine analyzes and converts into text. The output is text data.
[1273] Step 5:
[1274] The server translates text data into a foreign language using a translation engine. The input is text data generated by speech recognition, which the translation engine then translates. The output is the translated foreign language text.
[1275] Step 6:
[1276] The server converts translated text data into speech data using a speech synthesis engine. The input is translated foreign language text, which the speech synthesis engine then converts into speech data. The output is synthesized speech data.
[1277] Step 7:
[1278] The server compresses the synthesized speech data and sends it to the terminal via the internet. The input is synthesized speech data, which is sent to the terminal after compression. The output is the audio data received by the terminal.
[1279] Step 8:
[1280] The terminal decodes the received audio data and plays it back through the user's device. The input is the received audio data, which is decoded and then played back. The output is a foreign language audio heard by the user.
[1281] Step 9:
[1282] The server manages the user's learning progress based on their current status and provides personalized lessons. The input consists of the user's learning data and progress information, which are then analyzed to generate appropriate lessons. The output is a customized lesson plan.
[1283] Step 10:
[1284] The device performs a dialogue simulation based on a specific situation. The input is pre-configured situation information, and the dialogue content experienced by the user is generated through the simulation process. The output is the result of the dialogue simulation with the user.
[1285] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1286] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[1287] Device setup and connection
[1288] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[1289] Voice acquisition and transmission
[1290] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[1291] Speech recognition and translation
[1292] The server receives the audio data and converts it into text format using a speech recognition engine. The converted Japanese text data is then passed to a translation engine and translated into English text data.
[1293] Emotion recognition and analysis
[1294] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. The analysis results are added as supplementary information to the text data.
[1295] Speech synthesis and emotion regulation
[1296] The translated English text and added sentiment information are passed to the speech synthesis engine. Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[1297] Audio transmission and playback
[1298] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[1299] Specific example
[1300] The following are some specific situations.
[1301] The user's parent asks in Japanese, "What did you study at school today?"
[1302] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[1303] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[1304] The server uses a translation engine to translate "What did you study at school today?" into English.
[1305] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[1306] The English text, with added emotional information, is passed to the speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[1307] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[1308] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[1309] This system allows children to hear everyday conversations in a foreign language in real time, while simultaneously providing natural translations that help them understand the speaker's emotions.
[1310] The following describes the processing flow.
[1311] Step 1:
[1312] The user puts on audio glasses or bone conduction earphones, and then the device (smartphone or tablet) pairs with the device via Bluetooth or Wi-Fi.
[1313] Step 2:
[1314] The device launches a dedicated app and establishes an internet connection.
[1315] Step 3:
[1316] The device acquires surrounding conversational audio in real time via its microphone. Noise reduction is also applied to improve the quality of the audio data.
[1317] Step 4:
[1318] The terminal buffers the audio data, divides it into packets of a fixed size, and sends them to the server.
[1319] Step 5:
[1320] The server inputs the received audio data into the speech recognition engine, which then converts it into Japanese text data.
[1321] Step 6:
[1322] The server inputs Japanese text data into a translation engine, which then translates it into English text. Dictionaries and contextual analysis are used to improve translation accuracy.
[1323] Step 7:
[1324] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition. Emotional information is added to the text data as supplementary information.
[1325] Step 8:
[1326] The server passes the translated English text and added sentiment information to the speech synthesis engine, which converts it into English speech data. The sentiment engine adjusts the tone and pitch of the speech synthesis based on the sentiment information.
[1327] Step 9:
[1328] The server compresses the generated audio data, packets it, and sends it to the terminal.
[1329] Step 10:
[1330] The device decodes the received English audio data and prepares to transmit it to audio glasses or bone conduction earphones.
[1331] Step 11:
[1332] The device transmits English audio data to the user's device and plays it back in real time. The user can hear the surrounding Japanese conversation as natural English that reflects emotions.
[1333] As a concrete example, if the user's parent says in Japanese, "What did you study at school today?", the device captures this and sends it to the server. After the server performs translation and speech synthesis, it uses emotion recognition to analyze the user's parent's emotions (for example, "joy of asking"), and sends audio data reflecting the results back to the device. The device then plays the English audio, allowing the user to hear the English phrase "What did you study at school today?" in real time, conveying the emotion behind it.
[1334] (Example 2)
[1335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1336] Conventional speech translation systems simply translate speech without considering emotional information, resulting in unnatural-sounding translations that fail to convey the speaker's emotions. Furthermore, there has been no system that provides natural-sounding translations that reflect the speaker's emotions, which is crucial for children learning foreign languages in their daily lives. This can lead to low user satisfaction and a decrease in the efficiency of foreign language learning.
[1337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1338] In this invention, the server includes speech recognition means, translation means, emotion recognition means, emotion information addition means, and speech synthesis means. This makes it possible to consider emotion information when translating speech and generate natural speech that reflects the speaker's emotions.
[1339] "Voice input means" refers to a device or function for acquiring ambient sounds.
[1340] "Voice data transmission means" refers to a means for transmitting acquired voice data to a specific server or terminal.
[1341] "Speech recognition means" refers to a technology or device for converting speech data into text format.
[1342] "Translation means" refers to a technology or device for translating converted text data into another language.
[1343] "Emotion recognition means" refers to a technology or device that analyzes a speaker's emotions from audio data.
[1344] A "means for adding emotional information" refers to a means for adding analyzed emotional information to text data.
[1345] "Speech synthesis means" refers to a technology or device for converting text data into speech data.
[1346] A "synthesized speech transmission means" is a means for transmitting generated speech data to a specific terminal or device.
[1347] A "translation audio playback means" is a function for playing back audio data received by a terminal or device.
[1348] This invention provides a system that allows children to naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system consists of the following elements:
[1349] Device setup and connection
[1350] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones). Next, a terminal (e.g., a smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Subsequently, the terminal launches a dedicated application and establishes an internet connection.
[1351] Specific example:
[1352] The user puts on the bone conduction earphones, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[1353] Voice acquisition and transmission
[1354] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[1355] Specific example:
[1356] When the user's parent asks, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to a server via the internet.
[1357] Speech recognition and translation
[1358] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[1359] Specific example:
[1360] The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[1361] Emotion recognition and analysis
[1362] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[1363] Specific example:
[1364] The server inputs the audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[1365] Speech synthesis and emotion regulation
[1366] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[1367] Specific example:
[1368] The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[1369] Audio transmission and playback
[1370] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[1371] Specific example:
[1372] The translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is then sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[1373] Example of a prompt
[1374] "When a parent asks, 'What did you study at school today?', please translate this into English and play an audio recording that reflects emotional information."
[1375] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1376] Step 1: Device setup and connection
[1377] The user wears an audio device that provides voice input and output (e.g., audio glasses or bone conduction earphones), and the terminal (e.g., smartphone or tablet) pairs with the audio device via Bluetooth or Wi-Fi. Next, the terminal launches a dedicated application and establishes an internet connection.
[1378] Specific steps: The user places the bone conduction earphones in their ears, opens the Bluetooth settings on their smartphone, and pairs the earphones. Then they launch the dedicated app and connect to the internet. The app displays a connection confirmation message.
[1379] Input: The user wears an audio device and pairs it with the terminal.
[1380] Output: The terminal's dedicated application launches and an internet connection is established.
[1381] Step 2: Acquiring and transmitting audio
[1382] The device uses its built-in microphone to capture surrounding conversational audio in real time. The captured audio data is buffered and sent to a server via the internet.
[1383] Specific operation: When the user's parent says, "What did you study at school today?", the device's microphone picks up the audio. The audio data is buffered and sent to the server via the internet.
[1384] Input: Audio recording of the user's parent's conversation.
[1385] Output: Audio data sent to the server.
[1386] Step 3: Speech Recognition and Translation
[1387] The server receives the audio data and converts it into text format using a speech recognition engine (e.g., speech recognition engine). The converted Japanese text data is then passed to a translation engine (e.g., translation engine) to be translated into English text data.
[1388] Specific operation: The server receives the audio data "What did you study at school today?" and inputs it into the speech recognition engine. Within a few seconds, it is converted into text "What did you study at school today?" and passed to the translation engine, which converts it into the English text "What did you study at school today?".
[1389] Input: Acquired audio data.
[1390] Output: Translated English text data.
[1391] Step 4: Emotion Recognition and Analysis
[1392] The server uses speech input and translated text data to analyze the speaker's emotions using emotion recognition means (e.g., emotion recognition engine). The analysis results are added to the text data as supplementary information.
[1393] Specific operation: The server inputs audio data and the text "What did you study at school today?" into the emotion recognition engine. As a result, the emotion "joy of asking" is detected, and this emotion information is added to the English text.
[1394] Input: Audio data and translated English text data.
[1395] Output: English text data with added sentiment information.
[1396] Step 5: Speech synthesis and emotion regulation
[1397] The translated English text and sentiment information are passed to a speech synthesis engine (e.g., a speech synthesis engine). The sentiment engine adjusts the tone and pitch of the speech based on the sentiment information to generate natural-sounding speech data.
[1398] Specific process: The English text "What did you study at school today?" and emotional information are passed to the speech synthesis engine. Tone and pitch adjustments are made based on the emotion, and natural-sounding speech data is generated.
[1399] Input: English text data with added sentiment information.
[1400] Output: Voice data adjusted based on emotion.
[1401] Step 6: Sending and Playing Audio
[1402] The server sends the generated audio data to the terminal. The terminal decodes the audio data and sends it to the user's device for real-time playback.
[1403] Specific operation: Translated audio data is sent from the server to the terminal and immediately decoded. The decoded audio is sent to bone conduction earphones, allowing the user to hear the message "What did you study at school today?" in real time.
[1404] Input: Generated audio data.
[1405] Output: Audio played from the user's device.
[1406] (Application Example 2)
[1407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1408] Modern children do not have sufficient opportunities to naturally learn foreign languages in their daily lives. Furthermore, there is a lack of systems that provide natural-sounding translations that take emotional states into account in real time. This creates challenges in improving learning efficiency and the depth of foreign language comprehension.
[1409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1410] In this invention, the server includes a voice input means, an emotion recognition means, and a speech synthesis means. This allows children to listen to everyday conversations in real time as a foreign language while also understanding the emotions of the speakers.
[1411] "Voice input means" refers to hardware and software for acquiring ambient sounds.
[1412] "Voice data transmission means" refers to a means for transmitting acquired voice data to a server.
[1413] A "speech recognition means" is an engine for converting speech data into text format.
[1414] A "translation tool" is an engine used to translate text data into another language.
[1415] An "emotion recognition tool" is an algorithm used to analyze the emotional state of the speaker.
[1416] A "speech synthesis means" is an engine for generating speech from text data.
[1417] A "synthesized voice transmission means" is a means for transmitting generated voice data to a terminal.
[1418] A "translation audio playback means" is a means for playing back audio transmitted to a terminal.
[1419] This invention aims to provide an environment in which children can naturally learn a foreign language in their daily lives while recognizing the user's emotional state and adjusting the system's operation accordingly. This system is implemented as follows:
[1420] Device setup and connection
[1421] The user wears audio glasses or bone conduction earphones for voice input and output. Next, the device (smartphone or tablet) is paired with the device via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[1422] Voice acquisition and transmission
[1423] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[1424] Speech recognition and translation
[1425] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[1426] Emotion recognition and analysis
[1427] The server uses voice data acquired via voice input and translated text data to analyze the user's emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[1428] Speech synthesis and emotion regulation
[1429] The translated English text and added sentiment information are passed to a speech synthesis engine (e.g., pyttsx3). Based on the sentiment information, the sentiment engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[1430] Audio transmission and playback
[1431] The generated audio data is sent from the server to the terminal, decoded, and then sent to the user's device for real-time playback. This allows the user to hear surrounding Japanese conversations as natural-sounding English that reflects emotions.
[1432] Specific example
[1433] The following are some specific situations.
[1434] The user's parent asks in Japanese, "What did you study at school today?"
[1435] The device uses a microphone to pick up the parent's voice and transmits it to the server in real time.
[1436] The server uses a speech recognition engine to convert the sentence into Japanese text: "What did you study at school today?"
[1437] The server uses a translation engine to translate "What did you study at school today?" into English.
[1438] The server uses emotion recognition to analyze the user's parent's emotion as "the joy of asking a question."
[1439] English text with added emotional information is passed to a speech synthesis engine, which then adjusts the tone and pitch of the speech based on that information.
[1440] The adjusted audio data is sent to the terminal, which then sends it to the user's device for playback.
[1441] Users can hear in real time English audio that reflects the parent's emotions, such as "What did you study at school today?"
[1442] As a concrete example of using a generative AI model to perform emotion recognition and reflecting the results in speech synthesis, the following prompt sentence is used:
[1443] Analyze the speaker's emotions based on the question, "What did you study at school today?"
[1444] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1445] Step 1:
[1446] Device setup and connection
[1447] The user wears audio glasses or bone conduction earphones for voice input and output. Next, a device (smartphone or tablet) is paired with these devices via Bluetooth or Wi-Fi. The device launches a dedicated application and establishes an internet connection.
[1448] Input: Audio glasses or bone conduction earphones, device (smartphone or tablet)
[1449] Output: Paired device, launched dedicated application
[1450] Step 2:
[1451] Voice acquisition and transmission
[1452] The device acquires surrounding conversational audio in real time through its microphone. The acquired audio data is buffered and sent to a server via the internet.
[1453] Input: Surrounding conversation audio
[1454] Output: Audio data sent to the server (after buffering)
[1455] Step 3:
[1456] Speech recognition and translation
[1457] The server converts the received audio data into text format using a speech recognition engine (e.g., speech_recognition). The converted Japanese text data is then passed to a translation engine (e.g., googletrans) and translated into English text data.
[1458] Input: Audio data sent to the server
[1459] Output: Translated English text data
[1460] Step 4:
[1461] Emotion recognition and analysis
[1462] The server uses voice data acquired via voice input and translated text data to analyze the user's speech emotions using emotion recognition equipment (e.g., EmotionRecognizer). The analysis results are added as supplementary information to the text data.
[1463] Input: Translated English text data, acquired audio data
[1464] Output: Text data with emotional information added.
[1465] Step 5:
[1466] Speech synthesis and emotion regulation
[1467] The English text data, which includes emotional information, is passed to a speech synthesis engine (e.g., PyttsX3). Based on the emotional information, the emotion engine adjusts the tone and pitch of the speech synthesis engine to generate natural-sounding speech data that corresponds to the user's emotions.
[1468] Input: English text data with emotional information attached
[1469] Output: Adjusted natural audio data
[1470] Step 6:
[1471] Audio transmission and playback
[1472] The generated audio data is sent from the server to the terminal. The decoded audio data is sent to the user's device and played back in real time. This allows the user to hear surrounding Japanese conversations as natural-sounding English audio with corresponding emotions.
[1473] Input: Adjusted natural voice data
[1474] Output: Audio played on the user's device
[1475] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1476] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1477] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1478] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1479] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1480] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1481] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1482] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1483] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1484] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1485] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1486] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1487] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1488] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1489] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1490] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1491] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1492] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1493] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1494] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1495] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1496] The following is further disclosed regarding the embodiments described above.
[1497] (Claim 1)
[1498] Voice input method,
[1499] A means for transmitting voice data,
[1500] Voice recognition means and
[1501] Translation methods and,
[1502] A speech synthesis method,
[1503] A means for transmitting synthesized speech,
[1504] Translation audio playback means,
[1505] A system that includes this.
[1506] (Claim 2)
[1507] The system according to claim 1, wherein the voice input means acquires ambient sound through a microphone.
[1508] (Claim 3)
[1509] The system according to claim 1, which converts Japanese audio data acquired by a translation means into English using a translation engine.
[1510] "Example 1"
[1511] (Claim 1)
[1512] Voice input method,
[1513] A means for transmitting voice data,
[1514] Voice recognition means and
[1515] Translation methods and,
[1516] A speech synthesis method,
[1517] A means for transmitting synthesized speech,
[1518] Translation audio playback means,
[1519] A method for pairing an audio input device with a terminal,
[1520] Real-time audio data buffering means,
[1521] A method for packetizing voice data,
[1522] Natural language processing means,
[1523] Audio data compression means,
[1524] A system that includes this.
[1525] (Claim 2)
[1526] The system according to claim 1, wherein the voice input means acquires ambient sound through a microphone.
[1527] (Claim 3)
[1528] The system according to claim 1, wherein a translation means converts Japanese audio data acquired by the translation means into English using a translation engine, and a natural language processing means verifies the accuracy of grammar and semantics.
[1529] "Application Example 1"
[1530] (Claim 1)
[1531] Voice input method,
[1532] A means for transmitting voice data,
[1533] Voice recognition means and
[1534] Translation methods and,
[1535] A speech synthesis method,
[1536] A means for transmitting synthesized speech,
[1537] Translation audio playback means,
[1538] Educational real-time translation tools,
[1539] Learning progress management methods,
[1540] Dialogue simulation means,
[1541] A system that includes this.
[1542] (Claim 2)
[1543] The system according to claim 1, wherein the voice input means acquires ambient sound through a microphone.
[1544] (Claim 3)
[1545] The system according to claim 1, which converts Japanese audio data acquired by a translation means into English using a translation engine.
[1546] (Claim 4)
[1547] The system according to claim 1, wherein the educational real-time translation means translates surrounding Japanese conversations into a foreign language in real time.
[1548] (Claim 5)
[1549] The system according to claim 1, wherein the learning progress management means tracks the user's learning progress and provides customized lessons.
[1550] (Claim 6)
[1551] The system according to claim 1, wherein the dialogue simulation means provides a conversation simulation based on a specific situation.
[1552] "Example 2 of combining an emotion engine"
[1553] (Claim 1)
[1554] Voice input method,
[1555] A means for transmitting voice data,
[1556] Voice recognition means and
[1557] Translation methods and,
[1558] Means of recognizing emotions,
[1559] Means for adding emotional information,
[1560] A speech synthesis method,
[1561] A means for transmitting synthesized speech,
[1562] Translation audio playback means,
[1563] A system that includes this.
[1564] (Claim 2)
[1565] The system according to claim 1, wherein the voice input means acquires ambient sound through a microphone.
[1566] (Claim 3)
[1567] The system according to claim 1, which converts Japanese audio data acquired by a translation means into English using a translation engine.
[1568] "Application example 2 when combining with an emotional engine"
[1569] (Claim 1)
[1570] Voice input method,
[1571] A means for transmitting voice data,
[1572] Voice recognition means and
[1573] Translation methods and,
[1574] Means of recognizing emotions,
[1575] A speech synthesis method,
[1576] A means for transmitting synthesized speech,
[1577] Translation audio playback means,
[1578] A system that includes this.
[1579] (Claim 2)
[1580] The system according to claim 1, wherein a voice input means acquires ambient sound through a microphone, and an emotion recognition means analyzes the emotions of the speaker.
[1581] (Claim 3)
[1582] The system according to claim 1, wherein a translation means converts Japanese speech data acquired by a translation means into English using a translation engine, an emotion recognition means analyzes emotions based on the converted text, and a speech synthesis means generates natural speech using the emotion information. [Explanation of Symbols]
[1583] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A voice input means for acquiring ambient sounds, A voice data transmission means for sending acquired voice data to a remote server, A speech recognition means that converts acquired audio data into text format, A translation means that translates text obtained by a speech recognition means into another language, A speech synthesis means that converts text generated by a translation means into speech, A synthesized speech transmission means that transmits the generated audio data to a user device, A translated audio playback means for playing back transmitted audio data, A system that includes this.
2. The system according to claim 1, wherein the voice input means acquires ambient sound through a microphone.
3. The system according to claim 1, wherein Japanese audio data acquired by the translation means is converted into English using a translation engine.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A