system
A system for real-time multilingual interpretation using speech recognition, translation, and synthesis modules addresses the language barrier by enabling efficient and accurate communication across languages, reducing costs and improving operational efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
The language barrier poses a significant obstacle in situations requiring rapid and accurate communication, such as business and emergency scenarios, due to the limited availability and high cost of translators, leading to potential misunderstandings and lost opportunities.
A system comprising speech recognition, language translation, speech synthesis, and network communication modules that enable real-time conversion of audio to text, translation between languages, and output of translated audio, supported by a server for efficient interpretation.
Facilitates smooth, rapid, and cost-effective communication across multiple languages, reducing the need for interpreters and enhancing operational efficiency.
Smart Images

Figure 2026062222000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to the description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern global society, communication between different languages is important, but on the other hand, the language barrier is a major obstacle. Especially in situations where rapid and accurate communication is required, such as in business scenes, travel, and emergencies, multilingual translation is essential. However, the number of translators is limited, and ensuring them and the cost are problems. Also, inappropriate understanding or misunderstanding can lead to loss of business opportunities and cause troubles. To solve such problems, an advanced solution for instant translation between 32 major languages is required.
Means for Solving the Problems
[0005] The present invention provides a system comprising means for inputting audio, means for converting the input audio into text data, and means for translating the text data from one language to another. Furthermore, by including means for converting the translated text data into audio data and outputting the audio data, it enables real-time interpretation between multiple languages. In addition, by providing network communication means for transmitting and receiving audio data, and a server that receives audio and text data and performs translation processing based on them, it provides a real-time and efficient interpretation service. This system can solve the problems of securing interpreters and costs, and support smooth communication between different languages.
[0006] "Means of inputting voice" refers to a function that captures the user's speech as voice data using a device such as a microphone.
[0007] "Means of converting to text data" refers to a function that uses speech recognition technology to analyze audio data and convert it into text data as a string of characters.
[0008] "Translation means" refers to a function that uses language translation algorithms and technologies to convert text data expressed in the original language into another specified language.
[0009] "Means of converting to audio data" refers to a function that uses speech synthesis technology to convert translated text data back into audio data.
[0010] "Means for outputting audio data" refers to a function for playing back synthesized audio data to the user through an output device such as a speaker.
[0011] "Network communication means" refers to the function of using communication infrastructure and protocols to send and receive voice data and text data between a terminal and a server.
[0012] A "server" is a computing system that receives audio and text data, performs translation and speech synthesis as needed, and sends the processing results to a terminal. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention provides a system for instantaneous multilingual interpretation. This system consists of a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0035] Speech recognition module
[0036] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0037] Language translation module
[0038] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. The text data is sent from the terminal to the server, which uses the language translation module to perform the translation. For example, the text "こんにちは" (konnichiwa) is translated into the English "Hello".
[0039] Speech synthesis module
[0040] The speech synthesis module converts the translated text data back into speech data. The server receives the translated text and uses the speech synthesis module to generate speech data. For example, the translated text "Hello" is generated as speech data.
[0041] Network communication module
[0042] The network communication module enables the transmission and reception of text and voice data. Text data generated by the speech recognition module is sent to the server via the network, and voice data, after translation and speech synthesis, is sent back to the terminal.
[0043] Specific example
[0044] For example, in a meeting setting, it works as follows:
[0045] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?"
[0046] 2. Terminal A captures the audio and sends it to the server as text: "What do you think about this project?"
[0047] 3. The server's language translation module translates the text into English: "What do you think about this project?".
[0048] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" and sends it to terminal B.
[0049] 5. Terminal B plays the audio data, and User B (an English speaker) is able to understand the question.
[0050] This system enables smooth communication between multiple languages, significantly reducing the time and cost of interpretation. It also leads to increased operational efficiency and prevention of problems, and is expected to have diverse applications.
[0051] The following describes the processing flow.
[0052] Step 1:
[0053] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0054] Step 2:
[0055] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0056] Step 3:
[0057] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0058] Step 4:
[0059] The terminal sends the converted text data to the server via the network communication module. The transmitted data also includes input language information.
[0060] Step 5:
[0061] The server analyzes the received text data and input language information. For example, it analyzes the received text "Hello" and its language information "Japanese".
[0062] Step 6:
[0063] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0064] Step 7:
[0065] The server passes the translated text data to the speech synthesis module, which converts it from text to speech data. For example, the text "Hello" is generated as speech data.
[0066] Step 8:
[0067] The server sends the generated audio data to the terminal via the network communication module.
[0068] Step 9:
[0069] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0070] Step 10:
[0071] The user listens to the played audio and understands the translated content. In this way, rapid and accurate communication between different languages is achieved.
[0072] (Example 1)
[0073] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0074] In real-time multilingual communication, conventional systems have problems with the rapid and accurate recognition, translation, and output of speech data. Furthermore, the processing of speech and text data is distributed, which can lead to network communication delays and errors, thus compromising the user experience.
[0075] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0076] In this invention, the server includes means for converting speech to text data, means for transmitting and receiving the text data and speech data via network communication means, means for translating the text data from one language to another, means for converting the translated text data to speech data, and means for outputting the received speech data. This enables real-time, rapid, and accurate communication between multiple languages.
[0077] A "means of inputting voice" refers to a device that captures the voice spoken by a user as a digital signal.
[0078] "Means for converting audio to text data" refers to devices or programs for converting captured audio signals into text data in string format.
[0079] "Means of translating text data from one language to another" refers to software or algorithms that automatically convert input text data into another specified language.
[0080] "Means for converting translated text data into audio data" refers to software or hardware for generating translated text data as an audio signal.
[0081] "Network communication means" refers to devices and technologies for transmitting and receiving data (text data and audio data) via the internet or local networks.
[0082] A "server" is a computer system that receives data via a network, processes it, and transmits the results.
[0083] "Means for outputting audio data" refers to a device or program for playing back the generated audio data from an output device such as a speaker.
[0084] A "device" is hardware or electronic equipment used to perform a specific purpose or function.
[0085] This invention is a system for real-time multilingual interpretation. This system consists of the following main modules: a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0086] Speech recognition module
[0087] The voice spoken by the user into the microphone is converted into text data by the speech recognition module installed in the device. Specifically, speech recognition software such as Google® Speech-to-Text API is used.
[0088] Language translation module
[0089] The text data generated by the speech recognition module is sent from the terminal to the server via the network communication module. The server is equipped with a language translation module and uses translation software such as the DeepL API or Google Translate API to translate the text data into the specified language.
[0090] Speech synthesis module
[0091] The translated text data is then converted into speech data by a speech synthesis module on the server. Specifically, speech synthesis software such as the Google Text-to-Speech API or Amazon Polly is used.
[0092] Network communication module
[0093] The generated audio data is then transmitted back to the terminal via the network communication module. The terminal receives this audio data and plays the sound through an output device such as a speaker.
[0094] Specific example
[0095] For example, in a meeting setting, it works as follows:
[0096] 1. User A (a Japanese speaker) asks into the microphone, "What do you think about this project?"
[0097] 2. Device A captures the audio and converts it into text data, "What do you think about this project?", using the Google Speech-to-Text API.
[0098] 3. Terminal A sends the generated text data to the server via the network communication module.
[0099] 4. The server receives the text data and uses the DeepL API to translate it into "What do you think about this project?".
[0100] 5. The server uses the Google Text-to-Speech API to convert "What do you think about this project?" into audio data.
[0101] 6. The server transmits the voice data to terminal B via the network communication module.
[0102] 7. Terminal B plays back the received audio data, allowing User B (an English speaker) to understand the question.
[0103] Example of a prompt
[0104] Here are examples of prompts when using a generative AI model in this system:
[0105] "This system uses a speech recognition module, a language translation module, a speech synthesis module, and a network communication module to perform real-time multilingual interpretation. When a user speaks into the microphone, their voice is converted into text data and translated into the specified language. The translated text data is then converted back into speech and played back. Please provide specific use scenarios."
[0106] This system enables smooth communication between multiple languages and will be extremely useful in meetings and international business operations.
[0107] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0108] Step 1:
[0109] The user inputs voice into the microphone. This input is the user's voice, for example, by saying "Hello."
[0110] Step 2:
[0111] The device uses a speech recognition module to capture the user's voice and convert it into text data. Specifically, the Google Speech-to-Text API converts the voice signal "hello" into the text data "hello". In this step, the input is the user's voice and the output is text data.
[0112] Step 3:
[0113] The terminal sends the generated text data to the server via the network communication module. In this step, the text data is sent using the HTTP protocol or WebSocket. The input is the text data "Hello", and the output is the text data sent to the server.
[0114] Step 4:
[0115] The server receives text data over the network. The received text data is "Hello". The input is text data sent from the terminal, which becomes the data for the next translation process.
[0116] Step 5:
[0117] The server uses a language translation module to translate the received text data into another language. Here, we use the DeepL API and the Google Translate API to translate the Japanese "こんにちは" (konnichiwa) into the English "Hello". The input is the Japanese text data "こんにちは", and the output is the English text data "Hello".
[0118] Step 6:
[0119] The server uses a text-to-speech module to convert translated text data into speech data. Specifically, it uses the Google Text-to-Speech API or Amazon Polly to convert the English text "Hello" into speech data. The input is the English text data "Hello," and the output is speech data.
[0120] Step 7:
[0121] The server sends the generated audio data to the terminal via a network communication module. In this step, the HTTP protocol or WebSocket is used to send the audio data. The input is the audio data, and the output is the audio data sent to the terminal.
[0122] Step 8:
[0123] The device receives audio data over the network and plays it back using its speakers or connected headset. Specifically, the device's audio device plays the audio data "Hello". The input is audio data sent from the server, and the output is the audio that the user can hear.
[0124] (Application Example 1)
[0125] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0126] In situations where smooth communication between multiple languages is difficult, especially in tourist areas, tourists often cannot understand the local language and have difficulty obtaining necessary information. Furthermore, insufficient translation of local guides and information signs makes efficient tourist guidance challenging.
[0127] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0128] In this invention, the server includes means for inputting voice, means for converting the voice into text data, means for translating the text data from one language to another, means for converting the translated text data into voice data, means for outputting the voice data, and means for being a tourist guide application including an assistive device. This enables tourists to instantly understand local information in their native language.
[0129] "Means of inputting voice" refers to a function that allows the system to capture voices emitted by the user using an input device such as a microphone.
[0130] "Means for converting the audio into text data" refers to a function for analyzing the input audio and converting it into data in the corresponding text format.
[0131] "Means for translating the text data from a specific language to another language" refers to a function for converting text data from its original language to another specified language.
[0132] "Means for converting the translated text data into audio data" refers to a function for analyzing the translated text data and converting it into data in the corresponding audio format.
[0133] "Means for outputting the audio data" refers to a function for playing back the converted audio data through an output device such as a speaker.
[0134] "Accessible devices" are electronic devices that users can carry with them and that provide information in an assistive way, such as smartphones, tablets, and smart glasses.
[0135] A "tourist guide application" is a software program that provides information in multiple languages to users in tourist destinations.
[0136] This invention is a system that enables instantaneous interpretation between multiple languages, and is specifically implemented as an application for tourist guides. The system and its operation will be described in detail below.
[0137] Hardware configuration
[0138] The server includes the following hardware:
[0139] High-performance processor
[0140] Large capacity memory
[0141] High-speed network connection interface
[0142] User terminals include the following devices:
[0143] smartphone
[0144] Speakers and microphones
[0145] display
[0146] Network communication function (Wi-Fi, 4G / 5G)
[0147] Software Configuration
[0148] The following software modules will be installed on the server:
[0149] Speech recognition module (e.g., Google Cloud Speech-to-Text API)
[0150] Language translation module (e.g., Google Translate API)
[0151] Text-to-speech modules (e.g., Google Text-to-Speech API)
[0152] The applications installed on the user's terminal include the following features:
[0153] Voice input and capture
[0154] Sending and receiving text data
[0155] Playback of translated audio data
[0156] Data processing and calculations
[0157] The server receives the audio data sent from the user terminal and processes it in the following steps:
[0158] 1. Speech Recognition: The speech recognition module is used to convert speech data into text data. For example, the speech "Tell me about this temple" is converted to the text "Tell me about this temple".
[0159] 2. Language Translation: Use the language translation module to translate the converted text data into the specified language. For example, "Please tell me about this temple" will be translated into English as "Please tell me about this temple."
[0160] 3. Speech Synthesis: The speech synthesis module is used to convert translated text data into speech data. For example, the text "Please tell me about this temple." is converted into corresponding English speech data.
[0161] 4. Data transmission: The generated audio data is sent back to the user's terminal, and the application plays it.
[0162] Specific example
[0163] For example, consider a situation where tourists are visiting a temple in Kyoto.
[0164] 1. A tourist speaks into a smartphone application and says, "Tell me about this temple."
[0165] 2. The application captures this audio and uses a speech recognition module to convert it into text: "Tell me about this temple."
[0166] 3. The text data is sent to the server via the network and translated into "Please tell me about this temple." using a language translation module.
[0167] 4. The translated text is converted into speech data using a speech synthesis module.
[0168] 5. Audio data is sent from the server to the user's terminal, and the English audio "Please tell me about this temple." is played for the tourist.
[0169] Example of a prompt:
[0170] The user speaks "Tell me about this temple" in Japanese into their smartphone. The application captures the audio, converts it to text using the Google Translate API, and translates it into English. Then, it converts it back into English speech using the Google Text-to-Speech library and plays it through the smartphone's speaker. Tourists can then hear the information in English.
[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0172] Step 1:
[0173] The user launches a smartphone application and asks a question or requests guidance in a specific language (e.g., Japanese) into the microphone. The input is voice data, and the output is captured voice data. At this stage, the application uses the smartphone's microphone to capture the voice and stores it in its internal memory.
[0174] Step 2:
[0175] The terminal application sends the captured audio data to the speech recognition module. The input is the audio data obtained in step 1, and the output is the corresponding text data. Specifically, the speech recognition module performs speech analysis, generates the text data "Tell me about this temple," and saves it to its internal memory.
[0176] Step 3:
[0177] The terminal sends the generated text data to the server. The input is the text data obtained from the speech recognition module, and the output is the text data received by the server. The terminal uses its network communication function to transfer the text data to the server.
[0178] Step 4:
[0179] The server's language translation module translates the received text data into the specified language. The input is the sent text data ("Please tell me about this temple"), and the output is the translated text data ("Please tell me about this temple."). The server uses the Google Translate API to convert the text data to English and stores the translation result in internal memory.
[0180] Step 5:
[0181] The server's text-to-speech module converts translated text data into speech data. The input is translated text data, and the output is the corresponding speech data ("Please tell me about this temple."). The server uses the Google Text-to-Speech API to convert the text data into speech data and stores that speech data in internal memory.
[0182] Step 6:
[0183] The server sends the generated audio data to the terminal. The input is the audio data generated by the server, and the output is the audio data received by the terminal. The server transmits the audio data to the terminal via the network.
[0184] Step 7:
[0185] The device plays the received audio data through its speaker. The input is the audio data received from the server, and the output is the audio the user hears. The device plays the audio data and provides the user with the English response, "Please tell me about this temple."
[0186] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0187] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module.
[0188] Speech recognition module
[0189] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0190] Emotional Engine
[0191] The emotion engine analyzes the user's emotions from captured audio data and extracts emotional information. This emotional information represents emotional states such as anger, joy, and sadness, and is reflected in the translation results. For example, if the audio says "hello," the emotion engine will recognize that the user is happy.
[0192] Language translation module
[0193] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. Furthermore, it takes into account emotional information recognized by the emotion engine when performing the translation. For example, the Japanese word "konnichiwa" is translated into English as "Hello" along with the emotional information "joy".
[0194] Speech synthesis module
[0195] The speech synthesis module generates speech data with appropriate tone and intonation based on translated text data and emotional information. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0196] Network communication module
[0197] The network communication module enables the transmission and reception of text and voice data. The text data and sentiment information generated by the speech recognition module are sent to a server via the network, where translation and speech synthesis are performed, and then the voice data is sent back to the terminal.
[0198] Specific example
[0199] For example, in a meeting setting, it works as follows:
[0200] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?", and their voice contains an expression of anger.
[0201] 2. Terminal A captures the audio and sends the text "What do you think about this project?" and emotion information "Anger" to the server.
[0202] 3. The server's language translation module translates the text into English, "What do you think about this project?", and then adjusts the translation result to take into account the emotional information "anger".
[0203] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" in an angry tone and sends it to terminal B.
[0204] 5. Terminal B plays the audio data, and user B (an English speaker) understands the question and recognizes the emotional state of the user.
[0205] This system enables rapid and accurate communication across multiple languages, and can even convey emotional nuances, resulting in more human-like communication. It is useful not only in business settings but also in personal interactions and emergency situations.
[0206] The following describes the processing flow.
[0207] Step 1:
[0208] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0209] Step 2:
[0210] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0211] Step 3:
[0212] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0213] Step 4:
[0214] The device passes voice data to the emotion engine, which analyzes the user's emotions from the voice. For example, it might recognize "hello" as an emotion of joy.
[0215] Step 5:
[0216] The terminal transmits the converted text data and sentiment information to the server via the network communication module. The transmitted data also includes input language information.
[0217] Step 6:
[0218] The server analyzes the received text data, sentiment information, and input language information. For example, it analyzes the received text "Hello," the sentiment information "Joy," and the language information "Japanese."
[0219] Step 7:
[0220] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0221] Step 8:
[0222] The server adjusts the translation results by taking emotional information into account. For example, "Hello" will be output with an emotion that reflects happiness.
[0223] Step 9:
[0224] The server passes the translated text data and emotional information to the speech synthesis module, which then converts the text into speech data. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0225] Step 10:
[0226] The server sends the generated audio data to the terminal via the network communication module.
[0227] Step 11:
[0228] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0229] Step 12:
[0230] The user listens to the played audio and understands the translated content and its emotions. In this way, rapid, accurate, and emotionally responsive communication between different languages is achieved.
[0231] (Example 2)
[0232] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0233] Multilingual speech translation technology faces the challenge of simultaneously conveying both accuracy and emotional nuances. Conventional systems simply translate text, making it difficult to generate speech with appropriate tone and intonation that reflects the user's emotions. Furthermore, communication with remote users requires high-speed and reliable data communication methods. To solve these challenges, a system integrating speech recognition, sentiment analysis, translation, speech synthesis, and network communication is necessary.
[0234] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0235] In this invention, the server includes means for receiving voice data, text data, and emotional information, and performing translation and speech synthesis processing based thereon; network communication means for transmitting and receiving the voice data, text data, and emotional information; and means for extracting emotional information from text data. This makes it possible to convey not only rapid and accurate communication between multiple languages, but also emotional nuances.
[0236] "Means for inputting voice" refers to a device or method for a user to provide voice data to a system using an input device such as a microphone.
[0237] "Means for converting speech to text data" refers to software or hardware modules for converting captured speech data into text data. For example, speech recognition technology can be used.
[0238] "Means for extracting emotional information from text data" refers to software or hardware modules that analyze the user's emotions from converted text data and its audio data, and extract that emotional information.
[0239] A "means for translating text data and sentiment information from one language to another" refers to a software or hardware module that performs translation work based on text data and sentiment information. In this process, sentiment information is also reflected in the translation.
[0240] "Means for converting translated text data and emotional information into audio data" refers to software or hardware modules for generating audio data with appropriate tone and intonation based on translated text data and emotional information.
[0241] "Means for outputting audio data" refers to output devices such as speakers or headphones that provide the generated audio data to the user.
[0242] "Network communication means for transmitting and receiving voice data, text data, and emotional information" refers to communication infrastructure such as the internet or a local network for sending and receiving voice data, text data, and emotional information between modules within a system.
[0243] "Means for receiving audio data, text data, and emotional information, and performing translation and speech synthesis processing based thereon" refers to engines or software modules that perform translation and speech synthesis processing based on the various types of data received.
[0244] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module. The functions of each module, and how they process and calculate data, will be described in detail below.
[0245] Speech recognition module
[0246] The device uses its microphone to capture the user's voice and sends it to a speech recognition module. This speech recognition module then uses speech recognition technology to convert the voice data into text data. For example, if the user says "hello," the device captures this voice, and the speech recognition module converts it into the text data "hello."
[0247] Emotional Engine
[0248] The server receives text data generated by the speech recognition module and sends it to the emotion engine. This emotion engine extracts user emotion information from the text and speech data, for example, using emotion recognition technology. For example, it recognizes the emotion of "joy" in response to "hello."
[0249] Language translation module
[0250] The server sends text data containing emotional information to a language translation module, which then uses multilingual translation technology to translate it into another specified language. Furthermore, the emotional information is also reflected in the translation. For example, the Japanese "こんにちは" (konnichiwa) is translated into English as "Hello," and the emotion of "joy" is added.
[0251] Speech synthesis module
[0252] The server sends translated text data and emotional information to a speech synthesis module, which then generates speech data considering appropriate tone and intonation. For example, speech synthesis technology is used to generate speech data for "Hello" with a joyful tone.
[0253] Network communication module
[0254] The server and terminal use a network communication module to send and receive voice data, text data, and sentiment information. For example, data is sent over the internet or a local network, and appropriate processing is performed after it is received.
[0255] Specific example
[0256] For example, the actions taken at an international conference are as follows:
[0257] 1. User A (a Japanese speaker) speaks into the microphone and asks, "What do you think about this project?" At this time, User A is feeling angry.
[0258] 2. Terminal A uses its microphone to capture user A's voice.
[0259] 3. Terminal A's speech recognition module converts the speech data into text data and generates the text, "What do you think about this project?"
[0260] 4. Terminal A sends the converted text data to the emotion engine, which analyzes the emotion of anger and assigns the emotion information for "anger".
[0261] 5. Terminal A sends the generated text data and emotion information "anger" to the server via the network communication module.
[0262] 6. The server sends the received text data to the language translation module, which translates "What do you think about this project?" into English: "What do you think about this project?".
[0263] 7. The server reflects the emotion information "anger" in the translated text data.
[0264] 8. The server sends the translated text data and the emotion information "anger" to the speech synthesis module, which generates "What do you think about this project?" in an angry tone.
[0265] 9. The server transmits the generated audio data to terminal B via the network communication module.
[0266] 10. Play back the audio data received by terminal B, "What do you think about this project?", in an angry tone.
[0267] 11. User B (an English speaker) listens to the audio and understands the question and the emotion of anger.
[0268] Example of a prompt
[0269] For example, by inputting the following prompt sentence into the generative AI model, it is possible to generate translations and sentiment-reflecting speech through the steps described above:
[0270] "Translate the following Japanese into English and generate an audio recording that reflects the emotion. Japanese: 'What do you think about this project?' Emotion: 'Anger'"
[0271] This enables smooth multilingual communication between users and the system, and also allows for the transmission of emotional nuances. It is extremely useful in business settings, personal interactions, and emergency situations to achieve more human-like communication.
[0272] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0273] Step 1:
[0274] The user speaks into the microphone. For example, they might say "Hello." The input is audio data. The output is audio data captured through the microphone.
[0275] Step 2:
[0276] The device uses its microphone to capture the user's voice and sends that audio data to a speech recognition module. The input is the voice data from the user, and the output is the audio data sent to the speech recognition module. Specifically, the device's microphone captures the user's speech and saves it as an audio file.
[0277] Step 3:
[0278] The device's speech recognition module converts captured audio data into text data. For example, it converts the audio "Hello" into the text data "Hello". The input is the captured audio data, and the output is the converted text data. Specifically, the speech recognition software analyzes the audio waveform and generates the string data.
[0279] Step 4:
[0280] The device sends the converted text data to the emotion engine, which analyzes the text and the original audio data to extract the user's emotional information. For example, it might recognize the emotion of "joy" in response to "hello." The input consists of text data and audio data, and the output is the emotional information associated with the text data. Specifically, the emotion engine analyzes the text data and the tone of the voice and assigns an emotional label.
[0281] Step 5:
[0282] The terminal sends generated text data and sentiment information to the server via a network communication module. The input is text data and sentiment information, and the output is the data sent to the server. Specifically, the terminal's network module creates data packets and sends them to the server via the internet.
[0283] Step 6:
[0284] The server sends the received text data to the language translation module for translation into another specified language. For example, it translates "こんにちは" to "Hello". The input is text data, and the output is the translated text data. Specifically, the translation software on the server processes the text data and converts it into text in the corresponding other language.
[0285] Step 7:
[0286] The server reflects emotional information in the translated text data. For example, it imparts the emotion of "joy" to "Hello". The input is the translated text data and emotional information, and the output is the translated text with emotional information imparted. Specifically, an emotion label is applied to the translation result to generate the final translated text.
[0287] Step 8:
[0288] The server sends the translated text data and emotional information to the speech synthesis module to convert the text into speech data. For example, it generates "Hello" as speech data in a "joyful" tone. The input is the translated text data and emotional information, and the output is the synthesized speech data. Specifically, the speech synthesis software generates a speech waveform based on the emotion label and text data.
[0289] Step 9:
[0290] The server sends the generated speech data to terminal B through the network communication module. The input is the generated speech data, and the output is the data transmitted to terminal B. Specifically, the network module on the server converts the synthesized speech into data packets and transmits them to terminal B through the Internet. <00009Terminal B plays the audio data it receives. The input is the received audio data, and the output is user B listening to the audio. Specifically, terminal B's speaker plays the received audio data, and user B listens to the audio.
[0293] (Application Example 2)
[0294] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0295] This invention aims to provide a system for instantaneous multilingual interpretation that enables real-time speech recognition, translation, and display while reflecting emotions. Modern content delivery services require users to understand live streaming in various languages with rich emotional depth. However, current technology struggles to appropriately reflect emotional information in multilingual translation, potentially degrading the quality of communication. There is a need for technology that can solve this problem and enable more natural and human-like communication.
[0296] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice, means for converting voice into text data, means for translating text data from one language to another, means for adjusting the translated text data based on emotional information and converting it into voice data, means for outputting voice data via network communication means, and means for displaying it in real time as text data and voice data that reflect emotions. This enables users to communicate in real time, including emotions, between multiple languages.
[0297] "Means of inputting audio" refers to devices and software for capturing audio data.
[0298] "Means for converting the audio into text data" refers to speech recognition technology and software for converting captured audio data into text data.
[0299] "Means for translating the text data from a specific language to another language" refers to translation technologies and software for converting text data expressed in a specific language into another language.
[0300] "Means for adjusting the translated text data based on emotional information and converting it into audio data" refers to technology and software for generating audio data that reflects emotional information based on translated text data.
[0301] "Means for outputting the audio data via network communication means" refers to the technology and software for transmitting the generated audio data to other devices or servers via a network.
[0302] "Means of displaying text and audio data that reflect emotions in real time" refers to technologies and software for instantly displaying translation results that include emotional information visually and audibly.
[0303] "Network communication means" refers to communication infrastructure and protocols such as the internet and local networks used to send and receive data.
[0304] A "server" refers to a computer system used to process data and communicate with other devices.
[0305] This invention aims to translate and display streaming content in emotionally rich language using a system that integrates real-time multilingual translation and emotion recognition. The specific program and its processing are described below.
[0306] First, the terminal (smartphone, smart glasses, head-mounted display) used by the user captures the live streaming audio. This audio data is converted into text data by the speech recognition module. The speech recognition module uses Google's speech recognition API (e.g., the speech_recognition library in Python).
[0307] The text data generated by speech recognition is sent to the emotion engine to analyze the user's emotion. The Emotion Recognition library is used for this emotion engine. The emotion information represents emotional states such as anger, joy, sadness, etc., and this is reflected in the translation result.
[0308] Next, the text data is passed to the language translation module and translated into the specified other language. Google Translate API (e.g., the googletrans library in Python) is used for translation. Since the translation is performed taking into account the emotion information, for example, "こんにちは" is translated together with the emotion information "joy".
[0309] The translated text data and emotion information are converted into audio data with an appropriate tone and intonation by the text-to-speech module. Google Text-to-Speech (e.g., the gtts library in Python) is used for text-to-speech to generate audio data adjusted based on the emotion information.
[0310] The generated audio data is sent in real time to the user's terminal or other devices via the network communication module. As a result, the user can immediately visually and auditorily confirm the translation result reflecting the emotion.
[0311] A specific example is shown below.
[0312] User A speaks in Japanese, "What do you think about this project?", and the voice contains an emotion of anger. Terminal A captures the voice, and a speech recognition module converts it into text data, "What do you think about this project?". An emotion engine analyzes this and extracts the emotion information "anger". Then, a language translation module translates the text data into English, "What do you think about this project?", resulting in a translation that reflects the emotion information "anger". Finally, a speech synthesis module generates voice data with an angry tone based on this data and sends it to Terminal B via a network communication module. Terminal B plays the voice data, and User B understands the question and recognizes their emotional state.
[0313] Example of a prompt
[0314] "What do you think of this project? (with anger)"
[0315] This invention enables more natural and human-like real-time communication between multiple languages.
[0316] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0317] Step 1:
[0318] The user speaks into the live stream, and the device captures that audio. The input is the user's voice data, and the output is the captured audio data. Specifically, the audio is recorded using microphones on smartphones, smart glasses, or head-mounted displays.
[0319] Step 2:
[0320] The device passes the captured audio data to a speech recognition module, which converts it into text data. The input is the captured audio data, and the output is the converted text data. Specifically, Google's speech recognition API (Python's speech_recognition library) is used to convert the audio data into detailed text data.
[0321] Step 3:
[0322] The server receives text data from the speech recognition module and sends it to the emotion engine to extract emotional information. The input is text data, and the output is emotional information. Specifically, the Emotion Recognition library is used to perform emotional analysis on the text data. In this process, emotional states such as anger, joy, and sadness are evaluated.
[0323] Step 4:
[0324] The server passes text data and sentiment information to a language translation module for translation into another language. The input is text data and sentiment information, and the output is translated text data. Specifically, the Google Translate API (the googletrans library in Python) is used to perform accurate translations between multiple languages and to reflect sentiment information.
[0325] Step 5:
[0326] The server passes translated text data and sentiment information to a speech synthesis module, which then generates speech data with appropriate tone and intonation. The input is translated text data and sentiment information, and the output is the generated speech data. Specifically, Google Text-to-Speech (the gtts library in Python) is used to create speech data with a specific tone based on the sentiment information.
[0327] Step 6:
[0328] The server transmits the generated audio data to the user's terminal via network communication. The input is the generated audio data, and the output is the transmitted audio data. Specifically, data is transferred using the internet or a local network.
[0329] Step 7:
[0330] The system plays back audio data received by the user's device and displays a translation result that reflects emotions in real time. The input is the received audio data, and the output is the played audio data and the displayed text data. Specifically, it uses the speakers and displays of smartphones, smart glasses, and head-mounted displays to provide the user with emotionally rich audio and text.
[0331] Example of a prompt
[0332] "What do you think of this project? (with anger)"
[0333] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0334] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0335] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0336] [Second Embodiment]
[0337] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0338] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0339] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0340] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0341] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0342] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0343] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0344] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0345] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0346] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0347] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0348] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0349] This invention provides a system for instantaneous multilingual interpretation. This system consists of a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0350] Speech recognition module
[0351] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0352] Language translation module
[0353] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. The text data is sent from the terminal to the server, which uses the language translation module to perform the translation. For example, the text "こんにちは" (konnichiwa) is translated into English as "Hello".
[0354] Speech synthesis module
[0355] The speech synthesis module converts the translated text data back into speech data. The server receives the translated text and uses the speech synthesis module to generate speech data. For example, the translated text "Hello" is generated as speech data.
[0356] Network communication module
[0357] The network communication module enables the transmission and reception of text and voice data. Text data generated by the speech recognition module is sent to the server via the network, and voice data, after translation and speech synthesis, is sent back to the terminal.
[0358] Specific example
[0359] For example, in a meeting setting, it works as follows:
[0360] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?"
[0361] 2. Terminal A captures the audio and sends it to the server as text: "What do you think about this project?"
[0362] 3. The server's language translation module translates the text into English: "What do you think about this project?".
[0363] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" and sends it to terminal B.
[0364] 5. Terminal B plays the audio data, and User B (an English speaker) is able to understand the question.
[0365] This system enables smooth communication between multiple languages, significantly reducing the time and cost of interpretation. It also leads to increased operational efficiency and prevention of problems, and is expected to have diverse applications.
[0366] The following describes the processing flow.
[0367] Step 1:
[0368] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0369] Step 2:
[0370] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0371] Step 3:
[0372] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0373] Step 4:
[0374] The terminal sends the converted text data to the server via the network communication module. The transmitted data also includes input language information.
[0375] Step 5:
[0376] The server analyzes the received text data and input language information. For example, it analyzes the received text "Hello" and its language information "Japanese".
[0377] Step 6:
[0378] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0379] Step 7:
[0380] The server passes the translated text data to the speech synthesis module, which converts it from text to speech data. For example, the text "Hello" is generated as speech data.
[0381] Step 8:
[0382] The server sends the generated audio data to the terminal via the network communication module.
[0383] Step 9:
[0384] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0385] Step 10:
[0386] The user listens to the played audio and understands the translated content. In this way, rapid and accurate communication between different languages is achieved.
[0387] (Example 1)
[0388] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0389] In real-time multilingual communication, conventional systems have problems with the rapid and accurate recognition, translation, and output of speech data. Furthermore, the processing of speech and text data is distributed, which can lead to network communication delays and errors, thus compromising the user experience.
[0390] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0391] In this invention, the server includes means for converting speech to text data, means for transmitting and receiving the text data and speech data via network communication means, means for translating the text data from one language to another, means for converting the translated text data to speech data, and means for outputting the received speech data. This enables real-time, rapid, and accurate communication between multiple languages.
[0392] A "means of inputting voice" refers to a device that captures the voice spoken by a user as a digital signal.
[0393] "Means for converting audio to text data" refers to devices or programs that convert captured audio signals into text data in string format.
[0394] "Means of translating text data from one language to another" refers to software or algorithms that automatically convert input text data into another specified language.
[0395] "Means for converting translated text data into audio data" refers to software or hardware for generating translated text data as an audio signal.
[0396] "Network communication means" refers to devices and technologies for transmitting and receiving data (text data and audio data) via the internet or local networks.
[0397] A "server" is a computer system that receives data via a network, processes it, and transmits the results.
[0398] "Means for outputting audio data" refers to a device or program for playing back the generated audio data from an output device such as a speaker.
[0399] A "device" is hardware or electronic equipment used to perform a specific purpose or function.
[0400] This invention is a system for real-time multilingual interpretation. This system consists of the following main modules: a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0401] Speech recognition module
[0402] The voice spoken by the user into the microphone is converted into text data by the speech recognition module installed in the device. Specifically, speech recognition software such as the Google Speech-to-Text API is used.
[0403] Language translation module
[0404] The text data generated by the speech recognition module is sent from the terminal to the server via the network communication module. The server is equipped with a language translation module and uses translation software such as the DeepL API or Google Translate API to translate the text data into the specified language.
[0405] Speech synthesis module
[0406] The translated text data is then converted into speech data by a speech synthesis module on the server. Specifically, speech synthesis software such as the Google Text-to-Speech API or Amazon Polly is used.
[0407] Network communication module
[0408] The generated audio data is then transmitted back to the terminal via the network communication module. The terminal receives this audio data and plays the sound through an output device such as a speaker.
[0409] Specific example
[0410] For example, in a meeting setting, it works as follows:
[0411] 1. User A (a Japanese speaker) asks into the microphone, "What do you think about this project?"
[0412] 2. Device A captures the audio and converts it into text data, "What do you think about this project?", using the Google Speech-to-Text API.
[0413] 3. Terminal A sends the generated text data to the server via the network communication module.
[0414] 4. The server receives the text data and uses the DeepL API to translate it into "What do you think about this project?".
[0415] 5. The server uses the Google Text-to-Speech API to convert "What do you think about this project?" into audio data.
[0416] 6. The server transmits the voice data to terminal B via the network communication module.
[0417] 7. Terminal B plays back the received audio data, allowing User B (an English speaker) to understand the question.
[0418] Example of a prompt
[0419] Here are examples of prompts when using a generative AI model in this system:
[0420] "This system uses a speech recognition module, a language translation module, a speech synthesis module, and a network communication module to perform real-time multilingual interpretation. When a user speaks into the microphone, their voice is converted into text data and translated into the specified language. The translated text data is then converted back into speech and played back. Please provide specific use scenarios."
[0421] This system enables smooth communication between multiple languages and will be extremely useful in meetings and international business operations.
[0422] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0423] Step 1:
[0424] The user inputs voice into the microphone. This input is the user's voice, for example, by saying "Hello."
[0425] Step 2:
[0426] The device uses a speech recognition module to capture the user's voice and convert it into text data. Specifically, the Google Speech-to-Text API converts the voice signal "hello" into the text data "hello". In this step, the input is the user's voice and the output is text data.
[0427] Step 3:
[0428] The terminal sends the generated text data to the server via the network communication module. In this step, the text data is sent using the HTTP protocol or WebSocket. The input is the text data "Hello", and the output is the text data sent to the server.
[0429] Step 4:
[0430] The server receives text data over the network. The received text data is "Hello". The input is text data sent from the terminal, which becomes the data for the next translation process.
[0431] Step 5:
[0432] The server uses a language translation module to translate the received text data into another language. Here, we use the DeepL API and the Google Translate API to translate the Japanese "こんにちは" (konnichiwa) into the English "Hello". The input is the Japanese text data "こんにちは", and the output is the English text data "Hello".
[0433] Step 6:
[0434] The server uses a text-to-speech module to convert translated text data into speech data. Specifically, it uses the Google Text-to-Speech API or Amazon Polly to convert the English text "Hello" into speech data. The input is the English text data "Hello," and the output is speech data.
[0435] Step 7:
[0436] The server sends the generated audio data to the terminal via a network communication module. In this step, the HTTP protocol or WebSocket is used to send the audio data. The input is the audio data, and the output is the audio data sent to the terminal.
[0437] Step 8:
[0438] The device receives audio data over the network and plays it back using its speakers or connected headset. Specifically, the device's audio device plays the audio data "Hello". The input is audio data sent from the server, and the output is the audio that the user can hear.
[0439] (Application Example 1)
[0440] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0441] In situations where smooth communication between multiple languages is difficult, especially in tourist areas, tourists often cannot understand the local language and have difficulty obtaining necessary information. Furthermore, insufficient translation of local guides and information signs makes efficient tourist guidance challenging.
[0442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0443] In this invention, the server includes means for inputting voice, means for converting the voice into text data, means for translating the text data from one language to another, means for converting the translated text data into voice data, means for outputting the voice data, and means for being a tourist guide application including an assistive device. This enables tourists to instantly understand local information in their native language.
[0444] "Means of inputting voice" refers to a function that allows the system to capture voices emitted by the user using an input device such as a microphone.
[0445] "Means for converting the audio into text data" refers to a function for analyzing the input audio and converting it into data in the corresponding text format.
[0446] "Means for translating the text data from a specific language to another language" refers to a function for converting text data from its original language to another specified language.
[0447] "Means for converting the translated text data into audio data" refers to a function for analyzing the translated text data and converting it into data in the corresponding audio format.
[0448] "Means for outputting the audio data" refers to a function for playing back the converted audio data through an output device such as a speaker.
[0449] "Accessible devices" are electronic devices that users can carry with them and that provide information in an assistive way, such as smartphones, tablets, and smart glasses.
[0450] A "tourist guide application" is a software program that provides information in multiple languages to users in tourist destinations.
[0451] This invention is a system that enables instantaneous interpretation between multiple languages, and is specifically implemented as an application for tourist guides. The system and its operation will be described in detail below.
[0452] Hardware configuration
[0453] The server includes the following hardware:
[0454] High-performance processor
[0455] Large capacity memory
[0456] High-speed network connection interface
[0457] User terminals include the following devices:
[0458] smartphone
[0459] Speakers and microphones
[0460] display
[0461] Network communication function (Wi-Fi, 4G / 5G)
[0462] Software Configuration
[0463] The following software modules will be installed on the server:
[0464] Speech recognition module (e.g., Google Cloud Speech-to-Text API)
[0465] Language translation module (e.g., Google Translate API)
[0466] Text-to-speech modules (e.g., Google Text-to-Speech API)
[0467] The applications installed on the user's terminal include the following features:
[0468] Voice input and capture
[0469] Sending and receiving text data
[0470] Playback of translated audio data
[0471] Data processing and calculation
[0472] The server receives the audio data sent from the user terminal and processes it in the following steps:
[0473] 1. Speech Recognition: The speech recognition module is used to convert speech data into text data. For example, the speech "Tell me about this temple" is converted to the text "Tell me about this temple".
[0474] 2. Language Translation: Use the language translation module to translate the converted text data into the specified language. For example, "Please tell me about this temple" will be translated into English as "Please tell me about this temple."
[0475] 3. Speech Synthesis: The speech synthesis module is used to convert translated text data into speech data. For example, the text "Please tell me about this temple." is converted into corresponding English speech data.
[0476] 4. Data transmission: The generated audio data is sent back to the user's terminal, and the application plays it.
[0477] Specific example
[0478] For example, consider a situation where tourists are visiting a temple in Kyoto.
[0479] 1. A tourist speaks into a smartphone application and says, "Tell me about this temple."
[0480] 2. The application captures this audio and uses a speech recognition module to convert it into text: "Tell me about this temple."
[0481] 3. The text data is sent to the server via the network and translated into "Please tell me about this temple." using a language translation module.
[0482] 4. The translated text is converted into speech data using a speech synthesis module.
[0483] 5. Audio data is sent from the server to the user's terminal, and the English audio "Please tell me about this temple." is played for the tourist.
[0484] Example of a prompt:
[0485] The user speaks "Tell me about this temple" in Japanese into their smartphone. The application captures the audio, converts it to text using the Google Translate API, and translates it into English. Then, it converts it back into English speech using the Google Text-to-Speech library and plays it through the smartphone's speaker. Tourists can then hear the information in English.
[0486] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0487] Step 1:
[0488] The user launches a smartphone application and asks a question or requests guidance in a specific language (e.g., Japanese) into the microphone. The input is voice data, and the output is captured voice data. At this stage, the application uses the smartphone's microphone to capture the voice and stores it in its internal memory.
[0489] Step 2:
[0490] The terminal application sends the captured audio data to the speech recognition module. The input is the audio data obtained in step 1, and the output is the corresponding text data. Specifically, the speech recognition module performs speech analysis, generates the text data "Tell me about this temple," and saves it to its internal memory.
[0491] Step 3:
[0492] The terminal sends the generated text data to the server. The input is the text data obtained from the speech recognition module, and the output is the text data received by the server. The terminal uses its network communication function to transfer the text data to the server.
[0493] Step 4:
[0494] The server's language translation module translates the received text data into the specified language. The input is the sent text data ("Please tell me about this temple"), and the output is the translated text data ("Please tell me about this temple."). The server uses the Google Translate API to convert the text data to English and stores the translation result in internal memory.
[0495] Step 5:
[0496] The server's text-to-speech module converts translated text data into speech data. The input is translated text data, and the output is the corresponding speech data ("Please tell me about this temple."). The server uses the Google Text-to-Speech API to convert the text data into speech data and stores that speech data in internal memory.
[0497] Step 6:
[0498] The server sends the generated audio data to the terminal. The input is the audio data generated by the server, and the output is the audio data received by the terminal. The server transmits the audio data to the terminal via the network.
[0499] Step 7:
[0500] The device plays the received audio data through its speaker. The input is the audio data received from the server, and the output is the audio the user hears. The device plays the audio data and provides the user with the English response, "Please tell me about this temple."
[0501] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0502] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module.
[0503] Speech recognition module
[0504] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0505] Emotional Engine
[0506] The emotion engine analyzes the user's emotions from captured audio data and extracts emotional information. This emotional information represents emotional states such as anger, joy, and sadness, and is reflected in the translation results. For example, if the audio says "hello," the emotion engine will recognize that the user is happy.
[0507] Language translation module
[0508] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. Furthermore, it takes into account emotional information recognized by the emotion engine when performing the translation. For example, the Japanese word "konnichiwa" is translated into English as "Hello" along with the emotional information "joy".
[0509] Speech synthesis module
[0510] The speech synthesis module generates speech data with appropriate tone and intonation based on translated text data and emotional information. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0511] Network communication module
[0512] The network communication module enables the transmission and reception of text and voice data. The text data and sentiment information generated by the speech recognition module are sent to a server via the network, where translation and speech synthesis are performed, and then the voice data is sent back to the terminal.
[0513] Specific example
[0514] For example, in a meeting setting, it works as follows:
[0515] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?", and their voice contains an expression of anger.
[0516] 2. Terminal A captures the audio and sends the text "What do you think about this project?" and emotion information "Anger" to the server.
[0517] 3. The server's language translation module translates the text into English, "What do you think about this project?", and then adjusts the translation result to take into account the emotional information "anger".
[0518] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" in an angry tone and sends it to terminal B.
[0519] 5. Terminal B plays the audio data, and user B (an English speaker) understands the question and recognizes the emotional state of the user.
[0520] This system enables rapid and accurate communication across multiple languages, and can even convey emotional nuances, resulting in more human-like communication. It is useful not only in business settings but also in personal interactions and emergency situations.
[0521] The following describes the processing flow.
[0522] Step 1:
[0523] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0524] Step 2:
[0525] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0526] Step 3:
[0527] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0528] Step 4:
[0529] The device passes voice data to the emotion engine, which analyzes the user's emotions from the voice. For example, it might recognize "hello" as an emotion of joy.
[0530] Step 5:
[0531] The terminal transmits the converted text data and sentiment information to the server via the network communication module. The transmitted data also includes input language information.
[0532] Step 6:
[0533] The server analyzes the received text data, sentiment information, and input language information. For example, it analyzes the received text "Hello," sentiment information "Joy," and language information "Japanese."
[0534] Step 7:
[0535] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0536] Step 8:
[0537] The server adjusts the translation results by taking emotional information into account. For example, "Hello" will be output with an emotion that reflects happiness.
[0538] Step 9:
[0539] The server passes the translated text data and emotional information to the speech synthesis module, which then converts the text into speech data. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0540] Step 10:
[0541] The server sends the generated audio data to the terminal via the network communication module.
[0542] Step 11:
[0543] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0544] Step 12:
[0545] The user listens to the played audio and understands the translated content and its emotions. In this way, rapid, accurate, and emotionally responsive communication between different languages is achieved.
[0546] (Example 2)
[0547] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0548] Multilingual speech translation technology faces the challenge of simultaneously conveying both accuracy and emotional nuances. Conventional systems simply translate text, making it difficult to generate speech with appropriate tone and intonation that reflects the user's emotions. Furthermore, communication with remote users requires high-speed and reliable data communication methods. To solve these challenges, a system integrating speech recognition, sentiment analysis, translation, speech synthesis, and network communication is necessary.
[0549] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0550] In this invention, the server includes means for receiving voice data, text data, and emotional information, and performing translation and speech synthesis processing based thereon; network communication means for transmitting and receiving the voice data, text data, and emotional information; and means for extracting emotional information from text data. This makes it possible to convey not only rapid and accurate communication between multiple languages, but also emotional nuances.
[0551] "Means for inputting voice" refers to a device or method for a user to provide voice data to a system using an input device such as a microphone.
[0552] "Means for converting speech to text data" refers to software or hardware modules for converting captured speech data into text data. For example, speech recognition technology can be used.
[0553] "Means for extracting emotional information from text data" refers to software or hardware modules that analyze the user's emotions from converted text data and its audio data, and extract that emotional information.
[0554] A "means for translating text data and sentiment information from one language to another" refers to a software or hardware module that performs translation work based on text data and sentiment information. In this process, sentiment information is also reflected in the translation.
[0555] "Means for converting translated text data and emotional information into audio data" refers to software or hardware modules for generating audio data with appropriate tone and intonation based on translated text data and emotional information.
[0556] "Means for outputting audio data" refers to output devices such as speakers or headphones that provide the generated audio data to the user.
[0557] "Network communication means for transmitting and receiving voice data, text data, and emotional information" refers to communication infrastructure such as the internet or a local network for sending and receiving voice data, text data, and emotional information between modules within a system.
[0558] "Means for receiving audio data, text data, and emotional information, and performing translation and speech synthesis processing based thereon" refers to engines or software modules that perform translation and speech synthesis processing based on the various types of data received.
[0559] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module. The functions of each module, and how they process and calculate data, will be described in detail below.
[0560] Speech recognition module
[0561] The device uses its microphone to capture the user's voice and sends it to a speech recognition module. This speech recognition module then uses speech recognition technology to convert the voice data into text data. For example, if the user says "hello," the device captures this voice, and the speech recognition module converts it into the text data "hello."
[0562] Emotional Engine
[0563] The server receives text data generated by the speech recognition module and sends it to the emotion engine. This emotion engine extracts user emotion information from the text and speech data, for example, using emotion recognition technology. For example, it recognizes the emotion of "joy" in response to "hello."
[0564] Language translation module
[0565] The server sends text data containing emotional information to a language translation module, which then uses multilingual translation technology to translate it into another specified language. Furthermore, the emotional information is also reflected in the translation. For example, the Japanese "こんにちは" (konnichiwa) is translated into English as "Hello," and the emotion of "joy" is added.
[0566] Speech synthesis module
[0567] The server sends translated text data and emotional information to a speech synthesis module, which then generates speech data considering appropriate tone and intonation. For example, speech synthesis technology is used to generate speech data for "Hello" with a joyful tone.
[0568] Network communication module
[0569] The server and terminal use a network communication module to send and receive voice data, text data, and sentiment information. For example, data is sent over the internet or a local network, and appropriate processing is performed after it is received.
[0570] Specific example
[0571] For example, the actions taken at an international conference are as follows:
[0572] 1. User A (a Japanese speaker) speaks into the microphone and asks, "What do you think about this project?" At this time, User A is feeling angry.
[0573] 2. Terminal A uses its microphone to capture user A's voice.
[0574] 3. Terminal A's speech recognition module converts the speech data into text data and generates the text, "What do you think about this project?"
[0575] 4. Terminal A sends the converted text data to the emotion engine, which analyzes the emotion of anger and assigns the emotion information for "anger".
[0576] 5. Terminal A sends the generated text data and emotion information "anger" to the server via the network communication module.
[0577] 6. The server sends the received text data to the language translation module, which translates "What do you think about this project?" into English: "What do you think about this project?".
[0578] 7. The server reflects the emotion information "anger" in the translated text data.
[0579] 8. The server sends the translated text data and the emotion information "anger" to the speech synthesis module, which generates "What do you think about this project?" in an angry tone.
[0580] 9. The server transmits the generated audio data to terminal B via the network communication module.
[0581] 10. Terminal B plays the received audio data, "What do you think about this project?", in an angry tone.
[0582] 11. User B (an English speaker) listens to the audio and understands the question and the emotion of anger.
[0583] Example of a prompt
[0584] For example, by inputting the following prompt into the generative AI model, it is possible to generate translations and sentiment-reflecting speech through the steps described above:
[0585] "Translate the following Japanese into English and generate an audio file that reflects the emotion. Japanese: 'What do you think about this project?' Emotion: 'Anger'"
[0586] This enables smooth multilingual communication between users and systems, and also allows for the transmission of emotional nuances. It is extremely useful in business settings, personal interactions, and emergency situations to achieve more human-like communication.
[0587] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0588] Step 1:
[0589] The user speaks into the microphone. For example, they might say "Hello." The input is audio data. The output is audio data captured through the microphone.
[0590] Step 2:
[0591] The device uses its microphone to capture the user's voice and sends that audio data to a speech recognition module. The input is the voice data from the user, and the output is the audio data sent to the speech recognition module. Specifically, the device's microphone captures the user's speech and saves it as an audio file.
[0592] Step 3:
[0593] The device's speech recognition module converts captured audio data into text data. For example, it converts the audio "Hello" into the text data "Hello". The input is the captured audio data, and the output is the converted text data. Specifically, the speech recognition software analyzes the audio waveform and generates the string data.
[0594] Step 4:
[0595] The device sends the converted text data to the emotion engine, which analyzes the text and the original audio data to extract the user's emotional information. For example, it might recognize the emotion of "joy" in response to "hello." The input consists of text data and audio data, and the output is the emotional information associated with the text data. Specifically, the emotion engine analyzes the text data and the tone of the voice and assigns an emotional label.
[0596] Step 5:
[0597] The terminal sends generated text data and sentiment information to the server via a network communication module. The input is text data and sentiment information, and the output is the data sent to the server. Specifically, the terminal's network module creates data packets and sends them to the server via the internet.
[0598] Step 6:
[0599] The server receives text data and sends it to a language translation module for translation into the specified language. For example, it translates "こんにちは" to "Hello". The input is text data, and the output is translated text data. Specifically, the server's translation software processes the text data and converts it into text in the corresponding language.
[0600] Step 7:
[0601] The server incorporates sentiment information into the translated text data. For example, it assigns the sentiment "joy" to "Hello". The input is the translated text data and sentiment information, and the output is the translated text with the sentiment information added. Specifically, it applies the sentiment label to the translation result and generates the final translated text.
[0602] Step 8:
[0603] The server sends translated text data and sentiment information to a speech synthesis module, which converts the text into speech data. For example, it generates speech data for "Hello" with a "joyful" tone. The input is translated text data and sentiment information, and the output is synthesized speech data. Specifically, the speech synthesis software generates a speech waveform based on the sentiment label and text data.
[0604] Step 9:
[0605] The server transmits the generated voice data to terminal B via the network communication module. The input is the generated voice data, and the output is the data to be transmitted to terminal B. Specifically, the server's network module converts the synthesized voice into data packets and transmits them to terminal B via the internet.
[0606] Step 10:
[0607] Terminal B plays the audio data it receives. The input is the received audio data, and the output is user B listening to the audio. Specifically, terminal B's speaker plays the received audio data, and user B listens to the audio.
[0608] (Application Example 2)
[0609] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0610] This invention aims to provide a system for instantaneous multilingual interpretation that enables real-time speech recognition, translation, and display while reflecting emotions. Modern content delivery services require users to understand live streaming in various languages with rich emotional depth. However, current technology struggles to appropriately reflect emotional information in multilingual translation, potentially degrading the quality of communication. There is a need for technology that can solve this problem and enable more natural and human-like communication.
[0611] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice, means for converting voice into text data, means for translating text data from one language to another, means for adjusting the translated text data based on emotional information and converting it into voice data, means for outputting voice data via network communication means, and means for displaying it in real time as text data and voice data that reflect emotions. This enables users to communicate in real time, including emotions, between multiple languages.
[0612] "Means of inputting audio" refers to devices and software for capturing audio data.
[0613] "Means for converting the audio into text data" refers to speech recognition technology and software for converting captured audio data into text data.
[0614] "Means for translating the text data from a specific language to another language" refers to translation technologies and software for converting text data expressed in a specific language into another language.
[0615] "Means for adjusting the translated text data based on emotional information and converting it into audio data" refers to technology and software for generating audio data that reflects emotional information based on translated text data.
[0616] "Means for outputting the audio data via network communication means" refers to the technology and software for transmitting the generated audio data to other devices or servers via a network.
[0617] "Means of displaying text and audio data that reflect emotions in real time" refers to technologies and software for instantly displaying translation results that include emotional information visually and audibly.
[0618] "Network communication means" refers to communication infrastructure and protocols such as the internet and local networks used to send and receive data.
[0619] A "server" refers to a computer system used to process data and communicate with other devices.
[0620] This invention aims to translate and display streaming content in emotionally rich language using a system that integrates real-time multilingual translation and emotion recognition. The specific program and its processing are described below.
[0621] First, the user's device (smartphone, smart glasses, head-mounted display) captures the audio from the live stream. This audio data is then converted into text data by a speech recognition module. The speech recognition module used is Google's speech recognition API (e.g., Python's speech_recognition library).
[0622] The text data generated by speech recognition is sent to the emotion engine, which analyzes the user's emotions. This emotion engine utilizes the Emotion Recognition library. Emotional information represents emotional states such as anger, joy, and sadness, and this is reflected in the translation results.
[0623] Next, the text data is passed to a language translation module and translated into the specified language. The Google Translate API (e.g., the googletrans library in Python) is used for the translation. Sentiment information is also taken into account during the translation, so for example, "hello" will be translated along with the sentiment information "joy."
[0624] The translated text data and sentiment information are converted into speech data with appropriate tone and intonation by a speech synthesis module. Google Text-to-Speech (e.g., the gtts library in Python) is used for speech synthesis to generate speech data adjusted based on the sentiment information.
[0625] The generated audio data is transmitted in real time to the user's terminal or other devices via a network communication module. This allows the user to instantly see and hear the translated results, which reflect emotions.
[0626] Specific examples are given below.
[0627] User A speaks in Japanese, "What do you think about this project?", and the voice contains an emotion of anger. Terminal A captures the voice, and a speech recognition module converts it into text data, "What do you think about this project?". An emotion engine analyzes this and extracts the emotion information "anger". Then, a language translation module translates the text data into English, "What do you think about this project?", resulting in a translation that reflects the emotion information "anger". Finally, a speech synthesis module generates voice data with an angry tone based on this data and sends it to Terminal B via a network communication module. Terminal B plays the voice data, and User B understands the question and recognizes their emotional state.
[0628] Example of a prompt
[0629] "What do you think of this project? (with anger)"
[0630] This invention enables more natural and human-like real-time communication between multiple languages.
[0631] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0632] Step 1:
[0633] The user speaks into the live stream, and the device captures that audio. The input is the user's voice data, and the output is the captured audio data. Specifically, the audio is recorded using the microphones of a smartphone, smart glasses, or head-mounted display.
[0634] Step 2:
[0635] The device passes the captured audio data to a speech recognition module, which converts it into text data. The input is the captured audio data, and the output is the converted text data. Specifically, Google's speech recognition API (Python's speech_recognition library) is used to convert the audio data into detailed text data.
[0636] Step 3:
[0637] The server receives text data from the speech recognition module and sends it to the emotion engine to extract emotional information. The input is text data, and the output is emotional information. Specifically, the Emotion Recognition library is used to perform emotional analysis on the text data. In this process, emotional states such as anger, joy, and sadness are evaluated.
[0638] Step 4:
[0639] The server passes text data and sentiment information to a language translation module for translation into another language. The input is text data and sentiment information, and the output is translated text data. Specifically, the Google Translate API (the googletrans library in Python) is used to perform accurate translations between multiple languages and to reflect sentiment information.
[0640] Step 5:
[0641] The server passes translated text data and sentiment information to a speech synthesis module, which then generates speech data with appropriate tone and intonation. The input is translated text data and sentiment information, and the output is the generated speech data. Specifically, Google Text-to-Speech (the gtts library in Python) is used to create speech data with a specific tone based on the sentiment information.
[0642] Step 6:
[0643] The server transmits the generated audio data to the user's terminal via network communication. The input is the generated audio data, and the output is the transmitted audio data. Specifically, data is transferred using the internet or a local network.
[0644] Step 7:
[0645] The system plays back audio data received by the user's device and displays a translation result that reflects emotions in real time. The input is the received audio data, and the output is the played audio data and the displayed text data. Specifically, it uses the speakers and displays of smartphones, smart glasses, and head-mounted displays to provide the user with emotionally rich audio and text.
[0646] Example of a prompt
[0647] "What do you think of this project? (with anger)"
[0648] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0649] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0650] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0651] [Third Embodiment]
[0652] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0653] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0654] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0655] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0656] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0657] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0658] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0659] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0660] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0661] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0662] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0663] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0664] This invention provides a system for instantaneous multilingual interpretation. This system consists of a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0665] Speech recognition module
[0666] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0667] Language translation module
[0668] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. The text data is sent from the terminal to the server, which uses the language translation module to perform the translation. For example, the text "こんにちは" (konnichiwa) is translated into English as "Hello".
[0669] Speech synthesis module
[0670] The speech synthesis module converts the translated text data back into speech data. The server receives the translated text and uses the speech synthesis module to generate speech data. For example, the translated text "Hello" is generated as speech data.
[0671] Network communication module
[0672] The network communication module enables the transmission and reception of text and voice data. Text data generated by the speech recognition module is sent to the server via the network, and voice data, after translation and speech synthesis, is sent back to the terminal.
[0673] Specific example
[0674] For example, in a meeting setting, it works as follows:
[0675] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?"
[0676] 2. Terminal A captures the audio and sends it to the server as text: "What do you think about this project?"
[0677] 3. The server's language translation module translates the text into English: "What do you think about this project?".
[0678] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" and sends it to terminal B.
[0679] 5. Terminal B plays the audio data, and User B (an English speaker) is able to understand the question.
[0680] This system enables smooth communication between multiple languages, significantly reducing the time and cost of interpretation. It also leads to increased operational efficiency and prevention of problems, and is expected to have diverse applications.
[0681] The following describes the processing flow.
[0682] Step 1:
[0683] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0684] Step 2:
[0685] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0686] Step 3:
[0687] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0688] Step 4:
[0689] The terminal sends the converted text data to the server via the network communication module. The transmitted data also includes input language information.
[0690] Step 5:
[0691] The server analyzes the received text data and input language information. For example, it analyzes the received text "Hello" and its language information "Japanese".
[0692] Step 6:
[0693] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0694] Step 7:
[0695] The server passes the translated text data to the speech synthesis module, which converts it from text to speech data. For example, the text "Hello" is generated as speech data.
[0696] Step 8:
[0697] The server sends the generated audio data to the terminal via the network communication module.
[0698] Step 9:
[0699] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0700] Step 10:
[0701] The user listens to the played audio and understands the translated content. In this way, rapid and accurate communication between different languages is achieved.
[0702] (Example 1)
[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0704] In real-time multilingual communication, conventional systems have problems with the rapid and accurate recognition, translation, and output of speech data. Furthermore, the processing of speech and text data is distributed, which can lead to network communication delays and errors, thus compromising the user experience.
[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0706] In this invention, the server includes means for converting speech to text data, means for transmitting and receiving the text data and speech data via network communication means, means for translating the text data from one language to another, means for converting the translated text data to speech data, and means for outputting the received speech data. This enables real-time, rapid, and accurate communication between multiple languages.
[0707] A "means of inputting voice" refers to a device that captures the voice spoken by a user as a digital signal.
[0708] "Means for converting audio to text data" refers to devices or programs that convert captured audio signals into text data in string format.
[0709] "Means of translating text data from one language to another" refers to software or algorithms that automatically convert input text data into another specified language.
[0710] "Means for converting translated text data into audio data" refers to software or hardware for generating translated text data as an audio signal.
[0711] "Network communication means" refers to devices and technologies for transmitting and receiving data (text data and audio data) via the internet or local networks.
[0712] A "server" is a computer system that receives data via a network, processes it, and transmits the results.
[0713] "Means for outputting audio data" refers to a device or program for playing back the generated audio data from an output device such as a speaker.
[0714] A "device" is hardware or electronic equipment used to perform a specific purpose or function.
[0715] This invention is a system for real-time multilingual interpretation. This system consists of the following main modules: a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0716] Speech recognition module
[0717] The voice spoken by the user into the microphone is converted into text data by the speech recognition module installed in the device. Specifically, speech recognition software such as the Google Speech-to-Text API is used.
[0718] Language translation module
[0719] The text data generated by the speech recognition module is sent from the terminal to the server via the network communication module. The server is equipped with a language translation module and uses translation software such as the DeepL API or Google Translate API to translate the text data into the specified language.
[0720] Speech synthesis module
[0721] The translated text data is then converted into speech data by a speech synthesis module on the server. Specifically, speech synthesis software such as the Google Text-to-Speech API or Amazon Polly is used.
[0722] Network communication module
[0723] The generated audio data is then transmitted back to the terminal via the network communication module. The terminal receives this audio data and plays the sound through an output device such as a speaker.
[0724] Specific example
[0725] For example, in a meeting setting, it works as follows:
[0726] 1. User A (a Japanese speaker) asks into the microphone, "What do you think about this project?"
[0727] 2. Device A captures the audio and converts it into text data, "What do you think about this project?", using the Google Speech-to-Text API.
[0728] 3. Terminal A sends the generated text data to the server via the network communication module.
[0729] 4. The server receives the text data and uses the DeepL API to translate it into "What do you think about this project?".
[0730] 5. The server uses the Google Text-to-Speech API to convert "What do you think about this project?" into audio data.
[0731] 6. The server transmits the voice data to terminal B via the network communication module.
[0732] 7. Terminal B plays back the received audio data, allowing User B (an English speaker) to understand the question.
[0733] Example of a prompt
[0734] Here are examples of prompts when using a generative AI model in this system:
[0735] "This system uses a speech recognition module, a language translation module, a speech synthesis module, and a network communication module to perform real-time multilingual interpretation. When a user speaks into the microphone, their voice is converted into text data and translated into the specified language. The translated text data is then converted back into speech and played back. Please provide specific use scenarios."
[0736] This system enables smooth communication between multiple languages and will be extremely useful in meetings and international business operations.
[0737] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0738] Step 1:
[0739] The user inputs voice into the microphone. This input is the user's voice, for example, by saying "Hello."
[0740] Step 2:
[0741] The device uses a speech recognition module to capture the user's voice and convert it into text data. Specifically, the Google Speech-to-Text API converts the voice signal "hello" into the text data "hello". In this step, the input is the user's voice and the output is text data.
[0742] Step 3:
[0743] The terminal sends the generated text data to the server via the network communication module. In this step, the text data is sent using the HTTP protocol or WebSocket. The input is the text data "Hello", and the output is the text data sent to the server.
[0744] Step 4:
[0745] The server receives text data over the network. The received text data is "Hello". The input is text data sent from the terminal, which becomes the data for the next translation process.
[0746] Step 5:
[0747] The server uses a language translation module to translate the received text data into another language. Here, we use the DeepL API and the Google Translate API to translate the Japanese "こんにちは" (konnichiwa) into the English "Hello". The input is the Japanese text data "こんにちは", and the output is the English text data "Hello".
[0748] Step 6:
[0749] The server uses a text-to-speech module to convert translated text data into speech data. Specifically, it uses the Google Text-to-Speech API or Amazon Polly to convert the English text "Hello" into speech data. The input is the English text data "Hello," and the output is speech data.
[0750] Step 7:
[0751] The server sends the generated audio data to the terminal via a network communication module. In this step, the HTTP protocol or WebSocket is used to send the audio data. The input is the audio data, and the output is the audio data sent to the terminal.
[0752] Step 8:
[0753] The device receives audio data over the network and plays it back using its speakers or connected headset. Specifically, the device's audio device plays the audio data "Hello". The input is audio data sent from the server, and the output is the audio that the user can hear.
[0754] (Application Example 1)
[0755] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0756] In situations where smooth communication between multiple languages is difficult, especially in tourist areas, tourists often cannot understand the local language and have difficulty obtaining necessary information. Furthermore, insufficient translation of local guides and information signs makes efficient tourist guidance challenging.
[0757] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0758] In this invention, the server includes means for inputting voice, means for converting the voice into text data, means for translating the text data from one language to another, means for converting the translated text data into voice data, means for outputting the voice data, and means for being a tourist guide application including an assistive device. This enables tourists to instantly understand local information in their native language.
[0759] "Means of inputting voice" refers to a function that allows the system to capture voices emitted by the user using an input device such as a microphone.
[0760] "Means for converting the audio into text data" refers to a function for analyzing the input audio and converting it into data in the corresponding text format.
[0761] "Means for translating the text data from a specific language to another language" refers to a function for converting text data from its original language to another specified language.
[0762] "Means for converting the translated text data into audio data" refers to a function for analyzing the translated text data and converting it into data in the corresponding audio format.
[0763] "Means for outputting the audio data" refers to a function for playing back the converted audio data through an output device such as a speaker.
[0764] "Accessible devices" are electronic devices that users can carry with them and that provide information in an assistive way, such as smartphones, tablets, and smart glasses.
[0765] A "tourist guide application" is a software program that provides information in multiple languages to users in tourist destinations.
[0766] This invention is a system that enables instantaneous interpretation between multiple languages, and is specifically implemented as an application for tourist guides. The system and its operation will be described in detail below.
[0767] Hardware configuration
[0768] The server includes the following hardware:
[0769] High-performance processor
[0770] Large capacity memory
[0771] High-speed network connection interface
[0772] User terminals include the following devices:
[0773] smartphone
[0774] Speakers and microphones
[0775] display
[0776] Network communication function (Wi-Fi, 4G / 5G)
[0777] Software Configuration
[0778] The following software modules will be installed on the server:
[0779] Speech recognition module (e.g., Google Cloud Speech-to-Text API)
[0780] Language translation module (e.g., Google Translate API)
[0781] Text-to-speech modules (e.g., Google Text-to-Speech API)
[0782] The applications installed on the user's terminal include the following features:
[0783] Voice input and capture
[0784] Sending and receiving text data
[0785] Playback of translated audio data
[0786] Data processing and calculation
[0787] The server receives the audio data sent from the user terminal and processes it in the following steps:
[0788] 1. Speech Recognition: The speech recognition module is used to convert speech data into text data. For example, the speech "Tell me about this temple" is converted to the text "Tell me about this temple".
[0789] 2. Language Translation: Use the language translation module to translate the converted text data into the specified language. For example, "Please tell me about this temple" will be translated into English as "Please tell me about this temple."
[0790] 3. Speech Synthesis: The speech synthesis module is used to convert translated text data into speech data. For example, the text "Please tell me about this temple." is converted into corresponding English speech data.
[0791] 4. Data transmission: The generated audio data is sent back to the user's terminal, and the application plays it.
[0792] Specific example
[0793] For example, consider a situation where tourists are visiting a temple in Kyoto.
[0794] 1. A tourist speaks into a smartphone application and says, "Tell me about this temple."
[0795] 2. The application captures this audio and uses a speech recognition module to convert it into text: "Tell me about this temple."
[0796] 3. The text data is sent to the server via the network and translated into "Please tell me about this temple." using a language translation module.
[0797] 4. The translated text is converted into speech data using a speech synthesis module.
[0798] 5. Audio data is sent from the server to the user's terminal, and the English audio "Please tell me about this temple." is played for the tourist.
[0799] Example of a prompt:
[0800] The user speaks "Tell me about this temple" in Japanese into their smartphone. The application captures the audio, converts it to text using the Google Translate API, and translates it into English. Then, it converts it back into English speech using the Google Text-to-Speech library and plays it through the smartphone's speaker. Tourists can then hear the information in English.
[0801] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0802] Step 1:
[0803] The user launches a smartphone application and asks a question or requests guidance in a specific language (e.g., Japanese) into the microphone. The input is voice data, and the output is captured voice data. At this stage, the application uses the smartphone's microphone to capture the voice and stores it in its internal memory.
[0804] Step 2:
[0805] The terminal application sends the captured audio data to the speech recognition module. The input is the audio data obtained in step 1, and the output is the corresponding text data. Specifically, the speech recognition module performs speech analysis, generates the text data "Tell me about this temple," and saves it to its internal memory.
[0806] Step 3:
[0807] The terminal sends the generated text data to the server. The input is the text data obtained from the speech recognition module, and the output is the text data received by the server. The terminal uses its network communication function to transfer the text data to the server.
[0808] Step 4:
[0809] The server's language translation module translates the received text data into the specified language. The input is the sent text data ("Please tell me about this temple"), and the output is the translated text data ("Please tell me about this temple."). The server uses the Google Translate API to convert the text data to English and stores the translation result in internal memory.
[0810] Step 5:
[0811] The server's text-to-speech module converts translated text data into speech data. The input is translated text data, and the output is the corresponding speech data ("Please tell me about this temple."). The server uses the Google Text-to-Speech API to convert the text data into speech data and stores that speech data in internal memory.
[0812] Step 6:
[0813] The server sends the generated audio data to the terminal. The input is the audio data generated by the server, and the output is the audio data received by the terminal. The server transmits the audio data to the terminal via the network.
[0814] Step 7:
[0815] The device plays the received audio data through its speaker. The input is the audio data received from the server, and the output is the audio the user hears. The device plays the audio data and provides the user with the English response, "Please tell me about this temple."
[0816] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0817] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module.
[0818] Speech recognition module
[0819] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0820] Emotional Engine
[0821] The emotion engine analyzes the user's emotions from captured audio data and extracts emotional information. This emotional information represents emotional states such as anger, joy, and sadness, and is reflected in the translation results. For example, if the audio says "hello," the emotion engine will recognize that the user is happy.
[0822] Language translation module
[0823] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. Furthermore, it takes into account emotional information recognized by the emotion engine when performing the translation. For example, the Japanese word "konnichiwa" is translated into English as "Hello" along with the emotional information "joy".
[0824] Speech synthesis module
[0825] The speech synthesis module generates speech data with appropriate tone and intonation based on translated text data and emotional information. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0826] Network communication module
[0827] The network communication module enables the transmission and reception of text and voice data. The text data and sentiment information generated by the speech recognition module are sent to a server via the network, where translation and speech synthesis are performed, and then the voice data is sent back to the terminal.
[0828] Specific example
[0829] For example, in a meeting setting, it works as follows:
[0830] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?", and their voice contains an expression of anger.
[0831] 2. Terminal A captures the audio and sends the text "What do you think about this project?" and emotion information "Anger" to the server.
[0832] 3. The server's language translation module translates the text into English, "What do you think about this project?", and then adjusts the translation result to take into account the emotional information "anger".
[0833] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" in an angry tone and sends it to terminal B.
[0834] 5. Terminal B plays the audio data, and user B (an English speaker) understands the question and recognizes the emotional state of the user.
[0835] This system enables rapid and accurate communication across multiple languages, and can even convey emotional nuances, resulting in more human-like communication. It is useful not only in business settings but also in personal interactions and emergency situations.
[0836] The following describes the processing flow.
[0837] Step 1:
[0838] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[0839] Step 2:
[0840] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[0841] Step 3:
[0842] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[0843] Step 4:
[0844] The device passes voice data to the emotion engine, which analyzes the user's emotions from the voice. For example, it might recognize "hello" as an emotion of joy.
[0845] Step 5:
[0846] The terminal transmits the converted text data and sentiment information to the server via the network communication module. The transmitted data also includes input language information.
[0847] Step 6:
[0848] The server analyzes the received text data, sentiment information, and input language information. For example, it analyzes the received text "Hello," sentiment information "Joy," and language information "Japanese."
[0849] Step 7:
[0850] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[0851] Step 8:
[0852] The server adjusts the translation results by taking emotional information into account. For example, "Hello" will be output with an emotion that reflects happiness.
[0853] Step 9:
[0854] The server passes the translated text data and emotional information to the speech synthesis module, which then converts the text into speech data. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[0855] Step 10:
[0856] The server sends the generated audio data to the terminal via the network communication module.
[0857] Step 11:
[0858] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[0859] Step 12:
[0860] The user listens to the played audio and understands the translated content and its emotions. In this way, rapid, accurate, and emotionally responsive communication between different languages is achieved.
[0861] (Example 2)
[0862] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0863] Multilingual speech translation technology faces the challenge of simultaneously conveying both accuracy and emotional nuances. Conventional systems simply translate text, making it difficult to generate speech with appropriate tone and intonation that reflects the user's emotions. Furthermore, communication with remote users requires high-speed and reliable data communication methods. To solve these challenges, a system integrating speech recognition, sentiment analysis, translation, speech synthesis, and network communication is necessary.
[0864] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0865] In this invention, the server includes means for receiving voice data, text data, and emotional information, and performing translation and speech synthesis processing based thereon; network communication means for transmitting and receiving the voice data, text data, and emotional information; and means for extracting emotional information from text data. This makes it possible to convey not only rapid and accurate communication between multiple languages, but also emotional nuances.
[0866] "Means for inputting voice" refers to a device or method for a user to provide voice data to a system using an input device such as a microphone.
[0867] "Means for converting speech to text data" refers to software or hardware modules for converting captured speech data into text data. For example, speech recognition technology can be used.
[0868] "Means for extracting emotional information from text data" refers to software or hardware modules that analyze the user's emotions from converted text data and its audio data, and extract that emotional information.
[0869] A "means for translating text data and sentiment information from one language to another" refers to a software or hardware module that performs translation work based on text data and sentiment information. In this process, sentiment information is also reflected in the translation.
[0870] "Means for converting translated text data and emotional information into audio data" refers to software or hardware modules for generating audio data with appropriate tone and intonation based on translated text data and emotional information.
[0871] "Means for outputting audio data" refers to output devices such as speakers or headphones that provide the generated audio data to the user.
[0872] "Network communication means for transmitting and receiving voice data, text data, and emotional information" refers to communication infrastructure such as the internet or a local network for sending and receiving voice data, text data, and emotional information between modules within a system.
[0873] "Means for receiving audio data, text data, and emotional information, and performing translation and speech synthesis processing based thereon" refers to engines or software modules that perform translation and speech synthesis processing based on the various types of data received.
[0874] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module. The functions of each module, and how they process and calculate data, will be described in detail below.
[0875] Speech recognition module
[0876] The device uses its microphone to capture the user's voice and sends it to a speech recognition module. This speech recognition module then uses speech recognition technology to convert the voice data into text data. For example, if the user says "hello," the device captures this voice, and the speech recognition module converts it into the text data "hello."
[0877] Emotional Engine
[0878] The server receives text data generated by the speech recognition module and sends it to the emotion engine. This emotion engine extracts user emotion information from the text and speech data, for example, using emotion recognition technology. For example, it recognizes the emotion of "joy" in response to "hello."
[0879] Language translation module
[0880] The server sends text data containing emotional information to a language translation module, which then uses multilingual translation technology to translate it into another specified language. Furthermore, the emotional information is also reflected in the translation. For example, the Japanese "こんにちは" (konnichiwa) is translated into English as "Hello," and the emotion of "joy" is added.
[0881] Speech synthesis module
[0882] The server sends translated text data and emotional information to a speech synthesis module, which then generates speech data considering appropriate tone and intonation. For example, speech synthesis technology is used to generate speech data for "Hello" with a joyful tone.
[0883] Network communication module
[0884] The server and terminal use a network communication module to send and receive voice data, text data, and sentiment information. For example, data is sent over the internet or a local network, and appropriate processing is performed after it is received.
[0885] Specific example
[0886] For example, the actions taken at an international conference are as follows:
[0887] 1. User A (a Japanese speaker) speaks into the microphone and asks, "What do you think about this project?" At this time, User A is feeling angry.
[0888] 2. Terminal A uses its microphone to capture user A's voice.
[0889] 3. Terminal A's speech recognition module converts the speech data into text data and generates the text, "What do you think about this project?"
[0890] 4. Terminal A sends the converted text data to the emotion engine, which analyzes the emotion of anger and assigns the emotion information for "anger".
[0891] 5. Terminal A sends the generated text data and emotion information "anger" to the server via the network communication module.
[0892] 6. The server sends the received text data to the language translation module, which translates "What do you think about this project?" into English: "What do you think about this project?".
[0893] 7. The server reflects the emotion information "anger" in the translated text data.
[0894] 8. The server sends the translated text data and the emotion information "anger" to the speech synthesis module, which generates "What do you think about this project?" in an angry tone.
[0895] 9. The server transmits the generated audio data to terminal B via the network communication module.
[0896] 10. Terminal B plays the received audio data, "What do you think about this project?", in an angry tone.
[0897] 11. User B (an English speaker) listens to the audio and understands the question and the emotion of anger.
[0898] Example of a prompt
[0899] For example, by inputting the following prompt into the generative AI model, it is possible to generate translations and sentiment-reflecting speech through the steps described above:
[0900] "Translate the following Japanese into English and generate an audio file that reflects the emotion. Japanese: 'What do you think about this project?' Emotion: 'Anger'"
[0901] This enables smooth multilingual communication between users and systems, and also allows for the transmission of emotional nuances. It is extremely useful in business settings, personal interactions, and emergency situations to achieve more human-like communication.
[0902] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0903] Step 1:
[0904] The user speaks into the microphone. For example, they might say "Hello." The input is audio data. The output is audio data captured through the microphone.
[0905] Step 2:
[0906] The device uses its microphone to capture the user's voice and sends that audio data to a speech recognition module. The input is the voice data from the user, and the output is the audio data sent to the speech recognition module. Specifically, the device's microphone captures the user's speech and saves it as an audio file.
[0907] Step 3:
[0908] The device's speech recognition module converts captured audio data into text data. For example, it converts the audio "Hello" into the text data "Hello". The input is the captured audio data, and the output is the converted text data. Specifically, the speech recognition software analyzes the audio waveform and generates the string data.
[0909] Step 4:
[0910] The device sends the converted text data to the emotion engine, which analyzes the text and the original audio data to extract the user's emotional information. For example, it might recognize the emotion of "joy" in response to "hello." The input consists of text data and audio data, and the output is the emotional information associated with the text data. Specifically, the emotion engine analyzes the text data and the tone of the voice and assigns an emotional label.
[0911] Step 5:
[0912] The terminal sends generated text data and sentiment information to the server via a network communication module. The input is text data and sentiment information, and the output is the data sent to the server. Specifically, the terminal's network module creates data packets and sends them to the server via the internet.
[0913] Step 6:
[0914] The server receives text data and sends it to a language translation module for translation into the specified language. For example, it translates "こんにちは" to "Hello". The input is text data, and the output is translated text data. Specifically, the server's translation software processes the text data and converts it into text in the corresponding language.
[0915] Step 7:
[0916] The server incorporates sentiment information into the translated text data. For example, it assigns the sentiment "joy" to "Hello". The input is the translated text data and sentiment information, and the output is the translated text with the sentiment information added. Specifically, it applies the sentiment label to the translation result and generates the final translated text.
[0917] Step 8:
[0918] The server sends translated text data and sentiment information to a speech synthesis module, which converts the text into speech data. For example, it generates speech data for "Hello" with a "joyful" tone. The input is translated text data and sentiment information, and the output is synthesized speech data. Specifically, the speech synthesis software generates a speech waveform based on the sentiment label and text data.
[0919] Step 9:
[0920] The server transmits the generated voice data to terminal B via the network communication module. The input is the generated voice data, and the output is the data to be transmitted to terminal B. Specifically, the server's network module converts the synthesized voice into data packets and transmits them to terminal B via the internet.
[0921] Step 10:
[0922] Terminal B plays the audio data it receives. The input is the received audio data, and the output is user B listening to the audio. Specifically, terminal B's speaker plays the received audio data, and user B listens to the audio.
[0923] (Application Example 2)
[0924] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0925] This invention aims to provide a system for instantaneous multilingual interpretation that enables real-time speech recognition, translation, and display while reflecting emotions. Modern content delivery services require users to understand live streaming in various languages with rich emotional depth. However, current technology struggles to appropriately reflect emotional information in multilingual translation, potentially degrading the quality of communication. There is a need for technology that can solve this problem and enable more natural and human-like communication.
[0926] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice, means for converting voice into text data, means for translating text data from one language to another, means for adjusting the translated text data based on emotional information and converting it into voice data, means for outputting voice data via network communication means, and means for displaying it in real time as text data and voice data that reflect emotions. This enables users to communicate in real time, including emotions, between multiple languages.
[0927] "Means of inputting audio" refers to devices and software for capturing audio data.
[0928] "Means for converting the audio into text data" refers to speech recognition technology and software for converting captured audio data into text data.
[0929] "Means for translating the text data from a specific language to another language" refers to translation technologies and software for converting text data expressed in a specific language into another language.
[0930] "Means for adjusting the translated text data based on emotional information and converting it into audio data" refers to technology and software for generating audio data that reflects emotional information based on translated text data.
[0931] "Means for outputting the audio data via network communication means" refers to the technology and software for transmitting the generated audio data to other devices or servers via a network.
[0932] "Means of displaying text and audio data that reflect emotions in real time" refers to technologies and software for instantly displaying translation results that include emotional information visually and audibly.
[0933] "Network communication means" refers to communication infrastructure and protocols such as the internet and local networks used to send and receive data.
[0934] A "server" refers to a computer system used to process data and communicate with other devices.
[0935] This invention aims to translate and display streaming content in emotionally rich language using a system that integrates real-time multilingual translation and emotion recognition. The specific program and its processing are described below.
[0936] First, the user's device (smartphone, smart glasses, head-mounted display) captures the audio from the live stream. This audio data is then converted into text data by a speech recognition module. The speech recognition module used is Google's speech recognition API (e.g., Python's speech_recognition library).
[0937] The text data generated by speech recognition is sent to the emotion engine, which analyzes the user's emotions. This emotion engine utilizes the Emotion Recognition library. Emotional information represents emotional states such as anger, joy, and sadness, and this is reflected in the translation results.
[0938] Next, the text data is passed to a language translation module and translated into the specified language. The Google Translate API (e.g., the googletrans library in Python) is used for the translation. Sentiment information is also taken into account during the translation, so for example, "hello" will be translated along with the sentiment information "joy."
[0939] The translated text data and sentiment information are converted into speech data with appropriate tone and intonation by a speech synthesis module. Google Text-to-Speech (e.g., the gtts library in Python) is used for speech synthesis to generate speech data adjusted based on the sentiment information.
[0940] The generated audio data is transmitted in real time to the user's terminal or other devices via a network communication module. This allows the user to instantly see and hear the translated results, which reflect emotions.
[0941] Specific examples are given below.
[0942] User A speaks in Japanese, "What do you think about this project?", and the voice contains an emotion of anger. Terminal A captures the voice, and a speech recognition module converts it into text data, "What do you think about this project?". An emotion engine analyzes this and extracts the emotion information "anger". Then, a language translation module translates the text data into English, "What do you think about this project?", resulting in a translation that reflects the emotion information "anger". Finally, a speech synthesis module generates voice data with an angry tone based on this data and sends it to Terminal B via a network communication module. Terminal B plays the voice data, and User B understands the question and recognizes their emotional state.
[0943] Example of a prompt
[0944] "What do you think of this project? (with anger)"
[0945] This invention enables more natural and human-like real-time communication between multiple languages.
[0946] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0947] Step 1:
[0948] The user speaks into the live stream, and the device captures that audio. The input is the user's voice data, and the output is the captured audio data. Specifically, the audio is recorded using the microphones of a smartphone, smart glasses, or head-mounted display.
[0949] Step 2:
[0950] The device passes the captured audio data to a speech recognition module, which converts it into text data. The input is the captured audio data, and the output is the converted text data. Specifically, Google's speech recognition API (Python's speech_recognition library) is used to convert the audio data into detailed text data.
[0951] Step 3:
[0952] The server receives text data from the speech recognition module and sends it to the emotion engine to extract emotional information. The input is text data, and the output is emotional information. Specifically, the Emotion Recognition library is used to perform emotional analysis on the text data. In this process, emotional states such as anger, joy, and sadness are evaluated.
[0953] Step 4:
[0954] The server passes text data and sentiment information to a language translation module for translation into another language. The input is text data and sentiment information, and the output is translated text data. Specifically, the Google Translate API (the googletrans library in Python) is used to perform accurate translations between multiple languages and to reflect sentiment information.
[0955] Step 5:
[0956] The server passes translated text data and sentiment information to a speech synthesis module, which then generates speech data with appropriate tone and intonation. The input is translated text data and sentiment information, and the output is the generated speech data. Specifically, Google Text-to-Speech (the gtts library in Python) is used to create speech data with a specific tone based on the sentiment information.
[0957] Step 6:
[0958] The server transmits the generated audio data to the user's terminal via network communication. The input is the generated audio data, and the output is the transmitted audio data. Specifically, data is transferred using the internet or a local network.
[0959] Step 7:
[0960] The system plays back audio data received by the user's device and displays a translation result that reflects emotions in real time. The input is the received audio data, and the output is the played audio data and the displayed text data. Specifically, it uses the speakers and displays of smartphones, smart glasses, and head-mounted displays to provide the user with emotionally rich audio and text.
[0961] Example of a prompt
[0962] "What do you think of this project? (with anger)"
[0963] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0964] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0965] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0966] [Fourth Embodiment]
[0967] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0968] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0969] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0970] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0971] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0972] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0973] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0974] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0975] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0976] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0977] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0978] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0979] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0980] This invention provides a system for instantaneous multilingual interpretation. This system consists of a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[0981] Speech recognition module
[0982] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[0983] Language translation module
[0984] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. The text data is sent from the terminal to the server, which uses the language translation module to perform the translation. For example, the text "こんにちは" (konnichiwa) is translated into English as "Hello".
[0985] Speech synthesis module
[0986] The speech synthesis module converts the translated text data back into speech data. The server receives the translated text and uses the speech synthesis module to generate speech data. For example, the translated text "Hello" is generated as speech data.
[0987] Network communication module
[0988] The network communication module enables the transmission and reception of text and voice data. Text data generated by the speech recognition module is sent to the server via the network, and voice data, after translation and speech synthesis, is sent back to the terminal.
[0989] Specific example
[0990] For example, in a meeting setting, it works as follows:
[0991] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?"
[0992] 2. Terminal A captures the audio and sends it to the server as text: "What do you think about this project?"
[0993] 3. The server's language translation module translates the text into English: "What do you think about this project?".
[0994] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" and sends it to terminal B.
[0995] 5. Terminal B plays the audio data, and User B (an English speaker) is able to understand the question.
[0996] This system enables smooth communication between multiple languages, significantly reducing the time and cost of interpretation. It also leads to increased operational efficiency and prevention of problems, and is expected to have diverse applications.
[0997] The following describes the processing flow.
[0998] Step 1:
[0999] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[1000] Step 2:
[1001] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[1002] Step 3:
[1003] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[1004] Step 4:
[1005] The terminal sends the converted text data to the server via the network communication module. The transmitted data also includes input language information.
[1006] Step 5:
[1007] The server analyzes the received text data and input language information. For example, it analyzes the received text "Hello" and its language information "Japanese".
[1008] Step 6:
[1009] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[1010] Step 7:
[1011] The server passes the translated text data to the speech synthesis module, which converts it from text to speech data. For example, the text "Hello" is generated as speech data.
[1012] Step 8:
[1013] The server sends the generated audio data to the terminal via the network communication module.
[1014] Step 9:
[1015] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[1016] Step 10:
[1017] The user listens to the played audio and understands the translated content. In this way, rapid and accurate communication between different languages is achieved.
[1018] (Example 1)
[1019] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1020] In real-time multilingual communication, conventional systems have problems with the rapid and accurate recognition, translation, and output of speech data. Furthermore, the processing of speech and text data is distributed, which can lead to network communication delays and errors, thus compromising the user experience.
[1021] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1022] In this invention, the server includes means for converting speech to text data, means for transmitting and receiving the text data and speech data via network communication means, means for translating the text data from one language to another, means for converting the translated text data to speech data, and means for outputting the received speech data. This enables real-time, rapid, and accurate communication between multiple languages.
[1023] A "means of inputting voice" refers to a device that captures the voice spoken by a user as a digital signal.
[1024] "Means for converting audio to text data" refers to devices or programs that convert captured audio signals into text data in string format.
[1025] "Means of translating text data from one language to another" refers to software or algorithms that automatically convert input text data into another specified language.
[1026] "Means for converting translated text data into audio data" refers to software or hardware for generating translated text data as an audio signal.
[1027] "Network communication means" refers to devices and technologies for transmitting and receiving data (text data and audio data) via the internet or local networks.
[1028] A "server" is a computer system that receives data via a network, processes it, and transmits the results.
[1029] "Means for outputting audio data" refers to a device or program for playing back the generated audio data from an output device such as a speaker.
[1030] A "device" is hardware or electronic equipment used to perform a specific purpose or function.
[1031] This invention is a system for real-time multilingual interpretation. This system consists of the following main modules: a speech recognition module, a language translation module, a speech synthesis module, and a network communication module.
[1032] Speech recognition module
[1033] The voice spoken by the user into the microphone is converted into text data by the speech recognition module installed in the device. Specifically, speech recognition software such as the Google Speech-to-Text API is used.
[1034] Language translation module
[1035] The text data generated by the speech recognition module is sent from the terminal to the server via the network communication module. The server is equipped with a language translation module and uses translation software such as the DeepL API or Google Translate API to translate the text data into the specified language.
[1036] Speech synthesis module
[1037] The translated text data is then converted into speech data by a speech synthesis module on the server. Specifically, speech synthesis software such as the Google Text-to-Speech API or Amazon Polly is used.
[1038] Network communication module
[1039] The generated audio data is then transmitted back to the terminal via the network communication module. The terminal receives this audio data and plays the sound through an output device such as a speaker.
[1040] Specific example
[1041] For example, in a meeting setting, it works as follows:
[1042] 1. User A (a Japanese speaker) asks into the microphone, "What do you think about this project?"
[1043] 2. Device A captures the audio and converts it into text data, "What do you think about this project?", using the Google Speech-to-Text API.
[1044] 3. Terminal A sends the generated text data to the server via the network communication module.
[1045] 4. The server receives the text data and uses the DeepL API to translate it into "What do you think about this project?".
[1046] 5. The server uses the Google Text-to-Speech API to convert "What do you think about this project?" into audio data.
[1047] 6. The server transmits the voice data to terminal B via the network communication module.
[1048] 7. Terminal B plays back the received audio data, allowing User B (an English speaker) to understand the question.
[1049] Example of a prompt
[1050] Here are examples of prompts when using a generative AI model in this system:
[1051] "This system uses a speech recognition module, a language translation module, a speech synthesis module, and a network communication module to perform real-time multilingual interpretation. When a user speaks into the microphone, their voice is converted into text data and translated into the specified language. The translated text data is then converted back into speech and played back. Please provide specific use scenarios."
[1052] This system enables smooth communication between multiple languages and will be extremely useful in meetings and international business operations.
[1053] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1054] Step 1:
[1055] The user inputs voice into the microphone. This input is the user's voice, for example, by saying "Hello."
[1056] Step 2:
[1057] The device uses a speech recognition module to capture the user's voice and convert it into text data. Specifically, the Google Speech-to-Text API converts the voice signal "hello" into the text data "hello". In this step, the input is the user's voice and the output is text data.
[1058] Step 3:
[1059] The terminal sends the generated text data to the server via the network communication module. In this step, the text data is sent using the HTTP protocol or WebSocket. The input is the text data "Hello", and the output is the text data sent to the server.
[1060] Step 4:
[1061] The server receives text data over the network. The received text data is "Hello". The input is text data sent from the terminal, which becomes the data for the next translation process.
[1062] Step 5:
[1063] The server uses a language translation module to translate the received text data into another language. Here, we use the DeepL API and the Google Translate API to translate the Japanese "こんにちは" (konnichiwa) into the English "Hello". The input is the Japanese text data "こんにちは", and the output is the English text data "Hello".
[1064] Step 6:
[1065] The server uses a text-to-speech module to convert translated text data into speech data. Specifically, it uses the Google Text-to-Speech API or Amazon Polly to convert the English text "Hello" into speech data. The input is the English text data "Hello," and the output is speech data.
[1066] Step 7:
[1067] The server sends the generated audio data to the terminal via a network communication module. In this step, the HTTP protocol or WebSocket is used to send the audio data. The input is the audio data, and the output is the audio data sent to the terminal.
[1068] Step 8:
[1069] The device receives audio data over the network and plays it back using its speakers or connected headset. Specifically, the device's audio device plays the audio data "Hello". The input is audio data sent from the server, and the output is the audio that the user can hear.
[1070] (Application Example 1)
[1071] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1072] In situations where smooth communication between multiple languages is difficult, especially in tourist areas, tourists often cannot understand the local language and have difficulty obtaining necessary information. Furthermore, insufficient translation of local guides and information signs makes efficient tourist guidance challenging.
[1073] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1074] In this invention, the server includes means for inputting voice, means for converting the voice into text data, means for translating the text data from one language to another, means for converting the translated text data into voice data, means for outputting the voice data, and means for being a tourist guide application including an assistive device. This enables tourists to instantly understand local information in their native language.
[1075] "Means of inputting voice" refers to a function that allows the system to capture voices emitted by the user using an input device such as a microphone.
[1076] "Means for converting the audio into text data" refers to a function for analyzing the input audio and converting it into data in the corresponding text format.
[1077] "Means for translating the text data from a specific language to another language" refers to a function for converting text data from its original language to another specified language.
[1078] "Means for converting the translated text data into audio data" refers to a function for analyzing the translated text data and converting it into data in the corresponding audio format.
[1079] "Means for outputting the audio data" refers to a function for playing back the converted audio data through an output device such as a speaker.
[1080] "Accessible devices" are electronic devices that users can carry with them and that provide information in an assistive way, such as smartphones, tablets, and smart glasses.
[1081] A "tourist guide application" is a software program that provides information in multiple languages to users in tourist destinations.
[1082] This invention is a system that enables instantaneous interpretation between multiple languages, and is specifically implemented as an application for tourist guides. The system and its operation will be described in detail below.
[1083] Hardware configuration
[1084] The server includes the following hardware:
[1085] High-performance processor
[1086] Large capacity memory
[1087] High-speed network connection interface
[1088] User terminals include the following devices:
[1089] smartphone
[1090] Speakers and microphones
[1091] display
[1092] Network communication function (Wi-Fi, 4G / 5G)
[1093] Software Configuration
[1094] The following software modules will be installed on the server:
[1095] Speech recognition module (e.g., Google Cloud Speech-to-Text API)
[1096] Language translation module (e.g., Google Translate API)
[1097] Text-to-speech modules (e.g., Google Text-to-Speech API)
[1098] The applications installed on the user's terminal include the following features:
[1099] Voice input and capture
[1100] Sending and receiving text data
[1101] Playback of translated audio data
[1102] Data processing and calculation
[1103] The server receives the audio data sent from the user terminal and processes it in the following steps:
[1104] 1. Speech Recognition: The speech recognition module is used to convert speech data into text data. For example, the speech "Tell me about this temple" is converted to the text "Tell me about this temple".
[1105] 2. Language Translation: Use the language translation module to translate the converted text data into the specified language. For example, "Please tell me about this temple" will be translated into English as "Please tell me about this temple."
[1106] 3. Speech Synthesis: The speech synthesis module is used to convert translated text data into speech data. For example, the text "Please tell me about this temple." is converted into corresponding English speech data.
[1107] 4. Data transmission: The generated audio data is sent back to the user's terminal, and the application plays it.
[1108] Specific example
[1109] For example, consider a situation where tourists are visiting a temple in Kyoto.
[1110] 1. A tourist speaks into a smartphone application and says, "Tell me about this temple."
[1111] 2. The application captures this audio and uses a speech recognition module to convert it into text: "Tell me about this temple."
[1112] 3. The text data is sent to the server via the network and translated into "Please tell me about this temple." using a language translation module.
[1113] 4. The translated text is converted into speech data using a speech synthesis module.
[1114] 5. Audio data is sent from the server to the user's terminal, and the English audio "Please tell me about this temple." is played for the tourist.
[1115] Example of a prompt:
[1116] The user speaks "Tell me about this temple" in Japanese into their smartphone. The application captures the audio, converts it to text using the Google Translate API, and translates it into English. Then, it converts it back into English speech using the Google Text-to-Speech library and plays it through the smartphone's speaker. Tourists can then hear the information in English.
[1117] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1118] Step 1:
[1119] The user launches a smartphone application and asks a question or requests guidance in a specific language (e.g., Japanese) into the microphone. The input is voice data, and the output is captured voice data. At this stage, the application uses the smartphone's microphone to capture the voice and stores it in its internal memory.
[1120] Step 2:
[1121] The terminal application sends the captured audio data to the speech recognition module. The input is the audio data obtained in step 1, and the output is the corresponding text data. Specifically, the speech recognition module performs speech analysis, generates the text data "Tell me about this temple," and saves it to its internal memory.
[1122] Step 3:
[1123] The terminal sends the generated text data to the server. The input is the text data obtained from the speech recognition module, and the output is the text data received by the server. The terminal uses its network communication function to transfer the text data to the server.
[1124] Step 4:
[1125] The server's language translation module translates the received text data into the specified language. The input is the sent text data ("Please tell me about this temple"), and the output is the translated text data ("Please tell me about this temple."). The server uses the Google Translate API to convert the text data to English and stores the translation result in internal memory.
[1126] Step 5:
[1127] The server's text-to-speech module converts translated text data into speech data. The input is translated text data, and the output is the corresponding speech data ("Please tell me about this temple."). The server uses the Google Text-to-Speech API to convert the text data into speech data and stores that speech data in internal memory.
[1128] Step 6:
[1129] The server sends the generated audio data to the terminal. The input is the audio data generated by the server, and the output is the audio data received by the terminal. The server transmits the audio data to the terminal via the network.
[1130] Step 7:
[1131] The device plays the received audio data through its speaker. The input is the audio data received from the server, and the output is the audio the user hears. The device plays the audio data and provides the user with the English response, "Please tell me about this temple."
[1132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1133] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module.
[1134] Speech recognition module
[1135] The speech recognition module captures the voice a user speaks into the microphone and converts it into text. For example, if a user says "hello," the device captures this voice and uses the speech recognition module to convert it into the text data "hello."
[1136] Emotional Engine
[1137] The emotion engine analyzes the user's emotions from captured audio data and extracts emotional information. This emotional information represents emotional states such as anger, joy, and sadness, and is reflected in the translation results. For example, if the audio says "hello," the emotion engine will recognize that the user is happy.
[1138] Language translation module
[1139] The language translation module receives text data generated by the speech recognition module and translates it into the specified language. Furthermore, it takes into account emotional information recognized by the emotion engine when performing the translation. For example, the Japanese word "konnichiwa" is translated into English as "Hello" along with the emotional information "joy".
[1140] Speech synthesis module
[1141] The speech synthesis module generates speech data with appropriate tone and intonation based on translated text data and emotional information. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[1142] Network communication module
[1143] The network communication module enables the transmission and reception of text and voice data. The text data and sentiment information generated by the speech recognition module are sent to a server via the network, where translation and speech synthesis are performed, and then the voice data is sent back to the terminal.
[1144] Specific example
[1145] For example, in a meeting setting, it works as follows:
[1146] 1. User A (a Japanese person) asks in Japanese, "What do you think about this project?", and their voice contains an expression of anger.
[1147] 2. Terminal A captures the audio and sends the text "What do you think about this project?" and emotion information "Anger" to the server.
[1148] 3. The server's language translation module translates the text into English, "What do you think about this project?", and then adjusts the translation result to take into account the emotional information "anger".
[1149] 4. The server's speech synthesis module generates the English voice data "What do you think about this project?" in an angry tone and sends it to terminal B.
[1150] 5. Terminal B plays the audio data, and user B (an English speaker) understands the question and recognizes the emotional state of the user.
[1151] This system enables rapid and accurate communication across multiple languages, and can even convey emotional nuances, resulting in more human-like communication. It is useful not only in business settings but also in personal interactions and emergency situations.
[1152] The following describes the processing flow.
[1153] Step 1:
[1154] The user speaks into the microphone. For example, the user says "Konnichiwa" (hello) in Japanese.
[1155] Step 2:
[1156] The device captures the audio. The user's voice is captured as digital audio data using the microphone.
[1157] Step 3:
[1158] The device passes the voice data to the speech recognition module, which converts the voice into text data. For example, the voice "Hello" is converted to the text "Hello".
[1159] Step 4:
[1160] The device passes voice data to the emotion engine, which analyzes the user's emotions from the voice. For example, it might recognize "hello" as an emotion of joy.
[1161] Step 5:
[1162] The terminal transmits the converted text data and sentiment information to the server via the network communication module. The transmitted data also includes input language information.
[1163] Step 6:
[1164] The server analyzes the received text data, sentiment information, and input language information. For example, it analyzes the received text "Hello," sentiment information "Joy," and language information "Japanese."
[1165] Step 7:
[1166] The server uses a language translation module to translate the parsed text into the target language. For example, the Japanese "こんにちは" is translated into the English "Hello".
[1167] Step 8:
[1168] The server adjusts the translation results by taking emotional information into account. For example, "Hello" will be output with an emotion that reflects happiness.
[1169] Step 9:
[1170] The server passes the translated text data and emotional information to the speech synthesis module, which then converts the text into speech data. For example, "Hello" is generated as speech data with a tone that reflects the emotion of joy.
[1171] Step 10:
[1172] The server sends the generated audio data to the terminal via the network communication module.
[1173] Step 11:
[1174] The device plays the audio data it received. The audio data "Hello" is played using an output device such as a speaker.
[1175] Step 12:
[1176] The user listens to the played audio and understands the translated content and its emotions. In this way, rapid, accurate, and emotionally responsive communication between different languages is achieved.
[1177] (Example 2)
[1178] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1179] Multilingual speech translation technology faces the challenge of simultaneously conveying both accuracy and emotional nuances. Conventional systems simply translate text, making it difficult to generate speech with appropriate tone and intonation that reflects the user's emotions. Furthermore, communication with remote users requires high-speed and reliable data communication methods. To solve these challenges, a system integrating speech recognition, sentiment analysis, translation, speech synthesis, and network communication is necessary.
[1180] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1181] In this invention, the server includes means for receiving voice data, text data, and emotional information, and performing translation and speech synthesis processing based thereon; network communication means for transmitting and receiving the voice data, text data, and emotional information; and means for extracting emotional information from text data. This makes it possible to convey not only rapid and accurate communication between multiple languages, but also emotional nuances.
[1182] "Means for inputting voice" refers to a device or method for a user to provide voice data to a system using an input device such as a microphone.
[1183] "Means for converting speech to text data" refers to software or hardware modules for converting captured speech data into text data. For example, speech recognition technology can be used.
[1184] "Means for extracting emotional information from text data" refers to software or hardware modules that analyze the user's emotions from converted text data and its audio data, and extract that emotional information.
[1185] A "means for translating text data and sentiment information from one language to another" refers to a software or hardware module that performs translation work based on text data and sentiment information. In this process, sentiment information is also reflected in the translation.
[1186] "Means for converting translated text data and emotional information into audio data" refers to software or hardware modules for generating audio data with appropriate tone and intonation based on translated text data and emotional information.
[1187] "Means for outputting audio data" refers to output devices such as speakers or headphones that provide the generated audio data to the user.
[1188] "Network communication means for transmitting and receiving voice data, text data, and emotional information" refers to communication infrastructure such as the internet or a local network for sending and receiving voice data, text data, and emotional information between modules within a system.
[1189] "Means for receiving audio data, text data, and emotional information, and performing translation and speech synthesis processing based thereon" refers to engines or software modules that perform translation and speech synthesis processing based on the various types of data received.
[1190] This invention provides a system for instantaneous multilingual interpretation, further incorporating an emotion engine that recognizes the user's emotions and reflects them in the interpretation results. This system consists of a speech recognition module, an emotion engine, a language translation module, a speech synthesis module, and a network communication module. The functions of each module, and how they process and calculate data, will be described in detail below.
[1191] Speech recognition module
[1192] The device uses its microphone to capture the user's voice and sends it to a speech recognition module. This speech recognition module then uses speech recognition technology to convert the voice data into text data. For example, if the user says "hello," the device captures this voice, and the speech recognition module converts it into the text data "hello."
[1193] Emotional Engine
[1194] The server receives text data generated by the speech recognition module and sends it to the emotion engine. This emotion engine extracts user emotion information from the text and speech data, for example, using emotion recognition technology. For example, it recognizes the emotion of "joy" in response to "hello."
[1195] Language translation module
[1196] The server sends text data containing emotional information to a language translation module, which then uses multilingual translation technology to translate it into another specified language. Furthermore, the emotional information is also reflected in the translation. For example, the Japanese "こんにちは" (konnichiwa) is translated into English as "Hello," and the emotion of "joy" is added.
[1197] Speech synthesis module
[1198] The server sends translated text data and emotional information to a speech synthesis module, which then generates speech data considering appropriate tone and intonation. For example, speech synthesis technology is used to generate speech data for "Hello" with a joyful tone.
[1199] Network communication module
[1200] The server and terminal use a network communication module to send and receive voice data, text data, and sentiment information. For example, data is sent over the internet or a local network, and appropriate processing is performed after it is received.
[1201] Specific example
[1202] For example, the actions taken at an international conference are as follows:
[1203] 1. User A (a Japanese speaker) speaks into the microphone and asks, "What do you think about this project?" At this time, User A is feeling angry.
[1204] 2. Terminal A uses its microphone to capture user A's voice.
[1205] 3. Terminal A's speech recognition module converts the speech data into text data and generates the text, "What do you think about this project?"
[1206] 4. Terminal A sends the converted text data to the emotion engine, which analyzes the emotion of anger and assigns the emotion information for "anger".
[1207] 5. Terminal A sends the generated text data and emotion information "anger" to the server via the network communication module.
[1208] 6. The server sends the received text data to the language translation module, which translates "What do you think about this project?" into English: "What do you think about this project?".
[1209] 7. The server reflects the emotion information "anger" in the translated text data.
[1210] 8. The server sends the translated text data and the emotion information "anger" to the speech synthesis module, which generates "What do you think about this project?" in an angry tone.
[1211] 9. The server transmits the generated audio data to terminal B via the network communication module.
[1212] 10. Terminal B plays the received audio data, "What do you think about this project?", in an angry tone.
[1213] 11. User B (an English speaker) listens to the audio and understands the question and the emotion of anger.
[1214] Example of a prompt
[1215] For example, by inputting the following prompt into the generative AI model, it is possible to generate translations and sentiment-reflecting speech through the steps described above:
[1216] "Translate the following Japanese into English and generate an audio file that reflects the emotion. Japanese: 'What do you think about this project?' Emotion: 'Anger'"
[1217] This enables smooth multilingual communication between users and systems, and also allows for the transmission of emotional nuances. It is extremely useful in business settings, personal interactions, and emergency situations to achieve more human-like communication.
[1218] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1219] Step 1:
[1220] The user speaks into the microphone. For example, they might say "Hello." The input is audio data. The output is audio data captured through the microphone.
[1221] Step 2:
[1222] The device uses its microphone to capture the user's voice and sends that audio data to a speech recognition module. The input is the voice data from the user, and the output is the audio data sent to the speech recognition module. Specifically, the device's microphone captures the user's speech and saves it as an audio file.
[1223] Step 3:
[1224] The device's speech recognition module converts captured audio data into text data. For example, it converts the audio "Hello" into the text data "Hello". The input is the captured audio data, and the output is the converted text data. Specifically, the speech recognition software analyzes the audio waveform and generates the string data.
[1225] Step 4:
[1226] The device sends the converted text data to the emotion engine, which analyzes the text and the original audio data to extract the user's emotional information. For example, it might recognize the emotion of "joy" in response to "hello." The input consists of text data and audio data, and the output is the emotional information associated with the text data. Specifically, the emotion engine analyzes the text data and the tone of the voice and assigns an emotional label.
[1227] Step 5:
[1228] The terminal sends generated text data and sentiment information to the server via a network communication module. The input is text data and sentiment information, and the output is the data sent to the server. Specifically, the terminal's network module creates data packets and sends them to the server via the internet.
[1229] Step 6:
[1230] The server receives text data and sends it to a language translation module for translation into the specified language. For example, it translates "こんにちは" to "Hello". The input is text data, and the output is translated text data. Specifically, the server's translation software processes the text data and converts it into text in the corresponding language.
[1231] Step 7:
[1232] The server incorporates sentiment information into the translated text data. For example, it assigns the sentiment "joy" to "Hello". The input is the translated text data and sentiment information, and the output is the translated text with the sentiment information added. Specifically, it applies the sentiment label to the translation result and generates the final translated text.
[1233] Step 8:
[1234] The server sends translated text data and sentiment information to a speech synthesis module, which converts the text into speech data. For example, it generates speech data for "Hello" with a "joyful" tone. The input is translated text data and sentiment information, and the output is synthesized speech data. Specifically, the speech synthesis software generates a speech waveform based on the sentiment label and text data.
[1235] Step 9:
[1236] The server transmits the generated voice data to terminal B via the network communication module. The input is the generated voice data, and the output is the data to be transmitted to terminal B. Specifically, the server's network module converts the synthesized voice into data packets and transmits them to terminal B via the internet.
[1237] Step 10:
[1238] Terminal B plays the audio data it receives. The input is the received audio data, and the output is user B listening to the audio. Specifically, terminal B's speaker plays the received audio data, and user B listens to the audio.
[1239] (Application Example 2)
[1240] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1241] This invention aims to provide a system for instantaneous multilingual interpretation that enables real-time speech recognition, translation, and display while reflecting emotions. Modern content delivery services require users to understand live streaming in various languages with rich emotional depth. However, current technology struggles to appropriately reflect emotional information in multilingual translation, potentially degrading the quality of communication. There is a need for technology that can solve this problem and enable more natural and human-like communication.
[1242] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice, means for converting voice into text data, means for translating text data from one language to another, means for adjusting the translated text data based on emotional information and converting it into voice data, means for outputting voice data via network communication means, and means for displaying it in real time as text data and voice data that reflect emotions. This enables users to communicate in real time, including emotions, between multiple languages.
[1243] "Means of inputting audio" refers to devices and software for capturing audio data.
[1244] "Means for converting the audio into text data" refers to speech recognition technology and software for converting captured audio data into text data.
[1245] "Means for translating the text data from a specific language to another language" refers to translation technologies and software for converting text data expressed in a specific language into another language.
[1246] "Means for adjusting the translated text data based on emotional information and converting it into audio data" refers to technology and software for generating audio data that reflects emotional information based on translated text data.
[1247] "Means for outputting the audio data via network communication means" refers to the technology and software for transmitting the generated audio data to other devices or servers via a network.
[1248] "Means of displaying text and audio data that reflect emotions in real time" refers to technologies and software for instantly displaying translation results that include emotional information visually and audibly.
[1249] "Network communication means" refers to communication infrastructure and protocols such as the internet and local networks used to send and receive data.
[1250] A "server" refers to a computer system used to process data and communicate with other devices.
[1251] This invention aims to translate and display streaming content in emotionally rich language using a system that integrates real-time multilingual translation and emotion recognition. The specific program and its processing are described below.
[1252] First, the user's device (smartphone, smart glasses, head-mounted display) captures the audio from the live stream. This audio data is then converted into text data by a speech recognition module. The speech recognition module used is Google's speech recognition API (e.g., Python's speech_recognition library).
[1253] The text data generated by speech recognition is sent to the emotion engine, which analyzes the user's emotions. This emotion engine utilizes the Emotion Recognition library. Emotional information represents emotional states such as anger, joy, and sadness, and this is reflected in the translation results.
[1254] Next, the text data is passed to a language translation module and translated into the specified language. The Google Translate API (e.g., the googletrans library in Python) is used for the translation. Sentiment information is also taken into account during the translation, so for example, "hello" will be translated along with the sentiment information "joy."
[1255] The translated text data and sentiment information are converted into speech data with appropriate tone and intonation by a speech synthesis module. Google Text-to-Speech (e.g., the gtts library in Python) is used for speech synthesis to generate speech data adjusted based on the sentiment information.
[1256] The generated audio data is transmitted in real time to the user's terminal or other devices via a network communication module. This allows the user to instantly see and hear the translated results, which reflect emotions.
[1257] Specific examples are given below.
[1258] User A speaks in Japanese, "What do you think about this project?", and the voice contains an emotion of anger. Terminal A captures the voice, and a speech recognition module converts it into text data, "What do you think about this project?". An emotion engine analyzes this and extracts the emotion information "anger". Then, a language translation module translates the text data into English, "What do you think about this project?", resulting in a translation that reflects the emotion information "anger". Finally, a speech synthesis module generates voice data with an angry tone based on this data and sends it to Terminal B via a network communication module. Terminal B plays the voice data, and User B understands the question and recognizes their emotional state.
[1259] Example of a prompt
[1260] "What do you think of this project? (with anger)"
[1261] This invention enables more natural and human-like real-time communication between multiple languages.
[1262] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1263] Step 1:
[1264] The user speaks into the live stream, and the device captures that audio. The input is the user's voice data, and the output is the captured audio data. Specifically, the audio is recorded using the microphones of a smartphone, smart glasses, or head-mounted display.
[1265] Step 2:
[1266] The device passes the captured audio data to a speech recognition module, which converts it into text data. The input is the captured audio data, and the output is the converted text data. Specifically, Google's speech recognition API (Python's speech_recognition library) is used to convert the audio data into detailed text data.
[1267] Step 3:
[1268] The server receives text data from the speech recognition module and sends it to the emotion engine to extract emotional information. The input is text data, and the output is emotional information. Specifically, the Emotion Recognition library is used to perform emotional analysis on the text data. In this process, emotional states such as anger, joy, and sadness are evaluated.
[1269] Step 4:
[1270] The server passes text data and sentiment information to a language translation module for translation into another language. The input is text data and sentiment information, and the output is translated text data. Specifically, the Google Translate API (the googletrans library in Python) is used to perform accurate translations between multiple languages and to reflect sentiment information.
[1271] Step 5:
[1272] The server passes translated text data and sentiment information to a speech synthesis module, which then generates speech data with appropriate tone and intonation. The input is translated text data and sentiment information, and the output is the generated speech data. Specifically, Google Text-to-Speech (the gtts library in Python) is used to create speech data with a specific tone based on the sentiment information.
[1273] Step 6:
[1274] The server transmits the generated audio data to the user's terminal via network communication. The input is the generated audio data, and the output is the transmitted audio data. Specifically, data is transferred using the internet or a local network.
[1275] Step 7:
[1276] The system plays back audio data received by the user's device and displays a translation result that reflects emotions in real time. The input is the received audio data, and the output is the played audio data and the displayed text data. Specifically, it uses the speakers and displays of smartphones, smart glasses, and head-mounted displays to provide the user with emotionally rich audio and text.
[1277] Example of a prompt
[1278] "What do you think of this project? (with anger)"
[1279] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1280] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1281] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1282] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1283] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1284] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1285] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1286] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1287] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1288] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1289] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1290] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1291] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1292] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1293] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1294] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1295] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1296] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1297] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1298] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1299] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1300] The following is further disclosed regarding the embodiments described above.
[1301] (Claim 1)
[1302] A means of inputting voice,
[1303] means for converting the aforementioned audio into text data,
[1304] A means for translating the aforementioned text data from one language to another,
[1305] A means for converting the translated text data into audio data,
[1306] The means for outputting the aforementioned audio data,
[1307] A system that includes this.
[1308] (Claim 2)
[1309] The system according to claim 1, further comprising network communication means for transmitting and receiving the aforementioned audio data.
[1310] (Claim 3)
[1311] The system according to claim 1, further comprising a server that receives the aforementioned audio data and text data and performs translation processing based thereon.
[1312] "Example 1"
[1313] (Claim 1)
[1314] A means of inputting voice,
[1315] means for converting the aforementioned audio into text data,
[1316] A means for translating the aforementioned text data from one language to another,
[1317] A means for converting the translated text data into audio data,
[1318] Means for transmitting and receiving the text data and voice data via network communication means,
[1319] means for outputting the received audio data,
[1320] A system that includes this.
[1321] (Claim 2)
[1322] The system according to claim 1, further comprising a system for transmitting and receiving text data and voice data to and from a server using the aforementioned network communication means.
[1323] (Claim 3)
[1324] The system according to claim 1, further comprising a server that receives the text data and audio data and performs translation processing and speech synthesis processing based thereon.
[1325] "Application Example 1"
[1326] (Claim 1)
[1327] A means of inputting voice,
[1328] means for converting the aforementioned audio into text data,
[1329] A means for translating the aforementioned text data from one language to another,
[1330] A means for converting the translated text data into audio data,
[1331] The means for outputting the aforementioned audio data,
[1332] A means characterized by being a tourist guide application including an accessibility device,
[1333] A system that includes this.
[1334] (Claim 2)
[1335] The system according to claim 1, further comprising network communication means for transmitting and receiving the aforementioned audio data.
[1336] (Claim 3)
[1337] The system according to claim 1, further comprising a server that receives the aforementioned audio data and text data and performs translation processing based thereon.
[1338] "Example 2 of combining an emotion engine"
[1339] (Claim 1)
[1340] A means of inputting voice,
[1341] means for converting the aforementioned audio into text data,
[1342] A means for extracting emotional information from the aforementioned text data,
[1343] A means for translating the aforementioned text data and the aforementioned sentiment information from one language to another,
[1344] A means for converting the translated text data and the emotional information into audio data,
[1345] The means for outputting the aforementioned audio data,
[1346] A system that includes this.
[1347] (Claim 2)
[1348] The system according to claim 1, further comprising network communication means for transmitting and receiving the aforementioned voice data, text data, and emotional information.
[1349] (Claim 3)
[1350] The system according to claim 1, further comprising a server that receives the aforementioned voice data, text data, and emotion information, and performs translation and speech synthesis processing based thereon.
[1351] "Application example 2 when combining with an emotional engine"
[1352] (Claim 1)
[1353] A means of inputting voice,
[1354] means for converting the aforementioned audio into text data,
[1355] A means for translating the aforementioned text data from one language to another,
[1356] The means for adjusting the translated text data based on emotional information and converting it into audio data,
[1357] means for outputting the aforementioned audio data via network communication means,
[1358] A means of displaying text and audio data that reflect emotions in real time.
[1359] A system that includes this.
[1360] (Claim 2)
[1361] The system according to claim 1, further comprising network communication means for transmitting and receiving the aforementioned voice data and text data.
[1362] (Claim 3)
[1363] The system according to claim 1, further comprising a server that receives the aforementioned audio data and text data and performs translation processing and speech synthesis based thereon. [Explanation of Symbols]
[1364] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of inputting voice, means for converting the aforementioned audio into text data, A means for translating the aforementioned text data from one language to another, A means for converting the translated text data into audio data, The means for outputting the aforementioned audio data, A system that includes this.
2. The system according to claim 1, further comprising network communication means for transmitting and receiving the aforementioned audio data.
3. The system according to claim 1, further comprising a server that receives the aforementioned audio data and text data and performs translation processing based thereon.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A