system

The system addresses language barriers in real-time communication by using audio acquisition, conversion, analysis, and playback technologies to provide accurate and secure multilingual translation, improving communication in diverse settings.

JP2026062152APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing real-time communication systems face challenges due to language barriers, requiring significant resources for human interpreters and suffer from inaccuracies in automatic translation systems, leading to ineffective communication in multilingual settings.

Method used

A system that includes audio acquisition, conversion, transmission, analysis, translation, and playback components to translate and play back voices in multiple languages in real-time, utilizing devices like microphones, ADCs, communication modules, speech recognition engines, translation APIs, and speech synthesis APIs to ensure accurate and secure multilingual communication.

Benefits of technology

Enables real-time, accurate, and secure multilingual communication by overcoming language barriers, facilitating smooth conversations and enhancing communication efficiency in educational and industrial settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062152000001_ABST
    Figure 2026062152000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 An acquisition means for acquiring voice, A conversion means for converting the acquired voice data into a digital format, A transmission means for transmitting the converted voice data to a server, An analysis means for analyzing the transmitted voice data and identifying the language, A translation means for translating the identified language's voice data into a target language, A voice synthesis means for converting the translated text into voice data, A transmission means for transmitting the synthesized voice data to a terminal, A playback means for playing back the transmitted voice data, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In real-time communication between multiple languages, at present, it is difficult to have a smooth conversation due to language barriers. Especially when people speaking different languages study together in a specialized field such as an educational institution, simultaneous interpretation is required, but a large number of resources are needed for human interpreters to perform this. In addition, existing automatic translation systems often have problems with real-time performance and speech recognition accuracy, and there are many cases where they cannot withstand practical use. To solve this problem, a system that accurately translates and plays back voices in multiple languages in real-time is required.

Means for Solving the Problems

[0005] [[ID=4']] The present invention solves the problem by the following means in particular. First, it includes an acquisition means for acquiring audio. Next, it includes a conversion means for converting the acquired audio data into a digital format. Then, it includes a transmission means for sending the converted audio data to a server. On the server side, it includes an analysis means for analyzing the transmitted audio data and identifying the language. It includes a translation means for translating the audio data of the identified language into a target language, and further includes a speech synthesis means for converting the translated text into audio data. Finally, it includes a transmission means for sending the synthesized audio data to a terminal and a playback means for playing back the transmitted audio data. With a system configured in this way, it becomes possible to recognize, translate, and play back audio in real time, overcoming language barriers.

[0006] "Acquisition means" refers to a device that has the function of detecting the user's speech and surrounding sounds and capturing them as digital audio data.

[0007] A "conversion means" is a device that has the function of converting audio data captured by the acquisition means from analog to digital format.

[0008] "Transmission means" refers to a device that has communication capabilities for transmitting converted digital audio data or other data to a server or another device.

[0009] "Analysis means" refers to a device that analyzes transmitted audio data and has the function of identifying the language contained therein and performing speech recognition.

[0010] A "translation means" is a device that has the function of translating text in a language identified by an analysis means into a specified target language.

[0011] A "speech synthesis means" is a device that has the function of converting text data generated by a translation means into speech data.

[0012] A "playback device" is a device that has the function of playing back audio data sent via a transmission device or the like on a terminal. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Next, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. The configuration of this system and the processing of its program will be described in detail below.

[0035] System Configuration

[0036] This system primarily consists of wearable devices (such as glasses-type devices or earphone-type devices) and a server in the cloud. Users can wear the device and communicate via voice. Each device and server includes the following functional modules.

[0037] 1. The terminal is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the sound data into a digital format, a communication module to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0038] 2. The server includes a speech recognition module that analyzes the received audio data and identifies the language, a translation module that translates the identified language into the target language, a speech synthesis module that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0039] Program processing

[0040] The following is a natural language explanation of the program's processing logic.

[0041] Voice acquisition and transmission

[0042] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transfer is performed using a secure communication protocol.

[0043] Audio data analysis and translation

[0044] The server passes the received audio data to the analysis module. The analysis module uses speech recognition technology (e.g., a speech recognition API) to convert the audio to text. Next, it identifies the language used from the analyzed text. The text in the identified language is passed to the translation module, which translates it into the target language in real time. The translation technology used in this process is, for example, a machine translation engine (translation API, etc.).

[0045] Translated speech synthesis and transmission

[0046] The server passes the translated text to the speech synthesis module. The speech synthesis module converts the text in the target language into speech data (using a speech synthesis API, etc.). The generated speech data is sent to the terminal via the communication module.

[0047] Audio playback

[0048] The device receives audio data from the server and provides it to the user through a playback device (speaker or earphones). Through this process, the user can understand speech in different languages ​​in real time.

[0049] Specific usage examples

[0050] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0051] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics."

[0052] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0053] 3. The server analyzes the received audio data, converts it into Japanese text, and then uses a translation module to translate it into English as "Today's lecture is about the basics of quantum mechanics".

[0054] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0055] 5. The device plays the received English audio data, and User A understands it in real time.

[0056] This enables users who speak different languages ​​to communicate in real time, facilitating smooth lectures even in multinational classes in educational settings.

[0057] The following describes the processing flow.

[0058] Step 1:

[0059] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0060] Step 2:

[0061] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0062] Step 3:

[0063] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0064] Step 4:

[0065] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0066] Step 5:

[0067] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0068] Step 6:

[0069] The server passes text data to the translation module, which then translates it into the target language in real time.

[0070] Step 7:

[0071] The server passes the translated text data to the speech synthesis module, which then generates synthesized speech.

[0072] Step 8:

[0073] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0074] Step 9:

[0075] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0076] Step 10:

[0077] Users can listen to the played audio and understand speech in different languages ​​in real time.

[0078] (Example 1)

[0079] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0080] Current real-time speech translation systems suffer from insufficient security in the communication protocols they use, potentially leading to data eavesdropping during transmission. Furthermore, the accuracy of speech recognition and translation is often low, failing to meet user expectations. Additionally, complex and time-consuming intermediate processing can compromise real-time capabilities.

[0081] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0082] In this invention, the server includes means for analyzing voice data, means for translating, and means for synthesizing voice data. This enables efficient and secure execution of each process of analysis, translation, and voice synthesis, and provides real-time, highly accurate voice translation. Furthermore, by using an optimized communication protocol, the security of data transfer can be improved and user privacy can be protected.

[0083] "Acquisition means" refers to a device used to capture the voice spoken by the user, such as a microphone.

[0084] "Conversion means" refers to a device used to convert captured analog audio data into a digital format, specifically such as an ADC (Analog to Digital Converter).

[0085] "Transmission means" refers to devices and protocols used to send digital audio data to a server in the cloud, such as Wi-Fi modules and secure communication protocols like SSL / TLS.

[0086] "Analysis means" refers to technologies that convert audio data received by a server into text data and identify the language being used, specifically referring to speech recognition engines and speech recognition APIs.

[0087] "Translation methods" refer to technologies used to translate text data obtained through analysis into different languages, such as machine translation engines and translation APIs.

[0088] "Speech synthesis means" refers to technologies for converting translated text into speech data, such as speech synthesis APIs.

[0089] "Playback means" refers to a device that outputs the generated audio data so that the user can hear it, such as a speaker or earphones.

[0090] "Secure communication methods" refer to protocols and technologies for securely transferring voice data, including encrypted communication technologies such as SSL / TLS.

[0091] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. Specific embodiments of this system are described below.

[0092] System Configuration

[0093] This system primarily consists of wearable devices (such as glasses or earphones) and a server in the cloud. Users wear the device and can communicate using voice.

[0094] Device configuration

[0095] The device is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the audio data into a digital format, a communication module (such as a Wi-Fi module) to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0096] Server Configuration

[0097] The server includes a speech recognition module (e.g., a speech recognition API) that analyzes the received audio data and identifies the language, a translation module (e.g., a machine translation engine) that translates the identified language into the target language, a speech synthesis module (e.g., a speech synthesis API) that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0098] Specific example of processing

[0099] As a concrete example, consider a situation where a foreign language learner understands a conversation in real time.

[0100] Situation:

[0101] A scenario where an English-speaking tourist (User A) communicates with a Japanese-speaking guide (User B) at a tourist destination in Japan.

[0102] Situation handling procedure

[0103] 1. User B says in Japanese, "This is Senso-ji Temple."

[0104] 2. The terminal acquires this utterance and converts it into digital speech data using an ADC.

[0105] 3. The terminal sends the converted digital audio data to the server via the communication module. Secure protocols such as SSL / TLS are used for communication.

[0106] 4. The server passes the received audio data to the analysis module (speech recognition API) and converts it into the Japanese text "This is Senso-ji Temple."

[0107] 5. The server passes the analyzed text to a translation module (machine translation engine), which translates it into the English text "This is Sensoji Temple".

[0108] 6. The server passes the translated text to the speech synthesis module (speech synthesis API) and converts it into the English speech data "Hello".

[0109] 7. The server sends the generated English voice data to the terminal via the communication module. A secure protocol is used for communication again.

[0110] 8. The device plays the received English audio data through a playback device (earphones or speaker), and user A understands it as "This is Sensoji Temple."

[0111] This allows users to understand speech in different languages ​​in real time.

[0112] Example prompts for the generative AI model to be used

[0113] The following prompt can be entered using a generative AI model.

[0114] Please describe the specific processing flow of a system that, when a Japanese-speaking user says "Konnichiwa" (こんにちは), translates it into English in real time and provides the user with "Hello." Please also specify the technologies and devices used at each step.

[0115] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0116] Step 1:

[0117] The user says "Hello." The device's microphone captures this audio. The input is an analog audio signal, and the microphone captures it, resulting in an analog audio signal as output.

[0118] Step 2:

[0119] The terminal converts the analog audio signal acquired from the microphone into digital audio data using an ADC (Analog to Digital Converter). The input is an analog audio signal, and the output after conversion is digital audio data.

[0120] Step 3:

[0121] The terminal transmits the converted digital audio data to a server in the cloud via a communication module (e.g., a Wi-Fi module). Secure communication protocols such as SSL / TLS are used during this process. The input is digital audio data, and the output is the digital audio data transmitted to the server.

[0122] Step 4:

[0123] The server passes the received digital audio data to the speech recognition module (speech recognition API). The speech recognition module analyzes the input digital audio data and converts it into text data. Specifically, the output is the Japanese text data "こんにちは" (konnichiwa).

[0124] Step 5:

[0125] The server passes the text data obtained from the speech recognition module to the translation module (machine translation engine). The translation module translates the input Japanese text data "こんにちは" into the English text data "Hello". The output is the English text data "Hello".

[0126] Step 6:

[0127] The server passes the translated English text data "Hello" to the speech synthesis module (speech synthesis API). The speech synthesis module parses the input text data and converts it into speech data. The output is the English speech data "Hello".

[0128] Step 7:

[0129] The server sends the generated English audio data to the terminal via a communication module. Secure protocols such as SSL / TLS are again used for communication. The input is English audio data, and the output is the English audio data sent to the terminal.

[0130] Step 8:

[0131] The device receives English audio data and plays it back through a playback device (earphones or speaker) for the user to hear. The input is English audio data, and the output is the audio that the user can hear.

[0132] This allows users to understand speech in different languages ​​in real time.

[0133] (Application Example 1)

[0134] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0135] In factories, when multinational workers work together, language barriers can hinder smooth communication, leading to decreased production efficiency. Furthermore, important instructions and urgent announcements may not be accurately conveyed, potentially impacting safety. To resolve these issues and improve work efficiency and safety, measures to overcome language differences are necessary.

[0136] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0137] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio data into a digital format, and a transmission means for sending the converted audio data to the server. This enables a system, including a robot, for real-time communication with workers who speak multiple languages ​​in a factory environment.

[0138] "Acquisition means" refers to the devices or systems used to acquire sound.

[0139] "Conversion means" refers to a device or function for converting acquired audio data into a digital format.

[0140] "Transmission means" refers to communication functions or modules for sending the converted audio data to the server.

[0141] "Analysis means" refers to technologies and systems for analyzing transmitted audio data and identifying the language being used.

[0142] "Translation means" refers to technologies and devices for translating audio data of a specified language into a target language.

[0143] "Speech synthesis means" refers to technologies and systems for converting translated text into speech data.

[0144] "Playback means" refers to devices or functions for playing back audio data transmitted from a server.

[0145] A "robot" is a machine used in a factory environment to communicate in real time with multinational workers.

[0146] This invention relates to a multilingual voice translation system for facilitating communication among multinational workers in a factory. The embodiments for carrying out the invention are as follows:

[0147] System Configuration

[0148] This system primarily consists of wearable terminals (e.g., devices mounted on robots or work helmets) and cloud-based servers. Users can communicate via voice through the terminals. Each device and server includes the following functional modules:

[0149] terminal

[0150] Acquisition method: A microphone for acquiring sound. This microphone has a noise reduction function.

[0151] Conversion method: Analog to Digital Converter (ADC) that converts audio data acquired by a microphone into a digital format.

[0152] Transmission method: A communication module (e.g., Wi-Fi or 5G module) for securely transmitting the converted audio data to a server in the cloud.

[0153] Playback method: A high-quality speaker for playing back translated audio data received from the server.

[0154] server

[0155] Analysis method: A speech recognition engine (e.g., Google® Cloud Speech-to-Text) is used to analyze the audio data received by the server and identify the language being used.

[0156] Translation method: A machine translation engine (e.g., Microsoft® Translator) for real-time translation of text in a specified language into a target language.

[0157] Speech synthesis means: Speech synthesis technology for converting translated text into speech data (e.g., Amazon Polly).

[0158] Transmission method: A communication module for transmitting synthesized voice data to a terminal.

[0159] Explanation of program processing

[0160] The server receives the acquired audio data and passes it to a speech recognition engine to convert it into text. The analyzed text is then translated into the target language by a machine translation engine. The translated text is converted back into audio data using speech synthesis technology, and this audio data is transmitted to the terminal. The transmitted audio data is provided to the worker via a playback device, enabling real-time communication among multinational workers.

[0161] Specific example

[0162] For example, consider a scenario in a factory where a Japanese-speaking worker (User A) collaborates with a worker who only speaks English (User B).

[0163] 1. User A says in Japanese, "Please tell me the next steps."

[0164] 2. The terminal acquires this audio, converts it into digital audio data using an ADC, and sends it to the server.

[0165] 3. The server analyzes the received audio data, converts it to Japanese text, and then uses a translation module to translate it into English as "Please explain the next work step".

[0166] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0167] 5. The device plays the received English audio data, and user B understands it in real time.

[0168] Example of a prompt

[0169] Please capture the audio of a factory worker saying "Please tell me the next step" in Japanese. Translate it into English and convert it into audio data that can be understood by an English-speaking technician, then play it back.

[0170] This facilitates smoother communication among multinational workers within the factory, resulting in an efficient and safe working environment.

[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0172] Step 1:

[0173] The user speaks. The input is the user's voice, and the specific action is a factory worker saying "Please tell me the next work procedure" in Japanese. The output is an acoustic signal in the form of a speech waveform.

[0174] Step 2:

[0175] The device captures the user's voice. The input is an acoustic signal, and the specific operation involves a microphone acquiring the voice. The output is analog audio data.

[0176] Step 3:

[0177] The terminal's conversion mechanism converts analog audio data into a digital format. The input is analog audio data, and specifically, the ADC (Analog to Digital Converter) converts the signal into a digital format. The output is digital audio data.

[0178] Step 4:

[0179] The terminal's transmission method sends digital audio data to the server. The input is digital audio data, and specifically, the communication module (Wi-Fi or 5G) sends the data to a server in the cloud. The output is the digital audio data sent to the server.

[0180] Step 5:

[0181] The server's analysis tool analyzes the received audio data and converts it to text using a speech recognition engine. The input is the received digital audio data, and the specific operation is for the speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio into text data. The output is the analyzed text data.

[0182] Step 6:

[0183] The server's analysis tool identifies the language used in the text data. The input is the analyzed text data, and the specific operation is for the language model to identify the language from the text. The output is the identified language information.

[0184] Step 7:

[0185] The server's translation method translates text data in a specified language into the target language. The input consists of specified language information and text data; the specific operation involves a machine translation engine (e.g., Microsoft Translator) translating the text into the target language in real time. The output is the translated text data.

[0186] Step 8:

[0187] The server's speech synthesis system converts translated text data into speech data. The input is translated text data, and the specific operation involves speech synthesis technology (e.g., Amazon Polly) converting the text in the target language into speech data. The output is synthesized speech data.

[0188] Step 9:

[0189] The server's transmission mechanism sends synthesized audio data to the terminal. The input is synthesized audio data, and the specific operation is that the communication module sends the audio data to the terminal. The output is the audio data received by the terminal.

[0190] Step 10:

[0191] The device's playback mechanism plays the received audio data. The input is the audio data received by the device, and the specific operation is for the speaker to play the audio and provide it to the user. The output is audio in a language the user understands.

[0192] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0193] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. The configuration of this system and the processing of its program will be described in detail below.

[0194] System Configuration

[0195] This system consists of wearable devices (e.g., glasses-type devices or earphone-type devices) and a server in the cloud. By wearing the device and communicating via voice, the system enables the following functions. Its main components are as follows:

[0196] 1. The terminal is equipped with a microphone to capture user speech, an ADC (Analog to Digital Converter) to convert audio data into a digital format, a communication module to send data to the server, and a speaker or earphones to play back the received translated audio.

[0197] 2. The server includes a speech recognition module that analyzes received audio data and identifies the language, a translation module that translates the identified language into the target language, and a speech synthesis module that converts the translated text into speech. It also includes an emotion engine that recognizes the user's emotions and has the ability to adjust the tone and nuances of the voice based on those emotions.

[0198] 3. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, anger, surprise, etc.) from the voice data and uses this to influence the translation and speech synthesis processes.

[0199] Program processing

[0200] The following provides a detailed explanation of how this system's program works.

[0201] Voice acquisition and transmission

[0202] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol.

[0203] Audio data analysis and emotion recognition

[0204] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio into text. Simultaneously, an emotion engine operates to analyze the user's emotional state from the audio data.

[0205] Language identification and translation

[0206] The server identifies the language used from the analyzed text. Next, the text data and recognized sentiment data are passed to a translation module, which translates them into the target language in real time. This translation process uses the sentiment data to appropriately adjust tone and nuance.

[0207] Translated speech synthesis and transmission

[0208] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the terminal via the communication module.

[0209] Audio playback

[0210] The device receives audio data from the server and sends it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[0211] Specific usage examples

[0212] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0213] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0214] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0215] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0216] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0217] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0218] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0219] This enables users who speak different languages ​​to communicate in real time, including conveying emotions and nuances. It can be effectively used in educational settings and international business conferences.

[0220] The following describes the processing flow.

[0221] Step 1:

[0222] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0223] Step 2:

[0224] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0225] Step 3:

[0226] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0227] Step 4:

[0228] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0229] Step 5:

[0230] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0231] Step 6:

[0232] The server passes audio data to the emotion engine along with text data in the identified language to recognize the user's emotions. Specifically, it analyzes features such as tone, pace, and intonation of the voice to identify the emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0233] Step 7:

[0234] The server passes the text data and recognized sentiment data to the translation module, which then translates it into the target language. During this process, the translation module takes the sentiment data into account and generates translated text with appropriate tone and nuances.

[0235] Step 8:

[0236] The server passes the translated text data and sentiment data to the speech synthesis module to generate synthesized speech. Based on the sentiment data, the speech synthesis module adjusts the tone and intonation of the speech to produce natural-sounding speech.

[0237] Step 9:

[0238] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0239] Step 10:

[0240] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0241] Step 11:

[0242] Users can listen to the played audio and understand the nuances of speech and emotions in different languages ​​in real time.

[0243] (Example 2)

[0244] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0245] Conventional speech translation systems have a problem in that they fail to properly convey emotions and tone when translating user speech into other languages. Furthermore, because they do not consider the user's emotions during the speech recognition and translation process, nuances of communication are often lost. As a result, especially in educational and business settings, translated content is often not accurately conveyed, leading to misunderstandings.

[0246] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means for analyzing voice data and identifying language, an emotion recognition means for recognizing emotions from voice data, and an emotion adjustment means for adjusting translation and speech synthesis based on the recognized emotions. This enables real-time multilingual translation that reflects the user's emotions and nuances.

[0247] "Acquisition means" refers to a device or function used to capture a user's speech.

[0248] "Conversion means" refers to a device or function that converts acquired audio data from analog format to digital format.

[0249] "Transmission means" refers to a device or function for transmitting audio data converted into a digital format to a server.

[0250] "Analysis means" refers to a device or function that analyzes transmitted audio data and identifies the language.

[0251] "Translation means" refers to a device or function that translates audio data of a specified language into a target language.

[0252] "Speech synthesis means" refers to a device or function that converts translated text into speech data.

[0253] "Emotion recognition means" refers to a device or function that recognizes a user's emotions from audio data.

[0254] "Emotion adjustment means" refers to a device or function that adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0255] "Playback means" refers to a device or function for playing back transmitted audio data to the user.

[0256] A "server" is a central computer system used to perform voice data analysis, translation, speech synthesis, and emotion recognition.

[0257] A "terminal" is a device worn by a user that has the function of acquiring and playing back audio.

[0258] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. This system consists of a wearable device (e.g., glasses-type device or earphone-type device) and a server in the cloud. By wearing the device and communicating via voice, the system achieves the following functions:

[0259] System Configuration

[0260] The main components of this system are as follows:

[0261] terminal

[0262] Acquisition method: Equipped with a microphone for capturing user speech.

[0263] Conversion method: Includes an ADC (Analog to Digital Converter) that converts audio into a digital format.

[0264] Transmission method: It is equipped with a communication module that transmits the acquired digital audio data to a server in the cloud.

[0265] Playback method: It has a speaker or earphones that play back the translated audio received from the server.

[0266] server

[0267] Analysis method: The received audio data is analyzed and converted into text using a speech recognition engine (for example, Google Cloud Speech-to-Text API or Microsoft Azure® Speech Service).

[0268] Translation method: To translate text data into the target language, we use APIs such as Google Translate or DeepL.

[0269] Speech synthesis method: Amazon Polly or Google Cloud Text-to-Speech API are used to convert translated text into speech data.

[0270] Emotion recognition means: An emotion engine is used to recognize the user's emotions from voice data and evaluate tone and nuance.

[0271] Emotion adjustment mechanism: Adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0272] Examples of usage and prompt statements

[0273] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0274] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0275] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0276] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0277] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0278] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0279] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0280] Using a generative AI model, you can utilize the following example prompts:

[0281] Example prompt

[0282] In a university lecture, a professor who speaks Japanese said, "Today's lecture will be about the fundamentals of quantum mechanics," with a sense of questioning. Please explain in detail how to translate this into English in real time, while preserving the sense of questioning, and provide it to students who only speak English.

[0283] This prompt text is useful when specifically explaining the operation of the system for a particular scenario.

[0284] The flow of the specific process in Example 2 will be described using FIG. 13.

[0285] Step 1:

[0286] When the user speaks, the microphone of the terminal acquires the voice data. The acquired voice is in analog form and is converted into digital form by an ADC (Analog to Digital Converter). The input is the user's analog voice, and the output is digital voice data. Through this data conversion, digital voice data that can be used in subsequent processing steps is obtained.

[0287] Step 2:

[0288] The terminal transmits the converted digital voice data to a server on the cloud via a communication module. This transmission uses a secure communication protocol such as HTTPS. The input is digital voice data, and the output is the transmitted voice data. The communication module divides the data into packets and transmits it to ensure it reaches the server.

[0289] Step 3:

[0290] The server passes the received voice data to an analysis module. In the analysis module, a speech recognition engine (e.g., Google Cloud Speech-to-Text API) is used to convert the voice into text. The input is the digital voice data transmitted to the server, and the output is text data. The speech recognition engine performs phoneme and sound waveform analysis and generates appropriate text data.

[0291] Step 4:

[0292] The server's emotion recognition engine analyzes audio data in parallel to identify the user's emotional state. It evaluates the tone, pitch, and speed of the speech to recognize emotional states (such as joy, sadness, anger, or surprise). The input is digital audio data, and the output is emotion data. Emotion recognition is performed using machine learning algorithms.

[0293] Step 5:

[0294] The server passes the analyzed text data and sentiment data to the translation module. The translation module (e.g., Google Translate API or DeepL API) translates the text into the target language in real time. The input is text data and sentiment data, and the output is the translated text. During translation, the tone and nuances are also appropriately adjusted, taking sentiment data into account.

[0295] Step 6:

[0296] The server passes the translated text and sentiment data to a speech synthesis module. The speech synthesis module (e.g., Amazon Polly or Google Cloud Text-to-Speech API) converts the text into speech. The input is the translated text and sentiment data, and the output is the generated synthesized speech data. This process adjusts the tone and intonation of the speech based on the sentiment data.

[0297] Step 7:

[0298] The server sends the generated synthesized speech data to the terminal via a communication module. A secure protocol (e.g., HTTPS) is used for communication to ensure data integrity and confidentiality. The input is the synthesized speech data, and the output is the transmitted speech data.

[0299] Step 8:

[0300] The terminal receives audio data from the server and sends it to a playback device (speaker or earphones), providing it to the user as audio. The input is the received audio data, and the output is the played audio. This allows the user to understand speech in different languages ​​in real time, along with the emotions conveyed.

[0301] (Application Example 2)

[0302] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0303] In global manufacturing environments, language differences pose a significant obstacle when multinational staff collaborate. Furthermore, accurately conveying not only language but also the emotions and nuances of speech is crucial, but current systems fail to adequately address this. Therefore, there is a need for real-time multilingual translation and emotion recognition simultaneously to facilitate smooth communication on the factory floor.

[0304] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes emotion recognition means for analyzing emotional states from voice data, means for adjusting the tone and nuances of translated text based on the analyzed emotion data, analysis means for converting acquired voice data into text using a voice recognition engine and analyzing emotional states using an emotion engine, and translation means using a translation module that translates text data and emotion data in real time and adjusts the tone and nuances based on the emotion data. This enables accurate multilingual translation and emotion recognition in real time.

[0305] "Audio data" refers to data that represents audio information in a digital format.

[0306] A "server" is a computer system that provides data processing and storage services over a network.

[0307] The "emotion recognition means" is a function or device for analyzing and recognizing the speaker's emotion from voice data.

[0308] The "emotion data" is data indicating the emotional state analyzed by the emotion recognition means.

[0309] The "translation means" is a function or device for converting text or voice expressed in a certain language into another language.

[0310] The "translation module" is a software module for translating text data or voice into another language.

[0311] The "speech recognition engine" is a software engine for converting voice data into text.

[0312] "Tone and nuance" refers to the emotional expressions and subtle meanings of voice or text.

[0313] The "analysis means" is a function or device for analyzing the acquired voice data and extracting specific information or patterns.

[0314] The "acquisition means" is a function or device for collecting or capturing data such as voice and video.

[0315] The "reproduction means" is a function or device for reproducing data on a terminal.

[0316] [[ID=......]] Embodiments for Implementing the Invention

[0317] The present invention relates to a system that translates a user's speech into another language in real time, further recognizes the user's emotion, and provides an adapted translation based on that emotion. As an application example of this, it is assumed to realize an application for smart glasses to smooth communication in a global manufacturing site. Hereinafter, the system configuration of the present invention and the processing of the detailed program will be described.

[0318] System Configuration

[0319] This system consists of smart glasses-type devices and a server in the cloud. When the user wears the smart glasses and communicates via voice, the system enables the following functions.

[0320] Device (smart glasses)

[0321] Acquisition method: Includes a microphone for acquiring sound.

[0322] Conversion means: Includes an ADC (Analog to Digital Converter) that converts the acquired audio data into a digital format.

[0323] Transmission means: Includes a communication module that transmits the converted audio data to a server.

[0324] Playback means: Includes speakers or earphones that play audio data transmitted from the server.

[0325] server

[0326] Analysis method: The transmitted audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[0327] Emotion recognition method: An emotion engine (e.g., IBM Watson® Tone Analyzer) is used to analyze the emotional state from voice data.

[0328] Translation method: Text data and sentiment data are translated in real time and translated into the target language using a translation module (e.g., Microsoft Translator API).

[0329] Speech synthesis method: Based on translated text and sentiment data, a speech synthesis module (e.g., Amazon Polly) converts it into speech.

[0330] Program processing

[0331] Voice acquisition and transmission

[0332] When a user speaks, the microphone in the smart glasses captures their voice. The voice data is converted to a digital format by an ADC and sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol such as SSL.

[0333] Audio data analysis and emotion recognition

[0334] The audio data received by the server is converted into text in real time by a speech recognition engine. In parallel, an emotion recognition system operates to analyze the user's emotional state from the audio data.

[0335] Language identification and translation

[0336] The server identifies the language used from the analyzed text, passes the text data and recognized sentiment data to a translation module, and translates it into the target language in real time. This translation process uses sentiment data to appropriately adjust tone and nuances.

[0337] Translated speech synthesis and transmission

[0338] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the smart glasses via the communication module.

[0339] Audio playback

[0340] The smart glasses receive audio data from a server and send it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[0341] Specific example

[0342] Let's consider an example of use in a global factory setting. Suppose a team leader gives the instruction in Japanese: "Please install the next part." If the leader speaks in a calm tone, the smart glasses system will operate as follows:

[0343] 1. The reader's speech is captured by the microphone in the smart glasses and converted into a digital format.

[0344] 2. The digital audio data is sent to a cloud server and converted into text, "Please install the following parts," by a speech recognition engine.

[0345] 3. The emotion recognition system analyzes the calm emotion, and the translation module translates it into English as "Please install the next part." The tone and nuances are also adjusted based on the emotion data.

[0346] 4. The translated text is converted into English speech by a speech synthesis module and sent to the smart glasses.

[0347] 5. Translated English instructions are played through the smart glasses' speaker, allowing staff of other nationalities to understand them.

[0348] Example of a prompt

[0349] "Please provide prototype code for an application that translates user-initiated instructions into another language in real time and returns the translation with appropriate emotional nuances. Furthermore, please describe the hardware and software used, and the specific processing steps."

[0350] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0351] Step 1:

[0352] Acquiring audio

[0353] When a user speaks, the device's microphone captures the audio. The input is the user's spoken audio data, and the output is the audio data converted into a digital format. This audio data is converted into a digital format by an ADC (Audio-Digital Converter).

[0354] Step 2:

[0355] Sending audio data

[0356] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. The input is the audio data converted to digital format, and the output is the audio data sent to the server. This transmission is performed using a secure communication protocol such as SSL.

[0357] Step 3:

[0358] Analysis of audio data

[0359] The server passes the received audio data to the analysis device, which then converts it into text using a speech recognition engine. The input is the audio data sent to the server, and the output is text data. Specifically, the server's speech recognition engine analyzes the audio data as text data.

[0360] Step 4:

[0361] emotion recognition

[0362] The server simultaneously passes the audio data to the emotion recognition system, which uses an emotion engine to analyze the emotional state from the audio data. The inputs are text data and audio data, and the output is emotion data. The emotion engine identifies the emotional state from the text data and audio data.

[0363] Step 5:

[0364] translation

[0365] The server passes the analyzed text data and recognized sentiment data to the translation system, which translates it into the target language in real time. The input is text data and sentiment data, and the output is the translated text data in the target language. In this translation process, the sentiment data is used to adjust the tone and nuances.

[0366] Step 6:

[0367] Speech synthesis

[0368] The server passes the translated text data and sentiment data to the speech synthesis system, which uses the speech synthesis module to generate speech data. The input is the translated text data and sentiment data, and the output is the synthesized speech data. The speech synthesis module uses the sentiment data to adjust the tone and intonation of the speech.

[0369] Step 7:

[0370] Sending audio data

[0371] The server transmits the generated audio data to the terminal via a communication module. The input is synthesized speech data, and the output is the audio data transmitted to the terminal. A secure communication protocol is used to transmit the audio data to the terminal.

[0372] Step 8:

[0373] Audio playback

[0374] The terminal receives audio data from the server and sends it to the playback device (speaker or earphones). The input is the received audio data, and the output is the audio provided to the user. This allows the user to understand speech in different languages, including the emotions conveyed. The audio is then played back through the terminal's speaker or earphones.

[0375] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0376] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0377] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0378] [Second Embodiment]

[0379] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0380] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0381] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0382] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0383] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0384] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0385] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0386] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0387] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0388] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0389] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0390] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0391] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. The configuration of this system and the processing of its program will be described in detail below.

[0392] System Configuration

[0393] This system primarily consists of wearable devices (such as glasses-type devices or earphone-type devices) and a server in the cloud. Users can wear the device and communicate via voice. Each device and server includes the following functional modules.

[0394] 1. The terminal is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the sound data into a digital format, a communication module to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0395] 2. The server includes a speech recognition module that analyzes the received audio data and identifies the language, a translation module that translates the identified language into the target language, a speech synthesis module that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0396] Program processing

[0397] The following is a natural language explanation of the program's processing logic.

[0398] Voice acquisition and transmission

[0399] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transfer is performed using a secure communication protocol.

[0400] Audio data analysis and translation

[0401] The server passes the received audio data to the analysis module. The analysis module uses speech recognition technology (e.g., a speech recognition API) to convert the audio to text. Next, it identifies the language used from the analyzed text. The text in the identified language is passed to the translation module, which translates it into the target language in real time. The translation technology used in this process is, for example, a machine translation engine (translation API, etc.).

[0402] Translated speech synthesis and transmission

[0403] The server passes the translated text to the speech synthesis module. The speech synthesis module converts the text in the target language into speech data (using a speech synthesis API, etc.). The generated speech data is sent to the terminal via the communication module.

[0404] Audio playback

[0405] The device receives audio data from the server and provides it to the user through a playback device (speaker or earphones). Through this process, the user can understand speech in different languages ​​in real time.

[0406] Specific usage examples

[0407] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0408] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics."

[0409] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0410] 3. The server analyzes the received audio data, converts it into Japanese text, and then uses a translation module to translate it into English as "Today's lecture is about the basics of quantum mechanics".

[0411] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0412] 5. The device plays the received English audio data, and User A understands it in real time.

[0413] This enables users who speak different languages ​​to communicate in real time, facilitating smooth lectures even in multinational classes in educational settings.

[0414] The following describes the processing flow.

[0415] Step 1:

[0416] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0417] Step 2:

[0418] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0419] Step 3:

[0420] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0421] Step 4:

[0422] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0423] Step 5:

[0424] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0425] Step 6:

[0426] The server passes text data to the translation module, which then translates it into the target language in real time.

[0427] Step 7:

[0428] The server passes the translated text data to the speech synthesis module, which then generates synthesized speech.

[0429] Step 8:

[0430] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0431] Step 9:

[0432] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0433] Step 10:

[0434] Users can listen to the played audio and understand speech in different languages ​​in real time.

[0435] (Example 1)

[0436] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0437] Current real-time speech translation systems suffer from insufficient security in the communication protocols they use, potentially leading to data eavesdropping during transmission. Furthermore, the accuracy of speech recognition and translation is often low, failing to meet user expectations. Additionally, complex and time-consuming intermediate processing can compromise real-time capabilities.

[0438] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0439] In this invention, the server includes means for analyzing voice data, means for translating, and means for synthesizing voice data. This enables efficient and secure execution of each process of analysis, translation, and voice synthesis, and provides real-time, highly accurate voice translation. Furthermore, by using an optimized communication protocol, the security of data transfer can be improved and user privacy can be protected.

[0440] "Acquisition means" refers to a device used to capture the voice spoken by the user, such as a microphone.

[0441] "Conversion means" refers to a device used to convert captured analog audio data into a digital format, specifically such as an ADC (Analog to Digital Converter).

[0442] "Transmission means" refers to devices and protocols used to send digital audio data to a server in the cloud, such as Wi-Fi modules or secure communication protocols like SSL / TLS.

[0443] "Analysis means" refers to technologies that convert audio data received by a server into text data and identify the language being used, specifically referring to speech recognition engines and speech recognition APIs.

[0444] "Translation methods" refer to technologies used to translate text data obtained through analysis into different languages, such as machine translation engines and translation APIs.

[0445] "Speech synthesis means" refers to technologies for converting translated text into speech data, such as speech synthesis APIs.

[0446] "Playback means" refers to a device that outputs the generated audio data so that the user can hear it, such as a speaker or earphones.

[0447] "Secure communication methods" refer to protocols and technologies for securely transferring voice data, including encrypted communication technologies such as SSL / TLS.

[0448] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. Specific embodiments of this system are described below.

[0449] System Configuration

[0450] This system primarily consists of wearable devices (such as glasses or earphones) and a server in the cloud. Users wear the device and can communicate using voice.

[0451] Device configuration

[0452] The device is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the audio data into a digital format, a communication module (such as a Wi-Fi module) to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0453] Server Configuration

[0454] The server includes a speech recognition module (e.g., a speech recognition API) that analyzes the received audio data and identifies the language, a translation module (e.g., a machine translation engine) that translates the identified language into the target language, a speech synthesis module (e.g., a speech synthesis API) that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0455] Specific example of processing

[0456] As a concrete example, consider a situation where a foreign language learner understands a conversation in real time.

[0457] Situation:

[0458] A scenario where an English-speaking tourist (User A) communicates with a Japanese-speaking guide (User B) at a tourist destination in Japan.

[0459] Situation handling procedure

[0460] 1. User B says in Japanese, "This is Senso-ji Temple."

[0461] 2. The terminal acquires this utterance and converts it into digital speech data using an ADC.

[0462] 3. The terminal sends the converted digital audio data to the server via the communication module. Secure protocols such as SSL / TLS are used for communication.

[0463] 4. The server passes the received audio data to the analysis module (speech recognition API) and converts it into the Japanese text "This is Senso-ji Temple."

[0464] 5. The server passes the analyzed text to a translation module (machine translation engine), which translates it into the English text "This is Sensoji Temple".

[0465] 6. The server passes the translated text to the speech synthesis module (speech synthesis API) and converts it into the English speech data "Hello".

[0466] 7. The server sends the generated English voice data to the terminal via the communication module. A secure protocol is used for communication again.

[0467] 8. The device plays the received English audio data through a playback device (earphones or speaker), and user A understands it as "This is Sensoji Temple."

[0468] This allows users to understand speech in different languages ​​in real time.

[0469] Example prompts for the generative AI model to be used

[0470] The following prompt can be entered using a generative AI model.

[0471] Please describe the specific processing flow of a system that, when a Japanese-speaking user says "Konnichiwa" (こんにちは), translates it into English in real time and provides the user with "Hello." Please also specify the technologies and devices used at each step.

[0472] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0473] Step 1:

[0474] The user says "Hello." The device's microphone captures this audio. The input is an analog audio signal, and the microphone captures it, resulting in an analog audio signal as output.

[0475] Step 2:

[0476] The terminal converts the analog audio signal acquired from the microphone into digital audio data using an ADC (Analog to Digital Converter). The input is an analog audio signal, and the output after conversion is digital audio data.

[0477] Step 3:

[0478] The terminal transmits the converted digital audio data to a server in the cloud via a communication module (e.g., a Wi-Fi module). Secure communication protocols such as SSL / TLS are used during this process. The input is digital audio data, and the output is the digital audio data transmitted to the server.

[0479] Step 4:

[0480] The server passes the received digital audio data to the speech recognition module (speech recognition API). The speech recognition module analyzes the input digital audio data and converts it into text data. Specifically, the output is the Japanese text data "こんにちは" (konnichiwa).

[0481] Step 5:

[0482] The server passes the text data obtained from the speech recognition module to the translation module (machine translation engine). The translation module translates the input Japanese text data "こんにちは" into the English text data "Hello". The output is the English text data "Hello".

[0483] Step 6:

[0484] The server passes the translated English text data "Hello" to the speech synthesis module (speech synthesis API). The speech synthesis module parses the input text data and converts it into speech data. The output is the English speech data "Hello".

[0485] Step 7:

[0486] The server sends the generated English audio data to the terminal via a communication module. Secure protocols such as SSL / TLS are again used for communication. The input is English audio data, and the output is the English audio data sent to the terminal.

[0487] Step 8:

[0488] The device receives English audio data and plays it back through a playback device (earphones or speaker) for the user to hear. The input is English audio data, and the output is the audio that the user can hear.

[0489] This allows users to understand speech in different languages ​​in real time.

[0490] (Application Example 1)

[0491] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0492] In factories, when multinational workers work together, language barriers can hinder smooth communication, leading to decreased production efficiency. Furthermore, important instructions and urgent announcements may not be accurately conveyed, potentially impacting safety. To resolve these issues and improve work efficiency and safety, measures to overcome language differences are necessary.

[0493] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0494] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio data into a digital format, and a transmission means for sending the converted audio data to the server. This enables a system, including a robot, for real-time communication with workers who speak multiple languages ​​in a factory environment.

[0495] "Acquisition means" refers to the devices or systems used to acquire sound.

[0496] "Conversion means" refers to a device or function for converting acquired audio data into a digital format.

[0497] "Transmission means" refers to communication functions or modules for sending the converted audio data to the server.

[0498] "Analysis means" refers to technologies and systems for analyzing transmitted audio data and identifying the language being used.

[0499] "Translation means" refers to technologies and devices for translating audio data of a specified language into a target language.

[0500] "Speech synthesis means" refers to technologies and systems for converting translated text into speech data.

[0501] "Playback means" refers to devices or functions for playing back audio data transmitted from a server.

[0502] A "robot" is a machine used in a factory environment to communicate in real time with multinational workers.

[0503] This invention relates to a multilingual voice translation system for facilitating communication among multinational workers in a factory. The embodiments for carrying out the invention are as follows:

[0504] System Configuration

[0505] This system primarily consists of wearable terminals (e.g., devices mounted on robots or work helmets) and cloud-based servers. Users can communicate via voice through the terminals. Each device and server includes the following functional modules.

[0506] terminal

[0507] Acquisition method: A microphone for acquiring sound. This microphone has a noise reduction function.

[0508] Conversion method: Analog to Digital Converter (ADC) that converts audio data acquired by a microphone into a digital format.

[0509] Transmission method: A communication module (e.g., Wi-Fi or 5G module) for securely transmitting the converted audio data to a server in the cloud.

[0510] Playback method: A high-quality speaker for playing back translated audio data received from the server.

[0511] server

[0512] Analysis method: A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data received by the server and identify the language being used.

[0513] Translation method: A machine translation engine (e.g., Microsoft Translator) for translating text in a specified language into a target language in real time.

[0514] Speech synthesis means: Speech synthesis technology for converting translated text into speech data (e.g., Amazon Polly).

[0515] Transmission method: A communication module for transmitting synthesized voice data to a terminal.

[0516] Explanation of program processing

[0517] The server receives the acquired audio data and passes it to a speech recognition engine to convert it into text. The analyzed text is then translated into the target language by a machine translation engine. The translated text is converted back into audio data using speech synthesis technology, and this audio data is transmitted to the terminal. The transmitted audio data is provided to the worker via a playback device, enabling real-time communication among multinational workers.

[0518] Specific example

[0519] For example, consider a scenario in a factory where a Japanese-speaking worker (User A) collaborates with a worker who only speaks English (User B).

[0520] 1. User A says in Japanese, "Please tell me the next steps."

[0521] 2. The terminal acquires this audio, converts it into digital audio data using an ADC, and sends it to the server.

[0522] 3. The server analyzes the received audio data, converts it to Japanese text, and then uses a translation module to translate it into English as "Please explain the next work step".

[0523] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0524] 5. The device plays the received English audio data, and user B understands it in real time.

[0525] Example of a prompt

[0526] Please capture the audio of a factory worker saying "Please tell me the next step" in Japanese. Translate it into English and convert it into audio data that can be understood by an English-speaking technician, then play it back.

[0527] This facilitates smoother communication among multinational workers within the factory, resulting in an efficient and safe working environment.

[0528] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0529] Step 1:

[0530] The user speaks. The input is the user's voice, and the specific action is a factory worker saying "Please tell me the next work procedure" in Japanese. The output is an acoustic signal in the form of a speech waveform.

[0531] Step 2:

[0532] The device captures the user's voice. The input is an acoustic signal, and the specific operation involves a microphone acquiring the voice. The output is analog audio data.

[0533] Step 3:

[0534] The terminal's conversion mechanism converts analog audio data into a digital format. The input is analog audio data, and specifically, the ADC (Analog to Digital Converter) converts the signal into a digital format. The output is digital audio data.

[0535] Step 4:

[0536] The terminal's transmission method sends digital audio data to the server. The input is digital audio data, and specifically, the communication module (Wi-Fi or 5G) sends the data to a server in the cloud. The output is the digital audio data sent to the server.

[0537] Step 5:

[0538] The server's analysis tool analyzes the received audio data and converts it to text using a speech recognition engine. The input is the received digital audio data, and the specific operation is for the speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio into text data. The output is the analyzed text data.

[0539] Step 6:

[0540] The server's analysis tool identifies the language used in the text data. The input is the analyzed text data, and the specific operation is for the language model to identify the language from the text. The output is the identified language information.

[0541] Step 7:

[0542] The server's translation method translates text data in a specified language into the target language. The input consists of specified language information and text data; the specific operation involves a machine translation engine (e.g., Microsoft Translator) translating the text into the target language in real time. The output is the translated text data.

[0543] Step 8:

[0544] The server's speech synthesis system converts translated text data into speech data. The input is translated text data, and the specific operation involves speech synthesis technology (e.g., Amazon Polly) converting the text in the target language into speech data. The output is synthesized speech data.

[0545] Step 9:

[0546] The server's transmission mechanism sends synthesized audio data to the terminal. The input is synthesized audio data, and the specific operation is that the communication module sends the audio data to the terminal. The output is the audio data received by the terminal.

[0547] Step 10:

[0548] The device's playback mechanism plays the received audio data. The input is the audio data received by the device, and the specific operation is for the speaker to play the audio and provide it to the user. The output is audio in a language the user understands.

[0549] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0550] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. The configuration of this system and the processing of its program will be described in detail below.

[0551] System Configuration

[0552] This system consists of wearable devices (e.g., glasses-type devices or earphone-type devices) and a server in the cloud. By wearing the device and communicating via voice, the system enables the following functions. Its main components are as follows:

[0553] 1. The terminal is equipped with a microphone to capture user speech, an ADC (Analog to Digital Converter) to convert audio data into a digital format, a communication module to send data to the server, and a speaker or earphones to play back the received translated audio.

[0554] 2. The server includes a speech recognition module that analyzes received audio data and identifies the language, a translation module that translates the identified language into the target language, and a speech synthesis module that converts the translated text into speech. It also includes an emotion engine that recognizes the user's emotions and has the ability to adjust the tone and nuances of the voice based on those emotions.

[0555] 3. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, anger, surprise, etc.) from the voice data and uses this to influence the translation and speech synthesis processes.

[0556] Program processing

[0557] The following provides a detailed explanation of how this system's program processes information.

[0558] Voice acquisition and transmission

[0559] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol.

[0560] Audio data analysis and emotion recognition

[0561] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio into text. Simultaneously, an emotion engine operates to analyze the user's emotional state from the audio data.

[0562] Language identification and translation

[0563] The server identifies the language used from the analyzed text. Next, the text data and recognized sentiment data are passed to a translation module, which translates them into the target language in real time. This translation process uses the sentiment data to appropriately adjust tone and nuance.

[0564] Translated speech synthesis and transmission

[0565] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the terminal via the communication module.

[0566] Audio playback

[0567] The device receives audio data from the server and sends it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[0568] Specific usage examples

[0569] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0570] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0571] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0572] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0573] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0574] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0575] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0576] This enables users who speak different languages ​​to communicate in real time, including conveying emotions and nuances. It can be effectively used in educational settings and international business conferences.

[0577] The following describes the processing flow.

[0578] Step 1:

[0579] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0580] Step 2:

[0581] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0582] Step 3:

[0583] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0584] Step 4:

[0585] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0586] Step 5:

[0587] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0588] Step 6:

[0589] The server passes audio data to the emotion engine along with text data in the identified language to recognize the user's emotions. Specifically, it analyzes features such as tone, pace, and intonation of the voice to identify the emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0590] Step 7:

[0591] The server passes the text data and recognized sentiment data to the translation module, which then translates it into the target language. During this process, the translation module takes the sentiment data into account and generates translated text with appropriate tone and nuances.

[0592] Step 8:

[0593] The server passes the translated text data and sentiment data to the speech synthesis module to generate synthesized speech. Based on the sentiment data, the speech synthesis module adjusts the tone and intonation of the speech to produce natural-sounding speech.

[0594] Step 9:

[0595] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0596] Step 10:

[0597] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0598] Step 11:

[0599] Users can listen to the played audio and understand the nuances of speech and emotions in different languages ​​in real time.

[0600] (Example 2)

[0601] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0602] Conventional speech translation systems have a problem in that they fail to properly convey emotions and tone when translating user speech into other languages. Furthermore, because they do not consider the user's emotions during the speech recognition and translation process, nuances of communication are often lost. As a result, especially in educational and business settings, translated content is often not accurately conveyed, leading to misunderstandings.

[0603] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means for analyzing voice data and identifying language, an emotion recognition means for recognizing emotions from voice data, and an emotion adjustment means for adjusting translation and speech synthesis based on the recognized emotions. This enables real-time multilingual translation that reflects the user's emotions and nuances.

[0604] "Acquisition means" refers to a device or function used to capture a user's speech.

[0605] "Conversion means" refers to a device or function that converts acquired audio data from analog format to digital format.

[0606] "Transmission means" refers to a device or function for transmitting audio data converted into a digital format to a server.

[0607] "Analysis means" refers to a device or function that analyzes transmitted audio data and identifies the language.

[0608] "Translation means" refers to a device or function that translates audio data of a specified language into a target language.

[0609] "Speech synthesis means" refers to a device or function that converts translated text into speech data.

[0610] "Emotion recognition means" refers to a device or function that recognizes a user's emotions from audio data.

[0611] "Emotion adjustment means" refers to a device or function that adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0612] "Playback means" refers to a device or function for playing back transmitted audio data to the user.

[0613] A "server" is a central computer system used to perform voice data analysis, translation, speech synthesis, and emotion recognition.

[0614] A "terminal" is a device worn by a user that has the function of acquiring and playing back audio.

[0615] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. This system consists of a wearable device (e.g., glasses-type device or earphone-type device) and a server in the cloud. By wearing the device and communicating via voice, the system achieves the following functions:

[0616] System Configuration

[0617] The main components of this system are as follows:

[0618] terminal

[0619] Acquisition method: Equipped with a microphone for capturing user speech.

[0620] Conversion method: Includes an ADC (Analog to Digital Converter) that converts audio into a digital format.

[0621] Transmission method: It is equipped with a communication module that transmits the acquired digital audio data to a server in the cloud.

[0622] Playback method: It has a speaker or earphones that play back the translated audio received from the server.

[0623] server

[0624] Analysis method: The received audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API or Microsoft Azure Speech Service).

[0625] Translation method: To translate text data into the target language, we use APIs such as Google Translate or DeepL.

[0626] Speech synthesis method: Amazon Polly or Google Cloud Text-to-Speech API are used to convert translated text into speech data.

[0627] Emotion recognition means: An emotion engine is used to recognize the user's emotions from voice data and evaluate tone and nuance.

[0628] Emotion adjustment mechanism: Adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0629] Examples of usage and prompt statements

[0630] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0631] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0632] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0633] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0634] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0635] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0636] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0637] Using a generative AI model, you can utilize the following example prompts:

[0638] Example prompt

[0639] In a university lecture, a professor who speaks Japanese said, "Today's lecture will be about the fundamentals of quantum mechanics," with a sense of questioning. Please explain in detail how to translate this content into English in real time, while preserving the sense of questioning, and provide it to students who only speak English.

[0640] This prompt statement is useful for specifically describing the system's behavior in particular scenarios.

[0641] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0642] Step 1:

[0643] When a user speaks, the device's microphone captures the audio data. The captured audio is in analog format and is converted to digital format by an ADC (Analog to Digital Converter). The input is the user's analog voice, and the output is digital audio data. This data conversion provides digital audio data that can be used in subsequent processing steps.

[0644] Step 2:

[0645] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. This transmission uses a secure communication protocol such as HTTPS. The input is digital audio data, and the output is the transmitted audio data. The communication module divides the data into packets and transmits them to ensure they reach the server reliably.

[0646] Step 3:

[0647] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to convert the audio to text. The input is the digital audio data sent to the server, and the output is text data. The speech recognition engine performs phoneme and sound waveform analysis to generate appropriate text data.

[0648] Step 4:

[0649] The server's emotion recognition engine analyzes audio data in parallel to identify the user's emotional state. It evaluates the tone, pitch, and speed of the speech to recognize emotional states (such as joy, sadness, anger, or surprise). The input is digital audio data, and the output is emotion data. Emotion recognition is performed using machine learning algorithms.

[0650] Step 5:

[0651] The server passes the analyzed text data and sentiment data to the translation module. The translation module (e.g., Google Translate API or DeepL API) translates the text into the target language in real time. The input is text data and sentiment data, and the output is the translated text. During translation, the tone and nuances are also appropriately adjusted, taking sentiment data into account.

[0652] Step 6:

[0653] The server passes the translated text and sentiment data to a speech synthesis module. The speech synthesis module (e.g., Amazon Polly or Google Cloud Text-to-Speech API) converts the text into speech. The input is the translated text and sentiment data, and the output is the generated synthesized speech data. This process adjusts the tone and intonation of the speech based on the sentiment data.

[0654] Step 7:

[0655] The server sends the generated synthesized speech data to the terminal via a communication module. A secure protocol (e.g., HTTPS) is used for communication to ensure data integrity and confidentiality. The input is the synthesized speech data, and the output is the transmitted speech data.

[0656] Step 8:

[0657] The terminal receives audio data from the server, sends it to a playback device (speaker or earphones), and provides it to the user as audio. The input is the received audio data, and the output is the played audio. This allows the user to understand speech in different languages ​​in real time, along with the emotions conveyed.

[0658] (Application Example 2)

[0659] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0660] In global manufacturing environments, language differences pose a significant obstacle when multinational staff collaborate. Furthermore, accurately conveying not only language but also the emotions and nuances of speech is crucial, but current systems fail to adequately address this. Therefore, there is a need for real-time multilingual translation and emotion recognition simultaneously to facilitate smooth communication on the factory floor.

[0661] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes emotion recognition means for analyzing emotional states from voice data, means for adjusting the tone and nuances of translated text based on the analyzed emotion data, analysis means for converting acquired voice data into text using a voice recognition engine and analyzing emotional states using an emotion engine, and translation means using a translation module that translates text data and emotion data in real time and adjusts the tone and nuances based on the emotion data. This enables accurate multilingual translation and emotion recognition in real time.

[0662] "Audio data" refers to data that represents audio information in a digital format.

[0663] A "server" is a computer system that provides data processing and storage services over a network.

[0664] "Emotion recognition means" refers to functions or devices that analyze and recognize the speaker's emotions from audio data.

[0665] "Emotional data" refers to data that indicates the state of emotions as analyzed by emotion recognition tools.

[0666] A "translation tool" is a function or device used to convert text or audio expressed in one language into another language.

[0667] A "translation module" is a software module used to translate text data or audio into other languages.

[0668] A "speech recognition engine" is a software engine that converts speech data into text.

[0669] "Tone and nuance" refers to the emotional expression and subtle meanings conveyed by speech or text.

[0670] "Analysis means" refers to functions or devices for analyzing acquired audio data and extracting specific information or patterns.

[0671] "Acquisition means" refers to functions or devices for collecting or capturing data such as audio and video.

[0672] "Playback means" refers to functions or devices for playing back data on a terminal.

[0673] Modes for carrying out the invention

[0674] This invention relates to a system that translates user speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. As an example of its application, this system will be used to realize a smart glasses application that facilitates communication in global manufacturing environments. The system configuration and detailed program processing of this invention will be described below.

[0675] System Configuration

[0676] This system consists of smart glasses-type devices and a server in the cloud. When the user wears the smart glasses and communicates via voice, the system enables the following functions.

[0677] Device (smart glasses)

[0678] Acquisition method: Includes a microphone for acquiring sound.

[0679] Conversion means: Includes an ADC (Analog to Digital Converter) that converts the acquired audio data into a digital format.

[0680] Transmission means: Includes a communication module that transmits the converted audio data to a server.

[0681] Playback means: Includes speakers or earphones that play audio data transmitted from the server.

[0682] server

[0683] Analysis method: The transmitted audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[0684] Emotion recognition method: An emotion engine (e.g., IBM Watson Tone Analyzer) is used to analyze the emotional state from voice data.

[0685] Translation method: Text data and sentiment data are translated in real time and translated into the target language using a translation module (e.g., Microsoft Translator API).

[0686] Speech synthesis method: Based on translated text and sentiment data, a speech synthesis module (e.g., Amazon Polly) converts it into speech.

[0687] Program processing

[0688] Voice acquisition and transmission

[0689] When a user speaks, the microphone in the smart glasses captures their voice. The voice data is converted to a digital format by an ADC and sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol such as SSL.

[0690] Audio data analysis and emotion recognition

[0691] The audio data received by the server is converted into text in real time by a speech recognition engine. In parallel, an emotion recognition system operates to analyze the user's emotional state from the audio data.

[0692] Language identification and translation

[0693] The server identifies the language used from the analyzed text, passes the text data and recognized sentiment data to a translation module, and translates it into the target language in real time. This translation process uses sentiment data to appropriately adjust tone and nuances.

[0694] Translated speech synthesis and transmission

[0695] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the smart glasses via the communication module.

[0696] Audio playback

[0697] The smart glasses receive audio data from a server and send it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[0698] Specific example

[0699] Let's consider an example of use in a global factory setting. Suppose a team leader gives the instruction in Japanese: "Please install the next part." If the leader speaks in a calm tone, the smart glasses system will operate as follows:

[0700] 1. The reader's speech is captured by the microphone in the smart glasses and converted into a digital format.

[0701] 2. The digital audio data is sent to a cloud server and converted into text, "Please install the following parts," by a speech recognition engine.

[0702] 3. The emotion recognition system analyzes the calm emotion, and the translation module translates it into English as "Please install the next part." The tone and nuances are also adjusted based on the emotion data.

[0703] 4. The translated text is converted into English speech by a speech synthesis module and sent to the smart glasses.

[0704] 5. Translated English instructions are played through the smart glasses' speaker, allowing staff of other nationalities to understand them.

[0705] Example of a prompt

[0706] "Please provide prototype code for an application that translates user-initiated instructions into another language in real time and returns the translation with appropriate emotional nuances. Furthermore, please describe the hardware and software used, and the specific processing steps."

[0707] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0708] Step 1:

[0709] Acquiring audio

[0710] When a user speaks, the device's microphone captures the audio. The input is the user's spoken audio data, and the output is the audio data converted into a digital format. This audio data is converted into a digital format by an ADC (Audio-Digital Converter).

[0711] Step 2:

[0712] Sending audio data

[0713] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. The input is the audio data converted to digital format, and the output is the audio data sent to the server. This transmission is performed using a secure communication protocol such as SSL.

[0714] Step 3:

[0715] Analysis of audio data

[0716] The server passes the received audio data to the analysis device, which then converts it into text using a speech recognition engine. The input is the audio data sent to the server, and the output is text data. Specifically, the server's speech recognition engine analyzes the audio data as text data.

[0717] Step 4:

[0718] emotion recognition

[0719] The server simultaneously passes the audio data to the emotion recognition system, which uses an emotion engine to analyze the emotional state from the audio data. The inputs are text data and audio data, and the output is emotion data. The emotion engine identifies the emotional state from the text data and audio data.

[0720] Step 5:

[0721] translation

[0722] The server passes the analyzed text data and recognized sentiment data to the translation system, which translates it into the target language in real time. The input is text data and sentiment data, and the output is the translated text data in the target language. This translation process uses sentiment data to adjust tone and nuance.

[0723] Step 6:

[0724] Speech synthesis

[0725] The server passes the translated text data and sentiment data to the speech synthesis system, which uses the speech synthesis module to generate speech data. The input is the translated text data and sentiment data, and the output is the synthesized speech data. The speech synthesis module uses the sentiment data to adjust the tone and intonation of the speech.

[0726] Step 7:

[0727] Sending audio data

[0728] The server transmits the generated audio data to the terminal via a communication module. The input is synthesized speech data, and the output is the audio data transmitted to the terminal. A secure communication protocol is used to transmit the audio data to the terminal.

[0729] Step 8:

[0730] Audio playback

[0731] The terminal receives audio data from the server and sends it to the playback device (speaker or earphones). The input is the received audio data, and the output is the audio provided to the user. This allows the user to understand speech in different languages, including the emotions conveyed. The audio is then played back through the terminal's speaker or earphones.

[0732] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0733] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0734] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0735] [Third Embodiment]

[0736] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0737] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0738] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0739] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0740] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0741] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0742] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0743] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0744] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0745] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0746] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0747] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0748] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. The configuration of this system and the processing of its program will be described in detail below.

[0749] System Configuration

[0750] This system primarily consists of wearable devices (such as glasses-type devices or earphone-type devices) and a server in the cloud. Users can wear the device and communicate via voice. Each device and server includes the following functional modules.

[0751] 1. The terminal is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the sound data into a digital format, a communication module to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0752] 2. The server includes a speech recognition module that analyzes the received audio data and identifies the language, a translation module that translates the identified language into the target language, a speech synthesis module that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0753] Program processing

[0754] The following is a natural language explanation of the program's processing logic.

[0755] Voice acquisition and transmission

[0756] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transfer is performed using a secure communication protocol.

[0757] Audio data analysis and translation

[0758] The server passes the received audio data to the analysis module. The analysis module uses speech recognition technology (e.g., a speech recognition API) to convert the audio to text. Next, it identifies the language used from the analyzed text. The text in the identified language is passed to the translation module, which translates it into the target language in real time. The translation technology used in this process is, for example, a machine translation engine (translation API, etc.).

[0759] Translated speech synthesis and transmission

[0760] The server passes the translated text to the speech synthesis module. The speech synthesis module converts the text in the target language into speech data (using a speech synthesis API, etc.). The generated speech data is sent to the terminal via the communication module.

[0761] Audio playback

[0762] The device receives audio data from the server and provides it to the user through a playback device (speaker or earphones). Through this process, the user can understand speech in different languages ​​in real time.

[0763] Specific usage examples

[0764] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0765] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics."

[0766] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0767] 3. The server analyzes the received audio data, converts it into Japanese text, and then uses a translation module to translate it into English as "Today's lecture is about the basics of quantum mechanics".

[0768] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0769] 5. The device plays the received English audio data, and User A understands it in real time.

[0770] This enables users who speak different languages ​​to communicate in real time, facilitating smooth lectures even in multinational classes in educational settings.

[0771] The following describes the processing flow.

[0772] Step 1:

[0773] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0774] Step 2:

[0775] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0776] Step 3:

[0777] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0778] Step 4:

[0779] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0780] Step 5:

[0781] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0782] Step 6:

[0783] The server passes text data to the translation module, which then translates it into the target language in real time.

[0784] Step 7:

[0785] The server passes the translated text data to the speech synthesis module, which then generates synthesized speech.

[0786] Step 8:

[0787] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0788] Step 9:

[0789] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0790] Step 10:

[0791] Users can listen to the played audio and understand speech in different languages ​​in real time.

[0792] (Example 1)

[0793] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0794] Current real-time speech translation systems suffer from insufficient security in the communication protocols they use, potentially leading to data eavesdropping during transmission. Furthermore, the accuracy of speech recognition and translation is often low, failing to meet user expectations. Additionally, complex and time-consuming intermediate processing can compromise real-time capabilities.

[0795] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0796] In this invention, the server includes means for analyzing voice data, means for translating, and means for synthesizing voice data. This enables efficient and secure execution of each process of analysis, translation, and voice synthesis, and provides real-time, highly accurate voice translation. Furthermore, by using an optimized communication protocol, the security of data transfer can be improved and user privacy can be protected.

[0797] "Acquisition means" refers to a device used to capture the voice spoken by the user, such as a microphone.

[0798] "Conversion means" refers to a device used to convert captured analog audio data into a digital format, specifically such as an ADC (Analog to Digital Converter).

[0799] "Transmission means" refers to devices and protocols used to send digital audio data to a server in the cloud, such as Wi-Fi modules or secure communication protocols like SSL / TLS.

[0800] "Analysis means" refers to technologies that convert audio data received by a server into text data and identify the language being used, specifically referring to speech recognition engines and speech recognition APIs.

[0801] "Translation methods" refer to technologies used to translate text data obtained through analysis into different languages, such as machine translation engines and translation APIs.

[0802] "Speech synthesis means" refers to technologies for converting translated text into speech data, such as speech synthesis APIs.

[0803] "Playback means" refers to a device that outputs the generated audio data so that the user can hear it, such as a speaker or earphones.

[0804] "Secure communication methods" refer to protocols and technologies for securely transferring voice data, including encrypted communication technologies such as SSL / TLS.

[0805] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. Specific embodiments of this system are described below.

[0806] System Configuration

[0807] This system primarily consists of wearable devices (such as glasses or earphones) and a server in the cloud. Users wear the device and can communicate using voice.

[0808] Device configuration

[0809] The device is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the audio data into a digital format, a communication module (such as a Wi-Fi module) to send the data to the server, and a speaker or earphones to play back the received translated audio.

[0810] Server Configuration

[0811] The server includes a speech recognition module (e.g., a speech recognition API) that analyzes the received audio data and identifies the language, a translation module (e.g., a machine translation engine) that translates the identified language into the target language, a speech synthesis module (e.g., a speech synthesis API) that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[0812] Specific example of processing

[0813] As a concrete example, consider a situation where a foreign language learner understands a conversation in real time.

[0814] Situation:

[0815] A scenario where an English-speaking tourist (User A) communicates with a Japanese-speaking guide (User B) at a tourist destination in Japan.

[0816] Situation handling procedure

[0817] 1. User B says in Japanese, "This is Senso-ji Temple."

[0818] 2. The terminal acquires this utterance and converts it into digital speech data using an ADC.

[0819] 3. The terminal sends the converted digital audio data to the server via the communication module. Secure protocols such as SSL / TLS are used for communication.

[0820] 4. The server passes the received audio data to the analysis module (speech recognition API) and converts it into the Japanese text "This is Senso-ji Temple."

[0821] 5. The server passes the analyzed text to a translation module (machine translation engine), which translates it into the English text "This is Sensoji Temple".

[0822] 6. The server passes the translated text to the speech synthesis module (speech synthesis API) and converts it into the English speech data "Hello".

[0823] 7. The server sends the generated English voice data to the terminal via the communication module. A secure protocol is used for communication again.

[0824] 8. The device plays the received English audio data through a playback device (earphones or speaker), and user A understands it as "This is Sensoji Temple."

[0825] This allows users to understand speech in different languages ​​in real time.

[0826] Example prompts for the generative AI model to be used

[0827] The following prompt can be entered using a generative AI model.

[0828] Please describe the specific processing flow of a system that, when a Japanese-speaking user says "Konnichiwa" (こんにちは), translates it into English in real time and provides the user with "Hello." Please also specify the technologies and devices used at each step.

[0829] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0830] Step 1:

[0831] The user says "Hello." The device's microphone captures this audio. The input is an analog audio signal, and the microphone captures it, resulting in an analog audio signal as output.

[0832] Step 2:

[0833] The terminal converts the analog audio signal acquired from the microphone into digital audio data using an ADC (Analog to Digital Converter). The input is an analog audio signal, and the output after conversion is digital audio data.

[0834] Step 3:

[0835] The terminal transmits the converted digital audio data to a server in the cloud via a communication module (e.g., a Wi-Fi module). Secure communication protocols such as SSL / TLS are used during this process. The input is digital audio data, and the output is the digital audio data transmitted to the server.

[0836] Step 4:

[0837] The server passes the received digital audio data to the speech recognition module (speech recognition API). The speech recognition module analyzes the input digital audio data and converts it into text data. Specifically, the output is the Japanese text data "こんにちは" (konnichiwa).

[0838] Step 5:

[0839] The server passes the text data obtained from the speech recognition module to the translation module (machine translation engine). The translation module translates the input Japanese text data "こんにちは" into the English text data "Hello". The output is the English text data "Hello".

[0840] Step 6:

[0841] The server passes the translated English text data "Hello" to the speech synthesis module (speech synthesis API). The speech synthesis module parses the input text data and converts it into speech data. The output is the English speech data "Hello".

[0842] Step 7:

[0843] The server sends the generated English audio data to the terminal via a communication module. Secure protocols such as SSL / TLS are again used for communication. The input is English audio data, and the output is the English audio data sent to the terminal.

[0844] Step 8:

[0845] The device receives English audio data and plays it back through a playback device (earphones or speaker) for the user to hear. The input is English audio data, and the output is the audio that the user can hear.

[0846] This allows users to understand speech in different languages ​​in real time.

[0847] (Application Example 1)

[0848] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0849] In factories, when multinational workers work together, language barriers can hinder smooth communication, leading to decreased production efficiency. Furthermore, important instructions and urgent announcements may not be accurately conveyed, potentially impacting safety. To resolve these issues and improve work efficiency and safety, measures to overcome language differences are necessary.

[0850] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0851] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio data into a digital format, and a transmission means for sending the converted audio data to the server. This enables a system, including a robot, for real-time communication with workers who speak multiple languages ​​in a factory environment.

[0852] "Acquisition means" refers to the devices or systems used to acquire sound.

[0853] "Conversion means" refers to a device or function for converting acquired audio data into a digital format.

[0854] "Transmission means" refers to communication functions or modules for sending the converted audio data to the server.

[0855] "Analysis means" refers to technologies and systems for analyzing transmitted audio data and identifying the language being used.

[0856] "Translation means" refers to technologies and devices for translating audio data of a specified language into a target language.

[0857] "Speech synthesis means" refers to technologies and systems for converting translated text into speech data.

[0858] "Playback means" refers to devices or functions for playing back audio data transmitted from a server.

[0859] A "robot" is a machine used in a factory environment to communicate in real time with multinational workers.

[0860] This invention relates to a multilingual voice translation system for facilitating communication among multinational workers in a factory. The embodiments for carrying out the invention are as follows:

[0861] System Configuration

[0862] This system primarily consists of wearable terminals (e.g., devices mounted on robots or work helmets) and cloud-based servers. Users can communicate via voice through the terminals. Each device and server includes the following functional modules.

[0863] terminal

[0864] Acquisition method: A microphone for acquiring sound. This microphone has a noise reduction function.

[0865] Conversion method: Analog to Digital Converter (ADC) that converts audio data acquired by a microphone into a digital format.

[0866] Transmission method: A communication module (e.g., Wi-Fi or 5G module) for securely transmitting the converted audio data to a server in the cloud.

[0867] Playback method: A high-quality speaker for playing back translated audio data received from the server.

[0868] server

[0869] Analysis method: A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data received by the server and identify the language being used.

[0870] Translation method: A machine translation engine (e.g., Microsoft Translator) for translating text in a specified language into a target language in real time.

[0871] Speech synthesis means: Speech synthesis technology for converting translated text into speech data (e.g., Amazon Polly).

[0872] Transmission method: A communication module for transmitting synthesized voice data to a terminal.

[0873] Explanation of program processing

[0874] The server receives the acquired audio data and passes it to a speech recognition engine to convert it into text. The analyzed text is then translated into the target language by a machine translation engine. The translated text is converted back into audio data using speech synthesis technology, and this audio data is transmitted to the terminal. The transmitted audio data is provided to the worker via a playback device, enabling real-time communication among multinational workers.

[0875] Specific example

[0876] For example, consider a scenario in a factory where a Japanese-speaking worker (User A) collaborates with a worker who only speaks English (User B).

[0877] 1. User A says in Japanese, "Please tell me the next steps."

[0878] 2. The terminal acquires this audio, converts it into digital audio data using an ADC, and sends it to the server.

[0879] 3. The server analyzes the received audio data, converts it to Japanese text, and then uses a translation module to translate it into English as "Please explain the next work step".

[0880] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[0881] 5. The device plays the received English audio data, and user B understands it in real time.

[0882] Example of a prompt

[0883] Please capture the audio of a factory worker saying "Please tell me the next step" in Japanese. Translate it into English and convert it into audio data that can be understood by an English-speaking technician, then play it back.

[0884] This facilitates smoother communication among multinational workers within the factory, resulting in an efficient and safe working environment.

[0885] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0886] Step 1:

[0887] The user speaks. The input is the user's voice, and the specific action is a factory worker saying "Please tell me the next work procedure" in Japanese. The output is an acoustic signal in the form of a speech waveform.

[0888] Step 2:

[0889] The device captures the user's voice. The input is an acoustic signal, and the specific operation involves a microphone acquiring the voice. The output is analog audio data.

[0890] Step 3:

[0891] The terminal's conversion mechanism converts analog audio data into a digital format. The input is analog audio data, and specifically, the ADC (Analog to Digital Converter) converts the signal into a digital format. The output is digital audio data.

[0892] Step 4:

[0893] The terminal's transmission method sends digital audio data to the server. The input is digital audio data, and specifically, the communication module (Wi-Fi or 5G) sends the data to a server in the cloud. The output is the digital audio data sent to the server.

[0894] Step 5:

[0895] The server's analysis tool analyzes the received audio data and converts it to text using a speech recognition engine. The input is the received digital audio data, and the specific operation is for the speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio into text data. The output is the analyzed text data.

[0896] Step 6:

[0897] The server's analysis tool identifies the language used in the text data. The input is the analyzed text data, and the specific operation is for the language model to identify the language from the text. The output is the identified language information.

[0898] Step 7:

[0899] The server's translation method translates text data in a specified language into the target language. The input consists of specified language information and text data; the specific operation involves a machine translation engine (e.g., Microsoft Translator) translating the text into the target language in real time. The output is the translated text data.

[0900] Step 8:

[0901] The server's speech synthesis system converts translated text data into speech data. The input is translated text data, and the specific operation involves speech synthesis technology (e.g., Amazon Polly) converting the text in the target language into speech data. The output is synthesized speech data.

[0902] Step 9:

[0903] The server's transmission mechanism sends synthesized audio data to the terminal. The input is synthesized audio data, and the specific operation is that the communication module sends the audio data to the terminal. The output is the audio data received by the terminal.

[0904] Step 10:

[0905] The device's playback mechanism plays the received audio data. The input is the audio data received by the device, and the specific operation is for the speaker to play the audio and provide it to the user. The output is audio in a language the user understands.

[0906] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0907] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. The configuration of this system and the processing of its program will be described in detail below.

[0908] System Configuration

[0909] This system consists of wearable devices (e.g., glasses-type devices or earphone-type devices) and a server in the cloud. By wearing the device and communicating via voice, the system enables the following functions. Its main components are as follows:

[0910] 1. The terminal is equipped with a microphone to capture user speech, an ADC (Analog to Digital Converter) to convert audio data into a digital format, a communication module to send data to the server, and a speaker or earphones to play back the received translated audio.

[0911] 2. The server includes a speech recognition module that analyzes received audio data and identifies the language, a translation module that translates the identified language into the target language, and a speech synthesis module that converts the translated text into speech. It also includes an emotion engine that recognizes the user's emotions and has the ability to adjust the tone and nuances of the voice based on those emotions.

[0912] 3. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, anger, surprise, etc.) from the voice data and uses this to influence the translation and speech synthesis processes.

[0913] Program processing

[0914] The following provides a detailed explanation of how this system's program processes information.

[0915] Voice acquisition and transmission

[0916] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol.

[0917] Audio data analysis and emotion recognition

[0918] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio into text. Simultaneously, an emotion engine operates to analyze the user's emotional state from the audio data.

[0919] Language identification and translation

[0920] The server identifies the language used from the analyzed text. Next, the text data and recognized sentiment data are passed to a translation module, which translates them into the target language in real time. This translation process uses the sentiment data to appropriately adjust tone and nuance.

[0921] Translated speech synthesis and transmission

[0922] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the terminal via the communication module.

[0923] Audio playback

[0924] The device receives audio data from the server and sends it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[0925] Specific usage examples

[0926] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0927] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0928] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0929] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0930] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0931] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0932] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0933] This enables users who speak different languages ​​to communicate in real time, including conveying emotions and nuances. It can be effectively used in educational settings and international business conferences.

[0934] The following describes the processing flow.

[0935] Step 1:

[0936] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[0937] Step 2:

[0938] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[0939] Step 3:

[0940] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[0941] Step 4:

[0942] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[0943] Step 5:

[0944] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[0945] Step 6:

[0946] The server passes audio data to the emotion engine along with text data in the identified language to recognize the user's emotions. Specifically, it analyzes features such as tone, pace, and intonation of the voice to identify the emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0947] Step 7:

[0948] The server passes the text data and recognized sentiment data to the translation module, which then translates it into the target language. During this process, the translation module takes the sentiment data into account and generates translated text with appropriate tone and nuances.

[0949] Step 8:

[0950] The server passes the translated text data and sentiment data to the speech synthesis module to generate synthesized speech. Based on the sentiment data, the speech synthesis module adjusts the tone and intonation of the speech to produce natural-sounding speech.

[0951] Step 9:

[0952] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[0953] Step 10:

[0954] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[0955] Step 11:

[0956] Users can listen to the played audio and understand the nuances of speech and emotions in different languages ​​in real time.

[0957] (Example 2)

[0958] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0959] Conventional speech translation systems have a problem in that they fail to properly convey emotions and tone when translating user speech into other languages. Furthermore, because they do not consider the user's emotions during the speech recognition and translation process, nuances of communication are often lost. As a result, especially in educational and business settings, translated content is often not accurately conveyed, leading to misunderstandings.

[0960] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means for analyzing voice data and identifying language, an emotion recognition means for recognizing emotions from voice data, and an emotion adjustment means for adjusting translation and speech synthesis based on the recognized emotions. This enables real-time multilingual translation that reflects the user's emotions and nuances.

[0961] "Acquisition means" refers to a device or function used to capture a user's speech.

[0962] "Conversion means" refers to a device or function that converts acquired audio data from analog format to digital format.

[0963] "Transmission means" refers to a device or function for transmitting audio data converted into a digital format to a server.

[0964] "Analysis means" refers to a device or function that analyzes transmitted audio data and identifies the language.

[0965] "Translation means" refers to a device or function that translates audio data of a specified language into a target language.

[0966] "Speech synthesis means" refers to a device or function that converts translated text into speech data.

[0967] "Emotion recognition means" refers to a device or function that recognizes a user's emotions from audio data.

[0968] "Emotion adjustment means" refers to a device or function that adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0969] "Playback means" refers to a device or function for playing back transmitted audio data to the user.

[0970] A "server" is a central computer system used to perform voice data analysis, translation, speech synthesis, and emotion recognition.

[0971] A "terminal" is a device worn by a user that has the function of acquiring and playing back audio.

[0972] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. This system consists of a wearable device (e.g., glasses-type device or earphone-type device) and a server in the cloud. By wearing the device and communicating via voice, the system achieves the following functions:

[0973] System Configuration

[0974] The main components of this system are as follows:

[0975] terminal

[0976] Acquisition method: Equipped with a microphone for capturing user speech.

[0977] Conversion method: Includes an ADC (Analog to Digital Converter) that converts audio into a digital format.

[0978] Transmission method: It is equipped with a communication module that transmits the acquired digital audio data to a server in the cloud.

[0979] Playback method: It has a speaker or earphones that play back the translated audio received from the server.

[0980] server

[0981] Analysis method: The received audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API or Microsoft Azure Speech Service).

[0982] Translation method: To translate text data into the target language, we use APIs such as Google Translate or DeepL.

[0983] Speech synthesis method: Amazon Polly or Google Cloud Text-to-Speech API are used to convert translated text into speech data.

[0984] Emotion recognition means: An emotion engine is used to recognize the user's emotions from voice data and evaluate tone and nuance.

[0985] Emotion adjustment mechanism: Adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[0986] Examples of usage and prompt statements

[0987] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[0988] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[0989] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[0990] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[0991] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[0992] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[0993] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[0994] Using a generative AI model, you can utilize the following example prompts:

[0995] Example prompt

[0996] In a university lecture, a Japanese-speaking professor said, "Today's lecture will be about the fundamentals of quantum mechanics," with a hint of questioning. Please explain in detail how to translate this into English in real time, while preserving the sense of questioning, and provide it to English-only students.

[0997] This prompt statement is useful for specifically describing the system's behavior in particular scenarios.

[0998] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0999] Step 1:

[1000] When a user speaks, the device's microphone captures the audio data. The captured audio is in analog format and is converted to digital format by an ADC (Analog to Digital Converter). The input is the user's analog voice, and the output is digital audio data. This data conversion provides digital audio data that can be used in subsequent processing steps.

[1001] Step 2:

[1002] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. This transmission uses a secure communication protocol such as HTTPS. The input is digital audio data, and the output is the transmitted audio data. The communication module divides the data into packets and transmits them to ensure they reach the server reliably.

[1003] Step 3:

[1004] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to convert the audio to text. The input is the digital audio data sent to the server, and the output is text data. The speech recognition engine performs phoneme and sound waveform analysis to generate appropriate text data.

[1005] Step 4:

[1006] The server's emotion recognition engine analyzes audio data in parallel to identify the user's emotional state. It evaluates the tone, pitch, and speed of the speech to recognize emotional states (such as joy, sadness, anger, or surprise). The input is digital audio data, and the output is emotion data. Emotion recognition is performed using machine learning algorithms.

[1007] Step 5:

[1008] The server passes the analyzed text data and sentiment data to the translation module. The translation module (e.g., Google Translate API or DeepL API) translates the text into the target language in real time. The input is text data and sentiment data, and the output is the translated text. During translation, the tone and nuances are also appropriately adjusted, taking sentiment data into account.

[1009] Step 6:

[1010] The server passes the translated text and sentiment data to a speech synthesis module. The speech synthesis module (e.g., Amazon Polly or Google Cloud Text-to-Speech API) converts the text into speech. The input is the translated text and sentiment data, and the output is the generated synthesized speech data. This process adjusts the tone and intonation of the speech based on the sentiment data.

[1011] Step 7:

[1012] The server sends the generated synthesized speech data to the terminal via a communication module. A secure protocol (e.g., HTTPS) is used for communication to ensure data integrity and confidentiality. The input is the synthesized speech data, and the output is the transmitted speech data.

[1013] Step 8:

[1014] The terminal receives audio data from the server, sends it to a playback device (speaker or earphones), and provides it to the user as audio. The input is the received audio data, and the output is the played audio. This allows the user to understand speech in different languages ​​in real time, along with the emotions conveyed.

[1015] (Application Example 2)

[1016] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1017] In global manufacturing environments, language differences pose a significant obstacle when multinational staff collaborate. Furthermore, accurately conveying not only language but also the emotions and nuances of speech is crucial, but current systems fail to adequately address this. Therefore, there is a need for real-time multilingual translation and emotion recognition simultaneously to facilitate smooth communication on the factory floor.

[1018] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes emotion recognition means for analyzing emotional states from voice data, means for adjusting the tone and nuances of translated text based on the analyzed emotion data, analysis means for converting acquired voice data into text using a voice recognition engine and analyzing emotional states using an emotion engine, and translation means using a translation module that translates text data and emotion data in real time and adjusts the tone and nuances based on the emotion data. This enables accurate multilingual translation and emotion recognition in real time.

[1019] "Audio data" refers to data that represents audio information in a digital format.

[1020] A "server" is a computer system that provides data processing and storage services over a network.

[1021] "Emotion recognition means" refers to functions or devices that analyze and recognize the speaker's emotions from audio data.

[1022] "Emotional data" refers to data that indicates the state of emotions as analyzed by emotion recognition tools.

[1023] A "translation tool" is a function or device used to convert text or audio expressed in one language into another language.

[1024] A "translation module" is a software module used to translate text data or audio into other languages.

[1025] A "speech recognition engine" is a software engine that converts speech data into text.

[1026] "Tone and nuance" refers to the emotional expression and subtle meanings conveyed by speech or text.

[1027] "Analysis means" refers to functions or devices for analyzing acquired audio data and extracting specific information or patterns.

[1028] "Acquisition means" refers to functions or devices for collecting or capturing data such as audio and video.

[1029] "Playback means" refers to functions or devices for playing back data on a terminal.

[1030] Modes for carrying out the invention

[1031] This invention relates to a system that translates user speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. As an example of its application, this system will be used to realize a smart glasses application that facilitates communication in global manufacturing environments. The system configuration and detailed program processing of this invention will be described below.

[1032] System Configuration

[1033] This system consists of smart glasses-type devices and a server in the cloud. When the user wears the smart glasses and communicates via voice, the system enables the following functions.

[1034] Device (smart glasses)

[1035] Acquisition method: Includes a microphone for acquiring sound.

[1036] Conversion means: Includes an ADC (Analog to Digital Converter) that converts the acquired audio data into a digital format.

[1037] Transmission means: Includes a communication module that transmits the converted audio data to a server.

[1038] Playback means: Includes speakers or earphones that play audio data transmitted from the server.

[1039] server

[1040] Analysis method: The transmitted audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[1041] Emotion recognition method: An emotion engine (e.g., IBM Watson Tone Analyzer) is used to analyze the emotional state from voice data.

[1042] Translation method: Text data and sentiment data are translated in real time and translated into the target language using a translation module (e.g., Microsoft Translator API).

[1043] Speech synthesis method: Based on translated text and sentiment data, a speech synthesis module (e.g., Amazon Polly) converts it into speech.

[1044] Program processing

[1045] Voice acquisition and transmission

[1046] When a user speaks, the microphone in the smart glasses captures their voice. The voice data is converted to a digital format by an ADC and sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol such as SSL.

[1047] Audio data analysis and emotion recognition

[1048] The audio data received by the server is converted into text in real time by a speech recognition engine. In parallel, an emotion recognition system operates to analyze the user's emotional state from the audio data.

[1049] Language identification and translation

[1050] The server identifies the language used from the analyzed text, passes the text data and recognized sentiment data to a translation module, and translates it into the target language in real time. This translation process uses sentiment data to appropriately adjust tone and nuances.

[1051] Translated speech synthesis and transmission

[1052] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the smart glasses via the communication module.

[1053] Audio playback

[1054] The smart glasses receive audio data from a server and send it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[1055] Specific example

[1056] Let's consider an example of use in a global factory setting. Suppose a team leader gives the instruction in Japanese: "Please install the next part." If the leader speaks in a calm tone, the smart glasses system will operate as follows:

[1057] 1. The reader's speech is captured by the microphone in the smart glasses and converted into a digital format.

[1058] 2. The digital audio data is sent to a cloud server and converted into text, "Please install the following parts," by a speech recognition engine.

[1059] 3. The emotion recognition system analyzes the calm emotion, and the translation module translates it into English as "Please install the next part." The tone and nuances are also adjusted based on the emotion data.

[1060] 4. The translated text is converted into English speech by a speech synthesis module and sent to the smart glasses.

[1061] 5. Translated English instructions are played through the smart glasses' speaker, allowing staff of other nationalities to understand them.

[1062] Example of a prompt

[1063] "Please provide prototype code for an application that translates user-initiated instructions into another language in real time and returns the translation with appropriate emotional nuances. Furthermore, please describe the hardware and software used, and the specific processing steps."

[1064] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1065] Step 1:

[1066] Acquiring audio

[1067] When a user speaks, the device's microphone captures the audio. The input is the user's spoken audio data, and the output is the audio data converted into a digital format. This audio data is converted into a digital format by an ADC (Audio-Digital Converter).

[1068] Step 2:

[1069] Sending audio data

[1070] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. The input is the audio data converted to digital format, and the output is the audio data sent to the server. This transmission is performed using a secure communication protocol such as SSL.

[1071] Step 3:

[1072] Analysis of audio data

[1073] The server passes the received audio data to the analysis device, which then converts it into text using a speech recognition engine. The input is the audio data sent to the server, and the output is text data. Specifically, the server's speech recognition engine analyzes the audio data as text data.

[1074] Step 4:

[1075] emotion recognition

[1076] The server simultaneously passes the audio data to the emotion recognition system, which uses an emotion engine to analyze the emotional state from the audio data. The inputs are text data and audio data, and the output is emotion data. The emotion engine identifies the emotional state from the text data and audio data.

[1077] Step 5:

[1078] translation

[1079] The server passes the analyzed text data and recognized sentiment data to the translation system, which translates it into the target language in real time. The input is text data and sentiment data, and the output is the translated text data in the target language. This translation process uses sentiment data to adjust tone and nuance.

[1080] Step 6:

[1081] Speech synthesis

[1082] The server passes the translated text data and sentiment data to the speech synthesis system, which uses the speech synthesis module to generate speech data. The input is the translated text data and sentiment data, and the output is the synthesized speech data. The speech synthesis module uses the sentiment data to adjust the tone and intonation of the speech.

[1083] Step 7:

[1084] Sending audio data

[1085] The server transmits the generated audio data to the terminal via a communication module. The input is synthesized speech data, and the output is the audio data transmitted to the terminal. A secure communication protocol is used to transmit the audio data to the terminal.

[1086] Step 8:

[1087] Audio playback

[1088] The terminal receives audio data from the server and sends it to the playback device (speaker or earphones). The input is the received audio data, and the output is the audio provided to the user. This allows the user to understand speech in different languages, including the emotions conveyed. The audio is then played back through the terminal's speaker or earphones.

[1089] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1090] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1091] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1092] [Fourth Embodiment]

[1093] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1094] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1095] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1096] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1097] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1098] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1099] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1100] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1101] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1102] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1103] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1104] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1105] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1106] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. The configuration of this system and the processing of its program will be described in detail below.

[1107] System Configuration

[1108] This system primarily consists of wearable devices (such as glasses-type devices or earphone-type devices) and a server in the cloud. Users can wear the device and communicate via voice. Each device and server includes the following functional modules.

[1109] 1. The terminal is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the sound data into a digital format, a communication module to send the data to the server, and a speaker or earphones to play back the received translated audio.

[1110] 2. The server includes a speech recognition module that analyzes the received audio data and identifies the language, a translation module that translates the identified language into the target language, a speech synthesis module that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[1111] Program processing

[1112] The following is a natural language explanation of the program's processing logic.

[1113] Voice acquisition and transmission

[1114] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transfer is performed using a secure communication protocol.

[1115] Audio data analysis and translation

[1116] The server passes the received audio data to the analysis module. The analysis module uses speech recognition technology (e.g., a speech recognition API) to convert the audio to text. Next, it identifies the language used from the analyzed text. The text in the identified language is passed to the translation module, which translates it into the target language in real time. The translation technology used in this process is, for example, a machine translation engine (translation API, etc.).

[1117] Translated speech synthesis and transmission

[1118] The server passes the translated text to the speech synthesis module. The speech synthesis module converts the text in the target language into speech data (using a speech synthesis API, etc.). The generated speech data is sent to the terminal via the communication module.

[1119] Audio playback

[1120] The device receives audio data from the server and provides it to the user through a playback device (speaker or earphones). Through this process, the user can understand speech in different languages ​​in real time.

[1121] Specific usage examples

[1122] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[1123] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics."

[1124] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[1125] 3. The server analyzes the received audio data, converts it into Japanese text, and then uses a translation module to translate it into English as "Today's lecture is about the basics of quantum mechanics".

[1126] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[1127] 5. The device plays the received English audio data, and User A understands it in real time.

[1128] This enables users who speak different languages ​​to communicate in real time, facilitating smooth lectures even in multinational classes in educational settings.

[1129] The following describes the processing flow.

[1130] Step 1:

[1131] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[1132] Step 2:

[1133] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[1134] Step 3:

[1135] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[1136] Step 4:

[1137] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[1138] Step 5:

[1139] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[1140] Step 6:

[1141] The server passes text data to the translation module, which then translates it into the target language in real time.

[1142] Step 7:

[1143] The server passes the translated text data to the speech synthesis module, which then generates synthesized speech.

[1144] Step 8:

[1145] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[1146] Step 9:

[1147] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[1148] Step 10:

[1149] Users can listen to the played audio and understand speech in different languages ​​in real time.

[1150] (Example 1)

[1151] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1152] Current real-time speech translation systems suffer from insufficient security in the communication protocols they use, potentially leading to data eavesdropping during transmission. Furthermore, the accuracy of speech recognition and translation is often low, failing to meet user expectations. Additionally, complex and time-consuming intermediate processing can compromise real-time capabilities.

[1153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1154] In this invention, the server includes means for analyzing voice data, means for translating, and means for synthesizing voice data. This enables efficient and secure execution of each process of analysis, translation, and voice synthesis, and provides real-time, highly accurate voice translation. Furthermore, by using an optimized communication protocol, the security of data transfer can be improved and user privacy can be protected.

[1155] "Acquisition means" refers to a device used to capture the voice spoken by the user, such as a microphone.

[1156] "Conversion means" refers to a device used to convert captured analog audio data into a digital format, specifically such as an ADC (Analog to Digital Converter).

[1157] "Transmission means" refers to devices and protocols used to send digital audio data to a server in the cloud, such as Wi-Fi modules or secure communication protocols like SSL / TLS.

[1158] "Analysis means" refers to technologies that convert audio data received by a server into text data and identify the language being used, specifically referring to speech recognition engines and speech recognition APIs.

[1159] "Translation methods" refer to technologies used to translate text data obtained through analysis into different languages, such as machine translation engines and translation APIs.

[1160] "Speech synthesis means" refers to technologies for converting translated text into speech data, such as speech synthesis APIs.

[1161] "Playback means" refers to a device that outputs the generated audio data so that the user can hear it, such as a speaker or earphones.

[1162] "Secure communication methods" refer to protocols and technologies for securely transferring voice data, including encrypted communication technologies such as SSL / TLS.

[1163] This invention relates to a system that translates a user's speech into another language in real time and plays back the translated audio. Specific embodiments of this system are described below.

[1164] System Configuration

[1165] This system primarily consists of wearable devices (such as glasses or earphones) and a server in the cloud. Users wear the device and can communicate using voice.

[1166] Device configuration

[1167] The device is equipped with a microphone to acquire sound, an ADC (Analog to Digital Converter) to convert the audio data into a digital format, a communication module (such as a Wi-Fi module) to send the data to the server, and a speaker or earphones to play back the received translated audio.

[1168] Server Configuration

[1169] The server includes a speech recognition module (e.g., a speech recognition API) that analyzes the received audio data and identifies the language, a translation module (e.g., a machine translation engine) that translates the identified language into the target language, a speech synthesis module (e.g., a speech synthesis API) that converts the translated text into speech, and a communication module that sends the generated audio data back to the terminal.

[1170] Specific example of processing

[1171] As a concrete example, consider a situation where a foreign language learner understands a conversation in real time.

[1172] Situation:

[1173] A scenario where an English-speaking tourist (User A) communicates with a Japanese-speaking guide (User B) at a tourist destination in Japan.

[1174] Situation handling procedure

[1175] 1. User B says in Japanese, "This is Senso-ji Temple."

[1176] 2. The terminal acquires this utterance and converts it into digital speech data using an ADC.

[1177] 3. The terminal sends the converted digital audio data to the server via the communication module. Secure protocols such as SSL / TLS are used for communication.

[1178] 4. The server passes the received audio data to the analysis module (speech recognition API) and converts it into the Japanese text "This is Senso-ji Temple."

[1179] 5. The server passes the analyzed text to a translation module (machine translation engine), which translates it into the English text "This is Sensoji Temple".

[1180] 6. The server passes the translated text to the speech synthesis module (speech synthesis API) and converts it into the English speech data "Hello".

[1181] 7. The server sends the generated English voice data to the terminal via the communication module. A secure protocol is used for communication again.

[1182] 8. The device plays the received English audio data through a playback device (earphones or speaker), and user A understands it as "This is Sensoji Temple."

[1183] This allows users to understand speech in different languages ​​in real time.

[1184] Example prompts for the generative AI model to be used

[1185] The following prompt can be entered using a generative AI model.

[1186] Please describe the specific processing flow of a system that, when a Japanese-speaking user says "Konnichiwa" (こんにちは), translates it into English in real time and provides the user with "Hello." Please also specify the technologies and devices used at each step.

[1187] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1188] Step 1:

[1189] The user says "Hello." The device's microphone captures this audio. The input is an analog audio signal, and the microphone captures it, resulting in an analog audio signal as output.

[1190] Step 2:

[1191] The terminal converts the analog audio signal acquired from the microphone into digital audio data using an ADC (Analog to Digital Converter). The input is an analog audio signal, and the output after conversion is digital audio data.

[1192] Step 3:

[1193] The terminal transmits the converted digital audio data to a server in the cloud via a communication module (e.g., a Wi-Fi module). Secure communication protocols such as SSL / TLS are used during this process. The input is digital audio data, and the output is the digital audio data transmitted to the server.

[1194] Step 4:

[1195] The server passes the received digital audio data to the speech recognition module (speech recognition API). The speech recognition module analyzes the input digital audio data and converts it into text data. Specifically, the output is the Japanese text data "こんにちは" (konnichiwa).

[1196] Step 5:

[1197] The server passes the text data obtained from the speech recognition module to the translation module (machine translation engine). The translation module translates the input Japanese text data "こんにちは" into the English text data "Hello". The output is the English text data "Hello".

[1198] Step 6:

[1199] The server passes the translated English text data "Hello" to the speech synthesis module (speech synthesis API). The speech synthesis module parses the input text data and converts it into speech data. The output is the English speech data "Hello".

[1200] Step 7:

[1201] The server sends the generated English audio data to the terminal via a communication module. Secure protocols such as SSL / TLS are again used for communication. The input is English audio data, and the output is the English audio data sent to the terminal.

[1202] Step 8:

[1203] The device receives English audio data and plays it back through a playback device (earphones or speaker) for the user to hear. The input is English audio data, and the output is the audio that the user can hear.

[1204] This allows users to understand speech in different languages ​​in real time.

[1205] (Application Example 1)

[1206] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1207] In factories, when multinational workers work together, language barriers can hinder smooth communication, leading to decreased production efficiency. Furthermore, important instructions and urgent announcements may not be accurately conveyed, potentially impacting safety. To resolve these issues and improve work efficiency and safety, measures to overcome language differences are necessary.

[1208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1209] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio data into a digital format, and a transmission means for sending the converted audio data to the server. This enables a system, including a robot, for real-time communication with workers who speak multiple languages ​​in a factory environment.

[1210] "Acquisition means" refers to the devices or systems used to acquire sound.

[1211] "Conversion means" refers to a device or function for converting acquired audio data into a digital format.

[1212] "Transmission means" refers to communication functions or modules for sending the converted audio data to the server.

[1213] "Analysis means" refers to technologies and systems for analyzing transmitted audio data and identifying the language being used.

[1214] "Translation means" refers to technologies and devices for translating audio data of a specified language into a target language.

[1215] "Speech synthesis means" refers to technologies and systems for converting translated text into speech data.

[1216] "Playback means" refers to devices or functions for playing back audio data transmitted from a server.

[1217] A "robot" is a machine used in a factory environment to communicate in real time with multinational workers.

[1218] This invention relates to a multilingual voice translation system for facilitating communication among multinational workers in a factory. The embodiments for carrying out the invention are as follows:

[1219] System Configuration

[1220] This system primarily consists of wearable terminals (e.g., devices mounted on robots or work helmets) and cloud-based servers. Users can communicate via voice through the terminals. Each device and server includes the following functional modules.

[1221] terminal

[1222] Acquisition method: A microphone for acquiring sound. This microphone has a noise reduction function.

[1223] Conversion method: Analog to Digital Converter (ADC) that converts audio data acquired by a microphone into a digital format.

[1224] Transmission method: A communication module (e.g., Wi-Fi or 5G module) for securely transmitting the converted audio data to a server in the cloud.

[1225] Playback method: A high-quality speaker for playing back translated audio data received from the server.

[1226] server

[1227] Analysis method: A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data received by the server and identify the language being used.

[1228] Translation method: A machine translation engine (e.g., Microsoft Translator) for translating text in a specified language into a target language in real time.

[1229] Speech synthesis means: Speech synthesis technology for converting translated text into speech data (e.g., Amazon Polly).

[1230] Transmission method: A communication module for transmitting synthesized voice data to a terminal.

[1231] Explanation of program processing

[1232] The server receives the acquired audio data and passes it to a speech recognition engine to convert it into text. The analyzed text is then translated into the target language by a machine translation engine. The translated text is converted back into audio data using speech synthesis technology, and this audio data is transmitted to the terminal. The transmitted audio data is provided to the worker via a playback device, enabling real-time communication among multinational workers.

[1233] Specific example

[1234] For example, consider a scenario in a factory where a Japanese-speaking worker (User A) collaborates with a worker who only speaks English (User B).

[1235] 1. User A says in Japanese, "Please tell me the next steps."

[1236] 2. The terminal acquires this audio, converts it into digital audio data using an ADC, and sends it to the server.

[1237] 3. The server analyzes the received audio data, converts it to Japanese text, and then uses a translation module to translate it into English as "Please explain the next work step".

[1238] 4. The server converts the translated text into English audio data using a speech synthesis module and sends it to the terminal.

[1239] 5. The device plays the received English audio data, and user B understands it in real time.

[1240] Example of a prompt

[1241] Please capture the audio of a factory worker saying "Please tell me the next step" in Japanese. Translate it into English and convert it into audio data that can be understood by an English-speaking technician, then play it back.

[1242] This facilitates smoother communication among multinational workers within the factory, resulting in an efficient and safe working environment.

[1243] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1244] Step 1:

[1245] The user speaks. The input is the user's voice, and the specific action is a factory worker saying "Please tell me the next work procedure" in Japanese. The output is an acoustic signal in the form of a speech waveform.

[1246] Step 2:

[1247] The device captures the user's voice. The input is an acoustic signal, and the specific operation involves a microphone acquiring the voice. The output is analog audio data.

[1248] Step 3:

[1249] The terminal's conversion mechanism converts analog audio data into a digital format. The input is analog audio data, and specifically, the ADC (Analog to Digital Converter) converts the signal into a digital format. The output is digital audio data.

[1250] Step 4:

[1251] The terminal's transmission method sends digital audio data to the server. The input is digital audio data, and specifically, the communication module (Wi-Fi or 5G) sends the data to a server in the cloud. The output is the digital audio data sent to the server.

[1252] Step 5:

[1253] The server's analysis tool analyzes the received audio data and converts it to text using a speech recognition engine. The input is the received digital audio data, and the specific operation is for the speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio into text data. The output is the analyzed text data.

[1254] Step 6:

[1255] The server's analysis tool identifies the language used in the text data. The input is the analyzed text data, and the specific operation is for the language model to identify the language from the text. The output is the identified language information.

[1256] Step 7:

[1257] The server's translation method translates text data in a specified language into the target language. The input consists of specified language information and text data; the specific operation involves a machine translation engine (e.g., Microsoft Translator) translating the text into the target language in real time. The output is the translated text data.

[1258] Step 8:

[1259] The server's speech synthesis system converts translated text data into speech data. The input is translated text data, and the specific operation involves speech synthesis technology (e.g., Amazon Polly) converting the text in the target language into speech data. The output is synthesized speech data.

[1260] Step 9:

[1261] The server's transmission mechanism sends synthesized audio data to the terminal. The input is synthesized audio data, and the specific operation is that the communication module sends the audio data to the terminal. The output is the audio data received by the terminal.

[1262] Step 10:

[1263] The device's playback mechanism plays the received audio data. The input is the audio data received by the device, and the specific operation is for the speaker to play the audio and provide it to the user. The output is audio in a language the user understands.

[1264] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1265] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. The configuration of this system and the processing of its program will be described in detail below.

[1266] System Configuration

[1267] This system consists of wearable devices (e.g., glasses-type devices or earphone-type devices) and a server in the cloud. By wearing the device and communicating via voice, the system enables the following functions. Its main components are as follows:

[1268] 1. The terminal is equipped with a microphone to capture user speech, an ADC (Analog to Digital Converter) to convert audio data into a digital format, a communication module to send data to the server, and a speaker or earphones to play back the received translated audio.

[1269] 2. The server includes a speech recognition module that analyzes received audio data and identifies the language, a translation module that translates the identified language into the target language, and a speech synthesis module that converts the translated text into speech. It also includes an emotion engine that recognizes the user's emotions and has the ability to adjust the tone and nuances of the voice based on those emotions.

[1270] 3. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, anger, surprise, etc.) from the voice data and uses this to influence the translation and speech synthesis processes.

[1271] Program processing

[1272] The following provides a detailed explanation of how this system's program processes information.

[1273] Voice acquisition and transmission

[1274] When a user speaks, the device's microphone captures the audio. The audio data is converted to a digital format by an ADC, and the converted data is sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol.

[1275] Audio data analysis and emotion recognition

[1276] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio into text. Simultaneously, an emotion engine operates to analyze the user's emotional state from the audio data.

[1277] Language identification and translation

[1278] The server identifies the language used from the analyzed text. Next, the text data and recognized sentiment data are passed to a translation module, which translates them into the target language in real time. This translation process uses the sentiment data to appropriately adjust tone and nuance.

[1279] Translated speech synthesis and transmission

[1280] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the terminal via the communication module.

[1281] Audio playback

[1282] The device receives audio data from the server and sends it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[1283] Specific usage examples

[1284] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[1285] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[1286] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[1287] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[1288] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[1289] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[1290] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[1291] This enables users who speak different languages ​​to communicate in real time, including conveying emotions and nuances. It can be effectively used in educational settings and international business conferences.

[1292] The following describes the processing flow.

[1293] Step 1:

[1294] The user uses a wearable device and begins to speak. The device's microphone captures this speech.

[1295] Step 2:

[1296] The device receives the audio acquired by the microphone as analog data and converts it into digital audio data.

[1297] Step 3:

[1298] The device uses a communication module to send digital audio data to a server in the cloud. The transmission is performed using a secure communication protocol.

[1299] Step 4:

[1300] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine to convert the audio data into text data.

[1301] Step 5:

[1302] The server analyzes the text data generated by the speech recognition engine to identify the language being used.

[1303] Step 6:

[1304] The server passes audio data to the emotion engine along with text data in the identified language to recognize the user's emotions. Specifically, it analyzes features such as tone, pace, and intonation of the voice to identify the emotional state (e.g., joy, sadness, anger, surprise, etc.).

[1305] Step 7:

[1306] The server passes the text data and recognized sentiment data to the translation module, which then translates it into the target language. During this process, the translation module takes the sentiment data into account and generates translated text with appropriate tone and nuances.

[1307] Step 8:

[1308] The server passes the translated text data and sentiment data to the speech synthesis module to generate synthesized speech. Based on the sentiment data, the speech synthesis module adjusts the tone and intonation of the speech to produce natural-sounding speech.

[1309] Step 9:

[1310] The server sends the generated audio data back to the terminal. This transmission is also performed using a secure communication protocol.

[1311] Step 10:

[1312] The device receives audio data, sends it to a playback device (speaker or earphones), and provides it to the user as audio.

[1313] Step 11:

[1314] Users can listen to the played audio and understand the nuances of speech and emotions in different languages ​​in real time.

[1315] (Example 2)

[1316] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1317] Conventional speech translation systems have a problem in that they fail to properly convey emotions and tone when translating user speech into other languages. Furthermore, because they do not consider the user's emotions during the speech recognition and translation process, nuances of communication are often lost. As a result, especially in educational and business settings, translated content is often not accurately conveyed, leading to misunderstandings.

[1318] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means for analyzing voice data and identifying language, an emotion recognition means for recognizing emotions from voice data, and an emotion adjustment means for adjusting translation and speech synthesis based on the recognized emotions. This enables real-time multilingual translation that reflects the user's emotions and nuances.

[1319] "Acquisition means" refers to a device or function used to capture a user's speech.

[1320] "Conversion means" refers to a device or function that converts acquired audio data from analog format to digital format.

[1321] "Transmission means" refers to a device or function for transmitting audio data converted into a digital format to a server.

[1322] "Analysis means" refers to a device or function that analyzes transmitted audio data and identifies the language.

[1323] "Translation means" refers to a device or function that translates audio data of a specified language into a target language.

[1324] "Speech synthesis means" refers to a device or function that converts translated text into speech data.

[1325] "Emotion recognition means" refers to a device or function that recognizes a user's emotions from audio data.

[1326] "Emotion adjustment means" refers to a device or function that adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[1327] "Playback means" refers to a device or function for playing back transmitted audio data to the user.

[1328] A "server" is a central computer system used to perform voice data analysis, translation, speech synthesis, and emotion recognition.

[1329] A "terminal" is a device worn by a user that has the function of acquiring and playing back audio.

[1330] This invention relates to a system that translates a user's speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. This system consists of a wearable device (e.g., glasses-type device or earphone-type device) and a server in the cloud. By wearing the device and communicating via voice, the system achieves the following functions:

[1331] System Configuration

[1332] The main components of this system are as follows:

[1333] terminal

[1334] Acquisition method: Equipped with a microphone for capturing user speech.

[1335] Conversion method: Includes an ADC (Analog to Digital Converter) that converts audio into a digital format.

[1336] Transmission method: It is equipped with a communication module that transmits the acquired digital audio data to a server in the cloud.

[1337] Playback method: It has a speaker or earphones that play back the translated audio received from the server.

[1338] server

[1339] Analysis method: The received audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API or Microsoft Azure Speech Service).

[1340] Translation method: To translate text data into the target language, we use APIs such as Google Translate or DeepL.

[1341] Speech synthesis method: Amazon Polly or Google Cloud Text-to-Speech API are used to convert translated text into speech data.

[1342] Emotion recognition means: An emotion engine is used to recognize the user's emotions from voice data and evaluate tone and nuance.

[1343] Emotion adjustment mechanism: Adjusts the tone and nuances of translation and speech synthesis based on recognized emotions.

[1344] Examples of usage and prompt statements

[1345] As a concrete example, consider a university lecture. A professor who speaks Japanese (User B) begins a lecture in Japanese, while a student who only speaks English (User A) is wearing this system.

[1346] 1. User B says in Japanese, "Today's lecture is about the fundamentals of quantum mechanics." The professor speaks in a way that suggests he is questioning the statement.

[1347] 2. The terminal acquires this utterance, converts it into digital audio data, and sends it to the server.

[1348] 3. The server analyzes the received audio data, converts it into Japanese text, and uses an emotion engine to recognize the emotion behind the professor's question.

[1349] 4. The server passes the text data and sentiment data to the translation module, which translates "Today's lecture is about the basics of quantum mechanics" into English, while simultaneously reflecting the nuance of the question.

[1350] 5. The server converts the translated text and sentiment data into English speech data using a speech synthesis module and sends it to the terminal.

[1351] 6. The device plays the received English audio data, allowing user A to understand the nuances of the professor's questions in real time.

[1352] Using a generative AI model, you can utilize the following example prompts:

[1353] Example prompt

[1354] In a university lecture, a Japanese-speaking professor said, "Today's lecture will be about the fundamentals of quantum mechanics," with a hint of questioning. Please explain in detail how to translate this into English in real time, while preserving the sense of questioning, and provide it to English-only students.

[1355] This prompt statement is useful for specifically describing the system's behavior in particular scenarios.

[1356] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1357] Step 1:

[1358] When a user speaks, the device's microphone captures the audio data. The captured audio is in analog format and is converted to digital format by an ADC (Analog to Digital Converter). The input is the user's analog voice, and the output is digital audio data. This data conversion provides digital audio data that can be used in subsequent processing steps.

[1359] Step 2:

[1360] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. This transmission uses a secure communication protocol such as HTTPS. The input is digital audio data, and the output is the transmitted audio data. The communication module divides the data into packets and transmits them to ensure they reach the server reliably.

[1361] Step 3:

[1362] The server passes the received audio data to the analysis module. The analysis module uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to convert the audio to text. The input is the digital audio data sent to the server, and the output is text data. The speech recognition engine performs phoneme and sound waveform analysis to generate appropriate text data.

[1363] Step 4:

[1364] The server's emotion recognition engine analyzes audio data in parallel to identify the user's emotional state. It evaluates the tone, pitch, and speed of the speech to recognize emotional states (such as joy, sadness, anger, or surprise). The input is digital audio data, and the output is emotion data. Emotion recognition is performed using machine learning algorithms.

[1365] Step 5:

[1366] The server passes the analyzed text data and sentiment data to the translation module. The translation module (e.g., Google Translate API or DeepL API) translates the text into the target language in real time. The input is text data and sentiment data, and the output is the translated text. During translation, the tone and nuances are also appropriately adjusted, taking sentiment data into account.

[1367] Step 6:

[1368] The server passes the translated text and sentiment data to a speech synthesis module. The speech synthesis module (e.g., Amazon Polly or Google Cloud Text-to-Speech API) converts the text into speech. The input is the translated text and sentiment data, and the output is the generated synthesized speech data. This process adjusts the tone and intonation of the speech based on the sentiment data.

[1369] Step 7:

[1370] The server sends the generated synthesized speech data to the terminal via a communication module. A secure protocol (e.g., HTTPS) is used for communication to ensure data integrity and confidentiality. The input is the synthesized speech data, and the output is the transmitted speech data.

[1371] Step 8:

[1372] The terminal receives audio data from the server, sends it to a playback device (speaker or earphones), and provides it to the user as audio. The input is the received audio data, and the output is the played audio. This allows the user to understand speech in different languages ​​in real time, along with the emotions conveyed.

[1373] (Application Example 2)

[1374] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1375] In global manufacturing environments, language differences pose a significant obstacle when multinational staff collaborate. Furthermore, accurately conveying not only language but also the emotions and nuances of speech is crucial, but current systems fail to adequately address this. Therefore, there is a need for real-time multilingual translation and emotion recognition simultaneously to facilitate smooth communication on the factory floor.

[1376] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes emotion recognition means for analyzing emotional states from voice data, means for adjusting the tone and nuances of translated text based on the analyzed emotion data, analysis means for converting acquired voice data into text using a voice recognition engine and analyzing emotional states using an emotion engine, and translation means using a translation module that translates text data and emotion data in real time and adjusts the tone and nuances based on the emotion data. This enables accurate multilingual translation and emotion recognition in real time.

[1377] "Audio data" refers to data that represents audio information in a digital format.

[1378] A "server" is a computer system that provides data processing and storage services over a network.

[1379] "Emotion recognition means" refers to functions or devices that analyze and recognize the speaker's emotions from audio data.

[1380] "Emotional data" refers to data that indicates the state of emotions as analyzed by emotion recognition tools.

[1381] A "translation tool" is a function or device used to convert text or audio expressed in one language into another language.

[1382] A "translation module" is a software module used to translate text data or audio into other languages.

[1383] A "speech recognition engine" is a software engine that converts speech data into text.

[1384] "Tone and nuance" refers to the emotional expression and subtle meanings conveyed by speech or text.

[1385] "Analysis means" refers to functions or devices for analyzing acquired audio data and extracting specific information or patterns.

[1386] "Acquisition means" refers to functions or devices for collecting or capturing data such as audio and video.

[1387] "Playback means" refers to functions or devices for playing back data on a terminal.

[1388] Modes for carrying out the invention

[1389] This invention relates to a system that translates user speech into another language in real time, recognizes the user's emotions, and provides an adapted translation based on those emotions. As an example of its application, this system will be used to realize a smart glasses application that facilitates communication in global manufacturing environments. The system configuration and detailed program processing of this invention will be described below.

[1390] System Configuration

[1391] This system consists of smart glasses-type devices and a server in the cloud. When the user wears the smart glasses and communicates via voice, the system enables the following functions.

[1392] Device (smart glasses)

[1393] Acquisition method: Includes a microphone for acquiring sound.

[1394] Conversion means: Includes an ADC (Analog to Digital Converter) that converts the acquired audio data into a digital format.

[1395] Transmission means: Includes a communication module that transmits the converted audio data to a server.

[1396] Playback means: Includes speakers or earphones that play audio data transmitted from the server.

[1397] server

[1398] Analysis method: The transmitted audio data is analyzed and converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[1399] Emotion recognition method: An emotion engine (e.g., IBM Watson Tone Analyzer) is used to analyze the emotional state from voice data.

[1400] Translation method: Text data and sentiment data are translated in real time and translated into the target language using a translation module (e.g., Microsoft Translator API).

[1401] Speech synthesis method: Based on translated text and sentiment data, a speech synthesis module (e.g., Amazon Polly) converts it into speech.

[1402] Program processing

[1403] Voice acquisition and transmission

[1404] When a user speaks, the microphone in the smart glasses captures their voice. The voice data is converted to a digital format by an ADC and sent to a server in the cloud via a communication module. This transmission is performed using a secure communication protocol such as SSL.

[1405] Audio data analysis and emotion recognition

[1406] The audio data received by the server is converted into text in real time by a speech recognition engine. In parallel, an emotion recognition system operates to analyze the user's emotional state from the audio data.

[1407] Language identification and translation

[1408] The server identifies the language used from the analyzed text, passes the text data and recognized sentiment data to a translation module, and translates it into the target language in real time. This translation process uses sentiment data to appropriately adjust tone and nuances.

[1409] Translated speech synthesis and transmission

[1410] The server passes the translated text and sentiment data to the speech synthesis module to generate synthesized speech. The speech synthesis module adjusts the tone and intonation of the speech based on the sentiment data. The generated speech data is transmitted to the smart glasses via the communication module.

[1411] Audio playback

[1412] The smart glasses receive audio data from a server and send it to a playback device (speaker or earphones), providing it to the user as audio. This allows the user to understand speech in different languages, including the emotions behind it, in real time.

[1413] Specific example

[1414] Let's consider an example of use in a global factory setting. Suppose a team leader gives the instruction in Japanese: "Please install the next part." If the leader speaks in a calm tone, the smart glasses system will operate as follows:

[1415] 1. The reader's speech is captured by the microphone in the smart glasses and converted into a digital format.

[1416] 2. The digital audio data is sent to a cloud server and converted into text, "Please install the following parts," by a speech recognition engine.

[1417] 3. The emotion recognition system analyzes the calm emotion, and the translation module translates it into English as "Please install the next part." The tone and nuances are also adjusted based on the emotion data.

[1418] 4. The translated text is converted into English speech by a speech synthesis module and sent to the smart glasses.

[1419] 5. Translated English instructions are played through the smart glasses' speaker, allowing staff of other nationalities to understand them.

[1420] Example of a prompt

[1421] "Please provide prototype code for an application that translates user-initiated instructions into another language in real time and returns the translation with appropriate emotional nuances. Furthermore, please describe the hardware and software used, and the specific processing steps."

[1422] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1423] Step 1:

[1424] Acquiring audio

[1425] When a user speaks, the device's microphone captures the audio. The input is the user's spoken audio data, and the output is the audio data converted into a digital format. This audio data is converted into a digital format by an ADC (Audio-Digital Converter).

[1426] Step 2:

[1427] Sending audio data

[1428] The terminal transmits the converted digital audio data to a server in the cloud via a communication module. The input is the audio data converted to digital format, and the output is the audio data sent to the server. This transmission is performed using a secure communication protocol such as SSL.

[1429] Step 3:

[1430] Analysis of audio data

[1431] The server passes the received audio data to the analysis device, which then converts it into text using a speech recognition engine. The input is the audio data sent to the server, and the output is text data. Specifically, the server's speech recognition engine analyzes the audio data as text data.

[1432] Step 4:

[1433] emotion recognition

[1434] The server simultaneously passes the audio data to the emotion recognition system, which uses an emotion engine to analyze the emotional state from the audio data. The inputs are text data and audio data, and the output is emotion data. The emotion engine identifies the emotional state from the text data and audio data.

[1435] Step 5:

[1436] translation

[1437] The server passes the analyzed text data and recognized sentiment data to the translation system, which translates it into the target language in real time. The input is text data and sentiment data, and the output is the translated text data in the target language. This translation process uses sentiment data to adjust tone and nuance.

[1438] Step 6:

[1439] Speech synthesis

[1440] The server passes the translated text data and sentiment data to the speech synthesis system, which uses the speech synthesis module to generate speech data. The input is the translated text data and sentiment data, and the output is the synthesized speech data. The speech synthesis module uses the sentiment data to adjust the tone and intonation of the speech.

[1441] Step 7:

[1442] Sending audio data

[1443] The server transmits the generated audio data to the terminal via a communication module. The input is synthesized speech data, and the output is the audio data transmitted to the terminal. A secure communication protocol is used to transmit the audio data to the terminal.

[1444] Step 8:

[1445] Audio playback

[1446] The terminal receives audio data from the server and sends it to the playback device (speaker or earphones). The input is the received audio data, and the output is the audio provided to the user. This allows the user to understand speech in different languages, including the emotions conveyed. The audio is then played back through the terminal's speaker or earphones.

[1447] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1448] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1449] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1450] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1451] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1452] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1453] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1454] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1455] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1456] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1457] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1458] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1459] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1460] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1461] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1462] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1463] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1464] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1465] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1466] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1467] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1468] The following is further disclosed regarding the embodiments described above.

[1469] (Claim 1)

[1470] A means of acquiring sound,

[1471] A conversion means for converting acquired audio data into a digital format,

[1472] A transmission means for sending the converted audio data to a server,

[1473] An analysis means for analyzing transmitted audio data and identifying the language,

[1474] A translation method for translating audio data of a specified language into a target language,

[1475] A speech synthesis means that converts translated text into audio data,

[1476] A transmission means for sending synthesized audio data to a terminal,

[1477] A playback means for playing back transmitted audio data,

[1478] A system that includes this.

[1479] (Claim 2)

[1480] The system according to claim 1, wherein the analysis means converts the acquired audio data into text using a speech recognition engine.

[1481] (Claim 3)

[1482] The system according to claim 1, wherein the translation means uses a translation engine that translates text data in real time.

[1483] "Example 1"

[1484] (Claim 1)

[1485] means of acquisition,

[1486] A means of converting to a digital format,

[1487] A means of sending to the server,

[1488] The server analyzes and identifies the language using an analysis method,

[1489] A translation method that translates a specified language into a target language,

[1490] A speech synthesis means that converts translated text into audio data,

[1491] A means of transmission to send to the terminal,

[1492] A means of playback for playing audio data,

[1493] A secure communication means that securely transfers data using an optimized communication protocol,

[1494] A system that includes this.

[1495] (Claim 2)

[1496] The system according to claim 1, wherein the acquisition means includes a microphone for acquiring an audio waveform.

[1497] (Claim 3)

[1498] The system according to claim 1, wherein the analysis means uses speech recognition technology to convert an audio signal into text data.

[1499] "Application Example 1"

[1500] (Claim 1)

[1501] A means of acquiring sound,

[1502] A conversion means for converting acquired audio data into a digital format,

[1503] A transmission means for sending the converted audio data to a server,

[1504] An analysis means for analyzing transmitted audio data and identifying the language,

[1505] A translation method for translating audio data of a specified language into a target language,

[1506] A speech synthesis means that converts translated text into audio data,

[1507] A transmission means for sending synthesized audio data to a terminal,

[1508] A playback means for playing back transmitted audio data,

[1509] A system including robots for real-time communication with multilingual workers in a factory environment.

[1510] (Claim 2)

[1511] The system according to claim 1, wherein the analysis means converts the acquired audio data into text using a speech recognition engine.

[1512] (Claim 3)

[1513] The system according to claim 1, which uses a translation engine that translates text data in real time, and is used to facilitate communication between workers in a factory environment.

[1514] "Example 2 of combining an emotion engine"

[1515] (Claim 1)

[1516] A means of acquiring sound,

[1517] A conversion means for converting acquired audio data into a digital format,

[1518] A transmission means for sending the converted audio data to a server,

[1519] An analysis means for analyzing transmitted audio data and identifying the language,

[1520] A translation method for translating audio data of a specified language into a target language,

[1521] A speech synthesis means that converts translated text into audio data,

[1522] A transmission means for sending synthesized audio data to a terminal,

[1523] A playback means for playing back transmitted audio data,

[1524] An emotion recognition method that recognizes emotions from audio data,

[1525] An emotion adjustment means that adjusts translation and speech synthesis based on recognized emotions,

[1526] A system that includes this.

[1527] (Claim 2)

[1528] The system according to claim 1, wherein the analysis means converts the acquired audio data into text using a speech recognition engine, and the emotion recognition means evaluates the tone, pitch, and speed of the audio to recognize the emotional state.

[1529] (Claim 3)

[1530] The system according to claim 1, wherein the translation means uses a translation engine that translates text data and emotional states in real time, and the speech synthesis means adjusts the tone and intonation of the speech taking into account the recognized emotional state.

[1531] "Application example 2 when combining with an emotional engine"

[1532] (Claim 1)

[1533] A means of acquiring sound,

[1534] A conversion means for converting acquired audio data into a digital format,

[1535] A transmission means for sending the converted audio data to a server,

[1536] An analysis means for analyzing transmitted audio data and identifying the language,

[1537] A translation method for translating audio data of a specified language into a target language,

[1538] A speech synthesis means that converts translated text into audio data,

[1539] A transmission means for sending synthesized audio data to a terminal,

[1540] A playback means for playing back transmitted audio data,

[1541] An emotion recognition method that analyzes emotional states from audio data,

[1542] A means of adjusting the tone and nuances of translated text based on analyzed sentiment data,

[1543] A system that includes this.

[1544] (Claim 2)

[1545] The system according to claim 1, wherein the analysis means converts acquired audio data into text using a speech recognition engine, and the emotion recognition means analyzes the emotional state using an emotion engine.

[1546] (Claim 3)

[1547] The system according to claim 1, wherein the translation means uses a translation module that translates text data and sentiment data in real time and adjusts tone and nuance based on the sentiment data. [Explanation of Symbols]

[1548] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring sound, A conversion means for converting acquired audio data into a digital format, A transmission means for sending the converted audio data to a server, An analysis means for analyzing transmitted audio data and identifying the language, A translation method for translating audio data of a specified language into a target language, A speech synthesis means that converts translated text into audio data, A transmission means for sending synthesized audio data to a terminal, A playback means for playing back transmitted audio data, A system that includes this.

2. The system according to claim 1, wherein the analysis means converts the acquired audio data into text using a speech recognition engine.

3. The system according to claim 1, wherein the translation means uses a translation engine that translates text data in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A