System

The eyeglass-type device with integrated speech recognition and translation capabilities addresses the limitations of conventional interpretation devices by providing real-time, convenient, and private multilingual conversations.

JP2026019154APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120563
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional interpretation devices face challenges such as time loss during translation, disruption of natural conversation flow, privacy concerns due to the need for interpreters, and inconvenience due to large equipment size, making them unsuitable for easy portability.

Method used

A system comprising an eyeglass-type device with built-in microphone and speaker, a computing unit for speech recognition and translation, and a communication means for real-time data transfer, along with a dedicated application for managing settings and a display for subtitles, enabling seamless multilingual conversations.

Benefits of technology

Enables natural, time-efficient, and private multilingual conversations without the need for large equipment, maintaining conversation flow and enhancing user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019154000001_ABST
    Figure 2026019154000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining speech of another person by voice input means; voice recognition means for converting the obtained voice into text; translation means for translating the converted text into a target language set by a user; voice output means for converting the translated text into voice and outputting the voice; a glasses-type device including the voice input means and the voice output means; an arithmetic device including the voice recognition means and the translation means; and communication means for connecting the glasses-type device and the arithmetic device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When conducting a multilingual conversation, conventional interpretation devices and applications have the problem of time loss during translation, disrupting the natural flow of conversation. Another issue is that the need for an interpreter makes it difficult to protect the privacy of individuals. Furthermore, conventional voice interpretation devices require large equipment and are inconvenient to carry around. The purpose of this invention is to solve these problems and realize natural multilingual conversation without time loss. [Means for solving the problem]

[0005] The present invention is a system including a speech recognition unit that acquires speech from a user via a speech input unit and converts the acquired speech into text, a translation unit that translates the converted text into a target language set by the user, and a speech output unit that converts the translated text into speech and outputs it. The system includes an eyeglass-type device equipped with a speech input unit and a speech output unit, a computing unit including a speech recognition unit and a translation unit, and a communication unit for connecting the eyeglass-type device and the computing unit. Furthermore, the system includes a means for the user to manage settings using a dedicated application and a display unit that displays the translation results as subtitles, thereby enhancing user convenience. In this way, the present invention supports natural, time-saving conversations between multiple languages.

[0006] The "voice input means" refers to a microphone or other voice input device for capturing speech from the other party.

[0007] "Speech recognition means" refers to technology or software for converting captured speech into text.

[0008] "Translation tools" are techniques and software for translating the converted text into a target language selected by the user.

[0009] "Audio output means" refers to a speaker or other audio playback device that converts the translated text into audio and plays it aloud to the user.

[0010] An "eyeglasses-type device" is a device in the shape of glasses that physically incorporates an audio input means and an audio output means.

[0011] The "computing device" is a computer that includes a speech recognition means and a translation means and is responsible for processing speech data.

[0012] The "communication means" refers to wireless communication technology such as Bluetooth or Wi-Fi for connecting the eyeglass device and the computing unit.

[0013] A "dedicated application" is software that allows users to configure and manage the system.

[0014] The "display means" refers to a display or other display device for visually displaying the translation results as subtitles. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention provides a system for translating conversations between multiple languages ​​in real time and supporting natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, and a display means.

[0037] System configuration overview

[0038] 1. Eyeglass-type device worn by the user

[0039] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0040] 2. Arithmetic device

[0041] The apparatus includes a speech recognition unit for converting speech into text and a translation unit for translating the text into a target language. The computing unit receives the speech data obtained from the eyeglasses-type device and performs appropriate processing.

[0042] 3. Means of communication

[0043] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0044] 4. Dedicated application

[0045] Users use a dedicated application to manage the system's settings, such as the target language for translation and the voice speed of the audio output.

[0046] 5. Display means

[0047] A display means is provided to visually convey the translation results to the user, which makes it possible to display the translation content as subtitles.

[0048] System program processing

[0049] Acquiring voice input and converting it to text

[0050] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0051] Text translation

[0052] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0053] Voice output of translation results

[0054] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0055] Server and device communication

[0056] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0057] Specific examples

[0058] Scenario 1: English and Japanese Conversation

[0059] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text, "What time is the meeting?", which the translation means translates into "What time is the meeting?" The translated text is converted into speech and played back to User B.

[0060] Scenario 2: Configuration Management

[0061] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation.

[0062] In this way, the present invention enables natural conversation between multiple languages ​​and minimizes the time lost due to interpretation.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] The user puts on the audio glasses and starts the system, which enables the microphone and speaker in the audio glasses.

[0066] Step 2:

[0067] The device (Audio Glasses) picks up what the other person is saying with a microphone. This audio data is captured in real time and temporarily stored in the device.

[0068] Step 3:

[0069] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0070] Step 4:

[0071] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0072] Step 5:

[0073] The server receives the translated text data and converts it into audio data using a text-to-speech engine, which is in the target language rather than the original language.

[0074] Step 6:

[0075] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0076] Step 7:

[0077] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0078] Step 8:

[0079] Users can use a dedicated application to change settings such as the target language, speech output speed, and volume as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0080] Step 9:

[0081] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, and voice output, maintaining natural conversation between multiple languages.

[0082] Example 1

[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0084] In real-time communication between multiple languages, language barriers exist, making smooth conversation difficult. Existing translation systems have issues with translation accuracy and speed, making them insufficient for real-time conversation support. In particular, it is difficult to continue the flow of conversation without interruption, and this aspect needs improvement.

[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0086] In this invention, the server includes means for acquiring speech from the other party via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, means for converting the translated text into voice and outputting it, a wearable device equipped with the voice input means and the voice output means, an information processing device including the voice recognition means and the translation means, communication means for connecting the wearable device and the information processing device, and communication means for transmitting and receiving data in real time, thereby enabling users to have natural real-time conversations between multiple languages.

[0087] "Voice input means" is a device that has the function of acquiring the speech of the other party.

[0088] The "voice recognition means" is a device that has the function of converting acquired voice into text data.

[0089] The "translation means" is a device that has the function of translating text data into a target language set by the user.

[0090] "Speech output means" refers to a device that has the function of converting translated text into speech and outputting it.

[0091] A "wearable device" is a wearable device equipped with an audio input means and an audio output means.

[0092] The "information processing device" is a processing device that includes a speech recognition means and a translation means.

[0093] A "communication means" is a device that connects a wearable device to an information processing device and has the function of sending and receiving data.

[0094] "Communication means for transmitting and receiving data in real time" refers to a communication technology that enables the transmission and reception of data in real time.

[0095] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system includes a wearable device (e.g., a glasses-type device), an information processing device, communication means, a dedicated application for managing user settings, and display means. Each component and its operation are described in detail below.

[0096] Wearable devices

[0097] The wearable device is equipped with a voice input means and a voice output means. It is worn by the user and looks like a glasses-type device. This device has the following functions:

[0098] Voice input means: A microphone is built in to capture what the other person is saying.

[0099] Audio output means: A built-in speaker is included, allowing the user to hear the translated speech.

[0100] Information processing device

[0101] The information processing device includes a speech recognition unit and a translation unit. These units analyze the speech data acquired from the wearable device and perform translation processing. Specific components are as follows:

[0102] Speech recognition means: Converts acquired voice data into text data.

[0103] Translation method: Translates text data into the target language set by the user. The translation engine uses a generative AI model based on a neural network, for example.

[0104] communication means

[0105] The communication means connects the wearable device to the information processing device and transmits and receives data in real time. The communication means includes:

[0106] Bluetooth and Wi-Fi: Used to send and receive voice and translation data.

[0107] Dedicated application

[0108] A dedicated application allows users to manage the system settings. Through this application, the following settings can be configured:

[0109] Select target language: Set the language to translate into.

[0110] Audio output adjustment: Set the audio output speed, etc.

[0111] Display means

[0112] The display means is used to visually convey the translation results to the user. Specifically, it has the following functions:

[0113] Subtitle display: It is possible to display the translated text as subtitles.

[0114] Specific examples

[0115] Example 1: English-Japanese conversation

[0116] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the wearable device picks up the speech and sends it to an information processing device. The speech recognition means converts this speech into text data, "What time is the meeting?", which is then translated by the translation means into "What time is the meeting?" The translated text is converted into speech and played back to User B. The speaker in the wearable device outputs the speech, "What time is the meeting?"

[0117] Example 2: Configuration Management

[0118] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. For example, the target language is changed from Japanese to French, and the speech output speed is set to 1.25 times faster. This information is immediately sent to the information processing device, and the new settings are reflected from the next conversation. Every time the user starts a conversation, translation is performed based on the application settings.

[0119] Prompt Sentence Examples

[0120] Here are some examples of prompts to input to a generative AI model:

[0121] If user A speaks "Hello, how are you?" into a wearable device, how can I convert that speech into text and translate it into Japanese, the target language specified by user B?

[0122] In this way, the present invention enables natural conversations between multiple languages ​​and minimizes the time lost due to interpretation. Users can manage settings through a dedicated application, enabling smooth communication in real time.

[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0124] Step 1: Getting voice input

[0125] When a user starts a conversation, the microphone on the device (wearable device) picks up what the other person is saying. When User A says, "Hello, how are you?", the voice is input to the device through the microphone. The device converts this real-time voice input into digital voice data and prepares for the next step. Here, the input is analog voice, and the output is digital voice data.

[0126] Step 2: Sending audio data

[0127] The terminal (wearable device) transmits digital voice data to the server via Bluetooth or Wi-Fi. For example, voice data such as "Hello, how are you?" is delivered to the server via wireless communication. The input here is the digital voice data, and the output is the voice data transmitted to the server. During this transmission, the communication indicator LED on the terminal lights up to indicate the status of data transmission.

[0128] Step 3: Speech to Text

[0129] The server inputs the received voice data into a voice recognition means and converts it into text data. For example, the voice "Hello, how are you?" is converted into text "Hello, how are you?" The server uses a voice recognition algorithm to analyze the voice data and generate output in text format. Here, the input is the voice data sent to the server, and the output is text data. A processing progress bar on the server operates to display the progress of the conversion.

[0130] Step 4: Translate the text

[0131] The server passes the text data obtained by the speech recognition means to the translation means, which translates it into the target language set by the user. For example, if User B has set Japanese as the target language, "Hello, how are you?" will be translated into "Hello, how are you?" A generative AI model is used to perform highly accurate translation. The input here is the recognized text data, and the output is the translated text data. The original text and the translated text are recorded in the server log.

[0132] Step 5: Audio output of translation results

[0133] The server passes the translated text data to the voice output means, which converts it into voice data. The converted voice data is sent from the server to the terminal (wearable device). The converted voice is played from the terminal's speaker. For example, User B can hear the voice saying, "Hello, how are you?" Here, the input is the translated text data, and the output is playable voice data. An LED flashes to indicate that the terminal's speaker is outputting voice.

[0134] Step 6: Server and device communication

[0135] The server and the terminal (wearable device) are always in communication, sending and receiving data in real time. This allows each step of the conversation to be processed without delay. Every time the user speaks, the communication status can be confirmed by the communication indicator on the terminal lighting up. The input here is the user's voice and setting information, and the output is smooth conversation through real-time data transmission and reception.

[0136] Step 7: Configuration Management

[0137] The user opens a dedicated application to manage settings such as the target language of the conversation and the speech output speed. When the user changes the settings, the setting information is immediately sent to the server and applied throughout the system. For example, the user changes the target language from Japanese to French and sets the speech output speed to 1.25x. The new settings will be applied from the next conversation. The input here is the setting change information made by the user, and the output is the updated system settings.

[0138] In this way, the system translates conversations between multiple languages ​​in real time, providing smooth and natural communication.

[0139] (Application example 1)

[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] Many brick-and-mortar stores face the problem of inability to communicate smoothly with tourists who speak foreign languages. This communication barrier can make it difficult for store staff and tourists to exchange accurate information, resulting in a decline in service quality. Another issue is the time loss caused by the inability to obtain translation results immediately. To solve these problems, a system that supports multilingual conversations in real time is needed.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0143] In this invention, the server includes means for acquiring the other person's speech via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, voice output means for converting the translated text into voice and outputting it, a glasses-type device equipped with the voice input means and the voice output means, a computing device including the voice recognition means and the translation means, communication means for connecting the glasses-type device and a smartphone to the computing device, means for the user to manage settings using a dedicated application, and display means for displaying the translation results as subtitles or means for displaying them on the smartphone. This enables smooth communication between store staff and tourists across language barriers.

[0144] "Audio input means" refers to a device or technology for capturing audio.

[0145] A "speech recognition means" is a technique or device that converts captured speech into text.

[0146] A "translation means" is a technique or device that translates text into a target language set by the user.

[0147] "Audio output means" refers to a device or technology that converts translated text into audio and plays it back.

[0148] An "eyeglasses-type device" is an electronic device in the shape of glasses that has built-in audio input and audio output functions.

[0149] A "computing device" is an electronic device that includes a processor that performs speech recognition and translation.

[0150] "Communication means" refers to a technology or device for transmitting and receiving data between different devices, and includes wireless communication technologies such as Bluetooth and Wi-Fi.

[0151] A "dedicated application" is software that allows users to configure the system and runs on a smartphone or tablet.

[0152] "Display means" means a device or technology for visually displaying the translation result in text, including smart glasses or a smartphone display.

[0153] The present invention provides a real-time translation system that supports communication between users who speak different languages ​​in a physical store. The system specifically includes the following configuration and operation.

[0154] System configuration

[0155] 1. Voice input method

[0156] The server uses the microphone of the smart glasses or smartphone worn by the user to capture the speech of the interlocutor, thereby collecting the speech content as voice data.

[0157] 2. Voice Recognition Method

[0158] The captured voice data is sent to a server and converted into text data using voice recognition software, using the "speech_recognition" library.

[0159] 3. Translation Methods

[0160] The text data acquired by the speech recognition means is translated into the target language by a translation engine on the server. This translation process uses the "googletrans" library.

[0161] 4. Audio output means

[0162] The translated text is then converted back into audio data in the target language using speech synthesis technology, and the server uses the gTTS (Google Text-to-Speech) library to generate an audio file that is played on the user's smart glasses or smartphone speakers.

[0163] 5. Display means

[0164] The translation result is displayed as text on the smart glasses display or on the smartphone screen, allowing users to check the translation not only aloud but also visually.

[0165] Hardware and Software Use

[0166] The hardware used includes smart glasses and smartphones, which connect to the server via Bluetooth or Wi-Fi.

[0167] The software uses:

[0168] Speech Recognition: speech_recognition library

[0169] Translation: GoogleTrans Library

[0170] Speech synthesis: gTTS library

[0171] Audio playback: pygame library

[0172] Bluetooth communication: pybluez library

[0173] Specific examples

[0174] When a store staff member says in Japanese, "How do you use this product?", the microphone in the smart glasses or smartphone picks up the speech and converts it into text using speech recognition. This text is then translated into English by a translation engine into "How do you use this product?" The translation result is converted into speech and played back through the speaker of the staff member's smart glasses or smartphone, while also being displayed as text on the screen.

[0175] Prompt Sentence Examples

[0176] "Design an application for smart glasses that uses voice input to translate conversations between multiple languages ​​in real time, and outputs them in voice and text format."

[0177] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0178] Step 1:

[0179] The user makes a speech. The speech input means acquires the speech through the microphone of the smart glasses or smartphone. At this time, the acquired speech data is a raw speech signal. The server receives the corresponding speech data.

[0180] Step 2:

[0181] The server converts the received voice data into text data using speech recognition software (speech_recognition library). The input is the acquired voice data, and the output is the data converted from the speech into text format. Through this conversion, the voice information is expressed as text information.

[0182] Step 3:

[0183] The server inputs the converted text data into a translation engine (GoogleTrans library) and translates it into the target language. The input is text data, and the output is translated text data. This process converts the user's speech into a language the other person can understand.

[0184] Step 4:

[0185] The server converts the translated text data into audio data using speech synthesis technology (gTTS library). The input is the translated text data, and the output is an audio file. The server generates this audio file and prepares it for playback.

[0186] Step 5:

[0187] The server sends the generated audio file to the speaker of the smart glasses or smartphone and plays it through the audio output means, allowing the other person to hear the translated audio. At the same time, the translation result text is displayed on the display of the smart glasses or smartphone. The input is the audio file and text data, and the output is the playback of the translated audio and the text display.

[0188] Step 6:

[0189] Users manage system settings using a dedicated application. When a user changes the target language, for example, the setting information is sent to the server and is immediately reflected throughout the system. The input is the user's setting information, and the output is a reflection of the updated setting information.

[0190] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0191] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, a display means, and an emotion engine that recognizes the user's emotions.

[0192] System configuration overview

[0193] 1. Eyeglass-type device worn by the user

[0194] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0195] 2. Arithmetic device

[0196] The device includes a speech recognition unit that converts speech into text, a translation unit that translates the text into a target language, and an emotion engine that recognizes the user's emotions. The computing unit receives the speech data acquired from the glasses-type device and performs appropriate processing.

[0197] 3. Means of communication

[0198] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0199] 4. Dedicated application

[0200] Users can use a dedicated application to manage the system's settings, such as the target language for translation, voice output voice speed, volume, etc. There is also a training mode to improve the accuracy of emotion recognition.

[0201] 5. Display means

[0202] The system is equipped with a display means for visually conveying the translation results to the user, which makes it possible to display the translation content as subtitles.

[0203] 6. Emotion Engine

[0204] The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions, and adjusts the tone and speed of the translated speech based on the recognized emotion, ensuring natural conversation.

[0205] System program processing

[0206] Acquiring voice input and converting it to text

[0207] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0208] Text translation

[0209] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0210] Emotion recognition and voice output adjustment

[0211] The translated text data is converted into voice data by an emotion engine, taking into account the user's emotions. For example, if the user is excited, the translated voice will be played in a similarly excited tone. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and choice of words used.

[0212] Voice output of translation results

[0213] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0214] Server and device communication

[0215] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0216] Specific examples

[0217] Scenario 1: English and Japanese Conversation

[0218] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text "What time is the meeting?", and the translation means translates it into "What time is the meeting?" The emotion engine recognizes that User A's question contains tension and plays back the translated speech in a voice that reflects that tension.

[0219] Scenario 2: Configuration Management

[0220] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation. For example, if User B now selects French, "What time is the meeting?" is translated as "À quelle heure est la réunion?" and the emotion engine plays it in a calm, gentle tone.

[0221] In this way, the present invention provides an interpretation system that enables natural conversation between multiple languages ​​and is adaptable to various situations.

[0222] The processing flow will be explained below.

[0223] Step 1:

[0224] The user puts on the audio glasses and starts the system, which activates the microphone and speaker in the audio glasses and starts the emotion engine, preparing to analyze the user's facial expressions and tone of voice.

[0225] Step 2:

[0226] The device (Audio Glasses) picks up what the other person is saying with a microphone. This voice data is captured in real time and temporarily stored in the device. At the same time, the emotion engine records the user's facial expressions and voice tone and generates emotion data.

[0227] Step 3:

[0228] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0229] Step 4:

[0230] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0231] Step 5:

[0232] The server receives the translated text data and the emotion data generated by the emotion engine, which then adjusts the tone and speed of the translated speech based on that data.

[0233] Step 6:

[0234] The server uses a text-to-speech engine to convert the translated text data into speech data in the target language, rather than the original language, with tone and speed based on the emotion data.

[0235] Step 7:

[0236] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0237] Step 8:

[0238] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0239] Step 9:

[0240] Users can use a dedicated application to change the target language, speech output speed, volume, emotion recognition settings, etc. as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0241] Step 10:

[0242] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, sentiment analysis, and voice output, maintaining natural conversations between multiple languages.

[0243] Example 2

[0244] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0245] Technology for smooth, natural, real-time conversations between multiple languages ​​is important, especially when communicating between different languages. However, current interpretation systems produce mechanical translation results that make it difficult to reflect the user's feelings. Furthermore, there is a lack of ways for users to flexibly change system settings or visually check translated subtitles. This creates a problem of impeding the natural flow of conversation.

[0246] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice input means for acquiring the speech of the other party, a voice recognition means for converting the acquired voice into text, and a translation means for translating the converted text into a target language set by the user. This enables real-time conversation that reflects natural emotions.

[0247] The "voice input means for acquiring the other party's speech" is a device for capturing the other party's speech and acquiring the speech signal as digital data.

[0248] "Speech recognition means for converting acquired speech into text" refers to a technique or device for analyzing acquired speech data and converting it into text data.

[0249] The "translation means for translating the converted text into a target language set by the user" is a technique or device for automatically translating text data into another language set by the user.

[0250] The "audio output means for converting the translated text into audio and outputting it" is a device for converting the translated text data into an audio signal and outputting that audio to the user.

[0251] A "portable device" is an electronic device that can be easily worn by a user and is portable.

[0252] A "computing device" is a computer system that processes data and performs calculations.

[0253] The "communication means for connecting the portable device and the computing device" refers to a technique or device for transmitting and receiving data between the portable device and the computing device.

[0254] "Emotion recognition means for analyzing a user's facial expression and tone of voice to recognize emotions" refers to a technology or device for analyzing and recognizing emotions contained in a user's facial expression and tone of voice.

[0255] "Means for adjusting the tone and speed of the translated speech based on the recognized emotion" refers to a technology or device for appropriately adjusting the tone and speed of the speech based on the emotion recognized by the emotion recognition means.

[0256] This invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system translates what the other person is saying in real time and provides voice output that reflects their emotions, thereby realizing natural conversations between users.

[0257] The system includes the following components:

[0258] 1. Handheld devices:

[0259] It is a glasses-type device worn by the user that has built-in audio input and output means. This device can be audio glasses such as Bose Frames. The device captures the other person's speech, converts it into digital data, and sends it to a computing device.

[0260] 2. Computing equipment:

[0261] The device includes a speech recognition unit, a translation unit, and an emotion recognition unit. The computing unit receives and processes voice data from a mobile device via a communication means such as Bluetooth. The device uses the Google Cloud Speech-to-Text API for speech recognition, the DeepL API for translation, and the Microsoft Azure Emotion API for emotion recognition.

[0262] 3. Means of communication:

[0263] The portable device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data in real time.

[0264] 4. Dedicated applications:

[0265] Users can manage the system settings using a dedicated application installed on their smartphone or tablet, which allows them to select the target language and adjust the speech output speed and volume.

[0266] 5. Emotion recognition means:

[0267] The emotion recognition unit analyzes the user's facial expressions and tone of voice to recognize emotions, and has the function of adjusting the tone and speed of the translated speech based on the recognized emotions.

[0268] (Example)

[0269] Scenario 1: English and Japanese conversation:

[0270] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the server captures the voice through the microphone of the eyeglasses-type device and converts it into text data "What time is the meeting?" using the Google Cloud Speech-to-Text API. The DeepL API is used to translate this to "What time is the meeting?", and if the emotion recognition means recognizes that User A's question contains tension, it generates voice that reflects that tension. Finally, the generated voice data is transmitted to User B through the speaker of the eyeglasses-type device.

[0271] Scenario 2: Configuration Management:

[0272] User B opens the application, changes the target language to French, and adjusts the speech output speed to a calmer tone. This setting information is sent to the server in real time and applied to the entire system. In the next conversation, "What time is the meeting?" is translated to "À quelle heure est la réunion?" and the speech is output in a calm, gentle tone.

[0273] In this way, the present invention provides a system that enables natural conversation between multiple languages ​​and realizes smooth communication in a variety of situations.

[0274] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0275] Step 1:

[0276] Acquiring voice input

[0277] When a user initiates a conversation, a microphone on the mobile device picks up what the other person is saying, using audio glasses such as Bose Frames.

[0278] Input: Other person's speech (audio data)

[0279] Data processing: Acquisition of audio data

[0280] Output: Digital audio data

[0281] Specific operation: When user A says "What time is the meeting?", the voice is recorded as digital data via the microphone of the portable device.

[0282] Step 2:

[0283] Sending audio data

[0284] The digital audio data captured by the portable device is transmitted to the computing unit via Bluetooth.

[0285] Input: Digital audio data

[0286] Data processing: Data transmission

[0287] Output: Received digital audio data

[0288] Specific operation: A portable device transmits digital audio data to a computing device via Bluetooth.

[0289] Step 3:

[0290] Converting audio data to text

[0291] The computing device uses the Google Cloud Speech-to-Text API to analyze the voice data and convert it into text data.

[0292] Input: Digital audio data

[0293] Data processing: speech recognition and text conversion

[0294] Output: Text data

[0295] Specific operation: The computing device converts the voice data into text: "What time is the meeting?"

[0296] Step 4:

[0297] Text data translation

[0298] The computing device uses the DeepL API to translate the text data into the target language.

[0299] Input: Text data

[0300] Data processing: Text translation

[0301] Output: Text data in the target language

[0302] Specific behavior: The computing device translates "What time is the meeting?" to "What time is the meeting?"

[0303] Step 5:

[0304] Emotion recognition

[0305] The computing device's emotion recognition means (such as Microsoft Azure's Emotion API) analyzes the user's facial expressions and vocal tone to recognize their emotions.

[0306] Input: User's facial expression data and voice tone

[0307] Data processing: Emotion analysis and recognition

[0308] Output: Recognized emotion data

[0309] Specific operation: The emotion recognition means reads the user's level of tension from their facial expressions and voice.

[0310] Step 6:

[0311] Adjust the tone and speed of your voice

[0312] The computing device adjusts the tone and rate of the translated speech based on the recognized emotion.

[0313] Input: Translation text data and emotion data

[0314] Data processing: adjusting the tone and speed of the voice

[0315] Output: Modified audio data

[0316] What it does: The computing device converts the translated text into a voice tone that reflects tension.

[0317] Step 7:

[0318] Generate audio output

[0319] The computing device generates tailored voice data using a TTS engine such as Amazon Polly.

[0320] Input: Modified audio data

[0321] Data processing: voice synthesis

[0322] Output: Generated audio data

[0323] Specific operation: The computing device synthesizes the adjusted voice data and prepares it for output.

[0324] Step 8:

[0325] Audio data output

[0326] The server sends the generated audio data to the portable device via Bluetooth, where it is transmitted to the user through the speaker.

[0327] Input: Generated audio data

[0328] Data processing: Sending voice data

[0329] Output: Audio output

[0330] Specific Actions: User B hears the translated audio of "Hello, how are you?" through the speaker of their mobile device.

[0331] Through this series of processes, the system realizes natural dialogue between users.

[0332] (Application example 2)

[0333] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0334] Currently, achieving natural translation and emotionally-rich speech output in real time is a challenging task in multilingual conversations and content distribution. Furthermore, in live streaming and recorded content where multiple language-speaking viewers simultaneously participate, viewers often cannot accurately convey the emotions of the performer or the video while watching in their own language. Therefore, multilingual streaming services are required to ensure that viewers can enjoy a natural and emotionally-rich experience.

[0335] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for acquiring the other party's speech via a voice input means; voice recognition means for converting the acquired voice into text; translation means for translating the converted text into a target language set by the user; voice output means for converting the translated text into voice and outputting it; a glasses-type device equipped with the voice input means and the voice output means; a computing device including the voice recognition means and the translation means; communication means for connecting the glasses-type device and the computing device; display means for translating the conversations of performers and videos in real time and displaying subtitles in a language selected by the viewer; means for adjusting the translation results using emotion analysis and displaying them as natural conversational text; and a dedicated application for the viewer to manage settings. This enables viewers to enjoy a natural and emotionally rich experience in a multilingual distribution service.

[0336] The "means for acquiring the other party's speech by voice input means" is a technique for acquiring the other party's speech by using a voice input device such as a microphone.

[0337] "Speech recognition means" refers to software or algorithms that convert captured speech data into text.

[0338] The "translation means" is a technology that translates the text converted by the speech recognition means into a target language set by the user.

[0339] "Audio output means" refers to technology that converts the translated text back into audio and outputs it via a speaker or the like.

[0340] The "eyeglasses-type device" is an eyeglasses-type device equipped with an audio input means and an audio output means, and is a portable communication device.

[0341] The term "arithmetic unit" is a general term for hardware and software for performing various processes such as speech recognition means and translation means.

[0342] The "communication means" is a technology for connecting the eyeglass-type device and the computing device and transmitting and receiving data.

[0343] "Display means" refers to devices or technologies that translate the dialogue of the performers or in the video in real time and display subtitles in the language selected by the viewer.

[0344] "Emotion analysis means" is a technology that adjusts the translation results based on the user's emotions and outputs them in a natural conversational style.

[0345] "Dedicated application" is software that allows viewers to manage and customize system settings.

[0346] This invention provides a system for realizing real-time translation and emotion recognition in natural conversations between multiple languages ​​and content distribution. Specific embodiments of this system are described below.

[0347] System Configuration

[0348] The system includes the following major components:

[0349] 1. Voice input means: A device for capturing what the other person is saying. For example, a microphone.

[0350] 2. Speech recognition tool: Software or algorithms that convert captured voice data into text. For example, we use the Google Cloud Speech-to-Text API.

[0351] 3. Translation method: The technology that translates the text into the target language. For example, using the Amazon Translate API.

[0352] 4. Voice output method: Technology that converts the translated text into voice and outputs it through a speaker. For example, IBM Watson Text to Speech API is used.

[0353] 5. Glasses-type device: A portable device equipped with audio input and output means.

[0354] 6. Computing unit: An integrated system of hardware and software for performing various processes.

[0355] 7. Communication method: Technology that connects the glasses device to the computing device, such as Bluetooth or Wi-Fi.

[0356] 8. Display means: A device for displaying the translation results to the viewer as subtitles in real time.

[0357] 9. Sentiment analysis: Technology to analyze the user's emotions and output the translation results in a natural conversational style. For example, we use the Microsoft Azure Emotion API.

[0358] 10. Dedicated application: Software that allows viewers to manage their system settings. For example, using React Native as the platform.

[0359] System Operation

[0360] 1. Acquiring voice input and converting it to text

[0361] When a user starts talking, the microphone of the eyeglass device picks up what is being said. The picked-up voice data is sent to a computing device and converted into text data by a voice recognition means.

[0362] 2. Text Translation

[0363] The text converted by the speech recognition means is translated into a target language set by the user using the translation means.

[0364] 3. Emotion recognition and voice output adjustment

[0365] The translated text is converted into voice data by taking into account the user's emotions through emotion analysis, and this voice data is output in an appropriate tone and played to the audience.

[0366] 4. Displaying the translation results

[0367] The translated text is provided to the viewer in real time as subtitles using a display means.

[0368] Usage example

[0369] For example, if a speaker says "Hello, everyone! Welcome to our event!" during a live stream, the system will convert this into text in real time and translate it into the target language (e.g., Japanese). The translated "Hello, everyone! Welcome to our event!" will be played in an "excited" tone based on emotion analysis and simultaneously displayed as subtitles.

[0370] Prompt example

[0371] Below are some examples of prompts used in this system:

[0372] Design a real-time translation system for live streaming that provides subtitles and speech translation according to the viewer's language of choice and realizes natural conversational tone that takes into account the speaker's emotions. This system includes speech recognition, text translation, emotion recognition, and speech synthesis modules. Each module should be implemented using a specific API, and viewers should be able to change the settings.

[0373] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0374] Step 1:

[0375] Acquiring voice input:

[0376] When a user starts speaking, the microphone on the eyeglasses picks up what is being said. The input voice data is sent as a digital signal to a computing device, where it is converted into a digital audio format such as linear PCM.

[0377] Step 2:

[0378] Speech to text:

[0379] The server sends the voice data received by the computing device to the Google Cloud Speech-to-Text API. The API analyzes the voice data and converts it into corresponding text data. The input is digital voice data, and the output is text data that converts the voice into text.

[0380] Step 3:

[0381] Text Translation:

[0382] The server sends the converted data to the Amazon Translate API. The input is the text data obtained by the speech recognition means, and the output is the text data translated into the target language. Here, since the user has set the target language in advance, the data is translated into the appropriate language based on that setting.

[0383] Step 4:

[0384] Emotion recognition:

[0385] The server sends the translated text data and the original audio data to the Microsoft Azure Emotion API to analyze the user's emotions. The input is the translated text data and audio data, and the output is the emotion analysis results and the emotion labels and scores based on them.

[0386] Step 5:

[0387] Text-to-Speech:

[0388] The server uses the IBM Watson Text to Speech API to convert the translated text data into speech data that reflects the results of sentiment analysis. The input is the translated text data and emotion label, and the output is synthesized speech data. This speech data is generated based on the voice characteristics (speed, tone, etc.) set by the user.

[0389] Step 6:

[0390] Displaying subtitles:

[0391] The computing device transmits the translated text data to the display device in real time and displays it as subtitles. The input is the translated text data, and the output is the subtitles displayed on the screen in real time, allowing viewers to visually confirm the translation content.

[0392] Step 7:

[0393] Audio Output:

[0394] The translated and emotion-analyzed speech data is output to the user from the speaker of the eyeglasses via a computing device. The input is synthesized speech data, and the output is speech that reaches the user's ears. This allows the user to hear the translated speech in real time.

[0395] Step 8:

[0396] Settings management:

[0397] The user manages various system settings (target language, audio characteristics, subtitle display, etc.) through a dedicated application. The input is the user's setting information, and the output is the setting data sent to the computing device. This allows the translation environment desired by the user to be properly applied.

[0398] The above is an explanation of the processing steps of the system program that realizes the application example, as well as its specific operations and input / output.

[0399] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0400] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0401] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0402] [Second embodiment]

[0403] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0404] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0405] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0406] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0407] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0409] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0410] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0411] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0412] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0413] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0414] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0415] The present invention provides a system for translating conversations between multiple languages ​​in real time and supporting natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, and a display means.

[0416] System configuration overview

[0417] 1. Eyeglass-type device worn by the user

[0418] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0419] 2. Arithmetic device

[0420] The apparatus includes a speech recognition unit for converting speech into text and a translation unit for translating the text into a target language. The computing unit receives the speech data obtained from the eyeglasses-type device and performs appropriate processing.

[0421] 3. Means of communication

[0422] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0423] 4. Dedicated application

[0424] Users use a dedicated application to manage the system's settings, such as the target language for translation and the voice speed of the audio output.

[0425] 5. Display means

[0426] A display means is provided to visually convey the translation results to the user, which makes it possible to display the translation content as subtitles.

[0427] System program processing

[0428] Acquiring voice input and converting it to text

[0429] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0430] Text translation

[0431] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0432] Voice output of translation results

[0433] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0434] Server and device communication

[0435] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0436] Specific examples

[0437] Scenario 1: English and Japanese Conversation

[0438] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text, "What time is the meeting?", which the translation means translates into "What time is the meeting?" The translated text is converted into speech and played back to User B.

[0439] Scenario 2: Configuration Management

[0440] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation.

[0441] In this way, the present invention enables natural conversation between multiple languages ​​and minimizes the time lost due to interpretation.

[0442] The processing flow will be explained below.

[0443] Step 1:

[0444] The user puts on the audio glasses and starts the system, which enables the microphone and speaker in the audio glasses.

[0445] Step 2:

[0446] The device (Audio Glasses) picks up what the other person is saying with a microphone. This audio data is captured in real time and temporarily stored in the device.

[0447] Step 3:

[0448] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0449] Step 4:

[0450] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0451] Step 5:

[0452] The server receives the translated text data and converts it into audio data using a text-to-speech engine, which is in the target language rather than the original language.

[0453] Step 6:

[0454] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0455] Step 7:

[0456] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0457] Step 8:

[0458] Users can use a dedicated application to change settings such as the target language, speech output speed, and volume as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0459] Step 9:

[0460] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, and voice output, maintaining natural conversation between multiple languages.

[0461] Example 1

[0462] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0463] In real-time communication between multiple languages, language barriers exist, making smooth conversation difficult. Existing translation systems have issues with translation accuracy and speed, making them insufficient for real-time conversation support. In particular, it is difficult to continue the flow of conversation without interruption, and this aspect needs improvement.

[0464] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0465] In this invention, the server includes means for acquiring speech from the other party via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, means for converting the translated text into voice and outputting it, a wearable device equipped with the voice input means and the voice output means, an information processing device including the voice recognition means and the translation means, communication means for connecting the wearable device and the information processing device, and communication means for transmitting and receiving data in real time, thereby enabling users to have natural real-time conversations between multiple languages.

[0466] "Voice input means" is a device that has the function of acquiring the speech of the other party.

[0467] The "voice recognition means" is a device that has the function of converting acquired voice into text data.

[0468] The "translation means" is a device that has the function of translating text data into a target language set by the user.

[0469] "Speech output means" refers to a device that has the function of converting translated text into speech and outputting it.

[0470] A "wearable device" is a wearable device equipped with an audio input means and an audio output means.

[0471] The "information processing device" is a processing device that includes a speech recognition means and a translation means.

[0472] A "communication means" is a device that connects a wearable device to an information processing device and has the function of sending and receiving data.

[0473] "Communication means for transmitting and receiving data in real time" refers to a communication technology that enables the transmission and reception of data in real time.

[0474] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system includes a wearable device (e.g., a glasses-type device), an information processing device, communication means, a dedicated application for managing user settings, and display means. Each component and its operation are described in detail below.

[0475] Wearable devices

[0476] The wearable device is equipped with a voice input means and a voice output means. It is worn by the user and looks like a glasses-type device. This device has the following functions:

[0477] Voice input means: A microphone is built in to capture what the other person is saying.

[0478] Audio output means: A built-in speaker is included, allowing the user to hear the translated speech.

[0479] Information processing device

[0480] The information processing device includes a speech recognition unit and a translation unit. These units analyze the speech data acquired from the wearable device and perform translation processing. Specific components are as follows:

[0481] Speech recognition means: Converts acquired voice data into text data.

[0482] Translation method: Translates text data into the target language set by the user. The translation engine uses a generative AI model based on a neural network, for example.

[0483] communication means

[0484] The communication means connects the wearable device to the information processing device and transmits and receives data in real time. The communication means includes:

[0485] Bluetooth and Wi-Fi: Used to send and receive voice and translation data.

[0486] Dedicated application

[0487] A dedicated application allows users to manage the system settings. Through this application, the following settings can be configured:

[0488] Select target language: Set the language to translate into.

[0489] Audio output adjustment: Set the audio output speed, etc.

[0490] Display means

[0491] The display means is used to visually convey the translation results to the user. Specifically, it has the following functions:

[0492] Subtitle display: It is possible to display the translated text as subtitles.

[0493] Specific examples

[0494] Example 1: English-Japanese conversation

[0495] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the wearable device picks up the speech and sends it to an information processing device. The speech recognition means converts this speech into text data, "What time is the meeting?", which is then translated by the translation means into "What time is the meeting?" The translated text is converted into speech and played back to User B. The speaker in the wearable device outputs the speech, "What time is the meeting?"

[0496] Example 2: Configuration Management

[0497] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. For example, the target language is changed from Japanese to French, and the speech output speed is set to 1.25 times faster. This information is immediately sent to the information processing device, and the new settings are reflected from the next conversation. Every time the user starts a conversation, translation is performed based on the application settings.

[0498] Prompt Sentence Examples

[0499] Here are some examples of prompts to input to a generative AI model:

[0500] If user A speaks "Hello, how are you?" into a wearable device, how can I convert that speech into text and translate it into Japanese, the target language specified by user B?

[0501] In this way, the present invention enables natural conversations between multiple languages ​​and minimizes the time lost due to interpretation. Users can manage settings through a dedicated application, enabling smooth communication in real time.

[0502] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0503] Step 1: Getting voice input

[0504] When a user starts a conversation, the microphone on the device (wearable device) picks up what the other person is saying. When User A says, "Hello, how are you?", the voice is input to the device through the microphone. The device converts this real-time voice input into digital voice data and prepares for the next step. Here, the input is analog voice, and the output is digital voice data.

[0505] Step 2: Sending audio data

[0506] The terminal (wearable device) transmits digital voice data to the server via Bluetooth or Wi-Fi. For example, voice data such as "Hello, how are you?" is delivered to the server via wireless communication. The input here is the digital voice data, and the output is the voice data transmitted to the server. During this transmission, the communication indicator LED on the terminal lights up to indicate the status of data transmission.

[0507] Step 3: Speech to Text

[0508] The server inputs the received voice data into a voice recognition means and converts it into text data. For example, the voice "Hello, how are you?" is converted into text "Hello, how are you?" The server uses a voice recognition algorithm to analyze the voice data and generate output in text format. Here, the input is the voice data sent to the server, and the output is text data. A processing progress bar on the server operates to display the progress of the conversion.

[0509] Step 4: Translate the text

[0510] The server passes the text data obtained by the speech recognition means to the translation means, which translates it into the target language set by the user. For example, if User B has set Japanese as the target language, "Hello, how are you?" will be translated into "Hello, how are you?" A generative AI model is used to perform highly accurate translation. The input here is the recognized text data, and the output is the translated text data. The original text and the translated text are recorded in the server log.

[0511] Step 5: Audio output of translation results

[0512] The server passes the translated text data to the voice output means, which converts it into voice data. The converted voice data is sent from the server to the terminal (wearable device). The converted voice is played from the terminal's speaker. For example, User B can hear the voice saying, "Hello, how are you?" Here, the input is the translated text data, and the output is playable voice data. An LED flashes to indicate that the terminal's speaker is outputting voice.

[0513] Step 6: Server and device communication

[0514] The server and the terminal (wearable device) are always in communication, sending and receiving data in real time. This allows each step of the conversation to be processed without delay. Every time the user speaks, the communication status can be confirmed by the communication indicator on the terminal lighting up. The input here is the user's voice and setting information, and the output is smooth conversation through real-time data transmission and reception.

[0515] Step 7: Configuration Management

[0516] The user opens a dedicated application to manage settings such as the target language of the conversation and the speech output speed. When the user changes the settings, the setting information is immediately sent to the server and applied throughout the system. For example, the user changes the target language from Japanese to French and sets the speech output speed to 1.25x. The new settings will be applied from the next conversation. The input here is the setting change information made by the user, and the output is the updated system settings.

[0517] In this way, the system translates conversations between multiple languages ​​in real time, providing smooth and natural communication.

[0518] (Application example 1)

[0519] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] Many brick-and-mortar stores face the problem of inability to communicate smoothly with tourists who speak foreign languages. This communication barrier can make it difficult for store staff and tourists to exchange accurate information, resulting in a decline in service quality. Another issue is the time loss caused by the inability to obtain translation results immediately. To solve these problems, a system that supports multilingual conversations in real time is needed.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0522] In this invention, the server includes means for acquiring the other person's speech via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, voice output means for converting the translated text into voice and outputting it, a glasses-type device equipped with the voice input means and the voice output means, a computing device including the voice recognition means and the translation means, communication means for connecting the glasses-type device and a smartphone to the computing device, means for the user to manage settings using a dedicated application, and display means for displaying the translation results as subtitles or means for displaying them on the smartphone. This enables smooth communication between store staff and tourists across language barriers.

[0523] "Audio input means" refers to a device or technology for capturing audio.

[0524] A "speech recognition means" is a technique or device that converts captured speech into text.

[0525] A "translation means" is a technique or device that translates text into a target language set by the user.

[0526] "Audio output means" refers to a device or technology that converts translated text into audio and plays it back.

[0527] An "eyeglasses-type device" is an electronic device in the shape of glasses that has built-in audio input and audio output functions.

[0528] A "computing device" is an electronic device that includes a processor that performs speech recognition and translation.

[0529] "Communication means" refers to a technology or device for transmitting and receiving data between different devices, and includes wireless communication technologies such as Bluetooth and Wi-Fi.

[0530] A "dedicated application" is software that allows users to configure the system and runs on a smartphone or tablet.

[0531] "Display means" means a device or technology for visually displaying the translation result in text, including smart glasses or a smartphone display.

[0532] The present invention provides a real-time translation system that supports communication between users who speak different languages ​​in a physical store. The system specifically includes the following configuration and operation.

[0533] System configuration

[0534] 1. Voice input method

[0535] The server uses the microphone of the smart glasses or smartphone worn by the user to capture the speech of the interlocutor, thereby collecting the speech content as voice data.

[0536] 2. Voice Recognition Method

[0537] The captured voice data is sent to a server and converted into text data using voice recognition software, using the "speech_recognition" library.

[0538] 3. Translation Methods

[0539] The text data acquired by the speech recognition means is translated into the target language by a translation engine on the server. This translation process uses the "googletrans" library.

[0540] 4. Audio output means

[0541] The translated text is then converted back into audio data in the target language using speech synthesis technology, and the server uses the gTTS (Google Text-to-Speech) library to generate an audio file that is played on the user's smart glasses or smartphone speakers.

[0542] 5. Display means

[0543] The translation result is displayed as text on the smart glasses display or on the smartphone screen, allowing users to check the translation not only aloud but also visually.

[0544] Hardware and Software Use

[0545] The hardware used includes smart glasses and smartphones, which connect to the server via Bluetooth or Wi-Fi.

[0546] The software uses:

[0547] Speech Recognition: speech_recognition library

[0548] Translation: GoogleTrans Library

[0549] Speech synthesis: gTTS library

[0550] Audio playback: pygame library

[0551] Bluetooth communication: pybluez library

[0552] Specific examples

[0553] When a store staff member says in Japanese, "How do you use this product?", the microphone in the smart glasses or smartphone picks up the speech and converts it into text using speech recognition. This text is then translated into English by a translation engine into "How do you use this product?" The translation result is converted into speech and played back through the speaker of the staff member's smart glasses or smartphone, while also being displayed as text on the screen.

[0554] Prompt Sentence Examples

[0555] "Design an application for smart glasses that uses voice input to translate conversations between multiple languages ​​in real time, and outputs them in voice and text format."

[0556] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0557] Step 1:

[0558] The user makes a speech. The speech input means acquires the speech through the microphone of the smart glasses or smartphone. At this time, the acquired speech data is a raw speech signal. The server receives the corresponding speech data.

[0559] Step 2:

[0560] The server converts the received voice data into text data using speech recognition software (speech_recognition library). The input is the acquired voice data, and the output is the data converted from the speech into text format. Through this conversion, the voice information is expressed as text information.

[0561] Step 3:

[0562] The server inputs the converted text data into a translation engine (GoogleTrans library) and translates it into the target language. The input is text data, and the output is translated text data. This process converts the user's speech into a language the other person can understand.

[0563] Step 4:

[0564] The server converts the translated text data into audio data using speech synthesis technology (gTTS library). The input is the translated text data, and the output is an audio file. The server generates this audio file and prepares it for playback.

[0565] Step 5:

[0566] The server sends the generated audio file to the speaker of the smart glasses or smartphone and plays it through the audio output means, allowing the other person to hear the translated audio. At the same time, the translation result text is displayed on the display of the smart glasses or smartphone. The input is the audio file and text data, and the output is the playback of the translated audio and the text display.

[0567] Step 6:

[0568] Users manage system settings using a dedicated application. When a user changes the target language, for example, the setting information is sent to the server and is immediately reflected throughout the system. The input is the user's setting information, and the output is a reflection of the updated setting information.

[0569] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0570] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, a display means, and an emotion engine that recognizes the user's emotions.

[0571] System configuration overview

[0572] 1. Eyeglass-type device worn by the user

[0573] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0574] 2. Arithmetic device

[0575] The device includes a speech recognition unit that converts speech into text, a translation unit that translates the text into a target language, and an emotion engine that recognizes the user's emotions. The computing unit receives the speech data acquired from the glasses-type device and performs appropriate processing.

[0576] 3. Means of communication

[0577] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0578] 4. Dedicated application

[0579] Users can use a dedicated application to manage the system's settings, such as the target language for translation, voice output voice speed, volume, etc. There is also a training mode to improve the accuracy of emotion recognition.

[0580] 5. Display means

[0581] The system is equipped with a display means for visually conveying the translation results to the user, which makes it possible to display the translation content as subtitles.

[0582] 6. Emotion Engine

[0583] The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions, and adjusts the tone and speed of the translated speech based on the recognized emotion, ensuring natural conversation.

[0584] System program processing

[0585] Acquiring voice input and converting it to text

[0586] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0587] Text translation

[0588] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0589] Emotion recognition and voice output adjustment

[0590] The translated text data is converted into voice data by an emotion engine, taking into account the user's emotions. For example, if the user is excited, the translated voice will be played in a similarly excited tone. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and choice of words used.

[0591] Voice output of translation results

[0592] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0593] Server and device communication

[0594] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0595] Specific examples

[0596] Scenario 1: English and Japanese Conversation

[0597] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text "What time is the meeting?", and the translation means translates it into "What time is the meeting?" The emotion engine recognizes that User A's question contains tension and plays back the translated speech in a voice that reflects that tension.

[0598] Scenario 2: Configuration Management

[0599] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation. For example, if User B now selects French, "What time is the meeting?" is translated as "À quelle heure est la réunion?" and the emotion engine plays it in a calm, gentle tone.

[0600] In this way, the present invention provides an interpretation system that enables natural conversation between multiple languages ​​and is adaptable to various situations.

[0601] The processing flow will be explained below.

[0602] Step 1:

[0603] The user puts on the audio glasses and starts the system, which activates the microphone and speaker in the audio glasses and starts the emotion engine, preparing to analyze the user's facial expressions and tone of voice.

[0604] Step 2:

[0605] The device (Audio Glasses) picks up what the other person is saying with a microphone. This voice data is captured in real time and temporarily stored in the device. At the same time, the emotion engine records the user's facial expressions and voice tone and generates emotion data.

[0606] Step 3:

[0607] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0608] Step 4:

[0609] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0610] Step 5:

[0611] The server receives the translated text data and the emotion data generated by the emotion engine, which then adjusts the tone and speed of the translated speech based on that data.

[0612] Step 6:

[0613] The server uses a text-to-speech engine to convert the translated text data into speech data in the target language, rather than the original language, with tone and speed based on the emotion data.

[0614] Step 7:

[0615] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0616] Step 8:

[0617] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0618] Step 9:

[0619] Users can use a dedicated application to change the target language, speech output speed, volume, emotion recognition settings, etc. as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0620] Step 10:

[0621] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, sentiment analysis, and voice output, maintaining natural conversations between multiple languages.

[0622] Example 2

[0623] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0624] Technology for smooth, natural, real-time conversations between multiple languages ​​is important, especially when communicating between different languages. However, current interpretation systems produce mechanical translation results that make it difficult to reflect the user's feelings. Furthermore, there is a lack of ways for users to flexibly change system settings or visually check translated subtitles. This creates a problem of impeding the natural flow of conversation.

[0625] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice input means for acquiring the speech of the other party, a voice recognition means for converting the acquired voice into text, and a translation means for translating the converted text into a target language set by the user. This enables real-time conversation that reflects natural emotions.

[0626] The "voice input means for acquiring the other party's speech" is a device for capturing the other party's speech and acquiring the speech signal as digital data.

[0627] "Speech recognition means for converting acquired speech into text" refers to a technique or device for analyzing acquired speech data and converting it into text data.

[0628] The "translation means for translating the converted text into a target language set by the user" is a technique or device for automatically translating text data into another language set by the user.

[0629] The "audio output means for converting the translated text into audio and outputting it" is a device for converting the translated text data into an audio signal and outputting that audio to the user.

[0630] A "portable device" is an electronic device that can be easily worn by a user and is portable.

[0631] A "computing device" is a computer system that processes data and performs calculations.

[0632] The "communication means for connecting the portable device and the computing device" refers to a technique or device for transmitting and receiving data between the portable device and the computing device.

[0633] "Emotion recognition means for analyzing a user's facial expression and tone of voice to recognize emotions" refers to a technology or device for analyzing and recognizing emotions contained in a user's facial expression and tone of voice.

[0634] "Means for adjusting the tone and speed of the translated speech based on the recognized emotion" refers to a technology or device for appropriately adjusting the tone and speed of the speech based on the emotion recognized by the emotion recognition means.

[0635] This invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system translates what the other person is saying in real time and provides voice output that reflects their emotions, thereby realizing natural conversations between users.

[0636] The system includes the following components:

[0637] 1. Handheld devices:

[0638] It is a glasses-type device worn by the user that has built-in audio input and output means. This device can be audio glasses such as Bose Frames. The device captures the other person's speech, converts it into digital data, and sends it to a computing device.

[0639] 2. Computing equipment:

[0640] The device includes a speech recognition unit, a translation unit, and an emotion recognition unit. The computing unit receives and processes voice data from a mobile device via a communication means such as Bluetooth. The device uses the Google Cloud Speech-to-Text API for speech recognition, the DeepL API for translation, and the Microsoft Azure Emotion API for emotion recognition.

[0641] 3. Means of communication:

[0642] The portable device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data in real time.

[0643] 4. Dedicated applications:

[0644] Users can manage the system settings using a dedicated application installed on their smartphone or tablet, which allows them to select the target language and adjust the speech output speed and volume.

[0645] 5. Emotion recognition means:

[0646] The emotion recognition unit analyzes the user's facial expressions and tone of voice to recognize emotions, and has the function of adjusting the tone and speed of the translated speech based on the recognized emotions.

[0647] (Example)

[0648] Scenario 1: English and Japanese conversation:

[0649] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the server captures the voice through the microphone of the eyeglasses-type device and converts it into text data "What time is the meeting?" using the Google Cloud Speech-to-Text API. The DeepL API is used to translate this to "What time is the meeting?", and if the emotion recognition means recognizes that User A's question contains tension, it generates voice that reflects that tension. Finally, the generated voice data is transmitted to User B through the speaker of the eyeglasses-type device.

[0650] Scenario 2: Configuration Management:

[0651] User B opens the application, changes the target language to French, and adjusts the speech output speed to a calmer tone. This setting information is sent to the server in real time and applied to the entire system. In the next conversation, "What time is the meeting?" is translated to "À quelle heure est la réunion?" and the speech is output in a calm, gentle tone.

[0652] In this way, the present invention provides a system that enables natural conversation between multiple languages ​​and realizes smooth communication in a variety of situations.

[0653] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0654] Step 1:

[0655] Acquiring voice input

[0656] When a user initiates a conversation, a microphone on the mobile device picks up what the other person is saying, using audio glasses such as Bose Frames.

[0657] Input: Other person's speech (audio data)

[0658] Data processing: Acquisition of audio data

[0659] Output: Digital audio data

[0660] Specific operation: When user A says "What time is the meeting?", the voice is recorded as digital data via the microphone of the portable device.

[0661] Step 2:

[0662] Sending audio data

[0663] The digital audio data captured by the portable device is transmitted to the computing unit via Bluetooth.

[0664] Input: Digital audio data

[0665] Data processing: Data transmission

[0666] Output: Received digital audio data

[0667] Specific operation: A portable device transmits digital audio data to a computing device via Bluetooth.

[0668] Step 3:

[0669] Converting audio data to text

[0670] The computing device uses the Google Cloud Speech-to-Text API to analyze the voice data and convert it into text data.

[0671] Input: Digital audio data

[0672] Data processing: speech recognition and text conversion

[0673] Output: Text data

[0674] Specific operation: The computing device converts the voice data into text: "What time is the meeting?"

[0675] Step 4:

[0676] Text data translation

[0677] The computing device uses the DeepL API to translate the text data into the target language.

[0678] Input: Text data

[0679] Data processing: Text translation

[0680] Output: Text data in the target language

[0681] Specific behavior: The computing device translates "What time is the meeting?" to "What time is the meeting?"

[0682] Step 5:

[0683] Emotion recognition

[0684] The computing device's emotion recognition means (such as Microsoft Azure's Emotion API) analyzes the user's facial expressions and vocal tone to recognize their emotions.

[0685] Input: User's facial expression data and voice tone

[0686] Data processing: Emotion analysis and recognition

[0687] Output: Recognized emotion data

[0688] Specific operation: The emotion recognition means reads the user's level of tension from their facial expressions and voice.

[0689] Step 6:

[0690] Adjust the tone and speed of your voice

[0691] The computing device adjusts the tone and rate of the translated speech based on the recognized emotion.

[0692] Input: Translation text data and emotion data

[0693] Data processing: adjusting the tone and speed of the voice

[0694] Output: Modified audio data

[0695] What it does: The computing device converts the translated text into a voice tone that reflects tension.

[0696] Step 7:

[0697] Generate audio output

[0698] The computing device generates tailored voice data using a TTS engine such as Amazon Polly.

[0699] Input: Modified audio data

[0700] Data processing: voice synthesis

[0701] Output: Generated audio data

[0702] Specific operation: The computing device synthesizes the adjusted voice data and prepares it for output.

[0703] Step 8:

[0704] Audio data output

[0705] The server sends the generated audio data to the portable device via Bluetooth, where it is transmitted to the user through the speaker.

[0706] Input: Generated audio data

[0707] Data processing: Sending voice data

[0708] Output: Audio output

[0709] Specific Actions: User B hears the translated audio of "Hello, how are you?" through the speaker of their mobile device.

[0710] Through this series of processes, the system realizes natural dialogue between users.

[0711] (Application example 2)

[0712] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0713] Currently, achieving natural translation and emotionally-rich speech output in real time is a challenging task in multilingual conversations and content distribution. Furthermore, in live streaming and recorded content where multiple language-speaking viewers simultaneously participate, viewers often cannot accurately convey the emotions of the performer or the video while watching in their own language. Therefore, multilingual streaming services are required to ensure that viewers can enjoy a natural and emotionally-rich experience.

[0714] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for acquiring the other party's speech via a voice input means; voice recognition means for converting the acquired voice into text; translation means for translating the converted text into a target language set by the user; voice output means for converting the translated text into voice and outputting it; a glasses-type device equipped with the voice input means and the voice output means; a computing device including the voice recognition means and the translation means; communication means for connecting the glasses-type device and the computing device; display means for translating the conversations of performers and videos in real time and displaying subtitles in a language selected by the viewer; means for adjusting the translation results using emotion analysis and displaying them as natural conversational text; and a dedicated application for the viewer to manage settings. This enables viewers to enjoy a natural and emotionally rich experience in a multilingual distribution service.

[0715] The "means for acquiring the other party's speech by voice input means" is a technique for acquiring the other party's speech by using a voice input device such as a microphone.

[0716] "Speech recognition means" refers to software or algorithms that convert captured speech data into text.

[0717] The "translation means" is a technology that translates the text converted by the speech recognition means into a target language set by the user.

[0718] "Audio output means" refers to technology that converts the translated text back into audio and outputs it via a speaker or the like.

[0719] The "eyeglasses-type device" is an eyeglasses-type device equipped with an audio input means and an audio output means, and is a portable communication device.

[0720] The term "arithmetic unit" is a general term for hardware and software for performing various processes such as speech recognition means and translation means.

[0721] The "communication means" is a technology for connecting the eyeglass-type device and the computing device and transmitting and receiving data.

[0722] "Display means" refers to devices or technologies that translate the dialogue of the performers or in the video in real time and display subtitles in the language selected by the viewer.

[0723] "Emotion analysis means" is a technology that adjusts the translation results based on the user's emotions and outputs them in a natural conversational style.

[0724] "Dedicated application" is software that allows viewers to manage and customize system settings.

[0725] This invention provides a system for realizing real-time translation and emotion recognition in natural conversations between multiple languages ​​and content distribution. Specific embodiments of this system are described below.

[0726] System Configuration

[0727] The system includes the following major components:

[0728] 1. Voice input means: A device for capturing what the other person is saying. For example, a microphone.

[0729] 2. Speech recognition tool: Software or algorithms that convert captured voice data into text. For example, we use the Google Cloud Speech-to-Text API.

[0730] 3. Translation method: The technology that translates the text into the target language. For example, using the Amazon Translate API.

[0731] 4. Voice output method: Technology that converts the translated text into voice and outputs it through a speaker. For example, IBM Watson Text to Speech API is used.

[0732] 5. Glasses-type device: A portable device equipped with audio input and output means.

[0733] 6. Computing unit: An integrated system of hardware and software for performing various processes.

[0734] 7. Communication method: Technology that connects the glasses device to the computing device, such as Bluetooth or Wi-Fi.

[0735] 8. Display means: A device for displaying the translation results to the viewer as subtitles in real time.

[0736] 9. Sentiment analysis: Technology to analyze the user's emotions and output the translation results in a natural conversational style. For example, we use the Microsoft Azure Emotion API.

[0737] 10. Dedicated application: Software that allows viewers to manage their system settings. For example, using React Native as the platform.

[0738] System Operation

[0739] 1. Acquiring voice input and converting it to text

[0740] When a user starts talking, the microphone of the eyeglass device picks up what is being said. The picked-up voice data is sent to a computing device and converted into text data by a voice recognition means.

[0741] 2. Text Translation

[0742] The text converted by the speech recognition means is translated into a target language set by the user using the translation means.

[0743] 3. Emotion recognition and voice output adjustment

[0744] The translated text is converted into voice data by taking into account the user's emotions through emotion analysis, and this voice data is output in an appropriate tone and played to the audience.

[0745] 4. Displaying the translation results

[0746] The translated text is provided to the viewer in real time as subtitles using a display means.

[0747] Usage example

[0748] For example, if a speaker says "Hello, everyone! Welcome to our event!" during a live stream, the system will convert this into text in real time and translate it into the target language (e.g., Japanese). The translated "Hello, everyone! Welcome to our event!" will be played in an "excited" tone based on emotion analysis and simultaneously displayed as subtitles.

[0749] Prompt example

[0750] Below are some examples of prompts used in this system:

[0751] Design a real-time translation system for live streaming that provides subtitles and speech translation according to the viewer's language of choice and realizes natural conversational tone that takes into account the speaker's emotions. This system includes speech recognition, text translation, emotion recognition, and speech synthesis modules. Each module should be implemented using a specific API, and viewers should be able to change the settings.

[0752] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0753] Step 1:

[0754] Acquiring voice input:

[0755] When a user starts speaking, the microphone on the eyeglasses picks up what is being said. The input voice data is sent as a digital signal to a computing device, where it is converted into a digital audio format such as linear PCM.

[0756] Step 2:

[0757] Speech to text:

[0758] The server sends the voice data received by the computing device to the Google Cloud Speech-to-Text API. The API analyzes the voice data and converts it into corresponding text data. The input is digital voice data, and the output is text data that converts the voice into text.

[0759] Step 3:

[0760] Text Translation:

[0761] The server sends the converted data to the Amazon Translate API. The input is the text data obtained by the speech recognition means, and the output is the text data translated into the target language. Here, since the user has set the target language in advance, the data is translated into the appropriate language based on that setting.

[0762] Step 4:

[0763] Emotion recognition:

[0764] The server sends the translated text data and the original audio data to the Microsoft Azure Emotion API to analyze the user's emotions. The input is the translated text data and audio data, and the output is the emotion analysis results and the emotion labels and scores based on them.

[0765] Step 5:

[0766] Text-to-Speech:

[0767] The server uses the IBM Watson Text to Speech API to convert the translated text data into speech data that reflects the results of sentiment analysis. The input is the translated text data and emotion label, and the output is synthesized speech data. This speech data is generated based on the voice characteristics (speed, tone, etc.) set by the user.

[0768] Step 6:

[0769] Displaying subtitles:

[0770] The computing device transmits the translated text data to the display device in real time and displays it as subtitles. The input is the translated text data, and the output is the subtitles displayed on the screen in real time, allowing viewers to visually confirm the translation content.

[0771] Step 7:

[0772] Audio Output:

[0773] The translated and emotion-analyzed speech data is output to the user from the speaker of the eyeglasses via a computing device. The input is synthesized speech data, and the output is speech that reaches the user's ears. This allows the user to hear the translated speech in real time.

[0774] Step 8:

[0775] Settings management:

[0776] The user manages various system settings (target language, audio characteristics, subtitle display, etc.) through a dedicated application. The input is the user's setting information, and the output is the setting data sent to the computing device. This allows the translation environment desired by the user to be properly applied.

[0777] The above is an explanation of the processing steps of the system program that realizes the application example, as well as its specific operations and input / output.

[0778] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0779] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0780] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0781] [Third embodiment]

[0782] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0783] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0784] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0785] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0786] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0787] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0788] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0789] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0790] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0791] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0792] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0793] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0794] The present invention provides a system for translating conversations between multiple languages ​​in real time and supporting natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, and a display means.

[0795] System configuration overview

[0796] 1. Eyeglass-type device worn by the user

[0797] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0798] 2. Arithmetic device

[0799] The apparatus includes a speech recognition unit for converting speech into text and a translation unit for translating the text into a target language. The computing unit receives the speech data obtained from the eyeglasses-type device and performs appropriate processing.

[0800] 3. Means of communication

[0801] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0802] 4. Dedicated application

[0803] Users use a dedicated application to manage the system's settings, such as the target language for translation and the voice speed of the audio output.

[0804] 5. Display means

[0805] A display means is provided to visually convey the translation results to the user, which makes it possible to display the translation content as subtitles.

[0806] System program processing

[0807] Acquiring voice input and converting it to text

[0808] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0809] Text translation

[0810] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0811] Voice output of translation results

[0812] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0813] Server and device communication

[0814] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0815] Specific examples

[0816] Scenario 1: English and Japanese Conversation

[0817] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text, "What time is the meeting?", which the translation means translates into "What time is the meeting?" The translated text is converted into speech and played back to User B.

[0818] Scenario 2: Configuration Management

[0819] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation.

[0820] In this way, the present invention enables natural conversation between multiple languages ​​and minimizes the time lost due to interpretation.

[0821] The processing flow will be explained below.

[0822] Step 1:

[0823] The user puts on the audio glasses and starts the system, which enables the microphone and speaker in the audio glasses.

[0824] Step 2:

[0825] The device (Audio Glasses) picks up what the other person is saying with a microphone. This audio data is captured in real time and temporarily stored in the device.

[0826] Step 3:

[0827] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0828] Step 4:

[0829] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0830] Step 5:

[0831] The server receives the translated text data and converts it into audio data using a text-to-speech engine, which is in the target language rather than the original language.

[0832] Step 6:

[0833] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0834] Step 7:

[0835] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0836] Step 8:

[0837] Users can use a dedicated application to change settings such as the target language, speech output speed, and volume as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0838] Step 9:

[0839] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, and voice output, maintaining natural conversation between multiple languages.

[0840] Example 1

[0841] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0842] In real-time communication between multiple languages, language barriers exist, making smooth conversation difficult. Existing translation systems have issues with translation accuracy and speed, making them insufficient for real-time conversation support. In particular, it is difficult to continue the flow of conversation without interruption, and this aspect needs improvement.

[0843] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0844] In this invention, the server includes means for acquiring speech from the other party via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, means for converting the translated text into voice and outputting it, a wearable device equipped with the voice input means and the voice output means, an information processing device including the voice recognition means and the translation means, communication means for connecting the wearable device and the information processing device, and communication means for transmitting and receiving data in real time, thereby enabling users to have natural real-time conversations between multiple languages.

[0845] "Voice input means" is a device that has the function of acquiring the speech of the other party.

[0846] The "voice recognition means" is a device that has the function of converting acquired voice into text data.

[0847] The "translation means" is a device that has the function of translating text data into a target language set by the user.

[0848] "Speech output means" refers to a device that has the function of converting translated text into speech and outputting it.

[0849] A "wearable device" is a wearable device equipped with an audio input means and an audio output means.

[0850] The "information processing device" is a processing device that includes a speech recognition means and a translation means.

[0851] A "communication means" is a device that connects a wearable device to an information processing device and has the function of sending and receiving data.

[0852] "Communication means for transmitting and receiving data in real time" refers to a communication technology that enables the transmission and reception of data in real time.

[0853] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system includes a wearable device (e.g., a glasses-type device), an information processing device, communication means, a dedicated application for managing user settings, and display means. Each component and its operation are described in detail below.

[0854] Wearable devices

[0855] The wearable device is equipped with a voice input means and a voice output means. It is worn by the user and looks like a glasses-type device. This device has the following functions:

[0856] Voice input means: A microphone is built in to capture what the other person is saying.

[0857] Audio output means: A built-in speaker is included, allowing the user to hear the translated speech.

[0858] Information processing device

[0859] The information processing device includes a speech recognition unit and a translation unit. These units analyze the speech data acquired from the wearable device and perform translation processing. Specific components are as follows:

[0860] Speech recognition means: Converts acquired voice data into text data.

[0861] Translation method: Translates text data into the target language set by the user. The translation engine uses a generative AI model based on a neural network, for example.

[0862] communication means

[0863] The communication means connects the wearable device to the information processing device and transmits and receives data in real time. The communication means includes:

[0864] Bluetooth and Wi-Fi: Used to send and receive voice and translation data.

[0865] Dedicated application

[0866] A dedicated application allows users to manage the system settings. Through this application, the following settings can be configured:

[0867] Select target language: Set the language to translate into.

[0868] Audio output adjustment: Set the audio output speed, etc.

[0869] Display means

[0870] The display means is used to visually convey the translation results to the user. Specifically, it has the following functions:

[0871] Subtitle display: It is possible to display the translated text as subtitles.

[0872] Specific examples

[0873] Example 1: English-Japanese conversation

[0874] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the wearable device picks up the speech and sends it to an information processing device. The speech recognition means converts this speech into text data, "What time is the meeting?", which is then translated by the translation means into "What time is the meeting?" The translated text is converted into speech and played back to User B. The speaker in the wearable device outputs the speech, "What time is the meeting?"

[0875] Example 2: Configuration Management

[0876] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. For example, the target language is changed from Japanese to French, and the speech output speed is set to 1.25 times faster. This information is immediately sent to the information processing device, and the new settings are reflected from the next conversation. Every time the user starts a conversation, translation is performed based on the application settings.

[0877] Prompt Sentence Examples

[0878] Here are some examples of prompts to input to a generative AI model:

[0879] If user A speaks "Hello, how are you?" into a wearable device, how can I convert that speech into text and translate it into Japanese, the target language specified by user B?

[0880] In this way, the present invention enables natural conversations between multiple languages ​​and minimizes the time lost due to interpretation. Users can manage settings through a dedicated application, enabling smooth communication in real time.

[0881] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0882] Step 1: Getting voice input

[0883] When a user starts a conversation, the microphone on the device (wearable device) picks up what the other person is saying. When User A says, "Hello, how are you?", the voice is input to the device through the microphone. The device converts this real-time voice input into digital voice data and prepares for the next step. Here, the input is analog voice, and the output is digital voice data.

[0884] Step 2: Sending audio data

[0885] The terminal (wearable device) transmits digital voice data to the server via Bluetooth or Wi-Fi. For example, voice data such as "Hello, how are you?" is delivered to the server via wireless communication. The input here is the digital voice data, and the output is the voice data transmitted to the server. During this transmission, the communication indicator LED on the terminal lights up to indicate the status of data transmission.

[0886] Step 3: Speech to Text

[0887] The server inputs the received voice data into a voice recognition means and converts it into text data. For example, the voice "Hello, how are you?" is converted into text "Hello, how are you?" The server uses a voice recognition algorithm to analyze the voice data and generate output in text format. Here, the input is the voice data sent to the server, and the output is text data. A processing progress bar on the server operates to display the progress of the conversion.

[0888] Step 4: Translate the text

[0889] The server passes the text data obtained by the speech recognition means to the translation means, which translates it into the target language set by the user. For example, if User B has set Japanese as the target language, "Hello, how are you?" will be translated into "Hello, how are you?" A generative AI model is used to perform highly accurate translation. The input here is the recognized text data, and the output is the translated text data. The original text and the translated text are recorded in the server log.

[0890] Step 5: Audio output of translation results

[0891] The server passes the translated text data to the voice output means, which converts it into voice data. The converted voice data is sent from the server to the terminal (wearable device). The converted voice is played from the terminal's speaker. For example, User B can hear the voice saying, "Hello, how are you?" Here, the input is the translated text data, and the output is playable voice data. An LED flashes to indicate that the terminal's speaker is outputting voice.

[0892] Step 6: Server and device communication

[0893] The server and the terminal (wearable device) are always in communication, sending and receiving data in real time. This allows each step of the conversation to be processed without delay. Every time the user speaks, the communication status can be confirmed by the communication indicator on the terminal lighting up. The input here is the user's voice and setting information, and the output is smooth conversation through real-time data transmission and reception.

[0894] Step 7: Configuration Management

[0895] The user opens a dedicated application to manage settings such as the target language of the conversation and the speech output speed. When the user changes the settings, the setting information is immediately sent to the server and applied throughout the system. For example, the user changes the target language from Japanese to French and sets the speech output speed to 1.25x. The new settings will be applied from the next conversation. The input here is the setting change information made by the user, and the output is the updated system settings.

[0896] In this way, the system translates conversations between multiple languages ​​in real time, providing smooth and natural communication.

[0897] (Application example 1)

[0898] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0899] Many brick-and-mortar stores face the problem of inability to communicate smoothly with tourists who speak foreign languages. This communication barrier can make it difficult for store staff and tourists to exchange accurate information, resulting in a decline in service quality. Another issue is the time loss caused by the inability to obtain translation results immediately. To solve these problems, a system that supports multilingual conversations in real time is needed.

[0900] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0901] In this invention, the server includes means for acquiring the other person's speech via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, voice output means for converting the translated text into voice and outputting it, a glasses-type device equipped with the voice input means and the voice output means, a computing device including the voice recognition means and the translation means, communication means for connecting the glasses-type device and a smartphone to the computing device, means for the user to manage settings using a dedicated application, and display means for displaying the translation results as subtitles or means for displaying them on the smartphone. This enables smooth communication between store staff and tourists across language barriers.

[0902] "Audio input means" refers to a device or technology for capturing audio.

[0903] A "speech recognition means" is a technique or device that converts captured speech into text.

[0904] A "translation means" is a technique or device that translates text into a target language set by the user.

[0905] "Audio output means" refers to a device or technology that converts translated text into audio and plays it back.

[0906] An "eyeglasses-type device" is an electronic device in the shape of glasses that has built-in audio input and audio output functions.

[0907] A "computing device" is an electronic device that includes a processor that performs speech recognition and translation.

[0908] "Communication means" refers to a technology or device for transmitting and receiving data between different devices, and includes wireless communication technologies such as Bluetooth and Wi-Fi.

[0909] A "dedicated application" is software that allows users to configure the system and runs on a smartphone or tablet.

[0910] "Display means" means a device or technology for visually displaying the translation result in text, including smart glasses or a smartphone display.

[0911] The present invention provides a real-time translation system that supports communication between users who speak different languages ​​in a physical store. The system specifically includes the following configuration and operation.

[0912] System configuration

[0913] 1. Voice input method

[0914] The server uses the microphone of the smart glasses or smartphone worn by the user to capture the speech of the interlocutor, thereby collecting the speech content as voice data.

[0915] 2. Voice Recognition Method

[0916] The captured voice data is sent to a server and converted into text data using voice recognition software, using the "speech_recognition" library.

[0917] 3. Translation Methods

[0918] The text data acquired by the speech recognition means is translated into the target language by a translation engine on the server. This translation process uses the "googletrans" library.

[0919] 4. Audio output means

[0920] The translated text is then converted back into audio data in the target language using speech synthesis technology, and the server uses the gTTS (Google Text-to-Speech) library to generate an audio file that is played on the user's smart glasses or smartphone speakers.

[0921] 5. Display means

[0922] The translation result is displayed as text on the smart glasses display or on the smartphone screen, allowing users to check the translation not only aloud but also visually.

[0923] Hardware and Software Use

[0924] The hardware used includes smart glasses and smartphones, which connect to the server via Bluetooth or Wi-Fi.

[0925] The software uses:

[0926] Speech Recognition: speech_recognition library

[0927] Translation: GoogleTrans Library

[0928] Speech synthesis: gTTS library

[0929] Audio playback: pygame library

[0930] Bluetooth communication: pybluez library

[0931] Specific examples

[0932] When a store staff member says in Japanese, "How do you use this product?", the microphone in the smart glasses or smartphone picks up the speech and converts it into text using speech recognition. This text is then translated into English by a translation engine into "How do you use this product?" The translation result is converted into speech and played back through the speaker of the staff member's smart glasses or smartphone, while also being displayed as text on the screen.

[0933] Prompt Sentence Examples

[0934] "Design an application for smart glasses that uses voice input to translate conversations between multiple languages ​​in real time, and outputs them in voice and text format."

[0935] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0936] Step 1:

[0937] The user makes a speech. The speech input means acquires the speech through the microphone of the smart glasses or smartphone. At this time, the acquired speech data is a raw speech signal. The server receives the corresponding speech data.

[0938] Step 2:

[0939] The server converts the received voice data into text data using speech recognition software (speech_recognition library). The input is the acquired voice data, and the output is the data converted from the speech into text format. Through this conversion, the voice information is expressed as text information.

[0940] Step 3:

[0941] The server inputs the converted text data into a translation engine (GoogleTrans library) and translates it into the target language. The input is text data, and the output is translated text data. This process converts the user's speech into a language the other person can understand.

[0942] Step 4:

[0943] The server converts the translated text data into audio data using speech synthesis technology (gTTS library). The input is the translated text data, and the output is an audio file. The server generates this audio file and prepares it for playback.

[0944] Step 5:

[0945] The server sends the generated audio file to the speaker of the smart glasses or smartphone and plays it through the audio output means, allowing the other person to hear the translated audio. At the same time, the translation result text is displayed on the display of the smart glasses or smartphone. The input is the audio file and text data, and the output is the playback of the translated audio and the text display.

[0946] Step 6:

[0947] Users manage system settings using a dedicated application. When a user changes the target language, for example, the setting information is sent to the server and is immediately reflected throughout the system. The input is the user's setting information, and the output is a reflection of the updated setting information.

[0948] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0949] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, a display means, and an emotion engine that recognizes the user's emotions.

[0950] System configuration overview

[0951] 1. Eyeglass-type device worn by the user

[0952] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[0953] 2. Arithmetic device

[0954] The device includes a speech recognition unit that converts speech into text, a translation unit that translates the text into a target language, and an emotion engine that recognizes the user's emotions. The computing unit receives the speech data acquired from the glasses-type device and performs appropriate processing.

[0955] 3. Means of communication

[0956] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[0957] 4. Dedicated application

[0958] Users can use a dedicated application to manage the system's settings, such as the target language for translation, voice output voice speed, volume, etc. There is also a training mode to improve the accuracy of emotion recognition.

[0959] 5. Display means

[0960] The system is equipped with a display means for visually conveying the translation results to the user, which makes it possible to display the translation content as subtitles.

[0961] 6. Emotion Engine

[0962] The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions, and adjusts the tone and speed of the translated speech based on the recognized emotion, ensuring natural conversation.

[0963] System program processing

[0964] Acquiring voice input and converting it to text

[0965] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[0966] Text translation

[0967] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[0968] Emotion recognition and voice output adjustment

[0969] The translated text data is converted into voice data by an emotion engine, taking into account the user's emotions. For example, if the user is excited, the translated voice will be played in a similarly excited tone. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and choice of words used.

[0970] Voice output of translation results

[0971] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[0972] Server and device communication

[0973] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[0974] Specific examples

[0975] Scenario 1: English and Japanese Conversation

[0976] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text "What time is the meeting?", and the translation means translates it into "What time is the meeting?" The emotion engine recognizes that User A's question contains tension and plays back the translated speech in a voice that reflects that tension.

[0977] Scenario 2: Configuration Management

[0978] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation. For example, if User B now selects French, "What time is the meeting?" is translated as "À quelle heure est la réunion?" and the emotion engine plays it in a calm, gentle tone.

[0979] In this way, the present invention provides an interpretation system that enables natural conversation between multiple languages ​​and is adaptable to various situations.

[0980] The processing flow will be explained below.

[0981] Step 1:

[0982] The user puts on the audio glasses and starts the system, which activates the microphone and speaker in the audio glasses and starts the emotion engine, preparing to analyze the user's facial expressions and tone of voice.

[0983] Step 2:

[0984] The device (Audio Glasses) picks up what the other person is saying with a microphone. This voice data is captured in real time and temporarily stored in the device. At the same time, the emotion engine records the user's facial expressions and voice tone and generates emotion data.

[0985] Step 3:

[0986] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[0987] Step 4:

[0988] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[0989] Step 5:

[0990] The server receives the translated text data and the emotion data generated by the emotion engine, which then adjusts the tone and speed of the translated speech based on that data.

[0991] Step 6:

[0992] The server uses a text-to-speech engine to convert the translated text data into speech data in the target language, rather than the original language, with tone and speed based on the emotion data.

[0993] Step 7:

[0994] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[0995] Step 8:

[0996] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[0997] Step 9:

[0998] Users can use a dedicated application to change the target language, speech output speed, volume, emotion recognition settings, etc. as needed. This setting information is immediately sent to the server and reflected throughout the system.

[0999] Step 10:

[1000] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, sentiment analysis, and voice output, maintaining natural conversations between multiple languages.

[1001] Example 2

[1002] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1003] Technology for smooth, natural, real-time conversations between multiple languages ​​is important, especially when communicating between different languages. However, current interpretation systems produce mechanical translation results that make it difficult to reflect the user's feelings. Furthermore, there is a lack of ways for users to flexibly change system settings or visually check translated subtitles. This creates a problem of impeding the natural flow of conversation.

[1004] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice input means for acquiring the speech of the other party, a voice recognition means for converting the acquired voice into text, and a translation means for translating the converted text into a target language set by the user. This enables real-time conversation that reflects natural emotions.

[1005] The "voice input means for acquiring the other party's speech" is a device for capturing the other party's speech and acquiring the speech signal as digital data.

[1006] "Speech recognition means for converting acquired speech into text" refers to a technique or device for analyzing acquired speech data and converting it into text data.

[1007] The "translation means for translating the converted text into a target language set by the user" is a technique or device for automatically translating text data into another language set by the user.

[1008] The "audio output means for converting the translated text into audio and outputting it" is a device for converting the translated text data into an audio signal and outputting that audio to the user.

[1009] A "portable device" is an electronic device that can be easily worn by a user and is portable.

[1010] A "computing device" is a computer system that processes data and performs calculations.

[1011] The "communication means for connecting the portable device and the computing device" refers to a technique or device for transmitting and receiving data between the portable device and the computing device.

[1012] "Emotion recognition means for analyzing a user's facial expression and tone of voice to recognize emotions" refers to a technology or device for analyzing and recognizing emotions contained in a user's facial expression and tone of voice.

[1013] "Means for adjusting the tone and speed of the translated speech based on the recognized emotion" refers to a technology or device for appropriately adjusting the tone and speed of the speech based on the emotion recognized by the emotion recognition means.

[1014] This invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system translates what the other person is saying in real time and provides voice output that reflects their emotions, thereby realizing natural conversations between users.

[1015] The system includes the following components:

[1016] 1. Handheld devices:

[1017] It is a glasses-type device worn by the user that has built-in audio input and output means. This device can be audio glasses such as Bose Frames. The device captures the other person's speech, converts it into digital data, and sends it to a computing device.

[1018] 2. Computing equipment:

[1019] The device includes a speech recognition unit, a translation unit, and an emotion recognition unit. The computing unit receives and processes voice data from a mobile device via a communication means such as Bluetooth. The device uses the Google Cloud Speech-to-Text API for speech recognition, the DeepL API for translation, and the Microsoft Azure Emotion API for emotion recognition.

[1020] 3. Means of communication:

[1021] The portable device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data in real time.

[1022] 4. Dedicated applications:

[1023] Users can manage the system settings using a dedicated application installed on their smartphone or tablet, which allows them to select the target language and adjust the speech output speed and volume.

[1024] 5. Emotion recognition means:

[1025] The emotion recognition unit analyzes the user's facial expressions and tone of voice to recognize emotions, and has the function of adjusting the tone and speed of the translated speech based on the recognized emotions.

[1026] (Example)

[1027] Scenario 1: English and Japanese conversation:

[1028] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the server captures the voice through the microphone of the eyeglasses-type device and converts it into text data "What time is the meeting?" using the Google Cloud Speech-to-Text API. The DeepL API is used to translate this to "What time is the meeting?", and if the emotion recognition means recognizes that User A's question contains tension, it generates voice that reflects that tension. Finally, the generated voice data is transmitted to User B through the speaker of the eyeglasses-type device.

[1029] Scenario 2: Configuration Management:

[1030] User B opens the application, changes the target language to French, and adjusts the speech output speed to a calmer tone. This setting information is sent to the server in real time and applied to the entire system. In the next conversation, "What time is the meeting?" is translated to "À quelle heure est la réunion?" and the speech is output in a calm, gentle tone.

[1031] In this way, the present invention provides a system that enables natural conversation between multiple languages ​​and realizes smooth communication in a variety of situations.

[1032] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1033] Step 1:

[1034] Acquiring voice input

[1035] When a user initiates a conversation, a microphone on the mobile device picks up what the other person is saying, using audio glasses such as Bose Frames.

[1036] Input: Other person's speech (audio data)

[1037] Data processing: Acquisition of audio data

[1038] Output: Digital audio data

[1039] Specific operation: When user A says "What time is the meeting?", the voice is recorded as digital data via the microphone of the portable device.

[1040] Step 2:

[1041] Sending audio data

[1042] The digital audio data captured by the portable device is transmitted to the computing unit via Bluetooth.

[1043] Input: Digital audio data

[1044] Data processing: Data transmission

[1045] Output: Received digital audio data

[1046] Specific operation: A portable device transmits digital audio data to a computing device via Bluetooth.

[1047] Step 3:

[1048] Converting audio data to text

[1049] The computing device uses the Google Cloud Speech-to-Text API to analyze the voice data and convert it into text data.

[1050] Input: Digital audio data

[1051] Data processing: speech recognition and text conversion

[1052] Output: Text data

[1053] Specific operation: The computing device converts the voice data into text: "What time is the meeting?"

[1054] Step 4:

[1055] Text data translation

[1056] The computing device uses the DeepL API to translate the text data into the target language.

[1057] Input: Text data

[1058] Data processing: Text translation

[1059] Output: Text data in the target language

[1060] Specific behavior: The computing device translates "What time is the meeting?" to "What time is the meeting?"

[1061] Step 5:

[1062] Emotion recognition

[1063] The computing device's emotion recognition means (such as Microsoft Azure's Emotion API) analyzes the user's facial expressions and vocal tone to recognize their emotions.

[1064] Input: User's facial expression data and voice tone

[1065] Data processing: Emotion analysis and recognition

[1066] Output: Recognized emotion data

[1067] Specific operation: The emotion recognition means reads the user's level of tension from their facial expressions and voice.

[1068] Step 6:

[1069] Adjust the tone and speed of your voice

[1070] The computing device adjusts the tone and rate of the translated speech based on the recognized emotion.

[1071] Input: Translation text data and emotion data

[1072] Data processing: adjusting the tone and speed of the voice

[1073] Output: Modified audio data

[1074] What it does: The computing device converts the translated text into a voice tone that reflects tension.

[1075] Step 7:

[1076] Generate audio output

[1077] The computing device generates tailored voice data using a TTS engine such as Amazon Polly.

[1078] Input: Modified audio data

[1079] Data processing: voice synthesis

[1080] Output: Generated audio data

[1081] Specific operation: The computing device synthesizes the adjusted voice data and prepares it for output.

[1082] Step 8:

[1083] Audio data output

[1084] The server sends the generated audio data to the portable device via Bluetooth, where it is transmitted to the user through the speaker.

[1085] Input: Generated audio data

[1086] Data processing: Sending voice data

[1087] Output: Audio output

[1088] Specific Actions: User B hears the translated audio of "Hello, how are you?" through the speaker of their mobile device.

[1089] Through this series of processes, the system realizes natural dialogue between users.

[1090] (Application example 2)

[1091] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1092] Currently, achieving natural translation and emotionally-rich speech output in real time is a challenging task in multilingual conversations and content distribution. Furthermore, in live streaming and recorded content where multiple language-speaking viewers simultaneously participate, viewers often cannot accurately convey the emotions of the performer or the video while watching in their own language. Therefore, multilingual streaming services are required to ensure that viewers can enjoy a natural and emotionally-rich experience.

[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for acquiring the other party's speech via a voice input means; voice recognition means for converting the acquired voice into text; translation means for translating the converted text into a target language set by the user; voice output means for converting the translated text into voice and outputting it; a glasses-type device equipped with the voice input means and the voice output means; a computing device including the voice recognition means and the translation means; communication means for connecting the glasses-type device and the computing device; display means for translating the conversations of performers and videos in real time and displaying subtitles in a language selected by the viewer; means for adjusting the translation results using emotion analysis and displaying them as natural conversational text; and a dedicated application for the viewer to manage settings. This enables viewers to enjoy a natural and emotionally rich experience in a multilingual distribution service.

[1094] The "means for acquiring the other party's speech by voice input means" is a technique for acquiring the other party's speech by using a voice input device such as a microphone.

[1095] "Speech recognition means" refers to software or algorithms that convert captured speech data into text.

[1096] The "translation means" is a technology that translates the text converted by the speech recognition means into a target language set by the user.

[1097] "Audio output means" refers to technology that converts the translated text back into audio and outputs it via a speaker or the like.

[1098] The "eyeglasses-type device" is an eyeglasses-type device equipped with an audio input means and an audio output means, and is a portable communication device.

[1099] The term "arithmetic unit" is a general term for hardware and software for performing various processes such as speech recognition means and translation means.

[1100] The "communication means" is a technology for connecting the eyeglass-type device and the computing device and transmitting and receiving data.

[1101] "Display means" refers to devices or technologies that translate the dialogue of the performers or in the video in real time and display subtitles in the language selected by the viewer.

[1102] "Emotion analysis means" is a technology that adjusts the translation results based on the user's emotions and outputs them in a natural conversational style.

[1103] "Dedicated application" is software that allows viewers to manage and customize system settings.

[1104] This invention provides a system for realizing real-time translation and emotion recognition in natural conversations between multiple languages ​​and content distribution. Specific embodiments of this system are described below.

[1105] System Configuration

[1106] The system includes the following major components:

[1107] 1. Voice input means: A device for capturing what the other person is saying. For example, a microphone.

[1108] 2. Speech recognition tool: Software or algorithms that convert captured voice data into text. For example, we use the Google Cloud Speech-to-Text API.

[1109] 3. Translation method: The technology that translates the text into the target language. For example, using the Amazon Translate API.

[1110] 4. Voice output method: Technology that converts the translated text into voice and outputs it through a speaker. For example, IBM Watson Text to Speech API is used.

[1111] 5. Glasses-type device: A portable device equipped with audio input and output means.

[1112] 6. Computing unit: An integrated system of hardware and software for performing various processes.

[1113] 7. Communication method: Technology that connects the glasses device to the computing device, such as Bluetooth or Wi-Fi.

[1114] 8. Display means: A device for displaying the translation results to the viewer as subtitles in real time.

[1115] 9. Sentiment analysis: Technology to analyze the user's emotions and output the translation results in a natural conversational style. For example, we use the Microsoft Azure Emotion API.

[1116] 10. Dedicated application: Software that allows viewers to manage their system settings. For example, using React Native as the platform.

[1117] System Operation

[1118] 1. Acquiring voice input and converting it to text

[1119] When a user starts talking, the microphone of the eyeglass device picks up what is being said. The picked-up voice data is sent to a computing device and converted into text data by a voice recognition means.

[1120] 2. Text Translation

[1121] The text converted by the speech recognition means is translated into a target language set by the user using the translation means.

[1122] 3. Emotion recognition and voice output adjustment

[1123] The translated text is converted into voice data by taking into account the user's emotions through emotion analysis, and this voice data is output in an appropriate tone and played to the audience.

[1124] 4. Displaying the translation results

[1125] The translated text is provided to the viewer in real time as subtitles using a display means.

[1126] Usage example

[1127] For example, if a speaker says "Hello, everyone! Welcome to our event!" during a live stream, the system will convert this into text in real time and translate it into the target language (e.g., Japanese). The translated "Hello, everyone! Welcome to our event!" will be played in an "excited" tone based on emotion analysis and simultaneously displayed as subtitles.

[1128] Prompt example

[1129] Below are some examples of prompts used in this system:

[1130] Design a real-time translation system for live streaming that provides subtitles and speech translation according to the viewer's language of choice and realizes natural conversational tone that takes into account the speaker's emotions. This system includes speech recognition, text translation, emotion recognition, and speech synthesis modules. Each module should be implemented using a specific API, and viewers should be able to change the settings.

[1131] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1132] Step 1:

[1133] Acquiring voice input:

[1134] When a user starts speaking, the microphone on the eyeglasses picks up what is being said. The input voice data is sent as a digital signal to a computing device, where it is converted into a digital audio format such as linear PCM.

[1135] Step 2:

[1136] Speech to text:

[1137] The server sends the voice data received by the computing device to the Google Cloud Speech-to-Text API. The API analyzes the voice data and converts it into corresponding text data. The input is digital voice data, and the output is text data that converts the voice into text.

[1138] Step 3:

[1139] Text Translation:

[1140] The server sends the converted data to the Amazon Translate API. The input is the text data obtained by the speech recognition means, and the output is the text data translated into the target language. Here, since the user has set the target language in advance, the data is translated into the appropriate language based on that setting.

[1141] Step 4:

[1142] Emotion recognition:

[1143] The server sends the translated text data and the original audio data to the Microsoft Azure Emotion API to analyze the user's emotions. The input is the translated text data and audio data, and the output is the emotion analysis results and the emotion labels and scores based on them.

[1144] Step 5:

[1145] Text-to-Speech:

[1146] The server uses the IBM Watson Text to Speech API to convert the translated text data into speech data that reflects the results of sentiment analysis. The input is the translated text data and emotion label, and the output is synthesized speech data. This speech data is generated based on the voice characteristics (speed, tone, etc.) set by the user.

[1147] Step 6:

[1148] Displaying subtitles:

[1149] The computing device transmits the translated text data to the display device in real time and displays it as subtitles. The input is the translated text data, and the output is the subtitles displayed on the screen in real time, allowing viewers to visually confirm the translation content.

[1150] Step 7:

[1151] Audio Output:

[1152] The translated and emotion-analyzed speech data is output to the user from the speaker of the eyeglasses via a computing device. The input is synthesized speech data, and the output is speech that reaches the user's ears. This allows the user to hear the translated speech in real time.

[1153] Step 8:

[1154] Settings management:

[1155] The user manages various system settings (target language, audio characteristics, subtitle display, etc.) through a dedicated application. The input is the user's setting information, and the output is the setting data sent to the computing device. This allows the translation environment desired by the user to be properly applied.

[1156] The above is an explanation of the processing steps of the system program that realizes the application example, as well as its specific operations and input / output.

[1157] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1158] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1159] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1160] [Fourth embodiment]

[1161] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1162] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1163] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1164] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1165] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1166] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1167] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1168] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1169] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1170] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1171] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1172] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1173] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1174] The present invention provides a system for translating conversations between multiple languages ​​in real time and supporting natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, and a display means.

[1175] System configuration overview

[1176] 1. Eyeglass-type device worn by the user

[1177] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[1178] 2. Arithmetic device

[1179] The apparatus includes a speech recognition unit for converting speech into text and a translation unit for translating the text into a target language. The computing unit receives the speech data obtained from the eyeglasses-type device and performs appropriate processing.

[1180] 3. Means of communication

[1181] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[1182] 4. Dedicated application

[1183] Users use a dedicated application to manage the system's settings, such as the target language for translation and the voice speed of the audio output.

[1184] 5. Display means

[1185] A display means is provided to visually convey the translation results to the user, which makes it possible to display the translation content as subtitles.

[1186] System program processing

[1187] Acquiring voice input and converting it to text

[1188] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[1189] Text translation

[1190] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[1191] Voice output of translation results

[1192] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[1193] Server and device communication

[1194] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[1195] Specific examples

[1196] Scenario 1: English and Japanese Conversation

[1197] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text, "What time is the meeting?", which the translation means translates into "What time is the meeting?" The translated text is converted into speech and played back to User B.

[1198] Scenario 2: Configuration Management

[1199] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation.

[1200] In this way, the present invention enables natural conversation between multiple languages ​​and minimizes the time lost due to interpretation.

[1201] The processing flow will be explained below.

[1202] Step 1:

[1203] The user puts on the audio glasses and starts the system, which enables the microphone and speaker in the audio glasses.

[1204] Step 2:

[1205] The device (Audio Glasses) picks up what the other person is saying with a microphone. This audio data is captured in real time and temporarily stored in the device.

[1206] Step 3:

[1207] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[1208] Step 4:

[1209] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[1210] Step 5:

[1211] The server receives the translated text data and converts it into audio data using a text-to-speech engine, which is in the target language rather than the original language.

[1212] Step 6:

[1213] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[1214] Step 7:

[1215] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[1216] Step 8:

[1217] Users can use a dedicated application to change settings such as the target language, speech output speed, and volume as needed. This setting information is immediately sent to the server and reflected throughout the system.

[1218] Step 9:

[1219] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, and voice output, maintaining natural conversation between multiple languages.

[1220] Example 1

[1221] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1222] In real-time communication between multiple languages, language barriers exist, making smooth conversation difficult. Existing translation systems have issues with translation accuracy and speed, making them insufficient for real-time conversation support. In particular, it is difficult to continue the flow of conversation without interruption, and this aspect needs improvement.

[1223] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1224] In this invention, the server includes means for acquiring speech from the other party via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, means for converting the translated text into voice and outputting it, a wearable device equipped with the voice input means and the voice output means, an information processing device including the voice recognition means and the translation means, communication means for connecting the wearable device and the information processing device, and communication means for transmitting and receiving data in real time, thereby enabling users to have natural real-time conversations between multiple languages.

[1225] "Voice input means" is a device that has the function of acquiring the speech of the other party.

[1226] The "voice recognition means" is a device that has the function of converting acquired voice into text data.

[1227] The "translation means" is a device that has the function of translating text data into a target language set by the user.

[1228] "Speech output means" refers to a device that has the function of converting translated text into speech and outputting it.

[1229] A "wearable device" is a wearable device equipped with an audio input means and an audio output means.

[1230] The "information processing device" is a processing device that includes a speech recognition means and a translation means.

[1231] A "communication means" is a device that connects a wearable device to an information processing device and has the function of sending and receiving data.

[1232] "Communication means for transmitting and receiving data in real time" refers to a communication technology that enables the transmission and reception of data in real time.

[1233] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system includes a wearable device (e.g., a glasses-type device), an information processing device, communication means, a dedicated application for managing user settings, and display means. Each component and its operation are described in detail below.

[1234] Wearable devices

[1235] The wearable device is equipped with a voice input means and a voice output means. It is worn by the user and looks like a glasses-type device. This device has the following functions:

[1236] Voice input means: A microphone is built in to capture what the other person is saying.

[1237] Audio output means: A built-in speaker is included, allowing the user to hear the translated speech.

[1238] Information processing device

[1239] The information processing device includes a speech recognition unit and a translation unit. These units analyze the speech data acquired from the wearable device and perform translation processing. Specific components are as follows:

[1240] Speech recognition means: Converts acquired voice data into text data.

[1241] Translation method: Translates text data into the target language set by the user. The translation engine uses a generative AI model based on a neural network, for example.

[1242] communication means

[1243] The communication means connects the wearable device to the information processing device and transmits and receives data in real time. The communication means includes:

[1244] Bluetooth and Wi-Fi: Used to send and receive voice and translation data.

[1245] Dedicated application

[1246] A dedicated application allows users to manage the system settings. Through this application, the following settings can be configured:

[1247] Select target language: Set the language to translate into.

[1248] Audio output adjustment: Set the audio output speed, etc.

[1249] Display means

[1250] The display means is used to visually convey the translation results to the user. Specifically, it has the following functions:

[1251] Subtitle display: It is possible to display the translated text as subtitles.

[1252] Specific examples

[1253] Example 1: English-Japanese conversation

[1254] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the wearable device picks up the speech and sends it to an information processing device. The speech recognition means converts this speech into text data, "What time is the meeting?", which is then translated by the translation means into "What time is the meeting?" The translated text is converted into speech and played back to User B. The speaker in the wearable device outputs the speech, "What time is the meeting?"

[1255] Example 2: Configuration Management

[1256] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. For example, the target language is changed from Japanese to French, and the speech output speed is set to 1.25 times faster. This information is immediately sent to the information processing device, and the new settings are reflected from the next conversation. Every time the user starts a conversation, translation is performed based on the application settings.

[1257] Prompt Sentence Examples

[1258] Here are some examples of prompts to input to a generative AI model:

[1259] If user A speaks "Hello, how are you?" into a wearable device, how can I convert that speech into text and translate it into Japanese, the target language specified by user B?

[1260] In this way, the present invention enables natural conversations between multiple languages ​​and minimizes the time lost due to interpretation. Users can manage settings through a dedicated application, enabling smooth communication in real time.

[1261] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1262] Step 1: Getting voice input

[1263] When a user starts a conversation, the microphone on the device (wearable device) picks up what the other person is saying. When User A says, "Hello, how are you?", the voice is input to the device through the microphone. The device converts this real-time voice input into digital voice data and prepares for the next step. Here, the input is analog voice, and the output is digital voice data.

[1264] Step 2: Sending audio data

[1265] The terminal (wearable device) transmits digital voice data to the server via Bluetooth or Wi-Fi. For example, voice data such as "Hello, how are you?" is delivered to the server via wireless communication. The input here is the digital voice data, and the output is the voice data transmitted to the server. During this transmission, the communication indicator LED on the terminal lights up to indicate the status of data transmission.

[1266] Step 3: Speech to Text

[1267] The server inputs the received voice data into a voice recognition means and converts it into text data. For example, the voice "Hello, how are you?" is converted into text "Hello, how are you?" The server uses a voice recognition algorithm to analyze the voice data and generate output in text format. Here, the input is the voice data sent to the server, and the output is text data. A processing progress bar on the server operates to display the progress of the conversion.

[1268] Step 4: Translate the text

[1269] The server passes the text data obtained by the speech recognition means to the translation means, which translates it into the target language set by the user. For example, if User B has set Japanese as the target language, "Hello, how are you?" will be translated into "Hello, how are you?" A generative AI model is used to perform highly accurate translation. The input here is the recognized text data, and the output is the translated text data. The original text and the translated text are recorded in the server log.

[1270] Step 5: Audio output of translation results

[1271] The server passes the translated text data to the voice output means, which converts it into voice data. The converted voice data is sent from the server to the terminal (wearable device). The converted voice is played from the terminal's speaker. For example, User B can hear the voice saying, "Hello, how are you?" Here, the input is the translated text data, and the output is playable voice data. An LED flashes to indicate that the terminal's speaker is outputting voice.

[1272] Step 6: Server and device communication

[1273] The server and the terminal (wearable device) are always in communication, sending and receiving data in real time. This allows each step of the conversation to be processed without delay. Every time the user speaks, the communication status can be confirmed by the communication indicator on the terminal lighting up. The input here is the user's voice and setting information, and the output is smooth conversation through real-time data transmission and reception.

[1274] Step 7: Configuration Management

[1275] The user opens a dedicated application to manage settings such as the target language of the conversation and the speech output speed. When the user changes the settings, the setting information is immediately sent to the server and applied throughout the system. For example, the user changes the target language from Japanese to French and sets the speech output speed to 1.25x. The new settings will be applied from the next conversation. The input here is the setting change information made by the user, and the output is the updated system settings.

[1276] In this way, the system translates conversations between multiple languages ​​in real time, providing smooth and natural communication.

[1277] (Application example 1)

[1278] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1279] Many brick-and-mortar stores face the problem of inability to communicate smoothly with tourists who speak foreign languages. This communication barrier can make it difficult for store staff and tourists to exchange accurate information, resulting in a decline in service quality. Another issue is the time loss caused by the inability to obtain translation results immediately. To solve these problems, a system that supports multilingual conversations in real time is needed.

[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1281] In this invention, the server includes means for acquiring the other person's speech via a voice input means, voice recognition means for converting the acquired voice into text, translation means for translating the converted text into a target language set by the user, voice output means for converting the translated text into voice and outputting it, a glasses-type device equipped with the voice input means and the voice output means, a computing device including the voice recognition means and the translation means, communication means for connecting the glasses-type device and a smartphone to the computing device, means for the user to manage settings using a dedicated application, and display means for displaying the translation results as subtitles or means for displaying them on the smartphone. This enables smooth communication between store staff and tourists across language barriers.

[1282] "Audio input means" refers to a device or technology for capturing audio.

[1283] A "speech recognition means" is a technique or device that converts captured speech into text.

[1284] A "translation means" is a technique or device that translates text into a target language set by the user.

[1285] "Audio output means" refers to a device or technology that converts translated text into audio and plays it back.

[1286] An "eyeglasses-type device" is an electronic device in the shape of glasses that has built-in audio input and audio output functions.

[1287] A "computing device" is an electronic device that includes a processor that performs speech recognition and translation.

[1288] "Communication means" refers to a technology or device for transmitting and receiving data between different devices, and includes wireless communication technologies such as Bluetooth and Wi-Fi.

[1289] A "dedicated application" is software that allows users to configure the system and runs on a smartphone or tablet.

[1290] "Display means" means a device or technology for visually displaying the translation result in text, including smart glasses or a smartphone display.

[1291] The present invention provides a real-time translation system that supports communication between users who speak different languages ​​in a physical store. The system specifically includes the following configuration and operation.

[1292] System configuration

[1293] 1. Voice input method

[1294] The server uses the microphone of the smart glasses or smartphone worn by the user to capture the speech of the interlocutor, thereby collecting the speech content as voice data.

[1295] 2. Voice Recognition Method

[1296] The captured voice data is sent to a server and converted into text data using voice recognition software, using the "speech_recognition" library.

[1297] 3. Translation Methods

[1298] The text data acquired by the speech recognition means is translated into the target language by a translation engine on the server. This translation process uses the "googletrans" library.

[1299] 4. Audio output means

[1300] The translated text is then converted back into audio data in the target language using speech synthesis technology, and the server uses the gTTS (Google Text-to-Speech) library to generate an audio file that is played on the user's smart glasses or smartphone speakers.

[1301] 5. Display means

[1302] The translation result is displayed as text on the smart glasses display or on the smartphone screen, allowing users to check the translation not only aloud but also visually.

[1303] Hardware and Software Use

[1304] The hardware used includes smart glasses and smartphones, which connect to the server via Bluetooth or Wi-Fi.

[1305] The software uses:

[1306] Speech Recognition: speech_recognition library

[1307] Translation: GoogleTrans Library

[1308] Speech synthesis: gTTS library

[1309] Audio playback: pygame library

[1310] Bluetooth communication: pybluez library

[1311] Specific examples

[1312] When a store staff member says in Japanese, "How do you use this product?", the microphone in the smart glasses or smartphone picks up the speech and converts it into text using speech recognition. This text is then translated into English by a translation engine into "How do you use this product?" The translation result is converted into speech and played back through the speaker of the staff member's smart glasses or smartphone, while also being displayed as text on the screen.

[1313] Prompt Sentence Examples

[1314] "Design an application for smart glasses that uses voice input to translate conversations between multiple languages ​​in real time, and outputs them in voice and text format."

[1315] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1316] Step 1:

[1317] The user makes a speech. The speech input means acquires the speech through the microphone of the smart glasses or smartphone. At this time, the acquired speech data is a raw speech signal. The server receives the corresponding speech data.

[1318] Step 2:

[1319] The server converts the received voice data into text data using speech recognition software (speech_recognition library). The input is the acquired voice data, and the output is the data converted from the speech into text format. Through this conversion, the voice information is expressed as text information.

[1320] Step 3:

[1321] The server inputs the converted text data into a translation engine (GoogleTrans library) and translates it into the target language. The input is text data, and the output is translated text data. This process converts the user's speech into a language the other person can understand.

[1322] Step 4:

[1323] The server converts the translated text data into audio data using speech synthesis technology (gTTS library). The input is the translated text data, and the output is an audio file. The server generates this audio file and prepares it for playback.

[1324] Step 5:

[1325] The server sends the generated audio file to the speaker of the smart glasses or smartphone and plays it through the audio output means, allowing the other person to hear the translated audio. At the same time, the translation result text is displayed on the display of the smart glasses or smartphone. The input is the audio file and text data, and the output is the playback of the translated audio and the text display.

[1326] Step 6:

[1327] Users manage system settings using a dedicated application. When a user changes the target language, for example, the setting information is sent to the server and is immediately reflected throughout the system. The input is the user's setting information, and the output is a reflection of the updated setting information.

[1328] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1329] The present invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. The system includes a glasses-type device, a computing unit, a communication means, a dedicated application for managing user settings, a display means, and an emotion engine that recognizes the user's emotions.

[1330] System configuration overview

[1331] 1. Eyeglass-type device worn by the user

[1332] This device, also known as audio glasses, has a built-in microphone and speaker, which allows it to receive speech as a voice input and to play the translation results back to the user as a voice output.

[1333] 2. Arithmetic device

[1334] The device includes a speech recognition unit that converts speech into text, a translation unit that translates the text into a target language, and an emotion engine that recognizes the user's emotions. The computing unit receives the speech data acquired from the glasses-type device and performs appropriate processing.

[1335] 3. Means of communication

[1336] The glasses-type device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data.

[1337] 4. Dedicated application

[1338] Users can use a dedicated application to manage the system's settings, such as the target language for translation, voice output voice speed, volume, etc. There is also a training mode to improve the accuracy of emotion recognition.

[1339] 5. Display means

[1340] The system is equipped with a display means for visually conveying the translation results to the user, which makes it possible to display the translation content as subtitles.

[1341] 6. Emotion Engine

[1342] The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions, and adjusts the tone and speed of the translated speech based on the recognized emotion, ensuring natural conversation.

[1343] System program processing

[1344] Acquiring voice input and converting it to text

[1345] When a user starts a conversation, the microphone in the eyeglass device picks up what the other person is saying. The picked up voice is sent to the computing device and converted into text data by a voice recognition means. For example, if user A says "Hello, how are you?", this voice data is converted into the text "Hello, how are you?"

[1346] Text translation

[1347] The converted text data is translated into the target language set by the user through a translation means. If User B selects Japanese, "Hello, how are you?" is translated into "Hello, how are you?"

[1348] Emotion recognition and voice output adjustment

[1349] The translated text data is converted into voice data by an emotion engine, taking into account the user's emotions. For example, if the user is excited, the translated voice will be played in a similarly excited tone. The emotion engine recognizes emotions by analyzing the user's facial expressions, tone of voice, and choice of words used.

[1350] Voice output of translation results

[1351] The translated text is converted into speech by the speech output means and is played back to the user through the speaker of the eyeglasses-type device, allowing User B to hear the translated speech saying, "Hello, how are you?"

[1352] Server and device communication

[1353] The glasses-type device and the computing unit are constantly in communication, sending and receiving data in real time. This ensures that translation processing is carried out without delay. Furthermore, when a user changes settings through a dedicated application, the settings are immediately sent to the server and reflected throughout the system.

[1354] Specific examples

[1355] Scenario 1: English and Japanese Conversation

[1356] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the microphone in the eyeglass device picks up the speech and sends it to the computing device. The speech recognition means converts this speech into text "What time is the meeting?", and the translation means translates it into "What time is the meeting?" The emotion engine recognizes that User A's question contains tension and plays back the translated speech in a voice that reflects that tension.

[1357] Scenario 2: Configuration Management

[1358] User B opens the dedicated application, changes the target language, and adjusts the speech output speed. These settings are immediately sent to the server and applied across the entire system. The updated settings are reflected from the next conversation. For example, if User B now selects French, "What time is the meeting?" is translated as "À quelle heure est la réunion?" and the emotion engine plays it in a calm, gentle tone.

[1359] In this way, the present invention provides an interpretation system that enables natural conversation between multiple languages ​​and is adaptable to various situations.

[1360] The processing flow will be explained below.

[1361] Step 1:

[1362] The user puts on the audio glasses and starts the system, which activates the microphone and speaker in the audio glasses and starts the emotion engine, preparing to analyze the user's facial expressions and tone of voice.

[1363] Step 2:

[1364] The device (Audio Glasses) picks up what the other person is saying with a microphone. This voice data is captured in real time and temporarily stored in the device. At the same time, the emotion engine records the user's facial expressions and voice tone and generates emotion data.

[1365] Step 3:

[1366] The device converts the acquired voice data into text data through a voice recognition means. At this time, the device accesses a voice recognition API via the Internet, sends the voice data, and obtains the text data returned by the API.

[1367] Step 4:

[1368] The terminal sends the converted text data to the translation means, which translates the text data based on the target language preset by the user.

[1369] Step 5:

[1370] The server receives the translated text data and the emotion data generated by the emotion engine, which then adjusts the tone and speed of the translated speech based on that data.

[1371] Step 6:

[1372] The server uses a text-to-speech engine to convert the translated text data into speech data in the target language, rather than the original language, with tone and speed based on the emotion data.

[1373] Step 7:

[1374] The device automatically plays back the translated audio data returned from the server and lets the user hear it through the speakers in the audio glasses.

[1375] Step 8:

[1376] The user can then hear the translated audio through the audio glasses and understand what the other person is saying in their own language, and the process repeats as the conversation continues.

[1377] Step 9:

[1378] Users can use a dedicated application to change the target language, speech output speed, volume, emotion recognition settings, etc. as needed. This setting information is immediately sent to the server and reflected throughout the system.

[1379] Step 10:

[1380] Even after the user changes the settings, the device continues to perform the entire process of real-time voice input, text conversion, translation, sentiment analysis, and voice output, maintaining natural conversations between multiple languages.

[1381] Example 2

[1382] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1383] Technology for smooth, natural, real-time conversations between multiple languages ​​is important, especially when communicating between different languages. However, current interpretation systems produce mechanical translation results that make it difficult to reflect the user's feelings. Furthermore, there is a lack of ways for users to flexibly change system settings or visually check translated subtitles. This creates a problem of impeding the natural flow of conversation.

[1384] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a voice input means for acquiring the speech of the other party, a voice recognition means for converting the acquired voice into text, and a translation means for translating the converted text into a target language set by the user. This enables real-time conversation that reflects natural emotions.

[1385] The "voice input means for acquiring the other party's speech" is a device for capturing the other party's speech and acquiring the speech signal as digital data.

[1386] "Speech recognition means for converting acquired speech into text" refers to a technique or device for analyzing acquired speech data and converting it into text data.

[1387] The "translation means for translating the converted text into a target language set by the user" is a technique or device for automatically translating text data into another language set by the user.

[1388] The "audio output means for converting the translated text into audio and outputting it" is a device for converting the translated text data into an audio signal and outputting that audio to the user.

[1389] A "portable device" is an electronic device that can be easily worn by a user and is portable.

[1390] A "computing device" is a computer system that processes data and performs calculations.

[1391] The "communication means for connecting the portable device and the computing device" refers to a technique or device for transmitting and receiving data between the portable device and the computing device.

[1392] "Emotion recognition means for analyzing a user's facial expression and tone of voice to recognize emotions" refers to a technology or device for analyzing and recognizing emotions contained in a user's facial expression and tone of voice.

[1393] "Means for adjusting the tone and speed of the translated speech based on the recognized emotion" refers to a technology or device for appropriately adjusting the tone and speed of the speech based on the emotion recognized by the emotion recognition means.

[1394] This invention is a system that translates conversations between multiple languages ​​in real time and supports natural conversations. This system translates what the other person is saying in real time and provides voice output that reflects their emotions, thereby realizing natural conversations between users.

[1395] The system includes the following components:

[1396] 1. Handheld devices:

[1397] It is a glasses-type device worn by the user that has built-in audio input and output means. This device can be audio glasses such as Bose Frames. The device captures the other person's speech, converts it into digital data, and sends it to a computing device.

[1398] 2. Computing equipment:

[1399] The device includes a speech recognition unit, a translation unit, and an emotion recognition unit. The computing unit receives and processes voice data from a mobile device via a communication means such as Bluetooth. The device uses the Google Cloud Speech-to-Text API for speech recognition, the DeepL API for translation, and the Microsoft Azure Emotion API for emotion recognition.

[1400] 3. Means of communication:

[1401] The portable device and the computing unit are connected via communication methods such as Bluetooth and Wi-Fi, which allows for the sending and receiving of voice data and translation data in real time.

[1402] 4. Dedicated applications:

[1403] Users can manage the system settings using a dedicated application installed on their smartphone or tablet, which allows them to select the target language and adjust the speech output speed and volume.

[1404] 5. Emotion recognition means:

[1405] The emotion recognition unit analyzes the user's facial expressions and tone of voice to recognize emotions, and has the function of adjusting the tone and speed of the translated speech based on the recognized emotions.

[1406] (Example)

[1407] Scenario 1: English and Japanese conversation:

[1408] User A (English speaker) and User B (Japanese speaker) are having a conversation. When User A says, "What time is the meeting?", the server captures the voice through the microphone of the eyeglasses-type device and converts it into text data "What time is the meeting?" using the Google Cloud Speech-to-Text API. The DeepL API is used to translate this to "What time is the meeting?", and if the emotion recognition means recognizes that User A's question contains tension, it generates voice that reflects that tension. Finally, the generated voice data is transmitted to User B through the speaker of the eyeglasses-type device.

[1409] Scenario 2: Configuration Management:

[1410] User B opens the application, changes the target language to French, and adjusts the speech output speed to a calmer tone. This setting information is sent to the server in real time and applied to the entire system. In the next conversation, "What time is the meeting?" is translated to "À quelle heure est la réunion?" and the speech is output in a calm, gentle tone.

[1411] In this way, the present invention provides a system that enables natural conversation between multiple languages ​​and realizes smooth communication in a variety of situations.

[1412] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1413] Step 1:

[1414] Acquiring voice input

[1415] When a user initiates a conversation, a microphone on the mobile device picks up what the other person is saying, using audio glasses such as Bose Frames.

[1416] Input: Other person's speech (audio data)

[1417] Data processing: Acquisition of audio data

[1418] Output: Digital audio data

[1419] Specific operation: When user A says "What time is the meeting?", the voice is recorded as digital data via the microphone of the portable device.

[1420] Step 2:

[1421] Sending audio data

[1422] The digital audio data captured by the portable device is transmitted to the computing unit via Bluetooth.

[1423] Input: Digital audio data

[1424] Data processing: Data transmission

[1425] Output: Received digital audio data

[1426] Specific operation: A portable device transmits digital audio data to a computing device via Bluetooth.

[1427] Step 3:

[1428] Converting audio data to text

[1429] The computing device uses the Google Cloud Speech-to-Text API to analyze the voice data and convert it into text data.

[1430] Input: Digital audio data

[1431] Data processing: speech recognition and text conversion

[1432] Output: Text data

[1433] Specific operation: The computing device converts the voice data into text: "What time is the meeting?"

[1434] Step 4:

[1435] Text data translation

[1436] The computing device uses the DeepL API to translate the text data into the target language.

[1437] Input: Text data

[1438] Data processing: Text translation

[1439] Output: Text data in the target language

[1440] Specific behavior: The computing device translates "What time is the meeting?" to "What time is the meeting?"

[1441] Step 5:

[1442] Emotion recognition

[1443] The computing device's emotion recognition means (such as Microsoft Azure's Emotion API) analyzes the user's facial expressions and vocal tone to recognize their emotions.

[1444] Input: User's facial expression data and voice tone

[1445] Data processing: Emotion analysis and recognition

[1446] Output: Recognized emotion data

[1447] Specific operation: The emotion recognition means reads the user's level of tension from their facial expressions and voice.

[1448] Step 6:

[1449] Adjust the tone and speed of your voice

[1450] The computing device adjusts the tone and rate of the translated speech based on the recognized emotion.

[1451] Input: Translation text data and emotion data

[1452] Data processing: adjusting the tone and speed of the voice

[1453] Output: Modified audio data

[1454] What it does: The computing device converts the translated text into a voice tone that reflects tension.

[1455] Step 7:

[1456] Generate audio output

[1457] The computing device generates tailored voice data using a TTS engine such as Amazon Polly.

[1458] Input: Modified audio data

[1459] Data processing: voice synthesis

[1460] Output: Generated audio data

[1461] Specific operation: The computing device synthesizes the adjusted voice data and prepares it for output.

[1462] Step 8:

[1463] Audio data output

[1464] The server sends the generated audio data to the portable device via Bluetooth, where it is transmitted to the user through the speaker.

[1465] Input: Generated audio data

[1466] Data processing: Sending voice data

[1467] Output: Audio output

[1468] Specific Actions: User B hears the translated audio of "Hello, how are you?" through the speaker of their mobile device.

[1469] Through this series of processes, the system realizes natural dialogue between users.

[1470] (Application example 2)

[1471] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1472] Currently, achieving natural translation and emotionally-rich speech output in real time is a challenging task in multilingual conversations and content distribution. Furthermore, in live streaming and recorded content where multiple language-speaking viewers simultaneously participate, viewers often cannot accurately convey the emotions of the performer or the video while watching in their own language. Therefore, multilingual streaming services are required to ensure that viewers can enjoy a natural and emotionally-rich experience.

[1473] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for acquiring the other party's speech via a voice input means; voice recognition means for converting the acquired voice into text; translation means for translating the converted text into a target language set by the user; voice output means for converting the translated text into voice and outputting it; a glasses-type device equipped with the voice input means and the voice output means; a computing device including the voice recognition means and the translation means; communication means for connecting the glasses-type device and the computing device; display means for translating the conversations of performers and videos in real time and displaying subtitles in a language selected by the viewer; means for adjusting the translation results using emotion analysis and displaying them as natural conversational text; and a dedicated application for the viewer to manage settings. This enables viewers to enjoy a natural and emotionally rich experience in a multilingual distribution service.

[1474] The "means for acquiring the other party's speech by voice input means" is a technique for acquiring the other party's speech by using a voice input device such as a microphone.

[1475] "Speech recognition means" refers to software or algorithms that convert captured speech data into text.

[1476] The "translation means" is a technology that translates the text converted by the speech recognition means into a target language set by the user.

[1477] "Audio output means" refers to technology that converts the translated text back into audio and outputs it via a speaker or the like.

[1478] The "eyeglasses-type device" is an eyeglasses-type device equipped with an audio input means and an audio output means, and is a portable communication device.

[1479] The term "arithmetic unit" is a general term for hardware and software for performing various processes such as speech recognition means and translation means.

[1480] The "communication means" is a technology for connecting the eyeglass-type device and the computing device and transmitting and receiving data.

[1481] "Display means" refers to devices or technologies that translate the dialogue of the performers or in the video in real time and display subtitles in the language selected by the viewer.

[1482] "Emotion analysis means" is a technology that adjusts the translation results based on the user's emotions and outputs them in a natural conversational style.

[1483] "Dedicated application" is software that allows viewers to manage and customize system settings.

[1484] This invention provides a system for realizing real-time translation and emotion recognition in natural conversations between multiple languages ​​and content distribution. Specific embodiments of this system are described below.

[1485] System Configuration

[1486] The system includes the following major components:

[1487] 1. Voice input means: A device for capturing what the other person is saying. For example, a microphone.

[1488] 2. Speech recognition tool: Software or algorithms that convert captured voice data into text. For example, we use the Google Cloud Speech-to-Text API.

[1489] 3. Translation method: The technology that translates the text into the target language. For example, using the Amazon Translate API.

[1490] 4. Voice output method: Technology that converts the translated text into voice and outputs it through a speaker. For example, IBM Watson Text to Speech API is used.

[1491] 5. Glasses-type device: A portable device equipped with audio input and output means.

[1492] 6. Computing unit: An integrated system of hardware and software for performing various processes.

[1493] 7. Communication method: Technology that connects the glasses device to the computing device, such as Bluetooth or Wi-Fi.

[1494] 8. Display means: A device for displaying the translation results to the viewer as subtitles in real time.

[1495] 9. Sentiment analysis: Technology to analyze the user's emotions and output the translation results in a natural conversational style. For example, we use the Microsoft Azure Emotion API.

[1496] 10. Dedicated application: Software that allows viewers to manage their system settings. For example, using React Native as the platform.

[1497] System Operation

[1498] 1. Acquiring voice input and converting it to text

[1499] When a user starts talking, the microphone of the eyeglass device picks up what is being said. The picked-up voice data is sent to a computing device and converted into text data by a voice recognition means.

[1500] 2. Text Translation

[1501] The text converted by the speech recognition means is translated into a target language set by the user using the translation means.

[1502] 3. Emotion recognition and voice output adjustment

[1503] The translated text is converted into voice data by taking into account the user's emotions through emotion analysis, and this voice data is output in an appropriate tone and played to the audience.

[1504] 4. Displaying the translation results

[1505] The translated text is provided to the viewer in real time as subtitles using a display means.

[1506] Usage example

[1507] For example, if a speaker says "Hello, everyone! Welcome to our event!" during a live stream, the system will convert this into text in real time and translate it into the target language (e.g., Japanese). The translated "Hello, everyone! Welcome to our event!" will be played in an "excited" tone based on emotion analysis and simultaneously displayed as subtitles.

[1508] Prompt example

[1509] Below are some examples of prompts used in this system:

[1510] Design a real-time translation system for live streaming that provides subtitles and speech translation according to the viewer's language of choice and realizes natural conversational tone that takes into account the speaker's emotions. This system includes speech recognition, text translation, emotion recognition, and speech synthesis modules. Each module should be implemented using a specific API, and viewers should be able to change the settings.

[1511] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1512] Step 1:

[1513] Acquiring voice input:

[1514] When a user starts speaking, the microphone on the eyeglasses picks up what is being said. The input voice data is sent as a digital signal to a computing device, where it is converted into a digital audio format such as linear PCM.

[1515] Step 2:

[1516] Speech to text:

[1517] The server sends the voice data received by the computing device to the Google Cloud Speech-to-Text API. The API analyzes the voice data and converts it into corresponding text data. The input is digital voice data, and the output is text data that converts the voice into text.

[1518] Step 3:

[1519] Text Translation:

[1520] The server sends the converted data to the Amazon Translate API. The input is the text data obtained by the speech recognition means, and the output is the text data translated into the target language. Here, since the user has set the target language in advance, the data is translated into the appropriate language based on that setting.

[1521] Step 4:

[1522] Emotion recognition:

[1523] The server sends the translated text data and the original audio data to the Microsoft Azure Emotion API to analyze the user's emotions. The input is the translated text data and audio data, and the output is the emotion analysis results and the emotion labels and scores based on them.

[1524] Step 5:

[1525] Text-to-Speech:

[1526] The server uses the IBM Watson Text to Speech API to convert the translated text data into speech data that reflects the results of sentiment analysis. The input is the translated text data and emotion label, and the output is synthesized speech data. This speech data is generated based on the voice characteristics (speed, tone, etc.) set by the user.

[1527] Step 6:

[1528] Displaying subtitles:

[1529] The computing device transmits the translated text data to the display device in real time and displays it as subtitles. The input is the translated text data, and the output is the subtitles displayed on the screen in real time, allowing viewers to visually confirm the translation content.

[1530] Step 7:

[1531] Audio Output:

[1532] The translated and emotion-analyzed speech data is output to the user from the speaker of the eyeglasses via a computing device. The input is synthesized speech data, and the output is speech that reaches the user's ears. This allows the user to hear the translated speech in real time.

[1533] Step 8:

[1534] Settings management:

[1535] The user manages various system settings (target language, audio characteristics, subtitle display, etc.) through a dedicated application. The input is the user's setting information, and the output is the setting data sent to the computing device. This allows the translation environment desired by the user to be properly applied.

[1536] The above is an explanation of the processing steps of the system program that realizes the application example, as well as its specific operations and input / output.

[1537] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1538] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1539] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1540] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1541] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1542] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1543] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1544] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1545] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1546] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1547] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1548] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1549] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1550] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1551] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1552] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1553] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1554] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1555] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1556] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1557] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1558] The following is further disclosed regarding the above embodiment.

[1559] (Claim 1)

[1560] A means for acquiring the speech of the other party by a voice input means;

[1561] a speech recognition means for converting the captured speech into text;

[1562] a translation means for translating the converted text into a target language set by the user;

[1563] a voice output means for converting the translated text into voice and outputting it;

[1564] a glasses-type device equipped with a voice input means and a voice output means;

[1565] a computing device including the speech recognition means and translation means;

[1566] A system including a communication means for connecting the eyeglass-type device and a computing device.

[1567] (Claim 2)

[1568] 10. The system of claim 1, further comprising means for a user to manage settings using a dedicated application.

[1569] (Claim 3)

[1570] 10. The system of claim 1, further comprising a display means for displaying the translation result as subtitles.

[1571] "Example 1"

[1572] (Claim 1)

[1573] A means for acquiring the speech of the other party by a voice input means;

[1574] a speech recognition means for converting the captured speech into text;

[1575] a translation means for translating the converted text into a target language set by the user;

[1576] a voice output means for converting the translated text into voice and outputting it;

[1577] a wearable device equipped with a voice input means and a voice output means;

[1578] an information processing device including the speech recognition means and translation means;

[1579] a communication means for connecting the wearable device to an information processing device;

[1580] A system that includes a means of communication to send and receive data in real time.

[1581] (Claim 2)

[1582] 10. The system of claim 1, further comprising means for a user to manage settings using a dedicated application.

[1583] (Claim 3)

[1584] 10. The system of claim 1, further comprising: means for displaying the translation result as subtitles.

[1585] "Application Example 1"

[1586] Claims

[1587] (Claim 1)

[1588] A means for acquiring the speech of the other party by a voice input means;

[1589] a speech recognition means for converting the captured speech into text;

[1590] a translation means for translating the converted text into a target language set by the user;

[1591] a voice output means for converting the translated text into voice and outputting it;

[1592] a glasses-type device equipped with a voice input means and a voice output means;

[1593] a computing device including the speech recognition means and translation means;

[1594] A system including a communication means for connecting the glasses-type device, a smartphone, and a computing device.

[1595] (Claim 2)

[1596] 10. The system of claim 1, further comprising means for a user to manage settings using a dedicated application.

[1597] (Claim 3)

[1598] The system according to claim 1, further comprising a display means for displaying the translation result as subtitles or a display means for displaying the translation result on a smartphone.

[1599] "Example 2: Combining Emotion Engines"

[1600] (Claim 1)

[1601] A voice input means for acquiring the speech of the other party;

[1602] a speech recognition means for converting the captured speech into text;

[1603] a translation means for translating the converted text into a target language set by the user;

[1604] a voice output means for converting the translated text into voice and outputting it;

[1605] a portable device equipped with a voice input means and a voice output means;

[1606] a computing device including the speech recognition means and translation means;

[1607] a communication means for connecting the portable device and a computing unit;

[1608] emotion recognition means for analyzing a user's facial expression and tone of voice to recognize emotions;

[1609] means for adjusting the tone and rate of the translated speech based on the recognized emotion; and

[1610] A system including:

[1611] (Claim 2)

[1612] 10. The system of claim 1, further comprising means for a user to manage settings using a dedicated application.

[1613] (Claim 3)

[1614] 10. The system of claim 1, further comprising a display means for displaying the translation result as subtitles.

[1615] "Application example 2 when combining emotion engines"

[1616] (Claim 1)

[1617] A means for acquiring the speech of the other party by a voice input means;

[1618] a speech recognition means for converting the captured speech into text;

[1619] a translation means for translating the converted text into a target language set by the user;

[1620] a voice output means for converting the translated text into voice and outputting it;

[1621] a glasses-type device equipped with a voice input means and a voice output means;

[1622] a computing device including the speech recognition means and translation means;

[1623] a communication means for connecting the eyeglass-type device and a computing device;

[1624] A display means that translates the conversations of the speakers and the video in real time and displays subtitles in the language selected by the viewer;

[1625] A means to adjust the translation results using sentiment analysis and display them in a natural conversational style;

[1626] A system that includes a dedicated application for viewers to manage their settings.

[1627] (Claim 2)

[1628] 10. The system of claim 1, further comprising means for audibly outputting the translated text in real time into the viewer's selected language.

[1629] (Claim 3)

[1630] 10. The system of claim 1, further comprising means for a viewer to set subtitle display of the translation result. [Explanation of symbols]

[1631] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring the speech of the other party by a voice input means; a speech recognition means for converting the captured speech into text; a translation means for translating the converted text into a target language set by the user; a voice output means for converting the translated text into voice and outputting it; a glasses-type device equipped with a voice input means and a voice output means; a computing device including the speech recognition means and translation means; A system including a communication means for connecting the eyeglass-type device and a computing device.

2. The system of claim 1 further comprising means for a user to manage settings using a dedicated application.

3. 2. The system according to claim 1, further comprising display means for displaying the translation result as subtitles.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A