System

A system with voice capture, speech recognition, translation, and feedback mechanisms addresses language barriers by enabling real-time communication and improving translation accuracy through user feedback.

JP2026022342APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123859
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Language barriers create significant obstacles in international business, multilingual communities, and multinational events, leading to misunderstandings and miscommunication, and existing language learning methods are inadequate for immediate communication needs.

Method used

A system comprising an input means for voice capture, speech recognition, translation, speech synthesis, and a learning mechanism that receives user feedback to improve translation quality, enabling real-time communication across languages.

Benefits of technology

Facilitates smooth communication between individuals speaking different languages by providing real-time translation and continuous improvement of translation accuracy through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022342000001_ABST
    Figure 2026022342000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: an input device for inputting a voice; a voice recognizer for converting the voice inputted by the input device into text data; a translator for translating the text data generated by the voice recognizer into a designated target language; an speech synthesis device for converting the translated text data generated by the translator into voice data; an output device for outputting the voice data generated by the speech synthesis device; and a learner for receiving feedback from a user and storing the feedback as learner data for the translator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Language barriers are a major issue in international business, multilingual communities, international students, and multinational events. Misunderstandings and miscommunication due to language differences occur frequently, creating major obstacles in business and personal interactions. Furthermore, language learning requires time and significant resources, and does not meet the needs of immediate communication. For this reason, technology is needed to facilitate smooth communication between different languages ​​in real time. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means. First, an input means is provided for inputting a user's voice, and a speech recognition means is arranged for converting the voice data input by the input means into text data. Next, a translation means is provided for translating the text data generated by the speech recognition means into a specified target language. Furthermore, a speech synthesis means is used for converting the translated text data generated by the translation means into voice data, and an output means is used for outputting the generated voice data. Finally, a system is provided that includes a learning means for receiving feedback on translation quality provided by the user and utilizing the feedback as learning data for the translation means, thereby continuously improving the accuracy and quality of translation.

[0006] "Input means" refers to a device or function for acquiring voice as data.

[0007] "Speech recognition means" refers to a device or technology for converting captured voice data into text data.

[0008] "Translation means" refers to a device or technology for converting text data into a specified target language.

[0009] "Speech synthesis means" refers to a device or technology for converting translated text data into speech data.

[0010] The "output means" is a device or function for outputting the generated audio data in a form that can be heard by the user.

[0011] A "training means" is a device or technology that receives user-provided feedback and trains the translation means to improve translation quality.

[0012] "Audio data" refers to data obtained by converting audio information acquired by an input means into a digital signal.

[0013] "Text data" refers to data in which character information generated by a voice recognition means or a translation means is converted into a digital signal.

[0014] "Feedback" means any rating or opinion regarding the quality of a translation provided by a User.

[0015] "Target language" refers to the language into which text data is converted by the translation means. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] MODE FOR CARRYING OUT THE INVENTION

[0038] The present invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. The detailed configuration and operation of the system will be described below.

[0039] System Configuration

[0040] The system consists of the following main components:

[0041] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[0042] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[0043] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[0044] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[0045] 5. Output means: A device that allows the user to listen to the generated audio data. Specifically, it is a speaker built into an earphone.

[0046] 6. Learning tool: A device or technology that receives feedback from users and trains the translation tool. This includes a feedback processing system on the server.

[0047] System Operation

[0048] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[0049] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[0050] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to a translation engine on the server.

[0051] The server's AI translation engine receives the text data and translates it into the specified target language. During the translation process, it uses natural language processing technology to understand the context and select the appropriate translation.

[0052] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0053] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[0054] Specific examples

[0055] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[0056] 1. User A says "Hello" in English.

[0057] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0058] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0059] 4. The smartphone sends the text data to the server.

[0060] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0061] 6. The server sends the translation results back to the smartphone.

[0062] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[0063] 8. The earphone speaker will say "hello."

[0064] 9. User B hears the translated "hello" in real time.

[0065] This system allows users A and B to communicate smoothly across language barriers.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[0069] Step 2:

[0070] The user sets the input language and target language in the translation app, and also adjusts the volume and enables noise cancellation as needed.

[0071] Step 3:

[0072] When the user starts speaking, the microphone built into the earphones captures the user's voice, and the captured voice data is sent to the smartphone in real time.

[0073] Step 4:

[0074] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[0075] Step 5:

[0076] The terminal transmits the converted text data to a server via the Internet.

[0077] Step 6:

[0078] The server receives the text data and translates it into the specified target language using an AI translation engine. The translation engine uses natural language processing to understand the context and select the appropriate translation.

[0079] Step 7:

[0080] The server returns the translated text data to the device, and also prepares to receive feedback as learning data to improve translation accuracy.

[0081] Step 8:

[0082] The terminal passes the translated text data received from the server to a speech synthesis engine and converts it into voice data.

[0083] Step 9:

[0084] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[0085] Step 10:

[0086] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[0087] Step 11:

[0088] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[0089] Example 1

[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0091] It is difficult for people who speak different languages ​​to communicate smoothly in real time. Furthermore, the quality of translation must be improved. Therefore, there is a need for a system that can realize real-time translation in multiple languages ​​and continuously improve its translation performance by receiving feedback.

[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0093] In this invention, the server includes a connection means for automatically connecting to a user's device, an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, a translation means for translating the text data generated by the speech recognition means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by the user and saving the feedback as learning data for the translation means. This enables people speaking different languages ​​to communicate smoothly in real time. Furthermore, by utilizing user feedback, translation performance can be continuously improved.

[0094] "User" refers to an individual who uses the system to communicate with others who speak different languages.

[0095] "Device" refers to a device that includes hardware for use by a user, such as earphones or smartphones.

[0096] "Connection means" refers to a method or technology for connecting a user's device to each component of the system. Specifically, this applies to wireless communication technologies such as Bluetooth and Wi-Fi.

[0097] "Input means" refers to a device or technology for inputting sound into the system, specifically a microphone built into earphones.

[0098] "Voice data" refers to digital audio information that records a user's speech.

[0099] "Text data" refers to linguistic information in character format converted by a speech recognition means.

[0100] "Speech recognition means" refers to a technique or device for converting voice data into text data.

[0101] "Translation means" refers to technology or equipment for converting text data into a specified target language. Specifically, it includes an AI translation engine installed on a server.

[0102] "Speech synthesis means" refers to technology or equipment for converting translated text data into speech data.

[0103] "Output means" refers to a device or technology for transmitting synthesized voice data to the user. Specifically, this applies to the speaker built into the earphone.

[0104] "Feedback" refers to evaluation information about the quality of a translation provided by a user.

[0105] "Training means" refers to a technique or device for analyzing received feedback and improving the performance of the translation means.

[0106] MODE FOR CARRYING OUT THE INVENTION

[0107] The present invention relates to a system that enables people who speak different languages ​​to communicate smoothly in real time. The detailed configuration and operation of the system will be described below.

[0108] System Hardware

[0109] The system hardware consists of the following major components:

[0110] 1. Earphones: Earphones with built-in microphones and speakers that are responsible for inputting and outputting audio data.

[0111] 2. Smartphone: A translation application is installed and performs voice recognition, data transmission and reception, and voice synthesis.

[0112] 3. Server: Equipped with an AI translation engine, it translates text data.

[0113] Software used

[0114] The software used in this system is as follows:

[0115] 1. Smartphone application:

[0116] Speech recognition engine: Converts voice data into text data.

[0117] Speech synthesis engine: Converts translated text data into speech data.

[0118] Feedback feature: Receive feedback from users.

[0119] 2. AI translation engine on the server:

[0120] Uses natural language processing (NLP) technology to perform highly accurate translations.

[0121] System Operation

[0122] The user puts on the earphones and launches a translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. When the user begins to speak, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time.

[0123] The voice recognition engine of the smartphone terminal converts the received voice data into text data. For example, if a user says "Hello," the voice recognition engine generates the text data "Hello." The generated text data is then sent to the server.

[0124] The server's AI translation engine receives the text data and translates it into the specified target language. The server then uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The translation result is then sent back from the server to the smartphone.

[0125] The speech synthesis engine on the smartphone, which is the device that receives the translation results, converts the translated text data into voice data. This voice data is sent to the earphones. The earphone speakers pronounce the translated voice. User B can hear the voice saying "Hello" in real time.

[0126] After the conversation is over, the user provides feedback on the quality of the translation through the application, which then sends the feedback data to the server, where it is analyzed by the server's feedback processing system and used as training data for the AI ​​translation engine.

[0127] Specific examples

[0128] The following example shows how the system works when an English-speaking user A and a Japanese-speaking user B communicate in real time:

[0129] 1. User A says "Hello" in English.

[0130] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0131] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0132] 4. The smartphone sends the text data to the server.

[0133] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0134] 6. The server sends the translation results back to the smartphone.

[0135] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[0136] 8. The earphone speaker will say "hello."

[0137] 9. User B hears the translated "hello" in real time.

[0138] Prompt Sentence Examples

[0139] "Please explain how a system works, allowing users to communicate in real time with others who speak different languages. Please provide detailed examples of how data is processed using specific hardware (earphones, smartphone, server) and software (speech recognition engine, translation engine, speech synthesis engine)."

[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0141] Step 1: Connection and Setup

[0142] The user puts on the dedicated earphones and launches the translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. The input for this step is the user's language setting and earphone connection information. The output is the completed setting state.

[0143] Step 2: Capture audio

[0144] When a user starts talking, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time. The input of this step is the voice from the user, and the output is the voice data sent to the smartphone.

[0145] Step 3: Speech to text

[0146] The device (smartphone) converts the received voice data into text data using a voice recognition engine. For example, the voice "Hello" is converted into the text data "Hello." The input of this step is the captured voice data, and the output is the generated text data.

[0147] Step 4: Translate the text

[0148] The server's AI translation engine receives the text data and translates it into the specified target language. The server uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The input of this step is the text data, and the output is the translated text data.

[0149] Step 5: Speech synthesis of the translation result

[0150] The device (smartphone) converts the translated text data returned from the server into voice data using a speech synthesis engine. The text "Hello" is converted into voice data saying "Hello." The input of this step is the translated text data, and the output is the generated voice data.

[0151] Step 6: Output the translated audio

[0152] The earphone speaker provides the generated voice data to the user. User B can hear the voice saying "Hello" in real time. The input of this step is the synthesized voice data, and the output is the voice the user hears.

[0153] Step 7: Provide feedback

[0154] After the conversation is completed, the user provides feedback on the quality of the translation through the smartphone application, for example, by entering feedback such as "very satisfied." The input of this step is the feedback data from the user, and the output is the feedback data sent to the server.

[0155] Step 8: Analyze feedback and learn

[0156] The server analyzes the received feedback data and stores it as training data for the AI ​​translation engine, thereby continuously improving translation performance. The input of this step is the received feedback data, and the output is updated training data.

[0157] (Application example 1)

[0158] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0159] The goal of this project is to achieve efficient real-time communication between staff members who speak different languages ​​at a logistics center where multinational staff work. This will enable improved efficiency and accuracy of work. However, with the previous system, there were barriers to communication between staff members, which hindered the smooth execution of work.

[0160] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0161] In this invention, the server includes an input means for inputting speech, a speech recognition means, a translation means, a speech synthesis means, an output means, a learning means, and a means used to enable efficient communication among multinational staff at a logistics center, thereby enabling real-time speech translation between multiple languages ​​and communication based on the translation.

[0162] The "input means for inputting voice" is a device that captures the voice uttered by the user as digital data.

[0163] "Speech recognition means" is a technology for converting captured voice data into text data.

[0164] "Translation means" refers to a technique for converting text data generated by speech recognition means into a target language.

[0165] "Speech synthesis means" is a technology that converts text data translated into a target language into speech data.

[0166] The "output means for outputting voice data" is a device that allows the user to listen to the translated voice data.

[0167] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" refers to a database and processing system for receiving evaluations from users and improving translation accuracy based on them.

[0168] "Means used to enable multinational staff to communicate efficiently at logistics centers" refers to components of a system that enables multiple staff members who speak different languages ​​to communicate smoothly through a voice translation system.

[0169] System configuration

[0170] This invention is a system for realizing real-time communication among multinational staff in a logistics center. This system is composed of the following main elements:

[0171] 1. A microphone built into earphones worn by the user as an input means for inputting voice.

[0172] 2. A speech recognition engine built into a smartphone that converts voice data into text data as a means of speech recognition.

[0173] 3. As a translation tool, an AI translation engine on a server that translates text data into a specified target language.

[0174] 4. A smartphone speech synthesis engine that converts translated text data into speech data as a speech synthesis method.

[0175] 5. A speaker built into the earphone as an output means for outputting audio data.

[0176] 6. A feedback processing system on the server that receives feedback provided by users as a learning means and stores it as learning data for the translation means.

[0177] 7. A system configuration that combines the above elements as a means to enable multinational staff to communicate efficiently in a logistics center.

[0178] How it works

[0179] The user uses the system by putting on earphones and launching a dedicated translation application installed on their smartphone. The system operates as follows:

[0180] 1. Voice input:

[0181] When the user speaks, the microphone built into the earphones captures the sound, and the audio data is sent to the smartphone in real time.

[0182] 2. Speech Recognition:

[0183] The smartphone's voice recognition engine converts the received voice data into text data.

[0184] 3. Translation:

[0185] The text data is sent to a server, where an AI translation engine translates it into the target language.

[0186] 4. Speech synthesis:

[0187] The translated text data is then sent back to the smartphone, where it is converted into voice data by a voice synthesis engine.

[0188] 5. Audio output:

[0189] The converted audio data is output to the user through the earphone speaker.

[0190] 6. Feedback and learning:

[0191] After the conversation, the user provides feedback on the quality of the translation, which is sent to the server. The feedback processing system analyzes this and stores it as learning data for the AI ​​translation engine, contributing to future improvements in translation accuracy.

[0192] Hardware and software used

[0193] Hardware:

[0194] Earphones (with built-in microphone and speaker)

[0195] Smartphone

[0196] software:

[0197] SpeechRecognition (speech recognition library)

[0198] GoogleTrans (translation library)

[0199] gTTS (Text-to-Speech Library)

[0200] Playsound (sound playback library)

[0201] Specific examples

[0202] Logistics centers often employ foreign staff who do not understand Japanese. With this system, instructions such as "Please pick up the items from aisle 3" can be given in Japanese and instantly translated into the staff's native language, ensuring smooth communication.

[0203] Prompt Sentence Examples

[0204] "Write Python code to translate English audio into Japanese and play it back as audio."

[0205] This will enable smooth communication between multinational staff at the logistics center, significantly improving work efficiency.

[0206] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0207] Step 1:

[0208] The user puts on the earphones and launches the translation application installed on their smartphone. When the application is launched, the system automatically connects to the earphones via Bluetooth and is ready for voice input. The user then sets the input language and target language in the app.

[0209] Input: User utterance

[0210] Output: Ready notification, setting language information

[0211] Step 2:

[0212] When a user speaks, the microphone built into the earphones captures the voice. This voice data is sent to a smartphone in real time via Bluetooth. The user's voice is the input data.

[0213] Input: User's voice data

[0214] Output: Microphone captured audio data

[0215] Step 3:

[0216] The smartphone's voice recognition engine converts the received voice data into text data. Specifically, the SpeechRecognition library processes the voice data and generates text data.

[0217] Input: Audio data

[0218] Output: Converted text data

[0219] Step 4:

[0220] The text data is sent to the server via the internet. The server's AI translation engine uses the Google Translate API to translate the text into the target language. The translated text data is output.

[0221] Input: Text data, target language information

[0222] Output: Translated text data

[0223] Step 5:

[0224] The translated text data is returned to the smartphone and converted into audio data by the speech synthesis engine. Specifically, the gTTS library converts the text data into an audio file.

[0225] Input: Translated text data

[0226] Output: Synthesized speech data

[0227] Step 6:

[0228] The generated voice data is sent from the smartphone via Bluetooth to the earphones, which then output the audio through their speakers, allowing the user to hear the audio in the target language in real time.

[0229] Input: Audio data

[0230] Output: Audio output from earphones

[0231] Step 7:

[0232] After the conversation is over, the user provides feedback on the quality of the translation via the app. The feedback data is sent to the server and stored in the feedback processing system. This data is used as training data for the translation engine to improve future translation accuracy.

[0233] Input: User feedback data

[0234] Output: Improved learning model with saved feedback data

[0235] Through these steps, efficient real-time communication between multinational staff at a logistics center is achieved.

[0236] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0237] MODE FOR CARRYING OUT THE INVENTION

[0238] This invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide translation results that correspond to those emotions. The detailed configuration and operation of the system are described below.

[0239] System Configuration

[0240] The system consists of the following main components:

[0241] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[0242] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[0243] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[0244] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[0245] 5. Output means: A device for outputting the generated audio data in a form that can be heard by the user. Specifically, it is a speaker built into an earphone.

[0246] 6. Learning means: A device or technology for receiving feedback from users and storing it as learning data for the translation means. This includes a feedback processing system on a server.

[0247] 7. Emotion Engine: A device or technology that recognizes the user's emotions and reflects them in the translation process. Emotions are recognized by analyzing voice tone, speech rate, and language patterns.

[0248] System Operation

[0249] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[0250] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[0251] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to the emotion engine and the translation engine on the server.

[0252] The emotion engine analyzes voice tone, speech rate, and language patterns to recognize the user's emotion, and this emotion data is sent to the translation engine.

[0253] The server's AI translation engine receives the text data and translates it into the specified target language, taking into account the recognized emotional data. During the translation process, natural language processing technology is used to understand the context and emotions and select the appropriate translation.

[0254] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0255] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[0256] Specific examples

[0257] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[0258] 1. User A says "Hello" in English. His voice sounds cheerful.

[0259] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0260] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0261] 4. The smartphone sends the text data to the emotion engine and server.

[0262] 5. The emotion engine analyzes User A's tone and speech rate and recognizes the emotion "fun." This emotion data is sent to the server.

[0263] 6. The server's translation means translates "Hello" to the Japanese "Konnichiwa" and maintains a happy tone of voice.

[0264] 7. The server sends the translation results back to the smartphone.

[0265] 8. The smartphone's voice synthesis function converts "hello" into voice data in a pleasant voice tone.

[0266] 9. The earphone speaker pronounces "hello" in a cheerful voice.

[0267] 10. User B hears the translated "Hello" in real time in a happy voice.

[0268] This system allows users A and B to communicate smoothly and emotionally, overcoming language barriers.

[0269] The processing flow will be explained below.

[0270] Step 1:

[0271] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[0272] Step 2:

[0273] The user sets the input language and target language in the translation app, and also enables the emotion recognition function.

[0274] Step 3:

[0275] When the user starts speaking, the microphone built into the earphones captures the voice, and this voice data is sent to the smartphone in real time.

[0276] Step 4:

[0277] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[0278] Step 5:

[0279] The terminal sends the converted text data to the emotion engine, which analyzes the voice tone, speech rate, and language patterns to recognize the user's emotion, and the recognized emotion data is sent to the server.

[0280] Step 6:

[0281] The terminal transmits the text data to a server via the Internet.

[0282] Step 7:

[0283] The server receives the text data and emotional data, and the AI ​​translation engine translates it into the specified target language. The translation engine also takes into account the emotional data, understands the context, and selects the appropriate translation and tone of voice.

[0284] Step 8:

[0285] The server returns the translation results that reflect the emotional data to the terminal.

[0286] Step 9:

[0287] The device passes the translated text data received from the server to a speech synthesis engine and converts it into voice data, while maintaining the voice tone that reflects the emotional data.

[0288] Step 10:

[0289] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[0290] Step 11:

[0291] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[0292] Step 12:

[0293] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[0294] Example 2

[0295] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0296] There is a demand for systems that allow people who speak different languages ​​to communicate smoothly in real time. However, even if traditional translation systems provide accurate translations, they tend to fail to reflect emotional expressions, resulting in incomplete communication. Another issue is the lack of a mechanism for improving the system's accuracy through feedback.

[0297] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, an emotion analysis means for analyzing speech tone, speech rate, and language patterns to recognize a user's emotion, a translation means for translating the text data and emotion data generated by the speech recognition means and the emotion analysis means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by a user and saving the feedback as learning data for the translation means. This enables smooth real-time communication that incorporates emotions between users who speak different languages.

[0298] "Input means" refers to a device for acquiring audio, specifically a microphone built into an earphone.

[0299] The "voice recognition means" is a device or technology for converting voice data acquired by the input means into text data.

[0300] An "emotion analysis means" is a device or technology for recognizing a user's emotions by analyzing voice tone, speech rate, and language patterns.

[0301] The "translation means" is a device or technology for translating the text data and emotion data generated by the speech recognition means and emotion analysis means into a specified target language.

[0302] The "speech synthesis means" refers to a device or technology for converting the translated text data generated by the translation means into speech data.

[0303] The "output means" is a device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically refers to a speaker built into an earphone.

[0304] "Learning means" refers to a device or technology that receives feedback on translation quality provided by a user and stores it as training data for the translation means.

[0305] The present invention is a system that enables people who speak different languages ​​to communicate smoothly in real time by using advanced technologies that integrate speech recognition, emotion analysis, translation, speech synthesis, and feedback processing.

[0306] System configuration

[0307] The system consists of the following main components:

[0308] 1. Input means: A device for capturing audio, specifically a microphone built into earphones.

[0309] 2. Speech recognition means: A device for converting voice data acquired by the input means into text data, using a voice recognition engine built into the smartphone.

[0310] 3. Emotion analysis means: A device for analyzing voice tone, speech rate, and language patterns to recognize the user's emotions.

[0311] 4. Translation means: A device for translating text data and emotional data generated by the speech recognition means and emotional analysis means into a specified target language. It uses an AI translation engine on the server.

[0312] 5. Speech synthesis means: A device for converting the translated text data generated by the translation means into speech data, and uses a speech synthesis engine built into the smartphone.

[0313] 6. Output means: A device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically, a speaker built into the earphone is used.

[0314] 7. Learning means: A device that receives feedback on translation quality provided by users, stores it as learning data for the translation means, and improves the system's performance. A feedback processing system on the server is used.

[0315] System Operation

[0316] The user puts on the earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. The user sets the input language and target language to be used in the application. This setting information is saved on the smartphone.

[0317] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data in real time to the smartphone. The smartphone then uses a voice recognition engine to convert the voice data into text data. The conversion result is then sent to the emotion analysis means and the server.

[0318] The device's emotion analysis means analyzes voice tone, speech rate, and language patterns to recognize the user's emotions. This emotion data is sent to the translation means. The server's AI translation engine receives the text data and emotion data and translates it into the specified target language. The AI ​​translation engine understands the context and emotion and selects the appropriate translation.

[0319] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0320] After the conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server and analyzed by the server's learning means and used to improve the performance of the translation means.

[0321] Specific operation example

[0322] For example, when an English-speaking user A and a Japanese-speaking user B communicate in real time, the system operates as follows:

[0323] 1. User A launches the translation app and puts on earphones.

[0324] 2. User A sets English as the input language and Japanese as the target language on the app.

[0325] 3. User A says "Hello" happily.

[0326] 4. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0327] 5. The device uses a voice recognition means to convert the voice into text data such as "Hello."

[0328] 6. This data is sent to the server and sentiment analysis means.

[0329] 7. The device's emotion analysis means recognizes the emotion "fun" and sends this information to the server.

[0330] 8. The server's AI translation engine translates "Hello" into "Hello" taking into account the "joyful" emotion.

[0331] 9. The server sends the translation results back to the smartphone.

[0332] 10. The device's voice synthesis means converts "hello" into voice data in a "happy" voice tone.

[0333] 11. The earphone speaker says "Hello" to User B in a cheerful voice.

[0334] 12. User B hears "Hello" in real time and continues the conversation.

[0335] 13. After the conversation, User A provides feedback on the quality of the translation through the app, which is then sent to the server.

[0336] 14. The server's learning mechanism uses this feedback to improve the performance of the translation engine.

[0337] Prompt Sentence Examples

[0338] When User A says "Hello" in English, the earphone microphone captures the audio and sends it to the smartphone. The smartphone's speech recognition engine converts "Hello" into text, and the emotion analysis means recognizes the emotion as "happy." This information is sent to the translation engine on the server, and the audio converted to "Hello" in a happy voice tone is played from the earphone speaker.

[0339] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0340] Step 1:

[0341] The user puts on the system's earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. A connection confirmation message is displayed on the smartphone, and the user sets the input language and target language. This setting information is saved on the smartphone.

[0342] Specific behavior:

[0343] Input: User puts in earphones and launches translation app.

[0344] Data processing / calculation: The app checks the earphone connection and receives language setting information.

[0345] Output: A connection completion message and a setting confirmation message are displayed.

[0346] Step 2:

[0347] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data to the smartphone in real time.

[0348] Specific behavior:

[0349] Input: User's voice.

[0350] Data processing / calculation: The microphone captures the sound and converts it into digital sound data.

[0351] Output: Digital audio data is sent to the smartphone.

[0352] Step 3:

[0353] The terminal (smartphone) uses a voice recognition means to convert the voice data into text data, and the conversion result is sent to the emotion analysis means and the server.

[0354] Specific behavior:

[0355] Input: Digital audio data.

[0356] Data processing / calculation: The voice recognition engine analyzes the voice data and converts it into text data.

[0357] Output: The text data and the voice data are sent to the emotion analysis means and the server.

[0358] Step 4:

[0359] The emotion analysis means of the terminal analyzes the voice tone, speech rate, and language pattern to recognize the user's emotion, and this emotion data is sent to the translation means.

[0360] Specific behavior:

[0361] Input: Audio data.

[0362] Data processing / calculation: Emotion analysis means analyzes voice tone, speaking rate, and language patterns to extract emotions.

[0363] Output: The recognized emotion data is sent to the translation means.

[0364] Step 5:

[0365] The server's translation means receives the text data and emotion data and translates it into the target language. The AI ​​translation engine takes context and emotion into account when translating.

[0366] Specific behavior:

[0367] Input: Text data and emotion data.

[0368] Data processing / calculation: The translation engine analyzes the context and sentiment to generate an appropriate translation.

[0369] Output: The translated text data is generated and sent back to the smartphone.

[0370] Step 6:

[0371] The terminal receives the translation result returned from the server and converts the translated text data into voice data using a voice synthesis means.

[0372] Specific behavior:

[0373] Input: Translation text data.

[0374] Data processing / calculation: The speech synthesis engine converts text data into speech data.

[0375] Output: Audio data is generated.

[0376] Step 7:

[0377] The generated audio data is transmitted from the terminal to earphones, allowing the user to listen to it in real time.

[0378] Specific behavior:

[0379] Input: Audio data.

[0380] Data processing / calculation: Transfer data to earphones.

[0381] Output: Audio data is played from the earphone speaker.

[0382] Step 8:

[0383] After completing a conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server where it is analyzed by the server's learning means and used to improve the performance of the translation means.

[0384] Specific behavior:

[0385] Input: User feedback.

[0386] Data processing / calculation: Feedback data is sent to the server and analyzed.

[0387] Output: The analysis results are added to the translation engine's training data, improving translation performance.

[0388] (Application example 2)

[0389] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0390] There is a need for a method that allows people who speak different languages ​​to communicate smoothly in real time, including expressing their emotions. In particular, in the food delivery scene, a system is needed that allows delivery personnel and customers to communicate without worrying about language differences.

[0391] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for inputting voice, voice recognition means for converting voice data input by the input means into text data, translation means for translating the text data generated by the voice recognition means into a specified target language, voice synthesis means for converting the translated text data generated by the translation means into voice data, output means for outputting the voice data generated by the voice synthesis means, learning means for receiving feedback on translation quality provided by the user and saving it as learning data for the translation means, and an emotion engine that recognizes emotions by analyzing voice tone, speaking rate, and language patterns. This enables real-time communication incorporating emotions even when the delivery person and the customer speak different languages.

[0392] The "input means for inputting voice" is a device for capturing the user's speech.

[0393] "Speech recognition means for converting voice data input by an input means into text data" refers to a device or technology for converting voice input into data in text format.

[0394] "Translation means for translating text data generated by speech recognition means into a specified target language" refers to a device or technology for converting the generated text data into the target language.

[0395] "Speech synthesis means for converting translated text data generated by the translation means into speech data" refers to a device or technology for converting translated text into speech-format data.

[0396] The "output means for outputting the voice data generated by the voice synthesis means" is a device that enables the user to listen to the voice data.

[0397] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" is a technology for collecting user feedback and improving translation accuracy based on that data.

[0398] The "emotion engine that recognizes emotions by analyzing voice tone, speech rate, and language patterns" is a technology that analyzes the tone and rate of speech, the use of words, etc. to understand the speaker's emotions.

[0399] This invention is a system for realizing emotional real-time communication between delivery personnel and customers even when they speak different languages. The system includes the following elements:

[0400] 1. Input means: An input device for capturing the user's voice. Specifically, a microphone built into a smartphone or a microphone built into earphones can be used.

[0401] 2. Speech recognition means: Software for converting acquired voice data into text data. This is achieved by using the Python speech_recognition library.

[0402] 3. Translation tool: Software for translating the text data generated by the speech recognition tool into the target language. Translation is performed using Google's googletrans library.

[0403] 4. Speech synthesis means: Software for converting translated text data into speech data. This is achieved using the gTTS (Google Text-to-Speech) library.

[0404] 5. Output means: A device for transmitting the generated voice data to the user. Specifically, the speaker of a smartphone or a speaker of an earphone can be used.

[0405] 6. Learning method: A technology for collecting feedback data from users and using it to improve translation performance. The data is sent to the server for storage and analysis.

[0406] 7. Emotion Engine: Technology that analyzes voice tone, speaking rate, and language patterns to recognize user emotions, resulting in more natural and emotionally relevant translation results.

[0407] Hardware and Software Used

[0408] 1. Hardware:

[0409] Smartphone (microphone, speaker)

[0410] Earphones (microphone, speaker)

[0411] 2. Software:

[0412] Python

[0413] speech_recognition library (speech recognition)

[0414] googletrans library (translation)

[0415] gTTS library (speech synthesis)

[0416] Explanation of data processing and data calculation

[0417] The server receives voice data input via a smartphone or earphones and converts it into text data using a voice recognition means. This text data is then converted into the target language using a translation means. The translated text is converted back into voice data using a voice synthesis means and transmitted to the user via an output means. In addition, the learning means analyzes user feedback and uses it to improve system performance. During this series of processes, the emotion engine analyzes the user's emotions and reflects them in the translation results.

[0418] Specific examples

[0419] For example, consider the case where a delivery person says, "This is a delivery item." When the delivery person says, "I've brought your order" in English, this voice is captured by the smartphone's microphone. The speech recognition means converts this into text data, obtaining the string "I've brought your order." The translation means then translates this into Japanese, generating the text data "This is a delivery item." The speech synthesis means converts this Japanese text into speech data, which is finally conveyed to the customer via the smartphone or earphone speaker.

[0420] Example prompts to input to the generative AI model

[0421] Create a real-time translation assistant to help people who speak different languages ​​communicate smoothly. When a delivery person says "Delivery here" to a customer, it needs to translate from English to Japanese. This translation assistant will be realized using speech recognition, text translation, and speech synthesis.

[0422] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0423] Step 1:

[0424] Input and Output

[0425] A user starts the smartphone app and speaks into the microphone. The input is the user's voice. The output is recorded as audio data on the smartphone.

[0426] Specific actions

[0427] The user launches the smartphone app and says, "I've brought your order" in English. The smartphone's microphone captures the voice and generates voice data.

[0428] Step 2:

[0429] Input and Output

[0430] The voice data is sent to the voice recognition means. At this time, the input is the voice data generated in step 1. The output is the text data "I've brought your order" generated by the voice recognition means.

[0431] Specific actions

[0432] The smartphone's voice recognition engine (speech_recognition library) analyzes the voice data and converts it into corresponding text data.

[0433] Step 3:

[0434] Input and Output

[0435] The text data is sent to the translation means. At this time, the input is the text data "I've brought your order" generated in step 2. The output is the text data "Delivery item" in the target language generated by the translation means.

[0436] Specific actions

[0437] The smartphone's Google Trans library translates the text data into Japanese.

[0438] Step 4:

[0439] Input and Output

[0440] The text data in the target language is sent to the speech synthesis means. At this time, the input is the text data "Delivery item" generated in step 3. The output is speech data generated by the speech synthesis means.

[0441] Specific actions

[0442] The smartphone's gTTS library converts the text data into audio data and generates an audio file.

[0443] Step 5:

[0444] Input and Output

[0445] The audio data is sent to an output means, where the input is the audio data generated in step 4. The output is audio that can be heard by the user.

[0446] Specific actions

[0447] The translated voice message "Delivery here" is played to the customer through the speaker on their smartphone or earphones.

[0448] Step 6:

[0449] Input and Output

[0450] Feedback from the user is sent to the learning means. At this time, the input is the feedback on translation quality provided by the user. The output is the feedback data stored on the server and becomes learning data.

[0451] Specific actions

[0452] Users provide feedback on the quality of the translation through the app, which is then sent to the server, which analyzes the data and stores it in the translation engine's training database.

[0453] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0454] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0455] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0456] [Second embodiment]

[0457] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0458] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0459] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0460] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0461] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0462] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0463] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0464] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0465] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0466] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0467] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0468] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0469] MODE FOR CARRYING OUT THE INVENTION

[0470] The present invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. The detailed configuration and operation of the system will be described below.

[0471] System Configuration

[0472] The system consists of the following main components:

[0473] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[0474] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[0475] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[0476] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[0477] 5. Output means: A device that allows the user to listen to the generated audio data. Specifically, it is a speaker built into an earphone.

[0478] 6. Learning tool: A device or technology that receives feedback from users and trains the translation tool. This includes a feedback processing system on the server.

[0479] System Operation

[0480] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[0481] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[0482] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to a translation engine on the server.

[0483] The server's AI translation engine receives the text data and translates it into the specified target language. During the translation process, it uses natural language processing technology to understand the context and select the appropriate translation.

[0484] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0485] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[0486] Specific examples

[0487] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[0488] 1. User A says "Hello" in English.

[0489] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0490] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0491] 4. The smartphone sends the text data to the server.

[0492] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0493] 6. The server sends the translation results back to the smartphone.

[0494] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[0495] 8. The earphone speaker will say "hello."

[0496] 9. User B hears the translated "hello" in real time.

[0497] This system allows users A and B to communicate smoothly across language barriers.

[0498] The processing flow will be explained below.

[0499] Step 1:

[0500] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[0501] Step 2:

[0502] The user sets the input language and target language in the translation app, and also adjusts the volume and enables noise cancellation as needed.

[0503] Step 3:

[0504] When the user starts speaking, the microphone built into the earphones captures the user's voice, and the captured voice data is sent to the smartphone in real time.

[0505] Step 4:

[0506] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[0507] Step 5:

[0508] The terminal transmits the converted text data to a server via the Internet.

[0509] Step 6:

[0510] The server receives the text data and translates it into the specified target language using an AI translation engine. The translation engine uses natural language processing to understand the context and select the appropriate translation.

[0511] Step 7:

[0512] The server returns the translated text data to the device, and also prepares to receive feedback as learning data to improve translation accuracy.

[0513] Step 8:

[0514] The terminal passes the translated text data received from the server to a speech synthesis engine and converts it into voice data.

[0515] Step 9:

[0516] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[0517] Step 10:

[0518] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[0519] Step 11:

[0520] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[0521] Example 1

[0522] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0523] It is difficult for people who speak different languages ​​to communicate smoothly in real time. Furthermore, the quality of translation must be improved. Therefore, there is a need for a system that can realize real-time translation in multiple languages ​​and continuously improve its translation performance by receiving feedback.

[0524] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0525] In this invention, the server includes a connection means for automatically connecting to a user's device, an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, a translation means for translating the text data generated by the speech recognition means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by the user and saving the feedback as learning data for the translation means. This enables people speaking different languages ​​to communicate smoothly in real time. Furthermore, by utilizing user feedback, translation performance can be continuously improved.

[0526] "User" refers to an individual who uses the system to communicate with others who speak different languages.

[0527] "Device" refers to a device that includes hardware for use by a user, such as earphones or smartphones.

[0528] "Connection means" refers to a method or technology for connecting a user's device to each component of the system. Specifically, this applies to wireless communication technologies such as Bluetooth and Wi-Fi.

[0529] "Input means" refers to a device or technology for inputting sound into the system, specifically a microphone built into earphones.

[0530] "Voice data" refers to digital audio information that records a user's speech.

[0531] "Text data" refers to linguistic information in character format converted by a speech recognition means.

[0532] "Speech recognition means" refers to a technique or device for converting voice data into text data.

[0533] "Translation means" refers to technology or equipment for converting text data into a specified target language. Specifically, it includes an AI translation engine installed on a server.

[0534] "Speech synthesis means" refers to technology or equipment for converting translated text data into speech data.

[0535] "Output means" refers to a device or technology for transmitting synthesized voice data to the user. Specifically, this applies to the speaker built into the earphone.

[0536] "Feedback" refers to evaluation information about the quality of a translation provided by a user.

[0537] "Training means" refers to a technique or device for analyzing received feedback and improving the performance of the translation means.

[0538] MODE FOR CARRYING OUT THE INVENTION

[0539] The present invention relates to a system that enables people who speak different languages ​​to communicate smoothly in real time. The detailed configuration and operation of the system will be described below.

[0540] System Hardware

[0541] The system hardware consists of the following major components:

[0542] 1. Earphones: Earphones with built-in microphones and speakers that are responsible for inputting and outputting audio data.

[0543] 2. Smartphone: A translation application is installed and performs voice recognition, data transmission and reception, and voice synthesis.

[0544] 3. Server: Equipped with an AI translation engine, it translates text data.

[0545] Software used

[0546] The software used in this system is as follows:

[0547] 1. Smartphone application:

[0548] Speech recognition engine: Converts voice data into text data.

[0549] Speech synthesis engine: Converts translated text data into speech data.

[0550] Feedback feature: Receive feedback from users.

[0551] 2. AI translation engine on the server:

[0552] Uses natural language processing (NLP) technology to perform highly accurate translations.

[0553] System Operation

[0554] The user puts on the earphones and launches a translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. When the user begins to speak, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time.

[0555] The voice recognition engine of the smartphone terminal converts the received voice data into text data. For example, if a user says "Hello," the voice recognition engine generates the text data "Hello." The generated text data is then sent to the server.

[0556] The server's AI translation engine receives the text data and translates it into the specified target language. The server then uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The translation result is then sent back from the server to the smartphone.

[0557] The speech synthesis engine on the smartphone, which is the device that receives the translation results, converts the translated text data into voice data. This voice data is sent to the earphones. The earphone speakers pronounce the translated voice. User B can hear the voice saying "Hello" in real time.

[0558] After the conversation is over, the user provides feedback on the quality of the translation through the application, which then sends the feedback data to the server, where it is analyzed by the server's feedback processing system and used as training data for the AI ​​translation engine.

[0559] Specific examples

[0560] The following example shows how the system works when an English-speaking user A and a Japanese-speaking user B communicate in real time:

[0561] 1. User A says "Hello" in English.

[0562] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0563] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0564] 4. The smartphone sends the text data to the server.

[0565] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0566] 6. The server sends the translation results back to the smartphone.

[0567] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[0568] 8. The earphone speaker will say "hello."

[0569] 9. User B hears the translated "hello" in real time.

[0570] Prompt Sentence Examples

[0571] "Please explain how a system works, allowing users to communicate in real time with others who speak different languages. Please provide detailed examples of how data is processed using specific hardware (earphones, smartphone, server) and software (speech recognition engine, translation engine, speech synthesis engine)."

[0572] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0573] Step 1: Connection and Setup

[0574] The user puts on the dedicated earphones and launches the translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. The input for this step is the user's language setting and earphone connection information. The output is the completed setting state.

[0575] Step 2: Capture audio

[0576] When a user starts talking, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time. The input of this step is the voice from the user, and the output is the voice data sent to the smartphone.

[0577] Step 3: Speech to text

[0578] The device (smartphone) converts the received voice data into text data using a voice recognition engine. For example, the voice "Hello" is converted into the text data "Hello." The input of this step is the captured voice data, and the output is the generated text data.

[0579] Step 4: Translate the text

[0580] The server's AI translation engine receives the text data and translates it into the specified target language. The server uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The input of this step is the text data, and the output is the translated text data.

[0581] Step 5: Speech synthesis of the translation result

[0582] The device (smartphone) converts the translated text data returned from the server into voice data using a speech synthesis engine. The text "Hello" is converted into voice data saying "Hello." The input of this step is the translated text data, and the output is the generated voice data.

[0583] Step 6: Output the translated audio

[0584] The earphone speaker provides the generated voice data to the user. User B can hear the voice saying "Hello" in real time. The input of this step is the synthesized voice data, and the output is the voice the user hears.

[0585] Step 7: Provide feedback

[0586] After the conversation is completed, the user provides feedback on the quality of the translation through the smartphone application, for example, by entering feedback such as "very satisfied." The input of this step is the feedback data from the user, and the output is the feedback data sent to the server.

[0587] Step 8: Analyze feedback and learn

[0588] The server analyzes the received feedback data and stores it as training data for the AI ​​translation engine, thereby continuously improving translation performance. The input of this step is the received feedback data, and the output is updated training data.

[0589] (Application example 1)

[0590] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0591] The goal of this project is to achieve efficient real-time communication between staff members who speak different languages ​​at a logistics center where multinational staff work. This will enable improved efficiency and accuracy of work. However, with the previous system, there were barriers to communication between staff members, which hindered the smooth execution of work.

[0592] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0593] In this invention, the server includes an input means for inputting speech, a speech recognition means, a translation means, a speech synthesis means, an output means, a learning means, and a means used to enable efficient communication among multinational staff at a logistics center, thereby enabling real-time speech translation between multiple languages ​​and communication based on the translation.

[0594] The "input means for inputting voice" is a device that captures the voice uttered by the user as digital data.

[0595] "Speech recognition means" is a technology for converting captured voice data into text data.

[0596] "Translation means" refers to a technique for converting text data generated by speech recognition means into a target language.

[0597] "Speech synthesis means" is a technology that converts text data translated into a target language into speech data.

[0598] The "output means for outputting voice data" is a device that allows the user to listen to the translated voice data.

[0599] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" refers to a database and processing system for receiving evaluations from users and improving translation accuracy based on them.

[0600] "Means used to enable multinational staff to communicate efficiently at logistics centers" refers to components of a system that enables multiple staff members who speak different languages ​​to communicate smoothly through a voice translation system.

[0601] System configuration

[0602] This invention is a system for realizing real-time communication among multinational staff in a logistics center. This system is composed of the following main elements:

[0603] 1. A microphone built into earphones worn by the user as an input means for inputting voice.

[0604] 2. A speech recognition engine built into a smartphone that converts voice data into text data as a means of speech recognition.

[0605] 3. As a translation tool, an AI translation engine on a server that translates text data into a specified target language.

[0606] 4. A smartphone speech synthesis engine that converts translated text data into speech data as a speech synthesis method.

[0607] 5. A speaker built into the earphone as an output means for outputting audio data.

[0608] 6. A feedback processing system on the server that receives feedback provided by users as a learning means and stores it as learning data for the translation means.

[0609] 7. A system configuration that combines the above elements as a means to enable multinational staff to communicate efficiently in a logistics center.

[0610] How it works

[0611] The user uses the system by putting on earphones and launching a dedicated translation application installed on their smartphone. The system operates as follows:

[0612] 1. Voice input:

[0613] When the user speaks, the microphone built into the earphones captures the sound, and the audio data is sent to the smartphone in real time.

[0614] 2. Speech Recognition:

[0615] The smartphone's voice recognition engine converts the received voice data into text data.

[0616] 3. Translation:

[0617] The text data is sent to a server, where an AI translation engine translates it into the target language.

[0618] 4. Speech synthesis:

[0619] The translated text data is then sent back to the smartphone, where it is converted into voice data by a voice synthesis engine.

[0620] 5. Audio output:

[0621] The converted audio data is output to the user through the earphone speaker.

[0622] 6. Feedback and learning:

[0623] After the conversation, the user provides feedback on the quality of the translation, which is sent to the server. The feedback processing system analyzes this and stores it as learning data for the AI ​​translation engine, contributing to future improvements in translation accuracy.

[0624] Hardware and software used

[0625] Hardware:

[0626] Earphones (with built-in microphone and speaker)

[0627] Smartphone

[0628] software:

[0629] SpeechRecognition (speech recognition library)

[0630] GoogleTrans (translation library)

[0631] gTTS (Text-to-Speech Library)

[0632] Playsound (sound playback library)

[0633] Specific examples

[0634] Logistics centers often employ foreign staff who do not understand Japanese. With this system, instructions such as "Please pick up the items from aisle 3" can be given in Japanese and instantly translated into the staff's native language, ensuring smooth communication.

[0635] Prompt Sentence Examples

[0636] "Write Python code to translate English audio into Japanese and play it back as audio."

[0637] This will enable smooth communication between multinational staff at the logistics center, significantly improving work efficiency.

[0638] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0639] Step 1:

[0640] The user puts on the earphones and launches the translation application installed on their smartphone. When the application is launched, the system automatically connects to the earphones via Bluetooth and is ready for voice input. The user then sets the input language and target language in the app.

[0641] Input: User utterance

[0642] Output: Ready notification, setting language information

[0643] Step 2:

[0644] When a user speaks, the microphone built into the earphones captures the voice. This voice data is sent to a smartphone in real time via Bluetooth. The user's voice is the input data.

[0645] Input: User's voice data

[0646] Output: Microphone captured audio data

[0647] Step 3:

[0648] The smartphone's voice recognition engine converts the received voice data into text data. Specifically, the SpeechRecognition library processes the voice data and generates text data.

[0649] Input: Audio data

[0650] Output: Converted text data

[0651] Step 4:

[0652] The text data is sent to the server via the internet. The server's AI translation engine uses the Google Translate API to translate the text into the target language. The translated text data is output.

[0653] Input: Text data, target language information

[0654] Output: Translated text data

[0655] Step 5:

[0656] The translated text data is returned to the smartphone and converted into audio data by the speech synthesis engine. Specifically, the gTTS library converts the text data into an audio file.

[0657] Input: Translated text data

[0658] Output: Synthesized speech data

[0659] Step 6:

[0660] The generated voice data is sent from the smartphone via Bluetooth to the earphones, which then output the audio through their speakers, allowing the user to hear the audio in the target language in real time.

[0661] Input: Audio data

[0662] Output: Audio output from earphones

[0663] Step 7:

[0664] After the conversation is over, the user provides feedback on the quality of the translation via the app. The feedback data is sent to the server and stored in the feedback processing system. This data is used as training data for the translation engine to improve future translation accuracy.

[0665] Input: User feedback data

[0666] Output: Improved learning model with saved feedback data

[0667] Through these steps, efficient real-time communication between multinational staff at a logistics center is achieved.

[0668] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0669] MODE FOR CARRYING OUT THE INVENTION

[0670] This invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide translation results that correspond to those emotions. The detailed configuration and operation of the system are described below.

[0671] System Configuration

[0672] The system consists of the following main components:

[0673] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[0674] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[0675] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[0676] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[0677] 5. Output means: A device for outputting the generated audio data in a form that can be heard by the user. Specifically, it is a speaker built into an earphone.

[0678] 6. Learning means: A device or technology for receiving feedback from users and storing it as learning data for the translation means. This includes a feedback processing system on a server.

[0679] 7. Emotion Engine: A device or technology that recognizes the user's emotions and reflects them in the translation process. Emotions are recognized by analyzing voice tone, speech rate, and language patterns.

[0680] System Operation

[0681] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[0682] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[0683] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to the emotion engine and the translation engine on the server.

[0684] The emotion engine analyzes voice tone, speech rate, and language patterns to recognize the user's emotion, and this emotion data is sent to the translation engine.

[0685] The server's AI translation engine receives the text data and translates it into the specified target language, taking into account the recognized emotional data. During the translation process, natural language processing technology is used to understand the context and emotions and select the appropriate translation.

[0686] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0687] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[0688] Specific examples

[0689] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[0690] 1. User A says "Hello" in English. His voice sounds cheerful.

[0691] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0692] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0693] 4. The smartphone sends the text data to the emotion engine and server.

[0694] 5. The emotion engine analyzes User A's tone and speech rate and recognizes the emotion "fun." This emotion data is sent to the server.

[0695] 6. The server's translation means translates "Hello" to the Japanese "Konnichiwa" and maintains a happy tone of voice.

[0696] 7. The server sends the translation results back to the smartphone.

[0697] 8. The smartphone's voice synthesis function converts "hello" into voice data in a pleasant voice tone.

[0698] 9. The earphone speaker pronounces "hello" in a cheerful voice.

[0699] 10. User B hears the translated "Hello" in real time in a happy voice.

[0700] This system allows users A and B to communicate smoothly and emotionally, overcoming language barriers.

[0701] The processing flow will be explained below.

[0702] Step 1:

[0703] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[0704] Step 2:

[0705] The user sets the input language and target language in the translation app, and also enables the emotion recognition function.

[0706] Step 3:

[0707] When the user starts speaking, the microphone built into the earphones captures the voice, and this voice data is sent to the smartphone in real time.

[0708] Step 4:

[0709] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[0710] Step 5:

[0711] The terminal sends the converted text data to the emotion engine, which analyzes the voice tone, speech rate, and language patterns to recognize the user's emotion, and the recognized emotion data is sent to the server.

[0712] Step 6:

[0713] The terminal transmits the text data to a server via the Internet.

[0714] Step 7:

[0715] The server receives the text data and emotional data, and the AI ​​translation engine translates it into the specified target language. The translation engine also takes into account the emotional data, understands the context, and selects the appropriate translation and tone of voice.

[0716] Step 8:

[0717] The server returns the translation results that reflect the emotional data to the terminal.

[0718] Step 9:

[0719] The device passes the translated text data received from the server to a speech synthesis engine and converts it into voice data, while maintaining the voice tone that reflects the emotional data.

[0720] Step 10:

[0721] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[0722] Step 11:

[0723] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[0724] Step 12:

[0725] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[0726] Example 2

[0727] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0728] There is a demand for systems that allow people who speak different languages ​​to communicate smoothly in real time. However, even if traditional translation systems provide accurate translations, they tend to fail to reflect emotional expressions, resulting in incomplete communication. Another issue is the lack of a mechanism for improving the system's accuracy through feedback.

[0729] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, an emotion analysis means for analyzing speech tone, speech rate, and language patterns to recognize a user's emotion, a translation means for translating the text data and emotion data generated by the speech recognition means and the emotion analysis means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by a user and saving the feedback as learning data for the translation means. This enables smooth real-time communication that incorporates emotions between users who speak different languages.

[0730] "Input means" refers to a device for acquiring audio, specifically a microphone built into an earphone.

[0731] The "voice recognition means" is a device or technology for converting voice data acquired by the input means into text data.

[0732] An "emotion analysis means" is a device or technology for recognizing a user's emotions by analyzing voice tone, speech rate, and language patterns.

[0733] The "translation means" is a device or technology for translating the text data and emotion data generated by the speech recognition means and emotion analysis means into a specified target language.

[0734] The "speech synthesis means" refers to a device or technology for converting the translated text data generated by the translation means into speech data.

[0735] The "output means" is a device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically refers to a speaker built into an earphone.

[0736] "Learning means" refers to a device or technology that receives feedback on translation quality provided by a user and stores it as training data for the translation means.

[0737] The present invention is a system that enables people who speak different languages ​​to communicate smoothly in real time by using advanced technologies that integrate speech recognition, emotion analysis, translation, speech synthesis, and feedback processing.

[0738] System configuration

[0739] The system consists of the following main components:

[0740] 1. Input means: A device for capturing audio, specifically a microphone built into earphones.

[0741] 2. Speech recognition means: A device for converting voice data acquired by the input means into text data, using a voice recognition engine built into the smartphone.

[0742] 3. Emotion analysis means: A device for analyzing voice tone, speech rate, and language patterns to recognize the user's emotions.

[0743] 4. Translation means: A device for translating text data and emotional data generated by the speech recognition means and emotional analysis means into a specified target language. It uses an AI translation engine on the server.

[0744] 5. Speech synthesis means: A device for converting the translated text data generated by the translation means into speech data, and uses a speech synthesis engine built into the smartphone.

[0745] 6. Output means: A device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically, a speaker built into the earphone is used.

[0746] 7. Learning means: A device that receives feedback on translation quality provided by users, stores it as learning data for the translation means, and improves the system's performance. A feedback processing system on the server is used.

[0747] System Operation

[0748] The user puts on the earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. The user sets the input language and target language to be used in the application. This setting information is saved on the smartphone.

[0749] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data in real time to the smartphone. The smartphone then uses a voice recognition engine to convert the voice data into text data. The conversion result is then sent to the emotion analysis means and the server.

[0750] The device's emotion analysis means analyzes voice tone, speech rate, and language patterns to recognize the user's emotions. This emotion data is sent to the translation means. The server's AI translation engine receives the text data and emotion data and translates it into the specified target language. The AI ​​translation engine understands the context and emotion and selects the appropriate translation.

[0751] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0752] After the conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server and analyzed by the server's learning means and used to improve the performance of the translation means.

[0753] Specific operation example

[0754] For example, when an English-speaking user A and a Japanese-speaking user B communicate in real time, the system operates as follows:

[0755] 1. User A launches the translation app and puts on earphones.

[0756] 2. User A sets English as the input language and Japanese as the target language on the app.

[0757] 3. User A says "Hello" happily.

[0758] 4. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0759] 5. The device uses a voice recognition means to convert the voice into text data such as "Hello."

[0760] 6. This data is sent to the server and sentiment analysis means.

[0761] 7. The device's emotion analysis means recognizes the emotion "fun" and sends this information to the server.

[0762] 8. The server's AI translation engine translates "Hello" into "Hello" taking into account the "joyful" emotion.

[0763] 9. The server sends the translation results back to the smartphone.

[0764] 10. The device's voice synthesis means converts "hello" into voice data in a "happy" voice tone.

[0765] 11. The earphone speaker says "Hello" to User B in a cheerful voice.

[0766] 12. User B hears "Hello" in real time and continues the conversation.

[0767] 13. After the conversation, User A provides feedback on the quality of the translation through the app, which is then sent to the server.

[0768] 14. The server's learning mechanism uses this feedback to improve the performance of the translation engine.

[0769] Prompt Sentence Examples

[0770] When User A says "Hello" in English, the earphone microphone captures the audio and sends it to the smartphone. The smartphone's speech recognition engine converts "Hello" into text, and the emotion analysis means recognizes the emotion as "happy." This information is sent to the translation engine on the server, and the audio converted to "Hello" in a happy voice tone is played from the earphone speaker.

[0771] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0772] Step 1:

[0773] The user puts on the system's earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. A connection confirmation message is displayed on the smartphone, and the user sets the input language and target language. This setting information is saved on the smartphone.

[0774] Specific behavior:

[0775] Input: User puts in earphones and launches translation app.

[0776] Data processing / calculation: The app checks the earphone connection and receives language setting information.

[0777] Output: A connection completion message and a setting confirmation message are displayed.

[0778] Step 2:

[0779] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data to the smartphone in real time.

[0780] Specific behavior:

[0781] Input: User's voice.

[0782] Data processing / calculation: The microphone captures the sound and converts it into digital sound data.

[0783] Output: Digital audio data is sent to the smartphone.

[0784] Step 3:

[0785] The terminal (smartphone) uses a voice recognition means to convert the voice data into text data, and the conversion result is sent to the emotion analysis means and the server.

[0786] Specific behavior:

[0787] Input: Digital audio data.

[0788] Data processing / calculation: The voice recognition engine analyzes the voice data and converts it into text data.

[0789] Output: The text data and the voice data are sent to the emotion analysis means and the server.

[0790] Step 4:

[0791] The emotion analysis means of the terminal analyzes the voice tone, speech rate, and language pattern to recognize the user's emotion, and this emotion data is sent to the translation means.

[0792] Specific behavior:

[0793] Input: Audio data.

[0794] Data processing / calculation: Emotion analysis means analyzes voice tone, speaking rate, and language patterns to extract emotions.

[0795] Output: The recognized emotion data is sent to the translation means.

[0796] Step 5:

[0797] The server's translation means receives the text data and emotion data and translates it into the target language. The AI ​​translation engine takes context and emotion into account when translating.

[0798] Specific behavior:

[0799] Input: Text data and emotion data.

[0800] Data processing / calculation: The translation engine analyzes the context and sentiment to generate an appropriate translation.

[0801] Output: The translated text data is generated and sent back to the smartphone.

[0802] Step 6:

[0803] The terminal receives the translation result returned from the server and converts the translated text data into voice data using a voice synthesis means.

[0804] Specific behavior:

[0805] Input: Translation text data.

[0806] Data processing / calculation: The speech synthesis engine converts text data into speech data.

[0807] Output: Audio data is generated.

[0808] Step 7:

[0809] The generated audio data is transmitted from the terminal to earphones, allowing the user to listen to it in real time.

[0810] Specific behavior:

[0811] Input: Audio data.

[0812] Data processing / calculation: Transfer data to earphones.

[0813] Output: Audio data is played from the earphone speaker.

[0814] Step 8:

[0815] After completing a conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server where it is analyzed by the server's learning means and used to improve the performance of the translation means.

[0816] Specific behavior:

[0817] Input: User feedback.

[0818] Data processing / calculation: Feedback data is sent to the server and analyzed.

[0819] Output: The analysis results are added to the translation engine's training data, improving translation performance.

[0820] (Application example 2)

[0821] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0822] There is a need for a method that allows people who speak different languages ​​to communicate smoothly in real time, including expressing their emotions. In particular, in the food delivery scene, a system is needed that allows delivery personnel and customers to communicate without worrying about language differences.

[0823] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for inputting voice, voice recognition means for converting voice data input by the input means into text data, translation means for translating the text data generated by the voice recognition means into a specified target language, voice synthesis means for converting the translated text data generated by the translation means into voice data, output means for outputting the voice data generated by the voice synthesis means, learning means for receiving feedback on translation quality provided by the user and saving it as learning data for the translation means, and an emotion engine that recognizes emotions by analyzing voice tone, speaking rate, and language patterns. This enables real-time communication incorporating emotions even when the delivery person and the customer speak different languages.

[0824] The "input means for inputting voice" is a device for capturing the user's speech.

[0825] "Speech recognition means for converting voice data input by an input means into text data" refers to a device or technology for converting voice input into data in text format.

[0826] "Translation means for translating text data generated by speech recognition means into a specified target language" refers to a device or technology for converting the generated text data into the target language.

[0827] "Speech synthesis means for converting translated text data generated by the translation means into speech data" refers to a device or technology for converting translated text into speech-format data.

[0828] The "output means for outputting the voice data generated by the voice synthesis means" is a device that enables the user to listen to the voice data.

[0829] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" is a technology for collecting user feedback and improving translation accuracy based on that data.

[0830] The "emotion engine that recognizes emotions by analyzing voice tone, speech rate, and language patterns" is a technology that analyzes the tone and rate of speech, the use of words, etc. to understand the speaker's emotions.

[0831] This invention is a system for realizing emotional real-time communication between delivery personnel and customers even when they speak different languages. The system includes the following elements:

[0832] 1. Input means: An input device for capturing the user's voice. Specifically, a microphone built into a smartphone or a microphone built into earphones can be used.

[0833] 2. Speech recognition means: Software for converting acquired voice data into text data. This is achieved by using the Python speech_recognition library.

[0834] 3. Translation tool: Software for translating the text data generated by the speech recognition tool into the target language. Translation is performed using Google's googletrans library.

[0835] 4. Speech synthesis means: Software for converting translated text data into speech data. This is achieved using the gTTS (Google Text-to-Speech) library.

[0836] 5. Output means: A device for transmitting the generated voice data to the user. Specifically, the speaker of a smartphone or a speaker of an earphone can be used.

[0837] 6. Learning method: A technology for collecting feedback data from users and using it to improve translation performance. The data is sent to the server for storage and analysis.

[0838] 7. Emotion Engine: Technology that analyzes voice tone, speaking rate, and language patterns to recognize user emotions, resulting in more natural and emotionally relevant translation results.

[0839] Hardware and Software Used

[0840] 1. Hardware:

[0841] Smartphone (microphone, speaker)

[0842] Earphones (microphone, speaker)

[0843] 2. Software:

[0844] Python

[0845] speech_recognition library (speech recognition)

[0846] googletrans library (translation)

[0847] gTTS library (speech synthesis)

[0848] Explanation of data processing and data calculation

[0849] The server receives voice data input via a smartphone or earphones and converts it into text data using a voice recognition means. This text data is then converted into the target language using a translation means. The translated text is converted back into voice data using a voice synthesis means and transmitted to the user via an output means. In addition, the learning means analyzes user feedback and uses it to improve system performance. During this series of processes, the emotion engine analyzes the user's emotions and reflects them in the translation results.

[0850] Specific examples

[0851] For example, consider the case where a delivery person says, "This is a delivery item." When the delivery person says, "I've brought your order" in English, this voice is captured by the smartphone's microphone. The speech recognition means converts this into text data, obtaining the string "I've brought your order." The translation means then translates this into Japanese, generating the text data "This is a delivery item." The speech synthesis means converts this Japanese text into speech data, which is finally conveyed to the customer via the smartphone or earphone speaker.

[0852] Example prompts to input to the generative AI model

[0853] Create a real-time translation assistant to help people who speak different languages ​​communicate smoothly. When a delivery person says "Delivery here" to a customer, it needs to translate from English to Japanese. This translation assistant will be realized using speech recognition, text translation, and speech synthesis.

[0854] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0855] Step 1:

[0856] Input and Output

[0857] A user starts the smartphone app and speaks into the microphone. The input is the user's voice. The output is recorded as audio data on the smartphone.

[0858] Specific actions

[0859] The user launches the smartphone app and says, "I've brought your order" in English. The smartphone's microphone captures the voice and generates voice data.

[0860] Step 2:

[0861] Input and Output

[0862] The voice data is sent to the voice recognition means. At this time, the input is the voice data generated in step 1. The output is the text data "I've brought your order" generated by the voice recognition means.

[0863] Specific actions

[0864] The smartphone's voice recognition engine (speech_recognition library) analyzes the voice data and converts it into corresponding text data.

[0865] Step 3:

[0866] Input and Output

[0867] The text data is sent to the translation means. At this time, the input is the text data "I've brought your order" generated in step 2. The output is the text data "Delivery item" in the target language generated by the translation means.

[0868] Specific actions

[0869] The smartphone's Google Trans library translates the text data into Japanese.

[0870] Step 4:

[0871] Input and Output

[0872] The text data in the target language is sent to the speech synthesis means. At this time, the input is the text data "Delivery item" generated in step 3. The output is speech data generated by the speech synthesis means.

[0873] Specific actions

[0874] The smartphone's gTTS library converts the text data into audio data and generates an audio file.

[0875] Step 5:

[0876] Input and Output

[0877] The audio data is sent to an output means, where the input is the audio data generated in step 4. The output is audio that can be heard by the user.

[0878] Specific actions

[0879] The translated voice message "Delivery here" is played to the customer through the speaker on their smartphone or earphones.

[0880] Step 6:

[0881] Input and Output

[0882] Feedback from the user is sent to the learning means. At this time, the input is the feedback on translation quality provided by the user. The output is the feedback data stored on the server and becomes learning data.

[0883] Specific actions

[0884] Users provide feedback on the quality of the translation through the app, which is then sent to the server, which analyzes the data and stores it in the translation engine's training database.

[0885] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0886] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0887] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0888] [Third embodiment]

[0889] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0890] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0891] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0892] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0893] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0894] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0895] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0896] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0897] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0898] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0899] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0900] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0901] MODE FOR CARRYING OUT THE INVENTION

[0902] The present invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. The detailed configuration and operation of the system will be described below.

[0903] System Configuration

[0904] The system consists of the following main components:

[0905] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[0906] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[0907] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[0908] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[0909] 5. Output means: A device that allows the user to listen to the generated audio data. Specifically, it is a speaker built into an earphone.

[0910] 6. Learning tool: A device or technology that receives feedback from users and trains the translation tool. This includes a feedback processing system on the server.

[0911] System Operation

[0912] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[0913] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[0914] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to a translation engine on the server.

[0915] The server's AI translation engine receives the text data and translates it into the specified target language. During the translation process, it uses natural language processing technology to understand the context and select the appropriate translation.

[0916] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[0917] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[0918] Specific examples

[0919] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[0920] 1. User A says "Hello" in English.

[0921] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0922] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0923] 4. The smartphone sends the text data to the server.

[0924] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0925] 6. The server sends the translation results back to the smartphone.

[0926] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[0927] 8. The earphone speaker will say "hello."

[0928] 9. User B hears the translated "hello" in real time.

[0929] This system allows users A and B to communicate smoothly across language barriers.

[0930] The processing flow will be explained below.

[0931] Step 1:

[0932] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[0933] Step 2:

[0934] The user sets the input language and target language in the translation app, and also adjusts the volume and enables noise cancellation as needed.

[0935] Step 3:

[0936] When the user starts speaking, the microphone built into the earphones captures the user's voice, and the captured voice data is sent to the smartphone in real time.

[0937] Step 4:

[0938] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[0939] Step 5:

[0940] The terminal transmits the converted text data to a server via the Internet.

[0941] Step 6:

[0942] The server receives the text data and translates it into the specified target language using an AI translation engine. The translation engine uses natural language processing to understand the context and select the appropriate translation.

[0943] Step 7:

[0944] The server returns the translated text data to the device, and also prepares to receive feedback as learning data to improve translation accuracy.

[0945] Step 8:

[0946] The terminal passes the translated text data received from the server to a speech synthesis engine and converts it into voice data.

[0947] Step 9:

[0948] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[0949] Step 10:

[0950] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[0951] Step 11:

[0952] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[0953] Example 1

[0954] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0955] It is difficult for people who speak different languages ​​to communicate smoothly in real time. Furthermore, the quality of translation must be improved. Therefore, there is a need for a system that can realize real-time translation in multiple languages ​​and continuously improve its translation performance by receiving feedback.

[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0957] In this invention, the server includes a connection means for automatically connecting to a user's device, an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, a translation means for translating the text data generated by the speech recognition means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by the user and saving the feedback as learning data for the translation means. This enables people speaking different languages ​​to communicate smoothly in real time. Furthermore, by utilizing user feedback, translation performance can be continuously improved.

[0958] "User" refers to an individual who uses the system to communicate with others who speak different languages.

[0959] "Device" refers to a device that includes hardware for use by a user, such as earphones or smartphones.

[0960] "Connection means" refers to a method or technology for connecting a user's device to each component of the system. Specifically, this applies to wireless communication technologies such as Bluetooth and Wi-Fi.

[0961] "Input means" refers to a device or technology for inputting sound into the system, specifically a microphone built into earphones.

[0962] "Voice data" refers to digital audio information that records a user's speech.

[0963] "Text data" refers to linguistic information in character format converted by a speech recognition means.

[0964] "Speech recognition means" refers to a technique or device for converting voice data into text data.

[0965] "Translation means" refers to technology or equipment for converting text data into a specified target language. Specifically, it includes an AI translation engine installed on a server.

[0966] "Speech synthesis means" refers to technology or equipment for converting translated text data into speech data.

[0967] "Output means" refers to a device or technology for transmitting synthesized voice data to the user. Specifically, this applies to the speaker built into the earphone.

[0968] "Feedback" refers to evaluation information about the quality of a translation provided by a user.

[0969] "Training means" refers to a technique or device for analyzing received feedback and improving the performance of the translation means.

[0970] MODE FOR CARRYING OUT THE INVENTION

[0971] The present invention relates to a system that enables people who speak different languages ​​to communicate smoothly in real time. The detailed configuration and operation of the system will be described below.

[0972] System Hardware

[0973] The system hardware consists of the following major components:

[0974] 1. Earphones: Earphones with built-in microphones and speakers that are responsible for inputting and outputting audio data.

[0975] 2. Smartphone: A translation application is installed and performs voice recognition, data transmission and reception, and voice synthesis.

[0976] 3. Server: Equipped with an AI translation engine, it translates text data.

[0977] Software used

[0978] The software used in this system is as follows:

[0979] 1. Smartphone application:

[0980] Speech recognition engine: Converts voice data into text data.

[0981] Speech synthesis engine: Converts translated text data into speech data.

[0982] Feedback feature: Receive feedback from users.

[0983] 2. AI translation engine on the server:

[0984] Uses natural language processing (NLP) technology to perform highly accurate translations.

[0985] System Operation

[0986] The user puts on the earphones and launches a translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. When the user begins to speak, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time.

[0987] The voice recognition engine of the smartphone terminal converts the received voice data into text data. For example, if a user says "Hello," the voice recognition engine generates the text data "Hello." The generated text data is then sent to the server.

[0988] The server's AI translation engine receives the text data and translates it into the specified target language. The server then uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The translation result is then sent back from the server to the smartphone.

[0989] The speech synthesis engine on the smartphone, which is the device that receives the translation results, converts the translated text data into voice data. This voice data is sent to the earphones. The earphone speakers pronounce the translated voice. User B can hear the voice saying "Hello" in real time.

[0990] After the conversation is over, the user provides feedback on the quality of the translation through the application, which then sends the feedback data to the server, where it is analyzed by the server's feedback processing system and used as training data for the AI ​​translation engine.

[0991] Specific examples

[0992] The following example shows how the system works when an English-speaking user A and a Japanese-speaking user B communicate in real time:

[0993] 1. User A says "Hello" in English.

[0994] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[0995] 3. The smartphone's voice recognition function converts "Hello" into text data.

[0996] 4. The smartphone sends the text data to the server.

[0997] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[0998] 6. The server sends the translation results back to the smartphone.

[0999] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[1000] 8. The earphone speaker will say "hello."

[1001] 9. User B hears the translated "hello" in real time.

[1002] Prompt Sentence Examples

[1003] "Please explain how a system works, allowing users to communicate in real time with others who speak different languages. Please provide detailed examples of how data is processed using specific hardware (earphones, smartphone, server) and software (speech recognition engine, translation engine, speech synthesis engine)."

[1004] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1005] Step 1: Connection and Setup

[1006] The user puts on the dedicated earphones and launches the translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. The input for this step is the user's language setting and earphone connection information. The output is the completed setting state.

[1007] Step 2: Capture audio

[1008] When a user starts talking, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time. The input of this step is the voice from the user, and the output is the voice data sent to the smartphone.

[1009] Step 3: Speech to text

[1010] The device (smartphone) converts the received voice data into text data using a voice recognition engine. For example, the voice "Hello" is converted into the text data "Hello." The input of this step is the captured voice data, and the output is the generated text data.

[1011] Step 4: Translate the text

[1012] The server's AI translation engine receives the text data and translates it into the specified target language. The server uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The input of this step is the text data, and the output is the translated text data.

[1013] Step 5: Speech synthesis of the translation result

[1014] The device (smartphone) converts the translated text data returned from the server into voice data using a speech synthesis engine. The text "Hello" is converted into voice data saying "Hello." The input of this step is the translated text data, and the output is the generated voice data.

[1015] Step 6: Output the translated audio

[1016] The earphone speaker provides the generated voice data to the user. User B can hear the voice saying "Hello" in real time. The input of this step is the synthesized voice data, and the output is the voice the user hears.

[1017] Step 7: Provide feedback

[1018] After the conversation is completed, the user provides feedback on the quality of the translation through the smartphone application, for example, by entering feedback such as "very satisfied." The input of this step is the feedback data from the user, and the output is the feedback data sent to the server.

[1019] Step 8: Analyze feedback and learn

[1020] The server analyzes the received feedback data and stores it as training data for the AI ​​translation engine, thereby continuously improving translation performance. The input of this step is the received feedback data, and the output is updated training data.

[1021] (Application example 1)

[1022] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1023] The goal of this project is to achieve efficient real-time communication between staff members who speak different languages ​​at a logistics center where multinational staff work. This will enable improved efficiency and accuracy of work. However, with the previous system, there were barriers to communication between staff members, which hindered the smooth execution of work.

[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1025] In this invention, the server includes an input means for inputting speech, a speech recognition means, a translation means, a speech synthesis means, an output means, a learning means, and a means used to enable efficient communication among multinational staff at a logistics center, thereby enabling real-time speech translation between multiple languages ​​and communication based on the translation.

[1026] The "input means for inputting voice" is a device that captures the voice uttered by the user as digital data.

[1027] "Speech recognition means" is a technology for converting captured voice data into text data.

[1028] "Translation means" refers to a technique for converting text data generated by speech recognition means into a target language.

[1029] "Speech synthesis means" is a technology that converts text data translated into a target language into speech data.

[1030] The "output means for outputting voice data" is a device that allows the user to listen to the translated voice data.

[1031] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" refers to a database and processing system for receiving evaluations from users and improving translation accuracy based on them.

[1032] "Means used to enable multinational staff to communicate efficiently at logistics centers" refers to components of a system that enables multiple staff members who speak different languages ​​to communicate smoothly through a voice translation system.

[1033] System configuration

[1034] This invention is a system for realizing real-time communication among multinational staff in a logistics center. This system is composed of the following main elements:

[1035] 1. A microphone built into earphones worn by the user as an input means for inputting voice.

[1036] 2. A speech recognition engine built into a smartphone that converts voice data into text data as a means of speech recognition.

[1037] 3. As a translation tool, an AI translation engine on a server that translates text data into a specified target language.

[1038] 4. A smartphone speech synthesis engine that converts translated text data into speech data as a speech synthesis method.

[1039] 5. A speaker built into the earphone as an output means for outputting audio data.

[1040] 6. A feedback processing system on the server that receives feedback provided by users as a learning means and stores it as learning data for the translation means.

[1041] 7. A system configuration that combines the above elements as a means to enable multinational staff to communicate efficiently in a logistics center.

[1042] How it works

[1043] The user uses the system by putting on earphones and launching a dedicated translation application installed on their smartphone. The system operates as follows:

[1044] 1. Voice input:

[1045] When the user speaks, the microphone built into the earphones captures the sound, and the audio data is sent to the smartphone in real time.

[1046] 2. Speech Recognition:

[1047] The smartphone's voice recognition engine converts the received voice data into text data.

[1048] 3. Translation:

[1049] The text data is sent to a server, where an AI translation engine translates it into the target language.

[1050] 4. Speech synthesis:

[1051] The translated text data is then sent back to the smartphone, where it is converted into voice data by a voice synthesis engine.

[1052] 5. Audio output:

[1053] The converted audio data is output to the user through the earphone speaker.

[1054] 6. Feedback and learning:

[1055] After the conversation, the user provides feedback on the quality of the translation, which is sent to the server. The feedback processing system analyzes this and stores it as learning data for the AI ​​translation engine, contributing to future improvements in translation accuracy.

[1056] Hardware and software used

[1057] Hardware:

[1058] Earphones (with built-in microphone and speaker)

[1059] Smartphone

[1060] software:

[1061] SpeechRecognition (speech recognition library)

[1062] GoogleTrans (translation library)

[1063] gTTS (Text-to-Speech Library)

[1064] Playsound (sound playback library)

[1065] Specific examples

[1066] Logistics centers often employ foreign staff who do not understand Japanese. With this system, instructions such as "Please pick up the items from aisle 3" can be given in Japanese and instantly translated into the staff's native language, ensuring smooth communication.

[1067] Prompt Sentence Examples

[1068] "Write Python code to translate English audio into Japanese and play it back as audio."

[1069] This will enable smooth communication between multinational staff at the logistics center, significantly improving work efficiency.

[1070] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1071] Step 1:

[1072] The user puts on the earphones and launches the translation application installed on their smartphone. When the application is launched, the system automatically connects to the earphones via Bluetooth and is ready for voice input. The user then sets the input language and target language in the app.

[1073] Input: User utterance

[1074] Output: Ready notification, setting language information

[1075] Step 2:

[1076] When a user speaks, the microphone built into the earphones captures the voice. This voice data is sent to a smartphone in real time via Bluetooth. The user's voice is the input data.

[1077] Input: User's voice data

[1078] Output: Microphone captured audio data

[1079] Step 3:

[1080] The smartphone's voice recognition engine converts the received voice data into text data. Specifically, the SpeechRecognition library processes the voice data and generates text data.

[1081] Input: Audio data

[1082] Output: Converted text data

[1083] Step 4:

[1084] The text data is sent to the server via the internet. The server's AI translation engine uses the Google Translate API to translate the text into the target language. The translated text data is output.

[1085] Input: Text data, target language information

[1086] Output: Translated text data

[1087] Step 5:

[1088] The translated text data is returned to the smartphone and converted into audio data by the speech synthesis engine. Specifically, the gTTS library converts the text data into an audio file.

[1089] Input: Translated text data

[1090] Output: Synthesized speech data

[1091] Step 6:

[1092] The generated voice data is sent from the smartphone via Bluetooth to the earphones, which then output the audio through their speakers, allowing the user to hear the audio in the target language in real time.

[1093] Input: Audio data

[1094] Output: Audio output from earphones

[1095] Step 7:

[1096] After the conversation is over, the user provides feedback on the quality of the translation via the app. The feedback data is sent to the server and stored in the feedback processing system. This data is used as training data for the translation engine to improve future translation accuracy.

[1097] Input: User feedback data

[1098] Output: Improved learning model with saved feedback data

[1099] Through these steps, efficient real-time communication between multinational staff at a logistics center is achieved.

[1100] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1101] MODE FOR CARRYING OUT THE INVENTION

[1102] This invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide translation results that correspond to those emotions. The detailed configuration and operation of the system are described below.

[1103] System Configuration

[1104] The system consists of the following main components:

[1105] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[1106] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[1107] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[1108] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[1109] 5. Output means: A device for outputting the generated audio data in a form that can be heard by the user. Specifically, it is a speaker built into an earphone.

[1110] 6. Learning means: A device or technology for receiving feedback from users and storing it as learning data for the translation means. This includes a feedback processing system on a server.

[1111] 7. Emotion Engine: A device or technology that recognizes the user's emotions and reflects them in the translation process. Emotions are recognized by analyzing voice tone, speech rate, and language patterns.

[1112] System Operation

[1113] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[1114] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[1115] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to the emotion engine and the translation engine on the server.

[1116] The emotion engine analyzes voice tone, speech rate, and language patterns to recognize the user's emotion, and this emotion data is sent to the translation engine.

[1117] The server's AI translation engine receives the text data and translates it into the specified target language, taking into account the recognized emotional data. During the translation process, natural language processing technology is used to understand the context and emotions and select the appropriate translation.

[1118] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[1119] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[1120] Specific examples

[1121] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[1122] 1. User A says "Hello" in English. His voice sounds cheerful.

[1123] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1124] 3. The smartphone's voice recognition function converts "Hello" into text data.

[1125] 4. The smartphone sends the text data to the emotion engine and server.

[1126] 5. The emotion engine analyzes User A's tone and speech rate and recognizes the emotion "fun." This emotion data is sent to the server.

[1127] 6. The server's translation means translates "Hello" to the Japanese "Konnichiwa" and maintains a happy tone of voice.

[1128] 7. The server sends the translation results back to the smartphone.

[1129] 8. The smartphone's voice synthesis function converts "hello" into voice data in a pleasant voice tone.

[1130] 9. The earphone speaker pronounces "hello" in a cheerful voice.

[1131] 10. User B hears the translated "Hello" in real time in a happy voice.

[1132] This system allows users A and B to communicate smoothly and emotionally, overcoming language barriers.

[1133] The processing flow will be explained below.

[1134] Step 1:

[1135] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[1136] Step 2:

[1137] The user sets the input language and target language in the translation app, and also enables the emotion recognition function.

[1138] Step 3:

[1139] When the user starts speaking, the microphone built into the earphones captures the voice, and this voice data is sent to the smartphone in real time.

[1140] Step 4:

[1141] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[1142] Step 5:

[1143] The terminal sends the converted text data to the emotion engine, which analyzes the voice tone, speech rate, and language patterns to recognize the user's emotion, and the recognized emotion data is sent to the server.

[1144] Step 6:

[1145] The terminal transmits the text data to a server via the Internet.

[1146] Step 7:

[1147] The server receives the text data and emotional data, and the AI ​​translation engine translates it into the specified target language. The translation engine also takes into account the emotional data, understands the context, and selects the appropriate translation and tone of voice.

[1148] Step 8:

[1149] The server returns the translation results that reflect the emotional data to the terminal.

[1150] Step 9:

[1151] The device passes the translated text data received from the server to a speech synthesis engine and converts it into voice data, while maintaining the voice tone that reflects the emotional data.

[1152] Step 10:

[1153] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[1154] Step 11:

[1155] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[1156] Step 12:

[1157] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[1158] Example 2

[1159] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1160] There is a demand for systems that allow people who speak different languages ​​to communicate smoothly in real time. However, even if traditional translation systems provide accurate translations, they tend to fail to reflect emotional expressions, resulting in incomplete communication. Another issue is the lack of a mechanism for improving the system's accuracy through feedback.

[1161] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, an emotion analysis means for analyzing speech tone, speech rate, and language patterns to recognize a user's emotion, a translation means for translating the text data and emotion data generated by the speech recognition means and the emotion analysis means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by a user and saving the feedback as learning data for the translation means. This enables smooth real-time communication that incorporates emotions between users who speak different languages.

[1162] "Input means" refers to a device for acquiring audio, specifically a microphone built into an earphone.

[1163] The "voice recognition means" is a device or technology for converting voice data acquired by the input means into text data.

[1164] An "emotion analysis means" is a device or technology for recognizing a user's emotions by analyzing voice tone, speech rate, and language patterns.

[1165] The "translation means" is a device or technology for translating the text data and emotion data generated by the speech recognition means and emotion analysis means into a specified target language.

[1166] The "speech synthesis means" refers to a device or technology for converting the translated text data generated by the translation means into speech data.

[1167] The "output means" is a device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically refers to a speaker built into an earphone.

[1168] "Learning means" refers to a device or technology that receives feedback on translation quality provided by a user and stores it as training data for the translation means.

[1169] The present invention is a system that enables people who speak different languages ​​to communicate smoothly in real time by using advanced technologies that integrate speech recognition, emotion analysis, translation, speech synthesis, and feedback processing.

[1170] System configuration

[1171] The system consists of the following main components:

[1172] 1. Input means: A device for capturing audio, specifically a microphone built into earphones.

[1173] 2. Speech recognition means: A device for converting voice data acquired by the input means into text data, using a voice recognition engine built into the smartphone.

[1174] 3. Emotion analysis means: A device for analyzing voice tone, speech rate, and language patterns to recognize the user's emotions.

[1175] 4. Translation means: A device for translating text data and emotional data generated by the speech recognition means and emotional analysis means into a specified target language. It uses an AI translation engine on the server.

[1176] 5. Speech synthesis means: A device for converting the translated text data generated by the translation means into speech data, and uses a speech synthesis engine built into the smartphone.

[1177] 6. Output means: A device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically, a speaker built into the earphone is used.

[1178] 7. Learning means: A device that receives feedback on translation quality provided by users, stores it as learning data for the translation means, and improves the system's performance. A feedback processing system on the server is used.

[1179] System Operation

[1180] The user puts on the earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. The user sets the input language and target language to be used in the application. This setting information is saved on the smartphone.

[1181] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data in real time to the smartphone. The smartphone then uses a voice recognition engine to convert the voice data into text data. The conversion result is then sent to the emotion analysis means and the server.

[1182] The device's emotion analysis means analyzes voice tone, speech rate, and language patterns to recognize the user's emotions. This emotion data is sent to the translation means. The server's AI translation engine receives the text data and emotion data and translates it into the specified target language. The AI ​​translation engine understands the context and emotion and selects the appropriate translation.

[1183] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[1184] After the conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server and analyzed by the server's learning means and used to improve the performance of the translation means.

[1185] Specific operation example

[1186] For example, when an English-speaking user A and a Japanese-speaking user B communicate in real time, the system operates as follows:

[1187] 1. User A launches the translation app and puts on earphones.

[1188] 2. User A sets English as the input language and Japanese as the target language on the app.

[1189] 3. User A says "Hello" happily.

[1190] 4. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1191] 5. The device uses a voice recognition means to convert the voice into text data such as "Hello."

[1192] 6. This data is sent to the server and sentiment analysis means.

[1193] 7. The device's emotion analysis means recognizes the emotion "fun" and sends this information to the server.

[1194] 8. The server's AI translation engine translates "Hello" into "Hello" taking into account the "joyful" emotion.

[1195] 9. The server sends the translation results back to the smartphone.

[1196] 10. The device's voice synthesis means converts "hello" into voice data in a "happy" voice tone.

[1197] 11. The earphone speaker says "Hello" to User B in a cheerful voice.

[1198] 12. User B hears "Hello" in real time and continues the conversation.

[1199] 13. After the conversation, User A provides feedback on the quality of the translation through the app, which is then sent to the server.

[1200] 14. The server's learning mechanism uses this feedback to improve the performance of the translation engine.

[1201] Prompt Sentence Examples

[1202] When User A says "Hello" in English, the earphone microphone captures the audio and sends it to the smartphone. The smartphone's speech recognition engine converts "Hello" into text, and the emotion analysis means recognizes the emotion as "happy." This information is sent to the translation engine on the server, and the audio converted to "Hello" in a happy voice tone is played from the earphone speaker.

[1203] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1204] Step 1:

[1205] The user puts on the system's earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. A connection confirmation message is displayed on the smartphone, and the user sets the input language and target language. This setting information is saved on the smartphone.

[1206] Specific behavior:

[1207] Input: User puts in earphones and launches translation app.

[1208] Data processing / calculation: The app checks the earphone connection and receives language setting information.

[1209] Output: A connection completion message and a setting confirmation message are displayed.

[1210] Step 2:

[1211] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data to the smartphone in real time.

[1212] Specific behavior:

[1213] Input: User's voice.

[1214] Data processing / calculation: The microphone captures the sound and converts it into digital sound data.

[1215] Output: Digital audio data is sent to the smartphone.

[1216] Step 3:

[1217] The terminal (smartphone) uses a voice recognition means to convert the voice data into text data, and the conversion result is sent to the emotion analysis means and the server.

[1218] Specific behavior:

[1219] Input: Digital audio data.

[1220] Data processing / calculation: The voice recognition engine analyzes the voice data and converts it into text data.

[1221] Output: The text data and the voice data are sent to the emotion analysis means and the server.

[1222] Step 4:

[1223] The emotion analysis means of the terminal analyzes the voice tone, speech rate, and language pattern to recognize the user's emotion, and this emotion data is sent to the translation means.

[1224] Specific behavior:

[1225] Input: Audio data.

[1226] Data processing / calculation: Emotion analysis means analyzes voice tone, speaking rate, and language patterns to extract emotions.

[1227] Output: The recognized emotion data is sent to the translation means.

[1228] Step 5:

[1229] The server's translation means receives the text data and emotion data and translates it into the target language. The AI ​​translation engine takes context and emotion into account when translating.

[1230] Specific behavior:

[1231] Input: Text data and emotion data.

[1232] Data processing / calculation: The translation engine analyzes the context and sentiment to generate an appropriate translation.

[1233] Output: The translated text data is generated and sent back to the smartphone.

[1234] Step 6:

[1235] The terminal receives the translation result returned from the server and converts the translated text data into voice data using a voice synthesis means.

[1236] Specific behavior:

[1237] Input: Translation text data.

[1238] Data processing / calculation: The speech synthesis engine converts text data into speech data.

[1239] Output: Audio data is generated.

[1240] Step 7:

[1241] The generated audio data is transmitted from the terminal to earphones, allowing the user to listen to it in real time.

[1242] Specific behavior:

[1243] Input: Audio data.

[1244] Data processing / calculation: Transfer data to earphones.

[1245] Output: Audio data is played from the earphone speaker.

[1246] Step 8:

[1247] After completing a conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server where it is analyzed by the server's learning means and used to improve the performance of the translation means.

[1248] Specific behavior:

[1249] Input: User feedback.

[1250] Data processing / calculation: Feedback data is sent to the server and analyzed.

[1251] Output: The analysis results are added to the translation engine's training data, improving translation performance.

[1252] (Application example 2)

[1253] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1254] There is a need for a method that allows people who speak different languages ​​to communicate smoothly in real time, including expressing their emotions. In particular, in the food delivery scene, a system is needed that allows delivery personnel and customers to communicate without worrying about language differences.

[1255] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for inputting voice, voice recognition means for converting voice data input by the input means into text data, translation means for translating the text data generated by the voice recognition means into a specified target language, voice synthesis means for converting the translated text data generated by the translation means into voice data, output means for outputting the voice data generated by the voice synthesis means, learning means for receiving feedback on translation quality provided by the user and saving it as learning data for the translation means, and an emotion engine that recognizes emotions by analyzing voice tone, speaking rate, and language patterns. This enables real-time communication incorporating emotions even when the delivery person and the customer speak different languages.

[1256] The "input means for inputting voice" is a device for capturing the user's speech.

[1257] "Speech recognition means for converting voice data input by an input means into text data" refers to a device or technology for converting voice input into data in text format.

[1258] "Translation means for translating text data generated by speech recognition means into a specified target language" refers to a device or technology for converting the generated text data into the target language.

[1259] "Speech synthesis means for converting translated text data generated by the translation means into speech data" refers to a device or technology for converting translated text into speech-format data.

[1260] The "output means for outputting the voice data generated by the voice synthesis means" is a device that enables the user to listen to the voice data.

[1261] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" is a technology for collecting user feedback and improving translation accuracy based on that data.

[1262] The "emotion engine that recognizes emotions by analyzing voice tone, speech rate, and language patterns" is a technology that analyzes the tone and rate of speech, the use of words, etc. to understand the speaker's emotions.

[1263] This invention is a system for realizing emotional real-time communication between delivery personnel and customers even when they speak different languages. The system includes the following elements:

[1264] 1. Input means: An input device for capturing the user's voice. Specifically, a microphone built into a smartphone or a microphone built into earphones can be used.

[1265] 2. Speech recognition means: Software for converting acquired voice data into text data. This is achieved by using the Python speech_recognition library.

[1266] 3. Translation tool: Software for translating the text data generated by the speech recognition tool into the target language. Translation is performed using Google's googletrans library.

[1267] 4. Speech synthesis means: Software for converting translated text data into speech data. This is achieved using the gTTS (Google Text-to-Speech) library.

[1268] 5. Output means: A device for transmitting the generated voice data to the user. Specifically, the speaker of a smartphone or a speaker of an earphone can be used.

[1269] 6. Learning method: A technology for collecting feedback data from users and using it to improve translation performance. The data is sent to the server for storage and analysis.

[1270] 7. Emotion Engine: Technology that analyzes voice tone, speaking rate, and language patterns to recognize user emotions, resulting in more natural and emotionally relevant translation results.

[1271] Hardware and Software Used

[1272] 1. Hardware:

[1273] Smartphone (microphone, speaker)

[1274] Earphones (microphone, speaker)

[1275] 2. Software:

[1276] Python

[1277] speech_recognition library (speech recognition)

[1278] googletrans library (translation)

[1279] gTTS library (speech synthesis)

[1280] Explanation of data processing and data calculation

[1281] The server receives voice data input via a smartphone or earphones and converts it into text data using a voice recognition means. This text data is then converted into the target language using a translation means. The translated text is converted back into voice data using a voice synthesis means and transmitted to the user via an output means. In addition, the learning means analyzes user feedback and uses it to improve system performance. During this series of processes, the emotion engine analyzes the user's emotions and reflects them in the translation results.

[1282] Specific examples

[1283] For example, consider the case where a delivery person says, "This is a delivery item." When the delivery person says, "I've brought your order" in English, this voice is captured by the smartphone's microphone. The speech recognition means converts this into text data, obtaining the string "I've brought your order." The translation means then translates this into Japanese, generating the text data "This is a delivery item." The speech synthesis means converts this Japanese text into speech data, which is finally conveyed to the customer via the smartphone or earphone speaker.

[1284] Example prompts to input to the generative AI model

[1285] Create a real-time translation assistant to help people who speak different languages ​​communicate smoothly. When a delivery person says "Delivery here" to a customer, it needs to translate from English to Japanese. This translation assistant will be realized using speech recognition, text translation, and speech synthesis.

[1286] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1287] Step 1:

[1288] Input and Output

[1289] A user starts the smartphone app and speaks into the microphone. The input is the user's voice. The output is recorded as audio data on the smartphone.

[1290] Specific actions

[1291] The user launches the smartphone app and says, "I've brought your order" in English. The smartphone's microphone captures the voice and generates voice data.

[1292] Step 2:

[1293] Input and Output

[1294] The voice data is sent to the voice recognition means. At this time, the input is the voice data generated in step 1. The output is the text data "I've brought your order" generated by the voice recognition means.

[1295] Specific actions

[1296] The smartphone's voice recognition engine (speech_recognition library) analyzes the voice data and converts it into corresponding text data.

[1297] Step 3:

[1298] Input and Output

[1299] The text data is sent to the translation means. At this time, the input is the text data "I've brought your order" generated in step 2. The output is the text data "Delivery item" in the target language generated by the translation means.

[1300] Specific actions

[1301] The smartphone's Google Trans library translates the text data into Japanese.

[1302] Step 4:

[1303] Input and Output

[1304] The text data in the target language is sent to the speech synthesis means. At this time, the input is the text data "Delivery item" generated in step 3. The output is speech data generated by the speech synthesis means.

[1305] Specific actions

[1306] The smartphone's gTTS library converts the text data into audio data and generates an audio file.

[1307] Step 5:

[1308] Input and Output

[1309] The audio data is sent to an output means, where the input is the audio data generated in step 4. The output is audio that can be heard by the user.

[1310] Specific actions

[1311] The translated voice message "Delivery here" is played to the customer through the speaker on their smartphone or earphones.

[1312] Step 6:

[1313] Input and Output

[1314] Feedback from the user is sent to the learning means. At this time, the input is the feedback on translation quality provided by the user. The output is the feedback data stored on the server and becomes learning data.

[1315] Specific actions

[1316] Users provide feedback on the quality of the translation through the app, which is then sent to the server, which analyzes the data and stores it in the translation engine's training database.

[1317] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1318] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1319] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1320] [Fourth embodiment]

[1321] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1322] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1323] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1324] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1325] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1326] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1327] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1328] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1329] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1330] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1331] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1332] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1333] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1334] MODE FOR CARRYING OUT THE INVENTION

[1335] The present invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. The detailed configuration and operation of the system will be described below.

[1336] System Configuration

[1337] The system consists of the following main components:

[1338] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[1339] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[1340] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[1341] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[1342] 5. Output means: A device that allows the user to listen to the generated audio data. Specifically, it is a speaker built into an earphone.

[1343] 6. Learning tool: A device or technology that receives feedback from users and trains the translation tool. This includes a feedback processing system on the server.

[1344] System Operation

[1345] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[1346] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[1347] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to a translation engine on the server.

[1348] The server's AI translation engine receives the text data and translates it into the specified target language. During the translation process, it uses natural language processing technology to understand the context and select the appropriate translation.

[1349] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[1350] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[1351] Specific examples

[1352] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[1353] 1. User A says "Hello" in English.

[1354] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1355] 3. The smartphone's voice recognition function converts "Hello" into text data.

[1356] 4. The smartphone sends the text data to the server.

[1357] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[1358] 6. The server sends the translation results back to the smartphone.

[1359] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[1360] 8. The earphone speaker will say "hello."

[1361] 9. User B hears the translated "hello" in real time.

[1362] This system allows users A and B to communicate smoothly across language barriers.

[1363] The processing flow will be explained below.

[1364] Step 1:

[1365] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[1366] Step 2:

[1367] The user sets the input language and target language in the translation app, and also adjusts the volume and enables noise cancellation as needed.

[1368] Step 3:

[1369] When the user starts speaking, the microphone built into the earphones captures the user's voice, and the captured voice data is sent to the smartphone in real time.

[1370] Step 4:

[1371] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[1372] Step 5:

[1373] The terminal transmits the converted text data to a server via the Internet.

[1374] Step 6:

[1375] The server receives the text data and translates it into the specified target language using an AI translation engine. The translation engine uses natural language processing to understand the context and select the appropriate translation.

[1376] Step 7:

[1377] The server returns the translated text data to the device, and also prepares to receive feedback as learning data to improve translation accuracy.

[1378] Step 8:

[1379] The terminal passes the translated text data received from the server to a speech synthesis engine and converts it into voice data.

[1380] Step 9:

[1381] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[1382] Step 10:

[1383] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[1384] Step 11:

[1385] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[1386] Example 1

[1387] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1388] It is difficult for people who speak different languages ​​to communicate smoothly in real time. Furthermore, the quality of translation must be improved. Therefore, there is a need for a system that can realize real-time translation in multiple languages ​​and continuously improve its translation performance by receiving feedback.

[1389] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1390] In this invention, the server includes a connection means for automatically connecting to a user's device, an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, a translation means for translating the text data generated by the speech recognition means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by the user and saving the feedback as learning data for the translation means. This enables people speaking different languages ​​to communicate smoothly in real time. Furthermore, by utilizing user feedback, translation performance can be continuously improved.

[1391] "User" refers to an individual who uses the system to communicate with others who speak different languages.

[1392] "Device" refers to a device that includes hardware for use by a user, such as earphones or smartphones.

[1393] "Connection means" refers to a method or technology for connecting a user's device to each component of the system. Specifically, this applies to wireless communication technologies such as Bluetooth and Wi-Fi.

[1394] "Input means" refers to a device or technology for inputting sound into the system, specifically a microphone built into earphones.

[1395] "Voice data" refers to digital audio information that records a user's speech.

[1396] "Text data" refers to linguistic information in character format converted by a speech recognition means.

[1397] "Speech recognition means" refers to a technique or device for converting voice data into text data.

[1398] "Translation means" refers to technology or equipment for converting text data into a specified target language. Specifically, it includes an AI translation engine installed on a server.

[1399] "Speech synthesis means" refers to technology or equipment for converting translated text data into speech data.

[1400] "Output means" refers to a device or technology for transmitting synthesized voice data to the user. Specifically, this applies to the speaker built into the earphone.

[1401] "Feedback" refers to evaluation information about the quality of a translation provided by a user.

[1402] "Training means" refers to a technique or device for analyzing received feedback and improving the performance of the translation means.

[1403] MODE FOR CARRYING OUT THE INVENTION

[1404] The present invention relates to a system that enables people who speak different languages ​​to communicate smoothly in real time. The detailed configuration and operation of the system will be described below.

[1405] System Hardware

[1406] The system hardware consists of the following major components:

[1407] 1. Earphones: Earphones with built-in microphones and speakers that are responsible for inputting and outputting audio data.

[1408] 2. Smartphone: A translation application is installed and performs voice recognition, data transmission and reception, and voice synthesis.

[1409] 3. Server: Equipped with an AI translation engine, it translates text data.

[1410] Software used

[1411] The software used in this system is as follows:

[1412] 1. Smartphone application:

[1413] Speech recognition engine: Converts voice data into text data.

[1414] Speech synthesis engine: Converts translated text data into speech data.

[1415] Feedback feature: Receive feedback from users.

[1416] 2. AI translation engine on the server:

[1417] Uses natural language processing (NLP) technology to perform highly accurate translations.

[1418] System Operation

[1419] The user puts on the earphones and launches a translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. When the user begins to speak, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time.

[1420] The voice recognition engine of the smartphone terminal converts the received voice data into text data. For example, if a user says "Hello," the voice recognition engine generates the text data "Hello." The generated text data is then sent to the server.

[1421] The server's AI translation engine receives the text data and translates it into the specified target language. The server then uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The translation result is then sent back from the server to the smartphone.

[1422] The speech synthesis engine on the smartphone, which is the device that receives the translation results, converts the translated text data into voice data. This voice data is sent to the earphones. The earphone speakers pronounce the translated voice. User B can hear the voice saying "Hello" in real time.

[1423] After the conversation is over, the user provides feedback on the quality of the translation through the application, which then sends the feedback data to the server, where it is analyzed by the server's feedback processing system and used as training data for the AI ​​translation engine.

[1424] Specific examples

[1425] The following example shows how the system works when an English-speaking user A and a Japanese-speaking user B communicate in real time:

[1426] 1. User A says "Hello" in English.

[1427] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1428] 3. The smartphone's voice recognition function converts "Hello" into text data.

[1429] 4. The smartphone sends the text data to the server.

[1430] 5. The server's translation means translates "Hello" into Japanese "Konnichiwa".

[1431] 6. The server sends the translation results back to the smartphone.

[1432] 7. The smartphone's voice synthesis function converts "hello" into voice data.

[1433] 8. The earphone speaker will say "hello."

[1434] 9. User B hears the translated "hello" in real time.

[1435] Prompt Sentence Examples

[1436] "Please explain how a system works, allowing users to communicate in real time with others who speak different languages. Please provide detailed examples of how data is processed using specific hardware (earphones, smartphone, server) and software (speech recognition engine, translation engine, speech synthesis engine)."

[1437] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1438] Step 1: Connection and Setup

[1439] The user puts on the dedicated earphones and launches the translation app installed on their smartphone. When the app is launched, the earphones and smartphone are automatically connected via Bluetooth. Next, the user sets the input language (e.g., English) and target language (e.g., Japanese) in the app. The input for this step is the user's language setting and earphone connection information. The output is the completed setting state.

[1440] Step 2: Capture audio

[1441] When a user starts talking, the microphone built into the earphones captures the voice. This voice data is sent to the smartphone in real time. The input of this step is the voice from the user, and the output is the voice data sent to the smartphone.

[1442] Step 3: Speech to text

[1443] The device (smartphone) converts the received voice data into text data using a voice recognition engine. For example, the voice "Hello" is converted into the text data "Hello." The input of this step is the captured voice data, and the output is the generated text data.

[1444] Step 4: Translate the text

[1445] The server's AI translation engine receives the text data and translates it into the specified target language. The server uses natural language processing technology to analyze the meaning of the text and translate it into the target language. For example, "Hello" is translated into "Konnichiwa" in Japanese. The input of this step is the text data, and the output is the translated text data.

[1446] Step 5: Speech synthesis of the translation result

[1447] The device (smartphone) converts the translated text data returned from the server into voice data using a speech synthesis engine. The text "Hello" is converted into voice data saying "Hello." The input of this step is the translated text data, and the output is the generated voice data.

[1448] Step 6: Output the translated audio

[1449] The earphone speaker provides the generated voice data to the user. User B can hear the voice saying "Hello" in real time. The input of this step is the synthesized voice data, and the output is the voice the user hears.

[1450] Step 7: Provide feedback

[1451] After the conversation is completed, the user provides feedback on the quality of the translation through the smartphone application, for example, by entering feedback such as "very satisfied." The input of this step is the feedback data from the user, and the output is the feedback data sent to the server.

[1452] Step 8: Analyze feedback and learn

[1453] The server analyzes the received feedback data and stores it as training data for the AI ​​translation engine, thereby continuously improving translation performance. The input of this step is the received feedback data, and the output is updated training data.

[1454] (Application example 1)

[1455] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1456] The goal of this project is to achieve efficient real-time communication between staff members who speak different languages ​​at a logistics center where multinational staff work. This will enable improved efficiency and accuracy of work. However, with the previous system, there were barriers to communication between staff members, which hindered the smooth execution of work.

[1457] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1458] In this invention, the server includes an input means for inputting speech, a speech recognition means, a translation means, a speech synthesis means, an output means, a learning means, and a means used to enable efficient communication among multinational staff at a logistics center, thereby enabling real-time speech translation between multiple languages ​​and communication based on the translation.

[1459] The "input means for inputting voice" is a device that captures the voice uttered by the user as digital data.

[1460] "Speech recognition means" is a technology for converting captured voice data into text data.

[1461] "Translation means" refers to a technique for converting text data generated by speech recognition means into a target language.

[1462] "Speech synthesis means" is a technology that converts text data translated into a target language into speech data.

[1463] The "output means for outputting voice data" is a device that allows the user to listen to the translated voice data.

[1464] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" refers to a database and processing system for receiving evaluations from users and improving translation accuracy based on them.

[1465] "Means used to enable multinational staff to communicate efficiently at logistics centers" refers to components of a system that enables multiple staff members who speak different languages ​​to communicate smoothly through a voice translation system.

[1466] System configuration

[1467] This invention is a system for realizing real-time communication among multinational staff in a logistics center. This system is composed of the following main elements:

[1468] 1. A microphone built into earphones worn by the user as an input means for inputting voice.

[1469] 2. A speech recognition engine built into a smartphone that converts voice data into text data as a means of speech recognition.

[1470] 3. As a translation tool, an AI translation engine on a server that translates text data into a specified target language.

[1471] 4. A smartphone speech synthesis engine that converts translated text data into speech data as a speech synthesis method.

[1472] 5. A speaker built into the earphone as an output means for outputting audio data.

[1473] 6. A feedback processing system on the server that receives feedback provided by users as a learning means and stores it as learning data for the translation means.

[1474] 7. A system configuration that combines the above elements as a means to enable multinational staff to communicate efficiently in a logistics center.

[1475] How it works

[1476] The user uses the system by putting on earphones and launching a dedicated translation application installed on their smartphone. The system operates as follows:

[1477] 1. Voice input:

[1478] When the user speaks, the microphone built into the earphones captures the sound, and the audio data is sent to the smartphone in real time.

[1479] 2. Speech Recognition:

[1480] The smartphone's voice recognition engine converts the received voice data into text data.

[1481] 3. Translation:

[1482] The text data is sent to a server, where an AI translation engine translates it into the target language.

[1483] 4. Speech synthesis:

[1484] The translated text data is then sent back to the smartphone, where it is converted into voice data by a voice synthesis engine.

[1485] 5. Audio output:

[1486] The converted audio data is output to the user through the earphone speaker.

[1487] 6. Feedback and learning:

[1488] After the conversation, the user provides feedback on the quality of the translation, which is sent to the server. The feedback processing system analyzes this and stores it as learning data for the AI ​​translation engine, contributing to future improvements in translation accuracy.

[1489] Hardware and software used

[1490] Hardware:

[1491] Earphones (with built-in microphone and speaker)

[1492] Smartphone

[1493] software:

[1494] SpeechRecognition (speech recognition library)

[1495] GoogleTrans (translation library)

[1496] gTTS (Text-to-Speech Library)

[1497] Playsound (sound playback library)

[1498] Specific examples

[1499] Logistics centers often employ foreign staff who do not understand Japanese. With this system, instructions such as "Please pick up the items from aisle 3" can be given in Japanese and instantly translated into the staff's native language, ensuring smooth communication.

[1500] Prompt Sentence Examples

[1501] "Write Python code to translate English audio into Japanese and play it back as audio."

[1502] This will enable smooth communication between multinational staff at the logistics center, significantly improving work efficiency.

[1503] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1504] Step 1:

[1505] The user puts on the earphones and launches the translation application installed on their smartphone. When the application is launched, the system automatically connects to the earphones via Bluetooth and is ready for voice input. The user then sets the input language and target language in the app.

[1506] Input: User utterance

[1507] Output: Ready notification, setting language information

[1508] Step 2:

[1509] When a user speaks, the microphone built into the earphones captures the voice. This voice data is sent to a smartphone in real time via Bluetooth. The user's voice is the input data.

[1510] Input: User's voice data

[1511] Output: Microphone captured audio data

[1512] Step 3:

[1513] The smartphone's voice recognition engine converts the received voice data into text data. Specifically, the SpeechRecognition library processes the voice data and generates text data.

[1514] Input: Audio data

[1515] Output: Converted text data

[1516] Step 4:

[1517] The text data is sent to the server via the internet. The server's AI translation engine uses the Google Translate API to translate the text into the target language. The translated text data is output.

[1518] Input: Text data, target language information

[1519] Output: Translated text data

[1520] Step 5:

[1521] The translated text data is returned to the smartphone and converted into audio data by the speech synthesis engine. Specifically, the gTTS library converts the text data into an audio file.

[1522] Input: Translated text data

[1523] Output: Synthesized speech data

[1524] Step 6:

[1525] The generated voice data is sent from the smartphone via Bluetooth to the earphones, which then output the audio through their speakers, allowing the user to hear the audio in the target language in real time.

[1526] Input: Audio data

[1527] Output: Audio output from earphones

[1528] Step 7:

[1529] After the conversation is over, the user provides feedback on the quality of the translation via the app. The feedback data is sent to the server and stored in the feedback processing system. This data is used as training data for the translation engine to improve future translation accuracy.

[1530] Input: User feedback data

[1531] Output: Improved learning model with saved feedback data

[1532] Through these steps, efficient real-time communication between multinational staff at a logistics center is achieved.

[1533] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1534] MODE FOR CARRYING OUT THE INVENTION

[1535] This invention is a real-time translation system that enables people who speak different languages ​​to communicate smoothly. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide translation results that correspond to those emotions. The detailed configuration and operation of the system are described below.

[1536] System Configuration

[1537] The system consists of the following main components:

[1538] 1. Input means: A device for acquiring audio. Specifically, it is a microphone built into earphones.

[1539] 2. Speech recognition means: A device or technology for converting voice data into text data. A speech recognition engine built into a smartphone application falls into this category.

[1540] 3. Translation tool: A device or technology for converting text data into a target language. An AI translation engine on a server is an example of this.

[1541] 4. Speech synthesis means: A device or technology for converting translated text data into speech data. This applies to speech synthesis engines built into smartphone applications.

[1542] 5. Output means: A device for outputting the generated audio data in a form that can be heard by the user. Specifically, it is a speaker built into an earphone.

[1543] 6. Learning means: A device or technology for receiving feedback from users and storing it as learning data for the translation means. This includes a feedback processing system on a server.

[1544] 7. Emotion Engine: A device or technology that recognizes the user's emotions and reflects them in the translation process. Emotions are recognized by analyzing voice tone, speech rate, and language patterns.

[1545] System Operation

[1546] The user puts on the earphones and launches the translation app installed on their smartphone. The moment the app is launched, the earphones and smartphone are automatically connected via Bluetooth. The user then sets the input language and target language on the app.

[1547] When a user starts talking, the microphone built into the earphones captures the audio, which is then sent to a smartphone in real time.

[1548] The smartphone device converts the captured voice data into text data using a speech recognition means, which is then sent to the emotion engine and the translation engine on the server.

[1549] The emotion engine analyzes voice tone, speech rate, and language patterns to recognize the user's emotion, and this emotion data is sent to the translation engine.

[1550] The server's AI translation engine receives the text data and translates it into the specified target language, taking into account the recognized emotional data. During the translation process, natural language processing technology is used to understand the context and emotions and select the appropriate translation.

[1551] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[1552] After the conversation, the user provides feedback on the quality of the translation through the app, which is sent to the server, where the server's learning mechanism analyzes this data and uses it to improve the translation mechanism's performance.

[1553] Specific examples

[1554] The following example shows the operation of the system when an English-speaking user A and a Japanese-speaking user B communicate in real time.

[1555] 1. User A says "Hello" in English. His voice sounds cheerful.

[1556] 2. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1557] 3. The smartphone's voice recognition function converts "Hello" into text data.

[1558] 4. The smartphone sends the text data to the emotion engine and server.

[1559] 5. The emotion engine analyzes User A's tone and speech rate and recognizes the emotion "fun." This emotion data is sent to the server.

[1560] 6. The server's translation means translates "Hello" to the Japanese "Konnichiwa" and maintains a happy tone of voice.

[1561] 7. The server sends the translation results back to the smartphone.

[1562] 8. The smartphone's voice synthesis function converts "hello" into voice data in a pleasant voice tone.

[1563] 9. The earphone speaker pronounces "hello" in a cheerful voice.

[1564] 10. User B hears the translated "Hello" in real time in a happy voice.

[1565] This system allows users A and B to communicate smoothly and emotionally, overcoming language barriers.

[1566] The processing flow will be explained below.

[1567] Step 1:

[1568] The user puts on the earphones and launches the translation app installed on their smartphone. Once the app is launched, the earphones and smartphone are automatically connected via Bluetooth.

[1569] Step 2:

[1570] The user sets the input language and target language in the translation app, and also enables the emotion recognition function.

[1571] Step 3:

[1572] When the user starts speaking, the microphone built into the earphones captures the voice, and this voice data is sent to the smartphone in real time.

[1573] Step 4:

[1574] The device (smartphone) receives the voice data and starts processing it with its built-in voice recognition engine, which converts the voice data into text data.

[1575] Step 5:

[1576] The terminal sends the converted text data to the emotion engine, which analyzes the voice tone, speech rate, and language patterns to recognize the user's emotion, and the recognized emotion data is sent to the server.

[1577] Step 6:

[1578] The terminal transmits the text data to a server via the Internet.

[1579] Step 7:

[1580] The server receives the text data and emotional data, and the AI ​​translation engine translates it into the specified target language. The translation engine also takes into account the emotional data, understands the context, and selects the appropriate translation and tone of voice.

[1581] Step 8:

[1582] The server returns the translation results that reflect the emotional data to the terminal.

[1583] Step 9:

[1584] The device passes the translated text data received from the server to a speech synthesis engine and converts it into voice data, while maintaining the voice tone that reflects the emotional data.

[1585] Step 10:

[1586] The terminal transmits the generated voice data to the earphone, and the earphone provides the translation result to the user by voice.

[1587] Step 11:

[1588] After the conversation, the user provides feedback on the translation quality through the app, which is then sent to the server via the device.

[1589] Step 12:

[1590] The server receives feedback from users and uses it as learning data for the AI ​​translation engine to continuously improve translation quality.

[1591] Example 2

[1592] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1593] There is a demand for systems that allow people who speak different languages ​​to communicate smoothly in real time. However, even if traditional translation systems provide accurate translations, they tend to fail to reflect emotional expressions, resulting in incomplete communication. Another issue is the lack of a mechanism for improving the system's accuracy through feedback.

[1594] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an input means for inputting speech, a speech recognition means for converting speech data input by the input means into text data, an emotion analysis means for analyzing speech tone, speech rate, and language patterns to recognize a user's emotion, a translation means for translating the text data and emotion data generated by the speech recognition means and the emotion analysis means into a specified target language, a speech synthesis means for converting the translated text data generated by the translation means into speech data, an output means for outputting the speech data generated by the speech synthesis means, and a learning means for receiving feedback on translation quality provided by a user and saving the feedback as learning data for the translation means. This enables smooth real-time communication that incorporates emotions between users who speak different languages.

[1595] "Input means" refers to a device for acquiring audio, specifically a microphone built into an earphone.

[1596] The "voice recognition means" is a device or technology for converting voice data acquired by the input means into text data.

[1597] An "emotion analysis means" is a device or technology for recognizing a user's emotions by analyzing voice tone, speech rate, and language patterns.

[1598] The "translation means" is a device or technology for translating the text data and emotion data generated by the speech recognition means and emotion analysis means into a specified target language.

[1599] The "speech synthesis means" refers to a device or technology for converting the translated text data generated by the translation means into speech data.

[1600] The "output means" is a device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically refers to a speaker built into an earphone.

[1601] "Learning means" refers to a device or technology that receives feedback on translation quality provided by a user and stores it as training data for the translation means.

[1602] The present invention is a system that enables people who speak different languages ​​to communicate smoothly in real time by using advanced technologies that integrate speech recognition, emotion analysis, translation, speech synthesis, and feedback processing.

[1603] System configuration

[1604] The system consists of the following main components:

[1605] 1. Input means: A device for capturing audio, specifically a microphone built into earphones.

[1606] 2. Speech recognition means: A device for converting voice data acquired by the input means into text data, using a voice recognition engine built into the smartphone.

[1607] 3. Emotion analysis means: A device for analyzing voice tone, speech rate, and language patterns to recognize the user's emotions.

[1608] 4. Translation means: A device for translating text data and emotional data generated by the speech recognition means and emotional analysis means into a specified target language. It uses an AI translation engine on the server.

[1609] 5. Speech synthesis means: A device for converting the translated text data generated by the translation means into speech data, and uses a speech synthesis engine built into the smartphone.

[1610] 6. Output means: A device for outputting the voice data generated by the voice synthesis means in a form that can be heard by the user, and specifically, a speaker built into the earphone is used.

[1611] 7. Learning means: A device that receives feedback on translation quality provided by users, stores it as learning data for the translation means, and improves the system's performance. A feedback processing system on the server is used.

[1612] System Operation

[1613] The user puts on the earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. The user sets the input language and target language to be used in the application. This setting information is saved on the smartphone.

[1614] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data in real time to the smartphone. The smartphone then uses a voice recognition engine to convert the voice data into text data. The conversion result is then sent to the emotion analysis means and the server.

[1615] The device's emotion analysis means analyzes voice tone, speech rate, and language patterns to recognize the user's emotions. This emotion data is sent to the translation means. The server's AI translation engine receives the text data and emotion data and translates it into the specified target language. The AI ​​translation engine understands the context and emotion and selects the appropriate translation.

[1616] The translation result is sent back from the server to the smartphone and then sent to a speech synthesis means for conversion into audio data, which is then sent to earphones so the user can hear the translated audio in real time.

[1617] After the conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server and analyzed by the server's learning means and used to improve the performance of the translation means.

[1618] Specific operation example

[1619] For example, when an English-speaking user A and a Japanese-speaking user B communicate in real time, the system operates as follows:

[1620] 1. User A launches the translation app and puts on earphones.

[1621] 2. User A sets English as the input language and Japanese as the target language on the app.

[1622] 3. User A says "Hello" happily.

[1623] 4. The earphone's microphone captures the sound and sends the audio data to your smartphone.

[1624] 5. The device uses a voice recognition means to convert the voice into text data such as "Hello."

[1625] 6. This data is sent to the server and sentiment analysis means.

[1626] 7. The device's emotion analysis means recognizes the emotion "fun" and sends this information to the server.

[1627] 8. The server's AI translation engine translates "Hello" into "Hello" taking into account the "joyful" emotion.

[1628] 9. The server sends the translation results back to the smartphone.

[1629] 10. The device's voice synthesis means converts "hello" into voice data in a "happy" voice tone.

[1630] 11. The earphone speaker says "Hello" to User B in a cheerful voice.

[1631] 12. User B hears "Hello" in real time and continues the conversation.

[1632] 13. After the conversation, User A provides feedback on the quality of the translation through the app, which is then sent to the server.

[1633] 14. The server's learning mechanism uses this feedback to improve the performance of the translation engine.

[1634] Prompt Sentence Examples

[1635] When User A says "Hello" in English, the earphone microphone captures the audio and sends it to the smartphone. The smartphone's speech recognition engine converts "Hello" into text, and the emotion analysis means recognizes the emotion as "happy." This information is sent to the translation engine on the server, and the audio converted to "Hello" in a happy voice tone is played from the earphone speaker.

[1636] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1637] Step 1:

[1638] The user puts on the system's earphones and launches the translation application installed on their smartphone. When the application launches, the earphones and smartphone are automatically connected via Bluetooth. A connection confirmation message is displayed on the smartphone, and the user sets the input language and target language. This setting information is saved on the smartphone.

[1639] Specific behavior:

[1640] Input: User puts in earphones and launches translation app.

[1641] Data processing / calculation: The app checks the earphone connection and receives language setting information.

[1642] Output: A connection completion message and a setting confirmation message are displayed.

[1643] Step 2:

[1644] When the user starts speaking, the microphone built into the earphones captures the voice and transmits the voice data to the smartphone in real time.

[1645] Specific behavior:

[1646] Input: User's voice.

[1647] Data processing / calculation: The microphone captures the sound and converts it into digital sound data.

[1648] Output: Digital audio data is sent to the smartphone.

[1649] Step 3:

[1650] The terminal (smartphone) uses a voice recognition means to convert the voice data into text data, and the conversion result is sent to the emotion analysis means and the server.

[1651] Specific behavior:

[1652] Input: Digital audio data.

[1653] Data processing / calculation: The voice recognition engine analyzes the voice data and converts it into text data.

[1654] Output: The text data and the voice data are sent to the emotion analysis means and the server.

[1655] Step 4:

[1656] The emotion analysis means of the terminal analyzes the voice tone, speech rate, and language pattern to recognize the user's emotion, and this emotion data is sent to the translation means.

[1657] Specific behavior:

[1658] Input: Audio data.

[1659] Data processing / calculation: Emotion analysis means analyzes voice tone, speaking rate, and language patterns to extract emotions.

[1660] Output: The recognized emotion data is sent to the translation means.

[1661] Step 5:

[1662] The server's translation means receives the text data and emotion data and translates it into the target language. The AI ​​translation engine takes context and emotion into account when translating.

[1663] Specific behavior:

[1664] Input: Text data and emotion data.

[1665] Data processing / calculation: The translation engine analyzes the context and sentiment to generate an appropriate translation.

[1666] Output: The translated text data is generated and sent back to the smartphone.

[1667] Step 6:

[1668] The terminal receives the translation result returned from the server and converts the translated text data into voice data using a voice synthesis means.

[1669] Specific behavior:

[1670] Input: Translation text data.

[1671] Data processing / calculation: The speech synthesis engine converts text data into speech data.

[1672] Output: Audio data is generated.

[1673] Step 7:

[1674] The generated audio data is transmitted from the terminal to earphones, allowing the user to listen to it in real time.

[1675] Specific behavior:

[1676] Input: Audio data.

[1677] Data processing / calculation: Transfer data to earphones.

[1678] Output: Audio data is played from the earphone speaker.

[1679] Step 8:

[1680] After completing a conversation, the user provides feedback on the quality of the translation through the application, which is sent to the server where it is analyzed by the server's learning means and used to improve the performance of the translation means.

[1681] Specific behavior:

[1682] Input: User feedback.

[1683] Data processing / calculation: Feedback data is sent to the server and analyzed.

[1684] Output: The analysis results are added to the translation engine's training data, improving translation performance.

[1685] (Application example 2)

[1686] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1687] There is a need for a method that allows people who speak different languages ​​to communicate smoothly in real time, including expressing their emotions. In particular, in the food delivery scene, a system is needed that allows delivery personnel and customers to communicate without worrying about language differences.

[1688] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for inputting voice, voice recognition means for converting voice data input by the input means into text data, translation means for translating the text data generated by the voice recognition means into a specified target language, voice synthesis means for converting the translated text data generated by the translation means into voice data, output means for outputting the voice data generated by the voice synthesis means, learning means for receiving feedback on translation quality provided by the user and saving it as learning data for the translation means, and an emotion engine that recognizes emotions by analyzing voice tone, speaking rate, and language patterns. This enables real-time communication incorporating emotions even when the delivery person and the customer speak different languages.

[1689] The "input means for inputting voice" is a device for capturing the user's speech.

[1690] "Speech recognition means for converting voice data input by an input means into text data" refers to a device or technology for converting voice input into data in text format.

[1691] "Translation means for translating text data generated by speech recognition means into a specified target language" refers to a device or technology for converting the generated text data into the target language.

[1692] "Speech synthesis means for converting translated text data generated by the translation means into speech data" refers to a device or technology for converting translated text into speech-format data.

[1693] The "output means for outputting the voice data generated by the voice synthesis means" is a device that enables the user to listen to the voice data.

[1694] The "learning means for receiving feedback on translation quality provided by users and storing it as learning data for the translation means" is a technology for collecting user feedback and improving translation accuracy based on that data.

[1695] The "emotion engine that recognizes emotions by analyzing voice tone, speech rate, and language patterns" is a technology that analyzes the tone and rate of speech, the use of words, etc. to understand the speaker's emotions.

[1696] This invention is a system for realizing emotional real-time communication between delivery personnel and customers even when they speak different languages. The system includes the following elements:

[1697] 1. Input means: An input device for capturing the user's voice. Specifically, a microphone built into a smartphone or a microphone built into earphones can be used.

[1698] 2. Speech recognition means: Software for converting acquired voice data into text data. This is achieved by using the Python speech_recognition library.

[1699] 3. Translation tool: Software for translating the text data generated by the speech recognition tool into the target language. Translation is performed using Google's googletrans library.

[1700] 4. Speech synthesis means: Software for converting translated text data into speech data. This is achieved using the gTTS (Google Text-to-Speech) library.

[1701] 5. Output means: A device for transmitting the generated voice data to the user. Specifically, the speaker of a smartphone or a speaker of an earphone can be used.

[1702] 6. Learning method: A technology for collecting feedback data from users and using it to improve translation performance. The data is sent to the server for storage and analysis.

[1703] 7. Emotion Engine: Technology that analyzes voice tone, speaking rate, and language patterns to recognize user emotions, resulting in more natural and emotionally relevant translation results.

[1704] Hardware and Software Used

[1705] 1. Hardware:

[1706] Smartphone (microphone, speaker)

[1707] Earphones (microphone, speaker)

[1708] 2. Software:

[1709] Python

[1710] speech_recognition library (speech recognition)

[1711] googletrans library (translation)

[1712] gTTS library (speech synthesis)

[1713] Explanation of data processing and data calculation

[1714] The server receives voice data input via a smartphone or earphones and converts it into text data using a voice recognition means. This text data is then converted into the target language using a translation means. The translated text is converted back into voice data using a voice synthesis means and transmitted to the user via an output means. In addition, the learning means analyzes user feedback and uses it to improve system performance. During this series of processes, the emotion engine analyzes the user's emotions and reflects them in the translation results.

[1715] Specific examples

[1716] For example, consider the case where a delivery person says, "This is a delivery item." When the delivery person says, "I've brought your order" in English, this voice is captured by the smartphone's microphone. The speech recognition means converts this into text data, obtaining the string "I've brought your order." The translation means then translates this into Japanese, generating the text data "This is a delivery item." The speech synthesis means converts this Japanese text into speech data, which is finally conveyed to the customer via the smartphone or earphone speaker.

[1717] Example prompts to input to the generative AI model

[1718] Create a real-time translation assistant to help people who speak different languages ​​communicate smoothly. When a delivery person says "Delivery here" to a customer, it needs to translate from English to Japanese. This translation assistant will be realized using speech recognition, text translation, and speech synthesis.

[1719] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1720] Step 1:

[1721] Input and Output

[1722] A user starts the smartphone app and speaks into the microphone. The input is the user's voice. The output is recorded as audio data on the smartphone.

[1723] Specific actions

[1724] The user launches the smartphone app and says, "I've brought your order" in English. The smartphone's microphone captures the voice and generates voice data.

[1725] Step 2:

[1726] Input and Output

[1727] The voice data is sent to the voice recognition means. At this time, the input is the voice data generated in step 1. The output is the text data "I've brought your order" generated by the voice recognition means.

[1728] Specific actions

[1729] The smartphone's voice recognition engine (speech_recognition library) analyzes the voice data and converts it into corresponding text data.

[1730] Step 3:

[1731] Input and Output

[1732] The text data is sent to the translation means. At this time, the input is the text data "I've brought your order" generated in step 2. The output is the text data "Delivery item" in the target language generated by the translation means.

[1733] Specific actions

[1734] The smartphone's Google Trans library translates the text data into Japanese.

[1735] Step 4:

[1736] Input and Output

[1737] The text data in the target language is sent to the speech synthesis means. At this time, the input is the text data "Delivery item" generated in step 3. The output is speech data generated by the speech synthesis means.

[1738] Specific actions

[1739] The smartphone's gTTS library converts the text data into audio data and generates an audio file.

[1740] Step 5:

[1741] Input and Output

[1742] The audio data is sent to an output means, where the input is the audio data generated in step 4. The output is audio that can be heard by the user.

[1743] Specific actions

[1744] The translated voice message "Delivery here" is played to the customer through the speaker on their smartphone or earphones.

[1745] Step 6:

[1746] Input and Output

[1747] Feedback from the user is sent to the learning means. At this time, the input is the feedback on translation quality provided by the user. The output is the feedback data stored on the server and becomes learning data.

[1748] Specific actions

[1749] Users provide feedback on the quality of the translation through the app, which is then sent to the server, which analyzes the data and stores it in the translation engine's training database.

[1750] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1751] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1752] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1753] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1754] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1755] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1756] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1757] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1758] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1759] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1760] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1761] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1762] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1763] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1764] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1765] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1766] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1767] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1768] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1769] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1770] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1771] The following is further disclosed regarding the above embodiment.

[1772] (Claim 1)

[1773] an input means for inputting voice;

[1774] a voice recognition means for converting voice data input by the input means into text data;

[1775] a translation means for translating the text data generated by the speech recognition means into a specified target language;

[1776] a speech synthesis means for converting the translation text data generated by the translation means into speech data;

[1777] output means for outputting the voice data generated by the voice synthesis means;

[1778] a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means;

[1779] A system including:

[1780] (Claim 2)

[1781] 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

[1782] (Claim 3)

[1783] 2. The system according to claim 1, wherein the output means is a speaker built into the earphone.

[1784] "Example 1"

[1785] (Claim 1)

[1786] a connection means for automatically connecting with a user's device;

[1787] an input means for inputting voice;

[1788] a voice recognition means for converting voice data input by the input means into text data;

[1789] a translation means for translating the text data generated by the speech recognition means into a specified target language;

[1790] a speech synthesis means for converting the translation text data generated by the translation means into speech data;

[1791] output means for outputting the voice data generated by the voice synthesis means;

[1792] a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means;

[1793] A system including:

[1794] (Claim 2)

[1795] 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

[1796] (Claim 3)

[1797] 2. The system according to claim 1, wherein the output means is a speaker built into the earphone.

[1798] "Application Example 1"

[1799] (Claim 1)

[1800] an input means for inputting voice;

[1801] a voice recognition means for converting voice data input by the input means into text data;

[1802] a translation means for translating the text data generated by the speech recognition means into a specified target language;

[1803] a speech synthesis means for converting the translation text data generated by the translation means into speech data;

[1804] output means for outputting the voice data generated by the voice synthesis means;

[1805] a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means;

[1806] The methods used to ensure efficient communication among multinational staff at logistics centers;

[1807] A system including:

[1808] (Claim 2)

[1809] 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

[1810] (Claim 3)

[1811] 2. The system according to claim 1, wherein the output means is a speaker built into the earphone.

[1812] "Example 2: Combining Emotion Engines"

[1813] (Claim 1)

[1814] an input means for inputting voice;

[1815] a voice recognition means for converting voice data input by the input means into text data;

[1816] emotion analysis means for analyzing voice tone, speech rate, and language patterns to recognize the emotion of the user;

[1817] a translation means for translating the text data and emotion data generated by the speech recognition means and the emotion analysis means into a specified target language;

[1818] a speech synthesis means for converting the translation text data generated by the translation means into speech data;

[1819] output means for outputting the voice data generated by the voice synthesis means;

[1820] a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means;

[1821] A system including:

[1822] (Claim 2)

[1823] 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

[1824] (Claim 3)

[1825] 2. The system according to claim 1, wherein the output means is a speaker built into the earphone.

[1826] "Application example 2 when combining emotion engines"

[1827] (Claim 1)

[1828] an input means for inputting voice;

[1829] a voice recognition means for converting voice data input by the input means into text data;

[1830] a translation means for translating the text data generated by the speech recognition means into a specified target language;

[1831] a speech synthesis means for converting the translation text data generated by the translation means into speech data;

[1832] output means for outputting the voice data generated by the voice synthesis means;

[1833] a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means;

[1834] an emotion engine that analyzes voice tone, speech rate, and language patterns to recognize emotions;

[1835] A system including:

[1836] (Claim 2)

[1837] 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

[1838] (Claim 3)

[1839] 2. The system according to claim 1, wherein the output means is a speaker built into the earphone. [Explanation of symbols]

[1840] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. an input means for inputting voice; a voice recognition means for converting voice data input by the input means into text data; a translation means for translating the text data generated by the speech recognition means into a specified target language; a speech synthesis means for converting the translation text data generated by the translation means into speech data; output means for outputting the voice data generated by the voice synthesis means; a training means for receiving feedback on translation quality provided by a user and storing the feedback as training data for the translation means; A system including:

2. 2. The system according to claim 1, wherein the input means is a microphone built into the earphone.

3. 2. The system according to claim 1, wherein the output means is a speaker built into the earphone.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A