System
The system addresses language barriers by enabling real-time translation and interactive communication through voice and gesture recognition, facilitating seamless cross-language interactions.
Patent Information
- Application Number
- JP2024116458
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Existing translation applications fail to facilitate intimate communication across language barriers and lack real-time collaboration through screen sharing and interactive operations.
A system incorporating voice input, recognition, translation, synthesis, screen sharing, and gesture recognition and transmission to enable real-time language conversion and interactive communication.
Facilitates efficient and intimate cross-language communication by allowing real-time translation, screen sharing, and interactive gestures, overcoming language barriers.
Smart Images

Figure 2026014984000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With the increase in international exchange and travel in modern times, the demand for fast and accurate translation services is growing. However, existing translation applications are limited to mechanical translation and do not adequately promote intimate communication between users. As a result, issues remain, not only the language barrier but also the emotional distance between users. Furthermore, there is a lack of ways for users to collaborate in real time through screen sharing or interactive operation. The purpose of this invention is to solve these issues and provide a system that enables intimate communication that overcomes language barriers. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means: A system includes a voice input means for allowing a user to input voice data, and a voice recognition means for converting the voice input data into text data. By incorporating a translation means for translating the text data into another language, real-time language conversion is achieved. Furthermore, a voice synthesis means is provided for converting the translated text data into voice data, and a voice output means is included for outputting the synthesized voice data. The system also includes a screen sharing means for sharing a screen between users, and a display transmission means for sharing the display of the currently used terminal with the other party's terminal. A system including a gesture recognition means for recognizing touch gestures on the screen and a gesture transmission means for transmitting recognized gesture information to the other party's terminal can promote intimate communication not only through words but also through interactive elements. In this way, the present invention provides a means for making international communication efficient and intimate.
[0006] "Voice input means" refers to a device or software that converts the user's voice into electronic data and inputs it into the system.
[0007] "Speech recognition means" refers to a device or software that analyzes input voice data and converts it into corresponding text data.
[0008] "Translation means" refers to a device or software that converts recognized character data into another language.
[0009] "Speech synthesis means" refers to a device or software that converts translated text data into speech data.
[0010] The "audio output means" refers to a device or software for outputting synthesized audio data to the outside.
[0011] A "screen sharing means" is a device or software for sharing the display content of one terminal with other terminals.
[0012] The "display transmission means" refers to a device or software for transmitting screen data from a terminal to another terminal in real time.
[0013] A "gesture recognition means" is a device or software that detects touch operations on a screen and recognizes their shape and movement.
[0014] The "gesture transmitting means" is a device or software for transmitting recognized gesture data to another terminal.
[0015] The "gesture display means" is a device or software for displaying the received gesture data on a screen. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system that provides a function for a user to input voice, translate it into another language in real time, and output it as voice, as well as a system that allows users to share screens and use interactive touch gestures.
[0038] 1. System Overview
[0039] This system includes a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, and a gesture display means.
[0040] 2. Program Processing Overview
[0041] Voice input
[0042] When a user launches an application on a smartphone or tablet, a "Speak" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0043] Speech Recognition and Translation
[0044] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then converted into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[0045] Speech synthesis and output
[0046] The translated text data is converted back into audio data using a speech synthesis means. This audio data is sent to the terminal and played back by a speech output means. User B hears the audio in his or her own language, so he or she can understand what User A is saying.
[0047] screen sharing
[0048] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[0049] Gesture Recognition
[0050] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0051] Specific examples
[0052] Example 1: Communication at tourist spots
[0053] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0054] Example 2: Using an interactive tourist guide
[0055] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checkmarks the places they want to visit. The checkmarks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information across language barriers and communicate smoothly.
[0056] This system effectively removes language barriers in international exchanges and travel, enabling intimate communication.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user launches the "AI translation app" on their smartphone.
[0060] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[0061] Step 2:
[0062] The user taps the "Speak" button.
[0063] The terminal activates the voice input means and starts collecting voice data.
[0064] Step 3:
[0065] The user begins speaking.
[0066] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[0067] Step 4:
[0068] The server passes the received voice data to the voice recognition means.
[0069] A voice recognition means converts the voice data into text data.
[0070] Step 5:
[0071] The server passes the converted text data to the translation means.
[0072] A translation means translates the text data into another designated language.
[0073] Step 6:
[0074] The server passes the translated text data to the speech synthesis means.
[0075] A speech synthesis means converts the translated text into speech data.
[0076] Step 7:
[0077] The server transmits the generated voice data to User B's terminal.
[0078] User B's terminal receives the audio data and plays it back via the audio output means.
[0079] Step 8:
[0080] User B initiates reply.
[0081] User B's device collects the voice data and sends it to the server.
[0082] Step 9:
[0083] The server passes user B's voice data to a voice recognition means and converts it into text data.
[0084] The server translates the text data into User A's language.
[0085] Step 10:
[0086] The server converts the translated text data into audio data and sends it to User A's terminal.
[0087] User A's terminal receives the audio data and plays it back via the audio output means.
[0088] Step 11:
[0089] The user taps the "Screen Share" button.
[0090] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[0091] Step 12:
[0092] The device sends the screen capture data to the server.
[0093] The server transfers the received screen capture data to User B's terminal.
[0094] Step 13:
[0095] The terminal of user B displays the received screen capture data.
[0096] User A and User B share the same screen in real time.
[0097] Step 14:
[0098] User A performs a touch gesture on the screen.
[0099] The device detects the touch gesture and obtains its position and shape.
[0100] Step 15:
[0101] The terminal transmits the acquired gesture data to the server.
[0102] The server transfers the gesture data to User B's device.
[0103] Step 16:
[0104] The gesture data received by the terminal of User B is displayed on the screen in real time.
[0105] User A and User B can communicate interactively.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] There is a need for a system that can help users who speak different languages communicate smoothly in real time and achieve more effective communication by providing additional visual information. In particular, it is a challenge to achieve smooth communication across language barriers by simultaneously performing real-time speech translation, interactive screen sharing, and gesture recognition.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes a speech recognition unit, a translation unit, and a speech synthesis unit. This allows a user to input speech, have it translated into another language in real time, and then output it as speech. Furthermore, by utilizing a gesture recognition unit, a gesture display unit, and a screen sharing unit, it is possible for users to share screens and display interactive touch gestures. These units allow users to effectively share information across language barriers and realize intimate communication.
[0111] The "voice input means" is a means for collecting the voice uttered by the user as digital data.
[0112] "Speech recognition means" refers to means for converting voice data into character data.
[0113] A "translation means" is a means for converting character data from one language to another.
[0114] The "voice synthesis means" is a means for converting character data into voice data.
[0115] The "audio output means" is a means for outputting the generated audio data through a speaker, headphones, or the like.
[0116] A "screen sharing means" is a means for sharing the contents of a display between users in real time.
[0117] The "display transmission means" is a means for transmitting the display contents of the information display device in use to the information display device of the other party.
[0118] The "gesture recognition means" is a means for detecting an operation gesture made on the display and recognizing the information.
[0119] The "gesture transmitting means" is a means for transmitting the recognized gesture information to the information display device of the other party of the conversation.
[0120] The "buffering means" is a means for dividing the audio data into a certain buffer size, temporarily storing the data, and then transmitting the data to the processing device.
[0121] The "translation output means" is a means for translating character data received by the processing device and reproducing it as voice data.
[0122] MODE FOR CARRYING OUT THE INVENTION
[0123] This invention provides a system that allows users to input speech, translate it into other languages in real time, and output it as speech. It also allows users to share a display and provide interactive gesture control functions. This system includes a speech input unit, a speech recognition unit, a translation unit, a speech synthesis unit, a speech output unit, a screen sharing unit, a display transmission unit, a gesture recognition unit, a gesture transmission unit, and a gesture display unit.
[0124] System configuration
[0125] The system operates via a user's smartphone, tablet, or other information display device, which includes a microphone for collecting voice, a speaker for outputting voice, a display, and a touchscreen. A cloud-based server handles speech recognition, translation, and speech synthesis, using speech recognition software (e.g., Google Cloud Speech-to-Text API), translation software (e.g., Google Translate API), and speech synthesis software (e.g., Amazon Polly).
[0126] Program processing overview
[0127] Voice input
[0128] When a user launches an AI translation app on their smartphone or tablet, a "Speak" button appears on the home screen. When the user taps this button, the device activates the voice input method and begins collecting the user's speech.
[0129] Buffering and sending audio data
[0130] The collected voice data is buffered in the terminal, and when a certain buffer size is reached, the terminal transmits the voice data to the server. For this reason, a buffering means is installed.
[0131] Speech recognition and conversion to text data
[0132] The server converts the received voice data into text data using speech recognition software such as the Google Cloud Speech-to-Text API.
[0133] translation
[0134] The text data generated by speech recognition is translated into other languages using translation means on the server. Here, the Google Translate API is used. For example, if User A utters "Hello, where am I?", this is translated into French as "Bonjour, où sommes-nous?"
[0135] Speech synthesis and playback
[0136] The translated text data is converted back into voice data using a voice synthesis means, such as voice synthesis software like Amazon Polly, and the generated voice data is sent to the terminal and played back by the voice output means.
[0137] Specific examples
[0138] Use in tourist areas
[0139] For example, consider the case where User A (Japanese) is visiting a tourist spot in France. User A launches the "AI translation app," taps the "Speak" button, and asks the French guide (User B), "Hello, where is this?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0140] Interactive tourist guide
[0141] When User A and User B are making sightseeing plans while looking at a tourist guide map, User A taps the "Screen Share" button to share the map, and their device starts capturing the screen. The captured screen data is sent to User B's device via the server. Furthermore, when User A indicates places they want to visit using touch gestures, the gesture data is also displayed in real time on User B's device via the server. This allows them to make plans while sharing information across language barriers.
[0142] Prompt Sentence Examples
[0143] A description of this system can be generated by inputting the following prompt sentence into the generative AI model:
[0144] This system allows users to input speech, translate it into other languages in real time, and output it as speech, as well as provide screen sharing and interactive gesture control functions between users. Please explain the speech input, speech recognition, translation, speech synthesis, and gesture recognition and sharing functions.
[0145] In this way, this system can effectively remove language barriers in international exchanges and travel, enabling intimate communication.
[0146] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0147] Step 1:
[0148] A user launches an "AI translation app" on an information display device (smartphone or tablet) and taps the "Speak" button displayed on the home screen. This activates the device's voice input means. The input is the user's speech (voice data), and the collected voice data is obtained as output.
[0149] Step 2:
[0150] The device starts collecting the user's speech through the microphone. The collected speech data is buffered within the device, and when a certain buffer size is reached, the buffered speech data is sent to the server. The input is the raw speech data of the user's speech, and the output is the buffered speech data sent to the server.
[0151] Step 3:
[0152] The server processes the received voice data and converts it into text data using a speech recognition tool. The input is the voice data sent from the device, and the output is text data. In this case, speech recognition software (e.g., Google Cloud Speech-to-Text API) is used.
[0153] Step 4:
[0154] The server translates the voice-recognized text data into another language using a translation tool. The input is the text data processed by the server, and the output is the text data translated into another language. Here, translation software (e.g., Google Translate API) is used.
[0155] Step 5:
[0156] The server converts the translated text data back into voice data using a speech synthesis tool. The input is the translated text data, and the output is synthesized voice data. In this case, speech synthesis software (e.g., Amazon Polly) is used.
[0157] Step 6:
[0158] The server sends synthesized voice data to the terminal. The input is the voice data from the server, and the terminal receives it as output.
[0159] Step 7:
[0160] The terminal plays back the received voice data using the voice output means. The input is the voice data received from the server, and the output is the voice played back through the speaker. This allows user B to hear user A's speech in his own language.
[0161] Step 8:
[0162] The user taps the "Screen Share" button, and the device starts capturing the screen. The input is the user's screen, and the output is the captured screen data.
[0163] Step 9:
[0164] The terminal sends the captured screen data to the server. The input is the captured screen data, and the output is the data sent to the server.
[0165] Step 10:
[0166] The server transfers the received screen data to the other party's terminal. The input is the transmitted screen data, and the output is the data transmitted to the other party's terminal.
[0167] Step 11:
[0168] The screen data received by the other party's device is displayed. The input is the received screen data, and the output is the screen displayed on the display. This allows user B to see user A's screen in real time.
[0169] Step 12:
[0170] The user performs a touch gesture (e.g., drawing a heart symbol) on the screen. The input is the user's touch operation, and the output is gesture data.
[0171] Step 13:
[0172] The device recognizes the touch gesture and sends the data to the server. The input is the touch gesture operation data, and the output is the recognized gesture data.
[0173] Step 14:
[0174] The server transmits the received gesture data to the other device. The input is the recognized gesture data, and the output is the data transmitted to the other device.
[0175] Step 15:
[0176] The gesture data received by the other party's device is displayed in real time. The input is the received gesture data, and the output is the gesture displayed on the display. This allows the gesture drawn by User A to be displayed on User B's screen as well.
[0177] (Application example 1)
[0178] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0179] Currently, language barriers are a major obstacle in situations requiring real-time multilingual communication. Language differences in particular hinder smooth interactions between viewers and streamers during live streaming, as well as communication between viewers themselves. Furthermore, screen sharing and interactive gesture features are not fully integrated. This lack of an environment allows users of different languages to smoothly communicate with each other and enjoy live content with a sense of unity.
[0180] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0181] In this invention, the server includes a voice input means for allowing a user to input voice, a voice recognition means for converting the voice input data into character data, and a translation means for translating the character data into another language, thereby enabling multilingual live chat.
[0182] "Voice input means" refers to a device or technology that allows a user to input voice.
[0183] "Speech recognition means" refers to technology that analyzes input voice data and converts it into text data.
[0184] "Translation means" refers to a technology that converts text data written in one language into text data written in another language.
[0185] "Speech synthesis means" refers to technology that analyzes text data and converts it into voice data.
[0186] "Audio output means" refers to technology for playing audio data through a device such as a speaker.
[0187] "Screen sharing means" refers to a technology that allows users to share the same screen in real time.
[0188] "Display transmission means" refers to a technology for transmitting the display contents of a terminal in use to another terminal.
[0189] "Gesture recognition means" refers to technology that detects touch gestures made on the screen and recognizes their content.
[0190] The "gesture transmission means" refers to a technique for transmitting recognized gesture information to another terminal.
[0191] "Multilingual live chat means" refers to technology that enables viewers to communicate in different languages in real time during live streaming.
[0192] "Translation display means" refers to technology that translates information sent using the multilingual live chat means and displays it as audio or text.
[0193] The present invention provides a system for supporting smooth communication between users, particularly in situations where real-time communication in multiple languages is required. Specific embodiments of this system will be described below.
[0194] System configuration
[0195] This system is composed of a server, a user terminal, and various software modules. Specific components include the following:
[0196] Voice input means
[0197] Voice recognition means
[0198] Translation tools
[0199] Voice synthesis means
[0200] Audio output means
[0201] Screen sharing method
[0202] Display and transmission means
[0203] Gesture Recognition Method
[0204] Gesture sending method
[0205] Multilingual live chat options
[0206] Translation display means
[0207] System Operation
[0208] 1. Voice input and recognition:
[0209] A user inputs voice using a device such as a smartphone or tablet. The voice input means collects voice data and transmits it to the voice recognition means.
[0210] The speech recognition means converts the collected speech data into text data, which is processed in real time.
[0211] 2. Translation and speech synthesis:
[0212] The translation means converts the text data into a language specified by the user. For example, if a translation from Japanese to English is required, the translation means is applied.
[0213] The translated text data is converted back into voice data by the voice synthesis means, and the voice data is reproduced through the voice output means.
[0214] 3. Screen sharing and interactive gestures:
[0215] The user can use the screen sharing means to share the screen with the other party in real time, and the display transmission means transmits the display content of the terminal currently in use to the other party's terminal.
[0216] The gesture recognition means detects touch gestures made on the screen and transmits the information to the other party's terminal. This is handled by the gesture transmission means.
[0217] 4. Multilingual Live Chat:
[0218] The multilingual live chat means supports real-time communication between viewers during live streaming. Information sent by viewers is translated through the translation display means and displayed to the other person. This allows smooth communication even between viewers who speak different languages.
[0219] Hardware and software used
[0220] Hardware:
[0221] Server: Responsible for data processing, translation, etc.
[0222] User devices: smartphones, tablets, head-mounted displays, etc.
[0223] software:
[0224] Flask: Used as a web server
[0225] SpeechRecognition: Used for speech recognition
[0226] Googletrans: used for text translation
[0227] Pyttsx3: Used for speech synthesis
[0228] Specific example explanation
[0229] For example, imagine a situation where a question can be sent in Japanese during a live streaming event, and the question can be translated into English in real time and spoken to the streamer. In this way, viewers and streamers can communicate smoothly even if they speak different languages.
[0230] Example prompt sentence:
[0231] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0233] Step 1:
[0234] A user starts an application on a device such as a smartphone or tablet and inputs voice. The voice input means collects this voice data and buffers it in real time. The input is the user's voice, and it is temporarily stored in the device as voice data.
[0235] Step 2:
[0236] Voice data is sent to the server at fixed buffer intervals. The server passes this voice data to a voice recognition means and converts it into text data. The input is voice data and the output is text data. The server analyzes the voice data and outputs it as text data.
[0237] Step 3:
[0238] The server passes the recognized character data to the translation means and translates it into the specified language. The input is character data, and the output is translated character data. The server processes the character data to convert it into another language.
[0239] Step 4:
[0240] The server passes the translated text data to a speech synthesis means, which converts it into new speech data. The input is the translated text data, and the output is the new speech data. The speech synthesis means uses a specific speech synthesis algorithm to output the text data as speech data.
[0241] Step 5:
[0242] New audio data is sent from the server to the user's terminal and played through the audio output means. The input is the audio data and the output is the audio to be played. The terminal receives this audio data and outputs it as audio through a speaker.
[0243] Step 6:
[0244] When a user wishes to share their screen, the terminal starts capturing the screen using the screen sharing means and sends the captured data to the server. The input is the screen capture data, and the output is the screen data sent to the server. The server then relays this screen data to the terminals of other users.
[0245] Step 7:
[0246] The other user's device displays the screen data received from the server. The input is the received screen data, and the output is the screen display. The device analyzes the received data and displays it to the user.
[0247] Step 8:
[0248] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes its shape and movement. The input is touch gesture data, and the output is recognized gesture information. The device detects the touch operation and sends it to the server as gesture data.
[0249] Step 9:
[0250] The server transmits the recognized gesture information to the terminal of the other user and displays it in real time using a gesture transmission means. The input is the recognized gesture information and the output is the gesture to be displayed. The terminal displays the received gesture information on the screen.
[0251] Step 10:
[0252] A multilingual live chat means performs processing to support real-time communication between viewers. A translation display means translates messages sent by viewers in different languages and displays them on the terminal of the person interacting with them. The input is the viewer's message, and the output is a display of the translated message. The terminal presents this translated information to the viewer, enabling real-time communication.
[0253] When using a generative AI model for speech recognition or translation, the following example prompt can be set:
[0254] Example prompt sentence:
[0255] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[0256] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0257] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[0258] 1. System Overview
[0259] The system has the following elements:
[0260] Voice input means
[0261] Voice recognition means
[0262] Translation tools
[0263] Voice synthesis means
[0264] Audio output means
[0265] Screen sharing method
[0266] Display and transmission means
[0267] Gesture Recognition Method
[0268] Gesture sending method
[0269] Emotion Engine
[0270] Means of transmitting emotions
[0271] 2. Program Processing Overview
[0272] Voice input
[0273] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0274] Speech Recognition and Translation
[0275] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then translated into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[0276] Emotion recognition and reflection
[0277] In parallel with the voice input, the emotion engine analyzes the user's voice data and recognizes emotions. The recognized emotion information is sent to the server, and the translation means and speech synthesis means reflect this in the translation results and speech tone. For example, if the user's voice contains anger, the translated voice will also be adjusted to reflect an angry tone.
[0278] Speech synthesis and output
[0279] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the other party to hear and understand the translated voice.
[0280] Sharing emotional information
[0281] The emotional information recognized by the emotion engine is also sent to the other party's device via the emotion transmission means. The other party can visually check the other party's emotional state on the screen, which promotes understanding of emotions beyond language.
[0282] screen sharing
[0283] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[0284] Gesture Recognition
[0285] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0286] Specific examples
[0287] Example 1: Communication at tourist spots
[0288] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data and emotion data are sent to the server, where they are translated into French with the emotion reflected, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0289] Example 2: Using an interactive tourist guide
[0290] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checks off the places they want to visit. The check marks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information and communicate smoothly across language barriers. The emotion engine recognizes the user's emotions, and the emotional state is reflected in the displayed interface, allowing for a deeper understanding of each other's feelings.
[0291] This system effectively removes language barriers in international exchange and travel, enabling intimate communication that also includes emotions.
[0292] The processing flow will be explained below.
[0293] Step 1:
[0294] The user launches the "AI translation app" on their smartphone.
[0295] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[0296] Step 2:
[0297] The user taps the "Speak" button.
[0298] The terminal activates the voice input means and starts collecting voice data.
[0299] Step 3:
[0300] The user begins speaking.
[0301] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[0302] Step 4:
[0303] The server passes the received voice data to the voice recognition means.
[0304] The voice recognition means converts the voice data into character data.
[0305] Step 5:
[0306] The server passes the converted character data to the translation means.
[0307] A translation means translates the character data into another designated language.
[0308] Step 6:
[0309] The server passes the translated character data to the speech synthesis means.
[0310] A speech synthesis means converts the translated text into speech data.
[0311] Step 7:
[0312] The server transmits the generated voice data to User B's terminal.
[0313] User B's terminal receives the audio data and plays it back via the audio output means.
[0314] Step 8:
[0315] User B initiates reply.
[0316] User B's device collects the voice data and sends it to the server.
[0317] Step 9:
[0318] The server passes user B's voice data to a voice recognition means and converts it into text data.
[0319] The server translates the text data into User A's language.
[0320] Step 10:
[0321] The server converts the translated text data into audio data and sends it to User A's terminal.
[0322] User A's terminal receives the audio data and plays it back via the audio output means.
[0323] Step 11:
[0324] The user taps the "Screen Share" button.
[0325] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[0326] Step 12:
[0327] The device sends the screen capture data to the server.
[0328] The server transfers the received screen capture data to User B's terminal.
[0329] Step 13:
[0330] The terminal of user B displays the received screen capture data.
[0331] User A and User B share the same screen in real time.
[0332] Step 14:
[0333] User A performs a touch gesture on the screen.
[0334] The device detects the touch gesture and obtains its position and shape.
[0335] Step 15:
[0336] The terminal transmits the acquired gesture data to the server.
[0337] The server transfers the gesture data to User B's device.
[0338] Step 16:
[0339] The gesture data received by the terminal of User B is displayed on the screen in real time.
[0340] User A and User B can communicate interactively.
[0341] Step 17:
[0342] The terminal transmits the user's voice data to the emotion engine.
[0343] The emotion engine analyzes the voice data and recognizes the user's emotions.
[0344] Step 18:
[0345] The emotion engine transmits the recognized emotion information to the server.
[0346] The server reflects the emotional information in the translation means and the voice synthesis means.
[0347] Step 19:
[0348] The server converts the translated text data and emotional information into voice data.
[0349] The voice synthesis means generates voice data that reflects the emotion and transmits it to the user B's terminal.
[0350] Step 20:
[0351] The terminal of user B receives the voice data reflecting the emotion and plays it back via the voice output means.
[0352] User B can understand User A's emotions.
[0353] Step 21:
[0354] The server transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[0355] The emotion information received by the terminal of user B is visually displayed on the screen.
[0356] Example 2
[0357] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0358] While conventional speech translation systems provide the ability to translate speech into other languages, they lack the ability to recognize user emotions and reflect them in translation results and speech output. This makes it difficult for users who speak different languages to convey emotional nuances, making smooth communication difficult. Furthermore, screen sharing and interactive gesture functions are limited, making it difficult for users to intuitively share information and communicate with each other.
[0359] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0360] In this invention, the server includes a speech recognition unit, a translation unit, a speech synthesis unit, an emotion recognition unit, and an emotion transmission unit. This makes it possible to recognize emotions in a user's speech data in real time and reflect those emotions in the translation results and voice output. Furthermore, by combining a screen sharing unit and a gesture recognition unit, intuitive information sharing and interactive communication between users is realized.
[0361] The "voice input means" is a device for inputting the voice spoken by the user as digital data.
[0362] "Speech recognition means" is a technology that analyzes input voice data and converts it into text data.
[0363] "Translation means" is a technology that converts text data in one language into another language.
[0364] "Speech synthesis means" is a technology that converts text data into voice data and outputs it as synthesized voice.
[0365] The "audio output means" is a device for reproducing the synthesized audio data and letting the user hear it.
[0366] "Screen sharing means" is a technology that allows multiple users to view the same screen content at the same time.
[0367] The "display transmission means" is a technique for transmitting the display contents of the terminal in use to another terminal.
[0368] A "gesture recognition means" is a technology that detects touch gestures on a screen and recognizes their movements and shapes.
[0369] The "gesture transmitting means" is a technique for transmitting recognized gesture information to another terminal.
[0370] "Emotion recognition means" is a technology that analyzes the user's voice data and identifies emotions.
[0371] "Emotion transmission means" is a technology for transmitting recognized emotion information to another terminal.
[0372] The "gesture display means" is a technique for displaying gesture information received by the terminal of the other party on its screen.
[0373] A "buffer" is an area that temporarily stores data and is used to adjust speed differences when sending and receiving data.
[0374] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[0375] System configuration
[0376] The system has the following elements:
[0377] Voice input means
[0378] Voice recognition means
[0379] Translation tools
[0380] Voice synthesis means
[0381] Audio output means
[0382] Screen sharing method
[0383] Display and transmission means
[0384] Gesture Recognition Method
[0385] Gesture sending method
[0386] emotion recognition means
[0387] Means of transmitting emotions
[0388] Detailed Description
[0389] Voice input means
[0390] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0391] Voice recognition means
[0392] When the voice data arrives at the server, the server uses speech recognition software like the Google Cloud Speech-to-Text API to convert the speech into text. For example, a user might say, "Hello, where am I?" and it will be converted into text.
[0393] Translation tools
[0394] The server translates the obtained text data into other languages using a translation tool such as Google Translate API. For example, the Japanese phrase "Hello, where are you?" is translated into French as "Bonjour, où sommes-nous?"
[0395] Voice synthesis means
[0396] The translated text data is converted back into speech data using a speech synthesis tool (e.g., Amazon Polly), for example, to generate the translated French phrase "Bonjour, où sommes-nous?"
[0397] Audio output means
[0398] The voice data is sent to the other party's terminal and played back by the voice output means, allowing the other party to hear and understand the translated voice.
[0399] emotion recognition means
[0400] In parallel with the voice input, an emotion recognition means on the server (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotions. The recognized emotional information is reflected in the translation results and voice tone by the translation means and voice synthesis means. For example, if the user's voice contains anger, the translation results will also be adjusted to reflect an angry tone.
[0401] Means of transmitting emotions
[0402] The recognized emotional information is also sent to the other party's device via the emotion transmission means, allowing the other party to visually confirm the other party's emotional state on the screen, facilitating understanding of emotions beyond language.
[0403] Screen sharing and display transmission means
[0404] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This allows users to see each other's screens.
[0405] Gesture Recognition Means and Gesture Transmission Means
[0406] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes the shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if a user draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0407] Specific examples
[0408] Example 1: Communication at tourist spots
[0409] User A (Japanese) is visiting a tourist spot in a foreign country. He wants to communicate with a local guide (User B), but there is a language barrier. User A uses this system, taps the "Speak" button, and asks, "Hello, where am I?" This voice data and emotion data are sent to the server, where it is translated and the emotion is reflected in real time and sent to the guide's device. The guide listens to the voice and provides appropriate guidance.
[0410] Prompt Sentence Examples
[0411] "Please tell me how the voice input system works to communicate with foreigners at tourist spots. It also performs emotion recognition."
[0412] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0413] Step 1:
[0414] When a user launches an application on their smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. The user's voice data is used as input, and the voice data is saved in a buffer as output. Specifically, the device's microphone captures the user's voice and stores the data in a buffer.
[0415] Step 2:
[0416] The device sends audio data to the server in fixed buffers. The audio data in the buffer is used as input, and the audio data is sent to the server as output. Specifically, when the collected audio data reaches a certain volume, the device uploads the data to the server.
[0417] Step 3:
[0418] The server analyzes the received voice data and converts the voice into text data using a voice recognition tool. The voice data arriving at the server is used as input, and converted text data is generated as output. Specifically, the server calls a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data.
[0419] Step 4:
[0420] The server translates the obtained text data into another language. The text data generated by the speech recognition means is used as input, and the translated text data is generated as output. Specifically, the server calls a translation API (e.g., Google Translate) to translate from Japanese to French, etc.
[0421] Step 5:
[0422] The server's emotion recognition means analyzes the user's voice data and recognizes emotions. The received voice data is used as input, and recognized emotion information is generated as output. Specifically, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the voice data.
[0423] Step 6:
[0424] The recognized emotional information is reflected in the translation results and voice tone through the translation means and speech synthesis means. The emotional information generated by the emotion recognition means and the translated text data are used as input, and synthetic speech data reflecting the emotion is generated as output. Specifically, the server uses a speech synthesis API (e.g., Amazon Polly) to generate speech data based on the text data and emotional information.
[0425] Step 7:
[0426] The synthesized voice data is sent to the other party's terminal and played by the voice output means. The synthesized voice data is used as input, and the voice is played back in an audible form for the user as output. Specifically, the server sends the synthesized voice data to the other party's terminal, and the terminal plays the voice through a speaker.
[0427] Step 8:
[0428] The recognized emotional information is also sent to the conversation partner's terminal through the emotion transmission means. The emotional information is used as input, and the emotional state is visually displayed on the screen of the conversation partner's terminal as output. Specifically, the server transmits the emotional information to the conversation partner's terminal, and the emotional state is displayed on the terminal's display.
[0429] Step 9:
[0430] When a user taps the "Screen Share" button, the device starts capturing the screen. User operations and screen data are used as input, and screen capture data is generated as output. Specifically, the device periodically takes screenshots and sends the data to the server.
[0431] Step 10:
[0432] The screen data is sent to the other party's device via the server and displayed in real time. The screen capture data sent from the device is used as input, and the screen displayed on the other party's device is updated in real time as output. Specifically, the server relays the screen data to the other party's device, and the device displays the data on its display.
[0433] Step 11:
[0434] When a user makes a touch gesture during screen sharing, the gesture recognition means of the device detects the touch operation. Touch gesture data is used as input, and recognized gesture data is generated as output. Specifically, the device detects touch operations on the screen with a sensor and recognizes their shape and movement.
[0435] Step 12:
[0436] The gesture data is sent to the other party's device via the server and displayed in real time. The recognized gesture data is used as input to generate the gestures displayed on the other party's device as output. Specifically, the server relays the gesture data to the other party's device, and the device displays the data on its display.
[0437] (Application example 2)
[0438] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0439] In communication between workers and robots in factories, there is a problem that instructions cannot be transmitted quickly and accurately between different languages, resulting in reduced work efficiency. Furthermore, it is difficult to improve safety because the robot's behavior cannot be adjusted according to the worker's emotional state. Furthermore, it is difficult to share instructions and information between multiple workers using screen sharing and interactive touch gestures.
[0440] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0441] In this invention, the server includes a voice input device for allowing a user to input voice data, a voice recognition device for converting the voice input data into text data, a translation device for translating the text data into another language, a voice synthesis device for converting the translated text data into voice data, a voice output device for outputting the synthesized voice data, a screen sharing device for sharing a screen between users, a display transmission device for sharing the display of the currently used terminal with the other terminal, a gesture recognition device for recognizing touch gestures on the screen, a gesture transmission device for transmitting the recognized gesture information to the other terminal, an emotion engine for recognizing an emotional state from the voice data, an emotion transmission device for transmitting the recognized emotion information to the other terminal, and an instruction device for instructing the work machine using the translated text data and voice data. This enables quick and accurate instruction transmission between different languages. Furthermore, safety is improved by adjusting the robot's operation according to the worker's emotional state. Furthermore, interactive screen sharing and touch gesture sharing of instructions and information between multiple workers is possible.
[0442] "Voice input means" refers to a device or system that allows a user to input voice.
[0443] The "voice recognition means" is a device or system that converts voice data acquired by the voice input means into character data.
[0444] The "translation means" is a device or system that translates the character data generated by the speech recognition means into another language.
[0445] The "voice synthesis means" is a device or system that converts the translated text data generated by the translation means into voice data.
[0446] The "audio output means" is a device or system that outputs the audio data generated by the audio synthesis means to the user.
[0447] A "screen sharing means" is a device or system for sharing screen information among multiple users.
[0448] The "display transmission means" is a device or system that transmits the display of the terminal in use to the terminal of the other party.
[0449] A "gesture recognition means" is a device or system that recognizes touch gestures on a screen and processes them as data.
[0450] The "gesture transmitting means" is a device or system that transmits recognized gesture information to the terminal of the other party.
[0451] An "emotion engine" is a device or system for recognizing a user's emotional state from voice data.
[0452] The "emotion transmission means" is a device or system that transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[0453] The "instruction means" is a device or system that issues instructions to the work machine based on the translated character data and voice data.
[0454] This invention realizes a multilingual interactive interpretation system between workers and robots in a factory. This system is composed of a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, an emotion engine, an emotion transmission means, and an instruction means.
[0455] System Overview
[0456] Audio Input:
[0457] The server activates a voice input device, such as a smartphone or tablet, for the user to input voice. When the user starts speaking, this voice data is buffered in the device and sent to the server at regular intervals.
[0458] Speech Recognition and Translation:
[0459] The voice data received by the server is converted into text data using a voice recognition unit, and then this text data is translated into another language by a translation unit, for example, from Japanese to English.
[0460] Emotion recognition and reflection:
[0461] In parallel with voice input, the emotion engine analyzes the voice data and recognizes the user's emotional state. The recognized emotional information is sent to the server, where the translation means and voice synthesis means reflect it in the translation results and voice tone. If the user is tired, the robot's movement speed will be slowed down, for example.
[0462] Speech synthesis and output:
[0463] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the robot to understand and carry out the worker's instructions.
[0464] Sharing emotional information:
[0465] The emotion information recognized by the emotion engine is also sent to the terminal of the conversation partner through the emotion transmission means. The emotional state of the worker is displayed, promoting mutual understanding.
[0466] Screen sharing and gesture recognition:
[0467] To share a screen between users, a device initiates a screen capture and sends the captured data to the other device. Furthermore, touch gestures on the screen are recognized and the gesture data is sent to the other device, enabling interactive operation.
[0468] Specific examples
[0469] For example, in a factory, Worker A instructs a robot in Japanese, saying, "Please transport this part." This voice is translated into English in real time and transmitted to the robot. At the same time, the robot recognizes that Worker A is tired from his voice and switches its operating speed to slow mode. Worker B can then use the screen sharing function to instruct the robot on the specific destination with touch gestures, achieving efficient communication.
[0470] Example prompt sentence:
[0471] "Based on the invention below, we will design a multilingual interactive interpretation system for factory robots. It is a system that allows users to input speech and translate it into other languages in real time. Furthermore, it recognizes the user's emotions during the translation process and reflects those emotions in the translation results. It will also have functions that allow users to share screens and use interactive touch gestures. As a specific application example, when workers in a factory give instructions to a robot, it will use voice recognition and translation to make it multilingual. Also, please consider a system that recognizes the emotional state of the worker and adjusts the robot's movement speed and response content according to the situation."
[0472] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0473] Step 1:
[0474] The user activates the voice input means to input voice.
[0475] Input: User speech.
[0476] Output: The input audio data.
[0477] Specific operation: The user taps the "Speak" button on their smartphone or tablet to begin voice input. The device collects voice data through the microphone.
[0478] Step 2:
[0479] The device sends the collected voice data to the server.
[0480] Input: Audio data.
[0481] Output: The audio data sent to the server.
[0482] Specific operation: The collected audio data is sent from the terminal to the server at fixed buffer intervals.
[0483] Step 3:
[0484] The server converts the voice data into character data using a voice recognition means.
[0485] Input: The audio data sent to the server.
[0486] Output: Recognized speech text data.
[0487] Specific operation: The server calls a speech recognition API and converts the voice data into text data, for example, using the Google Speech-to-Text API.
[0488] Step 4:
[0489] The server uses a translation means to translate the character data into another language.
[0490] Input: Speech-recognized text data.
[0491] Output: The translated text data.
[0492] Specific operation: The server calls a translation API to translate the text data into another specified language. For example, it uses the Google Translate API.
[0493] Step 5:
[0494] The server uses an emotion engine to analyze the voice data and recognize the emotional state.
[0495] Input: The original audio data.
[0496] Output: Recognized emotion information.
[0497] Specific operation: Calls an emotion recognition API and recognizes emotions from features such as energy, pitch, and tempo. For example, it uses IBM Watson Tone Analyzer.
[0498] Step 6:
[0499] The server synthesizes voice data based on the translated text data and emotional information.
[0500] Input: translated text data, recognized emotion information.
[0501] Output: Synthesized audio data.
[0502] Specific operation: Calls a speech synthesis API and generates speech data that reflects emotional information. For example, it uses Amazon Polly.
[0503] Step 7:
[0504] The server transmits the synthesized voice data to the terminal.
[0505] Input: Synthesized speech data.
[0506] Output: The audio data sent to the device.
[0507] Specific operation: The server transmits the generated voice data to the other party's terminal.
[0508] Step 8:
[0509] The terminal uses the audio output means to play back the received audio data.
[0510] Input: Received audio data.
[0511] Output: The audio that is played.
[0512] Specific operation: The other party's device plays the synthesized voice through the speaker.
[0513] Step 9:
[0514] The user activates the screen sharing tool and starts capturing the screen.
[0515] Input: User operations, device screen data.
[0516] Output: Real-time captured screen data.
[0517] Specific operation: The user taps the "Screen Share" button and selects the screen they want to share. The device captures that screen in real time and sends it to the server.
[0518] Step 10:
[0519] The server transmits the screen data to the terminal of the other party.
[0520] Input: Screen data captured in real time.
[0521] Output: Screen data sent to the other device.
[0522] Specific operation: The server sends the captured screen data to the other device.
[0523] Step 11:
[0524] The terminal of the other party displays the received screen data.
[0525] Input: Received screen data.
[0526] Output: The screen that is displayed.
[0527] Specific operation: The screen data received by the other party's device is displayed on the screen.
[0528] Step 12:
[0529] A user performs a touch gesture using the gesture recognition means.
[0530] Input: User touch actions.
[0531] Output: Recognized gesture data.
[0532] Specific operation: When a user performs a touch gesture on the screen, the device recognizes the touch operation and generates gesture data.
[0533] Step 13:
[0534] The server transmits the recognized gesture data to the other terminal of the conversation partner.
[0535] Input: Recognized gesture data.
[0536] Output: Gesture data sent to the other device.
[0537] Specific operation: The server sends the recognized gesture data to the other device.
[0538] Step 14:
[0539] The gesture data received by the terminal of the conversation partner is displayed on the screen.
[0540] Input: Gesture data sent to the other device.
[0541] Output: Gesture data displayed on the screen.
[0542] Specific operation: The gesture data received by the other device is displayed on the screen in real time.
[0543] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0544] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0545] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0546] [Second embodiment]
[0547] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0548] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0549] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0550] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0551] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0552] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0553] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0554] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0555] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0556] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0557] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0558] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0559] This invention is a system that provides a function for a user to input voice, translate it into another language in real time, and output it as voice, as well as a system that allows users to share screens and use interactive touch gestures.
[0560] 1. System Overview
[0561] This system includes a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, and a gesture display means.
[0562] 2. Program Processing Overview
[0563] Voice input
[0564] When a user launches an application on a smartphone or tablet, a "Speak" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0565] Speech Recognition and Translation
[0566] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then converted into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[0567] Speech synthesis and output
[0568] The translated text data is converted back into audio data using a speech synthesis means. This audio data is sent to the terminal and played back by a speech output means. User B hears the audio in his or her own language, so he or she can understand what User A is saying.
[0569] screen sharing
[0570] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[0571] Gesture Recognition
[0572] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0573] Specific examples
[0574] Example 1: Communication at tourist spots
[0575] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0576] Example 2: Using an interactive tourist guide
[0577] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checkmarks the places they want to visit. The checkmarks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information across language barriers and communicate smoothly.
[0578] This system effectively removes language barriers in international exchanges and travel, enabling intimate communication.
[0579] The processing flow will be explained below.
[0580] Step 1:
[0581] The user launches the "AI translation app" on their smartphone.
[0582] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[0583] Step 2:
[0584] The user taps the "Speak" button.
[0585] The terminal activates the voice input means and starts collecting voice data.
[0586] Step 3:
[0587] The user begins speaking.
[0588] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[0589] Step 4:
[0590] The server passes the received voice data to the voice recognition means.
[0591] A voice recognition means converts the voice data into text data.
[0592] Step 5:
[0593] The server passes the converted text data to the translation means.
[0594] A translation means translates the text data into another designated language.
[0595] Step 6:
[0596] The server passes the translated text data to the speech synthesis means.
[0597] A speech synthesis means converts the translated text into speech data.
[0598] Step 7:
[0599] The server transmits the generated voice data to User B's terminal.
[0600] User B's terminal receives the audio data and plays it back via the audio output means.
[0601] Step 8:
[0602] User B initiates reply.
[0603] User B's device collects the voice data and sends it to the server.
[0604] Step 9:
[0605] The server passes user B's voice data to a voice recognition means and converts it into text data.
[0606] The server translates the text data into User A's language.
[0607] Step 10:
[0608] The server converts the translated text data into audio data and sends it to User A's terminal.
[0609] User A's terminal receives the audio data and plays it back via the audio output means.
[0610] Step 11:
[0611] The user taps the "Screen Share" button.
[0612] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[0613] Step 12:
[0614] The device sends the screen capture data to the server.
[0615] The server transfers the received screen capture data to User B's terminal.
[0616] Step 13:
[0617] The terminal of user B displays the received screen capture data.
[0618] User A and User B share the same screen in real time.
[0619] Step 14:
[0620] User A performs a touch gesture on the screen.
[0621] The device detects the touch gesture and obtains its position and shape.
[0622] Step 15:
[0623] The terminal transmits the acquired gesture data to the server.
[0624] The server transfers the gesture data to User B's device.
[0625] Step 16:
[0626] The gesture data received by the terminal of User B is displayed on the screen in real time.
[0627] User A and User B can communicate interactively.
[0628] Example 1
[0629] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0630] There is a need for a system that can help users who speak different languages communicate smoothly in real time and achieve more effective communication by providing additional visual information. In particular, it is a challenge to achieve smooth communication across language barriers by simultaneously performing real-time speech translation, interactive screen sharing, and gesture recognition.
[0631] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0632] In this invention, the server includes a speech recognition unit, a translation unit, and a speech synthesis unit. This allows a user to input speech, have it translated into another language in real time, and then output it as speech. Furthermore, by utilizing a gesture recognition unit, a gesture display unit, and a screen sharing unit, it is possible for users to share screens and display interactive touch gestures. These units allow users to effectively share information across language barriers and realize intimate communication.
[0633] The "voice input means" is a means for collecting the voice uttered by the user as digital data.
[0634] "Speech recognition means" refers to means for converting voice data into character data.
[0635] A "translation means" is a means for converting character data from one language to another.
[0636] The "voice synthesis means" is a means for converting character data into voice data.
[0637] The "audio output means" is a means for outputting the generated audio data through a speaker, headphones, or the like.
[0638] A "screen sharing means" is a means for sharing the contents of a display between users in real time.
[0639] The "display transmission means" is a means for transmitting the display contents of the information display device in use to the information display device of the other party.
[0640] The "gesture recognition means" is a means for detecting an operation gesture made on the display and recognizing the information.
[0641] The "gesture transmitting means" is a means for transmitting the recognized gesture information to the information display device of the other party of the conversation.
[0642] The "buffering means" is a means for dividing the audio data into a certain buffer size, temporarily storing the data, and then transmitting the data to the processing device.
[0643] The "translation output means" is a means for translating character data received by the processing device and reproducing it as voice data.
[0644] MODE FOR CARRYING OUT THE INVENTION
[0645] This invention provides a system that allows users to input speech, translate it into other languages in real time, and output it as speech. It also allows users to share a display and provide interactive gesture control functions. This system includes a speech input unit, a speech recognition unit, a translation unit, a speech synthesis unit, a speech output unit, a screen sharing unit, a display transmission unit, a gesture recognition unit, a gesture transmission unit, and a gesture display unit.
[0646] System configuration
[0647] The system operates via a user's smartphone, tablet, or other information display device, which includes a microphone for collecting voice, a speaker for outputting voice, a display, and a touchscreen. A cloud-based server handles speech recognition, translation, and speech synthesis, using speech recognition software (e.g., Google Cloud Speech-to-Text API), translation software (e.g., Google Translate API), and speech synthesis software (e.g., Amazon Polly).
[0648] Program processing overview
[0649] Voice input
[0650] When a user launches an AI translation app on their smartphone or tablet, a "Speak" button appears on the home screen. When the user taps this button, the device activates the voice input method and begins collecting the user's speech.
[0651] Buffering and sending audio data
[0652] The collected voice data is buffered in the terminal, and when a certain buffer size is reached, the terminal transmits the voice data to the server. For this reason, a buffering means is installed.
[0653] Speech recognition and conversion to text data
[0654] The server converts the received voice data into text data using speech recognition software such as the Google Cloud Speech-to-Text API.
[0655] translation
[0656] The text data generated by speech recognition is translated into other languages using translation means on the server. Here, the Google Translate API is used. For example, if User A utters "Hello, where am I?", this is translated into French as "Bonjour, où sommes-nous?"
[0657] Speech synthesis and playback
[0658] The translated text data is converted back into voice data using a voice synthesis means, such as voice synthesis software like Amazon Polly, and the generated voice data is sent to the terminal and played back by the voice output means.
[0659] Specific examples
[0660] Use in tourist areas
[0661] For example, consider the case where User A (Japanese) is visiting a tourist spot in France. User A launches the "AI translation app," taps the "Speak" button, and asks the French guide (User B), "Hello, where is this?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0662] Interactive tourist guide
[0663] When User A and User B are making sightseeing plans while looking at a tourist guide map, User A taps the "Screen Share" button to share the map, and their device starts capturing the screen. The captured screen data is sent to User B's device via the server. Furthermore, when User A indicates places they want to visit using touch gestures, the gesture data is also displayed in real time on User B's device via the server. This allows them to make plans while sharing information across language barriers.
[0664] Prompt Sentence Examples
[0665] A description of this system can be generated by inputting the following prompt sentence into the generative AI model:
[0666] This system allows users to input speech, translate it into other languages in real time, and output it as speech, as well as provide screen sharing and interactive gesture control functions between users. Please explain the speech input, speech recognition, translation, speech synthesis, and gesture recognition and sharing functions.
[0667] In this way, this system can effectively remove language barriers in international exchanges and travel, enabling intimate communication.
[0668] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0669] Step 1:
[0670] A user launches an "AI translation app" on an information display device (smartphone or tablet) and taps the "Speak" button displayed on the home screen. This activates the device's voice input means. The input is the user's speech (voice data), and the collected voice data is obtained as output.
[0671] Step 2:
[0672] The device starts collecting the user's speech through the microphone. The collected speech data is buffered within the device, and when a certain buffer size is reached, the buffered speech data is sent to the server. The input is the raw speech data of the user's speech, and the output is the buffered speech data sent to the server.
[0673] Step 3:
[0674] The server processes the received voice data and converts it into text data using a speech recognition tool. The input is the voice data sent from the device, and the output is text data. In this case, speech recognition software (e.g., Google Cloud Speech-to-Text API) is used.
[0675] Step 4:
[0676] The server translates the voice-recognized text data into another language using a translation tool. The input is the text data processed by the server, and the output is the text data translated into another language. Here, translation software (e.g., Google Translate API) is used.
[0677] Step 5:
[0678] The server converts the translated text data back into voice data using a speech synthesis tool. The input is the translated text data, and the output is synthesized voice data. In this case, speech synthesis software (e.g., Amazon Polly) is used.
[0679] Step 6:
[0680] The server sends synthesized voice data to the terminal. The input is the voice data from the server, and the terminal receives it as output.
[0681] Step 7:
[0682] The terminal plays back the received voice data using the voice output means. The input is the voice data received from the server, and the output is the voice played back through the speaker. This allows user B to hear user A's speech in his own language.
[0683] Step 8:
[0684] The user taps the "Screen Share" button, and the device starts capturing the screen. The input is the user's screen, and the output is the captured screen data.
[0685] Step 9:
[0686] The terminal sends the captured screen data to the server. The input is the captured screen data, and the output is the data sent to the server.
[0687] Step 10:
[0688] The server transfers the received screen data to the other party's terminal. The input is the transmitted screen data, and the output is the data transmitted to the other party's terminal.
[0689] Step 11:
[0690] The screen data received by the other party's device is displayed. The input is the received screen data, and the output is the screen displayed on the display. This allows user B to see user A's screen in real time.
[0691] Step 12:
[0692] The user performs a touch gesture (e.g., drawing a heart symbol) on the screen. The input is the user's touch operation, and the output is gesture data.
[0693] Step 13:
[0694] The device recognizes the touch gesture and sends the data to the server. The input is the touch gesture operation data, and the output is the recognized gesture data.
[0695] Step 14:
[0696] The server transmits the received gesture data to the other device. The input is the recognized gesture data, and the output is the data transmitted to the other device.
[0697] Step 15:
[0698] The gesture data received by the other party's device is displayed in real time. The input is the received gesture data, and the output is the gesture displayed on the display. This allows the gesture drawn by User A to be displayed on User B's screen as well.
[0699] (Application example 1)
[0700] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0701] Currently, language barriers are a major obstacle in situations requiring real-time multilingual communication. Language differences in particular hinder smooth interactions between viewers and streamers during live streaming, as well as communication between viewers themselves. Furthermore, screen sharing and interactive gesture features are not fully integrated. This lack of an environment allows users of different languages to smoothly communicate with each other and enjoy live content with a sense of unity.
[0702] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0703] In this invention, the server includes a voice input means for allowing a user to input voice, a voice recognition means for converting the voice input data into character data, and a translation means for translating the character data into another language, thereby enabling multilingual live chat.
[0704] "Voice input means" refers to a device or technology that allows a user to input voice.
[0705] "Speech recognition means" refers to technology that analyzes input voice data and converts it into text data.
[0706] "Translation means" refers to a technology that converts text data written in one language into text data written in another language.
[0707] "Speech synthesis means" refers to technology that analyzes text data and converts it into voice data.
[0708] "Audio output means" refers to technology for playing audio data through a device such as a speaker.
[0709] "Screen sharing means" refers to a technology that allows users to share the same screen in real time.
[0710] "Display transmission means" refers to a technology for transmitting the display contents of a terminal in use to another terminal.
[0711] "Gesture recognition means" refers to technology that detects touch gestures made on the screen and recognizes their content.
[0712] The "gesture transmission means" refers to a technique for transmitting recognized gesture information to another terminal.
[0713] "Multilingual live chat means" refers to technology that enables viewers to communicate in different languages in real time during live streaming.
[0714] "Translation display means" refers to technology that translates information sent using the multilingual live chat means and displays it as audio or text.
[0715] The present invention provides a system for supporting smooth communication between users, particularly in situations where real-time communication in multiple languages is required. Specific embodiments of this system will be described below.
[0716] System configuration
[0717] This system is composed of a server, a user terminal, and various software modules. Specific components include the following:
[0718] Voice input means
[0719] Voice recognition means
[0720] Translation tools
[0721] Voice synthesis means
[0722] Audio output means
[0723] Screen sharing method
[0724] Display and transmission means
[0725] Gesture Recognition Method
[0726] Gesture sending method
[0727] Multilingual live chat options
[0728] Translation display means
[0729] System Operation
[0730] 1. Voice input and recognition:
[0731] A user inputs voice using a device such as a smartphone or tablet. The voice input means collects voice data and transmits it to the voice recognition means.
[0732] The speech recognition means converts the collected speech data into text data, which is processed in real time.
[0733] 2. Translation and speech synthesis:
[0734] The translation means converts the text data into a language specified by the user. For example, if a translation from Japanese to English is required, the translation means is applied.
[0735] The translated text data is converted back into voice data by the voice synthesis means, and the voice data is reproduced through the voice output means.
[0736] 3. Screen sharing and interactive gestures:
[0737] The user can use the screen sharing means to share the screen with the other party in real time, and the display transmission means transmits the display content of the terminal currently in use to the other party's terminal.
[0738] The gesture recognition means detects touch gestures made on the screen and transmits the information to the other party's terminal. This is handled by the gesture transmission means.
[0739] 4. Multilingual Live Chat:
[0740] The multilingual live chat means supports real-time communication between viewers during live streaming. Information sent by viewers is translated through the translation display means and displayed to the other person. This allows smooth communication even between viewers who speak different languages.
[0741] Hardware and software used
[0742] Hardware:
[0743] Server: Responsible for data processing, translation, etc.
[0744] User devices: smartphones, tablets, head-mounted displays, etc.
[0745] software:
[0746] Flask: Used as a web server
[0747] SpeechRecognition: Used for speech recognition
[0748] Googletrans: used for text translation
[0749] Pyttsx3: Used for speech synthesis
[0750] Specific example explanation
[0751] For example, imagine a situation where a question can be sent in Japanese during a live streaming event, and the question can be translated into English in real time and spoken to the streamer. In this way, viewers and streamers can communicate smoothly even if they speak different languages.
[0752] Example prompt sentence:
[0753] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[0754] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0755] Step 1:
[0756] A user starts an application on a device such as a smartphone or tablet and inputs voice. The voice input means collects this voice data and buffers it in real time. The input is the user's voice, and it is temporarily stored in the device as voice data.
[0757] Step 2:
[0758] Voice data is sent to the server at fixed buffer intervals. The server passes this voice data to a voice recognition means and converts it into text data. The input is voice data and the output is text data. The server analyzes the voice data and outputs it as text data.
[0759] Step 3:
[0760] The server passes the recognized character data to the translation means and translates it into the specified language. The input is character data, and the output is translated character data. The server processes the character data to convert it into another language.
[0761] Step 4:
[0762] The server passes the translated text data to a speech synthesis means, which converts it into new speech data. The input is the translated text data, and the output is the new speech data. The speech synthesis means uses a specific speech synthesis algorithm to output the text data as speech data.
[0763] Step 5:
[0764] New audio data is sent from the server to the user's terminal and played through the audio output means. The input is the audio data and the output is the audio to be played. The terminal receives this audio data and outputs it as audio through a speaker.
[0765] Step 6:
[0766] When a user wishes to share their screen, the terminal starts capturing the screen using the screen sharing means and sends the captured data to the server. The input is the screen capture data, and the output is the screen data sent to the server. The server then relays this screen data to the terminals of other users.
[0767] Step 7:
[0768] The other user's device displays the screen data received from the server. The input is the received screen data, and the output is the screen display. The device analyzes the received data and displays it to the user.
[0769] Step 8:
[0770] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes its shape and movement. The input is touch gesture data, and the output is recognized gesture information. The device detects the touch operation and sends it to the server as gesture data.
[0771] Step 9:
[0772] The server transmits the recognized gesture information to the terminal of the other user and displays it in real time using a gesture transmission means. The input is the recognized gesture information and the output is the gesture to be displayed. The terminal displays the received gesture information on the screen.
[0773] Step 10:
[0774] A multilingual live chat means performs processing to support real-time communication between viewers. A translation display means translates messages sent by viewers in different languages and displays them on the terminal of the person interacting with them. The input is the viewer's message, and the output is a display of the translated message. The terminal presents this translated information to the viewer, enabling real-time communication.
[0775] When using a generative AI model for speech recognition or translation, the following example prompt can be set:
[0776] Example prompt sentence:
[0777] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[0778] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0779] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[0780] 1. System Overview
[0781] The system has the following elements:
[0782] Voice input means
[0783] Voice recognition means
[0784] Translation tools
[0785] Voice synthesis means
[0786] Audio output means
[0787] Screen sharing method
[0788] Display and transmission means
[0789] Gesture Recognition Method
[0790] Gesture sending method
[0791] Emotion Engine
[0792] Means of transmitting emotions
[0793] 2. Program Processing Overview
[0794] Voice input
[0795] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0796] Speech Recognition and Translation
[0797] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then translated into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[0798] Emotion recognition and reflection
[0799] In parallel with the voice input, the emotion engine analyzes the user's voice data and recognizes emotions. The recognized emotion information is sent to the server, and the translation means and speech synthesis means reflect this in the translation results and speech tone. For example, if the user's voice contains anger, the translated voice will also be adjusted to reflect an angry tone.
[0800] Speech synthesis and output
[0801] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the other party to hear and understand the translated voice.
[0802] Sharing emotional information
[0803] The emotional information recognized by the emotion engine is also sent to the other party's device via the emotion transmission means. The other party can visually check the other party's emotional state on the screen, which promotes understanding of emotions beyond language.
[0804] screen sharing
[0805] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[0806] Gesture Recognition
[0807] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0808] Specific examples
[0809] Example 1: Communication at tourist spots
[0810] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data and emotion data are sent to the server, where they are translated into French with the emotion reflected, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[0811] Example 2: Using an interactive tourist guide
[0812] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checks off the places they want to visit. The check marks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information and communicate smoothly across language barriers. The emotion engine recognizes the user's emotions, and the emotional state is reflected in the displayed interface, allowing for a deeper understanding of each other's feelings.
[0813] This system effectively removes language barriers in international exchange and travel, enabling intimate communication that also includes emotions.
[0814] The processing flow will be explained below.
[0815] Step 1:
[0816] The user launches the "AI translation app" on their smartphone.
[0817] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[0818] Step 2:
[0819] The user taps the "Speak" button.
[0820] The terminal activates the voice input means and starts collecting voice data.
[0821] Step 3:
[0822] The user begins speaking.
[0823] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[0824] Step 4:
[0825] The server passes the received voice data to the voice recognition means.
[0826] The voice recognition means converts the voice data into character data.
[0827] Step 5:
[0828] The server passes the converted character data to the translation means.
[0829] A translation means translates the character data into another designated language.
[0830] Step 6:
[0831] The server passes the translated character data to the speech synthesis means.
[0832] A speech synthesis means converts the translated text into speech data.
[0833] Step 7:
[0834] The server transmits the generated voice data to User B's terminal.
[0835] User B's terminal receives the audio data and plays it back via the audio output means.
[0836] Step 8:
[0837] User B initiates reply.
[0838] User B's device collects the voice data and sends it to the server.
[0839] Step 9:
[0840] The server passes user B's voice data to a voice recognition means and converts it into text data.
[0841] The server translates the text data into User A's language.
[0842] Step 10:
[0843] The server converts the translated text data into audio data and sends it to User A's terminal.
[0844] User A's terminal receives the audio data and plays it back via the audio output means.
[0845] Step 11:
[0846] The user taps the "Screen Share" button.
[0847] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[0848] Step 12:
[0849] The device sends the screen capture data to the server.
[0850] The server transfers the received screen capture data to User B's terminal.
[0851] Step 13:
[0852] The terminal of user B displays the received screen capture data.
[0853] User A and User B share the same screen in real time.
[0854] Step 14:
[0855] User A performs a touch gesture on the screen.
[0856] The device detects the touch gesture and obtains its position and shape.
[0857] Step 15:
[0858] The terminal transmits the acquired gesture data to the server.
[0859] The server transfers the gesture data to User B's device.
[0860] Step 16:
[0861] The gesture data received by the terminal of User B is displayed on the screen in real time.
[0862] User A and User B can communicate interactively.
[0863] Step 17:
[0864] The terminal transmits the user's voice data to the emotion engine.
[0865] The emotion engine analyzes the voice data and recognizes the user's emotions.
[0866] Step 18:
[0867] The emotion engine transmits the recognized emotion information to the server.
[0868] The server reflects the emotional information in the translation means and the voice synthesis means.
[0869] Step 19:
[0870] The server converts the translated text data and emotional information into voice data.
[0871] The voice synthesis means generates voice data that reflects the emotion and transmits it to the user B's terminal.
[0872] Step 20:
[0873] The terminal of user B receives the voice data reflecting the emotion and plays it back via the voice output means.
[0874] User B can understand User A's emotions.
[0875] Step 21:
[0876] The server transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[0877] The emotion information received by the terminal of user B is visually displayed on the screen.
[0878] Example 2
[0879] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0880] While conventional speech translation systems provide the ability to translate speech into other languages, they lack the ability to recognize user emotions and reflect them in translation results and speech output. This makes it difficult for users who speak different languages to convey emotional nuances, making smooth communication difficult. Furthermore, screen sharing and interactive gesture functions are limited, making it difficult for users to intuitively share information and communicate with each other.
[0881] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0882] In this invention, the server includes a speech recognition unit, a translation unit, a speech synthesis unit, an emotion recognition unit, and an emotion transmission unit. This makes it possible to recognize emotions in a user's speech data in real time and reflect those emotions in the translation results and voice output. Furthermore, by combining a screen sharing unit and a gesture recognition unit, intuitive information sharing and interactive communication between users is realized.
[0883] The "voice input means" is a device for inputting the voice spoken by the user as digital data.
[0884] "Speech recognition means" is a technology that analyzes input voice data and converts it into text data.
[0885] "Translation means" is a technology that converts text data in one language into another language.
[0886] "Speech synthesis means" is a technology that converts text data into voice data and outputs it as synthesized voice.
[0887] The "audio output means" is a device for reproducing the synthesized audio data and letting the user hear it.
[0888] "Screen sharing means" is a technology that allows multiple users to view the same screen content at the same time.
[0889] The "display transmission means" is a technique for transmitting the display contents of the terminal in use to another terminal.
[0890] A "gesture recognition means" is a technology that detects touch gestures on a screen and recognizes their movements and shapes.
[0891] The "gesture transmitting means" is a technique for transmitting recognized gesture information to another terminal.
[0892] "Emotion recognition means" is a technology that analyzes the user's voice data and identifies emotions.
[0893] "Emotion transmission means" is a technology for transmitting recognized emotion information to another terminal.
[0894] The "gesture display means" is a technique for displaying gesture information received by the terminal of the other party on its screen.
[0895] A "buffer" is an area that temporarily stores data and is used to adjust speed differences when sending and receiving data.
[0896] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[0897] System configuration
[0898] The system has the following elements:
[0899] Voice input means
[0900] Voice recognition means
[0901] Translation tools
[0902] Voice synthesis means
[0903] Audio output means
[0904] Screen sharing method
[0905] Display and transmission means
[0906] Gesture Recognition Method
[0907] Gesture sending method
[0908] emotion recognition means
[0909] Means of transmitting emotions
[0910] Detailed Description
[0911] Voice input means
[0912] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[0913] Voice recognition means
[0914] When the voice data arrives at the server, the server uses speech recognition software like the Google Cloud Speech-to-Text API to convert the speech into text. For example, a user might say, "Hello, where am I?" and it will be converted into text.
[0915] Translation tools
[0916] The server translates the obtained text data into other languages using a translation tool such as Google Translate API. For example, the Japanese phrase "Hello, where are you?" is translated into French as "Bonjour, où sommes-nous?"
[0917] Voice synthesis means
[0918] The translated text data is converted back into speech data using a speech synthesis tool (e.g., Amazon Polly), for example, to generate the translated French phrase "Bonjour, où sommes-nous?"
[0919] Audio output means
[0920] The voice data is sent to the other party's terminal and played back by the voice output means, allowing the other party to hear and understand the translated voice.
[0921] emotion recognition means
[0922] In parallel with the voice input, an emotion recognition means on the server (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotions. The recognized emotional information is reflected in the translation results and voice tone by the translation means and voice synthesis means. For example, if the user's voice contains anger, the translation results will also be adjusted to reflect an angry tone.
[0923] Means of transmitting emotions
[0924] The recognized emotional information is also sent to the other party's device via the emotion transmission means, allowing the other party to visually confirm the other party's emotional state on the screen, facilitating understanding of emotions beyond language.
[0925] Screen sharing and display transmission means
[0926] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This allows users to see each other's screens.
[0927] Gesture Recognition Means and Gesture Transmission Means
[0928] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes the shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if a user draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[0929] Specific examples
[0930] Example 1: Communication at tourist spots
[0931] User A (Japanese) is visiting a tourist spot in a foreign country. He wants to communicate with a local guide (User B), but there is a language barrier. User A uses this system, taps the "Speak" button, and asks, "Hello, where am I?" This voice data and emotion data are sent to the server, where it is translated and the emotion is reflected in real time and sent to the guide's device. The guide listens to the voice and provides appropriate guidance.
[0932] Prompt Sentence Examples
[0933] "Please tell me how the voice input system works to communicate with foreigners at tourist spots. It also performs emotion recognition."
[0934] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0935] Step 1:
[0936] When a user launches an application on their smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. The user's voice data is used as input, and the voice data is saved in a buffer as output. Specifically, the device's microphone captures the user's voice and stores the data in a buffer.
[0937] Step 2:
[0938] The device sends audio data to the server in fixed buffers. The audio data in the buffer is used as input, and the audio data is sent to the server as output. Specifically, when the collected audio data reaches a certain volume, the device uploads the data to the server.
[0939] Step 3:
[0940] The server analyzes the received voice data and converts the voice into text data using a voice recognition tool. The voice data arriving at the server is used as input, and converted text data is generated as output. Specifically, the server calls a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data.
[0941] Step 4:
[0942] The server translates the obtained text data into another language. The text data generated by the speech recognition means is used as input, and the translated text data is generated as output. Specifically, the server calls a translation API (e.g., Google Translate) to translate from Japanese to French, etc.
[0943] Step 5:
[0944] The server's emotion recognition means analyzes the user's voice data and recognizes emotions. The received voice data is used as input, and recognized emotion information is generated as output. Specifically, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the voice data.
[0945] Step 6:
[0946] The recognized emotional information is reflected in the translation results and voice tone through the translation means and speech synthesis means. The emotional information generated by the emotion recognition means and the translated text data are used as input, and synthetic speech data reflecting the emotion is generated as output. Specifically, the server uses a speech synthesis API (e.g., Amazon Polly) to generate speech data based on the text data and emotional information.
[0947] Step 7:
[0948] The synthesized voice data is sent to the other party's terminal and played by the voice output means. The synthesized voice data is used as input, and the voice is played back in an audible form for the user as output. Specifically, the server sends the synthesized voice data to the other party's terminal, and the terminal plays the voice through a speaker.
[0949] Step 8:
[0950] The recognized emotional information is also sent to the conversation partner's terminal through the emotion transmission means. The emotional information is used as input, and the emotional state is visually displayed on the screen of the conversation partner's terminal as output. Specifically, the server transmits the emotional information to the conversation partner's terminal, and the emotional state is displayed on the terminal's display.
[0951] Step 9:
[0952] When a user taps the "Screen Share" button, the device starts capturing the screen. User operations and screen data are used as input, and screen capture data is generated as output. Specifically, the device periodically takes screenshots and sends the data to the server.
[0953] Step 10:
[0954] The screen data is sent to the other party's device via the server and displayed in real time. The screen capture data sent from the device is used as input, and the screen displayed on the other party's device is updated in real time as output. Specifically, the server relays the screen data to the other party's device, and the device displays the data on its display.
[0955] Step 11:
[0956] When a user makes a touch gesture during screen sharing, the gesture recognition means of the device detects the touch operation. Touch gesture data is used as input, and recognized gesture data is generated as output. Specifically, the device detects touch operations on the screen with a sensor and recognizes their shape and movement.
[0957] Step 12:
[0958] The gesture data is sent to the other party's device via the server and displayed in real time. The recognized gesture data is used as input to generate the gestures displayed on the other party's device as output. Specifically, the server relays the gesture data to the other party's device, and the device displays the data on its display.
[0959] (Application example 2)
[0960] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0961] In communication between workers and robots in factories, there is a problem that instructions cannot be transmitted quickly and accurately between different languages, resulting in reduced work efficiency. Furthermore, it is difficult to improve safety because the robot's behavior cannot be adjusted according to the worker's emotional state. Furthermore, it is difficult to share instructions and information between multiple workers using screen sharing and interactive touch gestures.
[0962] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0963] In this invention, the server includes a voice input device for allowing a user to input voice data, a voice recognition device for converting the voice input data into text data, a translation device for translating the text data into another language, a voice synthesis device for converting the translated text data into voice data, a voice output device for outputting the synthesized voice data, a screen sharing device for sharing a screen between users, a display transmission device for sharing the display of the currently used terminal with the other terminal, a gesture recognition device for recognizing touch gestures on the screen, a gesture transmission device for transmitting the recognized gesture information to the other terminal, an emotion engine for recognizing an emotional state from the voice data, an emotion transmission device for transmitting the recognized emotion information to the other terminal, and an instruction device for instructing the work machine using the translated text data and voice data. This enables quick and accurate instruction transmission between different languages. Furthermore, safety is improved by adjusting the robot's operation according to the worker's emotional state. Furthermore, interactive screen sharing and touch gesture sharing of instructions and information between multiple workers is possible.
[0964] "Voice input means" refers to a device or system that allows a user to input voice.
[0965] The "voice recognition means" is a device or system that converts voice data acquired by the voice input means into character data.
[0966] The "translation means" is a device or system that translates the character data generated by the speech recognition means into another language.
[0967] The "voice synthesis means" is a device or system that converts the translated text data generated by the translation means into voice data.
[0968] The "audio output means" is a device or system that outputs the audio data generated by the audio synthesis means to the user.
[0969] A "screen sharing means" is a device or system for sharing screen information among multiple users.
[0970] The "display transmission means" is a device or system that transmits the display of the terminal in use to the terminal of the other party.
[0971] A "gesture recognition means" is a device or system that recognizes touch gestures on a screen and processes them as data.
[0972] The "gesture transmitting means" is a device or system that transmits recognized gesture information to the terminal of the other party.
[0973] An "emotion engine" is a device or system for recognizing a user's emotional state from voice data.
[0974] The "emotion transmission means" is a device or system that transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[0975] The "instruction means" is a device or system that issues instructions to the work machine based on the translated character data and voice data.
[0976] This invention realizes a multilingual interactive interpretation system between workers and robots in a factory. This system is composed of a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, an emotion engine, an emotion transmission means, and an instruction means.
[0977] System Overview
[0978] Audio Input:
[0979] The server activates a voice input device, such as a smartphone or tablet, for the user to input voice. When the user starts speaking, this voice data is buffered in the device and sent to the server at regular intervals.
[0980] Speech Recognition and Translation:
[0981] The voice data received by the server is converted into text data using a voice recognition unit, and then this text data is translated into another language by a translation unit, for example, from Japanese to English.
[0982] Emotion recognition and reflection:
[0983] In parallel with voice input, the emotion engine analyzes the voice data and recognizes the user's emotional state. The recognized emotional information is sent to the server, where the translation means and voice synthesis means reflect it in the translation results and voice tone. If the user is tired, the robot's movement speed will be slowed down, for example.
[0984] Speech synthesis and output:
[0985] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the robot to understand and carry out the worker's instructions.
[0986] Sharing emotional information:
[0987] The emotion information recognized by the emotion engine is also sent to the terminal of the conversation partner through the emotion transmission means. The emotional state of the worker is displayed, promoting mutual understanding.
[0988] Screen sharing and gesture recognition:
[0989] To share a screen between users, a device initiates a screen capture and sends the captured data to the other device. Furthermore, touch gestures on the screen are recognized and the gesture data is sent to the other device, enabling interactive operation.
[0990] Specific examples
[0991] For example, in a factory, Worker A instructs a robot in Japanese, saying, "Please transport this part." This voice is translated into English in real time and transmitted to the robot. At the same time, the robot recognizes that Worker A is tired from his voice and switches its operating speed to slow mode. Worker B can then use the screen sharing function to instruct the robot on the specific destination with touch gestures, achieving efficient communication.
[0992] Example prompt sentence:
[0993] "Based on the invention below, we will design a multilingual interactive interpretation system for factory robots. It is a system that allows users to input speech and translate it into other languages in real time. Furthermore, it recognizes the user's emotions during the translation process and reflects those emotions in the translation results. It will also have functions that allow users to share screens and use interactive touch gestures. As a specific application example, when workers in a factory give instructions to a robot, it will use voice recognition and translation to make it multilingual. Also, please consider a system that recognizes the emotional state of the worker and adjusts the robot's movement speed and response content according to the situation."
[0994] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0995] Step 1:
[0996] The user activates the voice input means to input voice.
[0997] Input: User speech.
[0998] Output: The input audio data.
[0999] Specific operation: The user taps the "Speak" button on their smartphone or tablet to begin voice input. The device collects voice data through the microphone.
[1000] Step 2:
[1001] The device sends the collected voice data to the server.
[1002] Input: Audio data.
[1003] Output: The audio data sent to the server.
[1004] Specific operation: The collected audio data is sent from the terminal to the server at fixed buffer intervals.
[1005] Step 3:
[1006] The server converts the voice data into character data using a voice recognition means.
[1007] Input: The audio data sent to the server.
[1008] Output: Recognized speech text data.
[1009] Specific operation: The server calls a speech recognition API and converts the voice data into text data, for example, using the Google Speech-to-Text API.
[1010] Step 4:
[1011] The server uses a translation means to translate the character data into another language.
[1012] Input: Speech-recognized text data.
[1013] Output: The translated text data.
[1014] Specific operation: The server calls a translation API to translate the text data into another specified language. For example, it uses the Google Translate API.
[1015] Step 5:
[1016] The server uses an emotion engine to analyze the voice data and recognize the emotional state.
[1017] Input: The original audio data.
[1018] Output: Recognized emotion information.
[1019] Specific operation: Calls an emotion recognition API and recognizes emotions from features such as energy, pitch, and tempo. For example, it uses IBM Watson Tone Analyzer.
[1020] Step 6:
[1021] The server synthesizes voice data based on the translated text data and emotional information.
[1022] Input: translated text data, recognized emotion information.
[1023] Output: Synthesized audio data.
[1024] Specific operation: Calls a speech synthesis API and generates speech data that reflects emotional information. For example, it uses Amazon Polly.
[1025] Step 7:
[1026] The server transmits the synthesized voice data to the terminal.
[1027] Input: Synthesized speech data.
[1028] Output: The audio data sent to the device.
[1029] Specific operation: The server transmits the generated voice data to the other party's terminal.
[1030] Step 8:
[1031] The terminal uses the audio output means to play back the received audio data.
[1032] Input: Received audio data.
[1033] Output: The audio that is played.
[1034] Specific operation: The other party's device plays the synthesized voice through the speaker.
[1035] Step 9:
[1036] The user activates the screen sharing tool and starts capturing the screen.
[1037] Input: User operations, device screen data.
[1038] Output: Real-time captured screen data.
[1039] Specific operation: The user taps the "Screen Share" button and selects the screen they want to share. The device captures that screen in real time and sends it to the server.
[1040] Step 10:
[1041] The server transmits the screen data to the terminal of the other party.
[1042] Input: Screen data captured in real time.
[1043] Output: Screen data sent to the other device.
[1044] Specific operation: The server sends the captured screen data to the other device.
[1045] Step 11:
[1046] The terminal of the other party displays the received screen data.
[1047] Input: Received screen data.
[1048] Output: The screen that is displayed.
[1049] Specific operation: The screen data received by the other party's device is displayed on the screen.
[1050] Step 12:
[1051] A user performs a touch gesture using the gesture recognition means.
[1052] Input: User touch actions.
[1053] Output: Recognized gesture data.
[1054] Specific operation: When a user performs a touch gesture on the screen, the device recognizes the touch operation and generates gesture data.
[1055] Step 13:
[1056] The server transmits the recognized gesture data to the other terminal of the conversation partner.
[1057] Input: Recognized gesture data.
[1058] Output: Gesture data sent to the other device.
[1059] Specific operation: The server sends the recognized gesture data to the other device.
[1060] Step 14:
[1061] The gesture data received by the terminal of the conversation partner is displayed on the screen.
[1062] Input: Gesture data sent to the other device.
[1063] Output: Gesture data displayed on the screen.
[1064] Specific operation: The gesture data received by the other device is displayed on the screen in real time.
[1065] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1066] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1067] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1068] [Third embodiment]
[1069] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1070] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1071] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1072] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1073] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1074] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1075] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1076] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1077] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1078] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1079] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1080] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1081] This invention is a system that provides a function for a user to input voice, translate it into another language in real time, and output it as voice, as well as a system that allows users to share screens and use interactive touch gestures.
[1082] 1. System Overview
[1083] This system includes a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, and a gesture display means.
[1084] 2. Program Processing Overview
[1085] Voice input
[1086] When a user launches an application on a smartphone or tablet, a "Speak" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1087] Speech Recognition and Translation
[1088] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then converted into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[1089] Speech synthesis and output
[1090] The translated text data is converted back into audio data using a speech synthesis means. This audio data is sent to the terminal and played back by a speech output means. User B hears the audio in his or her own language, so he or she can understand what User A is saying.
[1091] screen sharing
[1092] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[1093] Gesture Recognition
[1094] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1095] Specific examples
[1096] Example 1: Communication at tourist spots
[1097] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1098] Example 2: Using an interactive tourist guide
[1099] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checkmarks the places they want to visit. The checkmarks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information across language barriers and communicate smoothly.
[1100] This system effectively removes language barriers in international exchanges and travel, enabling intimate communication.
[1101] The processing flow will be explained below.
[1102] Step 1:
[1103] The user launches the "AI translation app" on their smartphone.
[1104] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[1105] Step 2:
[1106] The user taps the "Speak" button.
[1107] The terminal activates the voice input means and starts collecting voice data.
[1108] Step 3:
[1109] The user begins speaking.
[1110] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[1111] Step 4:
[1112] The server passes the received voice data to the voice recognition means.
[1113] A voice recognition means converts the voice data into text data.
[1114] Step 5:
[1115] The server passes the converted text data to the translation means.
[1116] A translation means translates the text data into another designated language.
[1117] Step 6:
[1118] The server passes the translated text data to the speech synthesis means.
[1119] A speech synthesis means converts the translated text into speech data.
[1120] Step 7:
[1121] The server transmits the generated voice data to User B's terminal.
[1122] User B's terminal receives the audio data and plays it back via the audio output means.
[1123] Step 8:
[1124] User B initiates reply.
[1125] User B's device collects the voice data and sends it to the server.
[1126] Step 9:
[1127] The server passes user B's voice data to a voice recognition means and converts it into text data.
[1128] The server translates the text data into User A's language.
[1129] Step 10:
[1130] The server converts the translated text data into audio data and sends it to User A's terminal.
[1131] User A's terminal receives the audio data and plays it back via the audio output means.
[1132] Step 11:
[1133] The user taps the "Screen Share" button.
[1134] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[1135] Step 12:
[1136] The device sends the screen capture data to the server.
[1137] The server transfers the received screen capture data to User B's terminal.
[1138] Step 13:
[1139] The terminal of user B displays the received screen capture data.
[1140] User A and User B share the same screen in real time.
[1141] Step 14:
[1142] User A performs a touch gesture on the screen.
[1143] The device detects the touch gesture and obtains its position and shape.
[1144] Step 15:
[1145] The terminal transmits the acquired gesture data to the server.
[1146] The server transfers the gesture data to User B's device.
[1147] Step 16:
[1148] The gesture data received by the terminal of User B is displayed on the screen in real time.
[1149] User A and User B can communicate interactively.
[1150] Example 1
[1151] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1152] There is a need for a system that can help users who speak different languages communicate smoothly in real time and achieve more effective communication by providing additional visual information. In particular, it is a challenge to achieve smooth communication across language barriers by simultaneously performing real-time speech translation, interactive screen sharing, and gesture recognition.
[1153] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1154] In this invention, the server includes a speech recognition unit, a translation unit, and a speech synthesis unit. This allows a user to input speech, have it translated into another language in real time, and then output it as speech. Furthermore, by utilizing a gesture recognition unit, a gesture display unit, and a screen sharing unit, it is possible for users to share screens and display interactive touch gestures. These units allow users to effectively share information across language barriers and realize intimate communication.
[1155] The "voice input means" is a means for collecting the voice uttered by the user as digital data.
[1156] "Speech recognition means" refers to means for converting voice data into character data.
[1157] A "translation means" is a means for converting character data from one language to another.
[1158] The "voice synthesis means" is a means for converting character data into voice data.
[1159] The "audio output means" is a means for outputting the generated audio data through a speaker, headphones, or the like.
[1160] A "screen sharing means" is a means for sharing the contents of a display between users in real time.
[1161] The "display transmission means" is a means for transmitting the display contents of the information display device in use to the information display device of the other party.
[1162] The "gesture recognition means" is a means for detecting an operation gesture made on the display and recognizing the information.
[1163] The "gesture transmitting means" is a means for transmitting the recognized gesture information to the information display device of the other party of the conversation.
[1164] The "buffering means" is a means for dividing the audio data into a certain buffer size, temporarily storing the data, and then transmitting the data to the processing device.
[1165] The "translation output means" is a means for translating character data received by the processing device and reproducing it as voice data.
[1166] MODE FOR CARRYING OUT THE INVENTION
[1167] This invention provides a system that allows users to input speech, translate it into other languages in real time, and output it as speech. It also allows users to share a display and provide interactive gesture control functions. This system includes a speech input unit, a speech recognition unit, a translation unit, a speech synthesis unit, a speech output unit, a screen sharing unit, a display transmission unit, a gesture recognition unit, a gesture transmission unit, and a gesture display unit.
[1168] System configuration
[1169] The system operates via a user's smartphone, tablet, or other information display device, which includes a microphone for collecting voice, a speaker for outputting voice, a display, and a touchscreen. A cloud-based server handles speech recognition, translation, and speech synthesis, using speech recognition software (e.g., Google Cloud Speech-to-Text API), translation software (e.g., Google Translate API), and speech synthesis software (e.g., Amazon Polly).
[1170] Program processing overview
[1171] Voice input
[1172] When a user launches an AI translation app on their smartphone or tablet, a "Speak" button appears on the home screen. When the user taps this button, the device activates the voice input method and begins collecting the user's speech.
[1173] Buffering and sending audio data
[1174] The collected voice data is buffered in the terminal, and when a certain buffer size is reached, the terminal transmits the voice data to the server. For this reason, a buffering means is installed.
[1175] Speech recognition and conversion to text data
[1176] The server converts the received voice data into text data using speech recognition software such as the Google Cloud Speech-to-Text API.
[1177] translation
[1178] The text data generated by speech recognition is translated into other languages using translation means on the server. Here, the Google Translate API is used. For example, if User A utters "Hello, where am I?", this is translated into French as "Bonjour, où sommes-nous?"
[1179] Speech synthesis and playback
[1180] The translated text data is converted back into voice data using a voice synthesis means, such as voice synthesis software like Amazon Polly, and the generated voice data is sent to the terminal and played back by the voice output means.
[1181] Specific examples
[1182] Use in tourist areas
[1183] For example, consider the case where User A (Japanese) is visiting a tourist spot in France. User A launches the "AI translation app," taps the "Speak" button, and asks the French guide (User B), "Hello, where is this?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1184] Interactive tourist guide
[1185] When User A and User B are making sightseeing plans while looking at a tourist guide map, User A taps the "Screen Share" button to share the map, and their device starts capturing the screen. The captured screen data is sent to User B's device via the server. Furthermore, when User A indicates places they want to visit using touch gestures, the gesture data is also displayed in real time on User B's device via the server. This allows them to make plans while sharing information across language barriers.
[1186] Prompt Sentence Examples
[1187] A description of this system can be generated by inputting the following prompt sentence into the generative AI model:
[1188] This system allows users to input speech, translate it into other languages in real time, and output it as speech, as well as provide screen sharing and interactive gesture control functions between users. Please explain the speech input, speech recognition, translation, speech synthesis, and gesture recognition and sharing functions.
[1189] In this way, this system can effectively remove language barriers in international exchanges and travel, enabling intimate communication.
[1190] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1191] Step 1:
[1192] A user launches an "AI translation app" on an information display device (smartphone or tablet) and taps the "Speak" button displayed on the home screen. This activates the device's voice input means. The input is the user's speech (voice data), and the collected voice data is obtained as output.
[1193] Step 2:
[1194] The device starts collecting the user's speech through the microphone. The collected speech data is buffered within the device, and when a certain buffer size is reached, the buffered speech data is sent to the server. The input is the raw speech data of the user's speech, and the output is the buffered speech data sent to the server.
[1195] Step 3:
[1196] The server processes the received voice data and converts it into text data using a speech recognition tool. The input is the voice data sent from the device, and the output is text data. In this case, speech recognition software (e.g., Google Cloud Speech-to-Text API) is used.
[1197] Step 4:
[1198] The server translates the voice-recognized text data into another language using a translation tool. The input is the text data processed by the server, and the output is the text data translated into another language. Here, translation software (e.g., Google Translate API) is used.
[1199] Step 5:
[1200] The server converts the translated text data back into voice data using a speech synthesis tool. The input is the translated text data, and the output is synthesized voice data. In this case, speech synthesis software (e.g., Amazon Polly) is used.
[1201] Step 6:
[1202] The server sends synthesized voice data to the terminal. The input is the voice data from the server, and the terminal receives it as output.
[1203] Step 7:
[1204] The terminal plays back the received voice data using the voice output means. The input is the voice data received from the server, and the output is the voice played back through the speaker. This allows user B to hear user A's speech in his own language.
[1205] Step 8:
[1206] The user taps the "Screen Share" button, and the device starts capturing the screen. The input is the user's screen, and the output is the captured screen data.
[1207] Step 9:
[1208] The terminal sends the captured screen data to the server. The input is the captured screen data, and the output is the data sent to the server.
[1209] Step 10:
[1210] The server transfers the received screen data to the other party's terminal. The input is the transmitted screen data, and the output is the data transmitted to the other party's terminal.
[1211] Step 11:
[1212] The screen data received by the other party's device is displayed. The input is the received screen data, and the output is the screen displayed on the display. This allows user B to see user A's screen in real time.
[1213] Step 12:
[1214] The user performs a touch gesture (e.g., drawing a heart symbol) on the screen. The input is the user's touch operation, and the output is gesture data.
[1215] Step 13:
[1216] The device recognizes the touch gesture and sends the data to the server. The input is the touch gesture operation data, and the output is the recognized gesture data.
[1217] Step 14:
[1218] The server transmits the received gesture data to the other device. The input is the recognized gesture data, and the output is the data transmitted to the other device.
[1219] Step 15:
[1220] The gesture data received by the other party's device is displayed in real time. The input is the received gesture data, and the output is the gesture displayed on the display. This allows the gesture drawn by User A to be displayed on User B's screen as well.
[1221] (Application example 1)
[1222] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1223] Currently, language barriers are a major obstacle in situations requiring real-time multilingual communication. Language differences in particular hinder smooth interactions between viewers and streamers during live streaming, as well as communication between viewers themselves. Furthermore, screen sharing and interactive gesture features are not fully integrated. This lack of an environment allows users of different languages to smoothly communicate with each other and enjoy live content with a sense of unity.
[1224] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1225] In this invention, the server includes a voice input means for allowing a user to input voice, a voice recognition means for converting the voice input data into character data, and a translation means for translating the character data into another language, thereby enabling multilingual live chat.
[1226] "Voice input means" refers to a device or technology that allows a user to input voice.
[1227] "Speech recognition means" refers to technology that analyzes input voice data and converts it into text data.
[1228] "Translation means" refers to a technology that converts text data written in one language into text data written in another language.
[1229] "Speech synthesis means" refers to technology that analyzes text data and converts it into voice data.
[1230] "Audio output means" refers to technology for playing audio data through a device such as a speaker.
[1231] "Screen sharing means" refers to a technology that allows users to share the same screen in real time.
[1232] "Display transmission means" refers to a technology for transmitting the display contents of a terminal in use to another terminal.
[1233] "Gesture recognition means" refers to technology that detects touch gestures made on the screen and recognizes their content.
[1234] The "gesture transmission means" refers to a technique for transmitting recognized gesture information to another terminal.
[1235] "Multilingual live chat means" refers to technology that enables viewers to communicate in different languages in real time during live streaming.
[1236] "Translation display means" refers to technology that translates information sent using the multilingual live chat means and displays it as audio or text.
[1237] The present invention provides a system for supporting smooth communication between users, particularly in situations where real-time communication in multiple languages is required. Specific embodiments of this system will be described below.
[1238] System configuration
[1239] This system is composed of a server, a user terminal, and various software modules. Specific components include the following:
[1240] Voice input means
[1241] Voice recognition means
[1242] Translation tools
[1243] Voice synthesis means
[1244] Audio output means
[1245] Screen sharing method
[1246] Display and transmission means
[1247] Gesture Recognition Method
[1248] Gesture sending method
[1249] Multilingual live chat options
[1250] Translation display means
[1251] System Operation
[1252] 1. Voice input and recognition:
[1253] A user inputs voice using a device such as a smartphone or tablet. The voice input means collects voice data and transmits it to the voice recognition means.
[1254] The speech recognition means converts the collected speech data into text data, which is processed in real time.
[1255] 2. Translation and speech synthesis:
[1256] The translation means converts the text data into a language specified by the user. For example, if a translation from Japanese to English is required, the translation means is applied.
[1257] The translated text data is converted back into voice data by the voice synthesis means, and the voice data is reproduced through the voice output means.
[1258] 3. Screen sharing and interactive gestures:
[1259] The user can use the screen sharing means to share the screen with the other party in real time, and the display transmission means transmits the display content of the terminal currently in use to the other party's terminal.
[1260] The gesture recognition means detects touch gestures made on the screen and transmits the information to the other party's terminal. This is handled by the gesture transmission means.
[1261] 4. Multilingual Live Chat:
[1262] The multilingual live chat means supports real-time communication between viewers during live streaming. Information sent by viewers is translated through the translation display means and displayed to the other person. This allows smooth communication even between viewers who speak different languages.
[1263] Hardware and software used
[1264] Hardware:
[1265] Server: Responsible for data processing, translation, etc.
[1266] User devices: smartphones, tablets, head-mounted displays, etc.
[1267] software:
[1268] Flask: Used as a web server
[1269] SpeechRecognition: Used for speech recognition
[1270] Googletrans: used for text translation
[1271] Pyttsx3: Used for speech synthesis
[1272] Specific example explanation
[1273] For example, imagine a situation where a question can be sent in Japanese during a live streaming event, and the question can be translated into English in real time and spoken to the streamer. In this way, viewers and streamers can communicate smoothly even if they speak different languages.
[1274] Example prompt sentence:
[1275] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[1276] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1277] Step 1:
[1278] A user starts an application on a device such as a smartphone or tablet and inputs voice. The voice input means collects this voice data and buffers it in real time. The input is the user's voice, and it is temporarily stored in the device as voice data.
[1279] Step 2:
[1280] Voice data is sent to the server at fixed buffer intervals. The server passes this voice data to a voice recognition means and converts it into text data. The input is voice data and the output is text data. The server analyzes the voice data and outputs it as text data.
[1281] Step 3:
[1282] The server passes the recognized character data to the translation means and translates it into the specified language. The input is character data, and the output is translated character data. The server processes the character data to convert it into another language.
[1283] Step 4:
[1284] The server passes the translated text data to a speech synthesis means, which converts it into new speech data. The input is the translated text data, and the output is the new speech data. The speech synthesis means uses a specific speech synthesis algorithm to output the text data as speech data.
[1285] Step 5:
[1286] New audio data is sent from the server to the user's terminal and played through the audio output means. The input is the audio data and the output is the audio to be played. The terminal receives this audio data and outputs it as audio through a speaker.
[1287] Step 6:
[1288] When a user wishes to share their screen, the terminal starts capturing the screen using the screen sharing means and sends the captured data to the server. The input is the screen capture data, and the output is the screen data sent to the server. The server then relays this screen data to the terminals of other users.
[1289] Step 7:
[1290] The other user's device displays the screen data received from the server. The input is the received screen data, and the output is the screen display. The device analyzes the received data and displays it to the user.
[1291] Step 8:
[1292] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes its shape and movement. The input is touch gesture data, and the output is recognized gesture information. The device detects the touch operation and sends it to the server as gesture data.
[1293] Step 9:
[1294] The server transmits the recognized gesture information to the terminal of the other user and displays it in real time using a gesture transmission means. The input is the recognized gesture information and the output is the gesture to be displayed. The terminal displays the received gesture information on the screen.
[1295] Step 10:
[1296] A multilingual live chat means performs processing to support real-time communication between viewers. A translation display means translates messages sent by viewers in different languages and displays them on the terminal of the person interacting with them. The input is the viewer's message, and the output is a display of the translated message. The terminal presents this translated information to the viewer, enabling real-time communication.
[1297] When using a generative AI model for speech recognition or translation, the following example prompt can be set:
[1298] Example prompt sentence:
[1299] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[1300] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1301] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[1302] 1. System Overview
[1303] The system has the following elements:
[1304] Voice input means
[1305] Voice recognition means
[1306] Translation tools
[1307] Voice synthesis means
[1308] Audio output means
[1309] Screen sharing method
[1310] Display and transmission means
[1311] Gesture Recognition Method
[1312] Gesture sending method
[1313] Emotion Engine
[1314] Means of transmitting emotions
[1315] 2. Program Processing Overview
[1316] Voice input
[1317] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1318] Speech Recognition and Translation
[1319] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then translated into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[1320] Emotion recognition and reflection
[1321] In parallel with the voice input, the emotion engine analyzes the user's voice data and recognizes emotions. The recognized emotion information is sent to the server, and the translation means and speech synthesis means reflect this in the translation results and speech tone. For example, if the user's voice contains anger, the translated voice will also be adjusted to reflect an angry tone.
[1322] Speech synthesis and output
[1323] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the other party to hear and understand the translated voice.
[1324] Sharing emotional information
[1325] The emotional information recognized by the emotion engine is also sent to the other party's device via the emotion transmission means. The other party can visually check the other party's emotional state on the screen, which promotes understanding of emotions beyond language.
[1326] screen sharing
[1327] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[1328] Gesture Recognition
[1329] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1330] Specific examples
[1331] Example 1: Communication at tourist spots
[1332] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data and emotion data are sent to the server, where they are translated into French with the emotion reflected, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1333] Example 2: Using an interactive tourist guide
[1334] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checks off the places they want to visit. The check marks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information and communicate smoothly across language barriers. The emotion engine recognizes the user's emotions, and the emotional state is reflected in the displayed interface, allowing for a deeper understanding of each other's feelings.
[1335] This system effectively removes language barriers in international exchange and travel, enabling intimate communication that also includes emotions.
[1336] The processing flow will be explained below.
[1337] Step 1:
[1338] The user launches the "AI translation app" on their smartphone.
[1339] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[1340] Step 2:
[1341] The user taps the "Speak" button.
[1342] The terminal activates the voice input means and starts collecting voice data.
[1343] Step 3:
[1344] The user begins speaking.
[1345] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[1346] Step 4:
[1347] The server passes the received voice data to the voice recognition means.
[1348] The voice recognition means converts the voice data into character data.
[1349] Step 5:
[1350] The server passes the converted character data to the translation means.
[1351] A translation means translates the character data into another designated language.
[1352] Step 6:
[1353] The server passes the translated character data to the speech synthesis means.
[1354] A speech synthesis means converts the translated text into speech data.
[1355] Step 7:
[1356] The server transmits the generated voice data to User B's terminal.
[1357] User B's terminal receives the audio data and plays it back via the audio output means.
[1358] Step 8:
[1359] User B initiates reply.
[1360] User B's device collects the voice data and sends it to the server.
[1361] Step 9:
[1362] The server passes user B's voice data to a voice recognition means and converts it into text data.
[1363] The server translates the text data into User A's language.
[1364] Step 10:
[1365] The server converts the translated text data into audio data and sends it to User A's terminal.
[1366] User A's terminal receives the audio data and plays it back via the audio output means.
[1367] Step 11:
[1368] The user taps the "Screen Share" button.
[1369] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[1370] Step 12:
[1371] The device sends the screen capture data to the server.
[1372] The server transfers the received screen capture data to User B's terminal.
[1373] Step 13:
[1374] The terminal of user B displays the received screen capture data.
[1375] User A and User B share the same screen in real time.
[1376] Step 14:
[1377] User A performs a touch gesture on the screen.
[1378] The device detects the touch gesture and obtains its position and shape.
[1379] Step 15:
[1380] The terminal transmits the acquired gesture data to the server.
[1381] The server transfers the gesture data to User B's device.
[1382] Step 16:
[1383] The gesture data received by the terminal of User B is displayed on the screen in real time.
[1384] User A and User B can communicate interactively.
[1385] Step 17:
[1386] The terminal transmits the user's voice data to the emotion engine.
[1387] The emotion engine analyzes the voice data and recognizes the user's emotions.
[1388] Step 18:
[1389] The emotion engine transmits the recognized emotion information to the server.
[1390] The server reflects the emotional information in the translation means and the voice synthesis means.
[1391] Step 19:
[1392] The server converts the translated text data and emotional information into voice data.
[1393] The voice synthesis means generates voice data that reflects the emotion and transmits it to the user B's terminal.
[1394] Step 20:
[1395] The terminal of user B receives the voice data reflecting the emotion and plays it back via the voice output means.
[1396] User B can understand User A's emotions.
[1397] Step 21:
[1398] The server transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[1399] The emotion information received by the terminal of user B is visually displayed on the screen.
[1400] Example 2
[1401] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1402] While conventional speech translation systems provide the ability to translate speech into other languages, they lack the ability to recognize user emotions and reflect them in translation results and speech output. This makes it difficult for users who speak different languages to convey emotional nuances, making smooth communication difficult. Furthermore, screen sharing and interactive gesture functions are limited, making it difficult for users to intuitively share information and communicate with each other.
[1403] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1404] In this invention, the server includes a speech recognition unit, a translation unit, a speech synthesis unit, an emotion recognition unit, and an emotion transmission unit. This makes it possible to recognize emotions in a user's speech data in real time and reflect those emotions in the translation results and voice output. Furthermore, by combining a screen sharing unit and a gesture recognition unit, intuitive information sharing and interactive communication between users is realized.
[1405] The "voice input means" is a device for inputting the voice spoken by the user as digital data.
[1406] "Speech recognition means" is a technology that analyzes input voice data and converts it into text data.
[1407] "Translation means" is a technology that converts text data in one language into another language.
[1408] "Speech synthesis means" is a technology that converts text data into voice data and outputs it as synthesized voice.
[1409] The "audio output means" is a device for reproducing the synthesized audio data and letting the user hear it.
[1410] "Screen sharing means" is a technology that allows multiple users to view the same screen content at the same time.
[1411] The "display transmission means" is a technique for transmitting the display contents of the terminal in use to another terminal.
[1412] A "gesture recognition means" is a technology that detects touch gestures on a screen and recognizes their movements and shapes.
[1413] The "gesture transmitting means" is a technique for transmitting recognized gesture information to another terminal.
[1414] "Emotion recognition means" is a technology that analyzes the user's voice data and identifies emotions.
[1415] "Emotion transmission means" is a technology for transmitting recognized emotion information to another terminal.
[1416] The "gesture display means" is a technique for displaying gesture information received by the terminal of the other party on its screen.
[1417] A "buffer" is an area that temporarily stores data and is used to adjust speed differences when sending and receiving data.
[1418] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[1419] System configuration
[1420] The system has the following elements:
[1421] Voice input means
[1422] Voice recognition means
[1423] Translation tools
[1424] Voice synthesis means
[1425] Audio output means
[1426] Screen sharing method
[1427] Display and transmission means
[1428] Gesture Recognition Method
[1429] Gesture sending method
[1430] emotion recognition means
[1431] Means of transmitting emotions
[1432] Detailed Description
[1433] Voice input means
[1434] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1435] Voice recognition means
[1436] When the voice data arrives at the server, the server uses speech recognition software like the Google Cloud Speech-to-Text API to convert the speech into text. For example, a user might say, "Hello, where am I?" and it will be converted into text.
[1437] Translation tools
[1438] The server translates the obtained text data into other languages using a translation tool such as Google Translate API. For example, the Japanese phrase "Hello, where are you?" is translated into French as "Bonjour, où sommes-nous?"
[1439] Voice synthesis means
[1440] The translated text data is converted back into speech data using a speech synthesis tool (e.g., Amazon Polly), for example, to generate the translated French phrase "Bonjour, où sommes-nous?"
[1441] Audio output means
[1442] The voice data is sent to the other party's terminal and played back by the voice output means, allowing the other party to hear and understand the translated voice.
[1443] emotion recognition means
[1444] In parallel with the voice input, an emotion recognition means on the server (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotions. The recognized emotional information is reflected in the translation results and voice tone by the translation means and voice synthesis means. For example, if the user's voice contains anger, the translation results will also be adjusted to reflect an angry tone.
[1445] Means of transmitting emotions
[1446] The recognized emotional information is also sent to the other party's device via the emotion transmission means, allowing the other party to visually confirm the other party's emotional state on the screen, facilitating understanding of emotions beyond language.
[1447] Screen sharing and display transmission means
[1448] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This allows users to see each other's screens.
[1449] Gesture Recognition Means and Gesture Transmission Means
[1450] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes the shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if a user draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1451] Specific examples
[1452] Example 1: Communication at tourist spots
[1453] User A (Japanese) is visiting a tourist spot in a foreign country. He wants to communicate with a local guide (User B), but there is a language barrier. User A uses this system, taps the "Speak" button, and asks, "Hello, where am I?" This voice data and emotion data are sent to the server, where it is translated and the emotion is reflected in real time and sent to the guide's device. The guide listens to the voice and provides appropriate guidance.
[1454] Prompt Sentence Examples
[1455] "Please tell me how the voice input system works to communicate with foreigners at tourist spots. It also performs emotion recognition."
[1456] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1457] Step 1:
[1458] When a user launches an application on their smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. The user's voice data is used as input, and the voice data is saved in a buffer as output. Specifically, the device's microphone captures the user's voice and stores the data in a buffer.
[1459] Step 2:
[1460] The device sends audio data to the server in fixed buffers. The audio data in the buffer is used as input, and the audio data is sent to the server as output. Specifically, when the collected audio data reaches a certain volume, the device uploads the data to the server.
[1461] Step 3:
[1462] The server analyzes the received voice data and converts the voice into text data using a voice recognition tool. The voice data arriving at the server is used as input, and converted text data is generated as output. Specifically, the server calls a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data.
[1463] Step 4:
[1464] The server translates the obtained text data into another language. The text data generated by the speech recognition means is used as input, and the translated text data is generated as output. Specifically, the server calls a translation API (e.g., Google Translate) to translate from Japanese to French, etc.
[1465] Step 5:
[1466] The server's emotion recognition means analyzes the user's voice data and recognizes emotions. The received voice data is used as input, and recognized emotion information is generated as output. Specifically, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the voice data.
[1467] Step 6:
[1468] The recognized emotional information is reflected in the translation results and voice tone through the translation means and speech synthesis means. The emotional information generated by the emotion recognition means and the translated text data are used as input, and synthetic speech data reflecting the emotion is generated as output. Specifically, the server uses a speech synthesis API (e.g., Amazon Polly) to generate speech data based on the text data and emotional information.
[1469] Step 7:
[1470] The synthesized voice data is sent to the other party's terminal and played by the voice output means. The synthesized voice data is used as input, and the voice is played back in an audible form for the user as output. Specifically, the server sends the synthesized voice data to the other party's terminal, and the terminal plays the voice through a speaker.
[1471] Step 8:
[1472] The recognized emotional information is also sent to the conversation partner's terminal through the emotion transmission means. The emotional information is used as input, and the emotional state is visually displayed on the screen of the conversation partner's terminal as output. Specifically, the server transmits the emotional information to the conversation partner's terminal, and the emotional state is displayed on the terminal's display.
[1473] Step 9:
[1474] When a user taps the "Screen Share" button, the device starts capturing the screen. User operations and screen data are used as input, and screen capture data is generated as output. Specifically, the device periodically takes screenshots and sends the data to the server.
[1475] Step 10:
[1476] The screen data is sent to the other party's device via the server and displayed in real time. The screen capture data sent from the device is used as input, and the screen displayed on the other party's device is updated in real time as output. Specifically, the server relays the screen data to the other party's device, and the device displays the data on its display.
[1477] Step 11:
[1478] When a user makes a touch gesture during screen sharing, the gesture recognition means of the device detects the touch operation. Touch gesture data is used as input, and recognized gesture data is generated as output. Specifically, the device detects touch operations on the screen with a sensor and recognizes their shape and movement.
[1479] Step 12:
[1480] The gesture data is sent to the other party's device via the server and displayed in real time. The recognized gesture data is used as input to generate the gestures displayed on the other party's device as output. Specifically, the server relays the gesture data to the other party's device, and the device displays the data on its display.
[1481] (Application example 2)
[1482] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1483] In communication between workers and robots in factories, there is a problem that instructions cannot be transmitted quickly and accurately between different languages, resulting in reduced work efficiency. Furthermore, it is difficult to improve safety because the robot's behavior cannot be adjusted according to the worker's emotional state. Furthermore, it is difficult to share instructions and information between multiple workers using screen sharing and interactive touch gestures.
[1484] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1485] In this invention, the server includes a voice input device for allowing a user to input voice data, a voice recognition device for converting the voice input data into text data, a translation device for translating the text data into another language, a voice synthesis device for converting the translated text data into voice data, a voice output device for outputting the synthesized voice data, a screen sharing device for sharing a screen between users, a display transmission device for sharing the display of the currently used terminal with the other terminal, a gesture recognition device for recognizing touch gestures on the screen, a gesture transmission device for transmitting the recognized gesture information to the other terminal, an emotion engine for recognizing an emotional state from the voice data, an emotion transmission device for transmitting the recognized emotion information to the other terminal, and an instruction device for instructing the work machine using the translated text data and voice data. This enables quick and accurate instruction transmission between different languages. Furthermore, safety is improved by adjusting the robot's operation according to the worker's emotional state. Furthermore, interactive screen sharing and touch gesture sharing of instructions and information between multiple workers is possible.
[1486] "Voice input means" refers to a device or system that allows a user to input voice.
[1487] The "voice recognition means" is a device or system that converts voice data acquired by the voice input means into character data.
[1488] The "translation means" is a device or system that translates the character data generated by the speech recognition means into another language.
[1489] The "voice synthesis means" is a device or system that converts the translated text data generated by the translation means into voice data.
[1490] The "audio output means" is a device or system that outputs the audio data generated by the audio synthesis means to the user.
[1491] A "screen sharing means" is a device or system for sharing screen information among multiple users.
[1492] The "display transmission means" is a device or system that transmits the display of the terminal in use to the terminal of the other party.
[1493] A "gesture recognition means" is a device or system that recognizes touch gestures on a screen and processes them as data.
[1494] The "gesture transmitting means" is a device or system that transmits recognized gesture information to the terminal of the other party.
[1495] An "emotion engine" is a device or system for recognizing a user's emotional state from voice data.
[1496] The "emotion transmission means" is a device or system that transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[1497] The "instruction means" is a device or system that issues instructions to the work machine based on the translated character data and voice data.
[1498] This invention realizes a multilingual interactive interpretation system between workers and robots in a factory. This system is composed of a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, an emotion engine, an emotion transmission means, and an instruction means.
[1499] System Overview
[1500] Audio Input:
[1501] The server activates a voice input device, such as a smartphone or tablet, for the user to input voice. When the user starts speaking, this voice data is buffered in the device and sent to the server at regular intervals.
[1502] Speech Recognition and Translation:
[1503] The voice data received by the server is converted into text data using a voice recognition unit, and then this text data is translated into another language by a translation unit, for example, from Japanese to English.
[1504] Emotion recognition and reflection:
[1505] In parallel with voice input, the emotion engine analyzes the voice data and recognizes the user's emotional state. The recognized emotional information is sent to the server, where the translation means and voice synthesis means reflect it in the translation results and voice tone. If the user is tired, the robot's movement speed will be slowed down, for example.
[1506] Speech synthesis and output:
[1507] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the robot to understand and carry out the worker's instructions.
[1508] Sharing emotional information:
[1509] The emotion information recognized by the emotion engine is also sent to the terminal of the conversation partner through the emotion transmission means. The emotional state of the worker is displayed, promoting mutual understanding.
[1510] Screen sharing and gesture recognition:
[1511] To share a screen between users, a device initiates a screen capture and sends the captured data to the other device. Furthermore, touch gestures on the screen are recognized and the gesture data is sent to the other device, enabling interactive operation.
[1512] Specific examples
[1513] For example, in a factory, Worker A instructs a robot in Japanese, saying, "Please transport this part." This voice is translated into English in real time and transmitted to the robot. At the same time, the robot recognizes that Worker A is tired from his voice and switches its operating speed to slow mode. Worker B can then use the screen sharing function to instruct the robot on the specific destination with touch gestures, achieving efficient communication.
[1514] Example prompt sentence:
[1515] "Based on the invention below, we will design a multilingual interactive interpretation system for factory robots. It is a system that allows users to input speech and translate it into other languages in real time. Furthermore, it recognizes the user's emotions during the translation process and reflects those emotions in the translation results. It will also have functions that allow users to share screens and use interactive touch gestures. As a specific application example, when workers in a factory give instructions to a robot, it will use voice recognition and translation to make it multilingual. Also, please consider a system that recognizes the emotional state of the worker and adjusts the robot's movement speed and response content according to the situation."
[1516] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1517] Step 1:
[1518] The user activates the voice input means to input voice.
[1519] Input: User speech.
[1520] Output: The input audio data.
[1521] Specific operation: The user taps the "Speak" button on their smartphone or tablet to begin voice input. The device collects voice data through the microphone.
[1522] Step 2:
[1523] The device sends the collected voice data to the server.
[1524] Input: Audio data.
[1525] Output: The audio data sent to the server.
[1526] Specific operation: The collected audio data is sent from the terminal to the server at fixed buffer intervals.
[1527] Step 3:
[1528] The server converts the voice data into character data using a voice recognition means.
[1529] Input: The audio data sent to the server.
[1530] Output: Recognized speech text data.
[1531] Specific operation: The server calls a speech recognition API and converts the voice data into text data, for example, using the Google Speech-to-Text API.
[1532] Step 4:
[1533] The server uses a translation means to translate the character data into another language.
[1534] Input: Speech-recognized text data.
[1535] Output: The translated text data.
[1536] Specific operation: The server calls a translation API to translate the text data into another specified language. For example, it uses the Google Translate API.
[1537] Step 5:
[1538] The server uses an emotion engine to analyze the voice data and recognize the emotional state.
[1539] Input: The original audio data.
[1540] Output: Recognized emotion information.
[1541] Specific operation: Calls an emotion recognition API and recognizes emotions from features such as energy, pitch, and tempo. For example, it uses IBM Watson Tone Analyzer.
[1542] Step 6:
[1543] The server synthesizes voice data based on the translated text data and emotional information.
[1544] Input: translated text data, recognized emotion information.
[1545] Output: Synthesized audio data.
[1546] Specific operation: Calls a speech synthesis API and generates speech data that reflects emotional information. For example, it uses Amazon Polly.
[1547] Step 7:
[1548] The server transmits the synthesized voice data to the terminal.
[1549] Input: Synthesized speech data.
[1550] Output: The audio data sent to the device.
[1551] Specific operation: The server transmits the generated voice data to the other party's terminal.
[1552] Step 8:
[1553] The terminal uses the audio output means to play back the received audio data.
[1554] Input: Received audio data.
[1555] Output: The audio that is played.
[1556] Specific operation: The other party's device plays the synthesized voice through the speaker.
[1557] Step 9:
[1558] The user activates the screen sharing tool and starts capturing the screen.
[1559] Input: User operations, device screen data.
[1560] Output: Real-time captured screen data.
[1561] Specific operation: The user taps the "Screen Share" button and selects the screen they want to share. The device captures that screen in real time and sends it to the server.
[1562] Step 10:
[1563] The server transmits the screen data to the terminal of the other party.
[1564] Input: Screen data captured in real time.
[1565] Output: Screen data sent to the other device.
[1566] Specific operation: The server sends the captured screen data to the other device.
[1567] Step 11:
[1568] The terminal of the other party displays the received screen data.
[1569] Input: Received screen data.
[1570] Output: The screen that is displayed.
[1571] Specific operation: The screen data received by the other party's device is displayed on the screen.
[1572] Step 12:
[1573] A user performs a touch gesture using the gesture recognition means.
[1574] Input: User touch actions.
[1575] Output: Recognized gesture data.
[1576] Specific operation: When a user performs a touch gesture on the screen, the device recognizes the touch operation and generates gesture data.
[1577] Step 13:
[1578] The server transmits the recognized gesture data to the other terminal of the conversation partner.
[1579] Input: Recognized gesture data.
[1580] Output: Gesture data sent to the other device.
[1581] Specific operation: The server sends the recognized gesture data to the other device.
[1582] Step 14:
[1583] The gesture data received by the terminal of the conversation partner is displayed on the screen.
[1584] Input: Gesture data sent to the other device.
[1585] Output: Gesture data displayed on the screen.
[1586] Specific operation: The gesture data received by the other device is displayed on the screen in real time.
[1587] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1588] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1589] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1590] [Fourth embodiment]
[1591] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1592] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1593] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1594] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1595] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1596] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1597] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1598] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1599] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1600] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1601] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1602] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1603] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1604] This invention is a system that provides a function for a user to input voice, translate it into another language in real time, and output it as voice, as well as a system that allows users to share screens and use interactive touch gestures.
[1605] 1. System Overview
[1606] This system includes a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, and a gesture display means.
[1607] 2. Program Processing Overview
[1608] Voice input
[1609] When a user launches an application on a smartphone or tablet, a "Speak" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1610] Speech Recognition and Translation
[1611] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then converted into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[1612] Speech synthesis and output
[1613] The translated text data is converted back into audio data using a speech synthesis means. This audio data is sent to the terminal and played back by a speech output means. User B hears the audio in his or her own language, so he or she can understand what User A is saying.
[1614] screen sharing
[1615] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[1616] Gesture Recognition
[1617] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1618] Specific examples
[1619] Example 1: Communication at tourist spots
[1620] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1621] Example 2: Using an interactive tourist guide
[1622] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checkmarks the places they want to visit. The checkmarks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information across language barriers and communicate smoothly.
[1623] This system effectively removes language barriers in international exchanges and travel, enabling intimate communication.
[1624] The processing flow will be explained below.
[1625] Step 1:
[1626] The user launches the "AI translation app" on their smartphone.
[1627] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[1628] Step 2:
[1629] The user taps the "Speak" button.
[1630] The terminal activates the voice input means and starts collecting voice data.
[1631] Step 3:
[1632] The user begins speaking.
[1633] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[1634] Step 4:
[1635] The server passes the received voice data to the voice recognition means.
[1636] A voice recognition means converts the voice data into text data.
[1637] Step 5:
[1638] The server passes the converted text data to the translation means.
[1639] A translation means translates the text data into another designated language.
[1640] Step 6:
[1641] The server passes the translated text data to the speech synthesis means.
[1642] A speech synthesis means converts the translated text into speech data.
[1643] Step 7:
[1644] The server transmits the generated voice data to User B's terminal.
[1645] User B's terminal receives the audio data and plays it back via the audio output means.
[1646] Step 8:
[1647] User B initiates reply.
[1648] User B's device collects the voice data and sends it to the server.
[1649] Step 9:
[1650] The server passes user B's voice data to a voice recognition means and converts it into text data.
[1651] The server translates the text data into User A's language.
[1652] Step 10:
[1653] The server converts the translated text data into audio data and sends it to User A's terminal.
[1654] User A's terminal receives the audio data and plays it back via the audio output means.
[1655] Step 11:
[1656] The user taps the "Screen Share" button.
[1657] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[1658] Step 12:
[1659] The device sends the screen capture data to the server.
[1660] The server transfers the received screen capture data to User B's terminal.
[1661] Step 13:
[1662] The terminal of user B displays the received screen capture data.
[1663] User A and User B share the same screen in real time.
[1664] Step 14:
[1665] User A performs a touch gesture on the screen.
[1666] The device detects the touch gesture and obtains its position and shape.
[1667] Step 15:
[1668] The terminal transmits the acquired gesture data to the server.
[1669] The server transfers the gesture data to User B's device.
[1670] Step 16:
[1671] The gesture data received by the terminal of User B is displayed on the screen in real time.
[1672] User A and User B can communicate interactively.
[1673] Example 1
[1674] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1675] There is a need for a system that can help users who speak different languages communicate smoothly in real time and achieve more effective communication by providing additional visual information. In particular, it is a challenge to achieve smooth communication across language barriers by simultaneously performing real-time speech translation, interactive screen sharing, and gesture recognition.
[1676] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1677] In this invention, the server includes a speech recognition unit, a translation unit, and a speech synthesis unit. This allows a user to input speech, have it translated into another language in real time, and then output it as speech. Furthermore, by utilizing a gesture recognition unit, a gesture display unit, and a screen sharing unit, it is possible for users to share screens and display interactive touch gestures. These units allow users to effectively share information across language barriers and realize intimate communication.
[1678] The "voice input means" is a means for collecting the voice uttered by the user as digital data.
[1679] "Speech recognition means" refers to means for converting voice data into character data.
[1680] A "translation means" is a means for converting character data from one language to another.
[1681] The "voice synthesis means" is a means for converting character data into voice data.
[1682] The "audio output means" is a means for outputting the generated audio data through a speaker, headphones, or the like.
[1683] A "screen sharing means" is a means for sharing the contents of a display between users in real time.
[1684] The "display transmission means" is a means for transmitting the display contents of the information display device in use to the information display device of the other party.
[1685] The "gesture recognition means" is a means for detecting an operation gesture made on the display and recognizing the information.
[1686] The "gesture transmitting means" is a means for transmitting the recognized gesture information to the information display device of the other party of the conversation.
[1687] The "buffering means" is a means for dividing the audio data into a certain buffer size, temporarily storing the data, and then transmitting the data to the processing device.
[1688] The "translation output means" is a means for translating character data received by the processing device and reproducing it as voice data.
[1689] MODE FOR CARRYING OUT THE INVENTION
[1690] This invention provides a system that allows users to input speech, translate it into other languages in real time, and output it as speech. It also allows users to share a display and provide interactive gesture control functions. This system includes a speech input unit, a speech recognition unit, a translation unit, a speech synthesis unit, a speech output unit, a screen sharing unit, a display transmission unit, a gesture recognition unit, a gesture transmission unit, and a gesture display unit.
[1691] System configuration
[1692] The system operates via a user's smartphone, tablet, or other information display device, which includes a microphone for collecting voice, a speaker for outputting voice, a display, and a touchscreen. A cloud-based server handles speech recognition, translation, and speech synthesis, using speech recognition software (e.g., Google Cloud Speech-to-Text API), translation software (e.g., Google Translate API), and speech synthesis software (e.g., Amazon Polly).
[1693] Program processing overview
[1694] Voice input
[1695] When a user launches an AI translation app on their smartphone or tablet, a "Speak" button appears on the home screen. When the user taps this button, the device activates the voice input method and begins collecting the user's speech.
[1696] Buffering and sending audio data
[1697] The collected voice data is buffered in the terminal, and when a certain buffer size is reached, the terminal transmits the voice data to the server. For this reason, a buffering means is installed.
[1698] Speech recognition and conversion to text data
[1699] The server converts the received voice data into text data using speech recognition software such as the Google Cloud Speech-to-Text API.
[1700] translation
[1701] The text data generated by speech recognition is translated into other languages using translation means on the server. Here, the Google Translate API is used. For example, if User A utters "Hello, where am I?", this is translated into French as "Bonjour, où sommes-nous?"
[1702] Speech synthesis and playback
[1703] The translated text data is converted back into voice data using a voice synthesis means, such as voice synthesis software like Amazon Polly, and the generated voice data is sent to the terminal and played back by the voice output means.
[1704] Specific examples
[1705] Use in tourist areas
[1706] For example, consider the case where User A (Japanese) is visiting a tourist spot in France. User A launches the "AI translation app," taps the "Speak" button, and asks the French guide (User B), "Hello, where is this?" The voice data is sent to the server, translated into French, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1707] Interactive tourist guide
[1708] When User A and User B are making sightseeing plans while looking at a tourist guide map, User A taps the "Screen Share" button to share the map, and their device starts capturing the screen. The captured screen data is sent to User B's device via the server. Furthermore, when User A indicates places they want to visit using touch gestures, the gesture data is also displayed in real time on User B's device via the server. This allows them to make plans while sharing information across language barriers.
[1709] Prompt Sentence Examples
[1710] A description of this system can be generated by inputting the following prompt sentence into the generative AI model:
[1711] This system allows users to input speech, translate it into other languages in real time, and output it as speech, as well as provide screen sharing and interactive gesture control functions between users. Please explain the speech input, speech recognition, translation, speech synthesis, and gesture recognition and sharing functions.
[1712] In this way, this system can effectively remove language barriers in international exchanges and travel, enabling intimate communication.
[1713] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1714] Step 1:
[1715] A user launches an "AI translation app" on an information display device (smartphone or tablet) and taps the "Speak" button displayed on the home screen. This activates the device's voice input means. The input is the user's speech (voice data), and the collected voice data is obtained as output.
[1716] Step 2:
[1717] The device starts collecting the user's speech through the microphone. The collected speech data is buffered within the device, and when a certain buffer size is reached, the buffered speech data is sent to the server. The input is the raw speech data of the user's speech, and the output is the buffered speech data sent to the server.
[1718] Step 3:
[1719] The server processes the received voice data and converts it into text data using a speech recognition tool. The input is the voice data sent from the device, and the output is text data. In this case, speech recognition software (e.g., Google Cloud Speech-to-Text API) is used.
[1720] Step 4:
[1721] The server translates the voice-recognized text data into another language using a translation tool. The input is the text data processed by the server, and the output is the text data translated into another language. Here, translation software (e.g., Google Translate API) is used.
[1722] Step 5:
[1723] The server converts the translated text data back into voice data using a speech synthesis tool. The input is the translated text data, and the output is synthesized voice data. In this case, speech synthesis software (e.g., Amazon Polly) is used.
[1724] Step 6:
[1725] The server sends synthesized voice data to the terminal. The input is the voice data from the server, and the terminal receives it as output.
[1726] Step 7:
[1727] The terminal plays back the received voice data using the voice output means. The input is the voice data received from the server, and the output is the voice played back through the speaker. This allows user B to hear user A's speech in his own language.
[1728] Step 8:
[1729] The user taps the "Screen Share" button, and the device starts capturing the screen. The input is the user's screen, and the output is the captured screen data.
[1730] Step 9:
[1731] The terminal sends the captured screen data to the server. The input is the captured screen data, and the output is the data sent to the server.
[1732] Step 10:
[1733] The server transfers the received screen data to the other party's terminal. The input is the transmitted screen data, and the output is the data transmitted to the other party's terminal.
[1734] Step 11:
[1735] The screen data received by the other party's device is displayed. The input is the received screen data, and the output is the screen displayed on the display. This allows user B to see user A's screen in real time.
[1736] Step 12:
[1737] The user performs a touch gesture (e.g., drawing a heart symbol) on the screen. The input is the user's touch operation, and the output is gesture data.
[1738] Step 13:
[1739] The device recognizes the touch gesture and sends the data to the server. The input is the touch gesture operation data, and the output is the recognized gesture data.
[1740] Step 14:
[1741] The server transmits the received gesture data to the other device. The input is the recognized gesture data, and the output is the data transmitted to the other device.
[1742] Step 15:
[1743] The gesture data received by the other party's device is displayed in real time. The input is the received gesture data, and the output is the gesture displayed on the display. This allows the gesture drawn by User A to be displayed on User B's screen as well.
[1744] (Application example 1)
[1745] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1746] Currently, language barriers are a major obstacle in situations requiring real-time multilingual communication. Language differences in particular hinder smooth interactions between viewers and streamers during live streaming, as well as communication between viewers themselves. Furthermore, screen sharing and interactive gesture features are not fully integrated. This lack of an environment allows users of different languages to smoothly communicate with each other and enjoy live content with a sense of unity.
[1747] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1748] In this invention, the server includes a voice input means for allowing a user to input voice, a voice recognition means for converting the voice input data into character data, and a translation means for translating the character data into another language, thereby enabling multilingual live chat.
[1749] "Voice input means" refers to a device or technology that allows a user to input voice.
[1750] "Speech recognition means" refers to technology that analyzes input voice data and converts it into text data.
[1751] "Translation means" refers to a technology that converts text data written in one language into text data written in another language.
[1752] "Speech synthesis means" refers to technology that analyzes text data and converts it into voice data.
[1753] "Audio output means" refers to technology for playing audio data through a device such as a speaker.
[1754] "Screen sharing means" refers to a technology that allows users to share the same screen in real time.
[1755] "Display transmission means" refers to a technology for transmitting the display contents of a terminal in use to another terminal.
[1756] "Gesture recognition means" refers to technology that detects touch gestures made on the screen and recognizes their content.
[1757] The "gesture transmission means" refers to a technique for transmitting recognized gesture information to another terminal.
[1758] "Multilingual live chat means" refers to technology that enables viewers to communicate in different languages in real time during live streaming.
[1759] "Translation display means" refers to technology that translates information sent using the multilingual live chat means and displays it as audio or text.
[1760] The present invention provides a system for supporting smooth communication between users, particularly in situations where real-time communication in multiple languages is required. Specific embodiments of this system will be described below.
[1761] System configuration
[1762] This system is composed of a server, a user terminal, and various software modules. Specific components include the following:
[1763] Voice input means
[1764] Voice recognition means
[1765] Translation tools
[1766] Voice synthesis means
[1767] Audio output means
[1768] Screen sharing method
[1769] Display and transmission means
[1770] Gesture Recognition Method
[1771] Gesture sending method
[1772] Multilingual live chat options
[1773] Translation display means
[1774] System Operation
[1775] 1. Voice input and recognition:
[1776] A user inputs voice using a device such as a smartphone or tablet. The voice input means collects voice data and transmits it to the voice recognition means.
[1777] The speech recognition means converts the collected speech data into text data, which is processed in real time.
[1778] 2. Translation and speech synthesis:
[1779] The translation means converts the text data into a language specified by the user. For example, if a translation from Japanese to English is required, the translation means is applied.
[1780] The translated text data is converted back into voice data by the voice synthesis means, and the voice data is reproduced through the voice output means.
[1781] 3. Screen sharing and interactive gestures:
[1782] The user can use the screen sharing means to share the screen with the other party in real time, and the display transmission means transmits the display content of the terminal currently in use to the other party's terminal.
[1783] The gesture recognition means detects touch gestures made on the screen and transmits the information to the other party's terminal. This is handled by the gesture transmission means.
[1784] 4. Multilingual Live Chat:
[1785] The multilingual live chat means supports real-time communication between viewers during live streaming. Information sent by viewers is translated through the translation display means and displayed to the other person. This allows smooth communication even between viewers who speak different languages.
[1786] Hardware and software used
[1787] Hardware:
[1788] Server: Responsible for data processing, translation, etc.
[1789] User devices: smartphones, tablets, head-mounted displays, etc.
[1790] software:
[1791] Flask: Used as a web server
[1792] SpeechRecognition: Used for speech recognition
[1793] Googletrans: used for text translation
[1794] Pyttsx3: Used for speech synthesis
[1795] Specific example explanation
[1796] For example, imagine a situation where a question can be sent in Japanese during a live streaming event, and the question can be translated into English in real time and spoken to the streamer. In this way, viewers and streamers can communicate smoothly even if they speak different languages.
[1797] Example prompt sentence:
[1798] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[1799] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1800] Step 1:
[1801] A user starts an application on a device such as a smartphone or tablet and inputs voice. The voice input means collects this voice data and buffers it in real time. The input is the user's voice, and it is temporarily stored in the device as voice data.
[1802] Step 2:
[1803] Voice data is sent to the server at fixed buffer intervals. The server passes this voice data to a voice recognition means and converts it into text data. The input is voice data and the output is text data. The server analyzes the voice data and outputs it as text data.
[1804] Step 3:
[1805] The server passes the recognized character data to the translation means and translates it into the specified language. The input is character data, and the output is translated character data. The server processes the character data to convert it into another language.
[1806] Step 4:
[1807] The server passes the translated text data to a speech synthesis means, which converts it into new speech data. The input is the translated text data, and the output is the new speech data. The speech synthesis means uses a specific speech synthesis algorithm to output the text data as speech data.
[1808] Step 5:
[1809] New audio data is sent from the server to the user's terminal and played through the audio output means. The input is the audio data and the output is the audio to be played. The terminal receives this audio data and outputs it as audio through a speaker.
[1810] Step 6:
[1811] When a user wishes to share their screen, the terminal starts capturing the screen using the screen sharing means and sends the captured data to the server. The input is the screen capture data, and the output is the screen data sent to the server. The server then relays this screen data to the terminals of other users.
[1812] Step 7:
[1813] The other user's device displays the screen data received from the server. The input is the received screen data, and the output is the screen display. The device analyzes the received data and displays it to the user.
[1814] Step 8:
[1815] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes its shape and movement. The input is touch gesture data, and the output is recognized gesture information. The device detects the touch operation and sends it to the server as gesture data.
[1816] Step 9:
[1817] The server transmits the recognized gesture information to the terminal of the other user and displays it in real time using a gesture transmission means. The input is the recognized gesture information and the output is the gesture to be displayed. The terminal displays the received gesture information on the screen.
[1818] Step 10:
[1819] A multilingual live chat means performs processing to support real-time communication between viewers. A translation display means translates messages sent by viewers in different languages and displays them on the terminal of the person interacting with them. The input is the viewer's message, and the output is a display of the translated message. The terminal presents this translated information to the viewer, enabling real-time communication.
[1820] When using a generative AI model for speech recognition or translation, the following example prompt can be set:
[1821] Example prompt sentence:
[1822] "Generate a server-side script that translates Japanese text into English in real time and plays the translation result as audio. The user uploads an audio file, and the audio output is played in English."
[1823] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1824] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[1825] 1. System Overview
[1826] The system has the following elements:
[1827] Voice input means
[1828] Voice recognition means
[1829] Translation tools
[1830] Voice synthesis means
[1831] Audio output means
[1832] Screen sharing method
[1833] Display and transmission means
[1834] Gesture Recognition Method
[1835] Gesture sending method
[1836] Emotion Engine
[1837] Means of transmitting emotions
[1838] 2. Program Processing Overview
[1839] Voice input
[1840] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1841] Speech Recognition and Translation
[1842] When the voice data arrives at the server, the server converts the voice into text data using a voice recognition means. This text data is then translated into another language by a translation means. For example, if user A asks "Hello, where am I?", this voice is translated into French as "Bonjour, où sommes-nous?"
[1843] Emotion recognition and reflection
[1844] In parallel with the voice input, the emotion engine analyzes the user's voice data and recognizes emotions. The recognized emotion information is sent to the server, and the translation means and speech synthesis means reflect this in the translation results and speech tone. For example, if the user's voice contains anger, the translated voice will also be adjusted to reflect an angry tone.
[1845] Speech synthesis and output
[1846] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the other party to hear and understand the translated voice.
[1847] Sharing emotional information
[1848] The emotional information recognized by the emotion engine is also sent to the other party's device via the emotion transmission means. The other party can visually check the other party's emotional state on the screen, which promotes understanding of emotions beyond language.
[1849] screen sharing
[1850] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This function allows users to see each other's screens.
[1851] Gesture Recognition
[1852] When a user makes a touch gesture while sharing a screen, the gesture recognition means detects the touch operation and recognizes its shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if user A draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1853] Specific examples
[1854] Example 1: Communication at tourist spots
[1855] User A (Japanese) is visiting a tourist spot in France. He wants to communicate with a local French guide (User B), but there is a language barrier. User A launches the "AI translation app," taps the "Speak" button, and asks, "Hello, where am I?" The voice data and emotion data are sent to the server, where they are translated into French with the emotion reflected, and sent to the guide's device. The guide listens to the voice and replies, "This is the Eiffel Tower." A similar process occurs in the reverse direction.
[1856] Example 2: Using an interactive tourist guide
[1857] User A and User B are making a sightseeing plan while looking at a tourist guide map. User A taps the "Screen Share" button to share the map and checks off the places they want to visit. The check marks are also displayed in real time on User B's device, and they discuss where to go. In this way, they can share information and communicate smoothly across language barriers. The emotion engine recognizes the user's emotions, and the emotional state is reflected in the displayed interface, allowing for a deeper understanding of each other's feelings.
[1858] This system effectively removes language barriers in international exchange and travel, enabling intimate communication that also includes emotions.
[1859] The processing flow will be explained below.
[1860] Step 1:
[1861] The user launches the "AI translation app" on their smartphone.
[1862] The device displays the home screen and renders interface elements such as the "Talk" and "Screen Share" buttons.
[1863] Step 2:
[1864] The user taps the "Speak" button.
[1865] The terminal activates the voice input means and starts collecting voice data.
[1866] Step 3:
[1867] The user begins speaking.
[1868] The terminal buffers the collected audio data and transmits it to the server at fixed buffer size intervals.
[1869] Step 4:
[1870] The server passes the received voice data to the voice recognition means.
[1871] The voice recognition means converts the voice data into character data.
[1872] Step 5:
[1873] The server passes the converted character data to the translation means.
[1874] A translation means translates the character data into another designated language.
[1875] Step 6:
[1876] The server passes the translated character data to the speech synthesis means.
[1877] A speech synthesis means converts the translated text into speech data.
[1878] Step 7:
[1879] The server transmits the generated voice data to User B's terminal.
[1880] User B's terminal receives the audio data and plays it back via the audio output means.
[1881] Step 8:
[1882] User B initiates reply.
[1883] User B's device collects the voice data and sends it to the server.
[1884] Step 9:
[1885] The server passes user B's voice data to a voice recognition means and converts it into text data.
[1886] The server translates the text data into User A's language.
[1887] Step 10:
[1888] The server converts the translated text data into audio data and sends it to User A's terminal.
[1889] User A's terminal receives the audio data and plays it back via the audio output means.
[1890] Step 11:
[1891] The user taps the "Screen Share" button.
[1892] The device will display a dialog asking permission to share the screen, and once permission is granted, it will begin capturing the screen.
[1893] Step 12:
[1894] The device sends the screen capture data to the server.
[1895] The server transfers the received screen capture data to User B's terminal.
[1896] Step 13:
[1897] The terminal of user B displays the received screen capture data.
[1898] User A and User B share the same screen in real time.
[1899] Step 14:
[1900] User A performs a touch gesture on the screen.
[1901] The device detects the touch gesture and obtains its position and shape.
[1902] Step 15:
[1903] The terminal transmits the acquired gesture data to the server.
[1904] The server transfers the gesture data to User B's device.
[1905] Step 16:
[1906] The gesture data received by the terminal of User B is displayed on the screen in real time.
[1907] User A and User B can communicate interactively.
[1908] Step 17:
[1909] The terminal transmits the user's voice data to the emotion engine.
[1910] The emotion engine analyzes the voice data and recognizes the user's emotions.
[1911] Step 18:
[1912] The emotion engine transmits the recognized emotion information to the server.
[1913] The server reflects the emotional information in the translation means and the voice synthesis means.
[1914] Step 19:
[1915] The server converts the translated text data and emotional information into voice data.
[1916] The voice synthesis means generates voice data that reflects the emotion and transmits it to the user B's terminal.
[1917] Step 20:
[1918] The terminal of user B receives the voice data reflecting the emotion and plays it back via the voice output means.
[1919] User B can understand User A's emotions.
[1920] Step 21:
[1921] The server transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[1922] The emotion information received by the terminal of user B is visually displayed on the screen.
[1923] Example 2
[1924] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1925] While conventional speech translation systems provide the ability to translate speech into other languages, they lack the ability to recognize user emotions and reflect them in translation results and speech output. This makes it difficult for users who speak different languages to convey emotional nuances, making smooth communication difficult. Furthermore, screen sharing and interactive gesture functions are limited, making it difficult for users to intuitively share information and communicate with each other.
[1926] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1927] In this invention, the server includes a speech recognition unit, a translation unit, a speech synthesis unit, an emotion recognition unit, and an emotion transmission unit. This makes it possible to recognize emotions in a user's speech data in real time and reflect those emotions in the translation results and voice output. Furthermore, by combining a screen sharing unit and a gesture recognition unit, intuitive information sharing and interactive communication between users is realized.
[1928] The "voice input means" is a device for inputting the voice spoken by the user as digital data.
[1929] "Speech recognition means" is a technology that analyzes input voice data and converts it into text data.
[1930] "Translation means" is a technology that converts text data in one language into another language.
[1931] "Speech synthesis means" is a technology that converts text data into voice data and outputs it as synthesized voice.
[1932] The "audio output means" is a device for reproducing the synthesized audio data and letting the user hear it.
[1933] "Screen sharing means" is a technology that allows multiple users to view the same screen content at the same time.
[1934] The "display transmission means" is a technique for transmitting the display contents of the terminal in use to another terminal.
[1935] A "gesture recognition means" is a technology that detects touch gestures on a screen and recognizes their movements and shapes.
[1936] The "gesture transmitting means" is a technique for transmitting recognized gesture information to another terminal.
[1937] "Emotion recognition means" is a technology that analyzes the user's voice data and identifies emotions.
[1938] "Emotion transmission means" is a technology for transmitting recognized emotion information to another terminal.
[1939] The "gesture display means" is a technique for displaying gesture information received by the terminal of the other party on its screen.
[1940] A "buffer" is an area that temporarily stores data and is used to adjust speed differences when sending and receiving data.
[1941] This invention is a system that allows users to input voice and translate it into other languages in real time, while also providing the ability to recognize the user's emotions and reflect them in the translation results and voice output, and is equipped with screen sharing and interactive touch gesture functions between users.
[1942] System configuration
[1943] The system has the following elements:
[1944] Voice input means
[1945] Voice recognition means
[1946] Translation tools
[1947] Voice synthesis means
[1948] Audio output means
[1949] Screen sharing method
[1950] Display and transmission means
[1951] Gesture Recognition Method
[1952] Gesture sending method
[1953] emotion recognition means
[1954] Means of transmitting emotions
[1955] Detailed Description
[1956] Voice input means
[1957] When a user launches an application on a smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and starts collecting the user's speech. This voice data is buffered within the device and sent to the server at regular intervals.
[1958] Voice recognition means
[1959] When the voice data arrives at the server, the server uses speech recognition software like the Google Cloud Speech-to-Text API to convert the speech into text. For example, a user might say, "Hello, where am I?" and it will be converted into text.
[1960] Translation tools
[1961] The server translates the obtained text data into other languages using a translation tool such as Google Translate API. For example, the Japanese phrase "Hello, where are you?" is translated into French as "Bonjour, où sommes-nous?"
[1962] Voice synthesis means
[1963] The translated text data is converted back into speech data using a speech synthesis tool (e.g., Amazon Polly), for example, to generate the translated French phrase "Bonjour, où sommes-nous?"
[1964] Audio output means
[1965] The voice data is sent to the other party's terminal and played back by the voice output means, allowing the other party to hear and understand the translated voice.
[1966] emotion recognition means
[1967] In parallel with the voice input, an emotion recognition means on the server (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotions. The recognized emotional information is reflected in the translation results and voice tone by the translation means and voice synthesis means. For example, if the user's voice contains anger, the translation results will also be adjusted to reflect an angry tone.
[1968] Means of transmitting emotions
[1969] The recognized emotional information is also sent to the other party's device via the emotion transmission means, allowing the other party to visually confirm the other party's emotional state on the screen, facilitating understanding of emotions beyond language.
[1970] Screen sharing and display transmission means
[1971] When a user taps the "Screen Share" button, the device starts capturing the screen. The screen data is sent to the other device via the server and displayed in real time. This allows users to see each other's screens.
[1972] Gesture Recognition Means and Gesture Transmission Means
[1973] When a user makes a touch gesture during screen sharing, the gesture recognition means detects the touch operation and recognizes the shape and movement. This gesture data is sent to the other party's device via the server and displayed in real time. For example, if a user draws a heart symbol on the screen, the same heart symbol will appear on the other party's screen.
[1974] Specific examples
[1975] Example 1: Communication at tourist spots
[1976] User A (Japanese) is visiting a tourist spot in a foreign country. He wants to communicate with a local guide (User B), but there is a language barrier. User A uses this system, taps the "Speak" button, and asks, "Hello, where am I?" This voice data and emotion data are sent to the server, where it is translated and the emotion is reflected in real time and sent to the guide's device. The guide listens to the voice and provides appropriate guidance.
[1977] Prompt Sentence Examples
[1978] "Please tell me how the voice input system works to communicate with foreigners at tourist spots. It also performs emotion recognition."
[1979] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1980] Step 1:
[1981] When a user launches an application on their smartphone or tablet, a "Speak" button or a "Screen Share" button appears on the home screen. When the user taps the "Speak" button, the device activates the voice input means and begins collecting the user's speech. The user's voice data is used as input, and the voice data is saved in a buffer as output. Specifically, the device's microphone captures the user's voice and stores the data in a buffer.
[1982] Step 2:
[1983] The device sends audio data to the server in fixed buffers. The audio data in the buffer is used as input, and the audio data is sent to the server as output. Specifically, when the collected audio data reaches a certain volume, the device uploads the data to the server.
[1984] Step 3:
[1985] The server analyzes the received voice data and converts the voice into text data using a voice recognition tool. The voice data arriving at the server is used as input, and converted text data is generated as output. Specifically, the server calls a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data.
[1986] Step 4:
[1987] The server translates the obtained text data into another language. The text data generated by the speech recognition means is used as input, and the translated text data is generated as output. Specifically, the server calls a translation API (e.g., Google Translate) to translate from Japanese to French, etc.
[1988] Step 5:
[1989] The server's emotion recognition means analyzes the user's voice data and recognizes emotions. The received voice data is used as input, and recognized emotion information is generated as output. Specifically, the server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the voice data.
[1990] Step 6:
[1991] The recognized emotional information is reflected in the translation results and voice tone through the translation means and speech synthesis means. The emotional information generated by the emotion recognition means and the translated text data are used as input, and synthetic speech data reflecting the emotion is generated as output. Specifically, the server uses a speech synthesis API (e.g., Amazon Polly) to generate speech data based on the text data and emotional information.
[1992] Step 7:
[1993] The synthesized voice data is sent to the other party's terminal and played by the voice output means. The synthesized voice data is used as input, and the voice is played back in an audible form for the user as output. Specifically, the server sends the synthesized voice data to the other party's terminal, and the terminal plays the voice through a speaker.
[1994] Step 8:
[1995] The recognized emotional information is also sent to the conversation partner's terminal through the emotion transmission means. The emotional information is used as input, and the emotional state is visually displayed on the screen of the conversation partner's terminal as output. Specifically, the server transmits the emotional information to the conversation partner's terminal, and the emotional state is displayed on the terminal's display.
[1996] Step 9:
[1997] When a user taps the "Screen Share" button, the device starts capturing the screen. User operations and screen data are used as input, and screen capture data is generated as output. Specifically, the device periodically takes screenshots and sends the data to the server.
[1998] Step 10:
[1999] The screen data is sent to the other party's device via the server and displayed in real time. The screen capture data sent from the device is used as input, and the screen displayed on the other party's device is updated in real time as output. Specifically, the server relays the screen data to the other party's device, and the device displays the data on its display.
[2000] Step 11:
[2001] When a user makes a touch gesture during screen sharing, the gesture recognition means of the device detects the touch operation. Touch gesture data is used as input, and recognized gesture data is generated as output. Specifically, the device detects touch operations on the screen with a sensor and recognizes their shape and movement.
[2002] Step 12:
[2003] The gesture data is sent to the other party's device via the server and displayed in real time. The recognized gesture data is used as input to generate the gestures displayed on the other party's device as output. Specifically, the server relays the gesture data to the other party's device, and the device displays the data on its display.
[2004] (Application example 2)
[2005] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2006] In communication between workers and robots in factories, there is a problem that instructions cannot be transmitted quickly and accurately between different languages, resulting in reduced work efficiency. Furthermore, it is difficult to improve safety because the robot's behavior cannot be adjusted according to the worker's emotional state. Furthermore, it is difficult to share instructions and information between multiple workers using screen sharing and interactive touch gestures.
[2007] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2008] In this invention, the server includes a voice input device for allowing a user to input voice data, a voice recognition device for converting the voice input data into text data, a translation device for translating the text data into another language, a voice synthesis device for converting the translated text data into voice data, a voice output device for outputting the synthesized voice data, a screen sharing device for sharing a screen between users, a display transmission device for sharing the display of the currently used terminal with the other terminal, a gesture recognition device for recognizing touch gestures on the screen, a gesture transmission device for transmitting the recognized gesture information to the other terminal, an emotion engine for recognizing an emotional state from the voice data, an emotion transmission device for transmitting the recognized emotion information to the other terminal, and an instruction device for instructing the work machine using the translated text data and voice data. This enables quick and accurate instruction transmission between different languages. Furthermore, safety is improved by adjusting the robot's operation according to the worker's emotional state. Furthermore, interactive screen sharing and touch gesture sharing of instructions and information between multiple workers is possible.
[2009] "Voice input means" refers to a device or system that allows a user to input voice.
[2010] The "voice recognition means" is a device or system that converts voice data acquired by the voice input means into character data.
[2011] The "translation means" is a device or system that translates the character data generated by the speech recognition means into another language.
[2012] The "voice synthesis means" is a device or system that converts the translated text data generated by the translation means into voice data.
[2013] The "audio output means" is a device or system that outputs the audio data generated by the audio synthesis means to the user.
[2014] A "screen sharing means" is a device or system for sharing screen information among multiple users.
[2015] The "display transmission means" is a device or system that transmits the display of the terminal in use to the terminal of the other party.
[2016] A "gesture recognition means" is a device or system that recognizes touch gestures on a screen and processes them as data.
[2017] The "gesture transmitting means" is a device or system that transmits recognized gesture information to the terminal of the other party.
[2018] An "emotion engine" is a device or system for recognizing a user's emotional state from voice data.
[2019] The "emotion transmission means" is a device or system that transmits the emotion information recognized by the emotion engine to the terminal of the conversation partner.
[2020] The "instruction means" is a device or system that issues instructions to the work machine based on the translated character data and voice data.
[2021] This invention realizes a multilingual interactive interpretation system between workers and robots in a factory. This system is composed of a voice input means, a voice recognition means, a translation means, a voice synthesis means, a voice output means, a screen sharing means, a display transmission means, a gesture recognition means, a gesture transmission means, an emotion engine, an emotion transmission means, and an instruction means.
[2022] System Overview
[2023] Audio Input:
[2024] The server activates a voice input device, such as a smartphone or tablet, for the user to input voice. When the user starts speaking, this voice data is buffered in the device and sent to the server at regular intervals.
[2025] Speech Recognition and Translation:
[2026] The voice data received by the server is converted into text data using a voice recognition unit, and then this text data is translated into another language by a translation unit, for example, from Japanese to English.
[2027] Emotion recognition and reflection:
[2028] In parallel with voice input, the emotion engine analyzes the voice data and recognizes the user's emotional state. The recognized emotional information is sent to the server, where the translation means and voice synthesis means reflect it in the translation results and voice tone. If the user is tired, the robot's movement speed will be slowed down, for example.
[2029] Speech synthesis and output:
[2030] The translated text data is converted back into voice data using a voice synthesis means. This voice data is sent to the other party's terminal and played back by a voice output means. This allows the robot to understand and carry out the worker's instructions.
[2031] Sharing emotional information:
[2032] The emotion information recognized by the emotion engine is also sent to the terminal of the conversation partner through the emotion transmission means. The emotional state of the worker is displayed, promoting mutual understanding.
[2033] Screen sharing and gesture recognition:
[2034] To share a screen between users, a device initiates a screen capture and sends the captured data to the other device. Furthermore, touch gestures on the screen are recognized and the gesture data is sent to the other device, enabling interactive operation.
[2035] Specific examples
[2036] For example, in a factory, Worker A instructs a robot in Japanese, saying, "Please transport this part." This voice is translated into English in real time and transmitted to the robot. At the same time, the robot recognizes that Worker A is tired from his voice and switches its operating speed to slow mode. Worker B can then use the screen sharing function to instruct the robot on the specific destination with touch gestures, achieving efficient communication.
[2037] Example prompt sentence:
[2038] "Based on the invention below, we will design a multilingual interactive interpretation system for factory robots. It is a system that allows users to input speech and translate it into other languages in real time. Furthermore, it recognizes the user's emotions during the translation process and reflects those emotions in the translation results. It will also have functions that allow users to share screens and use interactive touch gestures. As a specific application example, when workers in a factory give instructions to a robot, it will use voice recognition and translation to make it multilingual. Also, please consider a system that recognizes the emotional state of the worker and adjusts the robot's movement speed and response content according to the situation."
[2039] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2040] Step 1:
[2041] The user activates the voice input means to input voice.
[2042] Input: User speech.
[2043] Output: The input audio data.
[2044] Specific operation: The user taps the "Speak" button on their smartphone or tablet to begin voice input. The device collects voice data through the microphone.
[2045] Step 2:
[2046] The device sends the collected voice data to the server.
[2047] Input: Audio data.
[2048] Output: The audio data sent to the server.
[2049] Specific operation: The collected audio data is sent from the terminal to the server at fixed buffer intervals.
[2050] Step 3:
[2051] The server converts the voice data into character data using a voice recognition means.
[2052] Input: The audio data sent to the server.
[2053] Output: Recognized speech text data.
[2054] Specific operation: The server calls a speech recognition API and converts the voice data into text data, for example, using the Google Speech-to-Text API.
[2055] Step 4:
[2056] The server uses a translation means to translate the character data into another language.
[2057] Input: Speech-recognized text data.
[2058] Output: The translated text data.
[2059] Specific operation: The server calls a translation API to translate the text data into another specified language. For example, it uses the Google Translate API.
[2060] Step 5:
[2061] The server uses an emotion engine to analyze the voice data and recognize the emotional state.
[2062] Input: The original audio data.
[2063] Output: Recognized emotion information.
[2064] Specific operation: Calls an emotion recognition API and recognizes emotions from features such as energy, pitch, and tempo. For example, it uses IBM Watson Tone Analyzer.
[2065] Step 6:
[2066] The server synthesizes voice data based on the translated text data and emotional information.
[2067] Input: translated text data, recognized emotion information.
[2068] Output: Synthesized audio data.
[2069] Specific operation: Calls a speech synthesis API and generates speech data that reflects emotional information. For example, it uses Amazon Polly.
[2070] Step 7:
[2071] The server transmits the synthesized voice data to the terminal.
[2072] Input: Synthesized speech data.
[2073] Output: The audio data sent to the device.
[2074] Specific operation: The server transmits the generated voice data to the other party's terminal.
[2075] Step 8:
[2076] The terminal uses the audio output means to play back the received audio data.
[2077] Input: Received audio data.
[2078] Output: The audio that is played.
[2079] Specific operation: The other party's device plays the synthesized voice through the speaker.
[2080] Step 9:
[2081] The user activates the screen sharing tool and starts capturing the screen.
[2082] Input: User operations, device screen data.
[2083] Output: Real-time captured screen data.
[2084] Specific operation: The user taps the "Screen Share" button and selects the screen they want to share. The device captures that screen in real time and sends it to the server.
[2085] Step 10:
[2086] The server transmits the screen data to the terminal of the other party.
[2087] Input: Screen data captured in real time.
[2088] Output: Screen data sent to the other device.
[2089] Specific operation: The server sends the captured screen data to the other device.
[2090] Step 11:
[2091] The terminal of the other party displays the received screen data.
[2092] Input: Received screen data.
[2093] Output: The screen that is displayed.
[2094] Specific operation: The screen data received by the other party's device is displayed on the screen.
[2095] Step 12:
[2096] A user performs a touch gesture using the gesture recognition means.
[2097] Input: User touch actions.
[2098] Output: Recognized gesture data.
[2099] Specific operation: When a user performs a touch gesture on the screen, the device recognizes the touch operation and generates gesture data.
[2100] Step 13:
[2101] The server transmits the recognized gesture data to the other terminal of the conversation partner.
[2102] Input: Recognized gesture data.
[2103] Output: Gesture data sent to the other device.
[2104] Specific operation: The server sends the recognized gesture data to the other device.
[2105] Step 14:
[2106] The gesture data received by the terminal of the conversation partner is displayed on the screen.
[2107] Input: Gesture data sent to the other device.
[2108] Output: Gesture data displayed on the screen.
[2109] Specific operation: The gesture data received by the other device is displayed on the screen in real time.
[2110] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2111] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2112] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2113] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2114] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2115] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2116] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2117] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2118] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2119] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2120] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2121] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2122] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2123] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2124] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2125] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2126] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2127] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2128] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2129] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2130] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2131] The following is further disclosed regarding the above embodiment.
[2132] (Claim 1)
[2133] a voice input means for allowing a user to input voice;
[2134] a voice recognition means for converting voice input data into character data;
[2135] a translation means for translating the character data into another language;
[2136] a speech synthesis means for converting the translated text data into speech data;
[2137] an audio output means for outputting the synthesized audio data;
[2138] a screen sharing means for sharing a screen between users;
[2139] a display transmission means for sharing the display of the terminal being used with the terminal of the other party;
[2140] a gesture recognition means for recognizing a touch gesture on the screen;
[2141] a gesture transmitting means for transmitting the recognized gesture information to a terminal of a conversation partner;
[2142] A system including:
[2143] (Claim 2)
[2144] 2. The system according to claim 1, further comprising a gesture display means for displaying the gesture information received by the terminal of the conversation partner on a screen.
[2145] (Claim 3)
[2146] 2. The system according to claim 1, wherein the voice input means collects voice data in real time, and the voice recognition means transmits the collected voice data to the server at fixed buffer intervals.
[2147] (Claim 4)
[2148] 2. The system according to claim 1, wherein the translation means translates the data converted into character data by the speech recognition means into another designated language in real time.
[2149] (Claim 5)
[2150] 2. The system according to claim 1, wherein the voice output means reproduces the translated voice data in real time and allows the other person to hear it.
[2151] "Example 1"
[2152] (Claim 1)
[2153] a voice input means for allowing a user to input voice;
[2154] a voice recognition means for converting voice input data into character data;
[2155] a translation means for translating the character data into another language;
[2156] a speech synthesis means for converting the translated text data into speech data;
[2157] an audio output means for outputting the synthesized audio data;
[2158] a screen sharing means for sharing a display between users;
[2159] a display transmission means for sharing the display of the currently used information display device with the other party's information display device;
[2160] a gesture recognition means for recognizing an operation gesture on the display;
[2161] a gesture transmitting means for transmitting the recognized gesture information to the information display device of the other party;
[2162] a buffering means for collecting voice data in real time and transmitting the data to a processing unit in units of a certain number of buffers;
[2163] a translation output means for synchronously converting the voice data received by the processing device into character data and translated character data, and reproducing the data as voice data;
[2164] A system including:
[2165] (Claim 2)
[2166] 2. The system according to claim 1, further comprising a gesture display means for displaying the gesture information received by the information display device of the conversation partner on a display.
[2167] (Claim 3)
[2168] 2. The system according to claim 1, wherein the speech recognition means transmit...
Claims
1. a voice input means for allowing a user to input voice; a voice recognition means for converting voice input data into character data; a translation means for translating the character data into another language; a speech synthesis means for converting the translated text data into speech data; an audio output means for outputting the synthesized audio data; a screen sharing means for sharing a screen between users; a display transmission means for sharing the display of the terminal being used with the terminal of the other party; a gesture recognition means for recognizing a touch gesture on the screen; a gesture transmitting means for transmitting the recognized gesture information to a terminal of a conversation partner; A system including:
2. 2. The system according to claim 1, further comprising gesture display means for displaying the gesture information received by the terminal of the other party on a screen.
3. 2. The system according to claim 1, wherein the voice input means collects voice data in real time, and the voice recognition means transmits the collected voice data to the server at fixed buffer intervals.
4. 2. The system according to claim 1, wherein the translation means translates the data converted into character data by the speech recognition means into another designated language in real time.
5. 2. The system according to claim 1, wherein the voice output means reproduces the translated voice data in real time and allows the other person to hear it.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A