System
The system addresses the challenge of maintaining natural speech and detecting sensitive content in cross-language communication by translating user voice in real-time while preserving voice characteristics and providing alerts, facilitating smooth and efficient international conversations.
Patent Information
- Application Number
- JP2024119123
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional translation systems struggle with maintaining the naturalness of speech during translation, leading to incongruity and difficulty in international communication, and lack real-time detection of sensitive content and facilitation of smooth conversations.
A system that converts user voice into a digital signal, translates it in real-time to a specified language while preserving the original voice characteristics, detects sensitive content, and provides alerts or suggestions to facilitate communication.
Enables natural and smooth cross-language communication, supports the detection of sensitive content, and ensures efficient progress of meetings by maintaining the user's voice characteristics and providing real-time alerts and suggestions.
Smart Images

Figure 2026018062000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Describe the "problem that the invention aims to solve" and the "means for solving the problem."
[0005] ---
[0006] In modern international exchange and business settings, communication in multiple languages is required, but communication via an interpreter can be problematic, as it can be difficult to identify the speaker and time lags can occur. Furthermore, conventional translation systems have the problem of making natural conversation difficult because the audio produced does not sound like the speaker's own voice, creating a sense of incongruity. The objective of this invention is to solve these problems and realize natural international exchange and business communication without the need for language barriers. [Means for solving the problem]
[0007] In order to solve the above problems, the present invention provides the following means: a system including: means for converting a user's voice into a digital signal, means for transmitting the digital signal to a server, means for converting the voice digital signal into character string text on the server, means for translating the converted text into a specified language in real time, means for converting the translated text into voice data that retains the characteristics of the original voice, means for transmitting the voice data to a terminal, and means for playing the voice data on the terminal; and a system including means for detecting sensitive content, means for generating a sensitive alert and notifying the terminal, means for generating suggestions or instructions to support the progress of the conference, and means for transmitting the suggestions or instructions to the terminal and notifying the user.
[0008] ---
[0009] Understood. Below are definitions of important terms included in the claims.
[0010] ---
[0011] "User" refers to an individual or organization that uses the system to input speech and output translated results.
[0012] "Terminal" refers to a device that converts a user's voice into a digital signal and transmits it to a server, and a device that receives and plays back the voice data transmitted from the server.
[0013] "Server" refers to a central processing unit that processes voice data sent from a terminal and performs voice recognition, text translation, and voice synthesis.
[0014] "Audio digital signal" refers to a user's voice converted from an analog signal to a digital signal.
[0015] "Text string" refers to text data obtained by analyzing a digital audio signal.
[0016] "Translation" refers to the process of converting text entered in one language into another specified language.
[0017] "Audio Data" refers to the audio signal generated from the translated text using speech synthesis technology.
[0018] "Sensitive content" refers to confidential information, inappropriate language, or content that is sensitive due to language or culture.
[0019] "Sensitive alert" refers to a notification that alerts the user when sensitive content is detected.
[0020] "Suggestions and instructions to support the progress of the meeting" refers to guidance and instructions generated to facilitate the smooth progress of the meeting.
[0021] --- [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] Understood. Below is the "Mode for carrying out the invention".
[0044] ---
[0045] The present invention relates to a translation multi-assist AI system that overcomes language barriers in international communication and enables natural conversation. Specific embodiments of the present invention will be described below.
[0046] Overall system configuration
[0047] This system consists of a user, a terminal, and a server. The user inputs speech, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into speech data that retains the user's speech characteristics and sent to the terminal. The terminal plays back the received speech data and provides the translation result to the user.
[0048] Program processing explanation
[0049] User voice input
[0050] The user speaks into the device's microphone. For example, if the user says "Good morning" in Japanese, this speech is captured by the device.
[0051] Speech recognition and text conversion
[0052] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this example, the audio "Good morning" is converted into the text "Good morning."
[0053] Real-time translation
[0054] The server passes the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0055] Text-to-speech translation
[0056] The server analyzes the user's voice characteristics and uses speech synthesis technology to convert the translated text into audio data in the user's voice. The server synthesizes "Good morning" in English, preserving the pitch, tone, and speed of the user's voice.
[0057] Sending and outputting synthesized speech
[0058] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0059] Specific examples
[0060] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0061] Sensitive Alert Function
[0062] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[0063] Facilitator function
[0064] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0065] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[0066] ---
[0067] The processing flow will be explained below.
[0068] Understood. Below I will explain the process in concrete steps.
[0069] ---
[0070] Step 1:
[0071] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0072] Step 2:
[0073] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[0074] Step 3:
[0075] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text as "Good morning."
[0076] Step 4:
[0077] The server sends the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0078] Step 5:
[0079] The server converts the translated text into audio data that preserves the user's voice characteristics. Specifically, it analyzes the pitch, tone, and speed of the user's original speech and uses speech synthesis technology to generate "Good morning" in the user's voice.
[0080] Step 6:
[0081] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[0082] Step 7:
[0083] The device decodes the received audio data and plays it through the speaker, allowing the user to hear the translated "Good morning" in their own voice.
[0084] Step 8:
[0085] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[0086] Step 9:
[0087] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[0088] Step 10:
[0089] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[0090] Step 11:
[0091] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[0092] ---
[0093] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[0094] Example 1
[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0096] In international communication, it is difficult to achieve natural and smooth conversations between people who speak different languages. Furthermore, when conversation content contains sensitive information, it is important to have a system that can appropriately process and notify the content. Furthermore, a system that supports the progress of meetings and important communications is also needed.
[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0098] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server via a communication network; means for converting the voice digital signal into character string text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the characteristics of the original voice; means for transmitting the voice data to a terminal via a communication network; means for playing the voice data on the terminal; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of the conference; and means for transmitting the suggestions or instructions to the terminal and notifying the user. This enables natural and smooth conversations between people speaking different languages, enables appropriate processing and notification of sensitive content, and supports the progress of conferences and important communications.
[0099] Understood. Below are definitions of important terms included in the claims mentioned above.
[0100] ---
[0101] A "digital signal" refers to a signal that represents analog audio data in numerical form.
[0102] "Communication network" refers to communication infrastructure such as the Internet or local networks used to send and receive data.
[0103] A "server" refers to a computer system that receives and processes requests from users over a network.
[0104] "Speech recognition API" refers to a programming interface for receiving voice data as input and converting it into corresponding text data.
[0105] "Text string" refers to text data converted by speech recognition.
[0106] "Translation API" refers to a program interface for converting input text data into another specified language.
[0107] "Speech features" refer to individual characteristics of speech such as pitch, tone, and rate.
[0108] "Audio Data" means an audio file represented in digital form.
[0109] "Sensitive alert" refers to a warning that notifies the user when sensitive content is detected.
[0110] "Suggestions and Instructions" refers to recommendations and instructions generated to help facilitate conversation or communication.
[0111] "Terminal" refers to a device (e.g., a smartphone or PC) that a user directly operates and that communicates with a server.
[0112] ---
[0113] These definitions are intended to clarify the meaning of key terms in the claims.
[0114] The present invention is a system that translates a user's voice input in real time and outputs the translated text as voice data that retains the original voice characteristics. This system is composed of a user, a terminal, and a server.
[0115] The user speaks into the microphone of the device. For example, when the user says "Good morning," the voice is captured by the device. The device converts the voice into a digital signal through the microphone and transmits the digital signal to the server via a communication network.
[0116] The server converts the received digital signal into a string of text using a speech recognition API (e.g., Google Cloud Speech-to-Text API). In this example, the speech "Good morning" is converted into the text "Good morning."
[0117] The server then passes the converted text to a translation API (e.g., Google Cloud Translation API) for real-time translation into the specified language (e.g., English). Here, the text "Ohayo gozaimasu" is translated into English as "Good morning."
[0118] The server uses speech synthesis technology (e.g., Amazon Polly) to convert the translated text into speech that retains the characteristics of the original voice, so that the translated "Good morning" is converted into speech that reflects the pitch, tone, and rate of the user's voice.
[0119] The server then sends the generated voice data to the device via a communication network. The device then decodes the received voice data and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0120] Furthermore, this system has the function of detecting sensitive content. The server analyzes the conversation content, and if it detects sensitive content, it generates a sensitive alert and notifies the terminal. The terminal then presents this alert to the user in text or voice format, allowing the user to respond appropriately.
[0121] The server also has the ability to generate suggestions and instructions to help meetings and important communications proceed smoothly. For example, if a conversation stalls, the server will automatically generate a suggestion such as, "Shall we move on to the next topic?" This allows the device to notify the user of the suggestion or instruction, helping to ensure the smooth progress of the meeting.
[0122] As a concrete example, consider a situation in which a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0123] An example of a prompt sentence is, "A Japanese-speaking user says 'Good morning' into the terminal. Explain how this speech is converted into a digital signal, processed by the server, and finally played back as English speech."
[0124] As described above, the present invention enables natural and smooth conversations between people who speak different languages, and enables sensitive content to be appropriately handled and communicated, thereby supporting the progress of meetings and important communications.
[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0126] Understood. Below, we will divide the processing flow of the system program into specific processing steps and explain the specific operations including input and output.
[0127] ---
[0128] Step 1:
[0129] Capture audio input
[0130] The user speaks into the device's microphone.
[0131] Input: User speech (e.g. "Good morning").
[0132] Device behavior: The device uses a microphone to capture the user's voice.
[0133] Output: Analog audio signal.
[0134] Step 2:
[0135] Digital conversion of analog audio signals
[0136] The terminal converts the captured analog voice signal into a digital signal.
[0137] Input: Analog audio signal.
[0138] Terminal operation: The A / D converter inside the terminal converts the audio signal into a digital signal and records it as time-series data.
[0139] Output: Digital audio signal.
[0140] Step 3:
[0141] Transmitting digital audio signals
[0142] The terminal transmits the digital audio signal to the server over a communication network.
[0143] Input: Digital audio signal.
[0144] Terminal operation: The voice signal is converted into packet data and sent to a server via the Internet through a communications module.
[0145] Output: Audio data packets to the server.
[0146] Step 4:
[0147] Speech recognition and text conversion
[0148] The server inputs the received digital voice signal into a voice recognition API and converts it into text.
[0149] Input: Digital audio signal.
[0150] Server operation: A speech recognition API (e.g., a speech recognition cloud service) converts the speech data into text data.
[0151] Output: String text (e.g. "Good morning").
[0152] Step 5:
[0153] Text translation
[0154] The server passes the converted text to the translation API, which translates it into the specified language.
[0155] Input: String text (e.g. "Good morning").
[0156] Server action: Translate the text into the specified language (e.g., English) using a translation API (e.g., a translation cloud service).
[0157] Output: The translated text (e.g. "Good morning").
[0158] Step 6:
[0159] Speech synthesis of translated text
[0160] The server converts the translated text into voice data using voice synthesis technology.
[0161] Input: The translated text (e.g. "Good morning") and the user's voice characteristics.
[0162] Server operation: A speech synthesis API (e.g., a speech synthesis cloud service) converts text into speech data that retains the user's speech characteristics.
[0163] Output: Synthetic speech data.
[0164] Step 7:
[0165] Sending audio data
[0166] The server transmits the generated voice data to the terminal via a communication network.
[0167] Input: Synthetic speech data.
[0168] Server operation: Converts voice data into packet data and sends it to the terminal via the Internet through the communication module.
[0169] Output: Audio data packets destined for the device.
[0170] Step 8:
[0171] Playing audio data
[0172] The terminal plays the received audio data through the speaker.
[0173] Input: Synthetic speech data.
[0174] Device operation: The decoded audio data is converted into an analog signal using a digital-to-analog converter and played through the speaker.
[0175] Output: The speech the user hears (e.g. "Good morning").
[0176] ---
[0177] This concludes the explanation of the specific processing steps of the program. Each step clearly shows the type of data processing and calculation that will be performed based on the input data. This system makes it possible to have natural conversations that transcend language barriers.
[0178] (Application example 1)
[0179] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0180] In international communication, there is a demand for systems that support natural conversations between users who speak different languages. However, existing systems have difficulty preserving the characteristics of speech while translating, and lack technology to support communication between various languages in autonomous vehicles. Furthermore, they lack the ability to detect and notify sensitive content in real time, and the ability to support the progress of meetings.
[0181] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0182] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server; means for converting the voice digital signal into text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the original voice characteristics; means for transmitting the voice data to a terminal; means for playing the voice data on the terminal; means for supporting communication between users in various languages within an autonomous vehicle; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of a meeting; means for transmitting the suggestions or instructions to the terminal and notifying the user; and a function for translating voice in real time using a generative AI model, installed on a smartphone, smart glasses, or head-mounted display. This supports natural communication between users who speak different languages and enables smooth communication across language barriers within an autonomous vehicle. Furthermore, the functions for detecting and notifying sensitive content and supporting the progress of a meeting enable safe and efficient communication.
[0183] The "means for converting a user's voice into a digital signal" is a device or program for converting a user's voice from an analog signal into a digital signal.
[0184] The "means for transmitting the digital signal to the server" refers to a communication device or program for transmitting the converted digital signal to the server via a network.
[0185] The "means for converting the digital audio signal into a character string text on the server" refers to software or a program for converting the digital audio signal into a text format on the server using voice recognition technology.
[0186] "Means for translating converted text into a specified language in real time" means a technology or program for instantly translating text into another specified language.
[0187] "Means for converting translated text into speech data that preserves the characteristics of the original speech" refers to technology or a program for converting translated text into speech data in a form that preserves the user's speech characteristics.
[0188] The "means for transmitting the voice data to the terminal" is a communication device or a program for transmitting the generated voice data to the user's terminal via a network.
[0189] The "means for reproducing the audio data on the terminal" refers to hardware or software for reproducing the audio data received on the terminal.
[0190] "Means for supporting communication between users in various languages within an autonomous vehicle" refers to technology or a program that enables smooth communication between passengers who speak different languages within an autonomous vehicle.
[0191] "Sensitive content detection means" means an algorithm or program that identifies sensitive or inappropriate content in real time.
[0192] The "means for generating a sensitive alert and notifying the terminal" is a program for generating a warning for detected sensitive content and notifying the user's terminal of the warning.
[0193] The "means for generating suggestions and instructions to support the progress of a meeting" refers to an algorithm or program for generating appropriate suggestions and instructions so that the meeting can proceed smoothly.
[0194] The "means for transmitting the proposal or instruction to the terminal and notifying the user" refers to a communication device or program for transmitting the generated proposal or instruction to the user's terminal and notifying the user of it.
[0195] "A feature installed on a smartphone, smart glasses, or head-mounted display that uses a generative AI model to translate speech in real time" refers to software or an application that utilizes the AI model built into these devices to provide real-time speech translation functionality.
[0196] The present invention relates to a translation multi-assist AI system for supporting natural conversation between users who speak different languages. Hereinafter, an embodiment of the present invention will be specifically described.
[0197] Overall system configuration
[0198] This system consists of a user, a terminal, the environment inside the autonomous vehicle, and a server. The user inputs voice through the terminal's microphone, and the terminal converts the voice into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language in real time. The translated text is then converted back into voice data that retains the user's voice characteristics and sent to the terminal. Finally, the terminal plays back the received voice data and provides the translation result to the user.
[0199] Program processing explanation
[0200] Audio input and digital signal conversion
[0201] The user speaks into the microphone of the device. For example, when the user says, "Take me to the airport, please," this voice is captured by the microphone of the device. The device converts the voice into a digital signal, which is then transmitted to the server via the network.
[0202] Speech recognition and text conversion
[0203] The server uses a speech recognition API (Google's speech recognition API is a specific example) to convert digital voice signals into text. This technology converts the speech "Please take me to the airport" into the text "Please take me to the airport."
[0204] Real-time translation
[0205] The server passes the converted text to a translation API (e.g., Google Translate API) and translates it into the specified language (e.g., English) in real time. In this example, "Kuukou de onegaishimasu" is translated to "Please go to the airport."
[0206] Text-to-speech translation
[0207] The server analyzes the user's voice characteristics and converts the translated text into audio data in the user's voice using speech synthesis technology (e.g., Google Text-to-Speech). The server generates audio data saying "Please go to the airport" in English, while retaining the user's voice characteristics.
[0208] Sending and playing audio data
[0209] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Please go to the airport" translated in their own voice.
[0210] Sensitive Alert Function
[0211] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitive alert in real time and notifies the device. This alert is presented in text or audio format, allowing the user to respond appropriately.
[0212] Facilitator function
[0213] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0214] Specific examples
[0215] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Please go to the airport," this speech is sent to the server via the device. The server converts the speech into text "Please go to the airport" and translates it into English "Please go to the airport." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Please go to the airport" in user (A)'s voice.
[0216] Prompt Sentence Examples
[0217] "Create an AI model that automatically translates what is said in the car into English. Convert the incoming Japanese speech into text, translate it into English using an API, and then convert it back into speech and play it back."
[0218] Thus, the present invention is a system that provides a natural translation function between users who speak different languages, detects sensitive content, and supports the progress of meetings.
[0219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0220] Step 1:
[0221] The user speaks into a terminal inside the autonomous vehicle. The terminal's microphone captures the user's voice and converts the analog voice signal into a digital signal. The input is the user's voice, and the output is a digital voice signal. Specifically, the terminal's voice input device captures the voice, and the built-in digital signal processor digitizes it.
[0222] Step 2:
[0223] The terminal transmits the converted digital signal to the server via the network. The input is a digital audio signal, and the output is the transfer of the digital signal to the server. Specifically, the terminal's communication module (e.g., Wi-Fi or mobile data communication) transmits the digital signal to the server.
[0224] Step 3:
[0225] The server converts the received digital voice signal into a string of text using a speech recognition API. The input is a digital voice signal, and the output is text data such as "Take me to the airport, please." Specifically, the server calls speech recognition software (e.g., Google Speech-to-Text API) and converts the voice signal into a string of text.
[0226] Step 4:
[0227] The server translates text data into the specified language in real time using a translation API. The input is text data, and the output is translated text such as "Please go to the airport." Specifically, the server uses translation software (e.g., Google Translate API) to convert the text into the specified language.
[0228] Step 5:
[0229] The server converts the translated text into audio data using speech synthesis technology. The input is the translated text, and the output is audio data. Specifically, the server uses speech synthesis software (e.g., Google Text-to-Speech API) to convert the text into audio data that retains the user's voice characteristics.
[0230] Step 6:
[0231] The server transmits the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal. In concrete terms, the communication module of the server transmits the voice data to the terminal.
[0232] Step 7:
[0233] The device decodes the received audio data and plays it through the speaker. The input is audio data and the output is audio output from the speaker. Specifically, the device's audio playback device decodes the audio data and plays it as audio through the speaker.
[0234] Step 8:
[0235] The server analyzes the content of the conversation and detects sensitive content. The input is voice or text data, and the output is a sensitive alert. Specifically, the server uses natural language processing technology to analyze the content of the conversation and detect whether it contains sensitive information. If it detects sensitive information, it generates an alert.
[0236] Step 9:
[0237] When a sensitive alert is generated, the server notifies the terminal of the alert. The input is the sensitive alert, and the output is a notification to the terminal. Specifically, the server's notification system sends the alert to the terminal and causes the terminal to display the notification.
[0238] Step 10:
[0239] To help facilitate the progress of the meeting, the server analyzes the conversation and generates appropriate suggestions and instructions. The input is conversation data, and the output is suggestions and instructions. Specifically, the server uses a generative AI model to understand the context of the conversation and generate appropriate suggestions and instructions.
[0240] Step 11:
[0241] The server sends the generated suggestions and instructions to the terminal and notifies the user. The input is the suggestions and instructions, and the output is a presentation to the terminal. Specifically, the instruction data is sent from the server to the terminal, and a notification is displayed on the terminal to the user.
[0242] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0243] Understood. Below is the "Mode for carrying out the invention" regarding the invention that combines an emotion engine.
[0244] ---
[0245] The present invention relates to a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[0246] Overall system configuration
[0247] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[0248] Program processing explanation
[0249] User voice input
[0250] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0251] Speech recognition and text conversion
[0252] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this case, "Good morning" is converted into text.
[0253] Recognition of emotional states
[0254] The server uses an emotion engine to recognize the user's emotional state from the input speech. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes a "positive" emotional state.
[0255] Real-time translation
[0256] The server sends the converted text and the recognized emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0257] Text-to-speech translation
[0258] The server analyzes the user's voice characteristics and the recognized emotional state, and uses speech synthesis technology to convert the translated text into audio data in the user's voice and emotional state, where "Good morning" is generated in the user's voice and in the recognized "positive" tone.
[0259] Sending and outputting synthesized speech
[0260] The server sends the generated voice data to the device, which then decodes the received voice data and plays it back through the speaker. This allows the user to hear "Good morning" translated in their own voice and with their own emotions.
[0261] Specific examples
[0262] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. This is then translated into English as "Good morning," and the translated text is synthesized using speech synthesis based on the voice and emotional characteristics of user (A). The speech data sent from the server to the device is played back, and user (B) hears "Good morning" in user (A)'s voice with a "positive" tone.
[0263] Sensitive Alert Function
[0264] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[0265] Facilitator function
[0266] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0267] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice and emotions, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[0268] ---
[0269] The processing flow will be explained below.
[0270] Understood. Below, I will explain the process flow for the invention that combines the emotion engine, broken down into specific steps.
[0271] ---
[0272] Step 1:
[0273] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0274] Step 2:
[0275] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[0276] Step 3:
[0277] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text "Good morning."
[0278] Step 4:
[0279] The server passes the recognized text to the emotion engine to recognize the user's emotional state. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes the user's emotional state as "positive."
[0280] Step 5:
[0281] The server sends the text with the emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0282] Step 6:
[0283] The server generates voice data that reflects the user's voice characteristics and emotional state based on the translated text and the recognized emotional state. Specifically, it generates "Good morning" in the user's voice while maintaining a "positive" tone.
[0284] Step 7:
[0285] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[0286] Step 8:
[0287] The device decodes the received audio data and plays it through the speaker, allowing the user to hear "Good morning" translated in their own voice and emotional state.
[0288] Step 9:
[0289] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[0290] Step 10:
[0291] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[0292] Step 11:
[0293] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[0294] Step 12:
[0295] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[0296] ---
[0297] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[0298] Example 2
[0299] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0300] Conventional translation systems are capable of converting speech to text and translating it into another language, but they do not take the user's emotional state into account when outputting the speech. As a result, the translated speech may sound unnatural or may not properly convey the user's emotions. Furthermore, they lack the functionality to provide real-time warnings for sensitive content or to support the progress of meetings. This poses a challenge, making smooth communication difficult in international and business settings.
[0301] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing an emotional state from text data, a means for translating the converted text and emotional state data into a specified language in real time, and a means for converting the translated text and the recognized emotional state into voice data that retains the characteristics of the original voice. This enables natural voice translation that reflects the user's emotional state. Furthermore, the system also includes a function for detecting sensitive content in real time and issuing a warning, as well as a support function for facilitating the progress of meetings, thereby realizing safer and more efficient communication.
[0302] "User" refers to an individual or organization that uses the system.
[0303] "Terminal" refers to electronic devices used by users, such as computers, smartphones, and tablets.
[0304] A "server" is a computer system that processes and stores data, and refers to a device that communicates with terminals over a network.
[0305] "Audio digital signal" refers to audio captured by an audio input device such as a microphone and converted into digital data.
[0306] "Text string" refers to data consisting of characters such as alphabets and kanji characters converted from speech using speech recognition technology.
[0307] "Emotional state" refers to the speaker's emotion that can be read from speech or text, such as positive, negative, or neutral.
[0308] "Real-time translation" refers to the process by which input data (voice or text) is translated into another language within a short time of being received.
[0309] "Voice features" refer to the characteristics of a user's voice and speaking style, and refer to elements that reproduce the way the user is speaking.
[0310] "Audio Data" means data for storing, processing, and playing audio signals in digital form.
[0311] A "sensitive alert" is a message that warns users when content is deemed to contain inappropriate or sensitive content.
[0312] "Suggestions and instructions to support the progress of meetings" refers to organizational suggestions and instructions that are generated to ensure that meetings and conversations proceed smoothly.
[0313] The present invention relates to a system that translates a user's voice in real time and enables natural international communication by combining it with an emotion engine. First, the overall configuration of this system will be described.
[0314] Overall system configuration
[0315] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[0316] Hardware and software used
[0317] Specific hardware includes the smartphone, tablet, or computer used by the user, while the following APIs and services are used as software:
[0318] Speech Recognition API: Google Cloud Speech-to-Text API
[0319] Emotion Recognition Engine: Azure Cognitive Services
[0320] Translation API: Google Translate API
[0321] Speech synthesis technology: Amazon Polly
[0322] Operation flow
[0323] The specific operation flow of the system will be explained below.
[0324] User voice input
[0325] The user speaks into the device's microphone. For example, the user says "Good morning" in Japanese. The device's microphone captures this speech and converts it into a digital signal.
[0326] Sending voice data to the server
[0327] The device transmits the captured audio data to a server via the Internet.
[0328] Speech recognition and text conversion
[0329] The server uses the Google Cloud Speech-to-Text API to convert the audio data into a string of text. For example, "Good morning" is converted into "Good morning."
[0330] Recognition of emotional states
[0331] The server uses Azure Cognitive Services to recognize emotional states based on text data. For example, if "Good morning" is spoken in a cheerful tone, the emotion engine will recognize it as "positive."
[0332] Real-time text translation
[0333] The server translates this text and emotion information into the specified language (e.g., English) using the Google Translate API, for example, "Ohayo gozaimasu" is translated to "Good morning."
[0334] Text-to-speech translation
[0335] The server uses Amazon Polly to convert the translated text into speech data that preserves the characteristics and emotion of the original voice, for example, "Good morning" in the user's voice.
[0336] Sending and outputting synthesized speech
[0337] The server then sends the generated voice data to the device, which then decodes it and plays it back through the speaker, allowing the user to hear the translated voice with their own voice and emotions preserved.
[0338] Specific examples
[0339] For example, when a Japanese-speaking user (A) converses with an English-speaking user (B), the following specific operations take place: When user (A) says "Good morning," the device converts this speech into a digital signal and sends it to the server. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. Next, it translates this into English "Good morning," and generates speech data by adding user (A)'s speech features and emotional characteristics based on the translated text. The speech data sent from the server to the device is played back on the device, and user (B) can hear "Good morning" in user (A)'s voice in a cheerful tone.
[0340] Prompt Sentence Examples
[0341] Here are some example prompts to input to the generative AI model:
[0342] When user (A) says "Good morning" in a cheerful voice, the system converts this speech into text and recognizes positive emotions. The server then translates this text into "Good morning" and generates a cheerful voice data in user (A)'s voice. Finally, user (B) can hear "Good morning" in user (A)'s voice.
[0343] The above is a specific embodiment for carrying out the present invention. This system enables natural speech translation that reflects the user's emotions, making international communication smoother.
[0344] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0345] Step 1:
[0346] User voice input
[0347] The user speaks into the microphone of the terminal. The user's voice signal is input to the microphone. In concrete terms, the microphone converts this voice signal into digital voice data. For example, when the user says "Good morning," this voice data is input to the terminal.
[0348] Step 2:
[0349] Sending voice data to the server
[0350] The device sends the captured audio data to the server. As input, the digital audio data is used. As data processing, the device converts the audio data into packets that are sent to the server via the network. As output, packets containing the audio data are sent to the server. The specific operation is that the device transfers the audio data to the server via an Internet connection.
[0351] Step 3:
[0352] Speech recognition and text conversion
[0353] The server converts the received voice data into text. The digital voice data sent to the server is used as input. The Google Cloud Speech-to-Text API is used for data calculations. The server sends the voice data to the API and receives text data as output. Specifically, the server sends a request to the speech recognition API and receives the text data "Good morning" as a response.
[0354] Step 4:
[0355] Recognition of emotional states
[0356] The server recognizes the emotional state from the converted text data. The text data is used as input. Azure Cognitive Services is used for data calculations. The server sends the text to the emotion engine and receives data indicating the emotional state as output. Specifically, "Good morning" is analyzed as having a cheerful tone, and the emotional data "positive" is obtained.
[0357] Step 5:
[0358] Real-time text translation
[0359] The server translates the text and emotion data into the specified language. The text data and emotion data are used as input. The Google Translate API is used for data calculation. The server sends the text data and emotion data to the API and receives the translated text data as output. Specifically, "Ohayou gozaimasu" is translated into "Good morning," and the corresponding positive emotion data is also obtained.
[0360] Step 6:
[0361] Text-to-speech translation
[0362] The server converts the translated text data into speech data. The inputs are the translated text, emotion data, and the original speech feature data. Amazon Polly is used for data processing. The server sends this data to a speech synthesis engine, and the output is speech data with the user's speech features. Specifically, it generates speech data with a positive tone based on the text "Good morning."
[0363] Step 7:
[0364] Sending and outputting synthesized speech
[0365] The server sends the generated voice data to the terminal. The voice data generated by the server is used as input. The terminal receives the data and prepares to play it. The terminal decodes the received voice data and plays it from the speaker. The specific operation is that the terminal plays the voice data "Good morning" so that the user can hear the voice.
[0366] (Application example 2)
[0367] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0368] With modern globalization, there is an increasing need to communicate naturally with users who speak different languages. However, language barriers and misunderstandings of emotions are common, which can lead to poor service quality, especially in brick-and-mortar stores and service industries. Furthermore, many existing translation systems ignore emotional states, resulting in inaccurate communication of user intent. Furthermore, they lack the ability to automatically detect and warn against sensitive topics and language.
[0369] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting a user's voice into a digital signal, means for transmitting the digital signal to the server, means for converting the voice digital signal into character string text on the server, means for translating the converted text into a specified language in real time, means for converting the translated text into voice data that preserves the original voice characteristics and emotional state, means for transmitting the voice data to a terminal, means for playing the voice data on the terminal, and means for supporting multilingual communication between users and staff in a physical store. This enables users and staff who speak different languages to communicate naturally while preserving voice characteristics and emotions, improving the quality of service. Furthermore, the safety of conversations is ensured by detecting sensitive content and issuing alerts at appropriate times.
[0370] The "means for converting a user's voice into a digital signal" refers to a device or system that converts the voice emitted by the user from analog to digital.
[0371] The "means for transmitting a digital signal to a server" refers to a device or system for transmitting the converted digital signal data to a server via a network.
[0372] The "means for converting the digital audio signal into a character string text on the server" refers to software or hardware on the server side that executes the process of converting the transmitted digital audio signal into a character string.
[0373] The "means for translating converted text into a specified language in real time" is software or a system that translates a string of text into another specified language in real time.
[0374] "Means for converting translated text into speech data that retains the original speech characteristics and emotional state" refers to a device or system that converts translated text into speech data that reflects the user's speech characteristics and emotions.
[0375] The "means for transmitting voice data to a terminal" refers to a device or system that transmits the generated voice data to a user's terminal via a network.
[0376] "Means for reproducing audio data on a terminal" refers to a device or system for reproducing received audio data on a speaker or the like of the terminal.
[0377] "Means for supporting multilingual communication between users and staff in a physical store" refers to devices or systems that support natural communication between users and staff who speak different languages in a physical store environment.
[0378] "Sensitive content detection means" means software or systems that detect sensitive content in speech or text in real time.
[0379] "Means for generating a sensitive alert and notifying the terminal" refers to software or a system that generates an alert and notifies the user's terminal when sensitive content is detected.
[0380] The "means for generating suggestions and instructions to support the progress of a meeting" refers to software or a system that generates appropriate suggestions and instructions based on the progress of the meeting.
[0381] The "means for transmitting suggestions and instructions to a terminal and notifying the user" refers to software or a system that transmits the generated suggestions and instructions to the user's terminal and notifies the user.
[0382] The present invention is a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[0383] Overall system configuration
[0384] This system consists of a user, a terminal, a server, and an emotion engine. When the user speaks into the terminal, the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates the text into the specified language. The translated text is converted into audio data that also retains the emotional state and is sent to the terminal. The user can listen to the translated speech through the terminal's speaker.
[0385] Specific Embodiments
[0386] 1. Audio input and digital signal conversion
[0387] Voice input is captured when a user speaks into the device's microphone. For example, a user might say, "Welcome, what product are you looking for?" in a physical store. This speech is converted into digital signals using voice recognition software within the device.
[0388] 2. Digital signal transmission to server
[0389] The converted digital signal is then sent to a server via a network, using a smartphone or tablet as the main hardware and a data transfer protocol as the software.
[0390] 3. Text Conversion
[0391] The digital signal that reaches the server is converted into text using a speech recognition API (e.g., Google Cloud Speech-to-Text API) on the server, for example, "Welcome, what product are you looking for?"
[0392] 4. Recognizing emotional states
[0393] The textual content is then subjected to emotional recognition using an emotion engine (e.g., Emotion Recognition API). For example, the text is classified into an emotional state such as "welcome."
[0394] 5. Real-time translation
[0395] The server translates the recognized text and emotional state into a specified language (e.g., English) in real time using a translation API (e.g., Google Translate API). In this case, the translation is "Welcome, what kind of product are you looking for?"
[0396] 6. Speech synthesis of translated text
[0397] Based on the translated text and the recognized emotional state, the server generates voice data using speech synthesis technology (e.g., Text-to-Speech API). The voice is generated with a tone based on the emotional state (welcome).
[0398] 7. Sending and outputting synthesized speech
[0399] The server sends the generated voice data to the device, which then plays the received voice data through the speaker, allowing the user to listen to the translated voice in real time through the device.
[0400] Specific use cases
[0401] For example, consider a conversation between a Japanese-speaking staff member and an English-speaking customer in a physical store. When the staff member says in Japanese, "Welcome, what kind of product are you looking for?", the system translates this into English and tells the customer, "Welcome, what kind of product are you looking for?" in a voice that retains the emotional state of the customer. This allows for smooth communication that transcends language barriers.
[0402] Prompt Sentence Examples
[0403] "We are developing a multilingual customer service assistant app that allows users to communicate naturally across language barriers and maintains their emotional state. This app translates the user's voice in real time and reflects their emotional state using an emotion engine. The process involves the following steps: 1. Record the user's voice and convert it into a digital signal. 2. Convert the voice to text (e.g., 'Welcome, what kind of product are you looking for?' -> 'Welcome, what kind of product are you looking for?'). 3. Recognize the emotional state of the text (e.g., 'Welcome'). 4. Translate while maintaining the emotional state and perform speech synthesis. 5. Output the final voice data and provide it to the user. The generated program is as follows:"
[0404] In this way, the present invention enables users and staff who speak different languages to communicate naturally while maintaining their emotional state.
[0405] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0406] Step 1:
[0407] The user speaks into the device's microphone. The input is the user's voice, which is captured by the voice input software and converted into a digital signal. The data is then processed by sampling the voice signal and converting it into digital data. The output is a digital voice signal.
[0408] Step 2:
[0409] The terminal sends the converted digital voice signal to the server. The input is the digital voice signal generated in step 1, which is sent to the server via the Internet using the terminal's communication module. The data is processed in the process of packetizing the digital signal and sending it. The output is the digital voice signal that has arrived at the server.
[0410] Step 3:
[0411] The server receives the digital voice signal and converts it into text using a speech recognition API. The input is the digital voice signal, which is converted into text by a speech recognition engine (e.g., Google Cloud Speech-to-Text API). The data operations are speech signal processing and natural language processing. The output is text data generated from the voice signal.
[0412] Step 4:
[0413] The server sends text data to an emotion recognition engine to recognize the emotional state. The input is text data, which is analyzed for emotion by an emotion recognition engine (e.g., Emotion Recognition API). The data is processed by string analysis and emotion categorization. The output is the recognized emotional state.
[0414] Step 5:
[0415] The server sends the text data and the recognized emotional state to a translation API, which translates it into the specified language in real time. The input is the text data and the emotional state, and the translation is performed using a translation API (e.g., Google Translate API). The data processing is the language conversion of the text. The output is the translated text data.
[0416] Step 6:
[0417] The server uses the translated text data and emotional state to synthesize speech. The input is the translated text data and emotional state, which is converted into speech data by a speech synthesis engine (e.g., Text-to-Speech API). The data is processed to convert text to speech and set the emotional tone. The output is the synthesized speech data.
[0418] Step 7:
[0419] The server sends the synthesized voice data to the terminal. The input is the synthesized voice data, which is sent to the terminal via the Internet using a communication module. The data is processed by sending data packets. The output is the voice data that has arrived at the terminal.
[0420] Step 8:
[0421] The device decodes the audio data received and plays it from the speaker. The input is audio data, which is decoded using the device's audio decoder and output as audio from the speaker. The data processing involves converting the digital audio signal to analog. The output is synthesized speech that the user can hear.
[0422] This allows users to communicate naturally with other users who speak different languages.
[0423] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0424] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0425] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0426] [Second embodiment]
[0427] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0428] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0429] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0430] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0431] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0432] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0433] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0434] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0435] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0436] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0437] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0438] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0439] Understood. Below is the "Mode for carrying out the invention".
[0440] ---
[0441] The present invention relates to a translation multi-assist AI system that overcomes language barriers in international communication and enables natural conversation. Specific embodiments of the present invention will be described below.
[0442] Overall system configuration
[0443] This system consists of a user, a terminal, and a server. The user inputs speech, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into speech data that retains the user's speech characteristics and sent to the terminal. The terminal plays back the received speech data and provides the translation result to the user.
[0444] Program processing explanation
[0445] User voice input
[0446] The user speaks into the device's microphone. For example, if the user says "Good morning" in Japanese, this speech is captured by the device.
[0447] Speech recognition and text conversion
[0448] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this example, the audio "Good morning" is converted into the text "Good morning."
[0449] Real-time translation
[0450] The server passes the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0451] Text-to-speech translation
[0452] The server analyzes the user's voice characteristics and uses speech synthesis technology to convert the translated text into audio data in the user's voice. The server synthesizes "Good morning" in English, preserving the pitch, tone, and speed of the user's voice.
[0453] Sending and outputting synthesized speech
[0454] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0455] Specific examples
[0456] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0457] Sensitive Alert Function
[0458] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[0459] Facilitator function
[0460] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0461] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[0462] ---
[0463] The processing flow will be explained below.
[0464] Understood. Below I will explain the process in concrete steps.
[0465] ---
[0466] Step 1:
[0467] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0468] Step 2:
[0469] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[0470] Step 3:
[0471] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text as "Good morning."
[0472] Step 4:
[0473] The server sends the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0474] Step 5:
[0475] The server converts the translated text into audio data that preserves the user's voice characteristics. Specifically, it analyzes the pitch, tone, and speed of the user's original speech and uses speech synthesis technology to generate "Good morning" in the user's voice.
[0476] Step 6:
[0477] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[0478] Step 7:
[0479] The device decodes the received audio data and plays it through the speaker, allowing the user to hear the translated "Good morning" in their own voice.
[0480] Step 8:
[0481] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[0482] Step 9:
[0483] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[0484] Step 10:
[0485] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[0486] Step 11:
[0487] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[0488] ---
[0489] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[0490] Example 1
[0491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0492] In international communication, it is difficult to achieve natural and smooth conversations between people who speak different languages. Furthermore, when conversation content contains sensitive information, it is important to have a system that can appropriately process and notify the content. Furthermore, a system that supports the progress of meetings and important communications is also needed.
[0493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0494] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server via a communication network; means for converting the voice digital signal into character string text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the characteristics of the original voice; means for transmitting the voice data to a terminal via a communication network; means for playing the voice data on the terminal; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of the conference; and means for transmitting the suggestions or instructions to the terminal and notifying the user. This enables natural and smooth conversations between people speaking different languages, enables appropriate processing and notification of sensitive content, and supports the progress of conferences and important communications.
[0495] Understood. Below are definitions of important terms included in the claims mentioned above.
[0496] ---
[0497] A "digital signal" refers to a signal that represents analog audio data in numerical form.
[0498] "Communication network" refers to communication infrastructure such as the Internet or local networks used to send and receive data.
[0499] A "server" refers to a computer system that receives and processes requests from users over a network.
[0500] "Speech recognition API" refers to a programming interface for receiving voice data as input and converting it into corresponding text data.
[0501] "Text string" refers to text data converted by speech recognition.
[0502] "Translation API" refers to a program interface for converting input text data into another specified language.
[0503] "Speech features" refer to individual characteristics of speech such as pitch, tone, and rate.
[0504] "Audio Data" means an audio file represented in digital form.
[0505] "Sensitive alert" refers to a warning that notifies the user when sensitive content is detected.
[0506] "Suggestions and Instructions" refers to recommendations and instructions generated to help facilitate conversation or communication.
[0507] "Terminal" refers to a device (e.g., a smartphone or PC) that a user directly operates and that communicates with a server.
[0508] ---
[0509] These definitions are intended to clarify the meaning of key terms in the claims.
[0510] The present invention is a system that translates a user's voice input in real time and outputs the translated text as voice data that retains the original voice characteristics. This system is composed of a user, a terminal, and a server.
[0511] The user speaks into the microphone of the device. For example, when the user says "Good morning," the voice is captured by the device. The device converts the voice into a digital signal through the microphone and transmits the digital signal to the server via a communication network.
[0512] The server converts the received digital signal into a string of text using a speech recognition API (e.g., Google Cloud Speech-to-Text API). In this example, the speech "Good morning" is converted into the text "Good morning."
[0513] The server then passes the converted text to a translation API (e.g., Google Cloud Translation API) for real-time translation into the specified language (e.g., English). Here, the text "Ohayo gozaimasu" is translated into English as "Good morning."
[0514] The server uses speech synthesis technology (e.g., Amazon Polly) to convert the translated text into speech that retains the characteristics of the original voice, so that the translated "Good morning" is converted into speech that reflects the pitch, tone, and rate of the user's voice.
[0515] The server then sends the generated voice data to the device via a communication network. The device then decodes the received voice data and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0516] Furthermore, this system has the function of detecting sensitive content. The server analyzes the conversation content, and if it detects sensitive content, it generates a sensitive alert and notifies the terminal. The terminal then presents this alert to the user in text or voice format, allowing the user to respond appropriately.
[0517] The server also has the ability to generate suggestions and instructions to help meetings and important communications proceed smoothly. For example, if a conversation stalls, the server will automatically generate a suggestion such as, "Shall we move on to the next topic?" This allows the device to notify the user of the suggestion or instruction, helping to ensure the smooth progress of the meeting.
[0518] As a concrete example, consider a situation in which a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0519] An example of a prompt sentence is, "A Japanese-speaking user says 'Good morning' into the terminal. Explain how this speech is converted into a digital signal, processed by the server, and finally played back as English speech."
[0520] As described above, the present invention enables natural and smooth conversations between people who speak different languages, and enables sensitive content to be appropriately handled and communicated, thereby supporting the progress of meetings and important communications.
[0521] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0522] Understood. Below, we will divide the processing flow of the system program into specific processing steps and explain the specific operations including input and output.
[0523] ---
[0524] Step 1:
[0525] Capture audio input
[0526] The user speaks into the device's microphone.
[0527] Input: User speech (e.g. "Good morning").
[0528] Device behavior: The device uses a microphone to capture the user's voice.
[0529] Output: Analog audio signal.
[0530] Step 2:
[0531] Digital conversion of analog audio signals
[0532] The terminal converts the captured analog voice signal into a digital signal.
[0533] Input: Analog audio signal.
[0534] Terminal operation: The A / D converter inside the terminal converts the audio signal into a digital signal and records it as time-series data.
[0535] Output: Digital audio signal.
[0536] Step 3:
[0537] Transmitting digital audio signals
[0538] The terminal transmits the digital audio signal to the server over a communication network.
[0539] Input: Digital audio signal.
[0540] Terminal operation: The voice signal is converted into packet data and sent to a server via the Internet through a communications module.
[0541] Output: Audio data packets to the server.
[0542] Step 4:
[0543] Speech recognition and text conversion
[0544] The server inputs the received digital voice signal into a voice recognition API and converts it into text.
[0545] Input: Digital audio signal.
[0546] Server operation: A speech recognition API (e.g., a speech recognition cloud service) converts the speech data into text data.
[0547] Output: String text (e.g. "Good morning").
[0548] Step 5:
[0549] Text translation
[0550] The server passes the converted text to the translation API, which translates it into the specified language.
[0551] Input: String text (e.g. "Good morning").
[0552] Server action: Translate the text into the specified language (e.g., English) using a translation API (e.g., a translation cloud service).
[0553] Output: The translated text (e.g. "Good morning").
[0554] Step 6:
[0555] Speech synthesis of translated text
[0556] The server converts the translated text into voice data using voice synthesis technology.
[0557] Input: The translated text (e.g. "Good morning") and the user's voice characteristics.
[0558] Server operation: A speech synthesis API (e.g., a speech synthesis cloud service) converts text into speech data that retains the user's speech characteristics.
[0559] Output: Synthetic speech data.
[0560] Step 7:
[0561] Sending audio data
[0562] The server transmits the generated voice data to the terminal via a communication network.
[0563] Input: Synthetic speech data.
[0564] Server operation: Converts voice data into packet data and sends it to the terminal via the Internet through the communication module.
[0565] Output: Audio data packets destined for the device.
[0566] Step 8:
[0567] Playing audio data
[0568] The terminal plays the received audio data through the speaker.
[0569] Input: Synthetic speech data.
[0570] Device operation: The decoded audio data is converted into an analog signal using a digital-to-analog converter and played through the speaker.
[0571] Output: The speech the user hears (e.g. "Good morning").
[0572] ---
[0573] This concludes the explanation of the specific processing steps of the program. Each step clearly shows the type of data processing and calculation that will be performed based on the input data. This system makes it possible to have natural conversations that transcend language barriers.
[0574] (Application example 1)
[0575] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0576] In international communication, there is a demand for systems that support natural conversations between users who speak different languages. However, existing systems have difficulty preserving the characteristics of speech while translating, and lack technology to support communication between various languages in autonomous vehicles. Furthermore, they lack the ability to detect and notify sensitive content in real time, and the ability to support the progress of meetings.
[0577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0578] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server; means for converting the voice digital signal into text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the original voice characteristics; means for transmitting the voice data to a terminal; means for playing the voice data on the terminal; means for supporting communication between users in various languages within an autonomous vehicle; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of a meeting; means for transmitting the suggestions or instructions to the terminal and notifying the user; and a function for translating voice in real time using a generative AI model, installed on a smartphone, smart glasses, or head-mounted display. This supports natural communication between users who speak different languages and enables smooth communication across language barriers within an autonomous vehicle. Furthermore, the functions for detecting and notifying sensitive content and supporting the progress of a meeting enable safe and efficient communication.
[0579] The "means for converting a user's voice into a digital signal" is a device or program for converting a user's voice from an analog signal into a digital signal.
[0580] The "means for transmitting the digital signal to the server" refers to a communication device or program for transmitting the converted digital signal to the server via a network.
[0581] The "means for converting the digital audio signal into a character string text on the server" refers to software or a program for converting the digital audio signal into a text format on the server using voice recognition technology.
[0582] "Means for translating converted text into a specified language in real time" means a technology or program for instantly translating text into another specified language.
[0583] "Means for converting translated text into speech data that preserves the characteristics of the original speech" refers to technology or a program for converting translated text into speech data in a form that preserves the user's speech characteristics.
[0584] The "means for transmitting the voice data to the terminal" is a communication device or a program for transmitting the generated voice data to the user's terminal via a network.
[0585] The "means for reproducing the audio data on the terminal" refers to hardware or software for reproducing the audio data received on the terminal.
[0586] "Means for supporting communication between users in various languages within an autonomous vehicle" refers to technology or a program that enables smooth communication between passengers who speak different languages within an autonomous vehicle.
[0587] "Sensitive content detection means" means an algorithm or program that identifies sensitive or inappropriate content in real time.
[0588] The "means for generating a sensitive alert and notifying the terminal" is a program for generating a warning for detected sensitive content and notifying the user's terminal of the warning.
[0589] The "means for generating suggestions and instructions to support the progress of a meeting" refers to an algorithm or program for generating appropriate suggestions and instructions so that the meeting can proceed smoothly.
[0590] The "means for transmitting the proposal or instruction to the terminal and notifying the user" refers to a communication device or program for transmitting the generated proposal or instruction to the user's terminal and notifying the user of it.
[0591] "A feature installed on a smartphone, smart glasses, or head-mounted display that uses a generative AI model to translate speech in real time" refers to software or an application that utilizes the AI model built into these devices to provide real-time speech translation functionality.
[0592] The present invention relates to a translation multi-assist AI system for supporting natural conversation between users who speak different languages. Hereinafter, an embodiment of the present invention will be specifically described.
[0593] Overall system configuration
[0594] This system consists of a user, a terminal, the environment inside the autonomous vehicle, and a server. The user inputs voice through the terminal's microphone, and the terminal converts the voice into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language in real time. The translated text is then converted back into voice data that retains the user's voice characteristics and sent to the terminal. Finally, the terminal plays back the received voice data and provides the translation result to the user.
[0595] Program processing explanation
[0596] Audio input and digital signal conversion
[0597] The user speaks into the microphone of the device. For example, when the user says, "Take me to the airport, please," this voice is captured by the microphone of the device. The device converts the voice into a digital signal, which is then transmitted to the server via the network.
[0598] Speech recognition and text conversion
[0599] The server uses a speech recognition API (Google's speech recognition API is a specific example) to convert digital voice signals into text. This technology converts the speech "Please take me to the airport" into the text "Please take me to the airport."
[0600] Real-time translation
[0601] The server passes the converted text to a translation API (e.g., Google Translate API) and translates it into the specified language (e.g., English) in real time. In this example, "Kuukou de onegaishimasu" is translated to "Please go to the airport."
[0602] Text-to-speech translation
[0603] The server analyzes the user's voice characteristics and converts the translated text into audio data in the user's voice using speech synthesis technology (e.g., Google Text-to-Speech). The server generates audio data saying "Please go to the airport" in English, while retaining the user's voice characteristics.
[0604] Sending and playing audio data
[0605] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Please go to the airport" translated in their own voice.
[0606] Sensitive Alert Function
[0607] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitive alert in real time and notifies the device. This alert is presented in text or audio format, allowing the user to respond appropriately.
[0608] Facilitator function
[0609] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0610] Specific examples
[0611] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Please go to the airport," this speech is sent to the server via the device. The server converts the speech into text "Please go to the airport" and translates it into English "Please go to the airport." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Please go to the airport" in user (A)'s voice.
[0612] Prompt Sentence Examples
[0613] "Create an AI model that automatically translates what is said in the car into English. Convert the incoming Japanese speech into text, translate it into English using an API, and then convert it back into speech and play it back."
[0614] Thus, the present invention is a system that provides a natural translation function between users who speak different languages, detects sensitive content, and supports the progress of meetings.
[0615] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0616] Step 1:
[0617] The user speaks into a terminal inside the autonomous vehicle. The terminal's microphone captures the user's voice and converts the analog voice signal into a digital signal. The input is the user's voice, and the output is a digital voice signal. Specifically, the terminal's voice input device captures the voice, and the built-in digital signal processor digitizes it.
[0618] Step 2:
[0619] The terminal transmits the converted digital signal to the server via the network. The input is a digital audio signal, and the output is the transfer of the digital signal to the server. Specifically, the terminal's communication module (e.g., Wi-Fi or mobile data communication) transmits the digital signal to the server.
[0620] Step 3:
[0621] The server converts the received digital voice signal into a string of text using a speech recognition API. The input is a digital voice signal, and the output is text data such as "Take me to the airport, please." Specifically, the server calls speech recognition software (e.g., Google Speech-to-Text API) and converts the voice signal into a string of text.
[0622] Step 4:
[0623] The server translates text data into the specified language in real time using a translation API. The input is text data, and the output is translated text such as "Please go to the airport." Specifically, the server uses translation software (e.g., Google Translate API) to convert the text into the specified language.
[0624] Step 5:
[0625] The server converts the translated text into audio data using speech synthesis technology. The input is the translated text, and the output is audio data. Specifically, the server uses speech synthesis software (e.g., Google Text-to-Speech API) to convert the text into audio data that retains the user's voice characteristics.
[0626] Step 6:
[0627] The server transmits the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal. In concrete terms, the communication module of the server transmits the voice data to the terminal.
[0628] Step 7:
[0629] The device decodes the received audio data and plays it through the speaker. The input is audio data and the output is audio output from the speaker. Specifically, the device's audio playback device decodes the audio data and plays it as audio through the speaker.
[0630] Step 8:
[0631] The server analyzes the content of the conversation and detects sensitive content. The input is voice or text data, and the output is a sensitive alert. Specifically, the server uses natural language processing technology to analyze the content of the conversation and detect whether it contains sensitive information. If it detects sensitive information, it generates an alert.
[0632] Step 9:
[0633] When a sensitive alert is generated, the server notifies the terminal of the alert. The input is the sensitive alert, and the output is a notification to the terminal. Specifically, the server's notification system sends the alert to the terminal and causes the terminal to display the notification.
[0634] Step 10:
[0635] To help facilitate the progress of the meeting, the server analyzes the conversation and generates appropriate suggestions and instructions. The input is conversation data, and the output is suggestions and instructions. Specifically, the server uses a generative AI model to understand the context of the conversation and generate appropriate suggestions and instructions.
[0636] Step 11:
[0637] The server sends the generated suggestions and instructions to the terminal and notifies the user. The input is the suggestions and instructions, and the output is a presentation to the terminal. Specifically, the instruction data is sent from the server to the terminal, and a notification is displayed on the terminal to the user.
[0638] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0639] Understood. Below is the "Mode for carrying out the invention" regarding the invention that combines an emotion engine.
[0640] ---
[0641] The present invention relates to a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[0642] Overall system configuration
[0643] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[0644] Program processing explanation
[0645] User voice input
[0646] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0647] Speech recognition and text conversion
[0648] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this case, "Good morning" is converted into text.
[0649] Recognition of emotional states
[0650] The server uses an emotion engine to recognize the user's emotional state from the input speech. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes a "positive" emotional state.
[0651] Real-time translation
[0652] The server sends the converted text and the recognized emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0653] Text-to-speech translation
[0654] The server analyzes the user's voice characteristics and the recognized emotional state, and uses speech synthesis technology to convert the translated text into audio data in the user's voice and emotional state, where "Good morning" is generated in the user's voice and in the recognized "positive" tone.
[0655] Sending and outputting synthesized speech
[0656] The server sends the generated voice data to the device, which then decodes the received voice data and plays it back through the speaker. This allows the user to hear "Good morning" translated in their own voice and with their own emotions.
[0657] Specific examples
[0658] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. This is then translated into English as "Good morning," and the translated text is synthesized using speech synthesis based on the voice and emotional characteristics of user (A). The speech data sent from the server to the device is played back, and user (B) hears "Good morning" in user (A)'s voice with a "positive" tone.
[0659] Sensitive Alert Function
[0660] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[0661] Facilitator function
[0662] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0663] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice and emotions, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[0664] ---
[0665] The processing flow will be explained below.
[0666] Understood. Below, I will explain the process flow for the invention that combines the emotion engine, broken down into specific steps.
[0667] ---
[0668] Step 1:
[0669] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0670] Step 2:
[0671] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[0672] Step 3:
[0673] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text "Good morning."
[0674] Step 4:
[0675] The server passes the recognized text to the emotion engine to recognize the user's emotional state. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes the user's emotional state as "positive."
[0676] Step 5:
[0677] The server sends the text with the emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0678] Step 6:
[0679] The server generates voice data that reflects the user's voice characteristics and emotional state based on the translated text and the recognized emotional state. Specifically, it generates "Good morning" in the user's voice while maintaining a "positive" tone.
[0680] Step 7:
[0681] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[0682] Step 8:
[0683] The device decodes the received audio data and plays it through the speaker, allowing the user to hear "Good morning" translated in their own voice and emotional state.
[0684] Step 9:
[0685] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[0686] Step 10:
[0687] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[0688] Step 11:
[0689] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[0690] Step 12:
[0691] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[0692] ---
[0693] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[0694] Example 2
[0695] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0696] Conventional translation systems are capable of converting speech to text and translating it into another language, but they do not take the user's emotional state into account when outputting the speech. As a result, the translated speech may sound unnatural or may not properly convey the user's emotions. Furthermore, they lack the functionality to provide real-time warnings for sensitive content or to support the progress of meetings. This poses a challenge, making smooth communication difficult in international and business settings.
[0697] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing an emotional state from text data, a means for translating the converted text and emotional state data into a specified language in real time, and a means for converting the translated text and the recognized emotional state into voice data that retains the characteristics of the original voice. This enables natural voice translation that reflects the user's emotional state. Furthermore, the system also includes a function for detecting sensitive content in real time and issuing a warning, as well as a support function for facilitating the progress of meetings, thereby realizing safer and more efficient communication.
[0698] "User" refers to an individual or organization that uses the system.
[0699] "Terminal" refers to electronic devices used by users, such as computers, smartphones, and tablets.
[0700] A "server" is a computer system that processes and stores data, and refers to a device that communicates with terminals over a network.
[0701] "Audio digital signal" refers to audio captured by an audio input device such as a microphone and converted into digital data.
[0702] "Text string" refers to data consisting of characters such as alphabets and kanji characters converted from speech using speech recognition technology.
[0703] "Emotional state" refers to the speaker's emotion that can be read from speech or text, such as positive, negative, or neutral.
[0704] "Real-time translation" refers to the process by which input data (voice or text) is translated into another language within a short time of being received.
[0705] "Voice features" refer to the characteristics of a user's voice and speaking style, and refer to elements that reproduce the way the user is speaking.
[0706] "Audio Data" means data for storing, processing, and playing audio signals in digital form.
[0707] A "sensitive alert" is a message that warns users when content is deemed to contain inappropriate or sensitive content.
[0708] "Suggestions and instructions to support the progress of meetings" refers to organizational suggestions and instructions that are generated to ensure that meetings and conversations proceed smoothly.
[0709] The present invention relates to a system that translates a user's voice in real time and enables natural international communication by combining it with an emotion engine. First, the overall configuration of this system will be described.
[0710] Overall system configuration
[0711] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[0712] Hardware and software used
[0713] Specific hardware includes the smartphone, tablet, or computer used by the user, while the following APIs and services are used as software:
[0714] Speech Recognition API: Google Cloud Speech-to-Text API
[0715] Emotion Recognition Engine: Azure Cognitive Services
[0716] Translation API: Google Translate API
[0717] Speech synthesis technology: Amazon Polly
[0718] Operation flow
[0719] The specific operation flow of the system will be explained below.
[0720] User voice input
[0721] The user speaks into the device's microphone. For example, the user says "Good morning" in Japanese. The device's microphone captures this speech and converts it into a digital signal.
[0722] Sending voice data to the server
[0723] The device transmits the captured audio data to a server via the Internet.
[0724] Speech recognition and text conversion
[0725] The server uses the Google Cloud Speech-to-Text API to convert the audio data into a string of text. For example, "Good morning" is converted into "Good morning."
[0726] Recognition of emotional states
[0727] The server uses Azure Cognitive Services to recognize emotional states based on text data. For example, if "Good morning" is spoken in a cheerful tone, the emotion engine will recognize it as "positive."
[0728] Real-time text translation
[0729] The server translates this text and emotion information into the specified language (e.g., English) using the Google Translate API, for example, "Ohayo gozaimasu" is translated to "Good morning."
[0730] Text-to-speech translation
[0731] The server uses Amazon Polly to convert the translated text into speech data that preserves the characteristics and emotion of the original voice, for example, "Good morning" in the user's voice.
[0732] Sending and outputting synthesized speech
[0733] The server then sends the generated voice data to the device, which then decodes it and plays it back through the speaker, allowing the user to hear the translated voice with their own voice and emotions preserved.
[0734] Specific examples
[0735] For example, when a Japanese-speaking user (A) converses with an English-speaking user (B), the following specific operations take place: When user (A) says "Good morning," the device converts this speech into a digital signal and sends it to the server. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. Next, it translates this into English "Good morning," and generates speech data by adding user (A)'s speech features and emotional characteristics based on the translated text. The speech data sent from the server to the device is played back on the device, and user (B) can hear "Good morning" in user (A)'s voice in a cheerful tone.
[0736] Prompt Sentence Examples
[0737] Here are some example prompts to input to the generative AI model:
[0738] When user (A) says "Good morning" in a cheerful voice, the system converts this speech into text and recognizes positive emotions. The server then translates this text into "Good morning" and generates a cheerful voice data in user (A)'s voice. Finally, user (B) can hear "Good morning" in user (A)'s voice.
[0739] The above is a specific embodiment for carrying out the present invention. This system enables natural speech translation that reflects the user's emotions, making international communication smoother.
[0740] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0741] Step 1:
[0742] User voice input
[0743] The user speaks into the microphone of the terminal. The user's voice signal is input to the microphone. In concrete terms, the microphone converts this voice signal into digital voice data. For example, when the user says "Good morning," this voice data is input to the terminal.
[0744] Step 2:
[0745] Sending voice data to the server
[0746] The device sends the captured audio data to the server. As input, the digital audio data is used. As data processing, the device converts the audio data into packets that are sent to the server via the network. As output, packets containing the audio data are sent to the server. The specific operation is that the device transfers the audio data to the server via an Internet connection.
[0747] Step 3:
[0748] Speech recognition and text conversion
[0749] The server converts the received voice data into text. The digital voice data sent to the server is used as input. The Google Cloud Speech-to-Text API is used for data calculations. The server sends the voice data to the API and receives text data as output. Specifically, the server sends a request to the speech recognition API and receives the text data "Good morning" as a response.
[0750] Step 4:
[0751] Recognition of emotional states
[0752] The server recognizes the emotional state from the converted text data. The text data is used as input. Azure Cognitive Services is used for data calculations. The server sends the text to the emotion engine and receives data indicating the emotional state as output. Specifically, "Good morning" is analyzed as having a cheerful tone, and the emotional data "positive" is obtained.
[0753] Step 5:
[0754] Real-time text translation
[0755] The server translates the text and emotion data into the specified language. The text data and emotion data are used as input. The Google Translate API is used for data calculation. The server sends the text data and emotion data to the API and receives the translated text data as output. Specifically, "Ohayou gozaimasu" is translated into "Good morning," and the corresponding positive emotion data is also obtained.
[0756] Step 6:
[0757] Text-to-speech translation
[0758] The server converts the translated text data into speech data. The inputs are the translated text, emotion data, and the original speech feature data. Amazon Polly is used for data processing. The server sends this data to a speech synthesis engine, and the output is speech data with the user's speech features. Specifically, it generates speech data with a positive tone based on the text "Good morning."
[0759] Step 7:
[0760] Sending and outputting synthesized speech
[0761] The server sends the generated voice data to the terminal. The voice data generated by the server is used as input. The terminal receives the data and prepares to play it. The terminal decodes the received voice data and plays it from the speaker. The specific operation is that the terminal plays the voice data "Good morning" so that the user can hear the voice.
[0762] (Application example 2)
[0763] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0764] With modern globalization, there is an increasing need to communicate naturally with users who speak different languages. However, language barriers and misunderstandings of emotions are common, which can lead to poor service quality, especially in brick-and-mortar stores and service industries. Furthermore, many existing translation systems ignore emotional states, resulting in inaccurate communication of user intent. Furthermore, they lack the ability to automatically detect and warn against sensitive topics and language.
[0765] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting a user's voice into a digital signal, means for transmitting the digital signal to the server, means for converting the voice digital signal into character string text on the server, means for translating the converted text into a specified language in real time, means for converting the translated text into voice data that preserves the original voice characteristics and emotional state, means for transmitting the voice data to a terminal, means for playing the voice data on the terminal, and means for supporting multilingual communication between users and staff in a physical store. This enables users and staff who speak different languages to communicate naturally while preserving voice characteristics and emotions, improving the quality of service. Furthermore, the safety of conversations is ensured by detecting sensitive content and issuing alerts at appropriate times.
[0766] The "means for converting a user's voice into a digital signal" refers to a device or system that converts the voice emitted by the user from analog to digital.
[0767] The "means for transmitting a digital signal to a server" refers to a device or system for transmitting the converted digital signal data to a server via a network.
[0768] The "means for converting the digital audio signal into a character string text on the server" refers to software or hardware on the server side that executes the process of converting the transmitted digital audio signal into a character string.
[0769] The "means for translating converted text into a specified language in real time" is software or a system that translates a string of text into another specified language in real time.
[0770] "Means for converting translated text into speech data that retains the original speech characteristics and emotional state" refers to a device or system that converts translated text into speech data that reflects the user's speech characteristics and emotions.
[0771] The "means for transmitting voice data to a terminal" refers to a device or system that transmits the generated voice data to a user's terminal via a network.
[0772] "Means for reproducing audio data on a terminal" refers to a device or system for reproducing received audio data on a speaker or the like of the terminal.
[0773] "Means for supporting multilingual communication between users and staff in a physical store" refers to devices or systems that support natural communication between users and staff who speak different languages in a physical store environment.
[0774] "Sensitive content detection means" means software or systems that detect sensitive content in speech or text in real time.
[0775] "Means for generating a sensitive alert and notifying the terminal" refers to software or a system that generates an alert and notifies the user's terminal when sensitive content is detected.
[0776] The "means for generating suggestions and instructions to support the progress of a meeting" refers to software or a system that generates appropriate suggestions and instructions based on the progress of the meeting.
[0777] The "means for transmitting suggestions and instructions to a terminal and notifying the user" refers to software or a system that transmits the generated suggestions and instructions to the user's terminal and notifies the user.
[0778] The present invention is a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[0779] Overall system configuration
[0780] This system consists of a user, a terminal, a server, and an emotion engine. When the user speaks into the terminal, the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates the text into the specified language. The translated text is converted into audio data that also retains the emotional state and is sent to the terminal. The user can listen to the translated speech through the terminal's speaker.
[0781] Specific Embodiments
[0782] 1. Audio input and digital signal conversion
[0783] Voice input is captured when a user speaks into the device's microphone. For example, a user might say, "Welcome, what product are you looking for?" in a physical store. This speech is converted into digital signals using voice recognition software within the device.
[0784] 2. Digital signal transmission to server
[0785] The converted digital signal is then sent to a server via a network, using a smartphone or tablet as the main hardware and a data transfer protocol as the software.
[0786] 3. Text Conversion
[0787] The digital signal that reaches the server is converted into text using a speech recognition API (e.g., Google Cloud Speech-to-Text API) on the server, for example, "Welcome, what product are you looking for?"
[0788] 4. Recognizing emotional states
[0789] The textual content is then subjected to emotional recognition using an emotion engine (e.g., Emotion Recognition API). For example, the text is classified into an emotional state such as "welcome."
[0790] 5. Real-time translation
[0791] The server translates the recognized text and emotional state into a specified language (e.g., English) in real time using a translation API (e.g., Google Translate API). In this case, the translation is "Welcome, what kind of product are you looking for?"
[0792] 6. Speech synthesis of translated text
[0793] Based on the translated text and the recognized emotional state, the server generates voice data using speech synthesis technology (e.g., Text-to-Speech API). The voice is generated with a tone based on the emotional state (welcome).
[0794] 7. Sending and outputting synthesized speech
[0795] The server sends the generated voice data to the device, which then plays the received voice data through the speaker, allowing the user to listen to the translated voice in real time through the device.
[0796] Specific use cases
[0797] For example, consider a conversation between a Japanese-speaking staff member and an English-speaking customer in a physical store. When the staff member says in Japanese, "Welcome, what kind of product are you looking for?", the system translates this into English and tells the customer, "Welcome, what kind of product are you looking for?" in a voice that retains the emotional state of the customer. This allows for smooth communication that transcends language barriers.
[0798] Prompt Sentence Examples
[0799] "We are developing a multilingual customer service assistant app that allows users to communicate naturally across language barriers and maintains their emotional state. This app translates the user's voice in real time and reflects their emotional state using an emotion engine. The process involves the following steps: 1. Record the user's voice and convert it into a digital signal. 2. Convert the voice to text (e.g., 'Welcome, what kind of product are you looking for?' -> 'Welcome, what kind of product are you looking for?'). 3. Recognize the emotional state of the text (e.g., 'Welcome'). 4. Translate while maintaining the emotional state and perform speech synthesis. 5. Output the final voice data and provide it to the user. The generated program is as follows:"
[0800] In this way, the present invention enables users and staff who speak different languages to communicate naturally while maintaining their emotional state.
[0801] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0802] Step 1:
[0803] The user speaks into the device's microphone. The input is the user's voice, which is captured by the voice input software and converted into a digital signal. The data is then processed by sampling the voice signal and converting it into digital data. The output is a digital voice signal.
[0804] Step 2:
[0805] The terminal sends the converted digital voice signal to the server. The input is the digital voice signal generated in step 1, which is sent to the server via the Internet using the terminal's communication module. The data is processed in the process of packetizing the digital signal and sending it. The output is the digital voice signal that has arrived at the server.
[0806] Step 3:
[0807] The server receives the digital voice signal and converts it into text using a speech recognition API. The input is the digital voice signal, which is converted into text by a speech recognition engine (e.g., Google Cloud Speech-to-Text API). The data operations are speech signal processing and natural language processing. The output is text data generated from the voice signal.
[0808] Step 4:
[0809] The server sends text data to an emotion recognition engine to recognize the emotional state. The input is text data, which is analyzed for emotion by an emotion recognition engine (e.g., Emotion Recognition API). The data is processed by string analysis and emotion categorization. The output is the recognized emotional state.
[0810] Step 5:
[0811] The server sends the text data and the recognized emotional state to a translation API, which translates it into the specified language in real time. The input is the text data and the emotional state, and the translation is performed using a translation API (e.g., Google Translate API). The data processing is the language conversion of the text. The output is the translated text data.
[0812] Step 6:
[0813] The server uses the translated text data and emotional state to synthesize speech. The input is the translated text data and emotional state, which is converted into speech data by a speech synthesis engine (e.g., Text-to-Speech API). The data is processed to convert text to speech and set the emotional tone. The output is the synthesized speech data.
[0814] Step 7:
[0815] The server sends the synthesized voice data to the terminal. The input is the synthesized voice data, which is sent to the terminal via the Internet using a communication module. The data is processed by sending data packets. The output is the voice data that has arrived at the terminal.
[0816] Step 8:
[0817] The device decodes the audio data received and plays it from the speaker. The input is audio data, which is decoded using the device's audio decoder and output as audio from the speaker. The data processing involves converting the digital audio signal to analog. The output is synthesized speech that the user can hear.
[0818] This allows users to communicate naturally with other users who speak different languages.
[0819] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0820] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0821] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0822] [Third embodiment]
[0823] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0824] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0825] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0826] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0827] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0828] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0829] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0830] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0831] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0832] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0833] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0834] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0835] Understood. Below is the "Mode for carrying out the invention".
[0836] ---
[0837] The present invention relates to a translation multi-assist AI system that overcomes language barriers in international communication and enables natural conversation. Specific embodiments of the present invention will be described below.
[0838] Overall system configuration
[0839] This system consists of a user, a terminal, and a server. The user inputs speech, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into speech data that retains the user's speech characteristics and sent to the terminal. The terminal plays back the received speech data and provides the translation result to the user.
[0840] Program processing explanation
[0841] User voice input
[0842] The user speaks into the device's microphone. For example, if the user says "Good morning" in Japanese, this speech is captured by the device.
[0843] Speech recognition and text conversion
[0844] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this example, the audio "Good morning" is converted into the text "Good morning."
[0845] Real-time translation
[0846] The server passes the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0847] Text-to-speech translation
[0848] The server analyzes the user's voice characteristics and uses speech synthesis technology to convert the translated text into audio data in the user's voice. The server synthesizes "Good morning" in English, preserving the pitch, tone, and speed of the user's voice.
[0849] Sending and outputting synthesized speech
[0850] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0851] Specific examples
[0852] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0853] Sensitive Alert Function
[0854] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[0855] Facilitator function
[0856] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[0857] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[0858] ---
[0859] The processing flow will be explained below.
[0860] Understood. Below I will explain the process in concrete steps.
[0861] ---
[0862] Step 1:
[0863] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[0864] Step 2:
[0865] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[0866] Step 3:
[0867] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text as "Good morning."
[0868] Step 4:
[0869] The server sends the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[0870] Step 5:
[0871] The server converts the translated text into audio data that preserves the user's voice characteristics. Specifically, it analyzes the pitch, tone, and speed of the user's original speech and uses speech synthesis technology to generate "Good morning" in the user's voice.
[0872] Step 6:
[0873] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[0874] Step 7:
[0875] The device decodes the received audio data and plays it through the speaker, allowing the user to hear the translated "Good morning" in their own voice.
[0876] Step 8:
[0877] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[0878] Step 9:
[0879] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[0880] Step 10:
[0881] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[0882] Step 11:
[0883] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[0884] ---
[0885] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[0886] Example 1
[0887] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0888] In international communication, it is difficult to achieve natural and smooth conversations between people who speak different languages. Furthermore, when conversation content contains sensitive information, it is important to have a system that can appropriately process and notify the content. Furthermore, a system that supports the progress of meetings and important communications is also needed.
[0889] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0890] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server via a communication network; means for converting the voice digital signal into character string text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the characteristics of the original voice; means for transmitting the voice data to a terminal via a communication network; means for playing the voice data on the terminal; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of the conference; and means for transmitting the suggestions or instructions to the terminal and notifying the user. This enables natural and smooth conversations between people speaking different languages, enables appropriate processing and notification of sensitive content, and supports the progress of conferences and important communications.
[0891] Understood. Below are definitions of important terms included in the claims mentioned above.
[0892] ---
[0893] A "digital signal" refers to a signal that represents analog audio data in numerical form.
[0894] "Communication network" refers to communication infrastructure such as the Internet or local networks used to send and receive data.
[0895] A "server" refers to a computer system that receives and processes requests from users over a network.
[0896] "Speech recognition API" refers to a programming interface for receiving voice data as input and converting it into corresponding text data.
[0897] "Text string" refers to text data converted by speech recognition.
[0898] "Translation API" refers to a program interface for converting input text data into another specified language.
[0899] "Speech features" refer to individual characteristics of speech such as pitch, tone, and rate.
[0900] "Audio Data" means an audio file represented in digital form.
[0901] "Sensitive alert" refers to a warning that notifies the user when sensitive content is detected.
[0902] "Suggestions and Instructions" refers to recommendations and instructions generated to help facilitate conversation or communication.
[0903] "Terminal" refers to a device (e.g., a smartphone or PC) that a user directly operates and that communicates with a server.
[0904] ---
[0905] These definitions are intended to clarify the meaning of key terms in the claims.
[0906] The present invention is a system that translates a user's voice input in real time and outputs the translated text as voice data that retains the original voice characteristics. This system is composed of a user, a terminal, and a server.
[0907] The user speaks into the microphone of the device. For example, when the user says "Good morning," the voice is captured by the device. The device converts the voice into a digital signal through the microphone and transmits the digital signal to the server via a communication network.
[0908] The server converts the received digital signal into a string of text using a speech recognition API (e.g., Google Cloud Speech-to-Text API). In this example, the speech "Good morning" is converted into the text "Good morning."
[0909] The server then passes the converted text to a translation API (e.g., Google Cloud Translation API) for real-time translation into the specified language (e.g., English). Here, the text "Ohayo gozaimasu" is translated into English as "Good morning."
[0910] The server uses speech synthesis technology (e.g., Amazon Polly) to convert the translated text into speech that retains the characteristics of the original voice, so that the translated "Good morning" is converted into speech that reflects the pitch, tone, and rate of the user's voice.
[0911] The server then sends the generated voice data to the device via a communication network. The device then decodes the received voice data and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[0912] Furthermore, this system has the function of detecting sensitive content. The server analyzes the conversation content, and if it detects sensitive content, it generates a sensitive alert and notifies the terminal. The terminal then presents this alert to the user in text or voice format, allowing the user to respond appropriately.
[0913] The server also has the ability to generate suggestions and instructions to help meetings and important communications proceed smoothly. For example, if a conversation stalls, the server will automatically generate a suggestion such as, "Shall we move on to the next topic?" This allows the device to notify the user of the suggestion or instruction, helping to ensure the smooth progress of the meeting.
[0914] As a concrete example, consider a situation in which a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[0915] An example of a prompt sentence is, "A Japanese-speaking user says 'Good morning' into the terminal. Explain how this speech is converted into a digital signal, processed by the server, and finally played back as English speech."
[0916] As described above, the present invention enables natural and smooth conversations between people who speak different languages, and enables sensitive content to be appropriately handled and communicated, thereby supporting the progress of meetings and important communications.
[0917] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0918] Understood. Below, we will divide the processing flow of the system program into specific processing steps and explain the specific operations including input and output.
[0919] ---
[0920] Step 1:
[0921] Capture audio input
[0922] The user speaks into the device's microphone.
[0923] Input: User speech (e.g. "Good morning").
[0924] Device behavior: The device uses a microphone to capture the user's voice.
[0925] Output: Analog audio signal.
[0926] Step 2:
[0927] Digital conversion of analog audio signals
[0928] The terminal converts the captured analog voice signal into a digital signal.
[0929] Input: Analog audio signal.
[0930] Terminal operation: The A / D converter inside the terminal converts the audio signal into a digital signal and records it as time-series data.
[0931] Output: Digital audio signal.
[0932] Step 3:
[0933] Transmitting digital audio signals
[0934] The terminal transmits the digital audio signal to the server over a communication network.
[0935] Input: Digital audio signal.
[0936] Terminal operation: The voice signal is converted into packet data and sent to a server via the Internet through a communications module.
[0937] Output: Audio data packets to the server.
[0938] Step 4:
[0939] Speech recognition and text conversion
[0940] The server inputs the received digital voice signal into a voice recognition API and converts it into text.
[0941] Input: Digital audio signal.
[0942] Server operation: A speech recognition API (e.g., a speech recognition cloud service) converts the speech data into text data.
[0943] Output: String text (e.g. "Good morning").
[0944] Step 5:
[0945] Text translation
[0946] The server passes the converted text to the translation API, which translates it into the specified language.
[0947] Input: String text (e.g. "Good morning").
[0948] Server action: Translate the text into the specified language (e.g., English) using a translation API (e.g., a translation cloud service).
[0949] Output: The translated text (e.g. "Good morning").
[0950] Step 6:
[0951] Speech synthesis of translated text
[0952] The server converts the translated text into voice data using voice synthesis technology.
[0953] Input: The translated text (e.g. "Good morning") and the user's voice characteristics.
[0954] Server operation: A speech synthesis API (e.g., a speech synthesis cloud service) converts text into speech data that retains the user's speech characteristics.
[0955] Output: Synthetic speech data.
[0956] Step 7:
[0957] Sending audio data
[0958] The server transmits the generated voice data to the terminal via a communication network.
[0959] Input: Synthetic speech data.
[0960] Server operation: Converts voice data into packet data and sends it to the terminal via the Internet through the communication module.
[0961] Output: Audio data packets destined for the device.
[0962] Step 8:
[0963] Playing audio data
[0964] The terminal plays the received audio data through the speaker.
[0965] Input: Synthetic speech data.
[0966] Device operation: The decoded audio data is converted into an analog signal using a digital-to-analog converter and played through the speaker.
[0967] Output: The speech the user hears (e.g. "Good morning").
[0968] ---
[0969] This concludes the explanation of the specific processing steps of the program. Each step clearly shows the type of data processing and calculation that will be performed based on the input data. This system makes it possible to have natural conversations that transcend language barriers.
[0970] (Application example 1)
[0971] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0972] In international communication, there is a demand for systems that support natural conversations between users who speak different languages. However, existing systems have difficulty preserving the characteristics of speech while translating, and lack technology to support communication between various languages in autonomous vehicles. Furthermore, they lack the ability to detect and notify sensitive content in real time, and the ability to support the progress of meetings.
[0973] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0974] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server; means for converting the voice digital signal into text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the original voice characteristics; means for transmitting the voice data to a terminal; means for playing the voice data on the terminal; means for supporting communication between users in various languages within an autonomous vehicle; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of a meeting; means for transmitting the suggestions or instructions to the terminal and notifying the user; and a function for translating voice in real time using a generative AI model, installed on a smartphone, smart glasses, or head-mounted display. This supports natural communication between users who speak different languages and enables smooth communication across language barriers within an autonomous vehicle. Furthermore, the functions for detecting and notifying sensitive content and supporting the progress of a meeting enable safe and efficient communication.
[0975] The "means for converting a user's voice into a digital signal" is a device or program for converting a user's voice from an analog signal into a digital signal.
[0976] The "means for transmitting the digital signal to the server" refers to a communication device or program for transmitting the converted digital signal to the server via a network.
[0977] The "means for converting the digital audio signal into a character string text on the server" refers to software or a program for converting the digital audio signal into a text format on the server using voice recognition technology.
[0978] "Means for translating converted text into a specified language in real time" means a technology or program for instantly translating text into another specified language.
[0979] "Means for converting translated text into speech data that preserves the characteristics of the original speech" refers to technology or a program for converting translated text into speech data in a form that preserves the user's speech characteristics.
[0980] The "means for transmitting the voice data to the terminal" is a communication device or a program for transmitting the generated voice data to the user's terminal via a network.
[0981] The "means for reproducing the audio data on the terminal" refers to hardware or software for reproducing the audio data received on the terminal.
[0982] "Means for supporting communication between users in various languages within an autonomous vehicle" refers to technology or a program that enables smooth communication between passengers who speak different languages within an autonomous vehicle.
[0983] "Sensitive content detection means" means an algorithm or program that identifies sensitive or inappropriate content in real time.
[0984] The "means for generating a sensitive alert and notifying the terminal" is a program for generating a warning for detected sensitive content and notifying the user's terminal of the warning.
[0985] The "means for generating suggestions and instructions to support the progress of a meeting" refers to an algorithm or program for generating appropriate suggestions and instructions so that the meeting can proceed smoothly.
[0986] The "means for transmitting the proposal or instruction to the terminal and notifying the user" refers to a communication device or program for transmitting the generated proposal or instruction to the user's terminal and notifying the user of it.
[0987] "A feature installed on a smartphone, smart glasses, or head-mounted display that uses a generative AI model to translate speech in real time" refers to software or an application that utilizes the AI model built into these devices to provide real-time speech translation functionality.
[0988] The present invention relates to a translation multi-assist AI system for supporting natural conversation between users who speak different languages. Hereinafter, an embodiment of the present invention will be specifically described.
[0989] Overall system configuration
[0990] This system consists of a user, a terminal, the environment inside the autonomous vehicle, and a server. The user inputs voice through the terminal's microphone, and the terminal converts the voice into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language in real time. The translated text is then converted back into voice data that retains the user's voice characteristics and sent to the terminal. Finally, the terminal plays back the received voice data and provides the translation result to the user.
[0991] Program processing explanation
[0992] Audio input and digital signal conversion
[0993] The user speaks into the microphone of the device. For example, when the user says, "Take me to the airport, please," this voice is captured by the microphone of the device. The device converts the voice into a digital signal, which is then transmitted to the server via the network.
[0994] Speech recognition and text conversion
[0995] The server uses a speech recognition API (Google's speech recognition API is a specific example) to convert digital voice signals into text. This technology converts the speech "Please take me to the airport" into the text "Please take me to the airport."
[0996] Real-time translation
[0997] The server passes the converted text to a translation API (e.g., Google Translate API) and translates it into the specified language (e.g., English) in real time. In this example, "Kuukou de onegaishimasu" is translated to "Please go to the airport."
[0998] Text-to-speech translation
[0999] The server analyzes the user's voice characteristics and converts the translated text into audio data in the user's voice using speech synthesis technology (e.g., Google Text-to-Speech). The server generates audio data saying "Please go to the airport" in English, while retaining the user's voice characteristics.
[1000] Sending and playing audio data
[1001] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Please go to the airport" translated in their own voice.
[1002] Sensitive Alert Function
[1003] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitive alert in real time and notifies the device. This alert is presented in text or audio format, allowing the user to respond appropriately.
[1004] Facilitator function
[1005] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[1006] Specific examples
[1007] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Please go to the airport," this speech is sent to the server via the device. The server converts the speech into text "Please go to the airport" and translates it into English "Please go to the airport." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Please go to the airport" in user (A)'s voice.
[1008] Prompt Sentence Examples
[1009] "Create an AI model that automatically translates what is said in the car into English. Convert the incoming Japanese speech into text, translate it into English using an API, and then convert it back into speech and play it back."
[1010] Thus, the present invention is a system that provides a natural translation function between users who speak different languages, detects sensitive content, and supports the progress of meetings.
[1011] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1012] Step 1:
[1013] The user speaks into a terminal inside the autonomous vehicle. The terminal's microphone captures the user's voice and converts the analog voice signal into a digital signal. The input is the user's voice, and the output is a digital voice signal. Specifically, the terminal's voice input device captures the voice, and the built-in digital signal processor digitizes it.
[1014] Step 2:
[1015] The terminal transmits the converted digital signal to the server via the network. The input is a digital audio signal, and the output is the transfer of the digital signal to the server. Specifically, the terminal's communication module (e.g., Wi-Fi or mobile data communication) transmits the digital signal to the server.
[1016] Step 3:
[1017] The server converts the received digital voice signal into a string of text using a speech recognition API. The input is a digital voice signal, and the output is text data such as "Take me to the airport, please." Specifically, the server calls speech recognition software (e.g., Google Speech-to-Text API) and converts the voice signal into a string of text.
[1018] Step 4:
[1019] The server translates text data into the specified language in real time using a translation API. The input is text data, and the output is translated text such as "Please go to the airport." Specifically, the server uses translation software (e.g., Google Translate API) to convert the text into the specified language.
[1020] Step 5:
[1021] The server converts the translated text into audio data using speech synthesis technology. The input is the translated text, and the output is audio data. Specifically, the server uses speech synthesis software (e.g., Google Text-to-Speech API) to convert the text into audio data that retains the user's voice characteristics.
[1022] Step 6:
[1023] The server transmits the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal. In concrete terms, the communication module of the server transmits the voice data to the terminal.
[1024] Step 7:
[1025] The device decodes the received audio data and plays it through the speaker. The input is audio data and the output is audio output from the speaker. Specifically, the device's audio playback device decodes the audio data and plays it as audio through the speaker.
[1026] Step 8:
[1027] The server analyzes the content of the conversation and detects sensitive content. The input is voice or text data, and the output is a sensitive alert. Specifically, the server uses natural language processing technology to analyze the content of the conversation and detect whether it contains sensitive information. If it detects sensitive information, it generates an alert.
[1028] Step 9:
[1029] When a sensitive alert is generated, the server notifies the terminal of the alert. The input is the sensitive alert, and the output is a notification to the terminal. Specifically, the server's notification system sends the alert to the terminal and causes the terminal to display the notification.
[1030] Step 10:
[1031] To help facilitate the progress of the meeting, the server analyzes the conversation and generates appropriate suggestions and instructions. The input is conversation data, and the output is suggestions and instructions. Specifically, the server uses a generative AI model to understand the context of the conversation and generate appropriate suggestions and instructions.
[1032] Step 11:
[1033] The server sends the generated suggestions and instructions to the terminal and notifies the user. The input is the suggestions and instructions, and the output is a presentation to the terminal. Specifically, the instruction data is sent from the server to the terminal, and a notification is displayed on the terminal to the user.
[1034] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1035] Understood. Below is the "Mode for carrying out the invention" regarding the invention that combines an emotion engine.
[1036] ---
[1037] The present invention relates to a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[1038] Overall system configuration
[1039] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[1040] Program processing explanation
[1041] User voice input
[1042] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[1043] Speech recognition and text conversion
[1044] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this case, "Good morning" is converted into text.
[1045] Recognition of emotional states
[1046] The server uses an emotion engine to recognize the user's emotional state from the input speech. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes a "positive" emotional state.
[1047] Real-time translation
[1048] The server sends the converted text and the recognized emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1049] Text-to-speech translation
[1050] The server analyzes the user's voice characteristics and the recognized emotional state, and uses speech synthesis technology to convert the translated text into audio data in the user's voice and emotional state, where "Good morning" is generated in the user's voice and in the recognized "positive" tone.
[1051] Sending and outputting synthesized speech
[1052] The server sends the generated voice data to the device, which then decodes the received voice data and plays it back through the speaker. This allows the user to hear "Good morning" translated in their own voice and with their own emotions.
[1053] Specific examples
[1054] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. This is then translated into English as "Good morning," and the translated text is synthesized using speech synthesis based on the voice and emotional characteristics of user (A). The speech data sent from the server to the device is played back, and user (B) hears "Good morning" in user (A)'s voice with a "positive" tone.
[1055] Sensitive Alert Function
[1056] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[1057] Facilitator function
[1058] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[1059] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice and emotions, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[1060] ---
[1061] The processing flow will be explained below.
[1062] Understood. Below, I will explain the process flow for the invention that combines the emotion engine, broken down into specific steps.
[1063] ---
[1064] Step 1:
[1065] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[1066] Step 2:
[1067] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[1068] Step 3:
[1069] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text "Good morning."
[1070] Step 4:
[1071] The server passes the recognized text to the emotion engine to recognize the user's emotional state. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes the user's emotional state as "positive."
[1072] Step 5:
[1073] The server sends the text with the emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1074] Step 6:
[1075] The server generates voice data that reflects the user's voice characteristics and emotional state based on the translated text and the recognized emotional state. Specifically, it generates "Good morning" in the user's voice while maintaining a "positive" tone.
[1076] Step 7:
[1077] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[1078] Step 8:
[1079] The device decodes the received audio data and plays it through the speaker, allowing the user to hear "Good morning" translated in their own voice and emotional state.
[1080] Step 9:
[1081] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[1082] Step 10:
[1083] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[1084] Step 11:
[1085] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[1086] Step 12:
[1087] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[1088] ---
[1089] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[1090] Example 2
[1091] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1092] Conventional translation systems are capable of converting speech to text and translating it into another language, but they do not take the user's emotional state into account when outputting the speech. As a result, the translated speech may sound unnatural or may not properly convey the user's emotions. Furthermore, they lack the functionality to provide real-time warnings for sensitive content or to support the progress of meetings. This poses a challenge, making smooth communication difficult in international and business settings.
[1093] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing an emotional state from text data, a means for translating the converted text and emotional state data into a specified language in real time, and a means for converting the translated text and the recognized emotional state into voice data that retains the characteristics of the original voice. This enables natural voice translation that reflects the user's emotional state. Furthermore, the system also includes a function for detecting sensitive content in real time and issuing a warning, as well as a support function for facilitating the progress of meetings, thereby realizing safer and more efficient communication.
[1094] "User" refers to an individual or organization that uses the system.
[1095] "Terminal" refers to electronic devices used by users, such as computers, smartphones, and tablets.
[1096] A "server" is a computer system that processes and stores data, and refers to a device that communicates with terminals over a network.
[1097] "Audio digital signal" refers to audio captured by an audio input device such as a microphone and converted into digital data.
[1098] "Text string" refers to data consisting of characters such as alphabets and kanji characters converted from speech using speech recognition technology.
[1099] "Emotional state" refers to the speaker's emotion that can be read from speech or text, such as positive, negative, or neutral.
[1100] "Real-time translation" refers to the process by which input data (voice or text) is translated into another language within a short time of being received.
[1101] "Voice features" refer to the characteristics of a user's voice and speaking style, and refer to elements that reproduce the way the user is speaking.
[1102] "Audio Data" means data for storing, processing, and playing audio signals in digital form.
[1103] A "sensitive alert" is a message that warns users when content is deemed to contain inappropriate or sensitive content.
[1104] "Suggestions and instructions to support the progress of meetings" refers to organizational suggestions and instructions that are generated to ensure that meetings and conversations proceed smoothly.
[1105] The present invention relates to a system that translates a user's voice in real time and enables natural international communication by combining it with an emotion engine. First, the overall configuration of this system will be described.
[1106] Overall system configuration
[1107] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[1108] Hardware and software used
[1109] Specific hardware includes the smartphone, tablet, or computer used by the user, while the following APIs and services are used as software:
[1110] Speech Recognition API: Google Cloud Speech-to-Text API
[1111] Emotion Recognition Engine: Azure Cognitive Services
[1112] Translation API: Google Translate API
[1113] Speech synthesis technology: Amazon Polly
[1114] Operation flow
[1115] The specific operation flow of the system will be explained below.
[1116] User voice input
[1117] The user speaks into the device's microphone. For example, the user says "Good morning" in Japanese. The device's microphone captures this speech and converts it into a digital signal.
[1118] Sending voice data to the server
[1119] The device transmits the captured audio data to a server via the Internet.
[1120] Speech recognition and text conversion
[1121] The server uses the Google Cloud Speech-to-Text API to convert the audio data into a string of text. For example, "Good morning" is converted into "Good morning."
[1122] Recognition of emotional states
[1123] The server uses Azure Cognitive Services to recognize emotional states based on text data. For example, if "Good morning" is spoken in a cheerful tone, the emotion engine will recognize it as "positive."
[1124] Real-time text translation
[1125] The server translates this text and emotion information into the specified language (e.g., English) using the Google Translate API, for example, "Ohayo gozaimasu" is translated to "Good morning."
[1126] Text-to-speech translation
[1127] The server uses Amazon Polly to convert the translated text into speech data that preserves the characteristics and emotion of the original voice, for example, "Good morning" in the user's voice.
[1128] Sending and outputting synthesized speech
[1129] The server then sends the generated voice data to the device, which then decodes it and plays it back through the speaker, allowing the user to hear the translated voice with their own voice and emotions preserved.
[1130] Specific examples
[1131] For example, when a Japanese-speaking user (A) converses with an English-speaking user (B), the following specific operations take place: When user (A) says "Good morning," the device converts this speech into a digital signal and sends it to the server. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. Next, it translates this into English "Good morning," and generates speech data by adding user (A)'s speech features and emotional characteristics based on the translated text. The speech data sent from the server to the device is played back on the device, and user (B) can hear "Good morning" in user (A)'s voice in a cheerful tone.
[1132] Prompt Sentence Examples
[1133] Here are some example prompts to input to the generative AI model:
[1134] When user (A) says "Good morning" in a cheerful voice, the system converts this speech into text and recognizes positive emotions. The server then translates this text into "Good morning" and generates a cheerful voice data in user (A)'s voice. Finally, user (B) can hear "Good morning" in user (A)'s voice.
[1135] The above is a specific embodiment for carrying out the present invention. This system enables natural speech translation that reflects the user's emotions, making international communication smoother.
[1136] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1137] Step 1:
[1138] User voice input
[1139] The user speaks into the microphone of the terminal. The user's voice signal is input to the microphone. In concrete terms, the microphone converts this voice signal into digital voice data. For example, when the user says "Good morning," this voice data is input to the terminal.
[1140] Step 2:
[1141] Sending voice data to the server
[1142] The device sends the captured audio data to the server. As input, the digital audio data is used. As data processing, the device converts the audio data into packets that are sent to the server via the network. As output, packets containing the audio data are sent to the server. The specific operation is that the device transfers the audio data to the server via an Internet connection.
[1143] Step 3:
[1144] Speech recognition and text conversion
[1145] The server converts the received voice data into text. The digital voice data sent to the server is used as input. The Google Cloud Speech-to-Text API is used for data calculations. The server sends the voice data to the API and receives text data as output. Specifically, the server sends a request to the speech recognition API and receives the text data "Good morning" as a response.
[1146] Step 4:
[1147] Recognition of emotional states
[1148] The server recognizes the emotional state from the converted text data. The text data is used as input. Azure Cognitive Services is used for data calculations. The server sends the text to the emotion engine and receives data indicating the emotional state as output. Specifically, "Good morning" is analyzed as having a cheerful tone, and the emotional data "positive" is obtained.
[1149] Step 5:
[1150] Real-time text translation
[1151] The server translates the text and emotion data into the specified language. The text data and emotion data are used as input. The Google Translate API is used for data calculation. The server sends the text data and emotion data to the API and receives the translated text data as output. Specifically, "Ohayou gozaimasu" is translated into "Good morning," and the corresponding positive emotion data is also obtained.
[1152] Step 6:
[1153] Text-to-speech translation
[1154] The server converts the translated text data into speech data. The inputs are the translated text, emotion data, and the original speech feature data. Amazon Polly is used for data processing. The server sends this data to a speech synthesis engine, and the output is speech data with the user's speech features. Specifically, it generates speech data with a positive tone based on the text "Good morning."
[1155] Step 7:
[1156] Sending and outputting synthesized speech
[1157] The server sends the generated voice data to the terminal. The voice data generated by the server is used as input. The terminal receives the data and prepares to play it. The terminal decodes the received voice data and plays it from the speaker. The specific operation is that the terminal plays the voice data "Good morning" so that the user can hear the voice.
[1158] (Application example 2)
[1159] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1160] With modern globalization, there is an increasing need to communicate naturally with users who speak different languages. However, language barriers and misunderstandings of emotions are common, which can lead to poor service quality, especially in brick-and-mortar stores and service industries. Furthermore, many existing translation systems ignore emotional states, resulting in inaccurate communication of user intent. Furthermore, they lack the ability to automatically detect and warn against sensitive topics and language.
[1161] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting a user's voice into a digital signal, means for transmitting the digital signal to the server, means for converting the voice digital signal into character string text on the server, means for translating the converted text into a specified language in real time, means for converting the translated text into voice data that preserves the original voice characteristics and emotional state, means for transmitting the voice data to a terminal, means for playing the voice data on the terminal, and means for supporting multilingual communication between users and staff in a physical store. This enables users and staff who speak different languages to communicate naturally while preserving voice characteristics and emotions, improving the quality of service. Furthermore, the safety of conversations is ensured by detecting sensitive content and issuing alerts at appropriate times.
[1162] The "means for converting a user's voice into a digital signal" refers to a device or system that converts the voice emitted by the user from analog to digital.
[1163] The "means for transmitting a digital signal to a server" refers to a device or system for transmitting the converted digital signal data to a server via a network.
[1164] The "means for converting the digital audio signal into a character string text on the server" refers to software or hardware on the server side that executes the process of converting the transmitted digital audio signal into a character string.
[1165] The "means for translating converted text into a specified language in real time" is software or a system that translates a string of text into another specified language in real time.
[1166] "Means for converting translated text into speech data that retains the original speech characteristics and emotional state" refers to a device or system that converts translated text into speech data that reflects the user's speech characteristics and emotions.
[1167] The "means for transmitting voice data to a terminal" refers to a device or system that transmits the generated voice data to a user's terminal via a network.
[1168] "Means for reproducing audio data on a terminal" refers to a device or system for reproducing received audio data on a speaker or the like of the terminal.
[1169] "Means for supporting multilingual communication between users and staff in a physical store" refers to devices or systems that support natural communication between users and staff who speak different languages in a physical store environment.
[1170] "Sensitive content detection means" means software or systems that detect sensitive content in speech or text in real time.
[1171] "Means for generating a sensitive alert and notifying the terminal" refers to software or a system that generates an alert and notifies the user's terminal when sensitive content is detected.
[1172] The "means for generating suggestions and instructions to support the progress of a meeting" refers to software or a system that generates appropriate suggestions and instructions based on the progress of the meeting.
[1173] The "means for transmitting suggestions and instructions to a terminal and notifying the user" refers to software or a system that transmits the generated suggestions and instructions to the user's terminal and notifies the user.
[1174] The present invention is a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[1175] Overall system configuration
[1176] This system consists of a user, a terminal, a server, and an emotion engine. When the user speaks into the terminal, the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates the text into the specified language. The translated text is converted into audio data that also retains the emotional state and is sent to the terminal. The user can listen to the translated speech through the terminal's speaker.
[1177] Specific Embodiments
[1178] 1. Audio input and digital signal conversion
[1179] Voice input is captured when a user speaks into the device's microphone. For example, a user might say, "Welcome, what product are you looking for?" in a physical store. This speech is converted into digital signals using voice recognition software within the device.
[1180] 2. Digital signal transmission to server
[1181] The converted digital signal is then sent to a server via a network, using a smartphone or tablet as the main hardware and a data transfer protocol as the software.
[1182] 3. Text Conversion
[1183] The digital signal that reaches the server is converted into text using a speech recognition API (e.g., Google Cloud Speech-to-Text API) on the server, for example, "Welcome, what product are you looking for?"
[1184] 4. Recognizing emotional states
[1185] The textual content is then subjected to emotional recognition using an emotion engine (e.g., Emotion Recognition API). For example, the text is classified into an emotional state such as "welcome."
[1186] 5. Real-time translation
[1187] The server translates the recognized text and emotional state into a specified language (e.g., English) in real time using a translation API (e.g., Google Translate API). In this case, the translation is "Welcome, what kind of product are you looking for?"
[1188] 6. Speech synthesis of translated text
[1189] Based on the translated text and the recognized emotional state, the server generates voice data using speech synthesis technology (e.g., Text-to-Speech API). The voice is generated with a tone based on the emotional state (welcome).
[1190] 7. Sending and outputting synthesized speech
[1191] The server sends the generated voice data to the device, which then plays the received voice data through the speaker, allowing the user to listen to the translated voice in real time through the device.
[1192] Specific use cases
[1193] For example, consider a conversation between a Japanese-speaking staff member and an English-speaking customer in a physical store. When the staff member says in Japanese, "Welcome, what kind of product are you looking for?", the system translates this into English and tells the customer, "Welcome, what kind of product are you looking for?" in a voice that retains the emotional state of the customer. This allows for smooth communication that transcends language barriers.
[1194] Prompt Sentence Examples
[1195] "We are developing a multilingual customer service assistant app that allows users to communicate naturally across language barriers and maintains their emotional state. This app translates the user's voice in real time and reflects their emotional state using an emotion engine. The process involves the following steps: 1. Record the user's voice and convert it into a digital signal. 2. Convert the voice to text (e.g., 'Welcome, what kind of product are you looking for?' -> 'Welcome, what kind of product are you looking for?'). 3. Recognize the emotional state of the text (e.g., 'Welcome'). 4. Translate while maintaining the emotional state and perform speech synthesis. 5. Output the final voice data and provide it to the user. The generated program is as follows:"
[1196] In this way, the present invention enables users and staff who speak different languages to communicate naturally while maintaining their emotional state.
[1197] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1198] Step 1:
[1199] The user speaks into the device's microphone. The input is the user's voice, which is captured by the voice input software and converted into a digital signal. The data is then processed by sampling the voice signal and converting it into digital data. The output is a digital voice signal.
[1200] Step 2:
[1201] The terminal sends the converted digital voice signal to the server. The input is the digital voice signal generated in step 1, which is sent to the server via the Internet using the terminal's communication module. The data is processed in the process of packetizing the digital signal and sending it. The output is the digital voice signal that has arrived at the server.
[1202] Step 3:
[1203] The server receives the digital voice signal and converts it into text using a speech recognition API. The input is the digital voice signal, which is converted into text by a speech recognition engine (e.g., Google Cloud Speech-to-Text API). The data operations are speech signal processing and natural language processing. The output is text data generated from the voice signal.
[1204] Step 4:
[1205] The server sends text data to an emotion recognition engine to recognize the emotional state. The input is text data, which is analyzed for emotion by an emotion recognition engine (e.g., Emotion Recognition API). The data is processed by string analysis and emotion categorization. The output is the recognized emotional state.
[1206] Step 5:
[1207] The server sends the text data and the recognized emotional state to a translation API, which translates it into the specified language in real time. The input is the text data and the emotional state, and the translation is performed using a translation API (e.g., Google Translate API). The data processing is the language conversion of the text. The output is the translated text data.
[1208] Step 6:
[1209] The server uses the translated text data and emotional state to synthesize speech. The input is the translated text data and emotional state, which is converted into speech data by a speech synthesis engine (e.g., Text-to-Speech API). The data is processed to convert text to speech and set the emotional tone. The output is the synthesized speech data.
[1210] Step 7:
[1211] The server sends the synthesized voice data to the terminal. The input is the synthesized voice data, which is sent to the terminal via the Internet using a communication module. The data is processed by sending data packets. The output is the voice data that has arrived at the terminal.
[1212] Step 8:
[1213] The device decodes the audio data received and plays it from the speaker. The input is audio data, which is decoded using the device's audio decoder and output as audio from the speaker. The data processing involves converting the digital audio signal to analog. The output is synthesized speech that the user can hear.
[1214] This allows users to communicate naturally with other users who speak different languages.
[1215] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1216] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1217] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1218] [Fourth embodiment]
[1219] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1220] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1221] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1222] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1223] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1225] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1226] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1227] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1228] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1229] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1230] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1231] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1232] Understood. Below is the "Mode for carrying out the invention".
[1233] ---
[1234] The present invention relates to a translation multi-assist AI system that overcomes language barriers in international communication and enables natural conversation. Specific embodiments of the present invention will be described below.
[1235] Overall system configuration
[1236] This system consists of a user, a terminal, and a server. The user inputs speech, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into speech data that retains the user's speech characteristics and sent to the terminal. The terminal plays back the received speech data and provides the translation result to the user.
[1237] Program processing explanation
[1238] User voice input
[1239] The user speaks into the device's microphone. For example, if the user says "Good morning" in Japanese, this speech is captured by the device.
[1240] Speech recognition and text conversion
[1241] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this example, the audio "Good morning" is converted into the text "Good morning."
[1242] Real-time translation
[1243] The server passes the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1244] Text-to-speech translation
[1245] The server analyzes the user's voice characteristics and uses speech synthesis technology to convert the translated text into audio data in the user's voice. The server synthesizes "Good morning" in English, preserving the pitch, tone, and speed of the user's voice.
[1246] Sending and outputting synthesized speech
[1247] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[1248] Specific examples
[1249] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[1250] Sensitive Alert Function
[1251] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[1252] Facilitator function
[1253] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[1254] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[1255] ---
[1256] The processing flow will be explained below.
[1257] Understood. Below I will explain the process in concrete steps.
[1258] ---
[1259] Step 1:
[1260] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[1261] Step 2:
[1262] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[1263] Step 3:
[1264] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text as "Good morning."
[1265] Step 4:
[1266] The server sends the converted text to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1267] Step 5:
[1268] The server converts the translated text into audio data that preserves the user's voice characteristics. Specifically, it analyzes the pitch, tone, and speed of the user's original speech and uses speech synthesis technology to generate "Good morning" in the user's voice.
[1269] Step 6:
[1270] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[1271] Step 7:
[1272] The device decodes the received audio data and plays it through the speaker, allowing the user to hear the translated "Good morning" in their own voice.
[1273] Step 8:
[1274] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[1275] Step 9:
[1276] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[1277] Step 10:
[1278] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[1279] Step 11:
[1280] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[1281] ---
[1282] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[1283] Example 1
[1284] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1285] In international communication, it is difficult to achieve natural and smooth conversations between people who speak different languages. Furthermore, when conversation content contains sensitive information, it is important to have a system that can appropriately process and notify the content. Furthermore, a system that supports the progress of meetings and important communications is also needed.
[1286] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1287] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server via a communication network; means for converting the voice digital signal into character string text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the characteristics of the original voice; means for transmitting the voice data to a terminal via a communication network; means for playing the voice data on the terminal; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of the conference; and means for transmitting the suggestions or instructions to the terminal and notifying the user. This enables natural and smooth conversations between people speaking different languages, enables appropriate processing and notification of sensitive content, and supports the progress of conferences and important communications.
[1288] Understood. Below are definitions of important terms included in the claims mentioned above.
[1289] ---
[1290] A "digital signal" refers to a signal that represents analog audio data in numerical form.
[1291] "Communication network" refers to communication infrastructure such as the Internet or local networks used to send and receive data.
[1292] A "server" refers to a computer system that receives and processes requests from users over a network.
[1293] "Speech recognition API" refers to a programming interface for receiving voice data as input and converting it into corresponding text data.
[1294] "Text string" refers to text data converted by speech recognition.
[1295] "Translation API" refers to a program interface for converting input text data into another specified language.
[1296] "Speech features" refer to individual characteristics of speech such as pitch, tone, and rate.
[1297] "Audio Data" means an audio file represented in digital form.
[1298] "Sensitive alert" refers to a warning that notifies the user when sensitive content is detected.
[1299] "Suggestions and Instructions" refers to recommendations and instructions generated to help facilitate conversation or communication.
[1300] "Terminal" refers to a device (e.g., a smartphone or PC) that a user directly operates and that communicates with a server.
[1301] ---
[1302] These definitions are intended to clarify the meaning of key terms in the claims.
[1303] The present invention is a system that translates a user's voice input in real time and outputs the translated text as voice data that retains the original voice characteristics. This system is composed of a user, a terminal, and a server.
[1304] The user speaks into the microphone of the device. For example, when the user says "Good morning," the voice is captured by the device. The device converts the voice into a digital signal through the microphone and transmits the digital signal to the server via a communication network.
[1305] The server converts the received digital signal into a string of text using a speech recognition API (e.g., Google Cloud Speech-to-Text API). In this example, the speech "Good morning" is converted into the text "Good morning."
[1306] The server then passes the converted text to a translation API (e.g., Google Cloud Translation API) for real-time translation into the specified language (e.g., English). Here, the text "Ohayo gozaimasu" is translated into English as "Good morning."
[1307] The server uses speech synthesis technology (e.g., Amazon Polly) to convert the translated text into speech that retains the characteristics of the original voice, so that the translated "Good morning" is converted into speech that reflects the pitch, tone, and rate of the user's voice.
[1308] The server then sends the generated voice data to the device via a communication network. The device then decodes the received voice data and plays it back through the speaker. The user can then hear "Good morning" translated in their own voice.
[1309] Furthermore, this system has the function of detecting sensitive content. The server analyzes the conversation content, and if it detects sensitive content, it generates a sensitive alert and notifies the terminal. The terminal then presents this alert to the user in text or voice format, allowing the user to respond appropriately.
[1310] The server also has the ability to generate suggestions and instructions to help meetings and important communications proceed smoothly. For example, if a conversation stalls, the server will automatically generate a suggestion such as, "Shall we move on to the next topic?" This allows the device to notify the user of the suggestion or instruction, helping to ensure the smooth progress of the meeting.
[1311] As a concrete example, consider a situation in which a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and translates it into English "Good morning." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Good morning" in user (A)'s voice.
[1312] An example of a prompt sentence is, "A Japanese-speaking user says 'Good morning' into the terminal. Explain how this speech is converted into a digital signal, processed by the server, and finally played back as English speech."
[1313] As described above, the present invention enables natural and smooth conversations between people who speak different languages, and enables sensitive content to be appropriately handled and communicated, thereby supporting the progress of meetings and important communications.
[1314] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1315] Understood. Below, we will divide the processing flow of the system program into specific processing steps and explain the specific operations including input and output.
[1316] ---
[1317] Step 1:
[1318] Capture audio input
[1319] The user speaks into the device's microphone.
[1320] Input: User speech (e.g. "Good morning").
[1321] Device behavior: The device uses a microphone to capture the user's voice.
[1322] Output: Analog audio signal.
[1323] Step 2:
[1324] Digital conversion of analog audio signals
[1325] The terminal converts the captured analog voice signal into a digital signal.
[1326] Input: Analog audio signal.
[1327] Terminal operation: The A / D converter inside the terminal converts the audio signal into a digital signal and records it as time-series data.
[1328] Output: Digital audio signal.
[1329] Step 3:
[1330] Transmitting digital audio signals
[1331] The terminal transmits the digital audio signal to the server over a communication network.
[1332] Input: Digital audio signal.
[1333] Terminal operation: The voice signal is converted into packet data and sent to a server via the Internet through a communications module.
[1334] Output: Audio data packets to the server.
[1335] Step 4:
[1336] Speech recognition and text conversion
[1337] The server inputs the received digital voice signal into a voice recognition API and converts it into text.
[1338] Input: Digital audio signal.
[1339] Server operation: A speech recognition API (e.g., a speech recognition cloud service) converts the speech data into text data.
[1340] Output: String text (e.g. "Good morning").
[1341] Step 5:
[1342] Text translation
[1343] The server passes the converted text to the translation API, which translates it into the specified language.
[1344] Input: String text (e.g. "Good morning").
[1345] Server action: Translate the text into the specified language (e.g., English) using a translation API (e.g., a translation cloud service).
[1346] Output: The translated text (e.g. "Good morning").
[1347] Step 6:
[1348] Speech synthesis of translated text
[1349] The server converts the translated text into voice data using voice synthesis technology.
[1350] Input: The translated text (e.g. "Good morning") and the user's voice characteristics.
[1351] Server operation: A speech synthesis API (e.g., a speech synthesis cloud service) converts text into speech data that retains the user's speech characteristics.
[1352] Output: Synthetic speech data.
[1353] Step 7:
[1354] Sending audio data
[1355] The server transmits the generated voice data to the terminal via a communication network.
[1356] Input: Synthetic speech data.
[1357] Server operation: Converts voice data into packet data and sends it to the terminal via the Internet through the communication module.
[1358] Output: Audio data packets destined for the device.
[1359] Step 8:
[1360] Playing audio data
[1361] The terminal plays the received audio data through the speaker.
[1362] Input: Synthetic speech data.
[1363] Device operation: The decoded audio data is converted into an analog signal using a digital-to-analog converter and played through the speaker.
[1364] Output: The speech the user hears (e.g. "Good morning").
[1365] ---
[1366] This concludes the explanation of the specific processing steps of the program. Each step clearly shows the type of data processing and calculation that will be performed based on the input data. This system makes it possible to have natural conversations that transcend language barriers.
[1367] (Application example 1)
[1368] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1369] In international communication, there is a demand for systems that support natural conversations between users who speak different languages. However, existing systems have difficulty preserving the characteristics of speech while translating, and lack technology to support communication between various languages in autonomous vehicles. Furthermore, they lack the ability to detect and notify sensitive content in real time, and the ability to support the progress of meetings.
[1370] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1371] In this invention, the server includes: means for converting a user's voice into a digital signal; means for transmitting the digital signal to the server; means for converting the voice digital signal into text on the server; means for translating the converted text into a specified language in real time; means for converting the translated text into voice data that retains the original voice characteristics; means for transmitting the voice data to a terminal; means for playing the voice data on the terminal; means for supporting communication between users in various languages within an autonomous vehicle; means for detecting sensitive content; means for generating a sensitive alert and notifying the terminal; means for generating suggestions or instructions to support the progress of a meeting; means for transmitting the suggestions or instructions to the terminal and notifying the user; and a function for translating voice in real time using a generative AI model, installed on a smartphone, smart glasses, or head-mounted display. This supports natural communication between users who speak different languages and enables smooth communication across language barriers within an autonomous vehicle. Furthermore, the functions for detecting and notifying sensitive content and supporting the progress of a meeting enable safe and efficient communication.
[1372] The "means for converting a user's voice into a digital signal" is a device or program for converting a user's voice from an analog signal into a digital signal.
[1373] The "means for transmitting the digital signal to the server" refers to a communication device or program for transmitting the converted digital signal to the server via a network.
[1374] The "means for converting the digital audio signal into a character string text on the server" refers to software or a program for converting the digital audio signal into a text format on the server using voice recognition technology.
[1375] "Means for translating converted text into a specified language in real time" means a technology or program for instantly translating text into another specified language.
[1376] "Means for converting translated text into speech data that preserves the characteristics of the original speech" refers to technology or a program for converting translated text into speech data in a form that preserves the user's speech characteristics.
[1377] The "means for transmitting the voice data to the terminal" is a communication device or a program for transmitting the generated voice data to the user's terminal via a network.
[1378] The "means for reproducing the audio data on the terminal" refers to hardware or software for reproducing the audio data received on the terminal.
[1379] "Means for supporting communication between users in various languages within an autonomous vehicle" refers to technology or a program that enables smooth communication between passengers who speak different languages within an autonomous vehicle.
[1380] "Sensitive content detection means" means an algorithm or program that identifies sensitive or inappropriate content in real time.
[1381] The "means for generating a sensitive alert and notifying the terminal" is a program for generating a warning for detected sensitive content and notifying the user's terminal of the warning.
[1382] The "means for generating suggestions and instructions to support the progress of a meeting" refers to an algorithm or program for generating appropriate suggestions and instructions so that the meeting can proceed smoothly.
[1383] The "means for transmitting the proposal or instruction to the terminal and notifying the user" refers to a communication device or program for transmitting the generated proposal or instruction to the user's terminal and notifying the user of it.
[1384] "A feature installed on a smartphone, smart glasses, or head-mounted display that uses a generative AI model to translate speech in real time" refers to software or an application that utilizes the AI model built into these devices to provide real-time speech translation functionality.
[1385] The present invention relates to a translation multi-assist AI system for supporting natural conversation between users who speak different languages. Hereinafter, an embodiment of the present invention will be specifically described.
[1386] Overall system configuration
[1387] This system consists of a user, a terminal, the environment inside the autonomous vehicle, and a server. The user inputs voice through the terminal's microphone, and the terminal converts the voice into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language in real time. The translated text is then converted back into voice data that retains the user's voice characteristics and sent to the terminal. Finally, the terminal plays back the received voice data and provides the translation result to the user.
[1388] Program processing explanation
[1389] Audio input and digital signal conversion
[1390] The user speaks into the microphone of the device. For example, when the user says, "Take me to the airport, please," this voice is captured by the microphone of the device. The device converts the voice into a digital signal, which is then transmitted to the server via the network.
[1391] Speech recognition and text conversion
[1392] The server uses a speech recognition API (Google's speech recognition API is a specific example) to convert digital voice signals into text. This technology converts the speech "Please take me to the airport" into the text "Please take me to the airport."
[1393] Real-time translation
[1394] The server passes the converted text to a translation API (e.g., Google Translate API) and translates it into the specified language (e.g., English) in real time. In this example, "Kuukou de onegaishimasu" is translated to "Please go to the airport."
[1395] Text-to-speech translation
[1396] The server analyzes the user's voice characteristics and converts the translated text into audio data in the user's voice using speech synthesis technology (e.g., Google Text-to-Speech). The server generates audio data saying "Please go to the airport" in English, while retaining the user's voice characteristics.
[1397] Sending and playing audio data
[1398] The server sends the generated voice data to the device, which then decodes it and plays it back through the speaker. The user can then hear "Please go to the airport" translated in their own voice.
[1399] Sensitive Alert Function
[1400] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitive alert in real time and notifies the device. This alert is presented in text or audio format, allowing the user to respond appropriately.
[1401] Facilitator function
[1402] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[1403] Specific examples
[1404] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Please go to the airport," this speech is sent to the server via the device. The server converts the speech into text "Please go to the airport" and translates it into English "Please go to the airport." The server then synthesizes the translated text based on user (A)'s voice characteristics and sends this speech data to the device. The device plays the speech data, and user (B) can hear "Please go to the airport" in user (A)'s voice.
[1405] Prompt Sentence Examples
[1406] "Create an AI model that automatically translates what is said in the car into English. Convert the incoming Japanese speech into text, translate it into English using an API, and then convert it back into speech and play it back."
[1407] Thus, the present invention is a system that provides a natural translation function between users who speak different languages, detects sensitive content, and supports the progress of meetings.
[1408] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1409] Step 1:
[1410] The user speaks into a terminal inside the autonomous vehicle. The terminal's microphone captures the user's voice and converts the analog voice signal into a digital signal. The input is the user's voice, and the output is a digital voice signal. Specifically, the terminal's voice input device captures the voice, and the built-in digital signal processor digitizes it.
[1411] Step 2:
[1412] The terminal transmits the converted digital signal to the server via the network. The input is a digital audio signal, and the output is the transfer of the digital signal to the server. Specifically, the terminal's communication module (e.g., Wi-Fi or mobile data communication) transmits the digital signal to the server.
[1413] Step 3:
[1414] The server converts the received digital voice signal into a string of text using a speech recognition API. The input is a digital voice signal, and the output is text data such as "Take me to the airport, please." Specifically, the server calls speech recognition software (e.g., Google Speech-to-Text API) and converts the voice signal into a string of text.
[1415] Step 4:
[1416] The server translates text data into the specified language in real time using a translation API. The input is text data, and the output is translated text such as "Please go to the airport." Specifically, the server uses translation software (e.g., Google Translate API) to convert the text into the specified language.
[1417] Step 5:
[1418] The server converts the translated text into audio data using speech synthesis technology. The input is the translated text, and the output is audio data. Specifically, the server uses speech synthesis software (e.g., Google Text-to-Speech API) to convert the text into audio data that retains the user's voice characteristics.
[1419] Step 6:
[1420] The server transmits the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal. In concrete terms, the communication module of the server transmits the voice data to the terminal.
[1421] Step 7:
[1422] The device decodes the received audio data and plays it through the speaker. The input is audio data and the output is audio output from the speaker. Specifically, the device's audio playback device decodes the audio data and plays it as audio through the speaker.
[1423] Step 8:
[1424] The server analyzes the content of the conversation and detects sensitive content. The input is voice or text data, and the output is a sensitive alert. Specifically, the server uses natural language processing technology to analyze the content of the conversation and detect whether it contains sensitive information. If it detects sensitive information, it generates an alert.
[1425] Step 9:
[1426] When a sensitive alert is generated, the server notifies the terminal of the alert. The input is the sensitive alert, and the output is a notification to the terminal. Specifically, the server's notification system sends the alert to the terminal and causes the terminal to display the notification.
[1427] Step 10:
[1428] To help facilitate the progress of the meeting, the server analyzes the conversation and generates appropriate suggestions and instructions. The input is conversation data, and the output is suggestions and instructions. Specifically, the server uses a generative AI model to understand the context of the conversation and generate appropriate suggestions and instructions.
[1429] Step 11:
[1430] The server sends the generated suggestions and instructions to the terminal and notifies the user. The input is the suggestions and instructions, and the output is a presentation to the terminal. Specifically, the instruction data is sent from the server to the terminal, and a notification is displayed on the terminal to the user.
[1431] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1432] Understood. Below is the "Mode for carrying out the invention" regarding the invention that combines an emotion engine.
[1433] ---
[1434] The present invention relates to a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[1435] Overall system configuration
[1436] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[1437] Program processing explanation
[1438] User voice input
[1439] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[1440] Speech recognition and text conversion
[1441] The device converts the captured audio into a digital signal and sends it to the server. The server then uses a speech recognition API to convert the digital audio signal into a string of text. In this case, "Good morning" is converted into text.
[1442] Recognition of emotional states
[1443] The server uses an emotion engine to recognize the user's emotional state from the input speech. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes a "positive" emotional state.
[1444] Real-time translation
[1445] The server sends the converted text and the recognized emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1446] Text-to-speech translation
[1447] The server analyzes the user's voice characteristics and the recognized emotional state, and uses speech synthesis technology to convert the translated text into audio data in the user's voice and emotional state, where "Good morning" is generated in the user's voice and in the recognized "positive" tone.
[1448] Sending and outputting synthesized speech
[1449] The server sends the generated voice data to the device, which then decodes the received voice data and plays it back through the speaker. This allows the user to hear "Good morning" translated in their own voice and with their own emotions.
[1450] Specific examples
[1451] For example, consider the case where a Japanese-speaking user (A) is conversing with an English-speaking user (B). When user (A) says "Good morning," this speech is sent to the server via the device. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. This is then translated into English as "Good morning," and the translated text is synthesized using speech synthesis based on the voice and emotional characteristics of user (A). The speech data sent from the server to the device is played back, and user (B) hears "Good morning" in user (A)'s voice with a "positive" tone.
[1452] Sensitive Alert Function
[1453] This system has the ability to detect sensitive content. If the content of a conversation is deemed sensitive, the server generates a sensitivity alert in real time and notifies the device. This notification is presented in text or audio format, allowing the user to respond appropriately.
[1454] Facilitator function
[1455] During meetings or important communications, the server generates suggestions and instructions to help the conversation proceed. For example, if the conversation stalls, the server will suggest, "Shall we move on to the next topic?" This instruction is notified to the user via their device, helping to ensure the smooth progress of the meeting.
[1456] As described above, this invention enables natural conversations that transcend language barriers. Because users can speak different languages using their own voice and emotions, it is expected to play an important role in international exchange and business. Furthermore, the inclusion of sensitive alert and facilitator functions allows for safer and more efficient communication.
[1457] ---
[1458] The processing flow will be explained below.
[1459] Understood. Below, I will explain the process flow for the invention that combines the emotion engine, broken down into specific steps.
[1460] ---
[1461] Step 1:
[1462] The user speaks into the microphone of the device. For example, the user says "Good morning" in Japanese.
[1463] Step 2:
[1464] The device converts the user's voice into a digital signal through a microphone, which is then sent to the server as data packets.
[1465] Step 3:
[1466] The server receives the voice data sent from the device and converts the voice into a string of text using a speech recognition API. In this case, "Good morning" is converted into text "Good morning."
[1467] Step 4:
[1468] The server passes the recognized text to the emotion engine to recognize the user's emotional state. For example, if "Good morning" is spoken in a bright and cheerful tone, the emotion engine recognizes the user's emotional state as "positive."
[1469] Step 5:
[1470] The server sends the text with the emotional state to a translation API, which translates it into the specified language (e.g., English) in real time. In this case, "Ohayo gozaimasu" is translated to "Good morning."
[1471] Step 6:
[1472] The server generates voice data that reflects the user's voice characteristics and emotional state based on the translated text and the recognized emotional state. Specifically, it generates "Good morning" in the user's voice while maintaining a "positive" tone.
[1473] Step 7:
[1474] The server sends the generated voice data to the terminal, where it is encoded as data packets and sent to the terminal.
[1475] Step 8:
[1476] The device decodes the received audio data and plays it through the speaker, allowing the user to hear "Good morning" translated in their own voice and emotional state.
[1477] Step 9:
[1478] If a message contains sensitive content, the server will detect this in real time and generate a sensitivity alert, for example, a text or audio alert if confidential information is included.
[1479] Step 10:
[1480] The server sends a sensitive alert to the device, which then notifies the user of the received alert and warns them by displaying or playing a sound.
[1481] Step 11:
[1482] To help facilitate the progress of the meeting, the server analyzes the conversation and generates suggestions and instructions for progress, such as "Shall we move on to the next topic?"
[1483] Step 12:
[1484] The server sends the generated suggestions and instructions to the terminal, and the terminal notifies the user, allowing the user to smoothly proceed with the conference in accordance with the suggestions and instructions.
[1485] ---
[1486] The above are the specific processing steps in the system of the present invention, which realize natural translation communication and eliminate language barriers.
[1487] Example 2
[1488] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1489] Conventional translation systems are capable of converting speech to text and translating it into another language, but they do not take the user's emotional state into account when outputting the speech. As a result, the translated speech may sound unnatural or may not properly convey the user's emotions. Furthermore, they lack the functionality to provide real-time warnings for sensitive content or to support the progress of meetings. This poses a challenge, making smooth communication difficult in international and business settings.
[1490] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing an emotional state from text data, a means for translating the converted text and emotional state data into a specified language in real time, and a means for converting the translated text and the recognized emotional state into voice data that retains the characteristics of the original voice. This enables natural voice translation that reflects the user's emotional state. Furthermore, the system also includes a function for detecting sensitive content in real time and issuing a warning, as well as a support function for facilitating the progress of meetings, thereby realizing safer and more efficient communication.
[1491] "User" refers to an individual or organization that uses the system.
[1492] "Terminal" refers to electronic devices used by users, such as computers, smartphones, and tablets.
[1493] A "server" is a computer system that processes and stores data, and refers to a device that communicates with terminals over a network.
[1494] "Audio digital signal" refers to audio captured by an audio input device such as a microphone and converted into digital data.
[1495] "Text string" refers to data consisting of characters such as alphabets and kanji characters converted from speech using speech recognition technology.
[1496] "Emotional state" refers to the speaker's emotion that can be read from speech or text, such as positive, negative, or neutral.
[1497] "Real-time translation" refers to the process by which input data (voice or text) is translated into another language within a short time of being received.
[1498] "Voice features" refer to the characteristics of a user's voice and speaking style, and refer to elements that reproduce the way the user is speaking.
[1499] "Audio Data" means data for storing, processing, and playing audio signals in digital form.
[1500] A "sensitive alert" is a message that warns users when content is deemed to contain inappropriate or sensitive content.
[1501] "Suggestions and instructions to support the progress of meetings" refers to organizational suggestions and instructions that are generated to ensure that meetings and conversations proceed smoothly.
[1502] The present invention relates to a system that translates a user's voice in real time and enables natural international communication by combining it with an emotion engine. First, the overall configuration of this system will be described.
[1503] Overall system configuration
[1504] This system consists of a user, a terminal, a server, and an emotion engine. The user speaks into the terminal, and the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates it into the specified language. The translated text is then converted back into voice data that retains the user's voice characteristics and emotional state and sent to the terminal. The emotion engine is responsible for recognizing the user's emotional state from their speech and reflecting it in the translation and speech synthesis.
[1505] Hardware and software used
[1506] Specific hardware includes the smartphone, tablet, or computer used by the user, while the following APIs and services are used as software:
[1507] Speech Recognition API: Google Cloud Speech-to-Text API
[1508] Emotion Recognition Engine: Azure Cognitive Services
[1509] Translation API: Google Translate API
[1510] Speech synthesis technology: Amazon Polly
[1511] Operation flow
[1512] The specific operation flow of the system will be explained below.
[1513] User voice input
[1514] The user speaks into the device's microphone. For example, the user says "Good morning" in Japanese. The device's microphone captures this speech and converts it into a digital signal.
[1515] Sending voice data to the server
[1516] The device transmits the captured audio data to a server via the Internet.
[1517] Speech recognition and text conversion
[1518] The server uses the Google Cloud Speech-to-Text API to convert the audio data into a string of text. For example, "Good morning" is converted into "Good morning."
[1519] Recognition of emotional states
[1520] The server uses Azure Cognitive Services to recognize emotional states based on text data. For example, if "Good morning" is spoken in a cheerful tone, the emotion engine will recognize it as "positive."
[1521] Real-time text translation
[1522] The server translates this text and emotion information into the specified language (e.g., English) using the Google Translate API, for example, "Ohayo gozaimasu" is translated to "Good morning."
[1523] Text-to-speech translation
[1524] The server uses Amazon Polly to convert the translated text into speech data that preserves the characteristics and emotion of the original voice, for example, "Good morning" in the user's voice.
[1525] Sending and outputting synthesized speech
[1526] The server then sends the generated voice data to the device, which then decodes it and plays it back through the speaker, allowing the user to hear the translated voice with their own voice and emotions preserved.
[1527] Specific examples
[1528] For example, when a Japanese-speaking user (A) converses with an English-speaking user (B), the following specific operations take place: When user (A) says "Good morning," the device converts this speech into a digital signal and sends it to the server. The server converts the speech into text "Good morning" and recognizes it as "positive" using an emotion engine. Next, it translates this into English "Good morning," and generates speech data by adding user (A)'s speech features and emotional characteristics based on the translated text. The speech data sent from the server to the device is played back on the device, and user (B) can hear "Good morning" in user (A)'s voice in a cheerful tone.
[1529] Prompt Sentence Examples
[1530] Here are some example prompts to input to the generative AI model:
[1531] When user (A) says "Good morning" in a cheerful voice, the system converts this speech into text and recognizes positive emotions. The server then translates this text into "Good morning" and generates a cheerful voice data in user (A)'s voice. Finally, user (B) can hear "Good morning" in user (A)'s voice.
[1532] The above is a specific embodiment for carrying out the present invention. This system enables natural speech translation that reflects the user's emotions, making international communication smoother.
[1533] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1534] Step 1:
[1535] User voice input
[1536] The user speaks into the microphone of the terminal. The user's voice signal is input to the microphone. In concrete terms, the microphone converts this voice signal into digital voice data. For example, when the user says "Good morning," this voice data is input to the terminal.
[1537] Step 2:
[1538] Sending voice data to the server
[1539] The device sends the captured audio data to the server. As input, the digital audio data is used. As data processing, the device converts the audio data into packets that are sent to the server via the network. As output, packets containing the audio data are sent to the server. The specific operation is that the device transfers the audio data to the server via an Internet connection.
[1540] Step 3:
[1541] Speech recognition and text conversion
[1542] The server converts the received voice data into text. The digital voice data sent to the server is used as input. The Google Cloud Speech-to-Text API is used for data calculations. The server sends the voice data to the API and receives text data as output. Specifically, the server sends a request to the speech recognition API and receives the text data "Good morning" as a response.
[1543] Step 4:
[1544] Recognition of emotional states
[1545] The server recognizes the emotional state from the converted text data. The text data is used as input. Azure Cognitive Services is used for data calculations. The server sends the text to the emotion engine and receives data indicating the emotional state as output. Specifically, "Good morning" is analyzed as having a cheerful tone, and the emotional data "positive" is obtained.
[1546] Step 5:
[1547] Real-time text translation
[1548] The server translates the text and emotion data into the specified language. The text data and emotion data are used as input. The Google Translate API is used for data calculation. The server sends the text data and emotion data to the API and receives the translated text data as output. Specifically, "Ohayou gozaimasu" is translated into "Good morning," and the corresponding positive emotion data is also obtained.
[1549] Step 6:
[1550] Text-to-speech translation
[1551] The server converts the translated text data into speech data. The inputs are the translated text, emotion data, and the original speech feature data. Amazon Polly is used for data processing. The server sends this data to a speech synthesis engine, and the output is speech data with the user's speech features. Specifically, it generates speech data with a positive tone based on the text "Good morning."
[1552] Step 7:
[1553] Sending and outputting synthesized speech
[1554] The server sends the generated voice data to the terminal. The voice data generated by the server is used as input. The terminal receives the data and prepares to play it. The terminal decodes the received voice data and plays it from the speaker. The specific operation is that the terminal plays the voice data "Good morning" so that the user can hear the voice.
[1555] (Application example 2)
[1556] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1557] With modern globalization, there is an increasing need to communicate naturally with users who speak different languages. However, language barriers and misunderstandings of emotions are common, which can lead to poor service quality, especially in brick-and-mortar stores and service industries. Furthermore, many existing translation systems ignore emotional states, resulting in inaccurate communication of user intent. Furthermore, they lack the ability to automatically detect and warn against sensitive topics and language.
[1558] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting a user's voice into a digital signal, means for transmitting the digital signal to the server, means for converting the voice digital signal into character string text on the server, means for translating the converted text into a specified language in real time, means for converting the translated text into voice data that preserves the original voice characteristics and emotional state, means for transmitting the voice data to a terminal, means for playing the voice data on the terminal, and means for supporting multilingual communication between users and staff in a physical store. This enables users and staff who speak different languages to communicate naturally while preserving voice characteristics and emotions, improving the quality of service. Furthermore, the safety of conversations is ensured by detecting sensitive content and issuing alerts at appropriate times.
[1559] The "means for converting a user's voice into a digital signal" refers to a device or system that converts the voice emitted by the user from analog to digital.
[1560] The "means for transmitting a digital signal to a server" refers to a device or system for transmitting the converted digital signal data to a server via a network.
[1561] The "means for converting the digital audio signal into a character string text on the server" refers to software or hardware on the server side that executes the process of converting the transmitted digital audio signal into a character string.
[1562] The "means for translating converted text into a specified language in real time" is software or a system that translates a string of text into another specified language in real time.
[1563] "Means for converting translated text into speech data that retains the original speech characteristics and emotional state" refers to a device or system that converts translated text into speech data that reflects the user's speech characteristics and emotions.
[1564] The "means for transmitting voice data to a terminal" refers to a device or system that transmits the generated voice data to a user's terminal via a network.
[1565] "Means for reproducing audio data on a terminal" refers to a device or system for reproducing received audio data on a speaker or the like of the terminal.
[1566] "Means for supporting multilingual communication between users and staff in a physical store" refers to devices or systems that support natural communication between users and staff who speak different languages in a physical store environment.
[1567] "Sensitive content detection means" means software or systems that detect sensitive content in speech or text in real time.
[1568] "Means for generating a sensitive alert and notifying the terminal" refers to software or a system that generates an alert and notifies the user's terminal when sensitive content is detected.
[1569] The "means for generating suggestions and instructions to support the progress of a meeting" refers to software or a system that generates appropriate suggestions and instructions based on the progress of the meeting.
[1570] The "means for transmitting suggestions and instructions to a terminal and notifying the user" refers to software or a system that transmits the generated suggestions and instructions to the user's terminal and notifies the user.
[1571] The present invention is a system that translates a user's voice in real time and combines it with an emotion engine to enable natural international communication. Specific embodiments of the present invention will be described below.
[1572] Overall system configuration
[1573] This system consists of a user, a terminal, a server, and an emotion engine. When the user speaks into the terminal, the terminal converts the speech into a digital signal and sends it to the server. The server converts the received digital signal into text and translates the text into the specified language. The translated text is converted into audio data that also retains the emotional state and is sent to the terminal. The user can listen to the translated speech through the terminal's speaker.
[1574] Specific Embodiments
[1575] 1. Audio input and digital signal conversion
[1576] Voice input is captured when a user speaks into the device's microphone. For example, a user might say, "Welcome, what product are you looking for?" in a physical store. This speech is converted into digital signals using voice recognition software within the device.
[1577] 2. Digital signal transmission to server
[1578] The converted digital signal is then sent to a server via a network, using a smartphone or tablet as the main hardware and a data transfer protocol as the software.
[1579] 3. Text Conversion
[1580] The digital signal that reaches the server is converted into text using a speech recognition API (e.g., Google Cloud Speech-to-Text API) on the server, for example, "Welcome, what product are you looking for?"
[1581] 4. Recognizing emotional states
[1582] The textual content is then subjected to emotional recognition using an emotion engine (e.g., Emotion Recognition API). For example, the text is classified into an emotional state such as "welcome."
[1583] 5. Real-time translation
[1584] The server translates the recognized text and emotional state into a specified language (e.g., English) in real time using a translation API (e.g., Google Translate API). In this case, the translation is "Welcome, what kind of product are you looking for?"
[1585] 6. Speech synthesis of translated text
[1586] Based on the translated text and the recognized emotional state, the server generates voice data using speech synthesis technology (e.g., Text-to-Speech API). The voice is generated with a tone based on the emotional state (welcome).
[1587] 7. Sending and outputting synthesized speech
[1588] The server sends the generated voice data to the device, which then plays the received voice data through the speaker, allowing the user to listen to the translated voice in real time through the device.
[1589] Specific use cases
[1590] For example, consider a conversation between a Japanese-speaking staff member and an English-speaking customer in a physical store. When the staff member says in Japanese, "Welcome, what kind of product are you looking for?", the system translates this into English and tells the customer, "Welcome, what kind of product are you looking for?" in a voice that retains the emotional state of the customer. This allows for smooth communication that transcends language barriers.
[1591] Prompt Sentence Examples
[1592] "We are developing a multilingual customer service assistant app that allows users to communicate naturally across language barriers and maintains their emotional state. This app translates the user's voice in real time and reflects their emotional state using an emotion engine. The process involves the following steps: 1. Record the user's voice and convert it into a digital signal. 2. Convert the voice to text (e.g., 'Welcome, what kind of product are you looking for?' -> 'Welcome, what kind of product are you looking for?'). 3. Recognize the emotional state of the text (e.g., 'Welcome'). 4. Translate while maintaining the emotional state and perform speech synthesis. 5. Output the final voice data and provide it to the user. The generated program is as follows:"
[1593] In this way, the present invention enables users and staff who speak different languages to communicate naturally while maintaining their emotional state.
[1594] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1595] Step 1:
[1596] The user speaks into the device's microphone. The input is the user's voice, which is captured by the voice input software and converted into a digital signal. The data is then processed by sampling the voice signal and converting it into digital data. The output is a digital voice signal.
[1597] Step 2:
[1598] The terminal sends the converted digital voice signal to the server. The input is the digital voice signal generated in step 1, which is sent to the server via the Internet using the terminal's communication module. The data is processed in the process of packetizing the digital signal and sending it. The output is the digital voice signal that has arrived at the server.
[1599] Step 3:
[1600] The server receives the digital voice signal and converts it into text using a speech recognition API. The input is the digital voice signal, which is converted into text by a speech recognition engine (e.g., Google Cloud Speech-to-Text API). The data operations are speech signal processing and natural language processing. The output is text data generated from the voice signal.
[1601] Step 4:
[1602] The server sends text data to an emotion recognition engine to recognize the emotional state. The input is text data, which is analyzed for emotion by an emotion recognition engine (e.g., Emotion Recognition API). The data is processed by string analysis and emotion categorization. The output is the recognized emotional state.
[1603] Step 5:
[1604] The server sends the text data and the recognized emotional state to a translation API, which translates it into the specified language in real time. The input is the text data and the emotional state, and the translation is performed using a translation API (e.g., Google Translate API). The data processing is the language conversion of the text. The output is the translated text data.
[1605] Step 6:
[1606] The server uses the translated text data and emotional state to synthesize speech. The input is the translated text data and emotional state, which is converted into speech data by a speech synthesis engine (e.g., Text-to-Speech API). The data is processed to convert text to speech and set the emotional tone. The output is the synthesized speech data.
[1607] Step 7:
[1608] The server sends the synthesized voice data to the terminal. The input is the synthesized voice data, which is sent to the terminal via the Internet using a communication module. The data is processed by sending data packets. The output is the voice data that has arrived at the terminal.
[1609] Step 8:
[1610] The device decodes the audio data received and plays it from the speaker. The input is audio data, which is decoded using the device's audio decoder and output as audio from the speaker. The data processing involves converting the digital audio signal to analog. The output is synthesized speech that the user can hear.
[1611] This allows users to communicate naturally with other users who speak different languages.
[1612] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1613] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1614] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1615] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1616] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1617] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1618] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1619] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1620] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1621] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1622] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1623] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1624] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1625] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1626] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1627] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1628] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1629] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1630] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1631] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1632] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1633] The following is further disclosed regarding the above embodiment.
[1634] Understood. Below are proposed draft claims.
[1635] ---
[1636] (Claim 1)
[1637] means for converting a user's voice into a digital signal;
[1638] means for transmitting said digital signal to a server;
[1639] means for converting the digital audio signal into text on a server;
[1640] means for translating the converted text into a specified language in real time;
[1641] a means for converting the translated text into speech data that preserves the characteristics of the original speech;
[1642] means for transmitting the voice data to a terminal;
[1643] means for playing back the audio data on a terminal;
[1644] A system including:
[1645] (Claim 2)
[1646] a means for detecting sensitive content;
[1647] A means for generating a sensitive alert and notifying the terminal;
[1648] 10. The system of claim 1, comprising:
[1649] (Claim 3)
[1650] A means of generating suggestions and instructions to help facilitate the meeting;
[1651] means for transmitting the suggestions or instructions to a terminal and notifying the user;
[1652] 10. The system of claim 1, comprising:
[1653] ---
[1654] "Example 1"
[1655] Understood. Below are the patent claims that include the characteristic parts of the system.
[1656] ---
[1657] (Claim 1)
[1658] means for converting a user's voice into a digital signal;
[1659] means for transmitting the digital signal to a server via a communications network;
[1660] means for converting the digital audio signal into text on a server;
[1661] means for translating the converted text into a specified language in real time;
[1662] a means for converting the translated text into speech data that preserves the characteristics of the original speech;
[1663] means for transmitting the voice data to a terminal via a communication network;
[1664] means for playing back the audio data on a terminal;
[1665] A system including:
[1666] (Claim 2)
[1667] a means for detecting sensitive content;
[1668] A means for generating a sensitive alert and notifying the terminal;
[1669] 10. The system of claim 1, comprising:
[1670] (Claim 3)
[1671] A means of generating suggestions and instructions to help facilitate the meeting;
[1672] means for transmitting the suggestions or instructions to a terminal and notifying the user;
[1673] 10. The system of claim 1, comprising:
[1674] ---
[1675] "Application Example 1"
[1676] (Claim 1)
[1677] means for converting a user's voice into a digital signal;
[1678] means for transmitting said digital signal to a server;
[1679] means for converting the digital audio signal into text on a server;
[1680] means for translating the converted text into a specified language in real time;
[1681] a means for converting the translated text into speech data that preserves the characteristics of the original speech;
[1682] means for transmitting the voice data to a terminal;
[1683] means for playing back the audio data on a terminal;
[1684] A means for assisting a user in communicating between different languages within an automated driving vehicle;
[1685] A system including:
[1686] (Claim 2)
[1687] a means for detecting sensitive content;
[1688] A means for generating a sensitive alert and notifying the terminal;
[1689] A means of generating suggestions and instructions to help facilitate the meeting;
[1690] means for transmitting the suggestions or instructions to a terminal and notifying the user;
[1691] 10. The system of claim 1, comprising:
[1692] (Claim 3)
[1693] 10. The system of claim 1, wherein the system is installed on a smartphone, smart glasses, or head-mounted display and includes the capability to translate speech in real time using a generative AI model.
[1694] "Example 2: Combining Emotion Engines"
[1695] (Claim 1)
[1696] means for converting a user's voice into a digital signal;
[1697] means for transmitting said digital signal to a server;
[1698] means for converting the digital audio signal into text on a server;
[1699] means for recognizing an emotional state from text data;
[1700] means for translating the converted text and emotional state data into a specified language in real time;
[1701] A means for converting the translated text and the recognized emotional state into speech data that retains the characteristics of the original speech;
[1702] means for transmitting the voice data to a terminal;
[1703] means for playing back the audio data on a terminal;
[1704] A system including:
[1705] (Claim 2)
[1706] a means for detecting sensitive content;
[1707] A means for generating a sensitive alert and notifying the terminal;
[1708] 10. The system of claim 1, comprising:
[1709] (Claim 3)
[1710] A means of generating suggestions and instructions to help facilitate the meeting;
[1711] means for transmitting the suggestions or instructions to a terminal and notifying the user;
[1712] 10. The system of claim 1, comprising:
[1713] "Applic...
Claims
1. means for converting a user's voice into a digital signal; means for transmitting said digital signal to a server; means for converting the digital audio signal into text on a server; means for translating the converted text into a specified language in real time; a means for converting the translated text into speech data that preserves the characteristics of the original speech; means for transmitting the voice data to a terminal; means for playing back the audio data on a terminal; A system including:
2. a means for detecting sensitive content; A means for generating a sensitive alert and notifying the terminal; The system of claim 1 , comprising:
3. A means of generating suggestions and instructions to help facilitate the meeting; means for transmitting the suggestions or instructions to a terminal and notifying the user; The system of claim 1 , comprising: ---
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A