System
The system addresses communication challenges by converting voice to text, analyzing, and generating appropriate responses, ensuring effective communication even when speaking is difficult, with direct text input options.
Patent Information
- Application Number
- JP2024138772
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
In modern society, situations arise where voice communication is difficult due to travel, public places, or health issues, making smooth communication challenging, especially during online meetings and phone calls.
A system that converts voice data into text data using voice recognition technology, analyzes the text data using natural language processing to generate appropriate responses, displays options for selection, converts the selected response back into voice data, and transmits it to the other party, also allowing direct text input for responses.
Enables effective communication in situations where speaking is difficult, providing quick and appropriate responses through voice synthesis, enhancing flexibility and convenience.
Smart Images

Figure 2026036245000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, situations where it is difficult to speak while traveling or in public places, or when poor health or a chronic illness makes it difficult to speak, are becoming more common. Even under these circumstances, smooth communication must be maintained. While prompt and appropriate responses are also required during online meetings and phone calls, this is extremely difficult to achieve when voice communication is difficult. A system that can solve these problems is needed. [Means for solving the problem]
[0005] The system according to the invention comprises the following means:
[0006] 1. How to acquire voice data: Voice data is acquired by the user speaking into a device such as a smartphone or tablet.
[0007] 2. A means of converting voice data into text data: Converting the acquired voice data into text data makes it easier to analyze.
[0008] 3. Means for analyzing text data and generating appropriate responses: Analyze the converted text data and generate appropriate responses according to the context.
[0009] 4. Means for displaying generated reply options: The generated reply options are displayed to the user.
[0010] 5. Means for converting the selected response into voice data: The response selected by the user is converted into voice data.
[0011] 6. Means for transmitting voice data: Plays voice data and transmits it to the other party.
[0012] Furthermore, by including a means for the user to directly input text and convert the input text into voice data, smooth communication can be carried out even when there is no suitable response available. Furthermore, by presenting the generated responses to the user as multiple options, a quick and appropriate response can be made. By using the above means, a systematic voice communication support system that solves the above-mentioned problems is provided.
[0013] "Audio data" refers to a digital audio signal that is a recording of a speaker's speech.
[0014] "Text data" is voice data converted into character information.
[0015] "Means for acquiring" refers to a device or system for capturing and recording audio data through a microphone.
[0016] "Means of converting" refers to a process that uses voice recognition technology to convert voice data into text information, or voice synthesis technology to convert text data into voice data.
[0017] "Means of analysis" refers to the process of understanding text data using natural language processing technology and determining its intent and background.
[0018] "Means of generation" refers to the technology or algorithms that automatically generate appropriate responses based on analyzed text data.
[0019] The "means for displaying" refers to an interface that displays the generated answer candidates on the screen of the device so that the user can select one.
[0020] "Means of transmission" refers to a device or system that plays the generated voice data through a speaker and transmits it to the other party. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0023] First, the terms used in the following description will be explained.
[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0029] [First embodiment]
[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0042] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0043] System Configuration
[0044] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[0045] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[0046] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0047] Program processing explanation
[0048] Voice input and initial processing
[0049] Users can input voice by speaking into the microphone of their smartphone or tablet. The device captures this voice data and sends it to the server. This includes recording and data transmission functions.
[0050] Analysis of audio data
[0051] The server passes the received voice data to a speech recognition engine, which converts it into text data. It then uses natural language processing (NLP) technology to analyze the text and understand the context. Based on the analysis results, it generates appropriate response candidates.
[0052] Displaying response options and user selection
[0053] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. The selected answer is then sent back to the server via the terminal.
[0054] Text-to-speech and calling
[0055] The server receives the selected text data and converts it into voice data using a speech synthesis engine. This voice data is then sent to the device, which plays it back through the speaker and transmits it to the other party.
[0056] Specific examples
[0057] Example 1: Use in public places
[0058] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates response candidates such as "Progress is going well," "It's behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby communicating the response to the question to the other party.
[0059] Example 2: Use when feeling unwell
[0060] Consider a case where a user is unwell and has difficulty speaking. When an important call comes in, the other party says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. The user selects "I'll reschedule right away, so please wait" from the options presented, and the server converts this into voice data and transmits it to the other party via the device.
[0061] This system functions as a powerful tool for efficient communication even in situations where speech is difficult. In addition, when options are not appropriate, a free text input function is provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[0065] Step 2:
[0066] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[0067] Step 3:
[0068] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[0069] Step 4:
[0070] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[0071] Step 5:
[0072] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[0073] Step 6:
[0074] Based on the analysis results of the NLP engine, the server uses an AI model (e.g., generative AI) to generate optimal response candidates, such as "The next meeting is at 10:00 AM," "The next meeting is at 3:00 PM," and "The next meeting needs to be rescheduled."
[0075] Step 7:
[0076] The server sends the generated answer candidates to the terminal in JSON format.
[0077] Step 8:
[0078] The device displays the received reply candidates to the user using UI elements (e.g., buttons or lists) to allow the user to select one.
[0079] Step 9:
[0080] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[0081] Step 10:
[0082] The device sends the user's selection to the server via an HTTP request or WebSocket.
[0083] Step 11:
[0084] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[0085] Step 12:
[0086] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[0087] Step 13:
[0088] The terminal plays back the received audio data so that the user can hear the content.
[0089] Step 14:
[0090] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[0091] Step 15:
[0092] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[0093] Step 16:
[0094] The terminal transmits the input text data to the server.
[0095] Step 17:
[0096] The server runs the free text through a speech synthesis engine and converts it into voice data.
[0097] Step 18:
[0098] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[0099] Example 1
[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0101] Currently, many people face situations and environments that make it difficult to communicate through voice. For example, it is difficult to communicate verbally in public places or when feeling unwell. There is a need for a system that can solve these problems and enable everyone to communicate easily.
[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0103] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting the selected response into voice data]; and [means for transmitting the voice data]. This provides a system that performs a series of processes from capturing voice data to outputting the response audibly, enabling effective communication even in situations where speaking is difficult.
[0104] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0105] "Text data" is data in the form of a character string generated by analyzing voice data.
[0106] A "server" is a computer system that processes voice and text data over a network.
[0107] A "terminal" is an electronic device that a user uses to input voice and receive information.
[0108] A "voice recognition engine" is a technology or system that converts voice data into text data.
[0109] "Natural language processing technology" is a technology that analyzes text data, understands the context, and generates appropriate responses.
[0110] A "generative AI model" is an artificial intelligence algorithm that generates appropriate responses from text data.
[0111] A "speech synthesis engine" is a technology or system that converts text data into speech data.
[0112] A "user interface" refers to a screen or operating means for displaying information to a user on a terminal and accepting input.
[0113] A "prompt" is an instruction given to a generative AI model that is used to generate appropriate response candidates.
[0114] "Network communication" refers to the Internet and other internal and external data communication means for sending and receiving data.
[0115] "Real-time" refers to a processing method that minimizes delays and provides immediate response.
[0116] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0117] The terminal receives voice data input by the user and transmits it to a server, which then converts the voice data into text data using voice recognition technology. For example, a smartphone, tablet, or dedicated recording device can be used to acquire this voice data. The acquired voice data is then transmitted to the server using a network communication means.
[0118] The server receives the voice data and instantly converts it into text using a speech recognition engine, such as the Google® Speech-to-Text API. It then analyzes the text using natural language processing technology to generate the most appropriate response based on the context. Natural language processing engines such as Spacy and BERT are used for the analysis.
[0119] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. A speech synthesis engine such as Amazon Polly is used for the speech synthesis. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0120] Specific examples
[0121] Example 1: Use in public places
[0122] Suppose a user is participating in an important meeting on the subway and is asked, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server uses a voice recognition engine to convert the voice data into text data and analyzes it. It then generates response candidates such as "Progress is going well" or "We're behind schedule, but we're working on it." The user selects "Progress is going well," and the server converts this into voice data and sends it to the device. Finally, the device plays back this voice data and communicates it to the other party.
[0123] Example 2: Use when feeling unwell
[0124] Suppose a user is feeling unwell and has difficulty speaking, but receives an important call. If the other party says, "I'd like to reschedule our next meeting," the device captures the voice data and sends it to the server. The server converts the voice data into text data and generates appropriate reply candidates. For example, a reply candidate such as "I'll reschedule right away, so please wait." The user selects "I'll reschedule right away, so please wait," and the server converts this into voice data and sends it to the device. The device then plays this voice data and conveys it to the other party.
[0125] Prompt Sentence Examples
[0126] Prompt: Transcribe the following audio data into text and generate an appropriate response.
[0127] Audio data: "I'd like to schedule our next meeting. When would be convenient for you?"
[0128] This system is a powerful tool for efficient communication even in situations where speech is difficult, greatly improving flexibility and convenience. In addition, when the options are not appropriate, a free text input function is provided, allowing users to enter responses in their own words and play them back as audio data.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] A user speaks into the microphone of a smartphone or tablet to input voice. The device captures this voice data in real time. Specifically, the device's recording function is used to convert the voice signal collected from the microphone into digital voice data. The input is the user's voice, and the output is digital voice data.
[0132] Step 2:
[0133] The device sends the captured audio data to the server. Using the network communication function, the data is sent via a communication protocol such as an HTTP request or WebSocket. The input is the captured digital audio data, and the output is a transmission confirmation message to the server.
[0134] Step 3:
[0135] The server passes the received voice data to a voice recognition engine and converts it into text data. This conversion is performed using the Google Speech-to-Text API or similar. The server passes the voice data to the voice recognition engine and receives the returned text data. The input is voice data and the output is text data.
[0136] Step 4:
[0137] The server passes the converted text data to a natural language processing (NLP) engine to analyze the context of the text. NLP techniques such as Spacy and BERT are used for this analysis. The server passes the text data to the engine and receives the analysis results in return. The input is the text data, and the output is the analysis results and contextual information.
[0138] Step 5:
[0139] The server uses the analyzed text data to generate appropriate reply candidates using a generative AI model (e.g., OpenAI (registered trademark) GPT-3 (registered trademark)). The server passes the prompt sentence and the analysis result to the generative AI model and obtains the reply candidates to be generated. The input is the prompt sentence and the analysis result, and the output is multiple reply candidates.
[0140] Step 6:
[0141] The server sends the generated answer candidates to the terminal. Data is sent quickly using a communication protocol. The input is the generated answer candidates, and the output is a transmission confirmation message to the terminal.
[0142] Step 7:
[0143] The device displays the received reply candidates in a user interface. A UI library (e.g., React Native or Flutter®) is used to display the candidate list on a display or touch screen. The input is the generated reply candidates, and the output is the displayed list of options.
[0144] Step 8:
[0145] The user selects the most appropriate answer from the presented answer candidates. The selection is made by the user's touch or click. The input is the displayed list of options, and the output is the user's selection.
[0146] Step 9:
[0147] The terminal sends the user's selection to the server. The selection result is sent using network communication. The input is the user's selection result, and the output is a transmission confirmation message to the server.
[0148] Step 10:
[0149] The server passes the selected text data to a speech synthesis engine (e.g., Amazon Polly) and converts it into speech data. The speech synthesis engine converts the text data into speech waveforms and converts them into a playable format. The input is the selected text data, and the output is speech data.
[0150] Step 11:
[0151] The server sends the generated voice data to the terminal. The data is sent quickly using a communication protocol. The input is the generated voice data, and the output is a transmission confirmation message to the terminal.
[0152] Step 12:
[0153] The device plays the received audio data. The audio is output through the built-in speaker or connected earphones. A multimedia library (e.g., AVFoundation or MediaPlayer) is used for playback. The input is the generated audio data, and the output is the played audio.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] In autonomous vehicles, there is a demand for a system that allows drivers to safely and efficiently operate various functions, make emergency contacts, and use infotainment functions using only voice. Furthermore, there is a lack of a mechanism that allows smooth communication and maximizes the use of autonomous vehicle functions even when the driver cannot speak. Therefore, the present invention aims to solve these problems by providing a system that handles everything from voice data acquisition to response generation and emergency response.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting a selected response into voice data]; [means for transmitting the voice data]; [means for recognizing voice commands in an autonomously controlled vehicle and executing driving behaviors and settings]; [means for detecting an emergency and automatically contacting an emergency contact]; and [means for playing music, reading news, and checking vehicle status by voice as infotainment functions.] This enables the driver to safely and efficiently operate an autonomously driven vehicle via voice input, automatically take countermeasures in an emergency, and further utilize infotainment functions.
[0159] The "means for acquiring voice data" is a function for capturing voice uttered by a user using a microphone of an electronic device and transmitting the voice data to a server.
[0160] The "means for converting voice data into text data" is a function for converting acquired voice data into text data using voice recognition technology.
[0161] The "means for analyzing text data and generating an appropriate response" is a function that analyzes the converted text data using natural language processing technology and generates an appropriate response according to the context.
[0162] The "means for displaying generated reply options" is a function that displays multiple reply options generated by the server to the user on the terminal.
[0163] The "means for converting the selected response into voice data" is a function for converting the text-format response selected by the user into voice data using voice synthesis technology.
[0164] "Means for transmitting voice data" is a function that transmits converted voice data to the other party through a speaker.
[0165] "Means for recognizing voice commands within an automatically controlled vehicle and executing driving behaviors or settings" refers to a function that recognizes voice commands issued by a user within the vehicle and executes driving behaviors or setting changes of the vehicle accordingly.
[0166] "Means of detecting an emergency and automatically contacting emergency contacts" is a function that detects an emergency (e.g., accident or illness) while driving, and automatically contacts registered emergency contacts.
[0167] "Means for playing music, reading news, and checking vehicle status by voice as infotainment functions" refers to a function that uses voice input to execute infotainment functions such as playing music, reading news, and checking vehicle status.
[0168] This invention is a system that supports voice communication and is applied to autonomous vehicles. It is a system that covers everything from voice data acquisition and analysis to response generation and emergency response. The system mainly operates via a terminal and a server.
[0169] The terminal has the role of acquiring voice data emitted by the user. Specifically, it captures the voice data using a microphone built into a smartphone, tablet, or vehicle and transmits it to a server.
[0170] The server performs a series of processes, including the following steps:
[0171] 1. Converting audio data to text data:
[0172] Using voice recognition technology, the server converts the voice data received from the device into text data. To do this, the server uses various voice recognition software (e.g., Google Speech-to-Text API, etc.).
[0173] 2. Text data analysis and response generation:
[0174] Leverage natural language processing techniques to analyze text data and generate appropriate responses based on the context, for example, using generative AI models (e.g., OpenAI's GPT-3) to generate responses based on prompts.
[0175] 3. Displaying generated answer choices:
[0176] The generated multiple reply candidates are sent to the terminal, and the user selects the most appropriate reply on the terminal.
[0177] 4. Converting selected responses to audio data:
[0178] The server converts the selected text response into audio using a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API).
[0179] 5. Transmission of voice data:
[0180] The converted audio data is sent back to the device and played through the speaker.
[0181] Specific examples
[0182] Navigation function in autonomous vehicles
[0183] When a user says, "Tell me the shortest route," the device captures the voice data and sends it to the server. The server converts the voice data into text data and sends the following prompt sentence to the generative AI model:
[0184] Example prompt sentence:
[0185] "What is the shortest route from my current location to my destination?"
[0186] The generated multiple answer candidates (e.g., "Turn right after 3.2 km" or "Route A is the shortest route") are displayed on the user's device. The user selects the appropriate answer (e.g., "Route A is the shortest route"), and the server converts it into voice data. The device plays the converted voice data, guiding the user to the shortest route.
[0187] Emergency response features
[0188] If a user feels unwell while driving and instructs the device to "make an emergency call," the device captures the voice data and sends it to the server. The server converts the voice data into text data and automatically calls the emergency contact. The device also provides voice information about the vehicle's location and the user's condition. For example, information such as "The driver is currently feeling unwell. Their location is..." is automatically sent to the contact.
[0189] This will enable a variety of functions to be realized through voice commands within self-driving vehicles, improving safety and convenience.
[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0191] Step 1:
[0192] Acquiring audio data
[0193] Input: Voice command from the user
[0194] Specific behavior:
[0195] The user issues a voice command in the vehicle, and the device (a smartphone, tablet, or a microphone built into the vehicle) captures the voice data.
[0196] Step 2:
[0197] Sending voice data to the server
[0198] Input: Captured audio data
[0199] Output: Audio data sent to the server
[0200] Specific behavior:
[0201] The device sends the captured audio data to a server using an internet connection.
[0202] Step 3:
[0203] Converting audio data to text data
[0204] Input: Audio data sent to the server
[0205] Output: Converted text data
[0206] Specific behavior:
[0207] Using speech recognition technology (such as Google Speech-to-Text API), the server converts the received voice data into text data. This process uses a speech recognition engine.
[0208] Step 4:
[0209] Text data analysis and response generation
[0210] Input: Converted text data
[0211] Output: Generated answer candidates
[0212] Specific behavior:
[0213] The server analyzes the text data using natural language processing technology to understand its context, and uses a generative AI model (such as OpenAI GPT-3) to generate appropriate prompts and create candidate responses.
[0214] Step 5:
[0215] Server sending of possible replies
[0216] Input: Generated answer candidates
[0217] Output: Device where possible responses are displayed
[0218] Specific behavior:
[0219] The server then sends the generated reply candidates to the terminal.
[0220] Step 6:
[0221] Viewing and selecting possible responses
[0222] Input: Answer candidates sent by the server
[0223] Output: Selected response
[0224] Specific behavior:
[0225] The device displays possible responses to the user, who then selects the response that they think is most appropriate.
[0226] Step 7:
[0227] Converting selected responses to audio data
[0228] Input: Selected response (text format)
[0229] Output: Audio data
[0230] Specific behavior:
[0231] The server converts the selected text response into audio data using speech synthesis technology (e.g., Google Text-to-Speech API).
[0232] Step 8:
[0233] Sending and playing audio data
[0234] Input: Converted audio data
[0235] Output: Audio data transmitted through the speaker
[0236] Specific behavior:
[0237] The server then sends the final generated voice data to the device, which then plays the data through a speaker and transmits it to the other party.
[0238] By following these steps, users will be able to use voice commands to perform various operations within a self-driving vehicle.
[0239] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0240] The present invention relates to a system that assists voice communication, providing an effective solution, particularly for situations or environments where speech is difficult. This system uses smartphones, tablets, and other electronic devices to convert voice data into text data, generate responses based on the context, and output the responses as voice data. Furthermore, by combining it with an emotion engine that recognizes emotions from the user's voice, it is possible to generate more appropriate and human-like responses.
[0241] System Configuration
[0242] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[0243] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[0244] The emotion engine recognizes the user's emotions from the voice data and provides this emotion information to the server, which then adjusts the generated responses based on this emotion information to generate more appropriate response candidates.
[0245] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0246] Program processing explanation
[0247] Voice input and initial processing
[0248] A user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?" The device captures the user's voice and saves the audio data. The device then sends the audio data to the server.
[0249] Voice data analysis and emotion recognition
[0250] The server receives the voice data and converts it into text using a speech recognition engine. The converted text data is then passed to a natural language processing (NLP) engine to analyze the context and intent. At the same time, the server uses an emotion engine to recognize the user's emotions. For example, it detects emotions such as tension in the user's voice.
[0251] Generating responses that take emotions into account
[0252] The server uses an AI model to generate optimal response candidates based on the results of NLP analysis and the recognition results of the emotion engine. By taking into account the information from the emotion engine, more appropriate responses can be tailored. For example, if the user is nervous, a gentle response such as "Relax, the next meeting is at 10:00 AM" will be generated.
[0253] Displaying response options and user selection
[0254] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. For example, "The next meeting is at 10:00 AM."
[0255] Text-to-speech and calling
[0256] The server passes the selected text data to a speech synthesis engine, converts it into voice data, and sends the generated voice data to the device, which plays it back through the speaker and communicates the response to the other party.
[0257] Specific examples
[0258] Example 1: Use in public places
[0259] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates candidate responses. At the same time, if the emotion engine detects tension in the user's voice, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby conveying the answer to the question to the other party.
[0260] Example 2: Use when feeling unwell
[0261] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[0262] This system serves as a powerful tool for efficient communication even in situations where speech is difficult. It also takes into account the user's emotions to provide more appropriate and human-like responses. A free text input function is also provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[0263] The processing flow will be explained below.
[0264] Step 1:
[0265] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[0266] Step 2:
[0267] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[0268] Step 3:
[0269] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[0270] Step 4:
[0271] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[0272] Step 5:
[0273] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[0274] Step 6:
[0275] At the same time, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine detects emotions such as nervousness, fatigue, or happiness from the tone, speed, and volume of the voice.
[0276] Step 7:
[0277] The server uses an AI model (e.g., generative AI) to generate optimal reply candidates based on the analysis results of the NLP engine and the recognition results of the emotion engine. Taking into account the information from the emotion engine, the server adjusts the reply according to the user's current emotional state. For example, if the user is nervous, it generates a reply that will make them feel more relaxed.
[0278] Step 8:
[0279] The server sends the generated answer candidates to the terminal in JSON format.
[0280] Step 9:
[0281] The device displays the received reply candidates to the user. The reply candidates are displayed using UI elements (e.g., buttons and lists) to make it easy for the user to select one.
[0282] Step 10:
[0283] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[0284] Step 11:
[0285] The device sends the user's selection to the server via an HTTP request or WebSocket.
[0286] Step 12:
[0287] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[0288] Step 13:
[0289] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[0290] Step 14:
[0291] The terminal plays back the received audio data so that the user can hear the content.
[0292] Step 15:
[0293] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[0294] Step 16:
[0295] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[0296] Step 17:
[0297] The terminal transmits the input text data to the server.
[0298] Step 18:
[0299] The server runs the free text through a speech synthesis engine and converts it into voice data.
[0300] Step 19:
[0301] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[0302] By combining this system with an emotion engine, it becomes possible to respond in a way that takes into account the user's emotional state, resulting in more natural and friendly communication. Furthermore, by using an emotion engine, responses that suit the user's state of mind can be provided, which is expected to have the effect of improving user satisfaction.
[0303] Example 2
[0304] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0305] Conventional voice communication systems have difficulty communicating effectively in situations or environments where speech is difficult. Furthermore, they lack the functionality to provide responses that take the user's emotions into account, making it difficult to generate more appropriate and human-like responses. Furthermore, the interface for selecting the most appropriate response from multiple candidate responses is inconvenient, which can hinder smooth communication.
[0306] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0307] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and recognizing the user's emotions, and means for generating appropriate reply candidates based on the analysis results and the emotional information. This makes it possible to convert voice data into text data and generate replies that take the user's emotions into consideration.
[0308] "Voice data" refers to acoustic data generated by a user's speech.
[0309] "Acquiring" refers to capturing and recording audio data using the device's microphone.
[0310] A "server" refers to a computer system that receives voice data from a terminal via a network and processes the data.
[0311] "Send" refers to the transfer of collected voice data by the terminal to the server.
[0312] "Text data" refers to data obtained by converting voice data into character information.
[0313] "Converting" refers to the process of changing audio data into text data or text data into audio data.
[0314] "Analyzing" refers to interpreting the meaning, intent, and context of text data using natural language processing techniques.
[0315] "Emotion recognition" refers to detecting a user's emotional state from speech or text data.
[0316] "Generating reply candidates" refers to creating appropriate responses based on text data and emotional information.
[0317] "Presenting to the user as options" refers to displaying the generated reply candidates as multiple options on the user's device.
[0318] "Speech synthesis" refers to the technology of converting text data into voice data.
[0319] "Transmitting audio" refers to playing the generated audio data through the device's speaker.
[0320] This invention is a system that allows users to smoothly communicate through speech, particularly in situations and environments where speech is difficult. The system converts speech into text and then analyzes the text to generate responses. It also recognizes the user's emotional state and reflects it in the responses, enabling more appropriate and human-like communication.
[0321] System Configuration
[0322] This system is realized mainly using the following hardware and software.
[0323] Hardware
[0324] 1. Device: An electronic device such as a smartphone or tablet that includes a microphone for capturing the user's voice and a speaker for playing back responses. Examples of smartphones include the Apple iPhone (registered trademark) and the Samsung Galaxy series.
[0325] 2. Microphone: Reliably captures the user's voice through a high-sensitivity microphone, such as the Shure MV88.
[0326] software
[0327] 1. Speech recognition engine: Software for converting voice data into text data. For example, we use the Google Speech-to-Text API.
[0328] 2. Natural Language Processing (NLP) Engine: Software that analyzes the converted text data to understand intent and context. We use OpenAI GPT-3 for this.
[0329] 3. Emotion Engine: Software for recognizing emotions from user voice and text data. As an example, we use IBM Watson (registered trademark) Tone Analyzer.
[0330] 4. Speech synthesis engine: Software that converts selected text data into speech data. An example is Amazon Polly.
[0331] System Operation Overview
[0332] When a user speaks into the device, the device's microphone captures the voice and saves it as audio data. This audio data is sent to the server using a communication protocol. The server then converts the audio data into text data using the Google Speech-to-Text API. It then uses OpenAI GPT-3 to analyze the context and intent of the text data. At the same time, it uses IBM Watson Tone Analyzer to recognize the user's emotions. Based on the analysis results and emotional information, OpenAI GPT-3 generates appropriate response candidates. These response candidates are sent from the server to the device and displayed to the user as multiple options. The user selects the most appropriate response from the displayed options and sends it to the server. The server then converts the selected text data into audio data using Amazon Polly and sends it back to the device. Finally, the audio data is played back through the device's speaker.
[0333] Specific examples
[0334] Example 1: Use in public places
[0335] Consider a situation where a user is attending an important meeting on the subway. During the meeting, someone asks, "Please tell me about the current progress." The user speaks the question into their smartphone, and the voice data is sent to the server. The server converts the voice into text, analyzes it, and generates candidate responses. If the emotion engine detects that the user is tense, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," and the voice data is played.
[0336] Example 2: Use when feeling unwell
[0337] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice and sends it to the server. The server converts the voice into text and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[0338] Prompt Sentence Examples
[0339] The following is an example of a prompt sentence to input to the generative AI model:
[0340] "When is the next meeting? User sentiment: Nervous Generate possible responses."
[0341] This prompt helps the system achieve smooth communication even in situations where the user has difficulty speaking.
[0342] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0343] Step 1:
[0344] The user speaks into the microphone on their smartphone or tablet.
[0345] The input is the user's speech. The output is the audio data captured by the device. Specifically, the user speaks, "When is the next meeting?" The device's high-sensitivity microphone (e.g., Shure MV88) picks up the voice and temporarily stores the audio data in the device's built-in storage.
[0346] Step 2:
[0347] The device compresses the captured audio data and sends it to the server using the HTTPS protocol.
[0348] The input is the audio data acquired in step 1. The output is compressed audio data sent to the server. Specifically, the device compresses the audio data for efficient transmission and sends it to the server encrypted with SSL / TLS.
[0349] Step 3:
[0350] The server converts the received voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text API).
[0351] The input is the voice data sent from the device. The output is the converted text data. Specifically, the server sends the voice data to the Google Speech-to-Text API, which returns the text data. The converted text data is then sent to the next processing step within the server.
[0352] Step 4:
[0353] The server passes the text data to a natural language processing (NLP) engine (e.g., OpenAI GPT-3) to analyze the context and intent.
[0354] The input is the text data generated in step 3. The output is the analyzed context information. Specifically, the server sends the text data to OpenAI GPT-3, and obtains information about the context and intent as the analysis result. This analysis result is used for subsequent processing.
[0355] Step 5:
[0356] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions.
[0357] The input is the text data and speech parameters generated in step 3. The output is the recognized emotion information. Specifically, the server sends the speech and text parameters to IBM Watson Tone Analyzer to recognize the user's emotion (e.g., tension, fatigue). This information is used to generate a response.
[0358] Step 6:
[0359] The server uses a generative AI model (e.g., OpenAI GPT-3) based on the analysis results and emotional information to generate appropriate response candidates.
[0360] The input is the context analysis result obtained in step 4 and the emotional information obtained in step 5. The output is multiple response candidates. Specifically, the server sends the analysis results and emotional information as prompts to OpenAI GPT-3, which then generates appropriate response candidates.
[0361] Step 7:
[0362] The server sends the generated answer candidates to the terminal in JSON format.
[0363] The input is the candidate answers generated in step 6. The output is the candidate answer data sent to the device. Specifically, the server formats the candidate answers in JSON format and sends them to the device.
[0364] Step 8:
[0365] The terminal displays the reply candidates on a user interface.
[0366] The input is candidate reply data sent from the server. The output is multiple candidate replies displayed on the user interface. Specifically, the device analyzes the candidate replies and displays options such as "Next meeting is at 10:00 AM," "Relax, next meeting is at 10:00 AM," and "Checking next meeting time" on the touchscreen.
[0367] Step 9:
[0368] The user selects the most appropriate option from the displayed options and taps it.
[0369] The input is the reply candidates displayed on the terminal. The output is the selected reply data. In concrete terms, the user selects "Relax, the next meeting is at 10:00 AM," and the selection information is saved on the terminal.
[0370] Step 10:
[0371] The terminal transmits the selected response data to the server.
[0372] The input is the response data selected by the user. The output is the selected data sent to the server. Specifically, the terminal encrypts the selected data again using SSL / TLS and sends it to the server.
[0373] Step 11:
[0374] The server passes the selected response to a speech synthesis engine (e.g., Amazon Polly) and converts it into voice data.
[0375] The input is the selected response data sent from the device. The output is the generated voice data. Specifically, the server sends the selected response to Amazon Polly, which generates the voice data.
[0376] Step 12:
[0377] The server transmits the generated voice data to the terminal.
[0378] The input is the voice data generated in step 11. The output is the voice data sent to the terminal. In concrete operation, the server sends the generated voice data back to the terminal.
[0379] Step 13:
[0380] The terminal plays the received audio data through a speaker.
[0381] The input is the voice data sent from the server. The output is the voice response that is transmitted to the other party. Specifically, the device decodes the voice data and plays it using the built-in speaker (e.g., Bose SoundLink Mini) to transmit the response to the other party.
[0382] (Application example 2)
[0383] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0384] Conventional voice communication assistance systems convert voice data into text data and generate responses, but they are not sufficient in generating appropriate responses that take the user's emotional state into account or in presenting responses quickly using visual devices. Therefore, there is a need for a system that provides more human-like and considerate responses in situations and environments where voice communication is difficult.
[0385] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting voice data to text data, means for analyzing the text data and generating an appropriate response, means for displaying the generated response options, means for converting the selected response into voice data, means for transmitting the voice data, means for analyzing emotional data and adjusting the response, and means for displaying candidate responses on a visual device. This enables the generation of appropriate and prompt responses that take the user's emotional state into consideration.
[0386] "Voice data" refers to data in which the user's speech is recorded as an electronic signal.
[0387] "Text data" is data obtained by converting voice data into a character string format.
[0388] "Analysis" is the process of understanding the content of text data and grasping its context and intent.
[0389] A "reply" is the content of a response generated based on the analysis results.
[0390] "Choices" refer to multiple possible responses presented to the user.
[0391] "Display" refers to showing the generated responses or options on a visual device.
[0392] "Speech synthesis" is a technology that converts text data into voice data.
[0393] "Transmission" is a function for transmitting generated voice data to the outside.
[0394] "Emotion data" is data of emotions estimated from the user's voice, facial expressions, movements, etc.
[0395] A "visual device" is a device for visually presenting information to a user, and in this context refers to smart glasses and the like.
[0396] This invention relates to a system that supports voice communication, particularly to a customer service support application in brick-and-mortar stores using visual devices. This system converts voice data into text data and generates appropriate responses based on emotion data through collaboration between smart glasses and a server. The generated responses are displayed on the smart glasses' display and transmitted as voice data.
[0397] Components and hardware / software used
[0398] 1. Smart Glasses:
[0399] A device worn by the user, equipped with a microphone and a display, that captures voice data and provides a visual response.
[0400] 2. Server:
[0401] It processes voice data, generates text data, analyzes emotional data, generates responses, and synthesizes voice.
[0402] The software used includes a speech recognition engine, a natural language processing engine, a sentiment analysis engine, and a speech synthesis engine.
[0403] Program processing explanation
[0404] 1. Acquire audio data:
[0405] The user speaks a customer question into the microphone of the smart glasses, for example, "What cake do you recommend?"
[0406] The smart glasses capture this voice data and send it to a server.
[0407] 2. Convert to text data:
[0408] The server uses a voice recognition engine to convert the received voice data into text data.
[0409] The converted text data may be a string such as "What cake do you recommend?"
[0410] 3. Emotional Data Analysis:
[0411] The server uses an emotion analysis engine to extract emotion data from the user's speech.
[0412] For example, friendliness is detected from the user's tone of voice.
[0413] 4. Generate a response:
[0414] The natural language processing engine generates optimal responses based on text data and emotional data.
[0415] A possible reply such as "Today's recommendation is shortcake" is generated.
[0416] 5. Visual display of response:
[0417] The generated response is sent from the server to the smart glasses and displayed on the display.
[0418] The user checks the displayed reply candidates and selects an appropriate reply.
[0419] 6. Conversion to voice data and transmission:
[0420] The server passes the selected response to a speech synthesis engine and converts it into voice data.
[0421] The generated voice data is transmitted to the customer through the speaker of the smart glasses.
[0422] Specific examples
[0423] Consider a scenario where a cafe staff member is wearing smart glasses during a busy lunchtime. When a customer asks, "What cake do you recommend?", the staff member captures the voice data through the microphone in the smart glasses. The server analyzes the voice as text data and generates a response that takes emotional data into account. A response such as "Today's recommendation is shortcake" is displayed on the smart glasses' display, and the staff member relays this to the customer.
[0424] Prompt Sentence Examples
[0425] Examples of prompts for a generative AI model include "Emotion-recognizing customer service assistant at a cafe" and "Generate a friendly response when a customer asks, 'What cake do you recommend?'"
[0426] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0427] Step 1:
[0428] The user speaks into the microphone of the smart glasses. For example, they say, "What cake do you recommend?" The smart glasses capture this voice data and store it in the device's memory. At this point, the input is the user's voice data, and the output is the saved audio file.
[0429] Step 2:
[0430] The device sends the acquired voice data to the server. The input is the saved voice file, and the output is the voice data sent to the server. Specifically, the smart glasses upload the voice data to the server.
[0431] Step 3:
[0432] The server runs the received voice data through a voice recognition engine and converts it into text data. At this point, the voice data is output as character string data. Specifically, the text data obtained is "What cake do you recommend?" The input is voice data, and the output is text data.
[0433] Step 4:
[0434] At the same time, the server uses an emotion analysis engine to extract emotional data from the voice data. For example, the emotion "friendliness" can be detected from the tone of the user's voice. In this case, the input is voice data, and the output is data indicating the emotional state. Specifically, the emotion analysis engine performs a process to analyze the voice characteristics.
[0435] Step 5:
[0436] The server passes the text data and emotion data to a natural language processing engine, which then generates the optimal response. For example, the response generated might be, "Today's recommendation is shortcake." The input is text data and emotion data, and the output is the generated response data. Specifically, the generative AI model creates a response based on the prompt sentence.
[0437] Step 6:
[0438] The server prepares the generated responses as multiple options and sends them to the smart glasses. In this case, options such as "Today's recommendation is shortcake" are included. The input is the generated response data, and the output is the response options displayed on the smart glasses. Specifically, the option data is displayed on the smart glasses' display.
[0439] Step 7:
[0440] The user checks the display of the smart glasses and selects the most appropriate response. For example, they might select "Today's recommendation is shortcake." In this case, the input is the response options displayed on the display, and the output is the user's selection. Specifically, the user confirms the selection by touching the display.
[0441] Step 8:
[0442] The server passes the selected response to the speech synthesis engine and converts it into speech data. At this time, the text data is output as a speech file. Speech data such as "Today's recommendation is shortcake" is generated. The input is the selected response data, and the output is speech data. Specifically, the speech synthesis engine performs the process of converting text into speech.
[0443] Step 9:
[0444] The terminal transmits the generated voice data through the speaker of the smart glasses. The customer is provided with the reply, "Today's recommendation is shortcake." The input is the generated voice data, and the output is the voice that the customer hears through the speaker. The specific operation is that the voice is played back from the speaker of the smart glasses.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0446] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0448] [Second embodiment]
[0449] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0450] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0451] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0453] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0455] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0456] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0457] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0460] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0461] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0462] System Configuration
[0463] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[0464] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[0465] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0466] Program processing explanation
[0467] Voice input and initial processing
[0468] Users can input voice by speaking into the microphone of their smartphone or tablet. The device captures this voice data and sends it to the server. This includes recording and data transmission functions.
[0469] Analysis of audio data
[0470] The server passes the received voice data to a speech recognition engine, which converts it into text data. It then uses natural language processing (NLP) technology to analyze the text and understand the context. Based on the analysis results, it generates appropriate response candidates.
[0471] Displaying response options and user selection
[0472] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. The selected answer is then sent back to the server via the terminal.
[0473] Text-to-speech and calling
[0474] The server receives the selected text data and converts it into voice data using a speech synthesis engine. This voice data is then sent to the device, which plays it back through the speaker and transmits it to the other party.
[0475] Specific examples
[0476] Example 1: Use in public places
[0477] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates response candidates such as "Progress is going well," "It's behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby communicating the response to the question to the other party.
[0478] Example 2: Use when feeling unwell
[0479] Consider a case where a user is unwell and has difficulty speaking. When an important call comes in, the other party says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. The user selects "I'll reschedule right away, so please wait" from the options presented, and the server converts this into voice data and transmits it to the other party via the device.
[0480] This system functions as a powerful tool for efficient communication even in situations where speech is difficult. In addition, when options are not appropriate, a free text input function is provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[0481] The processing flow will be explained below.
[0482] Step 1:
[0483] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[0484] Step 2:
[0485] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[0486] Step 3:
[0487] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[0488] Step 4:
[0489] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[0490] Step 5:
[0491] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[0492] Step 6:
[0493] Based on the analysis results of the NLP engine, the server uses an AI model (e.g., generative AI) to generate optimal response candidates, such as "The next meeting is at 10:00 AM," "The next meeting is at 3:00 PM," and "The next meeting needs to be rescheduled."
[0494] Step 7:
[0495] The server sends the generated answer candidates to the terminal in JSON format.
[0496] Step 8:
[0497] The device displays the received reply candidates to the user using UI elements (e.g., buttons or lists) to allow the user to select one.
[0498] Step 9:
[0499] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[0500] Step 10:
[0501] The device sends the user's selection to the server via an HTTP request or WebSocket.
[0502] Step 11:
[0503] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[0504] Step 12:
[0505] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[0506] Step 13:
[0507] The terminal plays back the received audio data so that the user can hear the content.
[0508] Step 14:
[0509] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[0510] Step 15:
[0511] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[0512] Step 16:
[0513] The terminal transmits the input text data to the server.
[0514] Step 17:
[0515] The server runs the free text through a speech synthesis engine and converts it into voice data.
[0516] Step 18:
[0517] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[0518] Example 1
[0519] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0520] Currently, many people face situations and environments that make it difficult to communicate through voice. For example, it is difficult to communicate verbally in public places or when feeling unwell. There is a need for a system that can solve these problems and enable everyone to communicate easily.
[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0522] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting the selected response into voice data]; and [means for transmitting the voice data]. This provides a system that performs a series of processes from capturing voice data to outputting the response audibly, enabling effective communication even in situations where speaking is difficult.
[0523] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0524] "Text data" is data in the form of a character string generated by analyzing voice data.
[0525] A "server" is a computer system that processes voice and text data over a network.
[0526] A "terminal" is an electronic device that a user uses to input voice and receive information.
[0527] A "voice recognition engine" is a technology or system that converts voice data into text data.
[0528] "Natural language processing technology" is a technology that analyzes text data, understands the context, and generates appropriate responses.
[0529] A "generative AI model" is an artificial intelligence algorithm that generates appropriate responses from text data.
[0530] A "speech synthesis engine" is a technology or system that converts text data into speech data.
[0531] A "user interface" refers to a screen or operating means for displaying information to a user on a terminal and accepting input.
[0532] A "prompt" is an instruction given to a generative AI model that is used to generate appropriate response candidates.
[0533] "Network communication" refers to the Internet and other internal and external data communication means for sending and receiving data.
[0534] "Real-time" refers to a processing method that minimizes delays and provides immediate response.
[0535] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0536] The terminal receives voice data input by the user and transmits it to a server, which then converts the voice data into text data using voice recognition technology. For example, a smartphone, tablet, or dedicated recording device can be used to acquire this voice data. The acquired voice data is then transmitted to the server using a network communication means.
[0537] The server receives the voice data and instantly converts it into text using a speech recognition engine such as the Google Speech-to-Text API. It then analyzes the text using natural language processing technology to generate the most appropriate response based on the context. Natural language processing engines such as Spacy and BERT are used for the analysis.
[0538] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. A speech synthesis engine such as Amazon Polly is used for the speech synthesis. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0539] Specific examples
[0540] Example 1: Use in public places
[0541] Suppose a user is participating in an important meeting on the subway and is asked, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server uses a voice recognition engine to convert the voice data into text data and analyzes it. It then generates response candidates such as "Progress is going well" or "We're behind schedule, but we're working on it." The user selects "Progress is going well," and the server converts this into voice data and sends it to the device. Finally, the device plays back this voice data and communicates it to the other party.
[0542] Example 2: Use when feeling unwell
[0543] Suppose a user is feeling unwell and has difficulty speaking, but receives an important call. If the other party says, "I'd like to reschedule our next meeting," the device captures the voice data and sends it to the server. The server converts the voice data into text data and generates appropriate reply candidates. For example, a reply candidate such as "I'll reschedule right away, so please wait." The user selects "I'll reschedule right away, so please wait," and the server converts this into voice data and sends it to the device. The device then plays this voice data and conveys it to the other party.
[0544] Prompt Sentence Examples
[0545] Prompt: Transcribe the following audio data into text and generate an appropriate response.
[0546] Audio data: "I'd like to schedule our next meeting. When would be convenient for you?"
[0547] This system is a powerful tool for efficient communication even in situations where speech is difficult, greatly improving flexibility and convenience. In addition, when the options are not appropriate, a free text input function is provided, allowing users to enter responses in their own words and play them back as audio data.
[0548] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0549] Step 1:
[0550] A user speaks into the microphone of a smartphone or tablet to input voice. The device captures this voice data in real time. Specifically, the device's recording function is used to convert the voice signal collected from the microphone into digital voice data. The input is the user's voice, and the output is digital voice data.
[0551] Step 2:
[0552] The device sends the captured audio data to the server. Using the network communication function, the data is sent via a communication protocol such as an HTTP request or WebSocket. The input is the captured digital audio data, and the output is a transmission confirmation message to the server.
[0553] Step 3:
[0554] The server passes the received voice data to a voice recognition engine and converts it into text data. This conversion is performed using the Google Speech-to-Text API or similar. The server passes the voice data to the voice recognition engine and receives the returned text data. The input is voice data and the output is text data.
[0555] Step 4:
[0556] The server passes the converted text data to a natural language processing (NLP) engine to analyze the context of the text. NLP techniques such as Spacy and BERT are used for this analysis. The server passes the text data to the engine and receives the analysis results in return. The input is the text data, and the output is the analysis results and contextual information.
[0557] Step 5:
[0558] The server uses the parsed text data to generate appropriate reply candidates using a generative AI model (e.g., OpenAI GPT-3). The server passes the prompt sentence and the analysis result to the generative AI model and obtains the reply candidates that are generated. The input is the prompt sentence and the analysis result, and the output is multiple reply candidates.
[0559] Step 6:
[0560] The server sends the generated answer candidates to the terminal. Data is sent quickly using a communication protocol. The input is the generated answer candidates, and the output is a transmission confirmation message to the terminal.
[0561] Step 7:
[0562] The device displays the received reply candidates in a user interface. It uses a UI library (e.g., React Native or Flutter) to display the candidate list on a display or touchscreen. The input is the generated reply candidates, and the output is the displayed list of options.
[0563] Step 8:
[0564] The user selects the most appropriate answer from the presented answer candidates. The selection is made by the user's touch or click. The input is the displayed list of options, and the output is the user's selection.
[0565] Step 9:
[0566] The terminal sends the user's selection to the server. The selection result is sent using network communication. The input is the user's selection result, and the output is a transmission confirmation message to the server.
[0567] Step 10:
[0568] The server passes the selected text data to a speech synthesis engine (e.g., Amazon Polly) and converts it into speech data. The speech synthesis engine converts the text data into speech waveforms and converts them into a playable format. The input is the selected text data, and the output is speech data.
[0569] Step 11:
[0570] The server sends the generated voice data to the terminal. The data is sent quickly using a communication protocol. The input is the generated voice data, and the output is a transmission confirmation message to the terminal.
[0571] Step 12:
[0572] The device plays the received audio data. The audio is output through the built-in speaker or connected earphones. A multimedia library (e.g., AVFoundation or MediaPlayer) is used for playback. The input is the generated audio data, and the output is the played audio.
[0573] (Application example 1)
[0574] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0575] In autonomous vehicles, there is a demand for a system that allows drivers to safely and efficiently operate various functions, make emergency contacts, and use infotainment functions using only voice. Furthermore, there is a lack of a mechanism that allows smooth communication and maximizes the use of autonomous vehicle functions even when the driver cannot speak. Therefore, the present invention aims to solve these problems by providing a system that handles everything from voice data acquisition to response generation and emergency response.
[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0577] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting a selected response into voice data]; [means for transmitting the voice data]; [means for recognizing voice commands in an autonomously controlled vehicle and executing driving behaviors and settings]; [means for detecting an emergency and automatically contacting an emergency contact]; and [means for playing music, reading news, and checking vehicle status by voice as infotainment functions.] This enables the driver to safely and efficiently operate an autonomously driven vehicle via voice input, automatically take countermeasures in an emergency, and further utilize infotainment functions.
[0578] The "means for acquiring voice data" is a function for capturing voice uttered by a user using a microphone of an electronic device and transmitting the voice data to a server.
[0579] The "means for converting voice data into text data" is a function for converting acquired voice data into text data using voice recognition technology.
[0580] The "means for analyzing text data and generating an appropriate response" is a function that analyzes the converted text data using natural language processing technology and generates an appropriate response according to the context.
[0581] The "means for displaying generated reply options" is a function that displays multiple reply options generated by the server to the user on the terminal.
[0582] The "means for converting the selected response into voice data" is a function for converting the text-format response selected by the user into voice data using voice synthesis technology.
[0583] "Means for transmitting voice data" is a function that transmits converted voice data to the other party through a speaker.
[0584] "Means for recognizing voice commands within an automatically controlled vehicle and executing driving behaviors or settings" refers to a function that recognizes voice commands issued by a user within the vehicle and executes driving behaviors or setting changes of the vehicle accordingly.
[0585] "Means of detecting an emergency and automatically contacting emergency contacts" is a function that detects an emergency (e.g., accident or illness) while driving, and automatically contacts registered emergency contacts.
[0586] "Means for playing music, reading news, and checking vehicle status by voice as infotainment functions" refers to a function that uses voice input to execute infotainment functions such as playing music, reading news, and checking vehicle status.
[0587] This invention is a system that supports voice communication and is applied to autonomous vehicles. It is a system that covers everything from voice data acquisition and analysis to response generation and emergency response. The system mainly operates via a terminal and a server.
[0588] The terminal has the role of acquiring voice data emitted by the user. Specifically, it captures the voice data using a microphone built into a smartphone, tablet, or vehicle and transmits it to a server.
[0589] The server performs a series of processes, including the following steps:
[0590] 1. Converting audio data to text data:
[0591] Using voice recognition technology, the server converts the voice data received from the device into text data. To do this, the server uses various voice recognition software (e.g., Google Speech-to-Text API, etc.).
[0592] 2. Text data analysis and response generation:
[0593] Leverage natural language processing techniques to analyze text data and generate appropriate responses based on the context, for example, using generative AI models (e.g., OpenAI's GPT-3) to generate responses based on prompts.
[0594] 3. Displaying generated answer choices:
[0595] The generated multiple reply candidates are sent to the terminal, and the user selects the most appropriate reply on the terminal.
[0596] 4. Converting selected responses to audio data:
[0597] The server converts the selected text response into audio using a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API).
[0598] 5. Transmission of voice data:
[0599] The converted audio data is sent back to the device and played through the speaker.
[0600] Specific examples
[0601] Navigation function in autonomous vehicles
[0602] When a user says, "Tell me the shortest route," the device captures the voice data and sends it to the server. The server converts the voice data into text data and sends the following prompt sentence to the generative AI model:
[0603] Example prompt sentence:
[0604] "What is the shortest route from my current location to my destination?"
[0605] The generated multiple answer candidates (e.g., "Turn right after 3.2 km" or "Route A is the shortest route") are displayed on the user's device. The user selects the appropriate answer (e.g., "Route A is the shortest route"), and the server converts it into voice data. The device plays the converted voice data, guiding the user to the shortest route.
[0606] Emergency response features
[0607] If a user feels unwell while driving and instructs the device to "make an emergency call," the device captures the voice data and sends it to the server. The server converts the voice data into text data and automatically calls the emergency contact. The device also provides voice information about the vehicle's location and the user's condition. For example, information such as "The driver is currently feeling unwell. Their location is..." is automatically sent to the contact.
[0608] This will enable a variety of functions to be realized through voice commands within self-driving vehicles, improving safety and convenience.
[0609] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0610] Step 1:
[0611] Acquiring audio data
[0612] Input: Voice command from the user
[0613] Specific behavior:
[0614] The user issues a voice command in the vehicle, and the device (a smartphone, tablet, or a microphone built into the vehicle) captures the voice data.
[0615] Step 2:
[0616] Sending voice data to the server
[0617] Input: Captured audio data
[0618] Output: Audio data sent to the server
[0619] Specific behavior:
[0620] The device sends the captured audio data to a server using an internet connection.
[0621] Step 3:
[0622] Converting audio data to text data
[0623] Input: Audio data sent to the server
[0624] Output: Converted text data
[0625] Specific behavior:
[0626] Using speech recognition technology (such as Google Speech-to-Text API), the server converts the received voice data into text data. This process uses a speech recognition engine.
[0627] Step 4:
[0628] Text data analysis and response generation
[0629] Input: Converted text data
[0630] Output: Generated answer candidates
[0631] Specific behavior:
[0632] The server analyzes the text data using natural language processing technology to understand its context, and uses a generative AI model (such as OpenAI GPT-3) to generate appropriate prompts and create candidate responses.
[0633] Step 5:
[0634] Server sending of possible replies
[0635] Input: Generated answer candidates
[0636] Output: Device where possible responses are displayed
[0637] Specific behavior:
[0638] The server then sends the generated reply candidates to the terminal.
[0639] Step 6:
[0640] Viewing and selecting possible responses
[0641] Input: Answer candidates sent by the server
[0642] Output: Selected response
[0643] Specific behavior:
[0644] The device displays possible responses to the user, who then selects the response that they think is most appropriate.
[0645] Step 7:
[0646] Converting selected responses to audio data
[0647] Input: Selected response (text format)
[0648] Output: Audio data
[0649] Specific behavior:
[0650] The server converts the selected text response into audio data using speech synthesis technology (e.g., Google Text-to-Speech API).
[0651] Step 8:
[0652] Sending and playing audio data
[0653] Input: Converted audio data
[0654] Output: Audio data transmitted through the speaker
[0655] Specific behavior:
[0656] The server then sends the final generated voice data to the device, which then plays the data through a speaker and transmits it to the other party.
[0657] By following these steps, users will be able to use voice commands to perform various operations within a self-driving vehicle.
[0658] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0659] The present invention relates to a system that assists voice communication, providing an effective solution, particularly for situations or environments where speech is difficult. This system uses smartphones, tablets, and other electronic devices to convert voice data into text data, generate responses based on the context, and output the responses as voice data. Furthermore, by combining it with an emotion engine that recognizes emotions from the user's voice, it is possible to generate more appropriate and human-like responses.
[0660] System Configuration
[0661] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[0662] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[0663] The emotion engine recognizes the user's emotions from the voice data and provides this emotion information to the server, which then adjusts the generated responses based on this emotion information to generate more appropriate response candidates.
[0664] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0665] Program processing explanation
[0666] Voice input and initial processing
[0667] A user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?" The device captures the user's voice and saves the audio data. The device then sends the audio data to the server.
[0668] Voice data analysis and emotion recognition
[0669] The server receives the voice data and converts it into text using a speech recognition engine. The converted text data is then passed to a natural language processing (NLP) engine to analyze the context and intent. At the same time, the server uses an emotion engine to recognize the user's emotions. For example, it detects emotions such as tension in the user's voice.
[0670] Generating responses that take emotions into account
[0671] The server uses an AI model to generate optimal response candidates based on the results of NLP analysis and the recognition results of the emotion engine. By taking into account the information from the emotion engine, more appropriate responses can be tailored. For example, if the user is nervous, a gentle response such as "Relax, the next meeting is at 10:00 AM" will be generated.
[0672] Displaying response options and user selection
[0673] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. For example, "The next meeting is at 10:00 AM."
[0674] Text-to-speech and calling
[0675] The server passes the selected text data to a speech synthesis engine, converts it into voice data, and sends the generated voice data to the device, which plays it back through the speaker and communicates the response to the other party.
[0676] Specific examples
[0677] Example 1: Use in public places
[0678] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates candidate responses. At the same time, if the emotion engine detects tension in the user's voice, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby conveying the answer to the question to the other party.
[0679] Example 2: Use when feeling unwell
[0680] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[0681] This system serves as a powerful tool for efficient communication even in situations where speech is difficult. It also takes into account the user's emotions to provide more appropriate and human-like responses. A free text input function is also provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[0682] The processing flow will be explained below.
[0683] Step 1:
[0684] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[0685] Step 2:
[0686] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[0687] Step 3:
[0688] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[0689] Step 4:
[0690] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[0691] Step 5:
[0692] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[0693] Step 6:
[0694] At the same time, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine detects emotions such as nervousness, fatigue, or happiness from the tone, speed, and volume of the voice.
[0695] Step 7:
[0696] The server uses an AI model (e.g., generative AI) to generate optimal reply candidates based on the analysis results of the NLP engine and the recognition results of the emotion engine. Taking into account the information from the emotion engine, the server adjusts the reply according to the user's current emotional state. For example, if the user is nervous, it generates a reply that will make them feel more relaxed.
[0697] Step 8:
[0698] The server sends the generated answer candidates to the terminal in JSON format.
[0699] Step 9:
[0700] The device displays the received reply candidates to the user. The reply candidates are displayed using UI elements (e.g., buttons and lists) to make it easy for the user to select one.
[0701] Step 10:
[0702] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[0703] Step 11:
[0704] The device sends the user's selection to the server via an HTTP request or WebSocket.
[0705] Step 12:
[0706] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[0707] Step 13:
[0708] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[0709] Step 14:
[0710] The terminal plays back the received audio data so that the user can hear the content.
[0711] Step 15:
[0712] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[0713] Step 16:
[0714] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[0715] Step 17:
[0716] The terminal transmits the input text data to the server.
[0717] Step 18:
[0718] The server runs the free text through a speech synthesis engine and converts it into voice data.
[0719] Step 19:
[0720] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[0721] By combining this system with an emotion engine, it becomes possible to respond in a way that takes into account the user's emotional state, resulting in more natural and friendly communication. Furthermore, by using an emotion engine, responses that suit the user's state of mind can be provided, which is expected to have the effect of improving user satisfaction.
[0722] Example 2
[0723] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0724] Conventional voice communication systems have difficulty communicating effectively in situations or environments where speech is difficult. Furthermore, they lack the functionality to provide responses that take the user's emotions into account, making it difficult to generate more appropriate and human-like responses. Furthermore, the interface for selecting the most appropriate response from multiple candidate responses is inconvenient, which can hinder smooth communication.
[0725] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0726] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and recognizing the user's emotions, and means for generating appropriate reply candidates based on the analysis results and the emotional information. This makes it possible to convert voice data into text data and generate replies that take the user's emotions into consideration.
[0727] "Voice data" refers to acoustic data generated by a user's speech.
[0728] "Acquiring" refers to capturing and recording audio data using the device's microphone.
[0729] A "server" refers to a computer system that receives voice data from a terminal via a network and processes the data.
[0730] "Send" refers to the transfer of collected voice data by the terminal to the server.
[0731] "Text data" refers to data obtained by converting voice data into character information.
[0732] "Converting" refers to the process of changing audio data into text data or text data into audio data.
[0733] "Analyzing" refers to interpreting the meaning, intent, and context of text data using natural language processing techniques.
[0734] "Emotion recognition" refers to detecting a user's emotional state from speech or text data.
[0735] "Generating reply candidates" refers to creating appropriate responses based on text data and emotional information.
[0736] "Presenting to the user as options" refers to displaying the generated reply candidates as multiple options on the user's device.
[0737] "Speech synthesis" refers to the technology of converting text data into voice data.
[0738] "Transmitting audio" refers to playing the generated audio data through the device's speaker.
[0739] This invention is a system that allows users to smoothly communicate through speech, particularly in situations and environments where speech is difficult. The system converts speech into text and then analyzes the text to generate responses. It also recognizes the user's emotional state and reflects it in the responses, enabling more appropriate and human-like communication.
[0740] System Configuration
[0741] This system is realized mainly using the following hardware and software.
[0742] Hardware
[0743] 1. Device: An electronic device such as a smartphone or tablet that includes a microphone to capture the user's voice and a speaker to play back the response. Examples of smartphones include the Apple iPhone and Samsung Galaxy series.
[0744] 2. Microphone: Reliably captures the user's voice through a high-sensitivity microphone, such as the Shure MV88.
[0745] software
[0746] 1. Speech recognition engine: Software for converting voice data into text data. For example, we use the Google Speech-to-Text API.
[0747] 2. Natural Language Processing (NLP) Engine: Software that analyzes the converted text data to understand intent and context. We use OpenAI GPT-3 for this.
[0748] 3. Emotion Engine: Software for recognizing emotions from user voice and text data. As an example, we will use IBM Watson Tone Analyzer.
[0749] 4. Speech synthesis engine: Software that converts selected text data into speech data. An example is Amazon Polly.
[0750] System Operation Overview
[0751] When a user speaks into the device, the device's microphone captures the voice and saves it as audio data. This audio data is sent to the server using a communication protocol. The server then converts the audio data into text data using the Google Speech-to-Text API. It then uses OpenAI GPT-3 to analyze the context and intent of the text data. At the same time, it uses IBM Watson Tone Analyzer to recognize the user's emotions. Based on the analysis results and emotional information, OpenAI GPT-3 generates appropriate response candidates. These response candidates are sent from the server to the device and displayed to the user as multiple options. The user selects the most appropriate response from the displayed options and sends it to the server. The server then converts the selected text data into audio data using Amazon Polly and sends it back to the device. Finally, the audio data is played back through the device's speaker.
[0752] Specific examples
[0753] Example 1: Use in public places
[0754] Consider a situation where a user is attending an important meeting on the subway. During the meeting, someone asks, "Please tell me about the current progress." The user speaks the question into their smartphone, and the voice data is sent to the server. The server converts the voice into text, analyzes it, and generates candidate responses. If the emotion engine detects that the user is tense, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," and the voice data is played.
[0755] Example 2: Use when feeling unwell
[0756] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice and sends it to the server. The server converts the voice into text and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[0757] Prompt Sentence Examples
[0758] The following is an example of a prompt sentence to input to the generative AI model:
[0759] "When is the next meeting? User sentiment: Nervous Generate possible responses."
[0760] This prompt helps the system achieve smooth communication even in situations where the user has difficulty speaking.
[0761] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0762] Step 1:
[0763] The user speaks into the microphone on their smartphone or tablet.
[0764] The input is the user's speech. The output is the audio data captured by the device. Specifically, the user speaks, "When is the next meeting?" The device's high-sensitivity microphone (e.g., Shure MV88) picks up the voice and temporarily stores the audio data in the device's built-in storage.
[0765] Step 2:
[0766] The device compresses the captured audio data and sends it to the server using the HTTPS protocol.
[0767] The input is the audio data acquired in step 1. The output is compressed audio data sent to the server. Specifically, the device compresses the audio data for efficient transmission and sends it to the server encrypted with SSL / TLS.
[0768] Step 3:
[0769] The server converts the received voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text API).
[0770] The input is the voice data sent from the device. The output is the converted text data. Specifically, the server sends the voice data to the Google Speech-to-Text API, which returns the text data. The converted text data is then sent to the next processing step within the server.
[0771] Step 4:
[0772] The server passes the text data to a natural language processing (NLP) engine (e.g., OpenAI GPT-3) to analyze the context and intent.
[0773] The input is the text data generated in step 3. The output is the analyzed context information. Specifically, the server sends the text data to OpenAI GPT-3, and obtains information about the context and intent as the analysis result. This analysis result is used for subsequent processing.
[0774] Step 5:
[0775] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions.
[0776] The input is the text data and speech parameters generated in step 3. The output is the recognized emotion information. Specifically, the server sends the speech and text parameters to IBM Watson Tone Analyzer to recognize the user's emotion (e.g., tension, fatigue). This information is used to generate a response.
[0777] Step 6:
[0778] The server uses a generative AI model (e.g., OpenAI GPT-3) based on the analysis results and emotional information to generate appropriate response candidates.
[0779] The input is the context analysis result obtained in step 4 and the emotional information obtained in step 5. The output is multiple response candidates. Specifically, the server sends the analysis results and emotional information as prompts to OpenAI GPT-3, which then generates appropriate response candidates.
[0780] Step 7:
[0781] The server sends the generated answer candidates to the terminal in JSON format.
[0782] The input is the candidate answers generated in step 6. The output is the candidate answer data sent to the device. Specifically, the server formats the candidate answers in JSON format and sends them to the device.
[0783] Step 8:
[0784] The terminal displays the reply candidates on a user interface.
[0785] The input is candidate reply data sent from the server. The output is multiple candidate replies displayed on the user interface. Specifically, the device analyzes the candidate replies and displays options such as "Next meeting is at 10:00 AM," "Relax, next meeting is at 10:00 AM," and "Checking next meeting time" on the touchscreen.
[0786] Step 9:
[0787] The user selects the most appropriate option from the displayed options and taps it.
[0788] The input is the reply candidates displayed on the terminal. The output is the selected reply data. In concrete terms, the user selects "Relax, the next meeting is at 10:00 AM," and the selection information is saved on the terminal.
[0789] Step 10:
[0790] The terminal transmits the selected response data to the server.
[0791] The input is the response data selected by the user. The output is the selected data sent to the server. Specifically, the terminal encrypts the selected data again using SSL / TLS and sends it to the server.
[0792] Step 11:
[0793] The server passes the selected response to a speech synthesis engine (e.g., Amazon Polly) and converts it into voice data.
[0794] The input is the selected response data sent from the device. The output is the generated voice data. Specifically, the server sends the selected response to Amazon Polly, which generates the voice data.
[0795] Step 12:
[0796] The server transmits the generated voice data to the terminal.
[0797] The input is the voice data generated in step 11. The output is the voice data sent to the terminal. In concrete operation, the server sends the generated voice data back to the terminal.
[0798] Step 13:
[0799] The terminal plays the received audio data through a speaker.
[0800] The input is the voice data sent from the server. The output is the voice response that is transmitted to the other party. Specifically, the device decodes the voice data and plays it using the built-in speaker (e.g., Bose SoundLink Mini) to transmit the response to the other party.
[0801] (Application example 2)
[0802] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0803] Conventional voice communication assistance systems convert voice data into text data and generate responses, but they are not sufficient in generating appropriate responses that take the user's emotional state into account or in presenting responses quickly using visual devices. Therefore, there is a need for a system that provides more human-like and considerate responses in situations and environments where voice communication is difficult.
[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting voice data to text data, means for analyzing the text data and generating an appropriate response, means for displaying the generated response options, means for converting the selected response into voice data, means for transmitting the voice data, means for analyzing emotional data and adjusting the response, and means for displaying candidate responses on a visual device. This enables the generation of appropriate and prompt responses that take the user's emotional state into consideration.
[0805] "Voice data" refers to data in which the user's speech is recorded as an electronic signal.
[0806] "Text data" is data obtained by converting voice data into a character string format.
[0807] "Analysis" is the process of understanding the content of text data and grasping its context and intent.
[0808] A "reply" is the content of a response generated based on the analysis results.
[0809] "Choices" refer to multiple possible responses presented to the user.
[0810] "Display" refers to showing the generated responses or options on a visual device.
[0811] "Speech synthesis" is a technology that converts text data into voice data.
[0812] "Transmission" is a function for transmitting generated voice data to the outside.
[0813] "Emotion data" is data of emotions estimated from the user's voice, facial expressions, movements, etc.
[0814] A "visual device" is a device for visually presenting information to a user, and in this context refers to smart glasses and the like.
[0815] This invention relates to a system that supports voice communication, particularly to a customer service support application in brick-and-mortar stores using visual devices. This system converts voice data into text data and generates appropriate responses based on emotion data through collaboration between smart glasses and a server. The generated responses are displayed on the smart glasses' display and transmitted as voice data.
[0816] Components and hardware / software used
[0817] 1. Smart Glasses:
[0818] A device worn by the user, equipped with a microphone and a display, that captures voice data and provides a visual response.
[0819] 2. Server:
[0820] It processes voice data, generates text data, analyzes emotional data, generates responses, and synthesizes voice.
[0821] The software used includes a speech recognition engine, a natural language processing engine, a sentiment analysis engine, and a speech synthesis engine.
[0822] Program processing explanation
[0823] 1. Acquire audio data:
[0824] The user speaks a customer question into the microphone of the smart glasses, for example, "What cake do you recommend?"
[0825] The smart glasses capture this voice data and send it to a server.
[0826] 2. Convert to text data:
[0827] The server uses a voice recognition engine to convert the received voice data into text data.
[0828] The converted text data may be a string such as "What cake do you recommend?"
[0829] 3. Emotional Data Analysis:
[0830] The server uses an emotion analysis engine to extract emotion data from the user's speech.
[0831] For example, friendliness is detected from the user's tone of voice.
[0832] 4. Generate a response:
[0833] The natural language processing engine generates optimal responses based on text data and emotional data.
[0834] A possible reply such as "Today's recommendation is shortcake" is generated.
[0835] 5. Visual display of response:
[0836] The generated response is sent from the server to the smart glasses and displayed on the display.
[0837] The user checks the displayed reply candidates and selects an appropriate reply.
[0838] 6. Conversion to voice data and transmission:
[0839] The server passes the selected response to a speech synthesis engine and converts it into voice data.
[0840] The generated voice data is transmitted to the customer through the speaker of the smart glasses.
[0841] Specific examples
[0842] Consider a scenario where a cafe staff member is wearing smart glasses during a busy lunchtime. When a customer asks, "What cake do you recommend?", the staff member captures the voice data through the microphone in the smart glasses. The server analyzes the voice as text data and generates a response that takes emotional data into account. A response such as "Today's recommendation is shortcake" is displayed on the smart glasses' display, and the staff member relays this to the customer.
[0843] Prompt Sentence Examples
[0844] Examples of prompts for a generative AI model include "Emotion-recognizing customer service assistant at a cafe" and "Generate a friendly response when a customer asks, 'What cake do you recommend?'"
[0845] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0846] Step 1:
[0847] The user speaks into the microphone of the smart glasses. For example, they say, "What cake do you recommend?" The smart glasses capture this voice data and store it in the device's memory. At this point, the input is the user's voice data, and the output is the saved audio file.
[0848] Step 2:
[0849] The device sends the acquired voice data to the server. The input is the saved voice file, and the output is the voice data sent to the server. Specifically, the smart glasses upload the voice data to the server.
[0850] Step 3:
[0851] The server runs the received voice data through a voice recognition engine and converts it into text data. At this point, the voice data is output as character string data. Specifically, the text data obtained is "What cake do you recommend?" The input is voice data, and the output is text data.
[0852] Step 4:
[0853] At the same time, the server uses an emotion analysis engine to extract emotional data from the voice data. For example, the emotion "friendliness" can be detected from the tone of the user's voice. In this case, the input is voice data, and the output is data indicating the emotional state. Specifically, the emotion analysis engine performs a process to analyze the voice characteristics.
[0854] Step 5:
[0855] The server passes the text data and emotion data to a natural language processing engine, which then generates the optimal response. For example, the response generated might be, "Today's recommendation is shortcake." The input is text data and emotion data, and the output is the generated response data. Specifically, the generative AI model creates a response based on the prompt sentence.
[0856] Step 6:
[0857] The server prepares the generated responses as multiple options and sends them to the smart glasses. In this case, options such as "Today's recommendation is shortcake" are included. The input is the generated response data, and the output is the response options displayed on the smart glasses. Specifically, the option data is displayed on the smart glasses' display.
[0858] Step 7:
[0859] The user checks the display of the smart glasses and selects the most appropriate response. For example, they might select "Today's recommendation is shortcake." In this case, the input is the response options displayed on the display, and the output is the user's selection. Specifically, the user confirms the selection by touching the display.
[0860] Step 8:
[0861] The server passes the selected response to the speech synthesis engine and converts it into speech data. At this time, the text data is output as a speech file. Speech data such as "Today's recommendation is shortcake" is generated. The input is the selected response data, and the output is speech data. Specifically, the speech synthesis engine performs the process of converting text into speech.
[0862] Step 9:
[0863] The terminal transmits the generated voice data through the speaker of the smart glasses. The customer is provided with the reply, "Today's recommendation is shortcake." The input is the generated voice data, and the output is the voice that the customer hears through the speaker. The specific operation is that the voice is played back from the speaker of the smart glasses.
[0864] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0865] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0866] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0867] [Third embodiment]
[0868] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0869] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0870] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0871] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0872] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0873] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0874] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0875] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0876] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0877] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0878] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0879] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0880] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0881] System Configuration
[0882] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[0883] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[0884] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0885] Program processing explanation
[0886] Voice input and initial processing
[0887] Users can input voice by speaking into the microphone of their smartphone or tablet. The device captures this voice data and sends it to the server. This includes recording and data transmission functions.
[0888] Analysis of audio data
[0889] The server passes the received voice data to a speech recognition engine, which converts it into text data. It then uses natural language processing (NLP) technology to analyze the text and understand the context. Based on the analysis results, it generates appropriate response candidates.
[0890] Displaying response options and user selection
[0891] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. The selected answer is then sent back to the server via the terminal.
[0892] Text-to-speech and calling
[0893] The server receives the selected text data and converts it into voice data using a speech synthesis engine. This voice data is then sent to the device, which plays it back through the speaker and transmits it to the other party.
[0894] Specific examples
[0895] Example 1: Use in public places
[0896] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates response candidates such as "Progress is going well," "It's behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby communicating the response to the question to the other party.
[0897] Example 2: Use when feeling unwell
[0898] Consider a case where a user is unwell and has difficulty speaking. When an important call comes in, the other party says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. The user selects "I'll reschedule right away, so please wait" from the options presented, and the server converts this into voice data and transmits it to the other party via the device.
[0899] This system functions as a powerful tool for efficient communication even in situations where speech is difficult. In addition, when options are not appropriate, a free text input function is provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[0900] The processing flow will be explained below.
[0901] Step 1:
[0902] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[0903] Step 2:
[0904] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[0905] Step 3:
[0906] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[0907] Step 4:
[0908] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[0909] Step 5:
[0910] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[0911] Step 6:
[0912] Based on the analysis results of the NLP engine, the server uses an AI model (e.g., generative AI) to generate optimal response candidates, such as "The next meeting is at 10:00 AM," "The next meeting is at 3:00 PM," and "The next meeting needs to be rescheduled."
[0913] Step 7:
[0914] The server sends the generated answer candidates to the terminal in JSON format.
[0915] Step 8:
[0916] The device displays the received reply candidates to the user using UI elements (e.g., buttons or lists) to allow the user to select one.
[0917] Step 9:
[0918] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[0919] Step 10:
[0920] The device sends the user's selection to the server via an HTTP request or WebSocket.
[0921] Step 11:
[0922] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[0923] Step 12:
[0924] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[0925] Step 13:
[0926] The terminal plays back the received audio data so that the user can hear the content.
[0927] Step 14:
[0928] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[0929] Step 15:
[0930] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[0931] Step 16:
[0932] The terminal transmits the input text data to the server.
[0933] Step 17:
[0934] The server runs the free text through a speech synthesis engine and converts it into voice data.
[0935] Step 18:
[0936] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[0937] Example 1
[0938] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0939] Currently, many people face situations and environments that make it difficult to communicate through voice. For example, it is difficult to communicate verbally in public places or when feeling unwell. There is a need for a system that can solve these problems and enable everyone to communicate easily.
[0940] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0941] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting the selected response into voice data]; and [means for transmitting the voice data]. This provides a system that performs a series of processes from capturing voice data to outputting the response audibly, enabling effective communication even in situations where speaking is difficult.
[0942] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0943] "Text data" is data in the form of a character string generated by analyzing voice data.
[0944] A "server" is a computer system that processes voice and text data over a network.
[0945] A "terminal" is an electronic device that a user uses to input voice and receive information.
[0946] A "voice recognition engine" is a technology or system that converts voice data into text data.
[0947] "Natural language processing technology" is a technology that analyzes text data, understands the context, and generates appropriate responses.
[0948] A "generative AI model" is an artificial intelligence algorithm that generates appropriate responses from text data.
[0949] A "speech synthesis engine" is a technology or system that converts text data into speech data.
[0950] A "user interface" refers to a screen or operating means for displaying information to a user on a terminal and accepting input.
[0951] A "prompt" is an instruction given to a generative AI model that is used to generate appropriate response candidates.
[0952] "Network communication" refers to the Internet and other internal and external data communication means for sending and receiving data.
[0953] "Real-time" refers to a processing method that minimizes delays and provides immediate response.
[0954] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[0955] The terminal receives voice data input by the user and transmits it to a server, which then converts the voice data into text data using voice recognition technology. For example, a smartphone, tablet, or dedicated recording device can be used to acquire this voice data. The acquired voice data is then transmitted to the server using a network communication means.
[0956] The server receives the voice data and instantly converts it into text using a speech recognition engine such as the Google Speech-to-Text API. It then analyzes the text using natural language processing technology to generate the most appropriate response based on the context. Natural language processing engines such as Spacy and BERT are used for the analysis.
[0957] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. A speech synthesis engine such as Amazon Polly is used for the speech synthesis. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[0958] Specific examples
[0959] Example 1: Use in public places
[0960] Suppose a user is participating in an important meeting on the subway and is asked, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server uses a voice recognition engine to convert the voice data into text data and analyzes it. It then generates response candidates such as "Progress is going well" or "We're behind schedule, but we're working on it." The user selects "Progress is going well," and the server converts this into voice data and sends it to the device. Finally, the device plays back this voice data and communicates it to the other party.
[0961] Example 2: Use when feeling unwell
[0962] Suppose a user is feeling unwell and has difficulty speaking, but receives an important call. If the other party says, "I'd like to reschedule our next meeting," the device captures the voice data and sends it to the server. The server converts the voice data into text data and generates appropriate reply candidates. For example, a reply candidate such as "I'll reschedule right away, so please wait." The user selects "I'll reschedule right away, so please wait," and the server converts this into voice data and sends it to the device. The device then plays this voice data and conveys it to the other party.
[0963] Prompt Sentence Examples
[0964] Prompt: Transcribe the following audio data into text and generate an appropriate response.
[0965] Audio data: "I'd like to schedule our next meeting. When would be convenient for you?"
[0966] This system is a powerful tool for efficient communication even in situations where speech is difficult, greatly improving flexibility and convenience. In addition, when the options are not appropriate, a free text input function is provided, allowing users to enter responses in their own words and play them back as audio data.
[0967] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0968] Step 1:
[0969] A user speaks into the microphone of a smartphone or tablet to input voice. The device captures this voice data in real time. Specifically, the device's recording function is used to convert the voice signal collected from the microphone into digital voice data. The input is the user's voice, and the output is digital voice data.
[0970] Step 2:
[0971] The device sends the captured audio data to the server. Using the network communication function, the data is sent via a communication protocol such as an HTTP request or WebSocket. The input is the captured digital audio data, and the output is a transmission confirmation message to the server.
[0972] Step 3:
[0973] The server passes the received voice data to a voice recognition engine and converts it into text data. This conversion is performed using the Google Speech-to-Text API or similar. The server passes the voice data to the voice recognition engine and receives the returned text data. The input is voice data and the output is text data.
[0974] Step 4:
[0975] The server passes the converted text data to a natural language processing (NLP) engine to analyze the context of the text. NLP techniques such as Spacy and BERT are used for this analysis. The server passes the text data to the engine and receives the analysis results in return. The input is the text data, and the output is the analysis results and contextual information.
[0976] Step 5:
[0977] The server uses the parsed text data to generate appropriate reply candidates using a generative AI model (e.g., OpenAI GPT-3). The server passes the prompt sentence and the analysis result to the generative AI model and obtains the reply candidates that are generated. The input is the prompt sentence and the analysis result, and the output is multiple reply candidates.
[0978] Step 6:
[0979] The server sends the generated answer candidates to the terminal. Data is sent quickly using a communication protocol. The input is the generated answer candidates, and the output is a transmission confirmation message to the terminal.
[0980] Step 7:
[0981] The device displays the received reply candidates in a user interface. It uses a UI library (e.g., React Native or Flutter) to display the candidate list on a display or touchscreen. The input is the generated reply candidates, and the output is the displayed list of options.
[0982] Step 8:
[0983] The user selects the most appropriate answer from the presented answer candidates. The selection is made by the user's touch or click. The input is the displayed list of options, and the output is the user's selection.
[0984] Step 9:
[0985] The terminal sends the user's selection to the server. The selection result is sent using network communication. The input is the user's selection result, and the output is a transmission confirmation message to the server.
[0986] Step 10:
[0987] The server passes the selected text data to a speech synthesis engine (e.g., Amazon Polly) and converts it into speech data. The speech synthesis engine converts the text data into speech waveforms and converts them into a playable format. The input is the selected text data, and the output is speech data.
[0988] Step 11:
[0989] The server sends the generated voice data to the terminal. The data is sent quickly using a communication protocol. The input is the generated voice data, and the output is a transmission confirmation message to the terminal.
[0990] Step 12:
[0991] The device plays the received audio data. The audio is output through the built-in speaker or connected earphones. A multimedia library (e.g., AVFoundation or MediaPlayer) is used for playback. The input is the generated audio data, and the output is the played audio.
[0992] (Application example 1)
[0993] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0994] In autonomous vehicles, there is a demand for a system that allows drivers to safely and efficiently operate various functions, make emergency contacts, and use infotainment functions using only voice. Furthermore, there is a lack of a mechanism that allows smooth communication and maximizes the use of autonomous vehicle functions even when the driver cannot speak. Therefore, the present invention aims to solve these problems by providing a system that handles everything from voice data acquisition to response generation and emergency response.
[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0996] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting a selected response into voice data]; [means for transmitting the voice data]; [means for recognizing voice commands in an autonomously controlled vehicle and executing driving behaviors and settings]; [means for detecting an emergency and automatically contacting an emergency contact]; and [means for playing music, reading news, and checking vehicle status by voice as infotainment functions.] This enables the driver to safely and efficiently operate an autonomously driven vehicle via voice input, automatically take countermeasures in an emergency, and further utilize infotainment functions.
[0997] The "means for acquiring voice data" is a function for capturing voice uttered by a user using a microphone of an electronic device and transmitting the voice data to a server.
[0998] The "means for converting voice data into text data" is a function for converting acquired voice data into text data using voice recognition technology.
[0999] The "means for analyzing text data and generating an appropriate response" is a function that analyzes the converted text data using natural language processing technology and generates an appropriate response according to the context.
[1000] The "means for displaying generated reply options" is a function that displays multiple reply options generated by the server to the user on the terminal.
[1001] The "means for converting the selected response into voice data" is a function for converting the text-format response selected by the user into voice data using voice synthesis technology.
[1002] "Means for transmitting voice data" is a function that transmits converted voice data to the other party through a speaker.
[1003] "Means for recognizing voice commands within an automatically controlled vehicle and executing driving behaviors or settings" refers to a function that recognizes voice commands issued by a user within the vehicle and executes driving behaviors or setting changes of the vehicle accordingly.
[1004] "Means of detecting an emergency and automatically contacting emergency contacts" is a function that detects an emergency (e.g., accident or illness) while driving, and automatically contacts registered emergency contacts.
[1005] "Means for playing music, reading news, and checking vehicle status by voice as infotainment functions" refers to a function that uses voice input to execute infotainment functions such as playing music, reading news, and checking vehicle status.
[1006] This invention is a system that supports voice communication and is applied to autonomous vehicles. It is a system that covers everything from voice data acquisition and analysis to response generation and emergency response. The system mainly operates via a terminal and a server.
[1007] The terminal has the role of acquiring voice data emitted by the user. Specifically, it captures the voice data using a microphone built into a smartphone, tablet, or vehicle and transmits it to a server.
[1008] The server performs a series of processes, including the following steps:
[1009] 1. Converting audio data to text data:
[1010] Using voice recognition technology, the server converts the voice data received from the device into text data. To do this, the server uses various voice recognition software (e.g., Google Speech-to-Text API, etc.).
[1011] 2. Text data analysis and response generation:
[1012] Leverage natural language processing techniques to analyze text data and generate appropriate responses based on the context, for example, using generative AI models (e.g., OpenAI's GPT-3) to generate responses based on prompts.
[1013] 3. Displaying generated answer choices:
[1014] The generated multiple reply candidates are sent to the terminal, and the user selects the most appropriate reply on the terminal.
[1015] 4. Converting selected responses to audio data:
[1016] The server converts the selected text response into audio using a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API).
[1017] 5. Transmission of voice data:
[1018] The converted audio data is sent back to the device and played through the speaker.
[1019] Specific examples
[1020] Navigation function in autonomous vehicles
[1021] When a user says, "Tell me the shortest route," the device captures the voice data and sends it to the server. The server converts the voice data into text data and sends the following prompt sentence to the generative AI model:
[1022] Example prompt sentence:
[1023] "What is the shortest route from my current location to my destination?"
[1024] The generated multiple answer candidates (e.g., "Turn right after 3.2 km" or "Route A is the shortest route") are displayed on the user's device. The user selects the appropriate answer (e.g., "Route A is the shortest route"), and the server converts it into voice data. The device plays the converted voice data, guiding the user to the shortest route.
[1025] Emergency response features
[1026] If a user feels unwell while driving and instructs the device to "make an emergency call," the device captures the voice data and sends it to the server. The server converts the voice data into text data and automatically calls the emergency contact. The device also provides voice information about the vehicle's location and the user's condition. For example, information such as "The driver is currently feeling unwell. Their location is..." is automatically sent to the contact.
[1027] This will enable a variety of functions to be realized through voice commands within self-driving vehicles, improving safety and convenience.
[1028] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1029] Step 1:
[1030] Acquiring audio data
[1031] Input: Voice command from the user
[1032] Specific behavior:
[1033] The user issues a voice command in the vehicle, and the device (a smartphone, tablet, or a microphone built into the vehicle) captures the voice data.
[1034] Step 2:
[1035] Sending voice data to the server
[1036] Input: Captured audio data
[1037] Output: Audio data sent to the server
[1038] Specific behavior:
[1039] The device sends the captured audio data to a server using an internet connection.
[1040] Step 3:
[1041] Converting audio data to text data
[1042] Input: Audio data sent to the server
[1043] Output: Converted text data
[1044] Specific behavior:
[1045] Using speech recognition technology (such as Google Speech-to-Text API), the server converts the received voice data into text data. This process uses a speech recognition engine.
[1046] Step 4:
[1047] Text data analysis and response generation
[1048] Input: Converted text data
[1049] Output: Generated answer candidates
[1050] Specific behavior:
[1051] The server analyzes the text data using natural language processing technology to understand its context, and uses a generative AI model (such as OpenAI GPT-3) to generate appropriate prompts and create candidate responses.
[1052] Step 5:
[1053] Server sending of possible replies
[1054] Input: Generated answer candidates
[1055] Output: Device where possible responses are displayed
[1056] Specific behavior:
[1057] The server then sends the generated reply candidates to the terminal.
[1058] Step 6:
[1059] Viewing and selecting possible responses
[1060] Input: Answer candidates sent by the server
[1061] Output: Selected response
[1062] Specific behavior:
[1063] The device displays possible responses to the user, who then selects the response that they think is most appropriate.
[1064] Step 7:
[1065] Converting selected responses to audio data
[1066] Input: Selected response (text format)
[1067] Output: Audio data
[1068] Specific behavior:
[1069] The server converts the selected text response into audio data using speech synthesis technology (e.g., Google Text-to-Speech API).
[1070] Step 8:
[1071] Sending and playing audio data
[1072] Input: Converted audio data
[1073] Output: Audio data transmitted through the speaker
[1074] Specific behavior:
[1075] The server then sends the final generated voice data to the device, which then plays the data through a speaker and transmits it to the other party.
[1076] By following these steps, users will be able to use voice commands to perform various operations within a self-driving vehicle.
[1077] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1078] The present invention relates to a system that assists voice communication, providing an effective solution, particularly for situations or environments where speech is difficult. This system uses smartphones, tablets, and other electronic devices to convert voice data into text data, generate responses based on the context, and output the responses as voice data. Furthermore, by combining it with an emotion engine that recognizes emotions from the user's voice, it is possible to generate more appropriate and human-like responses.
[1079] System Configuration
[1080] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[1081] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[1082] The emotion engine recognizes the user's emotions from the voice data and provides this emotion information to the server, which then adjusts the generated responses based on this emotion information to generate more appropriate response candidates.
[1083] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[1084] Program processing explanation
[1085] Voice input and initial processing
[1086] A user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?" The device captures the user's voice and saves the audio data. The device then sends the audio data to the server.
[1087] Voice data analysis and emotion recognition
[1088] The server receives the voice data and converts it into text using a speech recognition engine. The converted text data is then passed to a natural language processing (NLP) engine to analyze the context and intent. At the same time, the server uses an emotion engine to recognize the user's emotions. For example, it detects emotions such as tension in the user's voice.
[1089] Generating responses that take emotions into account
[1090] The server uses an AI model to generate optimal response candidates based on the results of NLP analysis and the recognition results of the emotion engine. By taking into account the information from the emotion engine, more appropriate responses can be tailored. For example, if the user is nervous, a gentle response such as "Relax, the next meeting is at 10:00 AM" will be generated.
[1091] Displaying response options and user selection
[1092] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. For example, "The next meeting is at 10:00 AM."
[1093] Text-to-speech and calling
[1094] The server passes the selected text data to a speech synthesis engine, converts it into voice data, and sends the generated voice data to the device, which plays it back through the speaker and communicates the response to the other party.
[1095] Specific examples
[1096] Example 1: Use in public places
[1097] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates candidate responses. At the same time, if the emotion engine detects tension in the user's voice, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby conveying the answer to the question to the other party.
[1098] Example 2: Use when feeling unwell
[1099] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[1100] This system serves as a powerful tool for efficient communication even in situations where speech is difficult. It also takes into account the user's emotions to provide more appropriate and human-like responses. A free text input function is also provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[1101] The processing flow will be explained below.
[1102] Step 1:
[1103] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[1104] Step 2:
[1105] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[1106] Step 3:
[1107] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[1108] Step 4:
[1109] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[1110] Step 5:
[1111] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[1112] Step 6:
[1113] At the same time, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine detects emotions such as nervousness, fatigue, or happiness from the tone, speed, and volume of the voice.
[1114] Step 7:
[1115] The server uses an AI model (e.g., generative AI) to generate optimal reply candidates based on the analysis results of the NLP engine and the recognition results of the emotion engine. Taking into account the information from the emotion engine, the server adjusts the reply according to the user's current emotional state. For example, if the user is nervous, it generates a reply that will make them feel more relaxed.
[1116] Step 8:
[1117] The server sends the generated answer candidates to the terminal in JSON format.
[1118] Step 9:
[1119] The device displays the received reply candidates to the user. The reply candidates are displayed using UI elements (e.g., buttons and lists) to make it easy for the user to select one.
[1120] Step 10:
[1121] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[1122] Step 11:
[1123] The device sends the user's selection to the server via an HTTP request or WebSocket.
[1124] Step 12:
[1125] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[1126] Step 13:
[1127] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[1128] Step 14:
[1129] The terminal plays back the received audio data so that the user can hear the content.
[1130] Step 15:
[1131] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[1132] Step 16:
[1133] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[1134] Step 17:
[1135] The terminal transmits the input text data to the server.
[1136] Step 18:
[1137] The server runs the free text through a speech synthesis engine and converts it into voice data.
[1138] Step 19:
[1139] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[1140] By combining this system with an emotion engine, it becomes possible to respond in a way that takes into account the user's emotional state, resulting in more natural and friendly communication. Furthermore, by using an emotion engine, responses that suit the user's state of mind can be provided, which is expected to have the effect of improving user satisfaction.
[1141] Example 2
[1142] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1143] Conventional voice communication systems have difficulty communicating effectively in situations or environments where speech is difficult. Furthermore, they lack the functionality to provide responses that take the user's emotions into account, making it difficult to generate more appropriate and human-like responses. Furthermore, the interface for selecting the most appropriate response from multiple candidate responses is inconvenient, which can hinder smooth communication.
[1144] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1145] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and recognizing the user's emotions, and means for generating appropriate reply candidates based on the analysis results and the emotional information. This makes it possible to convert voice data into text data and generate replies that take the user's emotions into consideration.
[1146] "Voice data" refers to acoustic data generated by a user's speech.
[1147] "Acquiring" refers to capturing and recording audio data using the device's microphone.
[1148] A "server" refers to a computer system that receives voice data from a terminal via a network and processes the data.
[1149] "Send" refers to the transfer of collected voice data by the terminal to the server.
[1150] "Text data" refers to data obtained by converting voice data into character information.
[1151] "Converting" refers to the process of changing audio data into text data or text data into audio data.
[1152] "Analyzing" refers to interpreting the meaning, intent, and context of text data using natural language processing techniques.
[1153] "Emotion recognition" refers to detecting a user's emotional state from speech or text data.
[1154] "Generating reply candidates" refers to creating appropriate responses based on text data and emotional information.
[1155] "Presenting to the user as options" refers to displaying the generated reply candidates as multiple options on the user's device.
[1156] "Speech synthesis" refers to the technology of converting text data into voice data.
[1157] "Transmitting audio" refers to playing the generated audio data through the device's speaker.
[1158] This invention is a system that allows users to smoothly communicate through speech, particularly in situations and environments where speech is difficult. The system converts speech into text and then analyzes the text to generate responses. It also recognizes the user's emotional state and reflects it in the responses, enabling more appropriate and human-like communication.
[1159] System Configuration
[1160] This system is realized mainly using the following hardware and software.
[1161] Hardware
[1162] 1. Device: An electronic device such as a smartphone or tablet that includes a microphone to capture the user's voice and a speaker to play back the response. Examples of smartphones include the Apple iPhone and Samsung Galaxy series.
[1163] 2. Microphone: Reliably captures the user's voice through a high-sensitivity microphone, such as the Shure MV88.
[1164] software
[1165] 1. Speech recognition engine: Software for converting voice data into text data. For example, we use the Google Speech-to-Text API.
[1166] 2. Natural Language Processing (NLP) Engine: Software that analyzes the converted text data to understand intent and context. We use OpenAI GPT-3 for this.
[1167] 3. Emotion Engine: Software for recognizing emotions from user voice and text data. As an example, we will use IBM Watson Tone Analyzer.
[1168] 4. Speech synthesis engine: Software that converts selected text data into speech data. An example is Amazon Polly.
[1169] System Operation Overview
[1170] When a user speaks into the device, the device's microphone captures the voice and saves it as audio data. This audio data is sent to the server using a communication protocol. The server then converts the audio data into text data using the Google Speech-to-Text API. It then uses OpenAI GPT-3 to analyze the context and intent of the text data. At the same time, it uses IBM Watson Tone Analyzer to recognize the user's emotions. Based on the analysis results and emotional information, OpenAI GPT-3 generates appropriate response candidates. These response candidates are sent from the server to the device and displayed to the user as multiple options. The user selects the most appropriate response from the displayed options and sends it to the server. The server then converts the selected text data into audio data using Amazon Polly and sends it back to the device. Finally, the audio data is played back through the device's speaker.
[1171] Specific examples
[1172] Example 1: Use in public places
[1173] Consider a situation where a user is attending an important meeting on the subway. During the meeting, someone asks, "Please tell me about the current progress." The user speaks the question into their smartphone, and the voice data is sent to the server. The server converts the voice into text, analyzes it, and generates candidate responses. If the emotion engine detects that the user is tense, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," and the voice data is played.
[1174] Example 2: Use when feeling unwell
[1175] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice and sends it to the server. The server converts the voice into text and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[1176] Prompt Sentence Examples
[1177] The following is an example of a prompt sentence to input to the generative AI model:
[1178] "When is the next meeting? User sentiment: Nervous Generate possible responses."
[1179] This prompt helps the system achieve smooth communication even in situations where the user has difficulty speaking.
[1180] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1181] Step 1:
[1182] The user speaks into the microphone on their smartphone or tablet.
[1183] The input is the user's speech. The output is the audio data captured by the device. Specifically, the user speaks, "When is the next meeting?" The device's high-sensitivity microphone (e.g., Shure MV88) picks up the voice and temporarily stores the audio data in the device's built-in storage.
[1184] Step 2:
[1185] The device compresses the captured audio data and sends it to the server using the HTTPS protocol.
[1186] The input is the audio data acquired in step 1. The output is compressed audio data sent to the server. Specifically, the device compresses the audio data for efficient transmission and sends it to the server encrypted with SSL / TLS.
[1187] Step 3:
[1188] The server converts the received voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text API).
[1189] The input is the voice data sent from the device. The output is the converted text data. Specifically, the server sends the voice data to the Google Speech-to-Text API, which returns the text data. The converted text data is then sent to the next processing step within the server.
[1190] Step 4:
[1191] The server passes the text data to a natural language processing (NLP) engine (e.g., OpenAI GPT-3) to analyze the context and intent.
[1192] The input is the text data generated in step 3. The output is the analyzed context information. Specifically, the server sends the text data to OpenAI GPT-3, and obtains information about the context and intent as the analysis result. This analysis result is used for subsequent processing.
[1193] Step 5:
[1194] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions.
[1195] The input is the text data and speech parameters generated in step 3. The output is the recognized emotion information. Specifically, the server sends the speech and text parameters to IBM Watson Tone Analyzer to recognize the user's emotion (e.g., tension, fatigue). This information is used to generate a response.
[1196] Step 6:
[1197] The server uses a generative AI model (e.g., OpenAI GPT-3) based on the analysis results and emotional information to generate appropriate response candidates.
[1198] The input is the context analysis result obtained in step 4 and the emotional information obtained in step 5. The output is multiple response candidates. Specifically, the server sends the analysis results and emotional information as prompts to OpenAI GPT-3, which then generates appropriate response candidates.
[1199] Step 7:
[1200] The server sends the generated answer candidates to the terminal in JSON format.
[1201] The input is the candidate answers generated in step 6. The output is the candidate answer data sent to the device. Specifically, the server formats the candidate answers in JSON format and sends them to the device.
[1202] Step 8:
[1203] The terminal displays the reply candidates on a user interface.
[1204] The input is candidate reply data sent from the server. The output is multiple candidate replies displayed on the user interface. Specifically, the device analyzes the candidate replies and displays options such as "Next meeting is at 10:00 AM," "Relax, next meeting is at 10:00 AM," and "Checking next meeting time" on the touchscreen.
[1205] Step 9:
[1206] The user selects the most appropriate option from the displayed options and taps it.
[1207] The input is the reply candidates displayed on the terminal. The output is the selected reply data. In concrete terms, the user selects "Relax, the next meeting is at 10:00 AM," and the selection information is saved on the terminal.
[1208] Step 10:
[1209] The terminal transmits the selected response data to the server.
[1210] The input is the response data selected by the user. The output is the selected data sent to the server. Specifically, the terminal encrypts the selected data again using SSL / TLS and sends it to the server.
[1211] Step 11:
[1212] The server passes the selected response to a speech synthesis engine (e.g., Amazon Polly) and converts it into voice data.
[1213] The input is the selected response data sent from the device. The output is the generated voice data. Specifically, the server sends the selected response to Amazon Polly, which generates the voice data.
[1214] Step 12:
[1215] The server transmits the generated voice data to the terminal.
[1216] The input is the voice data generated in step 11. The output is the voice data sent to the terminal. In concrete operation, the server sends the generated voice data back to the terminal.
[1217] Step 13:
[1218] The terminal plays the received audio data through a speaker.
[1219] The input is the voice data sent from the server. The output is the voice response that is transmitted to the other party. Specifically, the device decodes the voice data and plays it using the built-in speaker (e.g., Bose SoundLink Mini) to transmit the response to the other party.
[1220] (Application example 2)
[1221] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1222] Conventional voice communication assistance systems convert voice data into text data and generate responses, but they are not sufficient in generating appropriate responses that take the user's emotional state into account or in presenting responses quickly using visual devices. Therefore, there is a need for a system that provides more human-like and considerate responses in situations and environments where voice communication is difficult.
[1223] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting voice data to text data, means for analyzing the text data and generating an appropriate response, means for displaying the generated response options, means for converting the selected response into voice data, means for transmitting the voice data, means for analyzing emotional data and adjusting the response, and means for displaying candidate responses on a visual device. This enables the generation of appropriate and prompt responses that take the user's emotional state into consideration.
[1224] "Voice data" refers to data in which the user's speech is recorded as an electronic signal.
[1225] "Text data" is data obtained by converting voice data into a character string format.
[1226] "Analysis" is the process of understanding the content of text data and grasping its context and intent.
[1227] A "reply" is the content of a response generated based on the analysis results.
[1228] "Choices" refer to multiple possible responses presented to the user.
[1229] "Display" refers to showing the generated responses or options on a visual device.
[1230] "Speech synthesis" is a technology that converts text data into voice data.
[1231] "Transmission" is a function for transmitting generated voice data to the outside.
[1232] "Emotion data" is data of emotions estimated from the user's voice, facial expressions, movements, etc.
[1233] A "visual device" is a device for visually presenting information to a user, and in this context refers to smart glasses and the like.
[1234] This invention relates to a system that supports voice communication, particularly to a customer service support application in brick-and-mortar stores using visual devices. This system converts voice data into text data and generates appropriate responses based on emotion data through collaboration between smart glasses and a server. The generated responses are displayed on the smart glasses' display and transmitted as voice data.
[1235] Components and hardware / software used
[1236] 1. Smart Glasses:
[1237] A device worn by the user, equipped with a microphone and a display, that captures voice data and provides a visual response.
[1238] 2. Server:
[1239] It processes voice data, generates text data, analyzes emotional data, generates responses, and synthesizes voice.
[1240] The software used includes a speech recognition engine, a natural language processing engine, a sentiment analysis engine, and a speech synthesis engine.
[1241] Program processing explanation
[1242] 1. Acquire audio data:
[1243] The user speaks a customer question into the microphone of the smart glasses, for example, "What cake do you recommend?"
[1244] The smart glasses capture this voice data and send it to a server.
[1245] 2. Convert to text data:
[1246] The server uses a voice recognition engine to convert the received voice data into text data.
[1247] The converted text data may be a string such as "What cake do you recommend?"
[1248] 3. Emotional Data Analysis:
[1249] The server uses an emotion analysis engine to extract emotion data from the user's speech.
[1250] For example, friendliness is detected from the user's tone of voice.
[1251] 4. Generate a response:
[1252] The natural language processing engine generates optimal responses based on text data and emotional data.
[1253] A possible reply such as "Today's recommendation is shortcake" is generated.
[1254] 5. Visual display of response:
[1255] The generated response is sent from the server to the smart glasses and displayed on the display.
[1256] The user checks the displayed reply candidates and selects an appropriate reply.
[1257] 6. Conversion to voice data and transmission:
[1258] The server passes the selected response to a speech synthesis engine and converts it into voice data.
[1259] The generated voice data is transmitted to the customer through the speaker of the smart glasses.
[1260] Specific examples
[1261] Consider a scenario where a cafe staff member is wearing smart glasses during a busy lunchtime. When a customer asks, "What cake do you recommend?", the staff member captures the voice data through the microphone in the smart glasses. The server analyzes the voice as text data and generates a response that takes emotional data into account. A response such as "Today's recommendation is shortcake" is displayed on the smart glasses' display, and the staff member relays this to the customer.
[1262] Prompt Sentence Examples
[1263] Examples of prompts for a generative AI model include "Emotion-recognizing customer service assistant at a cafe" and "Generate a friendly response when a customer asks, 'What cake do you recommend?'"
[1264] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1265] Step 1:
[1266] The user speaks into the microphone of the smart glasses. For example, they say, "What cake do you recommend?" The smart glasses capture this voice data and store it in the device's memory. At this point, the input is the user's voice data, and the output is the saved audio file.
[1267] Step 2:
[1268] The device sends the acquired voice data to the server. The input is the saved voice file, and the output is the voice data sent to the server. Specifically, the smart glasses upload the voice data to the server.
[1269] Step 3:
[1270] The server runs the received voice data through a voice recognition engine and converts it into text data. At this point, the voice data is output as character string data. Specifically, the text data obtained is "What cake do you recommend?" The input is voice data, and the output is text data.
[1271] Step 4:
[1272] At the same time, the server uses an emotion analysis engine to extract emotional data from the voice data. For example, the emotion "friendliness" can be detected from the tone of the user's voice. In this case, the input is voice data, and the output is data indicating the emotional state. Specifically, the emotion analysis engine performs a process to analyze the voice characteristics.
[1273] Step 5:
[1274] The server passes the text data and emotion data to a natural language processing engine, which then generates the optimal response. For example, the response generated might be, "Today's recommendation is shortcake." The input is text data and emotion data, and the output is the generated response data. Specifically, the generative AI model creates a response based on the prompt sentence.
[1275] Step 6:
[1276] The server prepares the generated responses as multiple options and sends them to the smart glasses. In this case, options such as "Today's recommendation is shortcake" are included. The input is the generated response data, and the output is the response options displayed on the smart glasses. Specifically, the option data is displayed on the smart glasses' display.
[1277] Step 7:
[1278] The user checks the display of the smart glasses and selects the most appropriate response. For example, they might select "Today's recommendation is shortcake." In this case, the input is the response options displayed on the display, and the output is the user's selection. Specifically, the user confirms the selection by touching the display.
[1279] Step 8:
[1280] The server passes the selected response to the speech synthesis engine and converts it into speech data. At this time, the text data is output as a speech file. Speech data such as "Today's recommendation is shortcake" is generated. The input is the selected response data, and the output is speech data. Specifically, the speech synthesis engine performs the process of converting text into speech.
[1281] Step 9:
[1282] The terminal transmits the generated voice data through the speaker of the smart glasses. The customer is provided with the reply, "Today's recommendation is shortcake." The input is the generated voice data, and the output is the voice that the customer hears through the speaker. The specific operation is that the voice is played back from the speaker of the smart glasses.
[1283] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1284] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1285] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1286] [Fourth embodiment]
[1287] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1288] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1289] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1290] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1291] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1292] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1293] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1294] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1295] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1296] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1297] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1298] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1299] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1300] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[1301] System Configuration
[1302] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[1303] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[1304] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[1305] Program processing explanation
[1306] Voice input and initial processing
[1307] Users can input voice by speaking into the microphone of their smartphone or tablet. The device captures this voice data and sends it to the server. This includes recording and data transmission functions.
[1308] Analysis of audio data
[1309] The server passes the received voice data to a speech recognition engine, which converts it into text data. It then uses natural language processing (NLP) technology to analyze the text and understand the context. Based on the analysis results, it generates appropriate response candidates.
[1310] Displaying response options and user selection
[1311] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. The selected answer is then sent back to the server via the terminal.
[1312] Text-to-speech and calling
[1313] The server receives the selected text data and converts it into voice data using a speech synthesis engine. This voice data is then sent to the device, which plays it back through the speaker and transmits it to the other party.
[1314] Specific examples
[1315] Example 1: Use in public places
[1316] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates response candidates such as "Progress is going well," "It's behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby communicating the response to the question to the other party.
[1317] Example 2: Use when feeling unwell
[1318] Consider a case where a user is unwell and has difficulty speaking. When an important call comes in, the other party says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. The user selects "I'll reschedule right away, so please wait" from the options presented, and the server converts this into voice data and transmits it to the other party via the device.
[1319] This system functions as a powerful tool for efficient communication even in situations where speech is difficult. In addition, when options are not appropriate, a free text input function is provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[1320] The processing flow will be explained below.
[1321] Step 1:
[1322] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[1323] Step 2:
[1324] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[1325] Step 3:
[1326] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[1327] Step 4:
[1328] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[1329] Step 5:
[1330] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[1331] Step 6:
[1332] Based on the analysis results of the NLP engine, the server uses an AI model (e.g., generative AI) to generate optimal response candidates, such as "The next meeting is at 10:00 AM," "The next meeting is at 3:00 PM," and "The next meeting needs to be rescheduled."
[1333] Step 7:
[1334] The server sends the generated answer candidates to the terminal in JSON format.
[1335] Step 8:
[1336] The device displays the received reply candidates to the user using UI elements (e.g., buttons or lists) to allow the user to select one.
[1337] Step 9:
[1338] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[1339] Step 10:
[1340] The device sends the user's selection to the server via an HTTP request or WebSocket.
[1341] Step 11:
[1342] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[1343] Step 12:
[1344] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[1345] Step 13:
[1346] The terminal plays back the received audio data so that the user can hear the content.
[1347] Step 14:
[1348] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[1349] Step 15:
[1350] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[1351] Step 16:
[1352] The terminal transmits the input text data to the server.
[1353] Step 17:
[1354] The server runs the free text through a speech synthesis engine and converts it into voice data.
[1355] Step 18:
[1356] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[1357] Example 1
[1358] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1359] Currently, many people face situations and environments that make it difficult to communicate through voice. For example, it is difficult to communicate verbally in public places or when feeling unwell. There is a need for a system that can solve these problems and enable everyone to communicate easily.
[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1361] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting the selected response into voice data]; and [means for transmitting the voice data]. This provides a system that performs a series of processes from capturing voice data to outputting the response audibly, enabling effective communication even in situations where speaking is difficult.
[1362] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[1363] "Text data" is data in the form of a character string generated by analyzing voice data.
[1364] A "server" is a computer system that processes voice and text data over a network.
[1365] A "terminal" is an electronic device that a user uses to input voice and receive information.
[1366] A "voice recognition engine" is a technology or system that converts voice data into text data.
[1367] "Natural language processing technology" is a technology that analyzes text data, understands the context, and generates appropriate responses.
[1368] A "generative AI model" is an artificial intelligence algorithm that generates appropriate responses from text data.
[1369] A "speech synthesis engine" is a technology or system that converts text data into speech data.
[1370] A "user interface" refers to a screen or operating means for displaying information to a user on a terminal and accepting input.
[1371] A "prompt" is an instruction given to a generative AI model that is used to generate appropriate response candidates.
[1372] "Network communication" refers to the Internet and other internal and external data communication means for sending and receiving data.
[1373] "Real-time" refers to a processing method that minimizes delays and provides immediate response.
[1374] The present invention relates to a system for assisting voice communication, providing an effective solution for situations or environments where speech is difficult. The system uses a smartphone, tablet, or other electronic device to convert voice data into text data, generate responses based on the context, and output the responses as voice data.
[1375] The terminal receives voice data input by the user and transmits it to a server, which then converts the voice data into text data using voice recognition technology. For example, a smartphone, tablet, or dedicated recording device can be used to acquire this voice data. The acquired voice data is then transmitted to the server using a network communication means.
[1376] The server receives the voice data and instantly converts it into text using a speech recognition engine such as the Google Speech-to-Text API. It then analyzes the text using natural language processing technology to generate the most appropriate response based on the context. Natural language processing engines such as Spacy and BERT are used for the analysis.
[1377] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. A speech synthesis engine such as Amazon Polly is used for the speech synthesis. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[1378] Specific examples
[1379] Example 1: Use in public places
[1380] Suppose a user is participating in an important meeting on the subway and is asked, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server uses a voice recognition engine to convert the voice data into text data and analyzes it. It then generates response candidates such as "Progress is going well" or "We're behind schedule, but we're working on it." The user selects "Progress is going well," and the server converts this into voice data and sends it to the device. Finally, the device plays back this voice data and communicates it to the other party.
[1381] Example 2: Use when feeling unwell
[1382] Suppose a user is feeling unwell and has difficulty speaking, but receives an important call. If the other party says, "I'd like to reschedule our next meeting," the device captures the voice data and sends it to the server. The server converts the voice data into text data and generates appropriate reply candidates. For example, a reply candidate such as "I'll reschedule right away, so please wait." The user selects "I'll reschedule right away, so please wait," and the server converts this into voice data and sends it to the device. The device then plays this voice data and conveys it to the other party.
[1383] Prompt Sentence Examples
[1384] Prompt: Transcribe the following audio data into text and generate an appropriate response.
[1385] Audio data: "I'd like to schedule our next meeting. When would be convenient for you?"
[1386] This system is a powerful tool for efficient communication even in situations where speech is difficult, greatly improving flexibility and convenience. In addition, when the options are not appropriate, a free text input function is provided, allowing users to enter responses in their own words and play them back as audio data.
[1387] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1388] Step 1:
[1389] A user speaks into the microphone of a smartphone or tablet to input voice. The device captures this voice data in real time. Specifically, the device's recording function is used to convert the voice signal collected from the microphone into digital voice data. The input is the user's voice, and the output is digital voice data.
[1390] Step 2:
[1391] The device sends the captured audio data to the server. Using the network communication function, the data is sent via a communication protocol such as an HTTP request or WebSocket. The input is the captured digital audio data, and the output is a transmission confirmation message to the server.
[1392] Step 3:
[1393] The server passes the received voice data to a voice recognition engine and converts it into text data. This conversion is performed using the Google Speech-to-Text API or similar. The server passes the voice data to the voice recognition engine and receives the returned text data. The input is voice data and the output is text data.
[1394] Step 4:
[1395] The server passes the converted text data to a natural language processing (NLP) engine to analyze the context of the text. NLP techniques such as Spacy and BERT are used for this analysis. The server passes the text data to the engine and receives the analysis results in return. The input is the text data, and the output is the analysis results and contextual information.
[1396] Step 5:
[1397] The server uses the parsed text data to generate appropriate reply candidates using a generative AI model (e.g., OpenAI GPT-3). The server passes the prompt sentence and the analysis result to the generative AI model and obtains the reply candidates that are generated. The input is the prompt sentence and the analysis result, and the output is multiple reply candidates.
[1398] Step 6:
[1399] The server sends the generated answer candidates to the terminal. Data is sent quickly using a communication protocol. The input is the generated answer candidates, and the output is a transmission confirmation message to the terminal.
[1400] Step 7:
[1401] The device displays the received reply candidates in a user interface. It uses a UI library (e.g., React Native or Flutter) to display the candidate list on a display or touchscreen. The input is the generated reply candidates, and the output is the displayed list of options.
[1402] Step 8:
[1403] The user selects the most appropriate answer from the presented answer candidates. The selection is made by the user's touch or click. The input is the displayed list of options, and the output is the user's selection.
[1404] Step 9:
[1405] The terminal sends the user's selection to the server. The selection result is sent using network communication. The input is the user's selection result, and the output is a transmission confirmation message to the server.
[1406] Step 10:
[1407] The server passes the selected text data to a speech synthesis engine (e.g., Amazon Polly) and converts it into speech data. The speech synthesis engine converts the text data into speech waveforms and converts them into a playable format. The input is the selected text data, and the output is speech data.
[1408] Step 11:
[1409] The server sends the generated voice data to the terminal. The data is sent quickly using a communication protocol. The input is the generated voice data, and the output is a transmission confirmation message to the terminal.
[1410] Step 12:
[1411] The device plays the received audio data. The audio is output through the built-in speaker or connected earphones. A multimedia library (e.g., AVFoundation or MediaPlayer) is used for playback. The input is the generated audio data, and the output is the played audio.
[1412] (Application example 1)
[1413] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1414] In autonomous vehicles, there is a demand for a system that allows drivers to safely and efficiently operate various functions, make emergency contacts, and use infotainment functions using only voice. Furthermore, there is a lack of a mechanism that allows smooth communication and maximizes the use of autonomous vehicle functions even when the driver cannot speak. Therefore, the present invention aims to solve these problems by providing a system that handles everything from voice data acquisition to response generation and emergency response.
[1415] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1416] In this invention, the server includes: [means for acquiring voice data]; [means for converting voice data into text data]; [means for analyzing the text data and generating an appropriate response]; [means for displaying options for the generated responses]; [means for converting a selected response into voice data]; [means for transmitting the voice data]; [means for recognizing voice commands in an autonomously controlled vehicle and executing driving behaviors and settings]; [means for detecting an emergency and automatically contacting an emergency contact]; and [means for playing music, reading news, and checking vehicle status by voice as infotainment functions.] This enables the driver to safely and efficiently operate an autonomously driven vehicle via voice input, automatically take countermeasures in an emergency, and further utilize infotainment functions.
[1417] The "means for acquiring voice data" is a function for capturing voice uttered by a user using a microphone of an electronic device and transmitting the voice data to a server.
[1418] The "means for converting voice data into text data" is a function for converting acquired voice data into text data using voice recognition technology.
[1419] The "means for analyzing text data and generating an appropriate response" is a function that analyzes the converted text data using natural language processing technology and generates an appropriate response according to the context.
[1420] The "means for displaying generated reply options" is a function that displays multiple reply options generated by the server to the user on the terminal.
[1421] The "means for converting the selected response into voice data" is a function for converting the text-format response selected by the user into voice data using voice synthesis technology.
[1422] "Means for transmitting voice data" is a function that transmits converted voice data to the other party through a speaker.
[1423] "Means for recognizing voice commands within an automatically controlled vehicle and executing driving behaviors or settings" refers to a function that recognizes voice commands issued by a user within the vehicle and executes driving behaviors or setting changes of the vehicle accordingly.
[1424] "Means of detecting an emergency and automatically contacting emergency contacts" is a function that detects an emergency (e.g., accident or illness) while driving, and automatically contacts registered emergency contacts.
[1425] "Means for playing music, reading news, and checking vehicle status by voice as infotainment functions" refers to a function that uses voice input to execute infotainment functions such as playing music, reading news, and checking vehicle status.
[1426] This invention is a system that supports voice communication and is applied to autonomous vehicles. It is a system that covers everything from voice data acquisition and analysis to response generation and emergency response. The system mainly operates via a terminal and a server.
[1427] The terminal has the role of acquiring voice data emitted by the user. Specifically, it captures the voice data using a microphone built into a smartphone, tablet, or vehicle and transmits it to a server.
[1428] The server performs a series of processes, including the following steps:
[1429] 1. Converting audio data to text data:
[1430] Using voice recognition technology, the server converts the voice data received from the device into text data. To do this, the server uses various voice recognition software (e.g., Google Speech-to-Text API, etc.).
[1431] 2. Text data analysis and response generation:
[1432] Leverage natural language processing techniques to analyze text data and generate appropriate responses based on the context, for example, using generative AI models (e.g., OpenAI's GPT-3) to generate responses based on prompts.
[1433] 3. Displaying generated answer choices:
[1434] The generated multiple reply candidates are sent to the terminal, and the user selects the most appropriate reply on the terminal.
[1435] 4. Converting selected responses to audio data:
[1436] The server converts the selected text response into audio using a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API).
[1437] 5. Transmission of voice data:
[1438] The converted audio data is sent back to the device and played through the speaker.
[1439] Specific examples
[1440] Navigation function in autonomous vehicles
[1441] When a user says, "Tell me the shortest route," the device captures the voice data and sends it to the server. The server converts the voice data into text data and sends the following prompt sentence to the generative AI model:
[1442] Example prompt sentence:
[1443] "What is the shortest route from my current location to my destination?"
[1444] The generated multiple answer candidates (e.g., "Turn right after 3.2 km" or "Route A is the shortest route") are displayed on the user's device. The user selects the appropriate answer (e.g., "Route A is the shortest route"), and the server converts it into voice data. The device plays the converted voice data, guiding the user to the shortest route.
[1445] Emergency response features
[1446] If a user feels unwell while driving and instructs the device to "make an emergency call," the device captures the voice data and sends it to the server. The server converts the voice data into text data and automatically calls the emergency contact. The device also provides voice information about the vehicle's location and the user's condition. For example, information such as "The driver is currently feeling unwell. Their location is..." is automatically sent to the contact.
[1447] This will enable a variety of functions to be realized through voice commands within self-driving vehicles, improving safety and convenience.
[1448] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1449] Step 1:
[1450] Acquiring audio data
[1451] Input: Voice command from the user
[1452] Specific behavior:
[1453] The user issues a voice command in the vehicle, and the device (a smartphone, tablet, or a microphone built into the vehicle) captures the voice data.
[1454] Step 2:
[1455] Sending voice data to the server
[1456] Input: Captured audio data
[1457] Output: Audio data sent to the server
[1458] Specific behavior:
[1459] The device sends the captured audio data to a server using an internet connection.
[1460] Step 3:
[1461] Converting audio data to text data
[1462] Input: Audio data sent to the server
[1463] Output: Converted text data
[1464] Specific behavior:
[1465] Using speech recognition technology (such as Google Speech-to-Text API), the server converts the received voice data into text data. This process uses a speech recognition engine.
[1466] Step 4:
[1467] Text data analysis and response generation
[1468] Input: Converted text data
[1469] Output: Generated answer candidates
[1470] Specific behavior:
[1471] The server analyzes the text data using natural language processing technology to understand its context, and uses a generative AI model (such as OpenAI GPT-3) to generate appropriate prompts and create candidate responses.
[1472] Step 5:
[1473] Server sending of possible replies
[1474] Input: Generated answer candidates
[1475] Output: Device where possible responses are displayed
[1476] Specific behavior:
[1477] The server then sends the generated reply candidates to the terminal.
[1478] Step 6:
[1479] Viewing and selecting possible responses
[1480] Input: Answer candidates sent by the server
[1481] Output: Selected response
[1482] Specific behavior:
[1483] The device displays possible responses to the user, who then selects the response that they think is most appropriate.
[1484] Step 7:
[1485] Converting selected responses to audio data
[1486] Input: Selected response (text format)
[1487] Output: Audio data
[1488] Specific behavior:
[1489] The server converts the selected text response into audio data using speech synthesis technology (e.g., Google Text-to-Speech API).
[1490] Step 8:
[1491] Sending and playing audio data
[1492] Input: Converted audio data
[1493] Output: Audio data transmitted through the speaker
[1494] Specific behavior:
[1495] The server then sends the final generated voice data to the device, which then plays the data through a speaker and transmits it to the other party.
[1496] By following these steps, users will be able to use voice commands to perform various operations within a self-driving vehicle.
[1497] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1498] The present invention relates to a system that assists voice communication, providing an effective solution, particularly for situations or environments where speech is difficult. This system uses smartphones, tablets, and other electronic devices to convert voice data into text data, generate responses based on the context, and output the responses as voice data. Furthermore, by combining it with an emotion engine that recognizes emotions from the user's voice, it is possible to generate more appropriate and human-like responses.
[1499] System Configuration
[1500] The terminal receives voice data input by the user and transmits it to the server, which then converts the voice data into text data using voice recognition technology.
[1501] The server receives the voice data and instantly converts it into text data, which is then analyzed using natural language processing technology to generate the most appropriate response based on the context.
[1502] The emotion engine recognizes the user's emotions from the voice data and provides this emotion information to the server, which then adjusts the generated responses based on this emotion information to generate more appropriate response candidates.
[1503] The generated reply candidates are sent to the device and displayed to the user. The user selects the best reply from the presented candidates and sends their selection to the server. The server receives the selected reply and converts it into voice data using speech synthesis technology. Finally, this voice data is sent to the device and transmitted to the other party through the speaker.
[1504] Program processing explanation
[1505] Voice input and initial processing
[1506] A user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?" The device captures the user's voice and saves the audio data. The device then sends the audio data to the server.
[1507] Voice data analysis and emotion recognition
[1508] The server receives the voice data and converts it into text using a speech recognition engine. The converted text data is then passed to a natural language processing (NLP) engine to analyze the context and intent. At the same time, the server uses an emotion engine to recognize the user's emotions. For example, it detects emotions such as tension in the user's voice.
[1509] Generating responses that take emotions into account
[1510] The server uses an AI model to generate optimal response candidates based on the results of NLP analysis and the recognition results of the emotion engine. By taking into account the information from the emotion engine, more appropriate responses can be tailored. For example, if the user is nervous, a gentle response such as "Relax, the next meeting is at 10:00 AM" will be generated.
[1511] Displaying response options and user selection
[1512] The generated answer candidates are sent from the server to the terminal and displayed as multiple options to the user. The user selects the most appropriate answer from these options. For example, "The next meeting is at 10:00 AM."
[1513] Text-to-speech and calling
[1514] The server passes the selected text data to a speech synthesis engine, converts it into voice data, and sends the generated voice data to the device, which plays it back through the speaker and communicates the response to the other party.
[1515] Specific examples
[1516] Example 1: Use in public places
[1517] Imagine a user is attending an important meeting on the subway. During the meeting, the other party asks, "Please tell me about the current progress." The device captures this question as voice data and sends it to the server. The server converts the question into text data, analyzes it, and generates candidate responses. At the same time, if the emotion engine detects tension in the user's voice, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," converts it into voice data, and plays it back, thereby conveying the answer to the question to the other party.
[1518] Example 2: Use when feeling unwell
[1519] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice data and sends it to the server. The server converts it into text data and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[1520] This system serves as a powerful tool for efficient communication even in situations where speech is difficult. It also takes into account the user's emotions to provide more appropriate and human-like responses. A free text input function is also provided, allowing users to input responses in their own words and play them back as audio data, greatly improving flexibility and convenience.
[1521] The processing flow will be explained below.
[1522] Step 1:
[1523] The user speaks into the microphone of their smartphone or tablet. For example, they might say, "When is the next meeting?"
[1524] Step 2:
[1525] The device captures the user's voice with a microphone and temporarily stores it as voice data in storage.
[1526] Step 3:
[1527] The device sends audio data to the server via HTTP requests or WebSockets in a common audio format (e.g., WAV or MP3).
[1528] Step 4:
[1529] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine uses a general voice recognition API (e.g., a voice recognition cloud service).
[1530] Step 5:
[1531] The server passes the text data returned by the speech recognition engine to a natural language processing (NLP) engine, which analyzes the text to understand its context and intent.
[1532] Step 6:
[1533] At the same time, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine detects emotions such as nervousness, fatigue, or happiness from the tone, speed, and volume of the voice.
[1534] Step 7:
[1535] The server uses an AI model (e.g., generative AI) to generate optimal reply candidates based on the analysis results of the NLP engine and the recognition results of the emotion engine. Taking into account the information from the emotion engine, the server adjusts the reply according to the user's current emotional state. For example, if the user is nervous, it generates a reply that will make them feel more relaxed.
[1536] Step 8:
[1537] The server sends the generated answer candidates to the terminal in JSON format.
[1538] Step 9:
[1539] The device displays the received reply candidates to the user. The reply candidates are displayed using UI elements (e.g., buttons and lists) to make it easy for the user to select one.
[1540] Step 10:
[1541] The user selects the best answer from the provided answer options, for example, "The next meeting is at 10:00 AM."
[1542] Step 11:
[1543] The device sends the user's selection to the server via an HTTP request or WebSocket.
[1544] Step 12:
[1545] The server passes the received selected response text data to a speech synthesis engine and converts it into voice data. The speech synthesis engine uses a general-purpose speech synthesis API (e.g., a speech synthesis cloud service).
[1546] Step 13:
[1547] The server sends the generated audio data to the terminal in an audio file format (e.g., MP3).
[1548] Step 14:
[1549] The terminal plays back the received audio data so that the user can hear the content.
[1550] Step 15:
[1551] The terminal immediately plays back the confirmed voice data and conveys the response to the other party.
[1552] Step 16:
[1553] If the user does not find a suitable response from the options presented, they can use the free text input screen to directly input text, for example, "The next meeting has not yet been scheduled."
[1554] Step 17:
[1555] The terminal transmits the input text data to the server.
[1556] Step 18:
[1557] The server runs the free text through a speech synthesis engine and converts it into voice data.
[1558] Step 19:
[1559] The server transmits the generated voice data to the terminal, which then plays it back and transmits it to the other party.
[1560] By combining this system with an emotion engine, it becomes possible to respond in a way that takes into account the user's emotional state, resulting in more natural and friendly communication. Furthermore, by using an emotion engine, responses that suit the user's state of mind can be provided, which is expected to have the effect of improving user satisfaction.
[1561] Example 2
[1562] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1563] Conventional voice communication systems have difficulty communicating effectively in situations or environments where speech is difficult. Furthermore, they lack the functionality to provide responses that take the user's emotions into account, making it difficult to generate more appropriate and human-like responses. Furthermore, the interface for selecting the most appropriate response from multiple candidate responses is inconvenient, which can hinder smooth communication.
[1564] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1565] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and recognizing the user's emotions, and means for generating appropriate reply candidates based on the analysis results and the emotional information. This makes it possible to convert voice data into text data and generate replies that take the user's emotions into consideration.
[1566] "Voice data" refers to acoustic data generated by a user's speech.
[1567] "Acquiring" refers to capturing and recording audio data using the device's microphone.
[1568] A "server" refers to a computer system that receives voice data from a terminal via a network and processes the data.
[1569] "Send" refers to the transfer of collected voice data by the terminal to the server.
[1570] "Text data" refers to data obtained by converting voice data into character information.
[1571] "Converting" refers to the process of changing audio data into text data or text data into audio data.
[1572] "Analyzing" refers to interpreting the meaning, intent, and context of text data using natural language processing techniques.
[1573] "Emotion recognition" refers to detecting a user's emotional state from speech or text data.
[1574] "Generating reply candidates" refers to creating appropriate responses based on text data and emotional information.
[1575] "Presenting to the user as options" refers to displaying the generated reply candidates as multiple options on the user's device.
[1576] "Speech synthesis" refers to the technology of converting text data into voice data.
[1577] "Transmitting audio" refers to playing the generated audio data through the device's speaker.
[1578] This invention is a system that allows users to smoothly communicate through speech, particularly in situations and environments where speech is difficult. The system converts speech into text and then analyzes the text to generate responses. It also recognizes the user's emotional state and reflects it in the responses, enabling more appropriate and human-like communication.
[1579] System Configuration
[1580] This system is realized mainly using the following hardware and software.
[1581] Hardware
[1582] 1. Device: An electronic device such as a smartphone or tablet that includes a microphone to capture the user's voice and a speaker to play back the response. Examples of smartphones include the Apple iPhone and Samsung Galaxy series.
[1583] 2. Microphone: Reliably captures the user's voice through a high-sensitivity microphone, such as the Shure MV88.
[1584] software
[1585] 1. Speech recognition engine: Software for converting voice data into text data. For example, we use the Google Speech-to-Text API.
[1586] 2. Natural Language Processing (NLP) Engine: Software that analyzes the converted text data to understand intent and context. We use OpenAI GPT-3 for this.
[1587] 3. Emotion Engine: Software for recognizing emotions from user voice and text data. As an example, we will use IBM Watson Tone Analyzer.
[1588] 4. Speech synthesis engine: Software that converts selected text data into speech data. An example is Amazon Polly.
[1589] System Operation Overview
[1590] When a user speaks into the device, the device's microphone captures the voice and saves it as audio data. This audio data is sent to the server using a communication protocol. The server then converts the audio data into text data using the Google Speech-to-Text API. It then uses OpenAI GPT-3 to analyze the context and intent of the text data. At the same time, it uses IBM Watson Tone Analyzer to recognize the user's emotions. Based on the analysis results and emotional information, OpenAI GPT-3 generates appropriate response candidates. These response candidates are sent from the server to the device and displayed to the user as multiple options. The user selects the most appropriate response from the displayed options and sends it to the server. The server then converts the selected text data into audio data using Amazon Polly and sends it back to the device. Finally, the audio data is played back through the device's speaker.
[1591] Specific examples
[1592] Example 1: Use in public places
[1593] Consider a situation where a user is attending an important meeting on the subway. During the meeting, someone asks, "Please tell me about the current progress." The user speaks the question into their smartphone, and the voice data is sent to the server. The server converts the voice into text, analyzes it, and generates candidate responses. If the emotion engine detects that the user is tense, the candidate responses include considerate responses such as "Please relax. Progress is going well," "We're behind schedule, but we're working on it," and "I'll report more details at the next meeting." The user selects "Progress is going well," and the voice data is played.
[1594] Example 2: Use when feeling unwell
[1595] Consider a case where a user is unwell and finds it difficult to speak. When an important call comes in, the caller says, "I'd like to reschedule our next meeting." The device captures the voice and sends it to the server. The server converts the voice into text and generates appropriate reply candidates. If the emotion engine detects the user's fatigue, the reply candidates include considerate responses such as "Don't push yourself. We'll reschedule your next meeting" or "We'll reschedule right away, so please wait." The user selects "We'll reschedule right away, so please wait," and the server converts this into voice data and transmits it to the caller via the device.
[1596] Prompt Sentence Examples
[1597] The following is an example of a prompt sentence to input to the generative AI model:
[1598] "When is the next meeting? User sentiment: Nervous Generate possible responses."
[1599] This prompt helps the system achieve smooth communication even in situations where the user has difficulty speaking.
[1600] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1601] Step 1:
[1602] The user speaks into the microphone on their smartphone or tablet.
[1603] The input is the user's speech. The output is the audio data captured by the device. Specifically, the user speaks, "When is the next meeting?" The device's high-sensitivity microphone (e.g., Shure MV88) picks up the voice and temporarily stores the audio data in the device's built-in storage.
[1604] Step 2:
[1605] The device compresses the captured audio data and sends it to the server using the HTTPS protocol.
[1606] The input is the audio data acquired in step 1. The output is compressed audio data sent to the server. Specifically, the device compresses the audio data for efficient transmission and sends it to the server encrypted with SSL / TLS.
[1607] Step 3:
[1608] The server converts the received voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text API).
[1609] The input is the voice data sent from the device. The output is the converted text data. Specifically, the server sends the voice data to the Google Speech-to-Text API, which returns the text data. The converted text data is then sent to the next processing step within the server.
[1610] Step 4:
[1611] The server passes the text data to a natural language processing (NLP) engine (e.g., OpenAI GPT-3) to analyze the context and intent.
[1612] The input is the text data generated in step 3. The output is the analyzed context information. Specifically, the server sends the text data to OpenAI GPT-3, and obtains information about the context and intent as the analysis result. This analysis result is used for subsequent processing.
[1613] Step 5:
[1614] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions.
[1615] The input is the text data and speech parameters generated in step 3. The output is the recognized emotion information. Specifically, the server sends the speech and text parameters to IBM Watson Tone Analyzer to recognize the user's emotion (e.g., tension, fatigue). This information is used to generate a response.
[1616] Step 6:
[1617] The server uses a generative AI model (e.g., OpenAI GPT-3) based on the analysis results and emotional information to generate appropriate response candidates.
[1618] The input is the context analysis result obtained in step 4 and the emotional information obtained in step 5. The output is multiple response candidates. Specifically, the server sends the analysis results and emotional information as prompts to OpenAI GPT-3, which then generates appropriate response candidates.
[1619] Step 7:
[1620] The server sends the generated answer candidates to the terminal in JSON format.
[1621] The input is the candidate answers generated in step 6. The output is the candidate answer data sent to the device. Specifically, the server formats the candidate answers in JSON format and sends them to the device.
[1622] Step 8:
[1623] The terminal displays the reply candidates on a user interface.
[1624] The input is candidate reply data sent from the server. The output is multiple candidate replies displayed on the user interface. Specifically, the device analyzes the candidate replies and displays options such as "Next meeting is at 10:00 AM," "Relax, next meeting is at 10:00 AM," and "Checking next meeting time" on the touchscreen.
[1625] Step 9:
[1626] The user selects the most appropriate option from the displayed options and taps it.
[1627] The input is the reply candidates displayed on the terminal. The output is the selected reply data. In concrete terms, the user selects "Relax, the next meeting is at 10:00 AM," and the selection information is saved on the terminal.
[1628] Step 10:
[1629] The terminal transmits the selected response data to the server.
[1630] The input is the response data selected by the user. The output is the selected data sent to the server. Specifically, the terminal encrypts the selected data again using SSL / TLS and sends it to the server.
[1631] Step 11:
[1632] The server passes the selected response to a speech synthesis engine (e.g., Amazon Polly) and converts it into voice data.
[1633] The input is the selected response data sent from the device. The output is the generated voice data. Specifically, the server sends the selected response to Amazon Polly, which generates the voice data.
[1634] Step 12:
[1635] The server transmits the generated voice data to the terminal.
[1636] The input is the voice data generated in step 11. The output is the voice data sent to the terminal. In concrete operation, the server sends the generated voice data back to the terminal.
[1637] Step 13:
[1638] The terminal plays the received audio data through a speaker.
[1639] The input is the voice data sent from the server. The output is the voice response that is transmitted to the other party. Specifically, the device decodes the voice data and plays it using the built-in speaker (e.g., Bose SoundLink Mini) to transmit the response to the other party.
[1640] (Application example 2)
[1641] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1642] Conventional voice communication assistance systems convert voice data into text data and generate responses, but they are not sufficient in generating appropriate responses that take the user's emotional state into account or in presenting responses quickly using visual devices. Therefore, there is a need for a system that provides more human-like and considerate responses in situations and environments where voice communication is difficult.
[1643] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting voice data to text data, means for analyzing the text data and generating an appropriate response, means for displaying the generated response options, means for converting the selected response into voice data, means for transmitting the voice data, means for analyzing emotional data and adjusting the response, and means for displaying candidate responses on a visual device. This enables the generation of appropriate and prompt responses that take the user's emotional state into consideration.
[1644] "Voice data" refers to data in which the user's speech is recorded as an electronic signal.
[1645] "Text data" is data obtained by converting voice data into a character string format.
[1646] "Analysis" is the process of understanding the content of text data and grasping its context and intent.
[1647] A "reply" is the content of a response generated based on the analysis results.
[1648] "Choices" refer to multiple possible responses presented to the user.
[1649] "Display" refers to showing the generated responses or options on a visual device.
[1650] "Speech synthesis" is a technology that converts text data into voice data.
[1651] "Transmission" is a function for transmitting generated voice data to the outside.
[1652] "Emotion data" is data of emotions estimated from the user's voice, facial expressions, movements, etc.
[1653] A "visual device" is a device for visually presenting information to a user, and in this context refers to smart glasses and the like.
[1654] This invention relates to a system that supports voice communication, particularly to a customer service support application in brick-and-mortar stores using visual devices. This system converts voice data into text data and generates appropriate responses based on emotion data through collaboration between smart glasses and a server. The generated responses are displayed on the smart glasses' display and transmitted as voice data.
[1655] Components and hardware / software used
[1656] 1. Smart Glasses:
[1657] A device worn by the user, equipped with a microphone and a display, that captures voice data and provides a visual response.
[1658] 2. Server:
[1659] It processes voice data, generates text data, analyzes emotional data, generates responses, and synthesizes voice.
[1660] The software used includes a speech recognition engine, a natural language processing engine, a sentiment analysis engine, and a speech synthesis engine.
[1661] Program processing explanation
[1662] 1. Acquire audio data:
[1663] The user speaks a customer question into the microphone of the smart glasses, for example, "What cake do you recommend?"
[1664] The smart glasses capture this voice data and send it to a server.
[1665] 2. Convert to text data:
[1666] The server uses a voice recognition engine to convert the received voice data into text data.
[1667] The converted text data may be a string such as "What cake do you recommend?"
[1668] 3. Emotional Data Analysis:
[1669] The server uses an emotion analysis engine to extract emotion data from the user's speech.
[1670] For example, friendliness is detected from the user's tone of voice.
[1671] 4. Generate a response:
[1672] The natural language processing engine generates optimal responses based on text data and emotional data.
[1673] A possible reply such as "Today's recommendation is shortcake" is generated.
[1674] 5. Visual display of response:
[1675] The generated response is sent from the server to the smart glasses and displayed on the display.
[1676] The user checks the displayed reply candidates and selects an appropriate reply.
[1677] 6. Conversion to voice data and transmission:
[1678] The server passes the selected response to a speech synthesis engine and converts it into voice data.
[1679] The generated voice data is transmitted to the customer through the speaker of the smart glasses.
[1680] Specific examples
[1681] Consider a scenario where a cafe staff member is wearing smart glasses during a busy lunchtime. When a customer asks, "What cake do you recommend?", the staff member captures the voice data through the microphone in the smart glasses. The server analyzes the voice as text data and generates a response that takes emotional data into account. A response such as "Today's recommendation is shortcake" is displayed on the smart glasses' display, and the staff member relays this to the customer.
[1682] Prompt Sentence Examples
[1683] Examples of prompts for a generative AI model include "Emotion-recognizing customer service assistant at a cafe" and "Generate a friendly response when a customer asks, 'What cake do you recommend?'"
[1684] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1685] Step 1:
[1686] The user speaks into the microphone of the smart glasses. For example, they say, "What cake do you recommend?" The smart glasses capture this voice data and store it in the device's memory. At this point, the input is the user's voice data, and the output is the saved audio file.
[1687] Step 2:
[1688] The device sends the acquired voice data to the server. The input is the saved voice file, and the output is the voice data sent to the server. Specifically, the smart glasses upload the voice data to the server.
[1689] Step 3:
[1690] The server runs the received voice data through a voice recognition engine and converts it into text data. At this point, the voice data is output as character string data. Specifically, the text data obtained is "What cake do you recommend?" The input is voice data, and the output is text data.
[1691] Step 4:
[1692] At the same time, the server uses an emotion analysis engine to extract emotional data from the voice data. For example, the emotion "friendliness" can be detected from the tone of the user's voice. In this case, the input is voice data, and the output is data indicating the emotional state. Specifically, the emotion analysis engine performs a process to analyze the voice characteristics.
[1693] Step 5:
[1694] The server passes the text data and emotion data to a natural language processing engine, which then generates the optimal response. For example, the response generated might be, "Today's recommendation is shortcake." The input is text data and emotion data, and the output is the generated response data. Specifically, the generative AI model creates a response based on the prompt sentence.
[1695] Step 6:
[1696] The server prepares the generated responses as multiple options and sends them to the smart glasses. In this case, options such as "Today's recommendation is shortcake" are included. The input is the generated response data, and the output is the response options displayed on the smart glasses. Specifically, the option data is displayed on the smart glasses' display.
[1697] Step 7:
[1698] The user checks the display of the smart glasses and selects the most appropriate response. For example, they might select "Today's recommendation is shortcake." In this case, the input is the response options displayed on the display, and the output is the user's selection. Specifically, the user confirms the selection by touching the display.
[1699] Step 8:
[1700] The server passes the selected response to the speech synthesis engine and converts it into speech data. At this time, the text data is output as a speech file. Speech data such as "Today's recommendation is shortcake" is generated. The input is the selected response data, and the output is speech data. Specifically, the speech synthesis engine performs the process of converting text into speech.
[1701] Step 9:
[1702] The terminal transmits the generated voice data through the speaker of the smart glasses. The customer is provided with the reply, "Today's recommendation is shortcake." The input is the generated voice data, and the output is the voice that the customer hears through the speaker. The specific operation is that the voice is played back from the speaker of the smart glasses.
[1703] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1704] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1705] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1706] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1707] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1708] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1709] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1710] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1711] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1712] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1713] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1714] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1715] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1716] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1717] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1718] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1719] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1720] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1721] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1722] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1723] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1724] The following is further disclosed regarding the above embodiment.
[1725] (Claim 1)
[1726] [Means for acquiring audio data;
[1727] [Means for converting voice data into text data;
[1728] [Means for analyzing text data and generating appropriate responses;
[1729] [means for displaying the generated response options; and
[1730] [means for converting the selected response into audio data;
[1731] [A system including a means for transmitting audio data.
[1732] (Claim 2)
[1733] The system according to claim 1, further comprising means for a user to directly input text and convert the input text into voice data.
[1734] (Claim 3)
[1735] The system of claim 1, further comprising means for presenting the generated responses to the user as multiple options.
[1736] "Example 1"
[1737] (Claim 1)
[1738] [Means for acquiring audio data;
[1739] [Means for converting voice data into text data;
[1740] [Means for analyzing text data and generating appropriate responses;
[1741] [means for displaying the generated response options; and
[1742] [means for converting the selected response into audio data;
[1743] [A system including a means for transmitting audio data.
[1744] (Claim 2)
[1745] The system according to claim 1, further comprising means for a user to directly input text and convert the input text into voice data.
[1746] (Claim 3)
[1747] The system of claim 1, further comprising means for presenting the generated responses to the user as a plurality of options, accepting the user's selection, and transmitting the selection to the server.
[1748] (Claim 4)
[1749] [The system of claim 1, further comprising means for capturing audio data in real time and transmitting the data to a server via a network.
[1750] (Claim 5)
[1751] [The system according to claim 1, further comprising means for analyzing text data using natural language processing technology and generating optimal reply candidates according to the context.]
[1752] (Claim 6)
[1753] The system according to claim 3, further comprising means for displaying the generated reply candidates on a user interface and accepting a user's selection by a touch operation or a click operation.
[1754] (Claim 7)
[1755] [The system of claim 1, further comprising means for playing the generated audio data through a speaker.
[1756] (Claim 8)
[1757] The system of claim 1 includes a generative AI model that generates appropriate responses from text data using prompt sentences.
[1758] "Application Example 1"
[1759] (Claim 1)
[1760] [Means for acquiring audio data;
[1761] [Means for converting voice data into text data;
[1762] [Means for analyzing text data and generating appropriate responses;
[1763] [means for displaying the generated response options; and
[1764] [means for converting the selected response into audio data;
[1765] [Means for transmitting audio data;
[1766] [Means for recognizing voice commands in an autonomous vehicle and executing driving actions or settings;
[1767] [Means of detecting an emergency and automatically contacting emergency contacts;
[1768] [The system includes infotainment functions such as voice playback of music, news reading, and vehicle status confirmation.
[1769] (Claim 2)
[1770] [The system according to claim 1, wherein a user directly inputs text and converts the input text into voice data.
[1771] (Claim 3)
[1772] [The system of claim 1, wherein the generated responses are presented to the user as multiple options.
[1773] "Example 2: Combining Emotion Engines"
[1774] (Claim 1)
[1775] [Means for acquiring audio data;
[1776] [Means for transmitting the acquired voice data to a server;
[1777] [Means for converting voice data into text data;
[1778] [Means for analyzing text data and recognizing user emotions;
[1779] [Means for generating appropriate response candidates based on analysis results and emotional information,
[1780] [Means for presenting the generated answer candidates to a user as multiple options;
[1781] [means for converting the user-selected response into audio data;
[1782] [A system including a means for transmitting audio data.
[1783] (Claim 2)
[1784] The system according to claim 1, further comprising means for a user to directly input text and convert the input text into voice data.
[1785] (Claim 3)
[1786] The system of claim 1, further comprising means for adjusting candidate replies based on emotion recognition.
[1787] "Application example 2 when combining emotion engines"
[1788] (Claim 1)
[1789] [Means for acquiring audio data;
[1790] [Means for converting voice data into text data;
[1791] [Means for analyzing text data and generating appropriate responses;
[1792] [means for displaying the generated response options; and
[1793] [means for converting the selected response into audio data;
[1794] [Means for transmitting audio data;
[1795] [Means of analyzing emotional data and adjusting responses;
[1796] [A system including means for displaying possible responses on a visual device.
[1797] (Claim 2)
[1798] The system according to claim 1, further comprising means for a user to directly input text and convert the input text into voice data.
[1799] (Claim 3)
[1800] The system of claim 1, further comprising means for presenting the generated responses to the user as multiple options. [Explanation of symbols]
[1801] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for acquiring audio data; means for converting voice data into text data; means for analyzing the text data and generating an appropriate response; means for displaying the generated response options; means for converting the selected response into voice data; A system including means for transmitting audio data.
2. 2. The system according to claim 1, further comprising means for allowing a user to directly input text and converting the input text into voice data.
3. 10. The system of claim 1, further comprising means for presenting the generated responses to the user as a plurality of options.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A