System
The system addresses user stress by converting voice or text input into automated phone calls, generating dialogue scripts, and summarizing interactions, allowing stress-free information retrieval.
Patent Information
- Application Number
- JP2024119103
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Many individuals, particularly the younger generation, experience stress and discomfort when making phone calls for everyday tasks like reservations or inquiries, leading to a demand for technology that allows information retrieval without direct communication.
A system that converts user input into text data, uses a generative model to generate a dialogue script, synthesizes voice data for real-time interaction, records and summarizes the conversation, and notifies the user of the results, eliminating the need for direct phone calls.
Enables users to obtain necessary information efficiently and stress-free by automating phone interactions, reducing the burden of making calls and providing concise summaries of the conversation outcomes.
Smart Images

Figure 2026018042000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Many people today are not comfortable communicating over the phone. This often leads to stress when making everyday phone calls, such as for making reservations or inquiries. This tendency is particularly pronounced among the younger generation, for whom making a phone call is a significant burden. Given this situation, there is a demand for technology that allows users to obtain the information they need without having to make a direct call. [Means for solving the problem]
[0005] To solve these problems, the present invention provides a means for a user to input what they want to communicate over the phone and convert that input content into text data. Furthermore, by combining a generative model means for receiving text data, analyzing the user's intentions, and generating a dialogue script, the content of a telephone dialogue can be programmed naturally. The generated dialogue script is converted into voice data, and a means for making a call and having a real-time dialogue with the other party is also included. Furthermore, a means is provided for recording the content of the dialogue, converting the recorded voice data into text, and summarizing it, allowing the user to easily check the results. This allows the user to obtain the necessary information without stress, without having to make a call themselves.
[0006] "User" refers to a person who operates the system and needs to make inquiries or reservations by phone.
[0007] "Content to be communicated over the phone" refers to information or requests that the user wishes to convey to the other party over the phone.
[0008] "Input means" refers to an interface that allows a user to input what they want to say over the phone by voice or text.
[0009] "Text data" refers to data that has been converted into a character string from what the user has entered.
[0010] "Means for converting into text data" refers to a process or device that converts the content input by voice by the user into text.
[0011] The "generative model means" refers to an artificial intelligence model for analyzing the contents of a user's input and generating a dialogue script.
[0012] A "dialogue script" refers to a scenario that describes the content of a conversation with a partner, generated based on the user's intentions.
[0013] The "voice synthesis means" refers to a process or device that creates voice data based on the generated dialogue script.
[0014] "Voice data" refers to data that expresses a dialogue script as voice.
[0015] "Means for making phone calls" refers to a process or device that uses voice data to place a call to a party and establish a connection.
[0016] "Means of real-time interaction" refers to a process or device that allows instant communication with another party via telephone.
[0017] "Means for recording conversation content" refers to a process or device that saves the contents of a telephone conversation as audio data.
[0018] "A voice-to-text converter" refers to a process or device that converts recorded voice data into a string of characters.
[0019] The "summarizing means" refers to a process or device that summarizes the converted text data, extracts important information, and summarizes it compactly.
[0020] "Terminal" refers to a device such as a smartphone or computer operated by a user.
[0021] The "notifying means" refers to a process or device that transmits the summarized text data to the terminal and notifies the user. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] The present invention is a system developed to reduce the stress users feel when making phone calls. This system uses AI to make calls on behalf of users, obtain necessary information, and notify the user of the results without the user having to make the calls themselves. The following describes in detail the embodiments of the present invention.
[0044] Input and Initial Processing
[0045] 1. User:
[0046] Users start up a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0047] Example: If a user enters "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0048] 2. Terminal:
[0049] The device receives user input. In the case of voice input, it uses a speech recognition engine to convert the voice data into text data, which is then sent to the server.
[0050] AI processing and phone generation
[0051] 3. Server:
[0052] The server receives text data from the device and passes it to a generative AI model (e.g., a model using natural language processing). The AI model analyzes the user's intent and generates an appropriate dialogue script.
[0053] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0054] 4. Server (Speech synthesis):
[0055] The generated dialogue script is input to a speech synthesis engine to generate voice data that resembles the user's voice. This voice data is then transmitted over the telephone system.
[0056] Real-time dialogue and results notification
[0057] 5. Server:
[0058] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0059] 6. Server (Text Conversion and Summarization):
[0060] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[0061] 7. Notification from the server to the device:
[0062] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[0063] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0064] Specific example explanation
[0065] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[0066] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0067] 2. The device converts the speech into text and sends it to the server.
[0068] 3. The server uses the generative AI model to analyze the input and generate an appropriate dialogue script (e.g., "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability.").
[0069] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[0070] 5. The server communicates with the other party in real time and records the conversation.
[0071] 6. The server converts the recorded audio data into text and creates a summary.
[0072] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0073] In this way, the system of the present invention allows the user to obtain the necessary information without stress, without having to make a phone call himself.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0077] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0078] Step 2:
[0079] The device receives user input. In the case of voice input, the device's voice recognition engine is used to convert the voice data into text data.
[0080] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0081] Step 3:
[0082] The terminal transmits the converted text data to the server.
[0083] Step 4:
[0084] The server receives the text data sent from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intent from the text data.
[0085] Example: Analyze the text data "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0086] Step 5:
[0087] The server generates an interactive script based on the analysis results.
[0088] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0089] Step 6:
[0090] The server inputs the generated dialogue script into a voice synthesis engine to generate voice data that resembles the user's voice.
[0091] Example: The dialogue script generates the following speech data: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0092] Step 7:
[0093] The server sends the voice data to the telephone system, which then places a call to the destination (e.g., a restaurant).
[0094] Step 8:
[0095] The server establishes a telephone connection and transmits the generated voice data to the other party to start the conversation.
[0096] Step 9:
[0097] The server analyzes the response from the other party in real time and uses AI to generate an appropriate response.
[0098] Example: If a restaurant employee replies, "We have a table available at 7:00 PM," the AI generates a response: "Thank you. We'll make it that time then."
[0099] Step 10:
[0100] The server records the entire phone conversation.
[0101] Step 11:
[0102] The server converts the recorded audio data into text.
[0103] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[0104] Step 12:
[0105] The server summarizes the converted text data, extracts important information, and summarizes it compactly.
[0106] Example: The summary you get is "A restaurant reservation has been made for two people at 7:00 PM."
[0107] Step 13:
[0108] The server transmits the summarized text data to the terminal.
[0109] Step 14:
[0110] The text data received by the device is notified to the user via push notification, email, etc.
[0111] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0112] This series of steps allows the user to efficiently obtain the necessary information without having to make a phone call.
[0113] Example 1
[0114] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0115] In modern society, making reservations and inquiries over the phone is common, but it can cause stress and anxiety for many users. In particular, for users who are not good at communicating over the phone or who are busy, there is a demand for a system that can handle phone calls efficiently and easily. Conventional systems require users to make the calls themselves, which is time-consuming and mentally taxing. To solve this problem, the present invention aims to provide an automated telephone answering system that eliminates the need for users to make phone calls.
[0116] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0117] In this invention, the server includes means for inputting the content that the user wants to communicate over the phone, means for converting the input content into text data, and generative model means for receiving the text data, analyzing the intention, and generating a dialogue script. This allows the system to automatically make calls on behalf of the user, obtain the necessary information, and notify the user of the results, without the user having to make the call directly.
[0118] A "user" is a person or individual who uses the system to input what they want to say over the phone.
[0119] A "means" is a device or method used to achieve a particular purpose or function.
[0120] "Text data" refers to character-based information converted from audio data or other formats.
[0121] A "generative model" is an algorithm that uses natural language processing or machine learning to generate sentences that meet a specific purpose from specific input.
[0122] "Speech synthesis" is a technology that generates speech that sounds like human speech based on text data.
[0123] An "information processing device" is an electronic device for receiving and processing data, and typically includes smartphones, tablets, personal computers, etc.
[0124] An "external system" is an object other than the user's system, such as a restaurant or an office.
[0125] This invention is an automated telephone answering system developed to reduce the stress of users when making phone calls and to efficiently obtain the necessary information. The system aims to have users input the information they want to convey over the phone, and then have AI make the call on their behalf, obtain the information, and notify the user of the results.
[0126] Input and Initial Processing
[0127] 1. The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input can be done either by voice or text. Consider the example where the user inputs, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0128] 2. The device receives the user's voice input and converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text API), which is then sent to the server.
[0129] AI processing and phone generation
[0130] 3. The server receives the text data sent from the device and passes it to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the user's intent and generates an appropriate dialogue script. The generated dialogue script includes the following content: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0131] 4. The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data that resembles the user's voice. This voice data is then transmitted through the telephone system.
[0132] Real-time dialogue and results notification
[0133] 5. The server uses the generated voice to make a call and engage in a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0134] 6. The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the important points.
[0135] 7. The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking this notification, the user can easily understand the content and results of the call. For example, a notification may appear on the device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0136] Specific examples
[0137] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[0138] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0139] 2. The device converts the speech into text and sends it to the server.
[0140] 3. The server analyzes the input using a generative AI model and generates an appropriate dialogue script.
[0141] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[0142] 5. The server communicates with the other party in real time and records the conversation.
[0143] 6. The server converts the recorded audio data into text and summarizes the content.
[0144] 7. The server notifies the terminal of the summary, and the user confirms the results.
[0145] An example of a prompt sentence could be, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM. Please let me know the availability." In this way, the user can obtain the necessary information without stress, without having to make a phone call themselves.
[0146] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0147] Step 1:
[0148] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input methods include both voice input and text input. For example, a user might say, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." This will result in either voice data or text data being acquired as input data.
[0149] Step 2:
[0150] The device receives the user's input data. The received voice data is converted into text data using a speech recognition engine such as the Google Speech-to-Text API. During this conversion process, the voice data is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM." Text data is generated as output, and this text data is sent to the server.
[0151] Step 3:
[0152] The server receives the text data sent from the device. The received text data is passed to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the intent of the input text data and generates an appropriate dialogue script. For example, from the input data "I would like to make a restaurant reservation for tomorrow at 7:00 PM," the dialogue script generated is "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability." The dialogue script is obtained as the output.
[0153] Step 4:
[0154] The server inputs the generated dialogue script into a speech synthesis engine (for example, Google Text-to-Speech API). The speech synthesis engine generates voice data that resembles the user's voice based on the text data. This voice data is transmitted through the telephone system. The generated voice data is obtained as output and is used as the transmitted voice.
[0155] Step 5:
[0156] The server uses the generated voice data to make a call. The call is made to an external system, such as a restaurant, and the server conducts the conversation in real time. The AI system within the server generates an appropriate response based on the other party's response and continues the conversation. For example, the AI might say, "Could I please make a restaurant reservation for two people tomorrow at 7:00 PM?" and generate the next response based on the restaurant's response. The entire conversation is recorded within the server. The other party's response is obtained as output.
[0157] Step 6:
[0158] The server converts the recorded voice data into text data. The content is then summarized by a generative AI model. For example, if the restaurant responds, "We can make a reservation for two people at 7:00 PM," the resulting summarized text is, "A restaurant reservation has been made for two people at 7:00 PM." The summarized text is generated as the output.
[0159] Step 7:
[0160] The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking the notification, the user can easily understand the content and results of the call. For example, a notification saying "A restaurant reservation has been made for two people at 7:00 PM" is displayed on the user's smartphone. The notification message is obtained as output.
[0161] (Application example 1)
[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0163] In modern society, users experience a great deal of stress when making reservations or inquiries at physical stores over the phone. This stress increases especially when the call is not answered or when appropriate communication is not possible. There is a need to solve this problem and provide a method that allows users to make reservations and inquiries more easily.
[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0165] In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generation AI model means for receiving the text data, analyzing the intention, and generating a dialogue script; voice synthesis means for converting the dialogue script into voice data; means for making a call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for a user to input a reservation or inquiry for a physical store; means for receiving input related to the physical store, generating a dialogue script, and making a call to collect necessary information; and means for notifying the user of the results of the reservation or inquiry for the physical store. This enables a user to easily make a reservation or inquiry for a physical store without having to make a call themselves.
[0166] definition statement
[0167] "Means for users to input information they want to communicate over the phone" refers to an interface that allows users to input details of reservations, inquiries, etc. by voice or text using a smartphone or other device.
[0168] "Means for converting input content into text data" refers to the process of generating text data from the voice or image input by the user using a voice recognition engine or OCR technology.
[0169] A "generative AI model means" is an algorithm or program that uses a generative AI model (e.g., GPT-4) to generate an appropriate dialogue script based on the user's intentions.
[0170] The "voice synthesis means for converting a dialogue script into voice data" is a process that uses voice synthesis technology (e.g., a TTS engine) to convert text data into voice data.
[0171] "Means for making a call using voice data and having a real-time conversation with the other party" refers to a communication technology that uses the generated voice data in an actual telephone conversation to conduct a real-time conversation with the other party.
[0172] "Means for recording the content of a conversation and converting the recorded voice data into text and summarizing it" refers to the process of recording the content of a conversation, converting the recorded data into text data using voice recognition technology, and creating a summary using AI technology.
[0173] The "means for notifying the user of the summarized text" refers to a technology for delivering the generated summarized text to the user's smartphone or other device as a push notification or message.
[0174] "A means by which a user can input reservations or inquiries for a physical store" is an interface that allows a user to input reservation or inquiry information about a specific store.
[0175] "Means of receiving input about a physical store, generating a dialogue script, and making a phone call to collect the necessary information" refers to the process in which the generative AI model creates a dialogue script based on the information about the physical store entered by the user, and makes a phone call to obtain the necessary information.
[0176] "Means for notifying users of the results of reservations and inquiries at physical stores" refers to technology for notifying users of the results of reservations and inquiries obtained via their terminals.
[0177] MODE FOR CARRYING OUT THE INVENTION
[0178] The present invention is a system that automates the process of making reservations or inquiries about physical stores, reducing stress for users. This system operates through a smartphone application, generates a dialogue script using the generated AI model, and makes phone calls to collect necessary information. The following describes an embodiment of the present invention.
[0179] 1. Basic configuration
[0180] This system consists of the following main hardware and software:
[0181] User terminal (smartphone): A device where a user enters input and receives results. It includes a speech recognition engine and a text input interface.
[0182] Server: Receives user input, generates a dialogue script using a generative AI model, creates voice data using a speech synthesis engine, and processes the call.
[0183] Generative AI model: A technology that generates dialogue scripts based on user requests. An example of this is GPT-4.
[0184] Speech synthesis engine: A technology that converts text data into voice data. An example is Google Text-to-Speech (TTS).
[0185] 2. System processing flow
[0186] A user starts a smartphone application and inputs reservation or inquiry details. For example, a user inputs a prompt such as "I would like to make a restaurant reservation for two people tomorrow at 7:00 PM." This input can be done either by voice or text.
[0187] When voice input is used, the device uses a voice recognition engine such as Google Speech Recognition to convert the voice into text data and then sends the text data to the server.
[0188] The server receives the text data and uses a generative AI model (e.g., GPT-4) to generate a dialogue script based on the user's intention. The generated dialogue script is in the form of "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM."
[0189] The server then uses Google Text-to-Speech (TTS) to convert the generated dialogue script into voice data, and makes an actual call through the telephone system to interact with the physical store in real time.
[0190] During the call process, the server records the conversation, and after the call is over, it converts the audio data back into text data and uses AI technology to create a summary, such as "A restaurant reservation has been made for two people at 7:00 PM."
[0191] The completed summary is sent from the server to the device, and the results are communicated to the user via push notification or message. By checking this notification, the user can immediately understand the reservation details and the results of the inquiry.
[0192] 3. Specific Examples
[0193] For example, if a user launches the "Store Assistant AI" app on their smartphone and voice-inputs, "I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM," the system will operate as follows:
[0194] 1. The user enters the prompt, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0195] 2. The device converts the voice into text and sends it to the server.
[0196] 3. The server uses the generative AI model to generate a dialogue script that reads, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0197] 4. The server uses a speech synthesis engine to convert the dialogue script into voice data.
[0198] 5. The server makes a call and communicates with the cafe in real time to make the reservation.
[0199] 6. The server records the conversation and creates a summary.
[0200] 7. A notification is sent to the device stating, "A cafe reservation has been made for two people at 7:00 PM."
[0201] This allows users to easily make reservations or inquiries without having to make a phone call themselves.
[0202] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0203] Program processing steps
[0204] Step 1:
[0205] The user launches the smartphone application and inputs the details of the reservation or inquiry. Input can be done by voice or text. For example, the user inputs, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0206] Input: User voice or text input
[0207] Output: Audio or text data
[0208] Step 2:
[0209] The device receives user input and, in the case of voice input, converts the speech into text data using a speech recognition engine such as Google Speech Recognition.
[0210] Input: Audio data
[0211] Output: Text data
[0212] Step 3:
[0213] The device sends the converted text data to a server, which receives the text data and uses a generative AI model (e.g., GPT-4) to analyze the user's intent.
[0214] Input: Text data
[0215] Output: Intention analysis results
[0216] Step 4:
[0217] The server generates a dialogue script based on the intent analysis results. The generated dialogue script might be in the form of, for example, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0218] Input: Intention analysis result
[0219] Output: Interactive script
[0220] Step 5:
[0221] The server converts the generated dialogue script into voice data using the Google Text-to-Speech (TTS) engine, which is then used in the actual phone call.
[0222] Input: Interactive script
[0223] Output: Audio data
[0224] Step 6:
[0225] The server uses the voice data to make calls and communicate with the physical store in real time, and the conversation is recorded by the server during the conversation.
[0226] Input: Audio data
[0227] Output: Recorded dialogue
[0228] Step 7:
[0229] The server uses voice recognition technology to convert the recorded conversation into text data, and then uses AI technology to create a summary. The summarized text data might be in the form of, for example, "A reservation has been made for two people at a cafe at 7:00 PM."
[0230] Input: Recorded dialogue
[0231] Output: Summary text
[0232] Step 8:
[0233] The server sends the summary text to the user's terminal, and the user can check the notification to understand the reservation details.
[0234] Input: Summary text
[0235] Output: Information message
[0236] The above is the flow of specific processing performed by the system, and the operations include data processing and data calculation performed at each step.
[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0238] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains the necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is sensitive to those emotions. The following describes in detail the embodiments of the invention.
[0239] Input and Initial Processing
[0240] 1. User:
[0241] Users launch a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries or reservations. Input can be done either by voice or text, and the emotion engine recognizes the user's emotions as they type.
[0242] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0243] 2. Terminal:
[0244] The device receives user input. In the case of voice input, the device converts the voice data into text data using a voice recognition engine. The converted text data and the user's emotion information are sent to the server.
[0245] AI processing and phone generation
[0246] 3. Server:
[0247] The server receives text data and emotional information from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intentions from the text data and generates a dialogue script that takes the emotional information into account.
[0248] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0249] 4. Server (Speech synthesis):
[0250] The generated dialogue script is input into a speech synthesis engine and converted into voice data that resembles the user's voice. At this time, the tone and expression of the voice are adjusted based on emotional information generated by the emotion engine.
[0251] Example: The generated speech data is, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0252] 5. Server (Calling):
[0253] The voice data is transmitted through a telephone system to place a call to a destination (for example, a restaurant).
[0254] Real-time dialogue and results notification
[0255] 6. Server:
[0256] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0257] 7. Server (Text Conversion and Summarization):
[0258] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[0259] 8. Notification from the server to the device:
[0260] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[0261] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0262] Specific example explanation
[0263] If a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the specific flow is as follows:
[0264] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." The system simultaneously recognizes the user's emotions (for example, if the user is feeling stressed).
[0265] 2. The device converts the speech into text and sends it to the server along with emotional information.
[0266] 3. The server uses the generative AI model to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0267] 4. The server uses a speech synthesis engine to generate voice data that reflects the user's voice and emotions based on the emotional information.
[0268] 5. The server makes the call and communicates with the other party in real time.
[0269] 6. The server records the conversation, converts it into text, and creates a summary.
[0270] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0271] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[0272] The processing flow will be explained below.
[0273] Step 1:
[0274] The user launches a dedicated smartphone app and inputs by voice or text what they want to communicate over the phone, such as an inquiry or reservation. Once input is complete, the emotion engine recognizes the user's emotions.
[0275] Example: A user speaks, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0276] Step 2:
[0277] The terminal receives input voice data and converts the voice data into text data using a voice recognition engine.
[0278] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0279] Step 3:
[0280] The terminal transmits the converted text data and the recognized emotion information to the server.
[0281] Step 4:
[0282] The server analyzes the text data and emotional information received from the device. It uses a generative AI model to understand the user's intentions and generate a dialogue script. This script is generated by reflecting the emotional information.
[0283] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0284] Step 5:
[0285] The server inputs the generated dialogue script into a speech synthesis engine to generate voice data that resembles the user's voice. The tone and expression of the voice are adjusted based on the emotions recognized by the emotion engine.
[0286] Example: Speech data: "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0287] Step 6:
[0288] The server transmits the generated voice data to a telephone system, which places a call to the destination (e.g., a restaurant).
[0289] Step 7:
[0290] The server establishes a telephone connection, transmits the generated voice data to the other party, and initiates a dialogue. The dialogue with the other party is conducted in real time, and an appropriate response is generated based on the other party's responses.
[0291] Example: If the reply is "I have availability at 7:00 PM," generate a response of "Thank you. I'd like to meet you at that time."
[0292] Step 8:
[0293] The server records the entire phone conversation.
[0294] Step 9:
[0295] The server converts the recorded audio data into text and creates a summary.
[0296] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[0297] Step 10:
[0298] The server transmits the summarized text data to the terminal.
[0299] Step 11:
[0300] The text data received by the device is notified to the user via push notification, email, etc.
[0301] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0302] Through the above series of processing steps, the user can efficiently obtain the necessary information in a manner that is in tune with their emotions, without having to make a phone call themselves.
[0303] Example 2
[0304] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0305] Conventional telephone reservation systems have problems such as the stress users feel when making phone calls themselves and the inability to efficiently obtain the necessary information. Furthermore, dialogue systems that proceed without taking the user's emotions into consideration often fail to provide the quality of dialogue users expect. There is a need for a system that can solve these problems and allow users to make reservations and inquiries more comfortably.
[0306] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for allowing a user to input the content of an inquiry or reservation by voice or text and analyzing the emotion at the time of input; means for converting the input voice data into text data; means for receiving the text data and emotion information, analyzing the intention using a generative AI model, and generating a dialogue script that takes emotion into consideration; voice synthesis means for converting the generated dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the content of the conversation, converting the recorded voice data into text, and summarizing it; and means for notifying the user's terminal of the summarized text. This reduces the stress of the user making a phone call and enables efficient information acquisition through a dialogue that is sensitive to the user's emotions.
[0307] "User" refers to a person who uses the system.
[0308] A "dedicated app" refers to application software that runs on devices such as smartphones and allows users to enter details of inquiries and reservations.
[0309] "Emotion engine" refers to software or algorithms for analyzing emotions from a user's voice or input data.
[0310] "Speech recognition engine" refers to software or algorithms for converting input voice data into text data.
[0311] A "generative AI model" refers to an artificial intelligence model that analyzes user intent based on text data and emotional information, and generates an appropriate dialogue script.
[0312] A "dialogue script" refers to a text-based scenario for progressing a dialogue, which is generated based on the user's intentions.
[0313] "Speech synthesis engine" refers to software or algorithms for converting text data into speech data.
[0314] "Telephone System" means the communications infrastructure and software used to make telephone calls based on generated voice data.
[0315] "Text conversion means" refers to a system or technology for converting voice data into text data.
[0316] "Summarization means" refers to a system or technology for concisely summarizing the converted text data.
[0317] "Push notification" refers to a technology that allows a server to send information to a user's device in real time.
[0318] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is considerate of those emotions. The following describes in detail the embodiments of the invention.
[0319] 1. User Input Procedure
[0320] Users launch a dedicated smartphone app and input the details of their inquiry or reservation. For voice input, users press the microphone button and start speaking. For text input, users type directly into the text field. As the user types, the app uses an emotion engine to analyze the user's emotions.
[0321] Specific hardware or software you will be using:
[0322] Dedicated app (smartphone app)
[0323] Emotion Engine
[0324] 2. Initial processing on the device
[0325] The device receives the input voice data and converts it into text data using a speech recognition engine. The converted text data and emotion information are then sent to the server. A speech recognition API such as Google Speech-to-Text is used here.
[0326] Specific hardware or software you will be using:
[0327] Speech recognition engine (Google Speech-to-Text)
[0328] 3. Data processing on the server
[0329] The server inputs the received text data and emotional information into a generative AI model to analyze the user's intentions. Generative AI models such as OpenAI GPT-4 are used here. The generated dialogue script takes the user's emotions into account.
[0330] Specific hardware or software you will be using:
[0331] Generative AI model (OpenAI GPT-4)
[0332] 4. Voice conversion of dialogue scripts
[0333] The generated dialogue script is passed to a speech synthesis engine, which converts it into voice data that reflects the user's voice and emotions. A speech synthesis engine such as Amazon Polly is used for this purpose.
[0334] Specific hardware or software you will be using:
[0335] Speech synthesis engine (Amazon Polly)
[0336] 5. Calling via the telephone system
[0337] Using voice data, calls are made through a telephone system and a real-time conversation is held with the other party. The server analyzes the other party's response in real time and uses a generative AI model to generate the next appropriate response and continue the conversation. Telephone systems such as Twilio are used.
[0338] Specific hardware or software you will be using:
[0339] Phone system (Twilio)
[0340] 6. Recording and summarizing conversation content
[0341] After the call is over, the server converts the recorded voice data into text and summarizes the key points, which are then stored in the system and later shared with the user.
[0342] Specific hardware or software you will be using:
[0343] Text conversion methods
[0344] Summary tools
[0345] 7. Notification of Results
[0346] The server sends the summarized text data to the device, and the user is notified of the results by push notification or other means. The user can check the call results and next steps through the app.
[0347] Specific hardware or software you will be using:
[0348] Push notification system
[0349] Specific examples
[0350] The specific flow when a user wants to make a restaurant reservation for tomorrow at 19:00 is shown below.
[0351] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." At this time, the system recognizes that the user is feeling stressed.
[0352] 2. The device converts the speech into text and sends it to the server along with emotional information.
[0353] 3. The server uses a generative AI model (e.g., GPT-4) to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know availability."
[0354] 4. The server uses a speech synthesis engine (e.g., Amazon Polly) to generate voice data and adjust it to reflect the emotion.
[0355] 5. The server uses a telephone system (e.g., Twilio) to make a call and communicate with the restaurant in real time.
[0356] 6. The server converts the call into text and creates a summary.
[0357] 7. The server sends the summary to the user's terminal and notifies them of the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0358] Example prompt sentence:
[0359] "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0360] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[0361] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0362] Step 1:
[0363] Users launch a dedicated smartphone app and input the details of their inquiry or reservation by voice or text. Once input is complete, the app uses an emotion engine to analyze the user's emotions. The input data is output as text data (in the case of text input) or voice data (in the case of voice input). Emotional data is also output at the same time.
[0364] Specific behavior:
[0365] The user speaks, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0366] The app receives the voice data and analyzes the emotions (e.g., tension or stress) present in the input.
[0367] Step 2:
[0368] The device receives the user's voice data and converts it into text data using a speech recognition engine. The text data and analyzed emotion information are sent to the server. The input is voice data, and the output is text data and emotion information.
[0369] Specific behavior:
[0370] The terminal converts the voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text).
[0371] The converted text data and emotion data are sent to the server.
[0372] Step 3:
[0373] The server receives the text data and emotional information sent from the device. The text data and emotional information are input into the generative AI model, which analyzes the user's intentions. The generated dialogue script is output. The input is text data and emotional information, and the output is a dialogue script.
[0374] Specific behavior:
[0375] The server inputs the data into a generative AI model such as GPT-4, which analyzes the user's intent.
[0376] The generative AI model generates a dialogue script that reads, "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0377] Step 4:
[0378] The server passes the generated dialogue script to a speech synthesis engine and converts it into voice data. The tone and expression of the voice are also adjusted based on the emotional information. The input is the dialogue script and emotional information, and the output is voice data.
[0379] Specific behavior:
[0380] The server inputs the dialogue script into a speech synthesis engine such as Amazon Polly.
[0381] A speech synthesis engine converts the dialogue script into voice data that reflects the user's voice and emotions.
[0382] Step 5:
[0383] The server uses the generated voice data to make a call through the telephone system, initiating a real-time dialogue with the other party, and the server analyzes the other party's response in real time and generates the next appropriate response using a generative AI model. The input is the voice data, and the output is a real-time response.
[0384] Specific behavior:
[0385] The server calls the restaurant using a phone system such as Twilio.
[0386] The server analyzes the other party's response and uses a generative AI model to generate an appropriate response and continue the dialogue.
[0387] Step 6:
[0388] After the call ends, the server converts the recorded voice data into text and summarizes the key points. The converted text data and the summary are output. The input is the voice data, and the output is the text data and the summary.
[0389] Specific behavior:
[0390] The server uses a speech recognition engine to convert the call into text.
[0391] The server uses a summarization algorithm to summarize the converted text.
[0392] Step 7:
[0393] The server sends the summarized text data to the terminal, and the user is notified of the results via push notification, etc. The input is the summarized text data, and the output is a notification to the user.
[0394] Specific behavior:
[0395] The server sends a summary to the terminal stating, "A restaurant reservation has been made for two people at 19:00."
[0396] The user receives a push notification and checks for details.
[0397] (Application example 2)
[0398] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0399] In conventional food delivery services, when users place orders over the phone, they are prone to making mistakes in how they communicate their orders. Furthermore, there are few ways to reduce the stress and tension users feel when ordering over the phone. Furthermore, there are no systems that can process orders in a way that takes into account the user's emotions. This creates a need for a user-friendly and efficient food delivery ordering system.
[0400] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generative model means for receiving the text data, analyzing the intention, and generating a dialogue script; speech synthesis means for converting the dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for recognizing the user's emotion and optimizing the dialogue script and voice data based on the emotion; input means for allowing the user to select either voice or character input; and means for acquiring emotion information and adjusting the tone and expression of the voice based on the emotion information. This enables the user to efficiently order food delivery in a manner that is sensitive to their emotion without having to make a phone call.
[0401] "User" refers to an individual or corporation that uses this system to make phone calls, make reservations, place orders, and perform other operations.
[0402] "Content to be communicated over the phone" refers to information that the user needs to communicate to the other party over the phone, such as an order, an inquiry, or a reservation.
[0403] "Input means" refers to an interface or device used by a user to input content by voice or text.
[0404] "Text data" refers to text information converted from voice input or text input.
[0405] "Generative model means" refers to the AI model or algorithm used to analyze user input and generate an appropriate dialogue script.
[0406] "Speech synthesis means" refers to a technology or device that converts the generated dialogue script into voice data.
[0407] "Voice data" refers to digital information of voice generated by a voice synthesis means.
[0408] "Means for making a call" refers to a system or device for making a call using the generated voice data.
[0409] "Means of real-time interaction" refers to technology or systems that allow you to have a direct conversation with the person you are calling.
[0410] "Means for recording conversation content" refers to technology or devices that save telephone conversations as audio data.
[0411] "Means of converting to text and summarizing" refers to the technology or algorithms that convert recorded audio data into text and summarize that text concisely.
[0412] "Means for notifying the device" refers to technology or systems that notify the user of the summarized text on their device, such as a smartphone or tablet.
[0413] "Means of recognizing emotions" refers to technologies and models for analyzing and extracting emotional information from user input and dialogue content.
[0414] "Means for optimizing dialogue scripts and voice data based on emotions" refers to technologies and algorithms that use emotional information to adjust the tone and expression of the dialogue scripts and voice data that are generated.
[0415] An input means that allows the user to select between "voice and text input" refers to an interface or device that allows the user to freely select either voice input or text input.
[0416] This invention provides a specific implementation method for recognizing user emotions and efficiently completing orders in a food delivery ordering system.
[0417] System Configuration
[0418] The present invention is mainly composed of a user terminal, a server, and an AI module.
[0419] User terminal
[0420] To order food delivery, a user first uses a user device such as a smartphone or tablet. The user then launches a dedicated app and can enter the order details by voice or text. The user's emotions are recognized as they are entered.
[0421] server
[0422] The server receives input data (voice or text) and emotional information sent from the user's device. In the case of voice input, the server converts the voice data into text data using a voice recognition engine. Based on this converted text data and emotional information, a generative AI model generates a dialogue script. The generated dialogue script is then input into a voice synthesis engine and converted into voice data. At this time, the tone and expression of the voice are optimized based on the emotional information.
[0423] AI Module
[0424] The generative AI model and speech synthesis engine are included in the AI module. The generative AI model analyzes the user's intentions and generates a dialogue script that takes emotional information into account. The speech synthesis engine converts this dialogue script into voice data.
[0425] communication means
[0426] The server uses the generated voice data to make calls to the corresponding restaurant or delivery provider and engages in real-time conversations, which are recorded in the server and converted into text data as needed.
[0427] User Notification
[0428] Once the telephone conversation is complete, the server summarizes the conversation and sends the summarized information to the user's device via a push notification or other method, allowing the user to confirm whether the order was placed successfully.
[0429] Specific examples
[0430] For example, if a user says "I would like to order a pizza from Restaurant X", the system will:
[0431] 1. The user device receives voice input, and the emotion recognition engine analyzes the user's emotions.
[0432] 2. The server converts the voice data into text, and the generative AI model generates a dialogue script. The generated dialogue script is, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita pizza, please."
[0433] 3. Based on this dialogue script, the speech synthesis engine optimizes the tone and expression of the voice according to the emotional information and generates the voice data.
[0434] 4. The server will make a call and automatically communicate the order details.
[0435] 5. The restaurant's response is recorded, and the content of the conversation is converted into text data and summarized.
[0436] 6. The summary is sent to the user's terminal, and the user is informed that "your order has been accepted. Delivery time is 18:00."
[0437] Prompt Sentence Examples
[0438] An example of a prompt to pass to a generative AI model would be:
[0439] "Hello, I'd like to order a pizza from XX Restaurant. I'd like one Margherita, please."
[0440] Emotion: Stress
[0441] This allows the system to efficiently order food delivery while being sensitive to the user's emotions.
[0442] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0443] Step 1:
[0444] The user launches a dedicated smartphone app and inputs their order details by voice or text. The emotion recognition engine analyzes the user's emotions as they input their order. For example, if a user inputs "I'd like to order pizza from Restaurant X," the device captures the voice data and analyzes and obtains emotional information.
[0445] Step 2:
[0446] The voice data acquired by the device is converted into text data. A voice recognition engine (e.g., Google Speech-to-Text API) is used to convert the voice data into text data. For example, a voice saying "I would like to order pizza from XX Restaurant" is converted into text data saying "I would like to order pizza from XX Restaurant."
[0447] Step 3:
[0448] Text data and emotional information are sent from the terminal to the server. The order details (text data) and emotional information (e.g., stress) entered by the user are sent to the server and used as input for the next process.
[0449] Step 4:
[0450] Based on the text data and emotional information received by the server, a generative AI model (e.g., GPT-3) is used to generate a dialogue script. The generative AI model analyzes the given text and emotional information and generates an appropriate dialogue script. For example, it generates a script that reads, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please."
[0451] Step 5:
[0452] The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data. At this time, the tone and expression of the voice are adjusted based on the emotional information. For example, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please" is generated in a soft voice that matches the emotion.
[0453] Step 6:
[0454] The server uses the generated voice data to place a call to the specified restaurant and execute the order. This telephone conversation takes place in real time, and the server accurately conveys the order details.
[0455] Step 7:
[0456] The server records the audio data during the conversation, including the restaurant's responses and questions.
[0457] Step 8:
[0458] The server converts the recorded voice data into text and summarizes it. The speech recognition engine is used again to convert the voice recording into text data, and the content is concisely summarized using a summarization algorithm. For example, a summary text such as "Your order has been received. Delivery time is 6:00 PM" is generated.
[0459] Step 9:
[0460] The server sends the summarized text to the user's device and sends a push notification with the order details, which displays a notification such as "Your order has been accepted. Delivery time is 6:00 PM."
[0461] This allows users to efficiently place food delivery orders in a way that is in tune with their emotions, without having to make a phone call themselves.
[0462] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0463] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0464] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0465] [Second embodiment]
[0466] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0467] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0468] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0469] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0470] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0471] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0472] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0473] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0474] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0475] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0476] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0477] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0478] The present invention is a system developed to reduce the stress users feel when making phone calls. This system uses AI to make calls on behalf of users, obtain necessary information, and notify the user of the results without the user having to make the calls themselves. The following describes in detail the embodiments of the present invention.
[0479] Input and Initial Processing
[0480] 1. User:
[0481] Users start up a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0482] Example: If a user enters "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0483] 2. Terminal:
[0484] The device receives user input. In the case of voice input, it uses a speech recognition engine to convert the voice data into text data, which is then sent to the server.
[0485] AI processing and phone generation
[0486] 3. Server:
[0487] The server receives text data from the device and passes it to a generative AI model (e.g., a model using natural language processing). The AI model analyzes the user's intent and generates an appropriate dialogue script.
[0488] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0489] 4. Server (Speech synthesis):
[0490] The generated dialogue script is input to a speech synthesis engine to generate voice data that resembles the user's voice. This voice data is then transmitted over the telephone system.
[0491] Real-time dialogue and results notification
[0492] 5. Server:
[0493] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0494] 6. Server (Text Conversion and Summarization):
[0495] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[0496] 7. Notification from the server to the device:
[0497] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[0498] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0499] Specific example explanation
[0500] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[0501] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0502] 2. The device converts the speech into text and sends it to the server.
[0503] 3. The server uses the generative AI model to analyze the input and generate an appropriate dialogue script (e.g., "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability.").
[0504] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[0505] 5. The server communicates with the other party in real time and records the conversation.
[0506] 6. The server converts the recorded audio data into text and creates a summary.
[0507] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0508] In this way, the system of the present invention allows the user to obtain the necessary information without stress, without having to make a phone call himself.
[0509] The processing flow will be explained below.
[0510] Step 1:
[0511] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0512] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0513] Step 2:
[0514] The device receives user input. In the case of voice input, the device's voice recognition engine is used to convert the voice data into text data.
[0515] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0516] Step 3:
[0517] The terminal transmits the converted text data to the server.
[0518] Step 4:
[0519] The server receives the text data sent from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intent from the text data.
[0520] Example: Analyze the text data "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0521] Step 5:
[0522] The server generates an interactive script based on the analysis results.
[0523] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0524] Step 6:
[0525] The server inputs the generated dialogue script into a voice synthesis engine to generate voice data that resembles the user's voice.
[0526] Example: The dialogue script generates the following speech data: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0527] Step 7:
[0528] The server sends the voice data to the telephone system, which then places a call to the destination (e.g., a restaurant).
[0529] Step 8:
[0530] The server establishes a telephone connection and transmits the generated voice data to the other party to start the conversation.
[0531] Step 9:
[0532] The server analyzes the response from the other party in real time and uses AI to generate an appropriate response.
[0533] Example: If a restaurant employee replies, "We have a table available at 7:00 PM," the AI generates a response: "Thank you. We'll make it that time then."
[0534] Step 10:
[0535] The server records the entire phone conversation.
[0536] Step 11:
[0537] The server converts the recorded audio data into text.
[0538] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[0539] Step 12:
[0540] The server summarizes the converted text data, extracts important information, and summarizes it compactly.
[0541] Example: The summary you get is "A restaurant reservation has been made for two people at 7:00 PM."
[0542] Step 13:
[0543] The server transmits the summarized text data to the terminal.
[0544] Step 14:
[0545] The text data received by the device is notified to the user via push notification, email, etc.
[0546] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0547] This series of steps allows the user to efficiently obtain the necessary information without having to make a phone call.
[0548] Example 1
[0549] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0550] In modern society, making reservations and inquiries over the phone is common, but it can cause stress and anxiety for many users. In particular, for users who are not good at communicating over the phone or who are busy, there is a demand for a system that can handle phone calls efficiently and easily. Conventional systems require users to make the calls themselves, which is time-consuming and mentally taxing. To solve this problem, the present invention aims to provide an automated telephone answering system that eliminates the need for users to make phone calls.
[0551] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0552] In this invention, the server includes means for inputting the content that the user wants to communicate over the phone, means for converting the input content into text data, and generative model means for receiving the text data, analyzing the intention, and generating a dialogue script. This allows the system to automatically make calls on behalf of the user, obtain the necessary information, and notify the user of the results, without the user having to make the call directly.
[0553] A "user" is a person or individual who uses the system to input what they want to say over the phone.
[0554] A "means" is a device or method used to achieve a particular purpose or function.
[0555] "Text data" refers to character-based information converted from audio data or other formats.
[0556] A "generative model" is an algorithm that uses natural language processing or machine learning to generate sentences that meet a specific purpose from specific input.
[0557] "Speech synthesis" is a technology that generates speech that sounds like human speech based on text data.
[0558] An "information processing device" is an electronic device for receiving and processing data, and typically includes smartphones, tablets, personal computers, etc.
[0559] An "external system" is an object other than the user's system, such as a restaurant or an office.
[0560] This invention is an automated telephone answering system developed to reduce the stress of users when making phone calls and to efficiently obtain the necessary information. The system aims to have users input the information they want to convey over the phone, and then have AI make the call on their behalf, obtain the information, and notify the user of the results.
[0561] Input and Initial Processing
[0562] 1. The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input can be done either by voice or text. Consider the example where the user inputs, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0563] 2. The device receives the user's voice input and converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text API), which is then sent to the server.
[0564] AI processing and phone generation
[0565] 3. The server receives the text data sent from the device and passes it to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the user's intent and generates an appropriate dialogue script. The generated dialogue script includes the following content: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0566] 4. The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data that resembles the user's voice. This voice data is then transmitted through the telephone system.
[0567] Real-time dialogue and results notification
[0568] 5. The server uses the generated voice to make a call and engage in a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0569] 6. The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the important points.
[0570] 7. The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking this notification, the user can easily understand the content and results of the call. For example, a notification may appear on the device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0571] Specific examples
[0572] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[0573] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0574] 2. The device converts the speech into text and sends it to the server.
[0575] 3. The server analyzes the input using a generative AI model and generates an appropriate dialogue script.
[0576] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[0577] 5. The server communicates with the other party in real time and records the conversation.
[0578] 6. The server converts the recorded audio data into text and summarizes the content.
[0579] 7. The server notifies the terminal of the summary, and the user confirms the results.
[0580] An example of a prompt sentence could be, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM. Please let me know the availability." In this way, the user can obtain the necessary information without stress, without having to make a phone call themselves.
[0581] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0582] Step 1:
[0583] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input methods include both voice input and text input. For example, a user might say, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." This will result in either voice data or text data being acquired as input data.
[0584] Step 2:
[0585] The device receives the user's input data. The received voice data is converted into text data using a speech recognition engine such as the Google Speech-to-Text API. During this conversion process, the voice data is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM." Text data is generated as output, and this text data is sent to the server.
[0586] Step 3:
[0587] The server receives the text data sent from the device. The received text data is passed to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the intent of the input text data and generates an appropriate dialogue script. For example, from the input data "I would like to make a restaurant reservation for tomorrow at 7:00 PM," the dialogue script generated is "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability." The dialogue script is obtained as the output.
[0588] Step 4:
[0589] The server inputs the generated dialogue script into a speech synthesis engine (for example, Google Text-to-Speech API). The speech synthesis engine generates voice data that resembles the user's voice based on the text data. This voice data is transmitted through the telephone system. The generated voice data is obtained as output and is used as the transmitted voice.
[0590] Step 5:
[0591] The server uses the generated voice data to make a call. The call is made to an external system, such as a restaurant, and the server conducts the conversation in real time. The AI system within the server generates an appropriate response based on the other party's response and continues the conversation. For example, the AI might say, "Could I please make a restaurant reservation for two people tomorrow at 7:00 PM?" and generate the next response based on the restaurant's response. The entire conversation is recorded within the server. The other party's response is obtained as output.
[0592] Step 6:
[0593] The server converts the recorded voice data into text data. The content is then summarized by a generative AI model. For example, if the restaurant responds, "We can make a reservation for two people at 7:00 PM," the resulting summarized text is, "A restaurant reservation has been made for two people at 7:00 PM." The summarized text is generated as the output.
[0594] Step 7:
[0595] The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking the notification, the user can easily understand the content and results of the call. For example, a notification saying "A restaurant reservation has been made for two people at 7:00 PM" is displayed on the user's smartphone. The notification message is obtained as output.
[0596] (Application example 1)
[0597] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0598] In modern society, users experience a great deal of stress when making reservations or inquiries at physical stores over the phone. This stress increases especially when the call is not answered or when appropriate communication is not possible. There is a need to solve this problem and provide a method that allows users to make reservations and inquiries more easily.
[0599] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0600] In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generation AI model means for receiving the text data, analyzing the intention, and generating a dialogue script; voice synthesis means for converting the dialogue script into voice data; means for making a call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for a user to input a reservation or inquiry for a physical store; means for receiving input related to the physical store, generating a dialogue script, and making a call to collect necessary information; and means for notifying the user of the results of the reservation or inquiry for the physical store. This enables a user to easily make a reservation or inquiry for a physical store without having to make a call themselves.
[0601] definition statement
[0602] "Means for users to input information they want to communicate over the phone" refers to an interface that allows users to input details of reservations, inquiries, etc. by voice or text using a smartphone or other device.
[0603] "Means for converting input content into text data" refers to the process of generating text data from the voice or image input by the user using a voice recognition engine or OCR technology.
[0604] A "generative AI model means" is an algorithm or program that uses a generative AI model (e.g., GPT-4) to generate an appropriate dialogue script based on the user's intentions.
[0605] The "voice synthesis means for converting a dialogue script into voice data" is a process that uses voice synthesis technology (e.g., a TTS engine) to convert text data into voice data.
[0606] "Means for making a call using voice data and having a real-time conversation with the other party" refers to a communication technology that uses the generated voice data in an actual telephone conversation to conduct a real-time conversation with the other party.
[0607] "Means for recording the content of a conversation and converting the recorded voice data into text and summarizing it" refers to the process of recording the content of a conversation, converting the recorded data into text data using voice recognition technology, and creating a summary using AI technology.
[0608] The "means for notifying the user of the summarized text" refers to a technology for delivering the generated summarized text to the user's smartphone or other device as a push notification or message.
[0609] "A means by which a user can input reservations or inquiries for a physical store" is an interface that allows a user to input reservation or inquiry information about a specific store.
[0610] "Means of receiving input about a physical store, generating a dialogue script, and making a phone call to collect the necessary information" refers to the process in which the generative AI model creates a dialogue script based on the information about the physical store entered by the user, and makes a phone call to obtain the necessary information.
[0611] "Means for notifying users of the results of reservations and inquiries at physical stores" refers to technology for notifying users of the results of reservations and inquiries obtained via their terminals.
[0612] MODE FOR CARRYING OUT THE INVENTION
[0613] The present invention is a system that automates the process of making reservations or inquiries about physical stores, reducing stress for users. This system operates through a smartphone application, generates a dialogue script using the generated AI model, and makes phone calls to collect necessary information. The following describes an embodiment of the present invention.
[0614] 1. Basic configuration
[0615] This system consists of the following main hardware and software:
[0616] User terminal (smartphone): A device where a user enters input and receives results. It includes a speech recognition engine and a text input interface.
[0617] Server: Receives user input, generates a dialogue script using a generative AI model, creates voice data using a speech synthesis engine, and processes the call.
[0618] Generative AI model: A technology that generates dialogue scripts based on user requests. An example of this is GPT-4.
[0619] Speech synthesis engine: A technology that converts text data into voice data. An example is Google Text-to-Speech (TTS).
[0620] 2. System processing flow
[0621] A user starts a smartphone application and inputs reservation or inquiry details. For example, a user inputs a prompt such as "I would like to make a restaurant reservation for two people tomorrow at 7:00 PM." This input can be done either by voice or text.
[0622] When voice input is used, the device uses a voice recognition engine such as Google Speech Recognition to convert the voice into text data and then sends the text data to the server.
[0623] The server receives the text data and uses a generative AI model (e.g., GPT-4) to generate a dialogue script based on the user's intention. The generated dialogue script is in the form of "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM."
[0624] The server then uses Google Text-to-Speech (TTS) to convert the generated dialogue script into voice data, and makes an actual call through the telephone system to interact with the physical store in real time.
[0625] During the call process, the server records the conversation, and after the call is over, it converts the audio data back into text data and uses AI technology to create a summary, such as "A restaurant reservation has been made for two people at 7:00 PM."
[0626] The completed summary is sent from the server to the device, and the results are communicated to the user via push notification or message. By checking this notification, the user can immediately understand the reservation details and the results of the inquiry.
[0627] 3. Specific Examples
[0628] For example, if a user launches the "Store Assistant AI" app on their smartphone and voice-inputs, "I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM," the system will operate as follows:
[0629] 1. The user enters the prompt, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0630] 2. The device converts the voice into text and sends it to the server.
[0631] 3. The server uses the generative AI model to generate a dialogue script that reads, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0632] 4. The server uses a speech synthesis engine to convert the dialogue script into voice data.
[0633] 5. The server makes a call and communicates with the cafe in real time to make the reservation.
[0634] 6. The server records the conversation and creates a summary.
[0635] 7. A notification is sent to the device stating, "A cafe reservation has been made for two people at 7:00 PM."
[0636] This allows users to easily make reservations or inquiries without having to make a phone call themselves.
[0637] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0638] Program processing steps
[0639] Step 1:
[0640] The user launches the smartphone application and inputs the details of the reservation or inquiry. Input can be done by voice or text. For example, the user inputs, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0641] Input: User voice or text input
[0642] Output: Audio or text data
[0643] Step 2:
[0644] The device receives user input and, in the case of voice input, converts the speech into text data using a speech recognition engine such as Google Speech Recognition.
[0645] Input: Audio data
[0646] Output: Text data
[0647] Step 3:
[0648] The device sends the converted text data to a server, which receives the text data and uses a generative AI model (e.g., GPT-4) to analyze the user's intent.
[0649] Input: Text data
[0650] Output: Intention analysis results
[0651] Step 4:
[0652] The server generates a dialogue script based on the intent analysis results. The generated dialogue script might be in the form of, for example, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[0653] Input: Intention analysis result
[0654] Output: Interactive script
[0655] Step 5:
[0656] The server converts the generated dialogue script into voice data using the Google Text-to-Speech (TTS) engine, which is then used in the actual phone call.
[0657] Input: Interactive script
[0658] Output: Audio data
[0659] Step 6:
[0660] The server uses the voice data to make calls and communicate with the physical store in real time, and the conversation is recorded by the server during the conversation.
[0661] Input: Audio data
[0662] Output: Recorded dialogue
[0663] Step 7:
[0664] The server uses voice recognition technology to convert the recorded conversation into text data, and then uses AI technology to create a summary. The summarized text data might be in the form of, for example, "A reservation has been made for two people at a cafe at 7:00 PM."
[0665] Input: Recorded dialogue
[0666] Output: Summary text
[0667] Step 8:
[0668] The server sends the summary text to the user's terminal, and the user can check the notification to understand the reservation details.
[0669] Input: Summary text
[0670] Output: Information message
[0671] The above is the flow of specific processing performed by the system, and the operations include data processing and data calculation performed at each step.
[0672] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0673] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains the necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is sensitive to those emotions. The following describes in detail the embodiments of the invention.
[0674] Input and Initial Processing
[0675] 1. User:
[0676] Users launch a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries or reservations. Input can be done either by voice or text, and the emotion engine recognizes the user's emotions as they type.
[0677] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0678] 2. Terminal:
[0679] The device receives user input. In the case of voice input, the device converts the voice data into text data using a voice recognition engine. The converted text data and the user's emotion information are sent to the server.
[0680] AI processing and phone generation
[0681] 3. Server:
[0682] The server receives text data and emotional information from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intentions from the text data and generates a dialogue script that takes the emotional information into account.
[0683] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0684] 4. Server (Speech synthesis):
[0685] The generated dialogue script is input into a speech synthesis engine and converted into voice data that resembles the user's voice. At this time, the tone and expression of the voice are adjusted based on emotional information generated by the emotion engine.
[0686] Example: The generated speech data is, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0687] 5. Server (Calling):
[0688] The voice data is transmitted through a telephone system to place a call to a destination (for example, a restaurant).
[0689] Real-time dialogue and results notification
[0690] 6. Server:
[0691] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0692] 7. Server (Text Conversion and Summarization):
[0693] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[0694] 8. Notification from the server to the device:
[0695] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[0696] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0697] Specific example explanation
[0698] If a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the specific flow is as follows:
[0699] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." The system simultaneously recognizes the user's emotions (for example, if the user is feeling stressed).
[0700] 2. The device converts the speech into text and sends it to the server along with emotional information.
[0701] 3. The server uses the generative AI model to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0702] 4. The server uses a speech synthesis engine to generate voice data that reflects the user's voice and emotions based on the emotional information.
[0703] 5. The server makes the call and communicates with the other party in real time.
[0704] 6. The server records the conversation, converts it into text, and creates a summary.
[0705] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0706] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[0707] The processing flow will be explained below.
[0708] Step 1:
[0709] The user launches a dedicated smartphone app and inputs by voice or text what they want to communicate over the phone, such as an inquiry or reservation. Once input is complete, the emotion engine recognizes the user's emotions.
[0710] Example: A user speaks, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0711] Step 2:
[0712] The terminal receives input voice data and converts the voice data into text data using a voice recognition engine.
[0713] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0714] Step 3:
[0715] The terminal transmits the converted text data and the recognized emotion information to the server.
[0716] Step 4:
[0717] The server analyzes the text data and emotional information received from the device. It uses a generative AI model to understand the user's intentions and generate a dialogue script. This script is generated by reflecting the emotional information.
[0718] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0719] Step 5:
[0720] The server inputs the generated dialogue script into a speech synthesis engine to generate voice data that resembles the user's voice. The tone and expression of the voice are adjusted based on the emotions recognized by the emotion engine.
[0721] Example: Speech data: "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0722] Step 6:
[0723] The server transmits the generated voice data to a telephone system, which places a call to the destination (e.g., a restaurant).
[0724] Step 7:
[0725] The server establishes a telephone connection, transmits the generated voice data to the other party, and initiates a dialogue. The dialogue with the other party is conducted in real time, and an appropriate response is generated based on the other party's responses.
[0726] Example: If the reply is "I have availability at 7:00 PM," generate a response of "Thank you. I'd like to meet you at that time."
[0727] Step 8:
[0728] The server records the entire phone conversation.
[0729] Step 9:
[0730] The server converts the recorded audio data into text and creates a summary.
[0731] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[0732] Step 10:
[0733] The server transmits the summarized text data to the terminal.
[0734] Step 11:
[0735] The text data received by the device is notified to the user via push notification, email, etc.
[0736] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0737] Through the above series of processing steps, the user can efficiently obtain the necessary information in a manner that is in tune with their emotions, without having to make a phone call themselves.
[0738] Example 2
[0739] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0740] Conventional telephone reservation systems have problems such as the stress users feel when making phone calls themselves and the inability to efficiently obtain the necessary information. Furthermore, dialogue systems that proceed without taking the user's emotions into consideration often fail to provide the quality of dialogue users expect. There is a need for a system that can solve these problems and allow users to make reservations and inquiries more comfortably.
[0741] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for allowing a user to input the content of an inquiry or reservation by voice or text and analyzing the emotion at the time of input; means for converting the input voice data into text data; means for receiving the text data and emotion information, analyzing the intention using a generative AI model, and generating a dialogue script that takes emotion into consideration; voice synthesis means for converting the generated dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the content of the conversation, converting the recorded voice data into text, and summarizing it; and means for notifying the user's terminal of the summarized text. This reduces the stress of the user making a phone call and enables efficient information acquisition through a dialogue that is sensitive to the user's emotions.
[0742] "User" refers to a person who uses the system.
[0743] A "dedicated app" refers to application software that runs on devices such as smartphones and allows users to enter details of inquiries and reservations.
[0744] "Emotion engine" refers to software or algorithms for analyzing emotions from a user's voice or input data.
[0745] "Speech recognition engine" refers to software or algorithms for converting input voice data into text data.
[0746] A "generative AI model" refers to an artificial intelligence model that analyzes user intent based on text data and emotional information, and generates an appropriate dialogue script.
[0747] A "dialogue script" refers to a text-based scenario for progressing a dialogue, which is generated based on the user's intentions.
[0748] "Speech synthesis engine" refers to software or algorithms for converting text data into speech data.
[0749] "Telephone System" means the communications infrastructure and software used to make telephone calls based on generated voice data.
[0750] "Text conversion means" refers to a system or technology for converting voice data into text data.
[0751] "Summarization means" refers to a system or technology for concisely summarizing the converted text data.
[0752] "Push notification" refers to a technology that allows a server to send information to a user's device in real time.
[0753] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is considerate of those emotions. The following describes in detail the embodiments of the invention.
[0754] 1. User Input Procedure
[0755] Users launch a dedicated smartphone app and input the details of their inquiry or reservation. For voice input, users press the microphone button and start speaking. For text input, users type directly into the text field. As the user types, the app uses an emotion engine to analyze the user's emotions.
[0756] Specific hardware or software you will be using:
[0757] Dedicated app (smartphone app)
[0758] Emotion Engine
[0759] 2. Initial processing on the device
[0760] The device receives the input voice data and converts it into text data using a speech recognition engine. The converted text data and emotion information are then sent to the server. A speech recognition API such as Google Speech-to-Text is used here.
[0761] Specific hardware or software you will be using:
[0762] Speech recognition engine (Google Speech-to-Text)
[0763] 3. Data processing on the server
[0764] The server inputs the received text data and emotional information into a generative AI model to analyze the user's intentions. Generative AI models such as OpenAI GPT-4 are used here. The generated dialogue script takes the user's emotions into account.
[0765] Specific hardware or software you will be using:
[0766] Generative AI model (OpenAI GPT-4)
[0767] 4. Voice conversion of dialogue scripts
[0768] The generated dialogue script is passed to a speech synthesis engine, which converts it into voice data that reflects the user's voice and emotions. A speech synthesis engine such as Amazon Polly is used for this purpose.
[0769] Specific hardware or software you will be using:
[0770] Speech synthesis engine (Amazon Polly)
[0771] 5. Calling via the telephone system
[0772] Using voice data, calls are made through a telephone system and a real-time conversation is held with the other party. The server analyzes the other party's response in real time and uses a generative AI model to generate the next appropriate response and continue the conversation. Telephone systems such as Twilio are used.
[0773] Specific hardware or software you will be using:
[0774] Phone system (Twilio)
[0775] 6. Recording and summarizing conversation content
[0776] After the call is over, the server converts the recorded voice data into text and summarizes the key points, which are then stored in the system and later shared with the user.
[0777] Specific hardware or software you will be using:
[0778] Text conversion methods
[0779] Summary tools
[0780] 7. Notification of Results
[0781] The server sends the summarized text data to the device, and the user is notified of the results by push notification or other means. The user can check the call results and next steps through the app.
[0782] Specific hardware or software you will be using:
[0783] Push notification system
[0784] Specific examples
[0785] The specific flow when a user wants to make a restaurant reservation for tomorrow at 19:00 is shown below.
[0786] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." At this time, the system recognizes that the user is feeling stressed.
[0787] 2. The device converts the speech into text and sends it to the server along with emotional information.
[0788] 3. The server uses a generative AI model (e.g., GPT-4) to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know availability."
[0789] 4. The server uses a speech synthesis engine (e.g., Amazon Polly) to generate voice data and adjust it to reflect the emotion.
[0790] 5. The server uses a telephone system (e.g., Twilio) to make a call and communicate with the restaurant in real time.
[0791] 6. The server converts the call into text and creates a summary.
[0792] 7. The server sends the summary to the user's terminal and notifies them of the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0793] Example prompt sentence:
[0794] "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0795] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[0796] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0797] Step 1:
[0798] Users launch a dedicated smartphone app and input the details of their inquiry or reservation by voice or text. Once input is complete, the app uses an emotion engine to analyze the user's emotions. The input data is output as text data (in the case of text input) or voice data (in the case of voice input). Emotional data is also output at the same time.
[0799] Specific behavior:
[0800] The user speaks, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0801] The app receives the voice data and analyzes the emotions (e.g., tension or stress) present in the input.
[0802] Step 2:
[0803] The device receives the user's voice data and converts it into text data using a speech recognition engine. The text data and analyzed emotion information are sent to the server. The input is voice data, and the output is text data and emotion information.
[0804] Specific behavior:
[0805] The terminal converts the voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text).
[0806] The converted text data and emotion data are sent to the server.
[0807] Step 3:
[0808] The server receives the text data and emotional information sent from the device. The text data and emotional information are input into the generative AI model, which analyzes the user's intentions. The generated dialogue script is output. The input is text data and emotional information, and the output is a dialogue script.
[0809] Specific behavior:
[0810] The server inputs the data into a generative AI model such as GPT-4, which analyzes the user's intent.
[0811] The generative AI model generates a dialogue script that reads, "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0812] Step 4:
[0813] The server passes the generated dialogue script to a speech synthesis engine and converts it into voice data. The tone and expression of the voice are also adjusted based on the emotional information. The input is the dialogue script and emotional information, and the output is voice data.
[0814] Specific behavior:
[0815] The server inputs the dialogue script into a speech synthesis engine such as Amazon Polly.
[0816] A speech synthesis engine converts the dialogue script into voice data that reflects the user's voice and emotions.
[0817] Step 5:
[0818] The server uses the generated voice data to make a call through the telephone system, initiating a real-time dialogue with the other party, and the server analyzes the other party's response in real time and generates the next appropriate response using a generative AI model. The input is the voice data, and the output is a real-time response.
[0819] Specific behavior:
[0820] The server calls the restaurant using a phone system such as Twilio.
[0821] The server analyzes the other party's response and uses a generative AI model to generate an appropriate response and continue the dialogue.
[0822] Step 6:
[0823] After the call ends, the server converts the recorded voice data into text and summarizes the key points. The converted text data and the summary are output. The input is the voice data, and the output is the text data and the summary.
[0824] Specific behavior:
[0825] The server uses a speech recognition engine to convert the call into text.
[0826] The server uses a summarization algorithm to summarize the converted text.
[0827] Step 7:
[0828] The server sends the summarized text data to the terminal, and the user is notified of the results via push notification, etc. The input is the summarized text data, and the output is a notification to the user.
[0829] Specific behavior:
[0830] The server sends a summary to the terminal stating, "A restaurant reservation has been made for two people at 19:00."
[0831] The user receives a push notification and checks for details.
[0832] (Application example 2)
[0833] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0834] In conventional food delivery services, when users place orders over the phone, they are prone to making mistakes in how they communicate their orders. Furthermore, there are few ways to reduce the stress and tension users feel when ordering over the phone. Furthermore, there are no systems that can process orders in a way that takes into account the user's emotions. This creates a need for a user-friendly and efficient food delivery ordering system.
[0835] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generative model means for receiving the text data, analyzing the intention, and generating a dialogue script; speech synthesis means for converting the dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for recognizing the user's emotion and optimizing the dialogue script and voice data based on the emotion; input means for allowing the user to select either voice or character input; and means for acquiring emotion information and adjusting the tone and expression of the voice based on the emotion information. This enables the user to efficiently order food delivery in a manner that is sensitive to their emotion without having to make a phone call.
[0836] "User" refers to an individual or corporation that uses this system to make phone calls, make reservations, place orders, and perform other operations.
[0837] "Content to be communicated over the phone" refers to information that the user needs to communicate to the other party over the phone, such as an order, an inquiry, or a reservation.
[0838] "Input means" refers to an interface or device used by a user to input content by voice or text.
[0839] "Text data" refers to text information converted from voice input or text input.
[0840] "Generative model means" refers to the AI model or algorithm used to analyze user input and generate an appropriate dialogue script.
[0841] "Speech synthesis means" refers to a technology or device that converts the generated dialogue script into voice data.
[0842] "Voice data" refers to digital information of voice generated by a voice synthesis means.
[0843] "Means for making a call" refers to a system or device for making a call using the generated voice data.
[0844] "Means of real-time interaction" refers to technology or systems that allow you to have a direct conversation with the person you are calling.
[0845] "Means for recording conversation content" refers to technology or devices that save telephone conversations as audio data.
[0846] "Means of converting to text and summarizing" refers to the technology or algorithms that convert recorded audio data into text and summarize that text concisely.
[0847] "Means for notifying the device" refers to technology or systems that notify the user of the summarized text on their device, such as a smartphone or tablet.
[0848] "Means of recognizing emotions" refers to technologies and models for analyzing and extracting emotional information from user input and dialogue content.
[0849] "Means for optimizing dialogue scripts and voice data based on emotions" refers to technologies and algorithms that use emotional information to adjust the tone and expression of the dialogue scripts and voice data that are generated.
[0850] An input means that allows the user to select between "voice and text input" refers to an interface or device that allows the user to freely select either voice input or text input.
[0851] This invention provides a specific implementation method for recognizing user emotions and efficiently completing orders in a food delivery ordering system.
[0852] System Configuration
[0853] The present invention is mainly composed of a user terminal, a server, and an AI module.
[0854] User terminal
[0855] To order food delivery, a user first uses a user device such as a smartphone or tablet. The user then launches a dedicated app and can enter the order details by voice or text. The user's emotions are recognized as they are entered.
[0856] server
[0857] The server receives input data (voice or text) and emotional information sent from the user's device. In the case of voice input, the server converts the voice data into text data using a voice recognition engine. Based on this converted text data and emotional information, a generative AI model generates a dialogue script. The generated dialogue script is then input into a voice synthesis engine and converted into voice data. At this time, the tone and expression of the voice are optimized based on the emotional information.
[0858] AI Module
[0859] The generative AI model and speech synthesis engine are included in the AI module. The generative AI model analyzes the user's intentions and generates a dialogue script that takes emotional information into account. The speech synthesis engine converts this dialogue script into voice data.
[0860] communication means
[0861] The server uses the generated voice data to make calls to the corresponding restaurant or delivery provider and engages in real-time conversations, which are recorded in the server and converted into text data as needed.
[0862] User Notification
[0863] Once the telephone conversation is complete, the server summarizes the conversation and sends the summarized information to the user's device via a push notification or other method, allowing the user to confirm whether the order was placed successfully.
[0864] Specific examples
[0865] For example, if a user says "I would like to order a pizza from Restaurant X", the system will:
[0866] 1. The user device receives voice input, and the emotion recognition engine analyzes the user's emotions.
[0867] 2. The server converts the voice data into text, and the generative AI model generates a dialogue script. The generated dialogue script is, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita pizza, please."
[0868] 3. Based on this dialogue script, the speech synthesis engine optimizes the tone and expression of the voice according to the emotional information and generates the voice data.
[0869] 4. The server will make a call and automatically communicate the order details.
[0870] 5. The restaurant's response is recorded, and the content of the conversation is converted into text data and summarized.
[0871] 6. The summary is sent to the user's terminal, and the user is informed that "your order has been accepted. Delivery time is 18:00."
[0872] Prompt Sentence Examples
[0873] An example of a prompt to pass to a generative AI model would be:
[0874] "Hello, I'd like to order a pizza from XX Restaurant. I'd like one Margherita, please."
[0875] Emotion: Stress
[0876] This allows the system to efficiently order food delivery while being sensitive to the user's emotions.
[0877] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0878] Step 1:
[0879] The user launches a dedicated smartphone app and inputs their order details by voice or text. The emotion recognition engine analyzes the user's emotions as they input their order. For example, if a user inputs "I'd like to order pizza from Restaurant X," the device captures the voice data and analyzes and obtains emotional information.
[0880] Step 2:
[0881] The voice data acquired by the device is converted into text data. A voice recognition engine (e.g., Google Speech-to-Text API) is used to convert the voice data into text data. For example, a voice saying "I would like to order pizza from XX Restaurant" is converted into text data saying "I would like to order pizza from XX Restaurant."
[0882] Step 3:
[0883] Text data and emotional information are sent from the terminal to the server. The order details (text data) and emotional information (e.g., stress) entered by the user are sent to the server and used as input for the next process.
[0884] Step 4:
[0885] Based on the text data and emotional information received by the server, a generative AI model (e.g., GPT-3) is used to generate a dialogue script. The generative AI model analyzes the given text and emotional information and generates an appropriate dialogue script. For example, it generates a script that reads, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please."
[0886] Step 5:
[0887] The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data. At this time, the tone and expression of the voice are adjusted based on the emotional information. For example, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please" is generated in a soft voice that matches the emotion.
[0888] Step 6:
[0889] The server uses the generated voice data to place a call to the specified restaurant and execute the order. This telephone conversation takes place in real time, and the server accurately conveys the order details.
[0890] Step 7:
[0891] The server records the audio data during the conversation, including the restaurant's responses and questions.
[0892] Step 8:
[0893] The server converts the recorded voice data into text and summarizes it. The speech recognition engine is used again to convert the voice recording into text data, and the content is concisely summarized using a summarization algorithm. For example, a summary text such as "Your order has been received. Delivery time is 6:00 PM" is generated.
[0894] Step 9:
[0895] The server sends the summarized text to the user's device and sends a push notification with the order details, which displays a notification such as "Your order has been accepted. Delivery time is 6:00 PM."
[0896] This allows users to efficiently place food delivery orders in a way that is in tune with their emotions, without having to make a phone call themselves.
[0897] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0898] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0899] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0900] [Third embodiment]
[0901] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0902] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0903] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0904] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0905] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0906] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0907] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0908] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0909] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0910] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0911] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0912] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0913] The present invention is a system developed to reduce the stress users feel when making phone calls. This system uses AI to make calls on behalf of users, obtain necessary information, and notify the user of the results without the user having to make the calls themselves. The following describes in detail the embodiments of the present invention.
[0914] Input and Initial Processing
[0915] 1. User:
[0916] Users start up a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0917] Example: If a user enters "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0918] 2. Terminal:
[0919] The device receives user input. In the case of voice input, it uses a speech recognition engine to convert the voice data into text data, which is then sent to the server.
[0920] AI processing and phone generation
[0921] 3. Server:
[0922] The server receives text data from the device and passes it to a generative AI model (e.g., a model using natural language processing). The AI model analyzes the user's intent and generates an appropriate dialogue script.
[0923] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0924] 4. Server (Speech synthesis):
[0925] The generated dialogue script is input to a speech synthesis engine to generate voice data that resembles the user's voice. This voice data is then transmitted through the telephone system.
[0926] Real-time dialogue and results notification
[0927] 5. Server:
[0928] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[0929] 6. Server (Text Conversion and Summarization):
[0930] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[0931] 7. Notification from the server to the device:
[0932] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[0933] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0934] Specific example explanation
[0935] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[0936] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0937] 2. The device converts the speech into text and sends it to the server.
[0938] 3. The server uses the generative AI model to analyze the input and generate an appropriate dialogue script (e.g., "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability.").
[0939] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[0940] 5. The server communicates with the other party in real time and records the conversation.
[0941] 6. The server converts the recorded audio data into text and creates a summary.
[0942] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[0943] In this way, the system of the present invention allows the user to obtain the necessary information without stress, without having to make a phone call himself.
[0944] The processing flow will be explained below.
[0945] Step 1:
[0946] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[0947] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0948] Step 2:
[0949] The device receives user input. In the case of voice input, the device's voice recognition engine is used to convert the voice data into text data.
[0950] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0951] Step 3:
[0952] The terminal transmits the converted text data to the server.
[0953] Step 4:
[0954] The server receives the text data sent from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intent from the text data.
[0955] Example: Analyze the text data "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[0956] Step 5:
[0957] The server generates an interactive script based on the analysis results.
[0958] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[0959] Step 6:
[0960] The server inputs the generated dialogue script into a voice synthesis engine to generate voice data that resembles the user's voice.
[0961] Example: The dialogue script generates the following speech data: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[0962] Step 7:
[0963] The server sends the voice data to the telephone system, which then places a call to the destination (e.g., a restaurant).
[0964] Step 8:
[0965] The server establishes a telephone connection and transmits the generated voice data to the other party to start the conversation.
[0966] Step 9:
[0967] The server analyzes the response from the other party in real time and uses AI to generate an appropriate response.
[0968] Example: If a restaurant employee replies, "We have a table available at 7:00 PM," the AI generates a response: "Thank you. We'll make it that time then."
[0969] Step 10:
[0970] The server records the entire phone conversation.
[0971] Step 11:
[0972] The server converts the recorded audio data into text.
[0973] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[0974] Step 12:
[0975] The server summarizes the converted text data, extracts important information, and summarizes it compactly.
[0976] Example: The summary you get is "A restaurant reservation has been made for two people at 7:00 PM."
[0977] Step 13:
[0978] The server transmits the summarized text data to the terminal.
[0979] Step 14:
[0980] The text data received by the device is notified to the user via push notification, email, etc.
[0981] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[0982] This series of steps allows the user to efficiently obtain the necessary information without having to make a phone call.
[0983] Example 1
[0984] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0985] In modern society, making reservations and inquiries over the phone is common, but it can cause stress and anxiety for many users. In particular, for users who are not good at communicating over the phone or who are busy, there is a demand for a system that can handle phone calls efficiently and easily. Conventional systems require users to make the calls themselves, which is time-consuming and mentally taxing. To solve this problem, the present invention aims to provide an automated telephone answering system that eliminates the need for users to make phone calls.
[0986] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0987] In this invention, the server includes means for inputting the content that the user wants to communicate over the phone, means for converting the input content into text data, and generative model means for receiving the text data, analyzing the intention, and generating a dialogue script. This allows the system to automatically make calls on behalf of the user, obtain the necessary information, and notify the user of the results, without the user having to make the call directly.
[0988] A "user" is a person or individual who uses the system to input what they want to say over the phone.
[0989] A "means" is a device or method used to achieve a particular purpose or function.
[0990] "Text data" refers to character-based information converted from audio data or other formats.
[0991] A "generative model" is an algorithm that uses natural language processing or machine learning to generate sentences that meet a specific purpose from specific input.
[0992] "Speech synthesis" is a technology that generates speech that sounds like human speech based on text data.
[0993] An "information processing device" is an electronic device for receiving and processing data, and typically includes smartphones, tablets, personal computers, etc.
[0994] An "external system" is an object other than the user's system, such as a restaurant or an office.
[0995] This invention is an automated telephone answering system developed to reduce the stress of users when making phone calls and to efficiently obtain the necessary information. The system aims to have users input the information they want to convey over the phone, and then have AI make the call on their behalf, obtain the information, and notify the user of the results.
[0996] Input and Initial Processing
[0997] 1. The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input can be done either by voice or text. Consider the example where the user inputs, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[0998] 2. The device receives the user's voice input and converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text API), which is then sent to the server.
[0999] AI processing and phone generation
[1000] 3. The server receives the text data sent from the device and passes it to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the user's intent and generates an appropriate dialogue script. The generated dialogue script includes the following content: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1001] 4. The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data that resembles the user's voice. This voice data is then transmitted through the telephone system.
[1002] Real-time dialogue and results notification
[1003] 5. The server uses the generated voice to make a call and engage in a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[1004] 6. The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the important points.
[1005] 7. The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking this notification, the user can easily understand the content and results of the call. For example, a notification may appear on the device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1006] Specific examples
[1007] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[1008] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1009] 2. The device converts the speech into text and sends it to the server.
[1010] 3. The server analyzes the input using a generative AI model and generates an appropriate dialogue script.
[1011] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[1012] 5. The server communicates with the other party in real time and records the conversation.
[1013] 6. The server converts the recorded audio data into text and summarizes the content.
[1014] 7. The server notifies the terminal of the summary, and the user confirms the results.
[1015] An example of a prompt sentence could be, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM. Please let me know the availability." In this way, the user can obtain the necessary information without stress, without having to make a phone call themselves.
[1016] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1017] Step 1:
[1018] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input methods include both voice input and text input. For example, a user might say, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." This will result in either voice data or text data being acquired as input data.
[1019] Step 2:
[1020] The device receives the user's input data. The received voice data is converted into text data using a speech recognition engine such as the Google Speech-to-Text API. During this conversion process, the voice data is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM." Text data is generated as output, and this text data is sent to the server.
[1021] Step 3:
[1022] The server receives the text data sent from the device. The received text data is passed to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the intent of the input text data and generates an appropriate dialogue script. For example, from the input data "I would like to make a restaurant reservation for tomorrow at 7:00 PM," the dialogue script generated is "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability." The dialogue script is obtained as the output.
[1023] Step 4:
[1024] The server inputs the generated dialogue script into a speech synthesis engine (for example, Google Text-to-Speech API). The speech synthesis engine generates voice data that resembles the user's voice based on the text data. This voice data is transmitted through the telephone system. The generated voice data is obtained as output and is used as the transmitted voice.
[1025] Step 5:
[1026] The server uses the generated voice data to make a call. The call is made to an external system, such as a restaurant, and the server conducts the conversation in real time. The AI system within the server generates an appropriate response based on the other party's response and continues the conversation. For example, the AI might say, "Could I please make a restaurant reservation for two people tomorrow at 7:00 PM?" and generate the next response based on the restaurant's response. The entire conversation is recorded within the server. The other party's response is obtained as output.
[1027] Step 6:
[1028] The server converts the recorded voice data into text data. The content is then summarized by a generative AI model. For example, if the restaurant responds, "We can make a reservation for two people at 7:00 PM," the resulting summarized text is, "A restaurant reservation has been made for two people at 7:00 PM." The summarized text is generated as the output.
[1029] Step 7:
[1030] The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking the notification, the user can easily understand the content and results of the call. For example, a notification saying "A restaurant reservation has been made for two people at 7:00 PM" is displayed on the user's smartphone. The notification message is obtained as output.
[1031] (Application example 1)
[1032] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1033] In modern society, users experience a great deal of stress when making reservations or inquiries at physical stores over the phone. This stress increases especially when the call is not answered or when appropriate communication is not possible. There is a need to solve this problem and provide a method that allows users to make reservations and inquiries more easily.
[1034] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1035] In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generation AI model means for receiving the text data, analyzing the intention, and generating a dialogue script; voice synthesis means for converting the dialogue script into voice data; means for making a call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for a user to input a reservation or inquiry for a physical store; means for receiving input related to the physical store, generating a dialogue script, and making a call to collect necessary information; and means for notifying the user of the results of the reservation or inquiry for the physical store. This enables a user to easily make a reservation or inquiry for a physical store without having to make a call themselves.
[1036] definition statement
[1037] "Means for users to input information they want to communicate over the phone" refers to an interface that allows users to input details of reservations, inquiries, etc. by voice or text using a smartphone or other device.
[1038] "Means for converting input content into text data" refers to the process of generating text data from the voice or image input by the user using a voice recognition engine or OCR technology.
[1039] A "generative AI model means" is an algorithm or program that uses a generative AI model (e.g., GPT-4) to generate an appropriate dialogue script based on the user's intentions.
[1040] The "voice synthesis means for converting a dialogue script into voice data" is a process that uses voice synthesis technology (e.g., a TTS engine) to convert text data into voice data.
[1041] "Means for making a call using voice data and having a real-time conversation with the other party" refers to a communication technology that uses the generated voice data in an actual telephone conversation to conduct a real-time conversation with the other party.
[1042] "Means for recording the content of a conversation and converting the recorded voice data into text and summarizing it" refers to the process of recording the content of a conversation, converting the recorded data into text data using voice recognition technology, and creating a summary using AI technology.
[1043] The "means for notifying the user of the summarized text" refers to a technology for delivering the generated summarized text to the user's smartphone or other device as a push notification or message.
[1044] "A means by which a user can input reservations or inquiries for a physical store" is an interface that allows a user to input reservation or inquiry information about a specific store.
[1045] "Means of receiving input about a physical store, generating a dialogue script, and making a phone call to collect the necessary information" refers to the process in which the generative AI model creates a dialogue script based on the information about the physical store entered by the user, and makes a phone call to obtain the necessary information.
[1046] "Means for notifying users of the results of reservations and inquiries at physical stores" refers to technology for notifying users of the results of reservations and inquiries obtained via their terminals.
[1047] MODE FOR CARRYING OUT THE INVENTION
[1048] The present invention is a system that automates the process of making reservations or inquiries about physical stores, reducing stress for users. This system operates through a smartphone application, generates a dialogue script using the generated AI model, and makes phone calls to collect necessary information. The following describes an embodiment of the present invention.
[1049] 1. Basic configuration
[1050] This system consists of the following main hardware and software:
[1051] User terminal (smartphone): A device where a user enters input and receives results. It includes a speech recognition engine and a text input interface.
[1052] Server: Receives user input, generates a dialogue script using a generative AI model, creates voice data using a speech synthesis engine, and processes the call.
[1053] Generative AI model: A technology that generates dialogue scripts based on user requests. An example of this is GPT-4.
[1054] Speech synthesis engine: A technology that converts text data into voice data. An example is Google Text-to-Speech (TTS).
[1055] 2. System processing flow
[1056] A user starts a smartphone application and inputs reservation or inquiry details. For example, a user inputs a prompt such as "I would like to make a restaurant reservation for two people tomorrow at 7:00 PM." This input can be done either by voice or text.
[1057] When voice input is used, the device uses a voice recognition engine such as Google Speech Recognition to convert the voice into text data and then sends the text data to the server.
[1058] The server receives the text data and uses a generative AI model (e.g., GPT-4) to generate a dialogue script based on the user's intention. The generated dialogue script is in the form of "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM."
[1059] The server then uses Google Text-to-Speech (TTS) to convert the generated dialogue script into voice data, and makes an actual call through the telephone system to interact with the physical store in real time.
[1060] During the call process, the server records the conversation, and after the call is over, it converts the audio data back into text data and uses AI technology to create a summary, such as "A restaurant reservation has been made for two people at 7:00 PM."
[1061] The completed summary is sent from the server to the device, and the results are communicated to the user via push notification or message. By checking this notification, the user can immediately understand the reservation details and the results of the inquiry.
[1062] 3. Specific Examples
[1063] For example, if a user launches the "Store Assistant AI" app on their smartphone and voice-inputs, "I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM," the system will operate as follows:
[1064] 1. The user enters the prompt, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1065] 2. The device converts the voice into text and sends it to the server.
[1066] 3. The server uses the generative AI model to generate a dialogue script that reads, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1067] 4. The server uses a speech synthesis engine to convert the dialogue script into voice data.
[1068] 5. The server makes a call and communicates with the cafe in real time to make the reservation.
[1069] 6. The server records the conversation and creates a summary.
[1070] 7. A notification is sent to the device stating, "A cafe reservation has been made for two people at 7:00 PM."
[1071] This allows users to easily make reservations or inquiries without having to make a phone call themselves.
[1072] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1073] Program processing steps
[1074] Step 1:
[1075] The user launches the smartphone application and inputs the details of the reservation or inquiry. Input can be done by voice or text. For example, the user inputs, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1076] Input: User voice or text input
[1077] Output: Audio or text data
[1078] Step 2:
[1079] The device receives user input and, in the case of voice input, converts the speech into text data using a speech recognition engine such as Google Speech Recognition.
[1080] Input: Audio data
[1081] Output: Text data
[1082] Step 3:
[1083] The device sends the converted text data to a server, which receives the text data and uses a generative AI model (e.g., GPT-4) to analyze the user's intent.
[1084] Input: Text data
[1085] Output: Intention analysis results
[1086] Step 4:
[1087] The server generates a dialogue script based on the intent analysis results. The generated dialogue script might be in the form of, for example, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1088] Input: Intention analysis result
[1089] Output: Interactive script
[1090] Step 5:
[1091] The server converts the generated dialogue script into voice data using the Google Text-to-Speech (TTS) engine, which is then used in the actual phone call.
[1092] Input: Interactive script
[1093] Output: Audio data
[1094] Step 6:
[1095] The server uses the voice data to make calls and communicate with the physical store in real time, and the conversation is recorded by the server during the conversation.
[1096] Input: Audio data
[1097] Output: Recorded dialogue
[1098] Step 7:
[1099] The server uses voice recognition technology to convert the recorded conversation into text data, and then uses AI technology to create a summary. The summarized text data might be in the form of, for example, "A reservation has been made for two people at a cafe at 7:00 PM."
[1100] Input: Recorded dialogue
[1101] Output: Summary text
[1102] Step 8:
[1103] The server sends the summary text to the user's terminal, and the user can check the notification to understand the reservation details.
[1104] Input: Summary text
[1105] Output: Information message
[1106] The above is the flow of specific processing performed by the system, and the operations include data processing and data calculation performed at each step.
[1107] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1108] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains the necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is sensitive to those emotions. The following describes in detail the embodiments of the invention.
[1109] Input and Initial Processing
[1110] 1. User:
[1111] Users launch a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries or reservations. Input can be done either by voice or text, and the emotion engine recognizes the user's emotions as they type.
[1112] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1113] 2. Terminal:
[1114] The device receives user input. In the case of voice input, the device converts the voice data into text data using a voice recognition engine. The converted text data and the user's emotion information are sent to the server.
[1115] AI processing and phone generation
[1116] 3. Server:
[1117] The server receives text data and emotional information from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intentions from the text data and generates a dialogue script that takes the emotional information into account.
[1118] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1119] 4. Server (Speech synthesis):
[1120] The generated dialogue script is input into a speech synthesis engine and converted into voice data that resembles the user's voice. At this time, the tone and expression of the voice are adjusted based on emotional information generated by the emotion engine.
[1121] Example: The generated speech data is, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1122] 5. Server (Calling):
[1123] The voice data is transmitted through a telephone system to place a call to a destination (for example, a restaurant).
[1124] Real-time dialogue and results notification
[1125] 6. Server:
[1126] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[1127] 7. Server (Text Conversion and Summarization):
[1128] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[1129] 8. Notification from the server to the device:
[1130] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[1131] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1132] Specific example explanation
[1133] If a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the specific flow is as follows:
[1134] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." The system simultaneously recognizes the user's emotions (for example, if the user is feeling stressed).
[1135] 2. The device converts the speech into text and sends it to the server along with emotional information.
[1136] 3. The server uses the generative AI model to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1137] 4. The server uses a speech synthesis engine to generate voice data that reflects the user's voice and emotions based on the emotional information.
[1138] 5. The server makes the call and communicates with the other party in real time.
[1139] 6. The server records the conversation, converts it into text, and creates a summary.
[1140] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[1141] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[1142] The processing flow will be explained below.
[1143] Step 1:
[1144] The user launches a dedicated smartphone app and inputs by voice or text what they want to communicate over the phone, such as an inquiry or reservation. Once input is complete, the emotion engine recognizes the user's emotions.
[1145] Example: A user speaks, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[1146] Step 2:
[1147] The terminal receives input voice data and converts the voice data into text data using a voice recognition engine.
[1148] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1149] Step 3:
[1150] The terminal transmits the converted text data and the recognized emotion information to the server.
[1151] Step 4:
[1152] The server analyzes the text data and emotional information received from the device. It uses a generative AI model to understand the user's intentions and generate a dialogue script. This script is generated by reflecting the emotional information.
[1153] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1154] Step 5:
[1155] The server inputs the generated dialogue script into a speech synthesis engine to generate voice data that resembles the user's voice. The tone and expression of the voice are adjusted based on the emotions recognized by the emotion engine.
[1156] Example: Speech data: "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1157] Step 6:
[1158] The server transmits the generated voice data to a telephone system, which places a call to the destination (e.g., a restaurant).
[1159] Step 7:
[1160] The server establishes a telephone connection, transmits the generated voice data to the other party, and initiates a dialogue. The dialogue with the other party is conducted in real time, and an appropriate response is generated based on the other party's responses.
[1161] Example: If the reply is "I have availability at 7:00 PM," generate a response of "Thank you. I'd like to meet you at that time."
[1162] Step 8:
[1163] The server records the entire phone conversation.
[1164] Step 9:
[1165] The server converts the recorded audio data into text and creates a summary.
[1166] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[1167] Step 10:
[1168] The server transmits the summarized text data to the terminal.
[1169] Step 11:
[1170] The text data received by the device is notified to the user via push notification, email, etc.
[1171] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1172] Through the above series of processing steps, the user can efficiently obtain the necessary information in a manner that is in tune with their emotions, without having to make a phone call themselves.
[1173] Example 2
[1174] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1175] Conventional telephone reservation systems have problems such as the stress users feel when making phone calls themselves and the inability to efficiently obtain the necessary information. Furthermore, dialogue systems that proceed without taking the user's emotions into consideration often fail to provide the quality of dialogue users expect. There is a need for a system that can solve these problems and allow users to make reservations and inquiries more comfortably.
[1176] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for allowing a user to input the content of an inquiry or reservation by voice or text and analyzing the emotion at the time of input; means for converting the input voice data into text data; means for receiving the text data and emotion information, analyzing the intention using a generative AI model, and generating a dialogue script that takes emotion into consideration; voice synthesis means for converting the generated dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the content of the conversation, converting the recorded voice data into text, and summarizing it; and means for notifying the user's terminal of the summarized text. This reduces the stress of the user making a phone call and enables efficient information acquisition through a dialogue that is sensitive to the user's emotions.
[1177] "User" refers to a person who uses the system.
[1178] A "dedicated app" refers to application software that runs on devices such as smartphones and allows users to enter details of inquiries and reservations.
[1179] "Emotion engine" refers to software or algorithms for analyzing emotions from a user's voice or input data.
[1180] "Speech recognition engine" refers to software or algorithms for converting input voice data into text data.
[1181] A "generative AI model" refers to an artificial intelligence model that analyzes user intent based on text data and emotional information, and generates an appropriate dialogue script.
[1182] A "dialogue script" refers to a text-based scenario for progressing a dialogue, which is generated based on the user's intentions.
[1183] "Speech synthesis engine" refers to software or algorithms for converting text data into speech data.
[1184] "Telephone System" means the communications infrastructure and software used to make telephone calls based on generated voice data.
[1185] "Text conversion means" refers to a system or technology for converting voice data into text data.
[1186] "Summarization means" refers to a system or technology for concisely summarizing the converted text data.
[1187] "Push notification" refers to a technology that allows a server to send information to a user's device in real time.
[1188] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is considerate of those emotions. The following describes in detail the embodiments of the invention.
[1189] 1. User Input Procedure
[1190] Users launch a dedicated smartphone app and input the details of their inquiry or reservation. For voice input, users press the microphone button and start speaking. For text input, users type directly into the text field. As the user types, the app uses an emotion engine to analyze the user's emotions.
[1191] Specific hardware or software you will be using:
[1192] Dedicated app (smartphone app)
[1193] Emotion Engine
[1194] 2. Initial processing on the device
[1195] The device receives the input voice data and converts it into text data using a speech recognition engine. The converted text data and emotion information are then sent to the server. A speech recognition API such as Google Speech-to-Text is used here.
[1196] Specific hardware or software you will be using:
[1197] Speech recognition engine (Google Speech-to-Text)
[1198] 3. Data processing on the server
[1199] The server inputs the received text data and emotional information into a generative AI model to analyze the user's intentions. Generative AI models such as OpenAI GPT-4 are used here. The generated dialogue script takes the user's emotions into account.
[1200] Specific hardware or software you will be using:
[1201] Generative AI model (OpenAI GPT-4)
[1202] 4. Voice conversion of dialogue scripts
[1203] The generated dialogue script is passed to a speech synthesis engine, which converts it into voice data that reflects the user's voice and emotions. A speech synthesis engine such as Amazon Polly is used for this purpose.
[1204] Specific hardware or software you will be using:
[1205] Speech synthesis engine (Amazon Polly)
[1206] 5. Calling via the telephone system
[1207] Using voice data, calls are made through a telephone system and a real-time conversation is held with the other party. The server analyzes the other party's response in real time and uses a generative AI model to generate the next appropriate response and continue the conversation. Telephone systems such as Twilio are used.
[1208] Specific hardware or software you will be using:
[1209] Phone system (Twilio)
[1210] 6. Recording and summarizing conversation content
[1211] After the call is over, the server converts the recorded voice data into text and summarizes the key points, which are then stored in the system and later shared with the user.
[1212] Specific hardware or software you will be using:
[1213] Text conversion methods
[1214] Summary tools
[1215] 7. Notification of Results
[1216] The server sends the summarized text data to the device, and the user is notified of the results by push notification or other means. The user can check the call results and next steps through the app.
[1217] Specific hardware or software you will be using:
[1218] Push notification system
[1219] Specific examples
[1220] The specific flow when a user wants to make a restaurant reservation for tomorrow at 19:00 is shown below.
[1221] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." At this time, the system recognizes that the user is feeling stressed.
[1222] 2. The device converts the speech into text and sends it to the server along with emotional information.
[1223] 3. The server uses a generative AI model (e.g., GPT-4) to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know availability."
[1224] 4. The server uses a speech synthesis engine (e.g., Amazon Polly) to generate voice data and adjust it to reflect the emotion.
[1225] 5. The server uses a telephone system (e.g., Twilio) to make a call and communicate with the restaurant in real time.
[1226] 6. The server converts the call into text and creates a summary.
[1227] 7. The server sends the summary to the user's terminal and notifies them of the result: "A restaurant reservation has been made for two people at 7:00 PM."
[1228] Example prompt sentence:
[1229] "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[1230] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[1231] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1232] Step 1:
[1233] Users launch a dedicated smartphone app and input the details of their inquiry or reservation by voice or text. Once input is complete, the app uses an emotion engine to analyze the user's emotions. The input data is output as text data (in the case of text input) or voice data (in the case of voice input). Emotional data is also output at the same time.
[1234] Specific behavior:
[1235] The user speaks, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1236] The app receives the voice data and analyzes the emotions (e.g., tension or stress) present in the input.
[1237] Step 2:
[1238] The device receives the user's voice data and converts it into text data using a speech recognition engine. The text data and analyzed emotion information are sent to the server. The input is voice data, and the output is text data and emotion information.
[1239] Specific behavior:
[1240] The terminal converts the voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text).
[1241] The converted text data and emotion data are sent to the server.
[1242] Step 3:
[1243] The server receives the text data and emotional information sent from the device. The text data and emotional information are input into the generative AI model, which analyzes the user's intentions. The generated dialogue script is output. The input is text data and emotional information, and the output is a dialogue script.
[1244] Specific behavior:
[1245] The server inputs the data into a generative AI model such as GPT-4, which analyzes the user's intent.
[1246] The generative AI model generates a dialogue script that reads, "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1247] Step 4:
[1248] The server passes the generated dialogue script to a speech synthesis engine and converts it into voice data. The tone and expression of the voice are also adjusted based on the emotional information. The input is the dialogue script and emotional information, and the output is voice data.
[1249] Specific behavior:
[1250] The server inputs the dialogue script into a speech synthesis engine such as Amazon Polly.
[1251] A speech synthesis engine converts the dialogue script into voice data that reflects the user's voice and emotions.
[1252] Step 5:
[1253] The server uses the generated voice data to make a call through the telephone system, initiating a real-time dialogue with the other party, and the server analyzes the other party's response in real time and generates the next appropriate response using a generative AI model. The input is the voice data, and the output is a real-time response.
[1254] Specific behavior:
[1255] The server calls the restaurant using a phone system such as Twilio.
[1256] The server analyzes the other party's response and uses a generative AI model to generate an appropriate response and continue the dialogue.
[1257] Step 6:
[1258] After the call ends, the server converts the recorded voice data into text and summarizes the key points. The converted text data and the summary are output. The input is the voice data, and the output is the text data and the summary.
[1259] Specific behavior:
[1260] The server uses a speech recognition engine to convert the call into text.
[1261] The server uses a summarization algorithm to summarize the converted text.
[1262] Step 7:
[1263] The server sends the summarized text data to the terminal, and the user is notified of the results via push notification, etc. The input is the summarized text data, and the output is a notification to the user.
[1264] Specific behavior:
[1265] The server sends a summary to the terminal stating, "A restaurant reservation has been made for two people at 19:00."
[1266] The user receives a push notification and checks for details.
[1267] (Application example 2)
[1268] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1269] In conventional food delivery services, when users place orders over the phone, they are prone to making mistakes in how they communicate their orders. Furthermore, there are few ways to reduce the stress and tension users feel when ordering over the phone. Furthermore, there are no systems that can process orders in a way that takes into account the user's emotions. This creates a need for a user-friendly and efficient food delivery ordering system.
[1270] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generative model means for receiving the text data, analyzing the intention, and generating a dialogue script; speech synthesis means for converting the dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for recognizing the user's emotion and optimizing the dialogue script and voice data based on the emotion; input means for allowing the user to select either voice or character input; and means for acquiring emotion information and adjusting the tone and expression of the voice based on the emotion information. This enables the user to efficiently order food delivery in a manner that is sensitive to their emotion without having to make a phone call.
[1271] "User" refers to an individual or corporation that uses this system to make phone calls, make reservations, place orders, and perform other operations.
[1272] "Content to be communicated over the phone" refers to information that the user needs to communicate to the other party over the phone, such as an order, an inquiry, or a reservation.
[1273] "Input means" refers to an interface or device used by a user to input content by voice or text.
[1274] "Text data" refers to text information converted from voice input or text input.
[1275] "Generative model means" refers to the AI model or algorithm used to analyze user input and generate an appropriate dialogue script.
[1276] "Speech synthesis means" refers to a technology or device that converts the generated dialogue script into voice data.
[1277] "Voice data" refers to digital information of voice generated by a voice synthesis means.
[1278] "Means for making a call" refers to a system or device for making a call using the generated voice data.
[1279] "Means of real-time interaction" refers to technology or systems that allow you to have a direct conversation with the person you are calling.
[1280] "Means for recording conversation content" refers to technology or devices that save telephone conversations as audio data.
[1281] "Means of converting to text and summarizing" refers to the technology or algorithms that convert recorded audio data into text and summarize that text concisely.
[1282] "Means for notifying the device" refers to technology or systems that notify the user of the summarized text on their device, such as a smartphone or tablet.
[1283] "Means of recognizing emotions" refers to technologies and models for analyzing and extracting emotional information from user input and dialogue content.
[1284] "Means for optimizing dialogue scripts and voice data based on emotions" refers to technologies and algorithms that use emotional information to adjust the tone and expression of the dialogue scripts and voice data that are generated.
[1285] An input means that allows the user to select between "voice and text input" refers to an interface or device that allows the user to freely select either voice input or text input.
[1286] This invention provides a specific implementation method for recognizing user emotions and efficiently completing orders in a food delivery ordering system.
[1287] System Configuration
[1288] The present invention is mainly composed of a user terminal, a server, and an AI module.
[1289] User terminal
[1290] To order food delivery, a user first uses a user device such as a smartphone or tablet. The user then launches a dedicated app and can enter the order details by voice or text. The user's emotions are recognized as they are entered.
[1291] server
[1292] The server receives input data (voice or text) and emotional information sent from the user's device. In the case of voice input, the server converts the voice data into text data using a voice recognition engine. Based on this converted text data and emotional information, a generative AI model generates a dialogue script. The generated dialogue script is then input into a voice synthesis engine and converted into voice data. At this time, the tone and expression of the voice are optimized based on the emotional information.
[1293] AI Module
[1294] The generative AI model and speech synthesis engine are included in the AI module. The generative AI model analyzes the user's intentions and generates a dialogue script that takes emotional information into account. The speech synthesis engine converts this dialogue script into voice data.
[1295] communication means
[1296] The server uses the generated voice data to make calls to the corresponding restaurant or delivery provider and engages in real-time conversations, which are recorded in the server and converted into text data as needed.
[1297] User Notification
[1298] Once the telephone conversation is complete, the server summarizes the conversation and sends the summarized information to the user's device via a push notification or other method, allowing the user to confirm whether the order was placed successfully.
[1299] Specific examples
[1300] For example, if a user says "I would like to order a pizza from Restaurant X", the system will:
[1301] 1. The user device receives voice input, and the emotion recognition engine analyzes the user's emotions.
[1302] 2. The server converts the voice data into text, and the generative AI model generates a dialogue script. The generated dialogue script is, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita pizza, please."
[1303] 3. Based on this dialogue script, the speech synthesis engine optimizes the tone and expression of the voice according to the emotional information and generates the voice data.
[1304] 4. The server will make a call and automatically communicate the order details.
[1305] 5. The restaurant's response is recorded, and the content of the conversation is converted into text data and summarized.
[1306] 6. The summary is sent to the user's terminal, and the user is informed that "your order has been accepted. Delivery time is 18:00."
[1307] Prompt Sentence Examples
[1308] An example of a prompt to pass to a generative AI model would be:
[1309] "Hello, I'd like to order a pizza from XX Restaurant. I'd like one Margherita, please."
[1310] Emotion: Stress
[1311] This allows the system to efficiently order food delivery while being sensitive to the user's emotions.
[1312] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1313] Step 1:
[1314] The user launches a dedicated smartphone app and inputs their order details by voice or text. The emotion recognition engine analyzes the user's emotions as they input their order. For example, if a user inputs "I'd like to order pizza from Restaurant X," the device captures the voice data and analyzes and obtains emotional information.
[1315] Step 2:
[1316] The voice data acquired by the device is converted into text data. A voice recognition engine (e.g., Google Speech-to-Text API) is used to convert the voice data into text data. For example, a voice saying "I would like to order pizza from XX Restaurant" is converted into text data saying "I would like to order pizza from XX Restaurant."
[1317] Step 3:
[1318] Text data and emotional information are sent from the terminal to the server. The order details (text data) and emotional information (e.g., stress) entered by the user are sent to the server and used as input for the next process.
[1319] Step 4:
[1320] Based on the text data and emotional information received by the server, a generative AI model (e.g., GPT-3) is used to generate a dialogue script. The generative AI model analyzes the given text and emotional information and generates an appropriate dialogue script. For example, it generates a script that reads, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please."
[1321] Step 5:
[1322] The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data. At this time, the tone and expression of the voice are adjusted based on the emotional information. For example, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please" is generated in a soft voice that matches the emotion.
[1323] Step 6:
[1324] The server uses the generated voice data to place a call to the specified restaurant and execute the order. This telephone conversation takes place in real time, and the server accurately conveys the order details.
[1325] Step 7:
[1326] The server records the audio data during the conversation, including the restaurant's responses and questions.
[1327] Step 8:
[1328] The server converts the recorded voice data into text and summarizes it. The speech recognition engine is used again to convert the voice recording into text data, and the content is concisely summarized using a summarization algorithm. For example, a summary text such as "Your order has been received. Delivery time is 6:00 PM" is generated.
[1329] Step 9:
[1330] The server sends the summarized text to the user's device and sends a push notification with the order details, which displays a notification such as "Your order has been accepted. Delivery time is 6:00 PM."
[1331] This allows users to efficiently place food delivery orders in a way that is in tune with their emotions, without having to make a phone call themselves.
[1332] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1334] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1335] [Fourth embodiment]
[1336] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1337] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1338] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1339] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1340] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1341] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1342] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1343] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1344] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1345] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1346] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1347] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1348] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1349] The present invention is a system developed to reduce the stress users feel when making phone calls. This system uses AI to make calls on behalf of users, obtain necessary information, and notify the user of the results without the user having to make the calls themselves. The following describes in detail the embodiments of the present invention.
[1350] Input and Initial Processing
[1351] 1. User:
[1352] Users start up a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[1353] Example: If a user enters "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1354] 2. Terminal:
[1355] The device receives user input. In the case of voice input, it uses a speech recognition engine to convert the voice data into text data, which is then sent to the server.
[1356] AI processing and phone generation
[1357] 3. Server:
[1358] The server receives text data from the device and passes it to a generative AI model (e.g., a model using natural language processing). The AI model analyzes the user's intent and generates an appropriate dialogue script.
[1359] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1360] 4. Server (Speech synthesis):
[1361] The generated dialogue script is input to a speech synthesis engine to generate voice data that resembles the user's voice. This voice data is then transmitted over the telephone system.
[1362] Real-time dialogue and results notification
[1363] 5. Server:
[1364] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[1365] 6. Server (Text Conversion and Summarization):
[1366] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[1367] 7. Notification from the server to the device:
[1368] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[1369] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1370] Specific example explanation
[1371] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[1372] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1373] 2. The device converts the speech into text and sends it to the server.
[1374] 3. The server uses the generative AI model to analyze the input and generate an appropriate dialogue script (e.g., "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability.").
[1375] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[1376] 5. The server communicates with the other party in real time and records the conversation.
[1377] 6. The server converts the recorded audio data into text and creates a summary.
[1378] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[1379] In this way, the system of the present invention allows the user to obtain the necessary information without stress, without having to make a phone call himself.
[1380] The processing flow will be explained below.
[1381] Step 1:
[1382] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as inquiries, reservations, etc. Input can be done either by voice or text.
[1383] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1384] Step 2:
[1385] The device receives user input. In the case of voice input, the device's voice recognition engine is used to convert the voice data into text data.
[1386] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1387] Step 3:
[1388] The terminal transmits the converted text data to the server.
[1389] Step 4:
[1390] The server receives the text data sent from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intent from the text data.
[1391] Example: Analyze the text data "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1392] Step 5:
[1393] The server generates an interactive script based on the analysis results.
[1394] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1395] Step 6:
[1396] The server inputs the generated dialogue script into a voice synthesis engine to generate voice data that resembles the user's voice.
[1397] Example: The dialogue script generates the following speech data: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1398] Step 7:
[1399] The server sends the voice data to the telephone system, which then places a call to the destination (e.g., a restaurant).
[1400] Step 8:
[1401] The server establishes a telephone connection and transmits the generated voice data to the other party to start the conversation.
[1402] Step 9:
[1403] The server analyzes the response from the other party in real time and uses AI to generate an appropriate response.
[1404] Example: If a restaurant employee replies, "We have a table available at 7:00 PM," the AI generates a response: "Thank you. We'll make it that time then."
[1405] Step 10:
[1406] The server records the entire phone conversation.
[1407] Step 11:
[1408] The server converts the recorded audio data into text.
[1409] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[1410] Step 12:
[1411] The server summarizes the converted text data, extracts important information, and summarizes it compactly.
[1412] Example: The summary you get is "A restaurant reservation has been made for two people at 7:00 PM."
[1413] Step 13:
[1414] The server transmits the summarized text data to the terminal.
[1415] Step 14:
[1416] The text data received by the device is notified to the user via push notification, email, etc.
[1417] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1418] This series of steps allows the user to efficiently obtain the necessary information without having to make a phone call.
[1419] Example 1
[1420] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1421] In modern society, making reservations and inquiries over the phone is common, but it can cause stress and anxiety for many users. In particular, for users who are not good at communicating over the phone or who are busy, there is a demand for a system that can handle phone calls efficiently and easily. Conventional systems require users to make the calls themselves, which is time-consuming and mentally taxing. To solve this problem, the present invention aims to provide an automated telephone answering system that eliminates the need for users to make phone calls.
[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1423] In this invention, the server includes means for inputting the content that the user wants to communicate over the phone, means for converting the input content into text data, and generative model means for receiving the text data, analyzing the intention, and generating a dialogue script. This allows the system to automatically make calls on behalf of the user, obtain the necessary information, and notify the user of the results, without the user having to make the call directly.
[1424] A "user" is a person or individual who uses the system to input what they want to say over the phone.
[1425] A "means" is a device or method used to achieve a particular purpose or function.
[1426] "Text data" refers to character-based information converted from audio data or other formats.
[1427] A "generative model" is an algorithm that uses natural language processing or machine learning to generate sentences that meet a specific purpose from specific input.
[1428] "Speech synthesis" is a technology that generates speech that sounds like human speech based on text data.
[1429] An "information processing device" is an electronic device for receiving and processing data, and typically includes smartphones, tablets, personal computers, etc.
[1430] An "external system" is an object other than the user's system, such as a restaurant or an office.
[1431] This invention is an automated telephone answering system developed to reduce the stress of users when making phone calls and to efficiently obtain the necessary information. The system aims to have users input the information they want to convey over the phone, and then have AI make the call on their behalf, obtain the information, and notify the user of the results.
[1432] Input and Initial Processing
[1433] 1. The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input can be done either by voice or text. Consider the example where the user inputs, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[1434] 2. The device receives the user's voice input and converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text API), which is then sent to the server.
[1435] AI processing and phone generation
[1436] 3. The server receives the text data sent from the device and passes it to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the user's intent and generates an appropriate dialogue script. The generated dialogue script includes the following content: "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1437] 4. The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data that resembles the user's voice. This voice data is then transmitted through the telephone system.
[1438] Real-time dialogue and results notification
[1439] 5. The server uses the generated voice to make a call and engage in a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[1440] 6. The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the important points.
[1441] 7. The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking this notification, the user can easily understand the content and results of the call. For example, a notification may appear on the device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1442] Specific examples
[1443] For example, if a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the flow would be as follows:
[1444] 1. The user speaks to the app, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1445] 2. The device converts the speech into text and sends it to the server.
[1446] 3. The server analyzes the input using a generative AI model and generates an appropriate dialogue script.
[1447] 4. The server converts the dialogue script into voice data and sends it through the telephone system.
[1448] 5. The server communicates with the other party in real time and records the conversation.
[1449] 6. The server converts the recorded audio data into text and summarizes the content.
[1450] 7. The server notifies the terminal of the summary, and the user confirms the results.
[1451] An example of a prompt sentence could be, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM. Please let me know the availability." In this way, the user can obtain the necessary information without stress, without having to make a phone call themselves.
[1452] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1453] Step 1:
[1454] The user launches a dedicated smartphone app and inputs the information they want to communicate over the phone, such as an inquiry or reservation. Input methods include both voice input and text input. For example, a user might say, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." This will result in either voice data or text data being acquired as input data.
[1455] Step 2:
[1456] The device receives the user's input data. The received voice data is converted into text data using a speech recognition engine such as the Google Speech-to-Text API. During this conversion process, the voice data is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM." Text data is generated as output, and this text data is sent to the server.
[1457] Step 3:
[1458] The server receives the text data sent from the device. The received text data is passed to a generative AI model (for example, OpenAI's GPT-3). The generative AI model analyzes the intent of the input text data and generates an appropriate dialogue script. For example, from the input data "I would like to make a restaurant reservation for tomorrow at 7:00 PM," the dialogue script generated is "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability." The dialogue script is obtained as the output.
[1459] Step 4:
[1460] The server inputs the generated dialogue script into a speech synthesis engine (for example, Google Text-to-Speech API). The speech synthesis engine generates voice data that resembles the user's voice based on the text data. This voice data is transmitted through the telephone system. The generated voice data is obtained as output and is used as the transmitted voice.
[1461] Step 5:
[1462] The server uses the generated voice data to make a call. The call is made to an external system, such as a restaurant, and the server conducts the conversation in real time. The AI system within the server generates an appropriate response based on the other party's response and continues the conversation. For example, the AI might say, "Could I please make a restaurant reservation for two people tomorrow at 7:00 PM?" and generate the next response based on the restaurant's response. The entire conversation is recorded within the server. The other party's response is obtained as output.
[1463] Step 6:
[1464] The server converts the recorded voice data into text data. The content is then summarized by a generative AI model. For example, if the restaurant responds, "We can make a reservation for two people at 7:00 PM," the resulting summarized text is, "A restaurant reservation has been made for two people at 7:00 PM." The summarized text is generated as the output.
[1465] Step 7:
[1466] The server sends the summarized text data to the device, and the user is notified of the results by means of a push notification or other method. By checking the notification, the user can easily understand the content and results of the call. For example, a notification saying "A restaurant reservation has been made for two people at 7:00 PM" is displayed on the user's smartphone. The notification message is obtained as output.
[1467] (Application example 1)
[1468] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1469] In modern society, users experience a great deal of stress when making reservations or inquiries at physical stores over the phone. This stress increases especially when the call is not answered or when appropriate communication is not possible. There is a need to solve this problem and provide a method that allows users to make reservations and inquiries more easily.
[1470] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1471] In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generation AI model means for receiving the text data, analyzing the intention, and generating a dialogue script; voice synthesis means for converting the dialogue script into voice data; means for making a call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for a user to input a reservation or inquiry for a physical store; means for receiving input related to the physical store, generating a dialogue script, and making a call to collect necessary information; and means for notifying the user of the results of the reservation or inquiry for the physical store. This enables a user to easily make a reservation or inquiry for a physical store without having to make a call themselves.
[1472] definition statement
[1473] "Means for users to input information they want to communicate over the phone" refers to an interface that allows users to input details of reservations, inquiries, etc. by voice or text using a smartphone or other device.
[1474] "Means for converting input content into text data" refers to the process of generating text data from the voice or image input by the user using a voice recognition engine or OCR technology.
[1475] A "generative AI model means" is an algorithm or program that uses a generative AI model (e.g., GPT-4) to generate an appropriate dialogue script based on the user's intentions.
[1476] The "voice synthesis means for converting a dialogue script into voice data" is a process that uses voice synthesis technology (e.g., a TTS engine) to convert text data into voice data.
[1477] "Means for making a call using voice data and having a real-time conversation with the other party" refers to a communication technology that uses the generated voice data in an actual telephone conversation to conduct a real-time conversation with the other party.
[1478] "Means for recording the content of a conversation and converting the recorded voice data into text and summarizing it" refers to the process of recording the content of a conversation, converting the recorded data into text data using voice recognition technology, and creating a summary using AI technology.
[1479] The "means for notifying the user of the summarized text" refers to a technology for delivering the generated summarized text to the user's smartphone or other device as a push notification or message.
[1480] "A means by which a user can input reservations or inquiries for a physical store" is an interface that allows a user to input reservation or inquiry information about a specific store.
[1481] "Means of receiving input about a physical store, generating a dialogue script, and making a phone call to collect the necessary information" refers to the process in which the generative AI model creates a dialogue script based on the information about the physical store entered by the user, and makes a phone call to obtain the necessary information.
[1482] "Means for notifying users of the results of reservations and inquiries at physical stores" refers to technology for notifying users of the results of reservations and inquiries they have made to their terminals.
[1483] MODE FOR CARRYING OUT THE INVENTION
[1484] The present invention is a system that automates the process of making reservations or inquiries about physical stores, reducing stress for users. This system operates through a smartphone application, generates a dialogue script using the generated AI model, and makes phone calls to collect necessary information. The following describes an embodiment of the present invention.
[1485] 1. Basic configuration
[1486] This system consists of the following main hardware and software:
[1487] User terminal (smartphone): A device where a user enters input and receives results. It includes a speech recognition engine and a text input interface.
[1488] Server: Receives user input, generates a dialogue script using a generative AI model, creates voice data using a speech synthesis engine, and processes the call.
[1489] Generative AI model: A technology that generates dialogue scripts based on user requests. An example of this is GPT-4.
[1490] Speech synthesis engine: A technology that converts text data into voice data. An example is Google Text-to-Speech (TTS).
[1491] 2. System processing flow
[1492] A user starts a smartphone application and inputs reservation or inquiry details. For example, a user inputs a prompt such as "I would like to make a restaurant reservation for two people tomorrow at 7:00 PM." This input can be done either by voice or text.
[1493] When voice input is used, the device uses a voice recognition engine such as Google Speech Recognition to convert the voice into text data and then sends the text data to the server.
[1494] The server receives the text data and uses a generative AI model (e.g., GPT-4) to generate a dialogue script based on the user's intention. The generated dialogue script is in the form of "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM."
[1495] The server then uses Google Text-to-Speech (TTS) to convert the generated dialogue script into voice data, and makes an actual call through the telephone system to interact with the physical store in real time.
[1496] During the call process, the server records the conversation, and after the call is over, it converts the audio data back into text data and uses AI technology to create a summary, such as "A restaurant reservation has been made for two people at 7:00 PM."
[1497] The completed summary is sent from the server to the device, and the results are communicated to the user via push notification or message. By checking this notification, the user can immediately understand the reservation details and the results of the inquiry.
[1498] 3. Specific Examples
[1499] For example, if a user launches the "Store Assistant AI" app on their smartphone and voice-inputs, "I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM," the system will operate as follows:
[1500] 1. The user enters the prompt, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1501] 2. The device converts the voice into text and sends it to the server.
[1502] 3. The server uses the generative AI model to generate a dialogue script that reads, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1503] 4. The server uses a speech synthesis engine to convert the dialogue script into voice data.
[1504] 5. The server makes a call and communicates with the cafe in real time to make the reservation.
[1505] 6. The server records the conversation and creates a summary.
[1506] 7. A notification is sent to the device stating, "A cafe reservation has been made for two people at 7:00 PM."
[1507] This allows users to easily make reservations or inquiries without having to make a phone call themselves.
[1508] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1509] Program processing steps
[1510] Step 1:
[1511] The user launches the smartphone application and inputs the details of the reservation or inquiry. Input can be done by voice or text. For example, the user inputs, "I would like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1512] Input: User voice or text input
[1513] Output: Audio or text data
[1514] Step 2:
[1515] The device receives user input and, in the case of voice input, converts the speech into text data using a speech recognition engine such as Google Speech Recognition.
[1516] Input: Audio data
[1517] Output: Text data
[1518] Step 3:
[1519] The device sends the converted text data to a server, which receives the text data and uses a generative AI model (e.g., GPT-4) to analyze the user's intent.
[1520] Input: Text data
[1521] Output: Intention analysis results
[1522] Step 4:
[1523] The server generates a dialogue script based on the intent analysis results. The generated dialogue script might be in the form of, for example, "Hello, I'd like to make a reservation at a cafe for two people tomorrow at 7:00 PM."
[1524] Input: Intention analysis result
[1525] Output: Interactive script
[1526] Step 5:
[1527] The server converts the generated dialogue script into voice data using the Google Text-to-Speech (TTS) engine, which is then used in the actual phone call.
[1528] Input: Interactive script
[1529] Output: Audio data
[1530] Step 6:
[1531] The server uses the voice data to make calls and communicate with the physical store in real time, and the conversation is recorded by the server during the conversation.
[1532] Input: Audio data
[1533] Output: Recorded dialogue
[1534] Step 7:
[1535] The server uses voice recognition technology to convert the recorded conversation into text data, and then uses AI technology to create a summary. The summarized text data might be in the form of, for example, "A reservation has been made for two people at a cafe at 7:00 PM."
[1536] Input: Recorded dialogue
[1537] Output: Summary text
[1538] Step 8:
[1539] The server sends the summary text to the user's terminal, and the user can check the notification to understand the reservation details.
[1540] Input: Summary text
[1541] Output: Information message
[1542] The above is the flow of specific processing performed by the system, and the operations include data processing and data calculation performed at each step.
[1543] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1544] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains the necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is sensitive to those emotions. The following describes in detail the embodiments of the invention.
[1545] Input and Initial Processing
[1546] 1. User:
[1547] Users launch a dedicated smartphone app and input the information they want to communicate over the phone, such as inquiries or reservations. Input can be done either by voice or text, and the emotion engine recognizes the user's emotions as they type.
[1548] Example: A user enters, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1549] 2. Terminal:
[1550] The device receives user input. In the case of voice input, the device converts the voice data into text data using a voice recognition engine. The converted text data and the user's emotion information are sent to the server.
[1551] AI processing and phone generation
[1552] 3. Server:
[1553] The server receives text data and emotional information from the device and passes it to a generative AI model (e.g., a model using natural language processing). The generative AI model analyzes the user's intentions from the text data and generates a dialogue script that takes the emotional information into account.
[1554] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1555] 4. Server (Speech synthesis):
[1556] The generated dialogue script is input into a speech synthesis engine and converted into voice data that resembles the user's voice. At this time, the tone and expression of the voice are adjusted based on emotional information generated by the emotion engine.
[1557] Example: The generated speech data is, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1558] 5. Server (Calling):
[1559] The voice data is transmitted through a telephone system to place a call to a destination (for example, a restaurant).
[1560] Real-time dialogue and results notification
[1561] 6. Server:
[1562] The server uses the generated voice to make a call and have a real-time conversation with the other party (e.g., a restaurant). Depending on the other party's response, the AI generates an appropriate response and moves the conversation forward. The entire conversation is recorded on the server.
[1563] 7. Server (Text Conversion and Summarization):
[1564] The server converts the completed conversation from audio data to text data and then summarizes the content. The summarized information is a concise summary of the necessary information.
[1565] 8. Notification from the server to the device:
[1566] The summarized text data is sent from the server to the device, and the results are communicated to the user via push notifications, etc. By checking these notifications, the user can easily understand the content and results of the call.
[1567] Example: A notification will appear on your device saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1568] Specific example explanation
[1569] If a user wants to make a restaurant reservation for tomorrow at 7:00 PM, the specific flow is as follows:
[1570] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." The system simultaneously recognizes the user's emotions (for example, if the user is feeling stressed).
[1571] 2. The device converts the speech into text and sends it to the server along with emotional information.
[1572] 3. The server uses the generative AI model to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1573] 4. The server uses a speech synthesis engine to generate voice data that reflects the user's voice and emotions based on the emotional information.
[1574] 5. The server makes the call and communicates with the other party in real time.
[1575] 6. The server records the conversation, converts it into text, and creates a summary.
[1576] 7. The server notifies the terminal of the summary, and the user confirms the result: "A restaurant reservation has been made for two people at 7:00 PM."
[1577] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[1578] The processing flow will be explained below.
[1579] Step 1:
[1580] The user launches a dedicated smartphone app and inputs by voice or text what they want to communicate over the phone, such as an inquiry or reservation. Once input is complete, the emotion engine recognizes the user's emotions.
[1581] Example: A user speaks, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[1582] Step 2:
[1583] The terminal receives input voice data and converts the voice data into text data using a voice recognition engine.
[1584] Example: Voice input is converted into text data such as "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1585] Step 3:
[1586] The terminal transmits the converted text data and the recognized emotion information to the server.
[1587] Step 4:
[1588] The server analyzes the text data and emotional information received from the device. It uses a generative AI model to understand the user's intentions and generate a dialogue script. This script is generated by reflecting the emotional information.
[1589] Example: A dialogue script will be generated that reads, "Hello, I would like to make a restaurant reservation for two people at 7:00 PM tomorrow. Please let me know the availability."
[1590] Step 5:
[1591] The server inputs the generated dialogue script into a speech synthesis engine to generate voice data that resembles the user's voice. The tone and expression of the voice are adjusted based on the emotions recognized by the emotion engine.
[1592] Example: Speech data: "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1593] Step 6:
[1594] The server transmits the generated voice data to a telephone system, which places a call to the destination (e.g., a restaurant).
[1595] Step 7:
[1596] The server establishes a telephone connection, transmits the generated voice data to the other party, and initiates a dialogue. The dialogue with the other party is conducted in real time, and an appropriate response is generated based on the other party's responses.
[1597] Example: If the reply is "I have availability at 7:00 PM," generate a response of "Thank you. I'd like to meet you at that time."
[1598] Step 8:
[1599] The server records the entire phone conversation.
[1600] Step 9:
[1601] The server converts the recorded audio data into text and creates a summary.
[1602] Example: The text data generated is "A restaurant reservation has been made for two people at 7:00 PM."
[1603] Step 10:
[1604] The server transmits the summarized text data to the terminal.
[1605] Step 11:
[1606] The text data received by the device is notified to the user via push notification, email, etc.
[1607] Example: A notification appears on the user's smartphone saying, "A restaurant reservation has been made for two people at 7:00 PM."
[1608] Through the above series of processing steps, the user can efficiently obtain the necessary information in a manner that is in tune with their emotions, without having to make a phone call themselves.
[1609] Example 2
[1610] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1611] Conventional telephone reservation systems have problems such as the stress users feel when making phone calls themselves and the inability to efficiently obtain the necessary information. Furthermore, dialogue systems that proceed without taking the user's emotions into consideration often fail to provide the quality of dialogue users expect. There is a need for a system that can solve these problems and allow users to make reservations and inquiries more comfortably.
[1612] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for allowing a user to input the content of an inquiry or reservation by voice or text and analyzing the emotion at the time of input; means for converting the input voice data into text data; means for receiving the text data and emotion information, analyzing the intention using a generative AI model, and generating a dialogue script that takes emotion into consideration; voice synthesis means for converting the generated dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the content of the conversation, converting the recorded voice data into text, and summarizing it; and means for notifying the user's terminal of the summarized text. This reduces the stress of the user making a phone call and enables efficient information acquisition through a dialogue that is sensitive to the user's emotions.
[1613] "User" refers to a person who uses the system.
[1614] A "dedicated app" refers to application software that runs on devices such as smartphones and allows users to enter details of inquiries and reservations.
[1615] "Emotion engine" refers to software or algorithms for analyzing emotions from a user's voice or input data.
[1616] "Speech recognition engine" refers to software or algorithms for converting input voice data into text data.
[1617] A "generative AI model" refers to an artificial intelligence model that analyzes user intent based on text data and emotional information, and generates an appropriate dialogue script.
[1618] A "dialogue script" refers to a text-based scenario for progressing a dialogue, which is generated based on the user's intentions.
[1619] "Speech synthesis engine" refers to software or algorithms for converting text data into speech data.
[1620] "Telephone System" means the communications infrastructure and software used to make telephone calls based on generated voice data.
[1621] "Text conversion means" refers to a system or technology for converting voice data into text data.
[1622] "Summarization means" refers to a system or technology for concisely summarizing the converted text data.
[1623] "Push notification" refers to a technology that allows a server to send information to a user's device in real time.
[1624] The present invention provides a system that reduces stress when a user makes a phone call and efficiently obtains necessary information. In particular, the system has the function of recognizing the user's emotions and proceeding with the conversation in a way that is considerate of those emotions. The following describes in detail the embodiments of the invention.
[1625] 1. User Input Procedure
[1626] Users launch a dedicated smartphone app and input the details of their inquiry or reservation. For voice input, users press the microphone button and start speaking. For text input, users type directly into the text field. As the user types, the app uses an emotion engine to analyze the user's emotions.
[1627] Specific hardware or software you will be using:
[1628] Dedicated app (smartphone app)
[1629] Emotion Engine
[1630] 2. Initial processing on the device
[1631] The device receives the input voice data and converts it into text data using a speech recognition engine. The converted text data and emotion information are then sent to the server. A speech recognition API such as Google Speech-to-Text is used here.
[1632] Specific hardware or software you will be using:
[1633] Speech recognition engine (Google Speech-to-Text)
[1634] 3. Data processing on the server
[1635] The server inputs the received text data and emotional information into a generative AI model to analyze the user's intentions. Generative AI models such as OpenAI GPT-4 are used here. The generated dialogue script takes the user's emotions into account.
[1636] Specific hardware or software you will be using:
[1637] Generative AI model (OpenAI GPT-4)
[1638] 4. Voice conversion of dialogue scripts
[1639] The generated dialogue script is passed to a speech synthesis engine, which converts it into voice data that reflects the user's voice and emotions. A speech synthesis engine such as Amazon Polly is used for this purpose.
[1640] Specific hardware or software you will be using:
[1641] Speech synthesis engine (Amazon Polly)
[1642] 5. Calling via the telephone system
[1643] Using voice data, calls are made through a telephone system and a real-time conversation is held with the other party. The server analyzes the other party's response in real time and uses a generative AI model to generate the next appropriate response and continue the conversation. Telephone systems such as Twilio are used.
[1644] Specific hardware or software you will be using:
[1645] Phone system (Twilio)
[1646] 6. Recording and summarizing conversation content
[1647] After the call is over, the server converts the recorded voice data into text and summarizes the key points, which are then stored in the system and later shared with the user.
[1648] Specific hardware or software you will be using:
[1649] Text conversion methods
[1650] Summary tools
[1651] 7. Notification of Results
[1652] The server sends the summarized text data to the device, and the user is notified of the results by push notification or other means. The user can check the call results and next steps through the app.
[1653] Specific hardware or software you will be using:
[1654] Push notification system
[1655] Specific examples
[1656] The specific flow when a user wants to make a restaurant reservation for tomorrow at 19:00 is shown below.
[1657] 1. A user speaks to the app, "I'd like to make a restaurant reservation for tomorrow at 7:00 PM." At this time, the system recognizes that the user is feeling stressed.
[1658] 2. The device converts the speech into text and sends it to the server along with emotional information.
[1659] 3. The server uses a generative AI model (e.g., GPT-4) to analyze the text and emotional information and generate a dialogue script that reads, "Hello, I'd like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know availability."
[1660] 4. The server uses a speech synthesis engine (e.g., Amazon Polly) to generate voice data and adjust it to reflect the emotion.
[1661] 5. The server uses a telephone system (e.g., Twilio) to make a call and communicate with the restaurant in real time.
[1662] 6. The server converts the call into text and creates a summary.
[1663] 7. The server sends the summary to the user's terminal and notifies them of the result: "A restaurant reservation has been made for two people at 7:00 PM."
[1664] Example prompt sentence:
[1665] "I'd like to make a restaurant reservation for tomorrow at 7:00 PM."
[1666] In this way, the system of the present invention allows users to efficiently obtain information in an emotionally relevant manner without having to make a phone call themselves.
[1667] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1668] Step 1:
[1669] Users launch a dedicated smartphone app and input the details of their inquiry or reservation by voice or text. Once input is complete, the app uses an emotion engine to analyze the user's emotions. The input data is output as text data (in the case of text input) or voice data (in the case of voice input). Emotional data is also output at the same time.
[1670] Specific behavior:
[1671] The user speaks, "I would like to make a restaurant reservation for tomorrow at 7:00 PM."
[1672] The app receives the voice data and analyzes the emotions (e.g., tension or stress) present in the input.
[1673] Step 2:
[1674] The device receives the user's voice data and converts it into text data using a speech recognition engine. The text data and analyzed emotion information are sent to the server. The input is voice data, and the output is text data and emotion information.
[1675] Specific behavior:
[1676] The terminal converts the voice data into text data using a voice recognition engine (e.g., Google Speech-to-Text).
[1677] The converted text data and emotion data are sent to the server.
[1678] Step 3:
[1679] The server receives the text data and emotional information sent from the device. The text data and emotional information are input into the generative AI model, which analyzes the user's intentions. The generated dialogue script is output. The input is text data and emotional information, and the output is a dialogue script.
[1680] Specific behavior:
[1681] The server inputs the data into a generative AI model such as GPT-4, which analyzes the user's intent.
[1682] The generative AI model generates a dialogue script that reads, "Hello, I would like to make a restaurant reservation for two people tomorrow at 7:00 PM. Please let me know the availability."
[1683] Step 4:
[1684] The server passes the generated dialogue script to a speech synthesis engine and converts it into voice data. The tone and expression of the voice are also adjusted based on the emotional information. The input is the dialogue script and emotional information, and the output is voice data.
[1685] Specific behavior:
[1686] The server inputs the dialogue script into a speech synthesis engine such as Amazon Polly.
[1687] A speech synthesis engine converts the dialogue script into voice data that reflects the user's voice and emotions.
[1688] Step 5:
[1689] The server uses the generated voice data to make a call through the telephone system, initiating a real-time dialogue with the other party, and the server analyzes the other party's response in real time and generates the next appropriate response using a generative AI model. The input is the voice data, and the output is a real-time response.
[1690] Specific behavior:
[1691] The server calls the restaurant using a phone system such as Twilio.
[1692] The server analyzes the other party's response and uses a generative AI model to generate an appropriate response and continue the dialogue.
[1693] Step 6:
[1694] After the call ends, the server converts the recorded voice data into text and summarizes the key points. The converted text data and the summary are output. The input is the voice data, and the output is the text data and the summary.
[1695] Specific behavior:
[1696] The server uses a speech recognition engine to convert the call into text.
[1697] The server uses a summarization algorithm to summarize the converted text.
[1698] Step 7:
[1699] The server sends the summarized text data to the terminal, and the user is notified of the results via push notification, etc. The input is the summarized text data, and the output is a notification to the user.
[1700] Specific behavior:
[1701] The server sends a summary to the terminal stating, "A restaurant reservation has been made for two people at 19:00."
[1702] The user receives a push notification and checks for details.
[1703] (Application example 2)
[1704] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1705] In conventional food delivery services, when users place orders over the phone, they are prone to making mistakes in how they communicate their orders. Furthermore, there are few ways to reduce the stress and tension users feel when ordering over the phone. Furthermore, there are no systems that can process orders in a way that takes into account the user's emotions. This creates a need for a user-friendly and efficient food delivery ordering system.
[1706] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for inputting content that a user wants to communicate over the phone; means for converting the input content into text data; generative model means for receiving the text data, analyzing the intention, and generating a dialogue script; speech synthesis means for converting the dialogue script into voice data; means for making a phone call using the voice data and having a real-time conversation with the other party; means for recording the dialogue content and converting the recorded voice data into text to summarize it; means for notifying the user's terminal of the summarized text; means for recognizing the user's emotion and optimizing the dialogue script and voice data based on the emotion; input means for allowing the user to select either voice or character input; and means for acquiring emotion information and adjusting the tone and expression of the voice based on the emotion information. This enables the user to efficiently order food delivery in a manner that is sensitive to their emotion without having to make a phone call.
[1707] "User" refers to an individual or corporation that uses this system to make phone calls, make reservations, place orders, and perform other operations.
[1708] "Content to be communicated over the phone" refers to information that the user needs to communicate to the other party over the phone, such as an order, an inquiry, or a reservation.
[1709] "Input means" refers to an interface or device used by a user to input content by voice or text.
[1710] "Text data" refers to text information converted from voice input or text input.
[1711] "Generative model means" refers to the AI model or algorithm used to analyze user input and generate an appropriate dialogue script.
[1712] "Speech synthesis means" refers to a technology or device that converts the generated dialogue script into voice data.
[1713] "Voice data" refers to digital information of voice generated by a voice synthesis means.
[1714] "Means for making a call" refers to a system or device for making a call using the generated voice data.
[1715] "Means of real-time interaction" refers to technology or systems that allow you to have a direct conversation with the person you are calling.
[1716] "Means for recording conversation content" refers to technology or devices that save telephone conversations as audio data.
[1717] "Means of converting to text and summarizing" refers to the technology or algorithms that convert recorded audio data into text and summarize that text concisely.
[1718] "Means for notifying the device" refers to technology or systems that notify the user of the summarized text on their device, such as a smartphone or tablet.
[1719] "Means of recognizing emotions" refers to technologies and models for analyzing and extracting emotional information from user input and dialogue content.
[1720] "Means for optimizing dialogue scripts and voice data based on emotions" refers to technologies and algorithms that use emotional information to adjust the tone and expression of the dialogue scripts and voice data that are generated.
[1721] An input means that allows the user to select between "voice and text input" refers to an interface or device that allows the user to freely select either voice input or text input.
[1722] This invention provides a specific implementation method for recognizing user emotions and efficiently completing orders in a food delivery ordering system.
[1723] System Configuration
[1724] The present invention is mainly composed of a user terminal, a server, and an AI module.
[1725] User terminal
[1726] To order food delivery, a user first uses a user device such as a smartphone or tablet. The user then launches a dedicated app and can enter the order details by voice or text. The user's emotions are recognized as they are entered.
[1727] server
[1728] The server receives input data (voice or text) and emotional information sent from the user's device. In the case of voice input, the server converts the voice data into text data using a voice recognition engine. Based on this converted text data and emotional information, a generative AI model generates a dialogue script. The generated dialogue script is then input into a voice synthesis engine and converted into voice data. At this time, the tone and expression of the voice are optimized based on the emotional information.
[1729] AI Module
[1730] The generative AI model and speech synthesis engine are included in the AI module. The generative AI model analyzes the user's intentions and generates a dialogue script that takes emotional information into account. The speech synthesis engine converts this dialogue script into voice data.
[1731] communication means
[1732] The server uses the generated voice data to make calls to the corresponding restaurant or delivery provider and engages in real-time conversations, which are recorded in the server and converted into text data as needed.
[1733] User Notification
[1734] Once the telephone conversation is complete, the server summarizes the conversation and sends the summarized information to the user's device via a push notification or other method, allowing the user to confirm whether the order was placed successfully.
[1735] Specific examples
[1736] For example, if a user says "I would like to order a pizza from Restaurant X", the system will:
[1737] 1. The user device receives voice input, and the emotion recognition engine analyzes the user's emotions.
[1738] 2. The server converts the voice data into text, and the generative AI model generates a dialogue script. The generated dialogue script is, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita pizza, please."
[1739] 3. Based on this dialogue script, the speech synthesis engine optimizes the tone and expression of the voice according to the emotional information and generates the voice data.
[1740] 4. The server will make a call and automatically communicate the order details.
[1741] 5. The restaurant's response is recorded, and the content of the conversation is converted into text data and summarized.
[1742] 6. The summary is sent to the user's terminal, and the user is informed that "your order has been accepted. Delivery time is 18:00."
[1743] Prompt Sentence Examples
[1744] An example of a prompt to pass to a generative AI model would be:
[1745] "Hello, I'd like to order a pizza from XX Restaurant. I'd like one Margherita, please."
[1746] Emotion: Stress
[1747] This allows the system to efficiently order food delivery while being sensitive to the user's emotions.
[1748] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1749] Step 1:
[1750] The user launches a dedicated smartphone app and inputs their order details by voice or text. The emotion recognition engine analyzes the user's emotions as they input their order. For example, if a user inputs "I'd like to order pizza from Restaurant X," the device captures the voice data and analyzes and obtains emotional information.
[1751] Step 2:
[1752] The voice data acquired by the device is converted into text data. A voice recognition engine (e.g., Google Speech-to-Text API) is used to convert the voice data into text data. For example, a voice saying "I would like to order pizza from XX Restaurant" is converted into text data saying "I would like to order pizza from XX Restaurant."
[1753] Step 3:
[1754] Text data and emotional information are sent from the terminal to the server. The order details (text data) and emotional information (e.g., stress) entered by the user are sent to the server and used as input for the next process.
[1755] Step 4:
[1756] Based on the text data and emotional information received by the server, a generative AI model (e.g., GPT-3) is used to generate a dialogue script. The generative AI model analyzes the given text and emotional information and generates an appropriate dialogue script. For example, it generates a script that reads, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please."
[1757] Step 5:
[1758] The server inputs the generated dialogue script into a speech synthesis engine (e.g., Google Text-to-Speech API) to generate voice data. At this time, the tone and expression of the voice are adjusted based on the emotional information. For example, "Hello, I'd like to order a pizza from Restaurant X. I'd like one Margherita, please" is generated in a soft voice that matches the emotion.
[1759] Step 6:
[1760] The server uses the generated voice data to place a call to the specified restaurant and execute the order. This telephone conversation takes place in real time, and the server accurately conveys the order details.
[1761] Step 7:
[1762] The server records the audio data during the conversation, including the restaurant's responses and questions.
[1763] Step 8:
[1764] The server converts the recorded voice data into text and summarizes it. The speech recognition engine is used again to convert the voice recording into text data, and the content is concisely summarized using a summarization algorithm. For example, a summary text such as "Your order has been received. Delivery time is 6:00 PM" is generated.
[1765] Step 9:
[1766] The server sends the summarized text to the user's device and sends a push notification with the order details, which displays a notification such as "Your order has been accepted. Delivery time is 6:00 PM."
[1767] This allows users to efficiently place food delivery orders in a way that is in tune with their emotions, without having to make a phone call themselves.
[1768] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1769] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1770] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1771] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1772] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1773] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1774] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1775] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1776] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1777] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1778] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1779] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1780] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1781] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1782] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1783] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1784] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1785] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1786] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1787] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1788] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1789] The following is further disclosed regarding the above embodiment.
[1790] (Claim 1)
[1791] a means for the user to input what they want to say over the phone;
[1792] A means for converting input content into text data;
[1793] a generative modeling means for receiving text data, analyzing the intention of the data, and generating a dialogue script;
[1794] a voice synthesis means for converting the dialogue script into voice data;
[1795] A means for making calls using voice data and communicating with other parties in real time;
[1796] a means for recording the content of the dialogue and converting the recorded voice data into text and summarizing it;
[1797] The system includes means for notifying the user of the summarized text at the user's terminal.
[1798] (Claim 2)
[1799] 10. The system of claim 1, wherein said system converts voice input into text.
[1800] (Claim 3)
[1801] 2. The system according to claim 1, wherein voice data for a conversation with a partner is generated based on the generated conversation script.
[1802] "Example 1"
[1803] (Claim 1)
[1804] a means for the user to input what they want to say over the phone;
[1805] A means for converting input content into text data;
[1806] a generative modeling means for receiving text data, analyzing the intention of the data, and generating a dialogue script;
[1807] a voice synthesis means for converting the dialogue script into voice data;
[1808] a means for using voice data to make calls and interact with external systems in real time;
[1809] a means for recording the content of the dialogue and converting the recorded voice data into text and summarizing it;
[1810] The system includes means for notifying a user's information processing device of the summarized text.
[1811] (Claim 2)
[1812] 10. The system of claim 1, wherein said system converts voice input into text.
[1813] (Claim 3)
[1814] 2. The system according to claim 1, wherein voice data for interacting with an external system is generated based on the generated dialogue script.
[1815] "Application Example 1"
[1816] New Claims
[1817] (Claim 1)
[1818] a means for the user to input what they want to say over the phone;
[1819] A means for converting input content into text data;
[1820] a generating AI model means for receiving text data, analyzing the intent, and generating a dialogue script;
[1821] a voice synthesis means for converting the dialogue script into voice data;
[1822] A means for making calls using voice data and communicating with other parties in real time;
[1823] a means for recording the content of the dialogue and converting the recorded voice data into text and summarizing it;
[1824] means for notifying a user of the summarized text at a terminal of the user;
[1825] A means for users to make reservations or inquiries for physical stores,
[1826] means for receiving input about the physical store, generating a dialogue script, and making a call to gather the required information;
[1827] A system that includes a means for notifying users of the results of reservations and inquiries at physical stores.
[1828] (Claim 2)
[1829] 10. The system of claim 1, wherein said system converts voice input into text.
[1830] (Claim 3)
[1831] The system according to claim 1, wherein the system receives input relating to a physical store and generates voice data for a conversation with a counterpart based on the generated conversation script.
[1832] "Example 2: Combining Emotion Engines"
[1833] (Claim 1)
[1834] A means for users to input the details of inquiries or reservations by voice or text and analyze emotions when inputting;
[1835] means for converting input voice data into text data;
[1836] a means for receiving the text data and the emotion information, analyzing the intent using a generative AI model, and generating a dialogue script that takes the emotion into account;
[1837] a voice synthesis means for converting the generated dialogue script into voice data;
[1838] A means for making calls using voice data and communicating with other parties in real time;
[1839] a means for recording the content of the dialogue, converting the recorded voice data into text, and summarizing the text;
[1840] means for notifying a user of the summarized text at a terminal of the user;
[1841] A system including:
[1842] (Claim 2)
[1843] 2. The system according to claim 1, wherein the system converts voice input content into text and adds emotional information.
[1844] (Claim 3)
[1845] 2. The system according to claim 1, wherein voice data reflecting the user's voice and emotions is generated based on the generated dialogue script.
[1846] "Application example 2 when combining emotion engines"
[1847] (Claim 1)
[1848] a means for the user to input what they want to say over the phone;
[1849] A means for converting input content into text data;
[1850] a generative modeling means for receiving text data, analyzing the intention of the data, and generating a dialogue script;
[1851] a voice synthesis means for converting the dialogue script into voice data;
[1852] A means for making calls using voice data and communicating with other parties in real time;
[1853] a means for recording the content of the dialogue and converting the recorded voice data into text and summarizing it;
[1854] means for notifying a user of the summarized text at a terminal of the user;
[1855] means for recognizing a user's emotion and optimizing the dialogue script and voice data based on the emotion;
[1856] an input means that allows a user to select either voice input or character input;
[1857] A system that includes a means for obtaining emotional information and adjusting the tone and expression of speech based thereon.
[1858] (Claim 2)
[1859] 10. The system of claim 1, which converts voice input into text and recognizes emotional information.
[1860] (Claim 3)
[1861] 2. The system according to claim 1, wherein voice data for dialogue with a partner is generated based on the generated dialogue script, and the tone and expression of the voice are adjusted based on emotional information. [Explanation of symbols]
[1862] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for the user to input what they want to say over the phone; A means for converting input content into text data; a generative modeling means for receiving text data, analyzing the intention of the data, and generating a dialogue script; a voice synthesis means for converting the dialogue script into voice data; A means for making calls using voice data and communicating with other parties in real time; a means for recording the content of the dialogue and converting the recorded voice data into text and summarizing it; The system includes means for notifying the user of the summarized text at the user's terminal.
2. 10. The system of claim 1, wherein said system converts spoken input into text.
3. 2. The system according to claim 1, wherein voice data for a conversation with a partner is generated based on the generated conversation script.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A