system
The system automates communication by processing user data to initiate voice calls and summarize results, addressing the inefficiencies and security issues of conventional telephone communication, providing a seamless and secure experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional telephone communication poses psychological and physical barriers, is inconvenient, and lacks security against fraudulent calls, making it difficult for users to make calls in emergencies or when they feel bothered.
A system that automates communication by receiving user data, performing natural language processing, generating conversation content, initiating voice calls, and summarizing results, allowing users to manage communications without direct calls.
Enables efficient, secure, and stress-free communication by automating various interactions based on pre-set conditions, reducing the need for direct phone calls and improving user experience.
Smart Images

Figure 2026070984000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional telephone communication can be a psychological and physical barrier for users, and there may be cases where it is difficult to make a phone call or there is no time to make a phone call in an emergency. Also, there is a need for a means to reduce the hassle for users who feel that direct calls are bothersome. Furthermore, improving security to prevent fraudulent and sales-related illegal calls is also an issue.
Means for Solving the Problems
[0005] The present invention provides a system comprising means for receiving and analyzing data from a user, means for performing natural language processing based on the analyzed data, means for generating conversation content using the processed data, means for initiating a voice call via an external communication network, and means for summarizing the results of the call and notifying the user. This system allows the user to automate and securely perform various necessary communications based on pre-set conditions without having to make a phone call directly.
[0006] A "user" is the entity that operates the system and creates requests such as phone calls and inquiries.
[0007] "Data" refers to a collection of information that a user inputs into a system, representing the content of a request.
[0008] "Means of analysis" refer to the technologies and methods used to understand data received from users and extract the content of requests.
[0009] "Natural language processing" refers to the technology that enables computers to understand and process human language, and is used to provide appropriate responses to user requests.
[0010] "Generation means" refers to the technology and methods for creating the conversation content necessary for a call based on the analyzed data.
[0011] An "external communication network" is the infrastructure that enables communication with other devices and systems through external connection means such as telephone lines.
[0012] A "voice call" is a method of communication that transmits the generated conversation content to the other party as actual voice.
[0013] "Call results" refer to the information and responses obtained during a call, which are compiled for feedback to the user.
[0014] "Notification means" refers to the technologies and methods used to inform the user of the results of a call. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention is a system that automatically handles various types of communication without requiring the user to make a phone call directly. This system is realized through the cooperation of user terminals, servers, and external communication networks.
[0037] Users operate a dedicated application on their device to create phone call or inquiry requests. Users input various request details, such as restaurant reservations, inquiries to companies, and even emergency calls. The device receives this information and sends it to the server.
[0038] The server analyzes the received user data and uses a natural language processing engine to analyze its content. Based on the analyzed information, a generative AI model generates conversational text to be used in the call. This generated conversational content is designed to provide optimal responses and dialogue.
[0039] Next, the server uses an external communication network to initiate a voice call to a designated contact. For example, it might call a restaurant the user wishes to book and proceed with the booking process based on the specified information. Once the call is complete, the server organizes the information gathered during the call and summarizes the results in a format requested by the user.
[0040] The user terminal ultimately receives these results from the server and notifies the user. In other words, the user only needs to confirm the results based on their pre-configured preferences without having to engage in direct conversation.
[0041] For example, if a user requests a reservation for three people at Restaurant X at 7 PM tomorrow, the server automatically calls Restaurant X based on this information and handles the request. The server then notifies the user that the reservation was successful, allowing the user to directly verify the result. This reduces the stress and inconvenience associated with regular phone calls, enabling efficient and secure communication via telephone.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] Users launch a dedicated application on their device and enter details such as phone reservations or inquiries. This information includes details such as the store name, reservation date and time, and number of people.
[0045] Step 2:
[0046] The terminal formats the data entered by the user and sends it to the server. The data is structured using a standard format suitable for communication, such as JSON.
[0047] Step 3:
[0048] The server receives data sent from the terminal and parses its contents. This parsing identifies what the request is asking for (for example, a restaurant reservation, a product inquiry, etc.).
[0049] Step 4:
[0050] The server passes the analyzed data to a natural language processing engine, which generates natural conversational sentences based on the user's requests. During this process, the generation AI model prepares appropriate responses based on the context.
[0051] Step 5:
[0052] Based on the generated conversation text, the server initiates a voice call to a specified telephone number via an external communication network. A communication module is then used to perform the actual dialing and establish the connection.
[0053] Step 6:
[0054] Once a call is connected, the server automatically initiates a conversation through AI. During the conversation, it adjusts the content of the discussion in real time based on the other party's responses and gathers the necessary information.
[0055] Step 7:
[0056] Once the call is complete, the server organizes the call content and clarifies the results as requested by the user. These results are presented in a format that is easy for the user to understand.
[0057] Step 8:
[0058] The server sends the organized call results to the terminal. The result data is then reorganized according to the format and sent to the user's terminal.
[0059] Step 9:
[0060] The terminal receives the results from the server and notifies the user. The user can then review the results displayed on the terminal screen and take further action or make decisions.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] Traditionally, making inquiries or reservations by phone has been time-consuming and cumbersome, often accompanied by anxiety and stress during the conversation. Furthermore, automated systems have suffered from a lack of immediacy and accuracy. The challenge lies in improving this situation and enabling users to engage in more comfortable and diverse forms of communication.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes means for receiving and analyzing information from a user device, means for performing language processing based on the analyzed information, and means for generating dialogue content using the processed information. This enables users to communicate quickly and accurately, such as making inquiries and reservations, automatically, without having to make phone calls themselves.
[0066] A "user device" is a terminal operated by the user, and is a device used for inputting and transmitting data.
[0067] "Means for receiving and analyzing information" refers to a mechanism that has the ability to acquire data transmitted from a user device and understand the content of that data.
[0068] "Means of language processing" are technologies that use analyzed information to determine the meaning and context in natural language and construct appropriate dialogue.
[0069] A "generation method for generating dialogue content" is a system that automatically creates conversations in response to user requests and provides them in a format that can be used for voice calls.
[0070] "External communication media" refers to communication methods used for making voice calls, including the internet and public telephone networks.
[0071] "Means for initiating voice communication" refers to a method that automatically starts a call with a designated party based on the generated dialogue content.
[0072] "Means of organizing communication results" refers to the process of organizing information obtained after a call and summarizing it in a way that is easily understandable to the user.
[0073] "Means of notifying the user" refers to a method of sending organized information to the user's device and informing the user of the results.
[0074] This invention provides an automated voice call system that combines a user, a terminal, and a server. This system begins with the user using a dedicated application on the terminal to input various requests. For example, the user might input, "I would like to make a reservation for four people at Restaurant Y for 7 PM tomorrow."
[0075] The terminal, upon receiving user input, transmits that information to the server. The terminal is a typical mobile device or computer, communicating with the server via the internet. This system utilizes HTTPS as its communication protocol to ensure secure data transmission.
[0076] The server analyzes the received data. This analysis uses a natural language processing engine such as Google Cloud Natural Language API. Based on the analyzed data, the server generates appropriate conversation content using a generative AI model. This generative AI model is designed to accurately and effectively utilize the user's intended content in the call.
[0077] Based on the generated conversation content, the server initiates voice communication via an external communication medium. Specifically, the server automatically dials the designated recipient using the Public Public Telephone Network (PSTN) or VoIP technology. Once the call is answered, the conversation proceeds using the automatically generated conversation content.
[0078] After the call ends, the server organizes the information gathered during the call. This organization includes transcribing and summarizing the conversation. The organized results are sent to the user's device, and the device notifies the user of the results. The notification is delivered via a message within a dedicated app or through push notifications on the device.
[0079] This entire process eliminates the need for users to make phone calls themselves, allowing them to enjoy automated and efficient communication. For example, a user could use a prompt like, "I'd like to make a reservation for 5 people at Izakaya Z at 8 PM on the weekend," and confirm the reservation using the same procedure. This allows users to save time and effort while achieving accurate communication.
[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0081] Step 1:
[0082] The user opens a dedicated application on their device and enters a specific request. For example, they might enter something like, "I would like to reserve a table for two at Restaurant W tomorrow night at 8 PM." This entered data is temporarily saved on the device as text.
[0083] Step 2:
[0084] The terminal sends the user's input request to the server. The terminal uses HTTP requests to send text data as packets to the server. In this process, the user request, as input, arrives at the server as text data.
[0085] Step 3:
[0086] The server parses the text data of the received request. Here, a natural language processing engine is used to extract the meaning of the text data and convert it into a structured format. The input to this process is a user request in text format, and the output is structured data.
[0087] Step 4:
[0088] The server uses a generative AI model based on structured data to generate appropriate conversation content. This generation process uses prompt sentences as input, and the generated conversation is stored on the server in text format. The output is a customized conversation based on the user's intent.
[0089] Step 5:
[0090] The server uses the generated conversation content to initiate voice communication with a designated recipient via an external communication medium. The server uses a telephone API to convert the generated conversation content into speech and transmit it to the recipient. The input for this step is the generated conversation content, and the output is a real-time voice call.
[0091] Step 6:
[0092] After the call is completed, the server organizes the results. It documents the information obtained during the call as text and extracts the key points. The input to this process is the audio data of the call result, and the output is the textual result.
[0093] Step 7:
[0094] The server sends the organized call results to the terminal, and the terminal notifies the user of the results. On the terminal, the results are displayed in the notification bar or app interface, allowing the user to easily check the content. The input for this step is the organized call results, and the output is the result notification presented to the user.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] Currently, many users experience the inconvenience of having to directly contact service providers when making electronic payments. Furthermore, manual payment procedures are prone to errors and omissions, resulting in wasted time and effort. There is a need to solve these problems and establish a method that allows users to complete payment procedures efficiently and without stress.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, means for generating dialogue content using the processed information, means for initiating voice communication via an external information transmission network, means for organizing the results of the communication and notifying the user of the organized results, and means for automatically contacting the payee specified by the user and completing the payment. This enables the user to complete the payment procedure quickly and accurately through an automated process.
[0100] "Means of receiving and analyzing information from users" refers to a method in which a server receives data sent by a user and performs analysis to understand it.
[0101] "Means of natural language processing" refers to technologies used by computers to understand human language and to provide appropriate responses and processing.
[0102] "A means for generating dialogue content" refers to a method of creating a conversation used for communication with the recipient, based on analyzed data.
[0103] "Means of initiating voice communication via an external information transmission network" refers to a method of initiating a voice call and having a conversation with another person through an external communication system.
[0104] "A means of organizing the results of communication and notifying the user of those organized results" refers to a method of summarizing the results obtained during the communication process and conveying them to the user in an easy-to-understand manner.
[0105] "A means of automatically contacting the payment recipient specified by the user and completing the payment" refers to a method of automatically contacting the payment recipient specified by the user and completing the payment procedure.
[0106] To implement this invention, the server receives payment information transmitted from the user's terminal. The received information is first processed as data by a dedicated analysis engine, and its contents are analyzed in detail. A natural language processing engine (e.g., Google Cloud Natural Language API) is used for the analysis. In this process, the server accurately understands the user's intent and determines the specific steps to be taken next.
[0107] Specifically, based on the results of analysis using natural language processing, a generative AI model (e.g., OpenAI® GPT series) generates the dialogue content necessary for the payment process. The generated dialogue content is then used to automatically inquire with the company or service provider specified by the user.
[0108] The communication method involves voice communication initiated via an external information transmission network. Specifically, this system, configured through the user's terminal, engages in real-time conversation based on payment information, obtains necessary information, and completes the payment. After communication is complete, the server organizes the results and notifies the user's terminal of the updated payment status.
[0109] For example, if a user wants to pay their monthly electricity bill to the power company, the user enters "Pay this month's bill to the power company" into the application. The system then automatically contacts the power company and initiates the payment process. An example of a prompt message would be, "Please obtain the procedure to complete this month's payment to the power company." In this way, the present invention provides a system that allows users to complete payments quickly and accurately without having to go through cumbersome procedures.
[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0111] Step 1:
[0112] The user launches the application on their device and enters payment details and amounts. The entered information is then sent to the server as electronic data.
[0113] Step 2:
[0114] The server analyzes the received data. Using a natural language processing engine, it analyzes the user's input from a linguistic perspective and extracts the elements necessary for the payment process. The input is raw data from the user, and the output is a structured analysis result.
[0115] Step 3:
[0116] Based on the analysis results, the server uses a generation AI model to generate payment-related inquiry conversations. It takes a prompt sentence (for example, "Please get instructions on how to complete this month's payment to the utility company") as input and outputs the generated sentence.
[0117] Step 4:
[0118] The server initiates voice communication with the designated payee via an external information transmission network. Using the generated query, it conducts a voice call and retrieves information in real time.
[0119] Step 5:
[0120] After the call ends, the server organizes the acquired information. Specifically, it summarizes information such as payment status and processing status, and constructs data to notify the user. The input is the raw data acquired from the call, and the output is the organized information for notification to the user.
[0121] Step 6:
[0122] The device receives notifications from the server and displays the payment status and results to the user. By checking this information, the user can know that the payment has been completed.
[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0124] This invention relates to a system that identifies a user's emotional state and optimizes the process of making phone calls based on that state. The system is implemented with a configuration including a user terminal, a server, and an emotion engine, and automatically provides appropriate communication.
[0125] Users input the necessary information through a dedicated application on their device. For example, if they wish to make a reservation or inquiry, they input the target location and detailed requests. The device collects this data and uses an emotion engine to determine the user's emotional state in the process. The emotion engine analyzes the user's emotions from the input text and voice, accurately capturing their state.
[0126] The server receives data and emotional state transmitted from the terminal. Based on the analyzed emotional information, the natural language processing engine generates appropriate conversation content. In this process, the tone and expressions are adapted to the user's emotional state. For example, if the user is nervous, a calm tone of voice will be generated.
[0127] Next, the server initiates a voice call to the designated recipient via an external communication network. This call process is dynamically managed by prioritizing based on the user's emotions. If a response is received from the recipient during the conversation, the server updates the generated conversation content in real time and records the response.
[0128] After the call ends, the server compiles the information received and notifies the user. This includes confirmation of the reservation and answers to inquiries. The user can then check the results on their device and take further action as needed.
[0129] For example, if a user is feeling stressed, the system will prioritize the call as an urgent matter and complete it quickly. Conversely, if the user is relaxed, the system will provide a calm and courteous response, ensuring a comfortable calling experience.
[0130] Thus, the present invention provides a means to realize communication that takes user emotions into consideration and to resolve the problems of conventional telephone communication.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] Users launch a dedicated application on their device and enter their phone reservation or inquiry details. They can also add custom comments or voice messages to express their emotions.
[0134] Step 2:
[0135] The device collects data entered by the user and analyzes the user's emotional state using an emotion engine. It analyzes text and voice data to identify emotional patterns.
[0136] Step 3:
[0137] The device sends the analysis results to the server. This includes data on the user's request, desired action, and emotional state. The data is structured in a standard format.
[0138] Step 4:
[0139] The server analyzes the data received from the terminal and uses a natural language processing engine to generate appropriate responses to requests. Based on sentiment data, it adjusts the tone and expression of the conversation.
[0140] Step 5:
[0141] The server considers the generated conversation content and emotional state, and initiates a voice call to the designated recipient via an external communication network. It is also possible to set call priorities based on emotional state.
[0142] Step 6:
[0143] The server updates the conversation content in real time during the call using a dialogue generation AI. Based on the other party's responses, it retrieves appropriate information as needed and continues the conversation.
[0144] Step 7:
[0145] After the call is completed, the server organizes the information obtained during the call and compiles the data in a format that aligns with the user's request.
[0146] Step 8:
[0147] The server sends the processed results to the user's terminal. This allows the user to check the call results and responses.
[0148] Step 9:
[0149] The device notifies the user of the received results and displays the details on the screen. The user can then decide on their next action based on the results.
[0150] (Example 2)
[0151] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0152] Traditional communication technologies often suffer from a decline in communication quality because they do not take into account the user's emotional state when responding or prioritizing calls. In particular, when users are experiencing stress or anxiety, appropriate responses and prompt communication are often lacking. This highlights the need for improved user experience.
[0153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0154] In this invention, the server includes means for receiving and collecting user input information, means for analyzing the user's emotional state from the collected information, and means for initiating voice communication via an external communication network using the generated conversation content. This makes it possible to set appropriate responses and priorities according to the user's emotional state.
[0155] "User input information" refers to data such as reservation and inquiry details, as well as any additional conditions, that users submit to the system.
[0156] "Emotional state" refers to the user's psychological state and reactions, and is identified through the analysis of text and voice.
[0157] "Natural language processing" is a technology that enables computers to understand and generate human language, and it is a method for creating conversational content based on analyzed data.
[0158] "Conversation content" refers to the voice or text-based responses generated by the system, which are adjusted to suit the user's emotional state.
[0159] "External communication networks" refer to communication infrastructure such as the internet and telephone lines that servers use to communicate.
[0160] "Call priority" is an indicator that shows the importance of a call based on the user's emotional state, and is dynamically set according to urgency.
[0161] "Response content" refers to real-time response information from the other party, and is data acquired during the course of the conversation.
[0162] The following describes embodiments for carrying out the present invention.
[0163] The system of the present invention mainly consists of a user terminal, a communication server, and an emotion analysis engine. The user uses a dedicated application on the terminal to input necessary information regarding reservations and inquiries. This application includes hardware (e.g., microphone, keyboard) to acquire information entered in the form of text or voice, and software (e.g., speech recognition software, text analysis engine) to process the data.
[0164] The device uses an emotion analysis engine to determine the user's emotional state from the input information. The emotion analysis engine combines natural language processing and speech analysis technologies, and can capture emotions by analyzing, for example, the tone and speed of voice, and the emotional vocabulary in text.
[0165] The analyzed emotion information and user input data are sent to the server. The server uses a natural language processing engine to generate conversation content adapted to the user's emotions. This engine can create multiple response patterns using a generative AI model and select the optimal one. The generated conversation content is transmitted to the designated recipient via an external communication network, and a voice call is initiated.
[0166] During a call, the server updates the conversation in real time based on the responses received. Furthermore, call priorities are dynamically managed based on the user's emotional state, ensuring that urgent calls are processed without delay.
[0167] For example, if a user enters the text "I want to make a dentist appointment for next Tuesday," the device's emotion analysis engine detects that the user is nervous. Based on this information, the server prepares a response in a relaxed tone and prioritizes initiating the call with the dentist.
[0168] An example of a prompt message is, "Generate a call message to confirm a reservation while the user is feeling stressed." This prompt allows the system to generate the most appropriate conversation for the situation. This prompt serves as input for the generation AI model to function correctly.
[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0170] Step 1:
[0171] Users enter reservation and inquiry information using a dedicated application on their terminal. This input data includes information in text or audio format. The terminal converts this information into digital signals and applies them as input to an emotion analysis engine. This engine analyzes the user's emotions from the tone of voice and the context of the text, and outputs their emotional state.
[0172] Step 2:
[0173] The terminal sends the emotional state obtained by the emotion analysis engine and the information entered by the user to the server. The server activates a natural language processing engine based on the received data. This engine processes the conversation data using a generative AI model to generate optimal conversation content that corresponds to the user's emotions. The generated conversation content is output and used for the next processing step.
[0174] Step 3:
[0175] The server uses the generated conversation content to initiate a voice call to the designated recipient via an external communication network. During this process, priority is initially set based on sentiment information, with urgent calls being processed first. The server establishes an external connection and prepares to begin the conversation.
[0176] Step 4:
[0177] During a call, the server updates the conversation content in real time based on information received from the other party. The input is response data from the other party, which the server analyzes and outputs newly generated response content using a generative AI model. The response is applied immediately, maintaining the quality of the call.
[0178] Step 5:
[0179] After the call ends, the server processes the call results and notifies the user. The acquired information is stored in a database and displayed to help the user easily decide on the next steps. As a final output, the user is provided with a reservation confirmation and answers to their inquiries.
[0180] (Application Example 2)
[0181] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0182] In modern information and communication technology, users can experience anxiety and stress in various situations. These emotional states can lead to inefficiencies and problems in communication. However, existing systems lack sufficient means to accurately monitor users' emotional states and provide support at the appropriate time. Therefore, there is a need to support users' mental well-being and facilitate smooth communication.
[0183] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0184] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, and means for monitoring the user's emotional state and providing prompt support when an anomaly is detected. This makes it possible to monitor the user's emotional state in real time, provide appropriate language responses as needed, and provide rapid support in emergencies.
[0185] "Means of receiving and analyzing information from users" refers to functions that acquire data provided by users, analyze that data, and understand the user's state and requests.
[0186] "Means of performing natural language processing" refers to a function that uses acquired user information to enable a computer to understand the context and derive an appropriate response.
[0187] "Generating means for generating response content" refers to a function that creates appropriate responses or messages for the user based on the results of analysis using natural language processing.
[0188] "Means of initiating a voice call via an external communication network" refers to a function that uses a communication network to initiate a voice-based conversation with a designated party.
[0189] "A means of organizing call results and notifying the user of those organized results" refers to a function that summarizes the content and results of a voice call and reports or provides feedback to the user.
[0190] "A means of monitoring the user's emotional state and providing rapid support when an anomaly is detected" refers to a function that monitors the user's emotions in real time and provides immediate and appropriate support when an anomaly such as stress or panic is observed.
[0191] The system for realizing this invention is designed to monitor the user's emotional state in real time and provide appropriate responses. Its main components include a terminal used by the user, a server containing an emotion engine for analyzing the user's emotions, and a communication network for voice communication.
[0192] The user accesses the system through a terminal and enters the necessary information. This terminal is an internet-connected device such as a smartphone or wearable device. User information is collected by the terminal and sent to the server via the emotion engine. The emotion engine analyzes the user's emotional state from voice and text messages, using, for example, Python or an emotion analysis library (such as Hugging Face's Transformers). Based on this analysis, a natural language processing engine generates a response appropriate to the user.
[0193] Based on the generated response, the server initiates a voice call with the specified party using an external communication network. During this process, the server can utilize a communication API (e.g., Twilio). Once the call ends, the server can organize the call content and notify the user of the results.
[0194] As a concrete example, when a user is feeling stressed, the device reports its emotional state to the server, which then uses a generative AI model to generate a calming response. For instance, the following prompt is entered into the generative AI model: "Analyze the user's voice input and generate a response message if a stressed state is detected. Generate a message in a calming, gentle tone saying, 'Take a deep breath and relax.'"
[0195] In this way, the system can provide quick and appropriate communication tailored to the user's emotional state, thereby improving the user's sense of security and satisfaction.
[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0197] Step 1:
[0198] Users enter the necessary information for reservations and inquiries via their device. This includes specific requirements such as the planned date and time of visit and the services they wish to receive. The entered data is temporarily stored on the device.
[0199] Step 2:
[0200] The device transmits the collected user information to the emotion engine. It receives voice and text messages as input and analyzes the user's emotional state using an emotion analysis library. The output is an emotional state (e.g., stress, relaxation).
[0201] Step 3:
[0202] The server receives the analyzed emotional state and user input information. The server uses a natural language processing engine to generate an appropriate response based on the input data. As part of the data processing, the tone and expression are adjusted according to the emotional state. The output is a specific response message.
[0203] Step 4:
[0204] The server sends the generated response content via an external communication network and initiates a voice call with the specified recipient. A communication API is used to establish the voice call connection. Voice communication takes place between the user and the recipient.
[0205] Step 5:
[0206] After the call ends, the server organizes the call content and creates a summary. This includes recording responses from the other party and processing the data to extract important information. The summary results are then notified to the user.
[0207] Step 6:
[0208] Users receive notifications on their devices and can choose the next action as needed. They can also check if they received satisfactory support and request further information if necessary.
[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0212] [Second Embodiment]
[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0225] This invention is a system that automatically handles various types of communication without requiring the user to make a phone call directly. This system is realized through the cooperation of user terminals, servers, and external communication networks.
[0226] Users operate a dedicated application on their device to create phone call or inquiry requests. Users input various request details, such as restaurant reservations, inquiries to companies, and even emergency calls. The device receives this information and sends it to the server.
[0227] The server analyzes the received user data and uses a natural language processing engine to analyze its content. Based on the analyzed information, a generative AI model generates conversational text to be used in the call. This generated conversational content is designed to provide optimal responses and dialogue.
[0228] Next, the server uses an external communication network to initiate a voice call to a designated contact. For example, it might call a restaurant the user wishes to book and proceed with the booking process based on the specified information. Once the call is complete, the server organizes the information gathered during the call and summarizes the results in a format requested by the user.
[0229] The user terminal ultimately receives these results from the server and notifies the user. In other words, the user only needs to confirm the results based on their pre-configured preferences without having to engage in direct conversation.
[0230] For example, if a user requests a reservation for three people at Restaurant X at 7 PM tomorrow, the server automatically calls Restaurant X based on this information and handles the request. The server then notifies the user that the reservation was successful, allowing the user to directly verify the result. This reduces the stress and inconvenience associated with regular phone calls, enabling efficient and secure communication via telephone.
[0231] The following describes the processing flow.
[0232] Step 1:
[0233] Users launch a dedicated application on their device and enter details such as phone reservations or inquiries. This information includes details such as the store name, reservation date and time, and number of people.
[0234] Step 2:
[0235] The terminal formats the data entered by the user and sends it to the server. The data is structured using a standard format suitable for communication, such as JSON.
[0236] Step 3:
[0237] The server receives data sent from the terminal and parses its contents. This parsing identifies what the request is asking for (for example, a restaurant reservation, a product inquiry, etc.).
[0238] Step 4:
[0239] The server passes the analyzed data to a natural language processing engine, which generates natural conversational sentences based on the user's requests. During this process, the generation AI model prepares appropriate responses based on the context.
[0240] Step 5:
[0241] Based on the generated conversation text, the server initiates a voice call to a specified telephone number via an external communication network. A communication module is then used to perform the actual dialing and establish the connection.
[0242] Step 6:
[0243] Once a call is connected, the server automatically initiates a conversation through AI. During the conversation, it adjusts the content of the discussion in real time based on the other party's responses and gathers the necessary information.
[0244] Step 7:
[0245] Once the call is complete, the server organizes the call content and clarifies the results as requested by the user. These results are presented in a format that is easy for the user to understand.
[0246] Step 8:
[0247] The server sends the organized call results to the terminal. The result data is then reorganized according to the format and sent to the user's terminal.
[0248] Step 9:
[0249] The terminal receives the results from the server and notifies the user. The user can then review the results displayed on the terminal screen and take further action or make decisions.
[0250] (Example 1)
[0251] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0252] Traditionally, making inquiries or reservations by phone has been time-consuming and cumbersome, often accompanied by anxiety and stress during the conversation. Furthermore, automated systems have suffered from a lack of immediacy and accuracy. The challenge lies in improving this situation and enabling users to engage in more comfortable and diverse forms of communication.
[0253] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0254] In this invention, the server includes means for receiving and analyzing information from a user device, means for performing language processing based on the analyzed information, and means for generating dialogue content using the processed information. This enables users to communicate quickly and accurately, such as making inquiries and reservations, automatically, without having to make phone calls themselves.
[0255] A "user device" is a terminal operated by the user, and is a device used for inputting and transmitting data.
[0256] "Means for receiving and analyzing information" refers to a mechanism that has the ability to acquire data transmitted from a user device and understand the content of that data.
[0257] "Means of language processing" are technologies that use analyzed information to determine the meaning and context in natural language and construct appropriate dialogue.
[0258] A "generation method for generating dialogue content" is a system that automatically creates conversations in response to user requests and provides them in a format that can be used for voice calls.
[0259] "External communication media" refers to communication methods used for making voice calls, including the internet and public telephone networks.
[0260] "Means for initiating voice communication" refers to a method that automatically starts a call with a designated party based on the generated dialogue content.
[0261] "Means of organizing communication results" refers to the process of organizing information obtained after a call and summarizing it in a way that is easily understandable to the user.
[0262] "Means of notifying the user" refers to a method of sending organized information to the user's device and informing the user of the results.
[0263] This invention provides an automated voice call system that combines a user, a terminal, and a server. This system begins with the user using a dedicated application on the terminal to input various requests. For example, the user might input, "I would like to make a reservation for four people at Restaurant Y for 7 PM tomorrow."
[0264] The terminal, upon receiving user input, transmits that information to the server. The terminal is a typical mobile device or computer, communicating with the server via the internet. This system utilizes HTTPS as its communication protocol to ensure secure data transmission.
[0265] The server analyzes the received data. This analysis uses a natural language processing engine such as the Google Cloud Natural Language API. Based on the analyzed data, the server uses a generative AI model to generate appropriate conversation content. This generative AI model is designed to accurately and effectively utilize the user's intended meaning in the call.
[0266] Based on the generated conversation content, the server initiates voice communication via an external communication medium. Specifically, the server automatically dials the designated recipient using the Public Public Telephone Network (PSTN) or VoIP technology. Once the call is answered, the conversation proceeds using the automatically generated conversation content.
[0267] After the call ends, the server organizes the information gathered during the call. This organization includes transcribing and summarizing the conversation. The organized results are sent to the user's device, and the device notifies the user of the results. The notification is delivered via a message within a dedicated app or through push notifications on the device.
[0268] This entire process eliminates the need for users to make phone calls themselves, allowing them to enjoy automated and efficient communication. For example, a user could use a prompt like, "I'd like to make a reservation for 5 people at Izakaya Z at 8 PM on the weekend," and confirm the reservation using the same procedure. This allows users to save time and effort while achieving accurate communication.
[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0270] Step 1:
[0271] The user opens a dedicated application on their device and enters a specific request. For example, they might enter something like, "I would like to reserve a table for two at Restaurant W tomorrow night at 8 PM." This entered data is temporarily saved on the device as text.
[0272] Step 2:
[0273] The terminal sends the user's input request to the server. The terminal uses HTTP requests to send text data as packets to the server. In this process, the user request, as input, arrives at the server as text data.
[0274] Step 3:
[0275] The server parses the text data of the received request. Here, a natural language processing engine is used to extract the meaning of the text data and convert it into a structured format. The input to this process is a user request in text format, and the output is structured data.
[0276] Step 4:
[0277] The server uses a generative AI model based on structured data to generate appropriate conversation content. This generation process uses prompt sentences as input, and the generated conversation is stored on the server in text format. The output is a customized conversation based on the user's intent.
[0278] Step 5:
[0279] The server uses the generated conversation content to initiate voice communication with a designated recipient via an external communication medium. The server uses a telephone API to convert the generated conversation content into speech and transmit it to the recipient. The input for this step is the generated conversation content, and the output is a real-time voice call.
[0280] Step 6:
[0281] After the call is completed, the server organizes the results. It documents the information obtained during the call as text and extracts the key points. The input to this process is the audio data of the call result, and the output is the textual result.
[0282] Step 7:
[0283] The server sends the sorted call results to the terminal, and the terminal notifies the user of the results. On the terminal, the results are displayed in the notification bar or the app interface, and the user can easily check the content. The input for this step is the sorted call results, and the output is the result notification presented to the user.
[0284] (Application Example 1)
[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0286] Currently, many users feel the inconvenience of having to directly contact service providers when making electronic payments. Also, manual payment procedures are prone to errors and omissions, which may result in wasted time and effort. There is a need to establish a method to solve such problems and enable users to efficiently perform payment procedures without stress.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0288] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, generating means for generating conversation content using the processed information, means for starting voice communication via an external information transmission network, means for sorting the results of the communication and notifying the sorted results to the user, and means for automatically inquiring the payment destination specified by the user and completing the payment. As a result, the user can quickly and accurately complete the payment procedure through an automated process.
[0289] "Means of receiving and analyzing information from users" refers to a method in which a server receives data sent by a user and performs analysis to understand it.
[0290] "Means of natural language processing" refers to technologies used by computers to understand human language and to provide appropriate responses and processing.
[0291] "A means for generating dialogue content" refers to a method of creating a conversation used for communication with the recipient, based on analyzed data.
[0292] "Means of initiating voice communication via an external information transmission network" refers to a method of initiating a voice call and having a conversation with another person through an external communication system.
[0293] "A means of organizing the results of communication and notifying the user of those organized results" refers to a method of summarizing the results obtained during the communication process and conveying them to the user in an easy-to-understand manner.
[0294] "A means of automatically contacting the payment recipient specified by the user and completing the payment" refers to a method of automatically contacting the payment recipient specified by the user and completing the payment procedure.
[0295] To implement this invention, the server receives payment information transmitted from the user's terminal. The received information is first processed as data by a dedicated analysis engine, and its contents are analyzed in detail. A natural language processing engine (e.g., Google Cloud Natural Language API) is used for the analysis. In this process, the server accurately understands the user's intent and determines the specific steps to be taken next.
[0296] Specifically, based on the results of analysis using natural language processing, a generative AI model (e.g., OpenAI GPT series) generates the dialogue content necessary for the payment process. The generated dialogue content is then used to automatically inquire with the company or service provider specified by the user.
[0297] The communication method involves voice communication initiated via an external information transmission network. Specifically, this system, configured through the user's terminal, engages in real-time conversation based on payment information, obtains necessary information, and completes the payment. After communication is complete, the server organizes the results and notifies the user's terminal of the updated payment status.
[0298] For example, if a user wants to pay their monthly electricity bill to the power company, the user enters "Pay this month's bill to the power company" into the application. The system then automatically contacts the power company and initiates the payment process. An example of a prompt message would be, "Please obtain the procedure to complete this month's payment to the power company." In this way, the present invention provides a system that allows users to complete payments quickly and accurately without having to go through cumbersome procedures.
[0299] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0300] Step 1:
[0301] The user launches the application on their device and enters payment details and amounts. The entered information is then sent to the server as electronic data.
[0302] Step 2:
[0303] The server analyzes the received data. Using a natural language processing engine, it analyzes the user's input from a linguistic perspective and extracts the elements necessary for the payment process. The input is raw data from the user, and the output is a structured analysis result.
[0304] Step 3:
[0305] Based on the analysis results, the server uses the generative AI model to generate payment-related inquiry conversations. Using a prompt sentence (e.g., "Please obtain the procedures for completing this month's payment to the power company") as the input, the generated sentence is output.
[0306] Step 4:
[0307] The server starts a voice communication with the designated payee via an external information transmission network. Using the generated inquiry sentence, a voice call is made to obtain information in real time.
[0308] Step 5:
[0309] After the call ends, the server organizes the obtained information. Specifically, it summarizes the payment approval or rejection and processing status, etc., and constructs data for notifying the user. The input is the raw data obtained from the call, and the output is the organized information for notifying the user.
[0310] Step 6:
[0311] The terminal receives the notification from the server and displays the payment status and results to the user. By checking this information, the user can know that the payment has been completed.
[0312] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0313] The present invention is a system that identifies the user's emotional state and optimizes the process of making a call based on it. This system is realized in a configuration including a user terminal, a server, and an emotion engine, and automatically provides appropriate communication.
[0314] Users input the necessary information through a dedicated application on their device. For example, if they wish to make a reservation or inquiry, they input the target location and detailed requests. The device collects this data and uses an emotion engine to determine the user's emotional state in the process. The emotion engine analyzes the user's emotions from the input text and voice, accurately capturing their state.
[0315] The server receives data and emotional state transmitted from the terminal. Based on the analyzed emotional information, the natural language processing engine generates appropriate conversation content. In this process, the tone and expressions are adapted to the user's emotional state. For example, if the user is nervous, a calm tone of voice will be generated.
[0316] Next, the server initiates a voice call to the designated recipient via an external communication network. This call process is dynamically managed by prioritizing based on the user's emotions. If a response is received from the recipient during the conversation, the server updates the generated conversation content in real time and records the response.
[0317] After the call ends, the server compiles the information received and notifies the user. This includes confirmation of the reservation and answers to inquiries. The user can then check the results on their device and take further action as needed.
[0318] For example, if a user is feeling stressed, the system will prioritize the call as an urgent matter and complete it quickly. Conversely, if the user is relaxed, the system will provide a calm and courteous response, ensuring a comfortable calling experience.
[0319] Thus, the present invention provides a means to realize communication that takes user emotions into consideration and to resolve the problems of conventional telephone communication.
[0320] The following describes the processing flow.
[0321] Step 1:
[0322] Users launch a dedicated application on their device and enter their phone reservation or inquiry details. They can also add custom comments or voice messages to express their emotions.
[0323] Step 2:
[0324] The device collects data entered by the user and analyzes the user's emotional state using an emotion engine. It analyzes text and voice data to identify emotional patterns.
[0325] Step 3:
[0326] The device sends the analysis results to the server. This includes data on the user's request, desired action, and emotional state. The data is structured in a standard format.
[0327] Step 4:
[0328] The server analyzes the data received from the terminal and uses a natural language processing engine to generate appropriate responses to requests. Based on sentiment data, it adjusts the tone and expression of the conversation.
[0329] Step 5:
[0330] The server considers the generated conversation content and emotional state, and initiates a voice call to the designated recipient via an external communication network. It is also possible to set call priorities based on emotional state.
[0331] Step 6:
[0332] The server updates the conversation content in real time during the call using a dialogue generation AI. Based on the other party's responses, it retrieves appropriate information as needed and continues the conversation.
[0333] Step 7:
[0334] After the call is completed, the server organizes the information obtained during the call and compiles the data in a format that aligns with the user's request.
[0335] Step 8:
[0336] The server sends the processed results to the user's terminal. This allows the user to check the call results and responses.
[0337] Step 9:
[0338] The device notifies the user of the received results and displays the details on the screen. The user can then decide on their next action based on the results.
[0339] (Example 2)
[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0341] Traditional communication technologies often suffer from a decline in communication quality because they do not take into account the user's emotional state when responding or prioritizing calls. In particular, when users are experiencing stress or anxiety, appropriate responses and prompt communication are often lacking. This highlights the need for improved user experience.
[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0343] In this invention, the server includes means for receiving and collecting user input information, means for analyzing the user's emotional state from the collected information, and means for initiating voice communication via an external communication network using the generated conversation content. This makes it possible to set appropriate responses and priorities according to the user's emotional state.
[0344] "User input information" refers to data such as reservation and inquiry details, as well as any additional conditions, that users submit to the system.
[0345] "Emotional state" refers to the user's psychological state and reactions, and is identified through the analysis of text and voice.
[0346] "Natural language processing" is a technology that enables computers to understand and generate human language, and it is a method for creating conversational content based on analyzed data.
[0347] "Conversation content" refers to the voice or text-based responses generated by the system, which are adjusted to suit the user's emotional state.
[0348] "External communication networks" refer to communication infrastructure such as the internet and telephone lines that servers use to communicate.
[0349] "Call priority" is an indicator that shows the importance of a call based on the user's emotional state, and is dynamically set according to urgency.
[0350] "Response content" refers to real-time response information from the other party, and is data acquired during the course of the conversation.
[0351] The following describes embodiments for carrying out the present invention.
[0352] The system of the present invention mainly consists of a user terminal, a communication server, and an emotion analysis engine. The user uses a dedicated application on the terminal to input necessary information regarding reservations and inquiries. This application includes hardware (e.g., microphone, keyboard) to acquire information entered in the form of text or voice, and software (e.g., speech recognition software, text analysis engine) to process the data.
[0353] The device uses an emotion analysis engine to determine the user's emotional state from the input information. The emotion analysis engine combines natural language processing and speech analysis technologies, and can capture emotions by analyzing, for example, the tone and speed of voice, and the emotional vocabulary in text.
[0354] The analyzed emotion information and user input data are sent to the server. The server uses a natural language processing engine to generate conversation content adapted to the user's emotions. This engine can create multiple response patterns using a generative AI model and select the optimal one. The generated conversation content is transmitted to the designated recipient via an external communication network, and a voice call is initiated.
[0355] During a call, the server updates the conversation in real time based on the responses received. Furthermore, call priorities are dynamically managed based on the user's emotional state, ensuring that urgent calls are processed without delay.
[0356] For example, if a user enters the text "I want to make a dentist appointment for next Tuesday," the device's emotion analysis engine detects that the user is nervous. Based on this information, the server prepares a response in a relaxed tone and prioritizes initiating the call with the dentist.
[0357] An example of a prompt message is, "Generate a call message to confirm a reservation while the user is feeling stressed." This prompt allows the system to generate the most appropriate conversation for the situation. This prompt serves as input for the generation AI model to function correctly.
[0358] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0359] Step 1:
[0360] Users enter reservation and inquiry information using a dedicated application on their terminal. This input data includes information in text or audio format. The terminal converts this information into digital signals and applies them as input to an emotion analysis engine. This engine analyzes the user's emotions from the tone of voice and the context of the text, and outputs their emotional state.
[0361] Step 2:
[0362] The terminal sends the emotional state obtained by the emotion analysis engine and the information entered by the user to the server. The server activates a natural language processing engine based on the received data. This engine processes the conversation data using a generative AI model to generate optimal conversation content that corresponds to the user's emotions. The generated conversation content is output and used for the next processing step.
[0363] Step 3:
[0364] The server uses the generated conversation content to initiate a voice call to the designated recipient via an external communication network. During this process, priority is initially set based on sentiment information, with urgent calls being processed first. The server establishes an external connection and prepares to begin the conversation.
[0365] Step 4:
[0366] During a call, the server updates the conversation content in real time based on information received from the other party. The input is response data from the other party, which the server analyzes and outputs newly generated response content using a generative AI model. The response is applied immediately, maintaining the quality of the call.
[0367] Step 5:
[0368] After the call ends, the server processes the call results and notifies the user. The acquired information is stored in a database and displayed to help the user easily decide on the next steps. As a final output, the user is provided with a reservation confirmation and answers to their inquiries.
[0369] (Application Example 2)
[0370] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0371] In modern information and communication technology, users can experience anxiety and stress in various situations. These emotional states can lead to inefficiencies and problems in communication. However, existing systems lack sufficient means to accurately monitor users' emotional states and provide support at the appropriate time. Therefore, there is a need to support users' mental well-being and facilitate smooth communication.
[0372] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0373] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, and means for monitoring the user's emotional state and providing prompt support when an anomaly is detected. This makes it possible to monitor the user's emotional state in real time, provide appropriate language responses as needed, and provide rapid support in emergencies.
[0374] "Means of receiving and analyzing information from users" refers to functions that acquire data provided by users, analyze that data, and understand the user's state and requests.
[0375] "Means of performing natural language processing" refers to a function that uses acquired user information to enable a computer to understand the context and derive an appropriate response.
[0376] "Generating means for generating response content" refers to a function that creates appropriate responses or messages for the user based on the results of analysis using natural language processing.
[0377] "Means of initiating a voice call via an external communication network" refers to a function that uses a communication network to initiate a voice-based conversation with a designated party.
[0378] "A means of organizing call results and notifying the user of those organized results" refers to a function that summarizes the content and results of a voice call and reports or provides feedback to the user.
[0379] "A means of monitoring the user's emotional state and providing rapid support when an anomaly is detected" refers to a function that monitors the user's emotions in real time and provides immediate and appropriate support when an anomaly such as stress or panic is observed.
[0380] The system for realizing this invention is designed to monitor the user's emotional state in real time and provide appropriate responses. Its main components include a terminal used by the user, a server containing an emotion engine for analyzing the user's emotions, and a communication network for voice communication.
[0381] The user accesses the system through a terminal and enters the necessary information. This terminal is an internet-connected device such as a smartphone or wearable device. User information is collected by the terminal and sent to the server via the emotion engine. The emotion engine analyzes the user's emotional state from voice and text messages, using, for example, Python or an emotion analysis library (such as Hugging Face's Transformers). Based on this analysis, a natural language processing engine generates a response appropriate to the user.
[0382] Based on the generated response, the server initiates a voice call with the specified party using an external communication network. During this process, the server can utilize a communication API (e.g., Twilio). Once the call ends, the server can organize the call content and notify the user of the results.
[0383] As a concrete example, when a user is feeling stressed, the device reports its emotional state to the server, which then uses a generative AI model to generate a calming response. For instance, the following prompt is entered into the generative AI model: "Analyze the user's voice input and generate a response message if a stressed state is detected. Generate a message in a calming, gentle tone saying, 'Take a deep breath and relax.'"
[0384] In this way, the system can provide quick and appropriate communication tailored to the user's emotional state, thereby improving the user's sense of security and satisfaction.
[0385] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0386] Step 1:
[0387] Users enter the necessary information for reservations and inquiries via their device. This includes specific requirements such as the planned date and time of visit and the services they wish to receive. The entered data is temporarily stored on the device.
[0388] Step 2:
[0389] The device transmits the collected user information to the emotion engine. It receives voice and text messages as input and analyzes the user's emotional state using an emotion analysis library. The output is an emotional state (e.g., stress, relaxation).
[0390] Step 3:
[0391] The server receives the analyzed emotional state and user input information. The server uses a natural language processing engine to generate an appropriate response based on the input data. As part of the data processing, the tone and expression are adjusted according to the emotional state. The output is a specific response message.
[0392] Step 4:
[0393] The server sends the generated response content via an external communication network and initiates a voice call with the specified recipient. A communication API is used to establish the voice call connection. Voice communication takes place between the user and the recipient.
[0394] Step 5:
[0395] After the call ends, the server organizes the call content and creates a summary. This includes recording responses from the other party and processing the data to extract important information. The summary results are then notified to the user.
[0396] Step 6:
[0397] Users receive notifications on their devices and can choose the next action as needed. They can also check if they received satisfactory support and request further information if necessary.
[0398] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0399] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0400] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0401] [Third Embodiment]
[0402] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0403] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0404] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0405] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0406] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0407] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0408] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0409] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0410] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0411] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0412] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0413] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0414] This invention is a system that automatically handles various types of communication without requiring the user to make a phone call directly. This system is realized through the cooperation of user terminals, servers, and external communication networks.
[0415] Users operate a dedicated application on their device to create phone call or inquiry requests. Users input various request details, such as restaurant reservations, inquiries to companies, and even emergency calls. The device receives this information and sends it to the server.
[0416] The server analyzes the received user data and uses a natural language processing engine to analyze its content. Based on the analyzed information, a generative AI model generates conversational text to be used in the call. This generated conversational content is designed to provide optimal responses and dialogue.
[0417] Next, the server uses an external communication network to initiate a voice call to a designated contact. For example, it might call a restaurant the user wishes to book and proceed with the booking process based on the specified information. Once the call is complete, the server organizes the information gathered during the call and summarizes the results in a format requested by the user.
[0418] The user terminal ultimately receives these results from the server and notifies the user. In other words, the user only needs to confirm the results based on their pre-configured preferences without having to engage in direct conversation.
[0419] For example, if a user requests a reservation for three people at Restaurant X at 7 PM tomorrow, the server automatically calls Restaurant X based on this information and handles the request. The server then notifies the user that the reservation was successful, allowing the user to directly verify the result. This reduces the stress and inconvenience associated with regular phone calls, enabling efficient and secure communication via telephone.
[0420] The following describes the processing flow.
[0421] Step 1:
[0422] Users launch a dedicated application on their device and enter details such as phone reservations or inquiries. This information includes details such as the store name, reservation date and time, and number of people.
[0423] Step 2:
[0424] The terminal formats the data entered by the user and sends it to the server. The data is structured using a standard format suitable for communication, such as JSON.
[0425] Step 3:
[0426] The server receives data sent from the terminal and parses its contents. This parsing identifies what the request is asking for (for example, a restaurant reservation, a product inquiry, etc.).
[0427] Step 4:
[0428] The server passes the analyzed data to a natural language processing engine, which generates natural conversational sentences based on the user's requests. During this process, the generation AI model prepares appropriate responses based on the context.
[0429] Step 5:
[0430] Based on the generated conversation text, the server initiates a voice call to a specified telephone number via an external communication network. A communication module is then used to perform the actual dialing and establish the connection.
[0431] Step 6:
[0432] Once a call is connected, the server automatically initiates a conversation through AI. During the conversation, it adjusts the content of the discussion in real time based on the other party's responses and gathers the necessary information.
[0433] Step 7:
[0434] Once the call is complete, the server organizes the call content and clarifies the results as requested by the user. These results are presented in a format that is easy for the user to understand.
[0435] Step 8:
[0436] The server sends the organized call results to the terminal. The result data is then reorganized according to the format and sent to the user's terminal.
[0437] Step 9:
[0438] The terminal receives the results from the server and notifies the user. The user can then review the results displayed on the terminal screen and take further action or make decisions.
[0439] (Example 1)
[0440] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0441] Traditionally, making inquiries or reservations by phone has been time-consuming and cumbersome, often accompanied by anxiety and stress during the conversation. Furthermore, automated systems have suffered from a lack of immediacy and accuracy. The challenge lies in improving this situation and enabling users to engage in more comfortable and diverse forms of communication.
[0442] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0443] In this invention, the server includes means for receiving and analyzing information from a user device, means for performing language processing based on the analyzed information, and means for generating dialogue content using the processed information. This enables users to communicate quickly and accurately, such as making inquiries and reservations, automatically, without having to make phone calls themselves.
[0444] A "user device" is a terminal operated by the user, and is a device used for inputting and transmitting data.
[0445] "Means for receiving and analyzing information" refers to a mechanism that has the ability to acquire data transmitted from a user device and understand the content of that data.
[0446] "Means of language processing" are technologies that use analyzed information to determine the meaning and context in natural language and construct appropriate dialogue.
[0447] A "generation method for generating dialogue content" is a system that automatically creates conversations in response to user requests and provides them in a format that can be used for voice calls.
[0448] "External communication media" refers to communication methods used for making voice calls, including the internet and public telephone networks.
[0449] "Means for initiating voice communication" refers to a method that automatically starts a call with a designated party based on the generated dialogue content.
[0450] "Means of organizing communication results" refers to the process of organizing information obtained after a call and summarizing it in a way that is easily understandable to the user.
[0451] "Means of notifying the user" refers to a method of sending organized information to the user's device and informing the user of the results.
[0452] This invention provides an automated voice call system that combines a user, a terminal, and a server. This system begins with the user using a dedicated application on the terminal to input various requests. For example, the user might input, "I would like to make a reservation for four people at Restaurant Y for 7 PM tomorrow."
[0453] The terminal, upon receiving user input, transmits that information to the server. The terminal is a typical mobile device or computer, communicating with the server via the internet. This system utilizes HTTPS as its communication protocol to ensure secure data transmission.
[0454] The server analyzes the received data. This analysis uses a natural language processing engine such as the Google Cloud Natural Language API. Based on the analyzed data, the server uses a generative AI model to generate appropriate conversation content. This generative AI model is designed to accurately and effectively utilize the user's intended meaning in the call.
[0455] Based on the generated conversation content, the server initiates voice communication via an external communication medium. Specifically, the server automatically dials the designated recipient using the Public Public Telephone Network (PSTN) or VoIP technology. Once the call is answered, the conversation proceeds using the automatically generated conversation content.
[0456] After the call ends, the server organizes the information gathered during the call. This organization includes transcribing and summarizing the conversation. The organized results are sent to the user's device, and the device notifies the user of the results. The notification is delivered via a message within a dedicated app or through push notifications on the device.
[0457] This entire process eliminates the need for users to make phone calls themselves, allowing them to enjoy automated and efficient communication. For example, a user could use a prompt like, "I'd like to make a reservation for 5 people at Izakaya Z at 8 PM on the weekend," and confirm the reservation using the same procedure. This allows users to save time and effort while achieving accurate communication.
[0458] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0459] Step 1:
[0460] The user opens a dedicated application on their device and enters a specific request. For example, they might enter something like, "I would like to reserve a table for two at Restaurant W tomorrow night at 8 PM." This entered data is temporarily saved on the device as text.
[0461] Step 2:
[0462] The terminal sends the user's input request to the server. The terminal uses HTTP requests to send text data as packets to the server. In this process, the user request, as input, arrives at the server as text data.
[0463] Step 3:
[0464] The server parses the text data of the received request. Here, a natural language processing engine is used to extract the meaning of the text data and convert it into a structured format. The input to this process is a user request in text format, and the output is structured data.
[0465] Step 4:
[0466] The server uses a generative AI model based on structured data to generate appropriate conversation content. This generation process uses prompt sentences as input, and the generated conversation is stored on the server in text format. The output is a customized conversation based on the user's intent.
[0467] Step 5:
[0468] The server uses the generated conversation content to initiate voice communication with a designated recipient via an external communication medium. The server uses a telephone API to convert the generated conversation content into speech and transmit it to the recipient. The input for this step is the generated conversation content, and the output is a real-time voice call.
[0469] Step 6:
[0470] After the call is completed, the server organizes the results. It documents the information obtained during the call as text and extracts the key points. The input to this process is the audio data of the call result, and the output is the textual result.
[0471] Step 7:
[0472] The server sends the organized call results to the terminal, and the terminal notifies the user of the results. On the terminal, the results are displayed in the notification bar or app interface, allowing the user to easily check the content. The input for this step is the organized call results, and the output is the result notification presented to the user.
[0473] (Application Example 1)
[0474] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0475] Currently, many users experience the inconvenience of having to directly contact service providers when making electronic payments. Furthermore, manual payment procedures are prone to errors and omissions, resulting in wasted time and effort. There is a need to solve these problems and establish a method that allows users to complete payment procedures efficiently and without stress.
[0476] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0477] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, means for generating dialogue content using the processed information, means for initiating voice communication via an external information transmission network, means for organizing the results of the communication and notifying the user of the organized results, and means for automatically contacting the payee specified by the user and completing the payment. This enables the user to complete the payment procedure quickly and accurately through an automated process.
[0478] "Means of receiving and analyzing information from users" refers to a method in which a server receives data sent by a user and performs analysis to understand it.
[0479] "Means of natural language processing" refers to technologies used by computers to understand human language and to provide appropriate responses and processing.
[0480] "A means for generating dialogue content" refers to a method of creating a conversation used for communication with the recipient, based on analyzed data.
[0481] "Means of initiating voice communication via an external information transmission network" refers to a method of initiating a voice call and having a conversation with another person through an external communication system.
[0482] "A means of organizing the results of communication and notifying the user of those organized results" refers to a method of summarizing the results obtained during the communication process and conveying them to the user in an easy-to-understand manner.
[0483] "A means of automatically contacting the payment recipient specified by the user and completing the payment" refers to a method of automatically contacting the payment recipient specified by the user and completing the payment procedure.
[0484] To implement this invention, the server receives payment information transmitted from the user's terminal. The received information is first processed as data by a dedicated analysis engine, and its contents are analyzed in detail. A natural language processing engine (e.g., Google Cloud Natural Language API) is used for the analysis. In this process, the server accurately understands the user's intent and determines the specific steps to be taken next.
[0485] Specifically, based on the results of analysis using natural language processing, a generative AI model (e.g., OpenAI GPT series) generates the dialogue content necessary for the payment process. The generated dialogue content is then used to automatically inquire with the company or service provider specified by the user.
[0486] The communication method involves voice communication initiated via an external information transmission network. Specifically, this system, configured through the user's terminal, engages in real-time conversation based on payment information, obtains necessary information, and completes the payment. After communication is complete, the server organizes the results and notifies the user's terminal of the updated payment status.
[0487] For example, if a user wants to pay their monthly electricity bill to the power company, the user enters "Pay this month's bill to the power company" into the application. The system then automatically contacts the power company and initiates the payment process. An example of a prompt message would be, "Please obtain the procedure to complete this month's payment to the power company." In this way, the present invention provides a system that allows users to complete payments quickly and accurately without having to go through cumbersome procedures.
[0488] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0489] Step 1:
[0490] The user launches the application on their device and enters payment details and amounts. The entered information is then sent to the server as electronic data.
[0491] Step 2:
[0492] The server analyzes the received data. Using a natural language processing engine, it analyzes the user's input from a linguistic perspective and extracts the elements necessary for the payment process. The input is raw data from the user, and the output is a structured analysis result.
[0493] Step 3:
[0494] Based on the analysis results, the server uses a generation AI model to generate payment-related inquiry conversations. It takes a prompt sentence (for example, "Please get instructions on how to complete this month's payment to the utility company") as input and outputs the generated sentence.
[0495] Step 4:
[0496] The server initiates voice communication with the designated payee via an external information transmission network. Using the generated query, it conducts a voice call and retrieves information in real time.
[0497] Step 5:
[0498] After the call ends, the server organizes the acquired information. Specifically, it summarizes information such as payment status and processing status, and constructs data to notify the user. The input is the raw data acquired from the call, and the output is the organized information for notification to the user.
[0499] Step 6:
[0500] The device receives notifications from the server and displays the payment status and results to the user. By checking this information, the user can know that the payment has been completed.
[0501] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0502] This invention relates to a system that identifies a user's emotional state and optimizes the process of making phone calls based on that state. The system is implemented with a configuration including a user terminal, a server, and an emotion engine, and automatically provides appropriate communication.
[0503] Users input the necessary information through a dedicated application on their device. For example, if they wish to make a reservation or inquiry, they input the target location and detailed requests. The device collects this data and uses an emotion engine to determine the user's emotional state in the process. The emotion engine analyzes the user's emotions from the input text and voice, accurately capturing their state.
[0504] The server receives data and emotional state transmitted from the terminal. Based on the analyzed emotional information, the natural language processing engine generates appropriate conversation content. In this process, the tone and expressions are adapted to the user's emotional state. For example, if the user is nervous, a calm tone of voice will be generated.
[0505] Next, the server initiates a voice call to the designated recipient via an external communication network. This call process is dynamically managed by prioritizing based on the user's emotions. If a response is received from the recipient during the conversation, the server updates the generated conversation content in real time and records the response.
[0506] After the call ends, the server compiles the information received and notifies the user. This includes confirmation of the reservation and answers to inquiries. The user can then check the results on their device and take further action as needed.
[0507] For example, if a user is feeling stressed, the system will prioritize the call as an urgent matter and complete it quickly. Conversely, if the user is relaxed, the system will provide a calm and courteous response, ensuring a comfortable calling experience.
[0508] Thus, the present invention provides a means to realize communication that takes user emotions into consideration and to resolve the problems of conventional telephone communication.
[0509] The following describes the processing flow.
[0510] Step 1:
[0511] Users launch a dedicated application on their device and enter their phone reservation or inquiry details. They can also add custom comments or voice messages to express their emotions.
[0512] Step 2:
[0513] The device collects data entered by the user and analyzes the user's emotional state using an emotion engine. It analyzes text and voice data to identify emotional patterns.
[0514] Step 3:
[0515] The device sends the analysis results to the server. This includes data on the user's request, desired action, and emotional state. The data is structured in a standard format.
[0516] Step 4:
[0517] The server analyzes the data received from the terminal and uses a natural language processing engine to generate appropriate responses to requests. Based on sentiment data, it adjusts the tone and expression of the conversation.
[0518] Step 5:
[0519] The server considers the generated conversation content and emotional state, and initiates a voice call to the designated recipient via an external communication network. It is also possible to set call priorities based on emotional state.
[0520] Step 6:
[0521] The server updates the conversation content in real time during the call using a dialogue generation AI. Based on the other party's responses, it retrieves appropriate information as needed and continues the conversation.
[0522] Step 7:
[0523] After the call is completed, the server organizes the information obtained during the call and compiles the data in a format that aligns with the user's request.
[0524] Step 8:
[0525] The server sends the processed results to the user's terminal. This allows the user to check the call results and responses.
[0526] Step 9:
[0527] The device notifies the user of the received results and displays the details on the screen. The user can then decide on their next action based on the results.
[0528] (Example 2)
[0529] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] Traditional communication technologies often suffer from a decline in communication quality because they do not take into account the user's emotional state when responding or prioritizing calls. In particular, when users are experiencing stress or anxiety, appropriate responses and prompt communication are often lacking. This highlights the need for improved user experience.
[0531] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0532] In this invention, the server includes means for receiving and collecting user input information, means for analyzing the user's emotional state from the collected information, and means for initiating voice communication via an external communication network using the generated conversation content. This makes it possible to set appropriate responses and priorities according to the user's emotional state.
[0533] "User input information" refers to data such as reservation and inquiry details, as well as any additional conditions, that users submit to the system.
[0534] "Emotional state" refers to the user's psychological state and reactions, and is identified through the analysis of text and voice.
[0535] "Natural language processing" is a technology that enables computers to understand and generate human language, and it is a method for creating conversational content based on analyzed data.
[0536] "Conversation content" refers to the voice or text-based responses generated by the system, which are adjusted to suit the user's emotional state.
[0537] "External communication networks" refer to communication infrastructure such as the internet and telephone lines that servers use to communicate.
[0538] "Call priority" is an indicator that shows the importance of a call based on the user's emotional state, and is dynamically set according to urgency.
[0539] "Response content" refers to real-time response information from the other party, and is data acquired during the course of the conversation.
[0540] The following describes embodiments for carrying out the present invention.
[0541] The system of the present invention mainly consists of a user terminal, a communication server, and an emotion analysis engine. The user uses a dedicated application on the terminal to input necessary information regarding reservations and inquiries. This application includes hardware (e.g., microphone, keyboard) to acquire information entered in the form of text or voice, and software (e.g., speech recognition software, text analysis engine) to process the data.
[0542] The device uses an emotion analysis engine to determine the user's emotional state from the input information. The emotion analysis engine combines natural language processing and speech analysis technologies, and can capture emotions by analyzing, for example, the tone and speed of voice, and the emotional vocabulary in text.
[0543] The analyzed emotion information and user input data are sent to the server. The server uses a natural language processing engine to generate conversation content adapted to the user's emotions. This engine can create multiple response patterns using a generative AI model and select the optimal one. The generated conversation content is transmitted to the designated recipient via an external communication network, and a voice call is initiated.
[0544] During a call, the server updates the conversation in real time based on the responses received. Furthermore, call priorities are dynamically managed based on the user's emotional state, ensuring that urgent calls are processed without delay.
[0545] For example, if a user enters the text "I want to make a dentist appointment for next Tuesday," the device's emotion analysis engine detects that the user is nervous. Based on this information, the server prepares a response in a relaxed tone and prioritizes initiating the call with the dentist.
[0546] An example of a prompt message is, "Generate a call message to confirm a reservation while the user is feeling stressed." This prompt allows the system to generate the most appropriate conversation for the situation. This prompt serves as input for the generation AI model to function correctly.
[0547] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0548] Step 1:
[0549] Users enter reservation and inquiry information using a dedicated application on their terminal. This input data includes information in text or audio format. The terminal converts this information into digital signals and applies them as input to an emotion analysis engine. This engine analyzes the user's emotions from the tone of voice and the context of the text, and outputs their emotional state.
[0550] Step 2:
[0551] The terminal sends the emotional state obtained by the emotion analysis engine and the information entered by the user to the server. The server activates a natural language processing engine based on the received data. This engine processes the conversation data using a generative AI model to generate optimal conversation content that corresponds to the user's emotions. The generated conversation content is output and used for the next processing step.
[0552] Step 3:
[0553] The server uses the generated conversation content to initiate a voice call to the designated recipient via an external communication network. During this process, priority is initially set based on sentiment information, with urgent calls being processed first. The server establishes an external connection and prepares to begin the conversation.
[0554] Step 4:
[0555] During a call, the server updates the conversation content in real time based on information received from the other party. The input is response data from the other party, which the server analyzes and outputs newly generated response content using a generative AI model. The response is applied immediately, maintaining the quality of the call.
[0556] Step 5:
[0557] After the call ends, the server processes the call results and notifies the user. The acquired information is stored in a database and displayed to help the user easily decide on the next steps. As a final output, the user is provided with a reservation confirmation and answers to their inquiries.
[0558] (Application Example 2)
[0559] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0560] In modern information and communication technology, users can experience anxiety and stress in various situations. These emotional states can lead to inefficiencies and problems in communication. However, existing systems lack sufficient means to accurately monitor users' emotional states and provide support at the appropriate time. Therefore, there is a need to support users' mental well-being and facilitate smooth communication.
[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0562] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, and means for monitoring the user's emotional state and providing prompt support when an anomaly is detected. This makes it possible to monitor the user's emotional state in real time, provide appropriate language responses as needed, and provide rapid support in emergencies.
[0563] "Means of receiving and analyzing information from users" refers to functions that acquire data provided by users, analyze that data, and understand the user's state and requests.
[0564] "Means of performing natural language processing" refers to a function that uses acquired user information to enable a computer to understand the context and derive an appropriate response.
[0565] "Generating means for generating response content" refers to a function that creates appropriate responses or messages for the user based on the results of analysis using natural language processing.
[0566] "Means of initiating a voice call via an external communication network" refers to a function that uses a communication network to initiate a voice-based conversation with a designated party.
[0567] "A means of organizing call results and notifying the user of those organized results" refers to a function that summarizes the content and results of a voice call and reports or provides feedback to the user.
[0568] "A means of monitoring the user's emotional state and providing rapid support when an anomaly is detected" refers to a function that monitors the user's emotions in real time and provides immediate and appropriate support when an anomaly such as stress or panic is observed.
[0569] The system for realizing this invention is designed to monitor the user's emotional state in real time and provide appropriate responses. Its main components include a terminal used by the user, a server containing an emotion engine for analyzing the user's emotions, and a communication network for voice communication.
[0570] The user accesses the system through a terminal and enters the necessary information. This terminal is an internet-connected device such as a smartphone or wearable device. User information is collected by the terminal and sent to the server via the emotion engine. The emotion engine analyzes the user's emotional state from voice and text messages, using, for example, Python or an emotion analysis library (such as Hugging Face's Transformers). Based on this analysis, a natural language processing engine generates a response appropriate to the user.
[0571] Based on the generated response, the server initiates a voice call with the specified party using an external communication network. During this process, the server can utilize a communication API (e.g., Twilio). Once the call ends, the server can organize the call content and notify the user of the results.
[0572] As a concrete example, when a user is feeling stressed, the device reports its emotional state to the server, which then uses a generative AI model to generate a calming response. For instance, the following prompt is entered into the generative AI model: "Analyze the user's voice input and generate a response message if a stressed state is detected. Generate a message in a calming, gentle tone saying, 'Take a deep breath and relax.'"
[0573] In this way, the system can provide quick and appropriate communication tailored to the user's emotional state, thereby improving the user's sense of security and satisfaction.
[0574] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0575] Step 1:
[0576] Users enter the necessary information for reservations and inquiries via their device. This includes specific requirements such as the planned date and time of visit and the services they wish to receive. The entered data is temporarily stored on the device.
[0577] Step 2:
[0578] The device transmits the collected user information to the emotion engine. It receives voice and text messages as input and analyzes the user's emotional state using an emotion analysis library. The output is an emotional state (e.g., stress, relaxation).
[0579] Step 3:
[0580] The server receives the analyzed emotional state and user input information. The server uses a natural language processing engine to generate an appropriate response based on the input data. As part of the data processing, the tone and expression are adjusted according to the emotional state. The output is a specific response message.
[0581] Step 4:
[0582] The server sends the generated response content via an external communication network and initiates a voice call with the specified recipient. A communication API is used to establish the voice call connection. Voice communication takes place between the user and the recipient.
[0583] Step 5:
[0584] After the call ends, the server organizes the call content and creates a summary. This includes recording responses from the other party and processing the data to extract important information. The summary results are then notified to the user.
[0585] Step 6:
[0586] Users receive notifications on their devices and can choose the next action as needed. They can also check if they received satisfactory support and request further information if necessary.
[0587] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0588] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0589] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0590] [Fourth Embodiment]
[0591] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0592] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0593] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0594] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0595] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0596] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0597] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0598] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0599] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0600] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0601] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0602] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0603] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0604] This invention is a system that automatically handles various types of communication without requiring the user to make a phone call directly. This system is realized through the cooperation of user terminals, servers, and external communication networks.
[0605] Users operate a dedicated application on their device to create phone call or inquiry requests. Users input various request details, such as restaurant reservations, inquiries to companies, and even emergency calls. The device receives this information and sends it to the server.
[0606] The server analyzes the received user data and uses a natural language processing engine to analyze its content. Based on the analyzed information, a generative AI model generates conversational text to be used in the call. This generated conversational content is designed to provide optimal responses and dialogue.
[0607] Next, the server uses an external communication network to initiate a voice call to a designated contact. For example, it might call a restaurant the user wishes to book and proceed with the booking process based on the specified information. Once the call is complete, the server organizes the information gathered during the call and summarizes the results in a format requested by the user.
[0608] The user terminal ultimately receives these results from the server and notifies the user. In other words, the user only needs to confirm the results based on their pre-configured preferences without having to engage in direct conversation.
[0609] For example, if a user requests a reservation for three people at Restaurant X at 7 PM tomorrow, the server automatically calls Restaurant X based on this information and handles the request. The server then notifies the user that the reservation was successful, allowing the user to directly verify the result. This reduces the stress and inconvenience associated with regular phone calls, enabling efficient and secure communication via telephone.
[0610] The following describes the processing flow.
[0611] Step 1:
[0612] Users launch a dedicated application on their device and enter details such as phone reservations or inquiries. This information includes details such as the store name, reservation date and time, and number of people.
[0613] Step 2:
[0614] The terminal formats the data entered by the user and sends it to the server. The data is structured using a standard format suitable for communication, such as JSON.
[0615] Step 3:
[0616] The server receives data sent from the terminal and parses its contents. This parsing identifies what the request is asking for (for example, a restaurant reservation, a product inquiry, etc.).
[0617] Step 4:
[0618] The server passes the analyzed data to a natural language processing engine, which generates natural conversational sentences based on the user's requests. During this process, the generation AI model prepares appropriate responses based on the context.
[0619] Step 5:
[0620] Based on the generated conversation text, the server initiates a voice call to a specified telephone number via an external communication network. A communication module is then used to perform the actual dialing and establish the connection.
[0621] Step 6:
[0622] Once a call is connected, the server automatically initiates a conversation through AI. During the conversation, it adjusts the content of the discussion in real time based on the other party's responses and gathers the necessary information.
[0623] Step 7:
[0624] Once the call is complete, the server organizes the call content and clarifies the results as requested by the user. These results are presented in a format that is easy for the user to understand.
[0625] Step 8:
[0626] The server sends the organized call results to the terminal. The result data is then reorganized according to the format and sent to the user's terminal.
[0627] Step 9:
[0628] The terminal receives the results from the server and notifies the user. The user can then review the results displayed on the terminal screen and take further action or make decisions.
[0629] (Example 1)
[0630] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0631] Traditionally, making inquiries or reservations by phone has been time-consuming and cumbersome, often accompanied by anxiety and stress during the conversation. Furthermore, automated systems have suffered from a lack of immediacy and accuracy. The challenge lies in improving this situation and enabling users to engage in more comfortable and diverse forms of communication.
[0632] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0633] In this invention, the server includes means for receiving and analyzing information from a user device, means for performing language processing based on the analyzed information, and means for generating dialogue content using the processed information. This enables users to communicate quickly and accurately, such as making inquiries and reservations, automatically, without having to make phone calls themselves.
[0634] A "user device" is a terminal operated by the user, and is a device used for inputting and transmitting data.
[0635] "Means for receiving and analyzing information" refers to a mechanism that has the ability to acquire data transmitted from a user device and understand the content of that data.
[0636] "Means of language processing" are technologies that use analyzed information to determine the meaning and context in natural language and construct appropriate dialogue.
[0637] A "generation method for generating dialogue content" is a system that automatically creates conversations in response to user requests and provides them in a format that can be used for voice calls.
[0638] "External communication media" refers to communication methods used for making voice calls, including the internet and public telephone networks.
[0639] "Means for initiating voice communication" refers to a method that automatically starts a call with a designated party based on the generated dialogue content.
[0640] "Means of organizing communication results" refers to the process of organizing information obtained after a call and summarizing it in a way that is easily understandable to the user.
[0641] "Means of notifying the user" refers to a method of sending organized information to the user's device and informing the user of the results.
[0642] This invention provides an automated voice call system that combines a user, a terminal, and a server. This system begins with the user using a dedicated application on the terminal to input various requests. For example, the user might input, "I would like to make a reservation for four people at Restaurant Y for 7 PM tomorrow."
[0643] The terminal, upon receiving user input, transmits that information to the server. The terminal is a typical mobile device or computer, communicating with the server via the internet. This system utilizes HTTPS as its communication protocol to ensure secure data transmission.
[0644] The server analyzes the received data. This analysis uses a natural language processing engine such as the Google Cloud Natural Language API. Based on the analyzed data, the server uses a generative AI model to generate appropriate conversation content. This generative AI model is designed to accurately and effectively utilize the user's intended meaning in the call.
[0645] Based on the generated conversation content, the server initiates voice communication via an external communication medium. Specifically, the server automatically dials the designated recipient using the Public Public Telephone Network (PSTN) or VoIP technology. Once the call is answered, the conversation proceeds using the automatically generated conversation content.
[0646] After the call ends, the server organizes the information gathered during the call. This organization includes transcribing and summarizing the conversation. The organized results are sent to the user's device, and the device notifies the user of the results. The notification is delivered via a message within a dedicated app or through push notifications on the device.
[0647] This entire process eliminates the need for users to make phone calls themselves, allowing them to enjoy automated and efficient communication. For example, a user could use a prompt like, "I'd like to make a reservation for 5 people at Izakaya Z at 8 PM on the weekend," and confirm the reservation using the same procedure. This allows users to save time and effort while achieving accurate communication.
[0648] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0649] Step 1:
[0650] The user opens a dedicated application on their device and enters a specific request. For example, they might enter something like, "I would like to reserve a table for two at Restaurant W tomorrow night at 8 PM." This entered data is temporarily saved on the device as text.
[0651] Step 2:
[0652] The terminal sends the user's input request to the server. The terminal uses HTTP requests to send text data as packets to the server. In this process, the user request, as input, arrives at the server as text data.
[0653] Step 3:
[0654] The server parses the text data of the received request. Here, a natural language processing engine is used to extract the meaning of the text data and convert it into a structured format. The input to this process is a user request in text format, and the output is structured data.
[0655] Step 4:
[0656] The server uses a generative AI model based on structured data to generate appropriate conversation content. This generation process uses prompt sentences as input, and the generated conversation is stored on the server in text format. The output is a customized conversation based on the user's intent.
[0657] Step 5:
[0658] The server uses the generated conversation content to initiate voice communication with a designated recipient via an external communication medium. The server uses a telephone API to convert the generated conversation content into speech and transmit it to the recipient. The input for this step is the generated conversation content, and the output is a real-time voice call.
[0659] Step 6:
[0660] After the call is completed, the server organizes the results. It documents the information obtained during the call as text and extracts the key points. The input to this process is the audio data of the call result, and the output is the textual result.
[0661] Step 7:
[0662] The server sends the organized call results to the terminal, and the terminal notifies the user of the results. On the terminal, the results are displayed in the notification bar or app interface, allowing the user to easily check the content. The input for this step is the organized call results, and the output is the result notification presented to the user.
[0663] (Application Example 1)
[0664] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0665] Currently, many users experience the inconvenience of having to directly contact service providers when making electronic payments. Furthermore, manual payment procedures are prone to errors and omissions, resulting in wasted time and effort. There is a need to solve these problems and establish a method that allows users to complete payment procedures efficiently and without stress.
[0666] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0667] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, means for generating dialogue content using the processed information, means for initiating voice communication via an external information transmission network, means for organizing the results of the communication and notifying the user of the organized results, and means for automatically contacting the payee specified by the user and completing the payment. This enables the user to complete the payment procedure quickly and accurately through an automated process.
[0668] "Means of receiving and analyzing information from users" refers to a method in which a server receives data sent by a user and performs analysis in order to understand it.
[0669] "Means of natural language processing" refers to technologies used by computers to understand human language and to provide appropriate responses and processing.
[0670] "A means for generating dialogue content" refers to a method of creating a conversation used for communication with the recipient, based on analyzed data.
[0671] "Means of initiating voice communication via an external information transmission network" refers to a method of initiating a voice call and having a conversation with another person through an external communication system.
[0672] "A means of organizing the results of communication and notifying the user of those organized results" refers to a method of summarizing the results obtained during the communication process and conveying them to the user in an easy-to-understand manner.
[0673] "A means of automatically contacting the payment recipient specified by the user and completing the payment" refers to a method of automatically contacting the payment recipient specified by the user and completing the payment process.
[0674] To implement this invention, the server receives payment information transmitted from the user's terminal. The received information is first processed as data by a dedicated analysis engine, and its contents are analyzed in detail. A natural language processing engine (e.g., Google Cloud Natural Language API) is used for the analysis. In this process, the server accurately understands the user's intent and determines the specific steps to be taken next.
[0675] Specifically, based on the results of analysis using natural language processing, a generative AI model (e.g., OpenAI GPT series) generates the dialogue content necessary for the payment process. The generated dialogue content is then used to automatically inquire with the company or service provider specified by the user.
[0676] The communication method involves voice communication initiated via an external information transmission network. Specifically, this system, configured through the user's terminal, engages in real-time conversation based on payment information, obtains necessary information, and completes the payment. After communication is complete, the server organizes the results and notifies the user's terminal of the updated payment status.
[0677] For example, if a user wants to pay their monthly electricity bill to the power company, the user enters "Pay this month's bill to the power company" into the application. The system then automatically contacts the power company and initiates the payment process. An example of a prompt message would be, "Please obtain the procedure to complete this month's payment to the power company." In this way, the present invention provides a system that allows users to complete payments quickly and accurately without having to go through cumbersome procedures.
[0678] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0679] Step 1:
[0680] The user launches the application on their device and enters payment details and amounts. The entered information is then sent to the server as electronic data.
[0681] Step 2:
[0682] The server analyzes the received data. Using a natural language processing engine, it analyzes the user's input from a linguistic perspective and extracts the elements necessary for the payment process. The input is raw data from the user, and the output is a structured analysis result.
[0683] Step 3:
[0684] Based on the analysis results, the server uses a generation AI model to generate payment-related inquiry conversations. It takes a prompt sentence (for example, "Please get instructions on how to complete this month's payment to the utility company") as input and outputs the generated sentence.
[0685] Step 4:
[0686] The server initiates voice communication with the designated payee via an external information transmission network. Using the generated query, it conducts a voice call and retrieves information in real time.
[0687] Step 5:
[0688] After the call ends, the server organizes the acquired information. Specifically, it summarizes information such as payment status and processing status, and constructs data to notify the user. The input is the raw data acquired from the call, and the output is the organized information for notification to the user.
[0689] Step 6:
[0690] The device receives notifications from the server and displays the payment status and results to the user. By checking this information, the user can know that the payment has been completed.
[0691] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0692] This invention relates to a system that identifies a user's emotional state and optimizes the process of making phone calls based on that state. The system is implemented with a configuration including a user terminal, a server, and an emotion engine, and automatically provides appropriate communication.
[0693] Users input the necessary information through a dedicated application on their device. For example, if they wish to make a reservation or inquiry, they input the target location and detailed requests. The device collects this data and uses an emotion engine to determine the user's emotional state in the process. The emotion engine analyzes the user's emotions from the input text and voice, accurately capturing their state.
[0694] The server receives data and emotional state transmitted from the terminal. Based on the analyzed emotional information, the natural language processing engine generates appropriate conversation content. In this process, the tone and expressions are adapted to the user's emotional state. For example, if the user is nervous, a calm tone of voice will be generated.
[0695] Next, the server initiates a voice call to the designated recipient via an external communication network. This call process is dynamically managed by prioritizing based on the user's emotions. If a response is received from the recipient during the conversation, the server updates the generated conversation content in real time and records the response.
[0696] After the call ends, the server compiles the information received and notifies the user. This includes confirmation of the reservation and answers to inquiries. The user can then check the results on their device and take further action as needed.
[0697] For example, if a user is feeling stressed, the system will prioritize the call as an urgent matter and complete it quickly. Conversely, if the user is relaxed, the system will provide a calm and courteous response, ensuring a comfortable calling experience.
[0698] Thus, the present invention provides a means to realize communication that takes user emotions into consideration and to resolve the problems of conventional telephone communication.
[0699] The following describes the processing flow.
[0700] Step 1:
[0701] Users launch a dedicated application on their device and enter their phone reservation or inquiry details. They can also add custom comments or voice messages to express their emotions.
[0702] Step 2:
[0703] The device collects data entered by the user and analyzes the user's emotional state using an emotion engine. It analyzes text and voice data to identify emotional patterns.
[0704] Step 3:
[0705] The device sends the analysis results to the server. This includes data on the user's request, desired action, and emotional state. The data is structured in a standard format.
[0706] Step 4:
[0707] The server analyzes the data received from the terminal and uses a natural language processing engine to generate appropriate responses to requests. Based on sentiment data, it adjusts the tone and expression of the conversation.
[0708] Step 5:
[0709] The server considers the generated conversation content and emotional state, and initiates a voice call to the designated recipient via an external communication network. It is also possible to set call priorities based on emotional state.
[0710] Step 6:
[0711] The server updates the conversation content in real time during the call using a dialogue generation AI. Based on the other party's responses, it retrieves appropriate information as needed and continues the conversation.
[0712] Step 7:
[0713] After the call is completed, the server organizes the information obtained during the call and compiles the data in a format that aligns with the user's request.
[0714] Step 8:
[0715] The server sends the processed results to the user's terminal. This allows the user to check the call results and responses.
[0716] Step 9:
[0717] The device notifies the user of the received results and displays the details on the screen. The user can then decide on their next action based on the results.
[0718] (Example 2)
[0719] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0720] Traditional communication technologies often suffer from a decline in communication quality because they do not take into account the user's emotional state when responding or prioritizing calls. In particular, when users are experiencing stress or anxiety, appropriate responses and prompt communication are often lacking. This highlights the need for improved user experience.
[0721] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0722] In this invention, the server includes means for receiving and collecting user input information, means for analyzing the user's emotional state from the collected information, and means for initiating voice communication via an external communication network using the generated conversation content. This makes it possible to set appropriate responses and priorities according to the user's emotional state.
[0723] "User input information" refers to data such as reservation and inquiry details, as well as any additional conditions, that users submit to the system.
[0724] "Emotional state" refers to the user's psychological state and reactions, and is identified through the analysis of text and voice.
[0725] "Natural language processing" is a technology that enables computers to understand and generate human language, and it is a method for creating conversational content based on analyzed data.
[0726] "Conversation content" refers to the voice or text-based responses generated by the system, which are adjusted to suit the user's emotional state.
[0727] "External communication networks" refer to communication infrastructure such as the internet and telephone lines that servers use to communicate.
[0728] "Call priority" is an indicator that shows the importance of a call based on the user's emotional state, and is dynamically set according to urgency.
[0729] "Response content" refers to real-time response information from the other party, and is data acquired during the course of the conversation.
[0730] The following describes embodiments for carrying out the present invention.
[0731] The system of the present invention mainly consists of a user terminal, a communication server, and an emotion analysis engine. The user uses a dedicated application on the terminal to input necessary information regarding reservations and inquiries. This application includes hardware (e.g., microphone, keyboard) to acquire information entered in the form of text or voice, and software (e.g., speech recognition software, text analysis engine) to process the data.
[0732] The device uses an emotion analysis engine to determine the user's emotional state from the input information. The emotion analysis engine combines natural language processing and speech analysis technologies, and can capture emotions by analyzing, for example, the tone and speed of voice, and the emotional vocabulary in text.
[0733] The analyzed emotion information and user input data are sent to the server. The server uses a natural language processing engine to generate conversation content adapted to the user's emotions. This engine can create multiple response patterns using a generative AI model and select the optimal one. The generated conversation content is transmitted to the designated recipient via an external communication network, and a voice call is initiated.
[0734] During a call, the server updates the conversation in real time based on the responses received. Furthermore, call priorities are dynamically managed based on the user's emotional state, ensuring that urgent calls are processed without delay.
[0735] For example, if a user enters the text "I want to make a dentist appointment for next Tuesday," the device's emotion analysis engine detects that the user is nervous. Based on this information, the server prepares a response in a relaxed tone and prioritizes initiating the call with the dentist.
[0736] An example of a prompt message is, "Generate a call message to confirm a reservation while the user is feeling stressed." This prompt allows the system to generate the most appropriate conversation for the situation. This prompt serves as input for the generation AI model to function correctly.
[0737] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0738] Step 1:
[0739] Users enter reservation and inquiry information using a dedicated application on their terminal. This input data includes information in text or audio format. The terminal converts this information into digital signals and applies them as input to an emotion analysis engine. This engine analyzes the user's emotions from the tone of voice and the context of the text, and outputs their emotional state.
[0740] Step 2:
[0741] The terminal sends the emotional state obtained by the emotion analysis engine and the information entered by the user to the server. The server activates a natural language processing engine based on the received data. This engine processes the conversation data using a generative AI model to generate optimal conversation content that corresponds to the user's emotions. The generated conversation content is output and used for the next processing step.
[0742] Step 3:
[0743] The server uses the generated conversation content to initiate a voice call to the designated recipient via an external communication network. During this process, priority is initially set based on sentiment information, with urgent calls being processed first. The server establishes an external connection and prepares to begin the conversation.
[0744] Step 4:
[0745] During a call, the server updates the conversation content in real time based on information received from the other party. The input is response data from the other party, which the server analyzes and outputs newly generated response content using a generative AI model. The response is applied immediately, maintaining the quality of the call.
[0746] Step 5:
[0747] After the call ends, the server processes the call results and notifies the user. The acquired information is stored in a database and displayed to help the user easily decide on the next steps. As a final output, the user is provided with a reservation confirmation and answers to their inquiries.
[0748] (Application Example 2)
[0749] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0750] In modern information and communication technology, users can experience anxiety and stress in various situations. These emotional states can lead to inefficiencies and problems in communication. However, existing systems lack sufficient means to accurately monitor users' emotional states and provide support at the appropriate time. Therefore, there is a need to support users' mental well-being and facilitate smooth communication.
[0751] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0752] In this invention, the server includes means for receiving and analyzing information from the user, means for performing natural language processing based on the analyzed information, and means for monitoring the user's emotional state and providing prompt support when an anomaly is detected. This makes it possible to monitor the user's emotional state in real time, provide appropriate language responses as needed, and provide rapid support in emergencies.
[0753] "Means of receiving and analyzing information from users" refers to functions that acquire data provided by users, analyze that data, and understand the user's state and requests.
[0754] "Means of performing natural language processing" refers to a function that uses acquired user information to enable a computer to understand the context and derive an appropriate response.
[0755] "Generating means for generating response content" refers to a function that creates appropriate responses or messages for the user based on the results of analysis using natural language processing.
[0756] "Means of initiating a voice call via an external communication network" refers to a function that uses a communication network to initiate a voice-based conversation with a designated party.
[0757] "A means of organizing call results and notifying the user of those organized results" refers to a function that summarizes the content and results of a voice call and reports or provides feedback to the user.
[0758] "A means of monitoring the user's emotional state and providing rapid support when an anomaly is detected" refers to a function that monitors the user's emotions in real time and provides immediate and appropriate support when an anomaly such as stress or panic is observed.
[0759] The system for realizing this invention is designed to monitor the user's emotional state in real time and provide appropriate responses. Its main components include a terminal used by the user, a server containing an emotion engine for analyzing the user's emotions, and a communication network for voice communication.
[0760] The user accesses the system through a terminal and enters the necessary information. This terminal is an internet-connected device such as a smartphone or wearable device. User information is collected by the terminal and sent to the server via the emotion engine. The emotion engine analyzes the user's emotional state from voice and text messages, using, for example, Python or an emotion analysis library (such as Hugging Face's Transformers). Based on this analysis, a natural language processing engine generates a response appropriate to the user.
[0761] Based on the generated response, the server initiates a voice call with the specified party using an external communication network. During this process, the server can utilize a communication API (e.g., Twilio). Once the call ends, the server can organize the call content and notify the user of the results.
[0762] As a concrete example, when a user is feeling stressed, the device reports its emotional state to the server, which then uses a generative AI model to generate a calming response. For instance, the following prompt is entered into the generative AI model: "Analyze the user's voice input and generate a response message if a stressed state is detected. Generate a message in a calming, gentle tone saying, 'Take a deep breath and relax.'"
[0763] In this way, the system can provide quick and appropriate communication tailored to the user's emotional state, thereby improving the user's sense of security and satisfaction.
[0764] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0765] Step 1:
[0766] Users enter the necessary information for reservations and inquiries via their device. This includes specific requirements such as the planned date and time of visit and the services they wish to receive. The entered data is temporarily stored on the device.
[0767] Step 2:
[0768] The device transmits the collected user information to the emotion engine. It receives voice and text messages as input and analyzes the user's emotional state using an emotion analysis library. The output is an emotional state (e.g., stress, relaxation).
[0769] Step 3:
[0770] The server receives the analyzed emotional state and user input information. The server uses a natural language processing engine to generate an appropriate response based on the input data. As part of the data processing, the tone and expression are adjusted according to the emotional state. The output is a specific response message.
[0771] Step 4:
[0772] The server sends the generated response content via an external communication network and initiates a voice call with the specified recipient. A communication API is used to establish the voice call connection. Voice communication takes place between the user and the recipient.
[0773] Step 5:
[0774] After the call ends, the server organizes the call content and creates a summary. This includes recording responses from the other party and processing the data to extract important information. The summary results are then notified to the user.
[0775] Step 6:
[0776] Users receive notifications on their devices and can choose the next action as needed. They can also check if they received satisfactory support and request further information if necessary.
[0777] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0778] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0779] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0780] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0781] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0782] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0783] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0784] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0785] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0786] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0787] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0788] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0789] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0790] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0791] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0792] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0793] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0794] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0795] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0796] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0797] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0798] The following is further disclosed regarding the embodiments described above.
[0799] (Claim 1)
[0800] A means of receiving and analyzing data from users,
[0801] A means of performing natural language processing based on analyzed data,
[0802] A generation means that generates conversation content using processed data,
[0803] A means of initiating a voice call via an external communication network,
[0804] A means of organizing the results of a call and notifying the user of those organized results,
[0805] A system that includes this.
[0806] (Claim 2)
[0807] The system according to claim 1, comprising means for generating context-appropriate responses when analyzing using natural language processing.
[0808] (Claim 3)
[0809] The system according to claim 1, further comprising means for updating the generated conversation content in real time based on external responses and for recording the response content.
[0810] "Example 1"
[0811] (Claim 1)
[0812] A means for receiving and analyzing information from a user device,
[0813] A means for performing language processing based on the analyzed information,
[0814] A generation means for generating dialogue content using processed information,
[0815] A means of initiating voice communication via an external communication medium,
[0816] A means of organizing the results of the communication and notifying the user of those organized results,
[0817] A device that includes this.
[0818] (Claim 2)
[0819] The apparatus according to claim 1, comprising means for generating an appropriate response based on the background when performing analysis using language processing.
[0820] (Claim 3)
[0821] The apparatus according to claim 1, further comprising means for immediately updating the generated dialogue content based on an external response and recording the response content.
[0822] "Application Example 1"
[0823] (Claim 1)
[0824] A means of receiving and analyzing information from users,
[0825] A means of performing natural language processing based on the analyzed information,
[0826] A generation means for generating dialogue content using processed information,
[0827] A means of initiating voice communication via an external information transmission network,
[0828] A means of organizing the results of the communication and notifying the user of those organized results,
[0829] A means of automatically contacting the payment recipient specified by the user and completing the payment,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, comprising means for generating context-appropriate responses when analyzing using natural language processing.
[0833] (Claim 3)
[0834] The system according to claim 1, further comprising means for immediately updating the generated dialogue content based on an external response and recording the response content.
[0835] "Example 2 of combining an emotion engine"
[0836] (Claim 1)
[0837] A means of receiving and collecting user input information,
[0838] A means of analyzing the emotional state of users from the collected information,
[0839] A means for generating conversation content by performing natural language processing based on analyzed emotional states,
[0840] A means of initiating voice communication via an external communication network using the generated conversation content,
[0841] A means of dynamically setting call priority based on emotional state,
[0842] A means of updating and recording the content of a call in real time,
[0843] A means of summarizing the results of a call and notifying the user,
[0844] A system that includes this.
[0845] (Claim 2)
[0846] The system according to claim 1, comprising means for adjusting the tone and expression of conversation according to the emotional state of the user.
[0847] (Claim 3)
[0848] The system according to claim 1, comprising means for automatically managing the urgency of a call based on the emotional state.
[0849] "Application example 2 when combining with an emotional engine"
[0850] (Claim 1)
[0851] A means of receiving and analyzing information from users,
[0852] A means of performing natural language processing based on the analyzed information,
[0853] A generation means that generates response content using processed information,
[0854] A means of initiating a voice call via an external communication network,
[0855] A means of organizing the call results and notifying the user of those organized results,
[0856] A means of monitoring the user's emotional state and providing rapid support when an anomaly is detected,
[0857] A system that includes this.
[0858] (Claim 2)
[0859] The system according to claim 1, comprising means for generating context-appropriate responses when performing analysis using natural language processing.
[0860] (Claim 3)
[0861] The system according to claim 1, further comprising means for updating the generated response content in real time based on an external response and for recording the response. [Explanation of Symbols]
[0862] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving and analyzing data from users, A means of performing natural language processing based on analyzed data, A generation means that generates conversation content using processed data, A means of initiating a voice call via an external communication network, A means of organizing the results of a call and notifying the user of those organized results, A system that includes this.
2. The system according to claim 1, comprising means for generating context-appropriate responses when performing analysis using natural language processing.
3. The system according to claim 1, further comprising means for updating the generated conversation content in real time based on external responses and for recording the response content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A