system

A system using a user terminal, server, and voice AI automates tasks like subscription cancellations and reservation changes, addressing the inefficiencies of conventional voice call methods by enabling efficient and stress-free task completion.

JP2026038170APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Conventional tasks such as canceling subscriptions or changing restaurant reservations often require users to make voice calls during limited hours, which is inconvenient and inefficient, especially for busy individuals.

Method used

A system utilizing a user terminal, server, and voice AI to autonomously execute tasks via voice calls, enabling users to input task details, analyze requests, and generate dialogue scripts for voice AI to complete tasks efficiently and accurately.

Benefits of technology

Allows users to complete tasks like subscription cancellations and reservation changes without personal involvement, reducing stress and time consumption by automating operations through natural language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026038170000001_ABST
    Figure 2026038170000001_ABST
Patent Text Reader

Abstract

Provide a system. A means for a user to input the content of a task and necessary information; means for the server to receive and analyze requests from said users; A means for voice AI to autonomously perform tasks and make calls; means for the server to record the result of the call and notify the user; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional tasks such as canceling subscriptions or changing restaurant reservations can often only be completed via voice call, and the limited hours of operation make these tasks extremely inconvenient for users who are busy with work or daily life. Furthermore, these tasks require users to set aside time to complete them, which is inefficient. The purpose of this invention is to eliminate these inconveniences and allow users to complete these tasks autonomously and efficiently. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides the following means: A system including a means for a user to input the content of a task and necessary information, a means for a server to receive and analyze the request from the user, a means for a voice AI to autonomously execute the task and make a call, and a means for the server to record the results of the call and notify the user. Furthermore, the voice AI includes a means for conducting a dialogue using natural language processing technology, and the server includes a means for generating a dialogue script for the voice AI according to the task content, thereby enabling tasks to be completed efficiently and accurately.

[0006] "User" refers to the end user who uses the system to input task details and operate the system.

[0007] A "task" is a specific action or transaction performed through the system, such as canceling a subscription or changing a restaurant reservation.

[0008] "Information" refers to specific data required by a user to perform a task, including, for example, a contract number or reservation number.

[0009] "Terminal" refers to the device a user uses to input tasks or send requests, such as a smartphone or computer.

[0010] "Server" refers to a computer system with a central function that receives and analyzes user requests and manages information necessary to complete tasks.

[0011] A "request" refers to instructions or information about a task that a user sends to a server.

[0012] "Voice AI" refers to artificial intelligence that makes voice calls on behalf of users and performs tasks autonomously.

[0013] "Natural language processing technology" refers to the technology that enables voice AI to understand human language and engage in dialogue.

[0014] A "dialogue script" refers to a set of instructions that define the series of questions and answers that a voice AI will use during a call.

[0015] "Call results" refers to the content and results of the call made by the voice AI.

[0016] "Notification" refers to a message sent by a server to inform a user of the progress or completion status of a task. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by its use of voice AI to automate operations via voice calls. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0039] System configuration

[0040] 1. User Device

[0041] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[0042] 2. Server

[0043] The server receives and analyzes requests from users, retrieves necessary information from a database, and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user.

[0044] 3. Voice AI

[0045] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[0046] What the program does

[0047] User submits a request

[0048] The user uses a dedicated application or web interface to input the details of the task (e.g., canceling a subscription or changing a reservation). For example, if the user wants to cancel a subscription, they enter the contract number and other necessary information and press the submit button. The device then sends this request to the server.

[0049] The server parses the request

[0050] The server analyzes the received request and understands the task. It checks the necessary information (e.g., subscription contract information and reservation number) and prepares the voice AI. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[0051] Voice AI performs tasks

[0052] The voice AI connected to the server dials a real phone number on behalf of the user. After dialing the specified phone number and receiving the answer, the voice AI uses natural language processing technology to carry out the dialogue. For example, the voice AI might say, "Hello, I would like to cancel my subscription for contract number 123456." As the dialogue progresses, the AI ​​provides necessary information to complete the task.

[0053] Call outcome feedback

[0054] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification such as "Subscription cancellation completed" will be sent to the user's device. This allows the user to complete the necessary task without having to spend time on their own.

[0055] Specific examples

[0056] Example 1: Canceling a subscription

[0057] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI calls customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[0058] Example 2: Changing a restaurant reservation

[0059] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI calls the restaurant and says, "Please change reservation number 789012." After confirming the change, the server notifies the user that "the reservation change has been completed."

[0060] In this way, this system uses voice AI to efficiently perform tasks on behalf of the user and provide feedback on the results.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[0064] Step 2:

[0065] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[0066] Step 3:

[0067] The device sends a user request to the server, including the type of task and the required details.

[0068] Step 4:

[0069] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[0070] Step 5:

[0071] The server analyzes the request and identifies the type of task (e.g., canceling a subscription). It then retrieves the information needed to execute the task (such as contract and reservation information) from the database.

[0072] Step 6:

[0073] The server calls the voice AI module and prepares to execute the task. It generates a dialogue script according to the task. For example, if it is a cancellation procedure, it prepares a script that says, "I would like to cancel contract number 123456."

[0074] Step 7:

[0075] The voice AI automatically dials the specified phone number, waits for the call to connect, and starts the conversation when the other party answers.

[0076] Step 8:

[0077] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[0078] Step 9:

[0079] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[0080] Step 10:

[0081] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[0082] Step 11:

[0083] The server analyzes the results of the call and notifies the user of the outcome, including information about task completion and next steps.

[0084] Step 12:

[0085] The user receives a notification on their device confirming that the task was completed as expected, for example, a message saying "Your subscription has been successfully canceled."

[0086] This process allows users to complete tasks efficiently without spending time on their own.

[0087] Example 1

[0088] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0089] In today's society, tasks users perform over the phone (e.g., canceling subscriptions or changing reservations) are time-consuming and cumbersome, necessitating efficient solutions. Furthermore, negotiating directly over the phone can be stressful for users, so technology is needed to simplify and automate these processes. Furthermore, conventional voice recognition systems struggle to deliver natural dialogue and appropriate responses, so improvements are needed.

[0090] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0091] In this invention, the server includes a means for receiving and analyzing requests from users, a means for calling a voice AI and sending instructions including necessary information, and a means for recording the results of the call and notifying the user. This allows the user to efficiently automate tasks without having to make a call, reducing stress and enabling appropriate responses.

[0092] "User" means any person or entity that intends to use the System to perform a specific task.

[0093] A "task" is a specific purpose or action a user attempts to perform through the system (e.g., canceling a subscription or changing a reservation).

[0094] "Information" refers to data required to perform a task (e.g., contract number or reservation number).

[0095] "Means" refers to a method or combination of methods for achieving a specific function or operation.

[0096] "Server" refers to the central processing unit that receives and analyzes requests from users and controls the voice AI.

[0097] A "request" refers to a communication that describes the task or request sent by a user to a system.

[0098] "Analysis" refers to the operation by the server to interpret the contents of the request and determine the necessary processing.

[0099] "Voice AI" refers to artificial intelligence that makes voice calls on behalf of users and uses natural language processing technology to guide the conversation.

[0100] "Calling" refers to the operation in which the server sends instructions to the voice AI to perform a specific task.

[0101] "Call" refers to the voice communication that the voice AI has with customer support or other parties.

[0102] "Results" refers to information showing the status and results after the voice AI has performed a task.

[0103] "Notification" refers to communication from the server to the user informing them of task completion and progress.

[0104] "Natural language processing technology" refers to technology for understanding and generating human language.

[0105] A "dialogue script" refers to a scenario of a series of responses and questions that a voice AI uses during a call.

[0106] This invention is a system that allows users to efficiently perform specific tasks, and in particular automates operations via voice calls using voice AI. The main elements of this system are the user device, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0107] User Device

[0108] A user terminal is a device such as a smartphone or computer. The user interacts with the system through a dedicated application or a web interface. Using this interface, the user enters the details of the task (e.g., canceling a subscription, changing a reservation) and provides the required information. For example, if the user wants to cancel a subscription, they enter their contract number and other required information and press the submit button. The terminal then sends this request to the server.

[0109] server

[0110] The server receives and analyzes the request sent from the user's device. This analysis allows it to understand the task content and retrieve the necessary information from the database. For example, the server may determine that the user wishes to cancel a subscription and gather the necessary data for cancellation. The server also issues instructions to the voice AI to execute the task. It passes the necessary information (e.g., contract number 123456) to the voice AI and instructs it to execute a specific task.

[0111] voice AI

[0112] The voice AI calls a designated phone number on behalf of the user, and after the caller answers, it uses natural language processing technology to continue the conversation. For example, when a customer support operator answers, they say, "Hello, I'd like to cancel my subscription for contract number 123456," and then proceeds with the necessary procedures. They ask questions and provide information as the conversation progresses.

[0113] Call outcome feedback

[0114] After the task is completed, the voice AI sends the call results to the server. The server analyzes the call results and records the necessary information (e.g., cancellation completed, reservation change completed). The server then notifies the user's device that the task has been completed. For example, the user's device may receive a notification saying, "Subscription cancellation has been completed."

[0115] Specific examples

[0116] Example 1: Canceling a subscription

[0117] 1. The user opens the application and selects "Cancel Subscription."

[0118] 2. Enter the contract number "123456" and other required information and submit.

[0119] 3. The server analyzes the request and the voice AI calls customer support.

[0120] 4. The voice AI says, "Please cancel contract number 123456," and continues the conversation.

[0121] 5. Once the cancellation procedure is complete, the server will notify the user that "Cancellation has been completed."

[0122] Example 2: Changing a restaurant reservation

[0123] 1. The user selects "Change Restaurant Reservation" in the application.

[0124] 2. Enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit.

[0125] 3. The server analyzes the request and the voice AI calls the restaurant.

[0126] 4. The voice AI says, "Please change reservation number 789012."

[0127] 5. After confirming the change, the server notifies the user that the reservation change has been completed.

[0128] Prompt Sentence Examples

[0129] Example 1: Subscription cancellation prompt

[0130] Prompt: Please cancel your subscription. Contract number is 123456.

[0131] Example 2: Prompt for changing a restaurant reservation

[0132] Prompt: I would like to change my restaurant reservation. The reservation number is 789012, and the desired date and time is October 10th at 7:00 PM.

[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0134] Step 1:

[0135] The user enters the task

[0136] The user opens a dedicated application or web interface on their smartphone or computer and enters the details of the task. For example, if they want to cancel a subscription, they enter the contract number "123456" and add any other necessary information. When the user presses the "Submit" button, the device collects this input data and sends it to the server.

[0137] Input: Task details and contract number entered by the user

[0138] Output: Request data sent to the server (detailed task information)

[0139] Specific behavior:

[0140] Open the application.

[0141] Select "Cancel Subscription."

[0142] Enter the contract number "123456".

[0143] Enter other necessary information.

[0144] Press the send button.

[0145] Step 2:

[0146] The server receives and analyzes the request

[0147] The server receives the request sent from the device, analyzes its contents, identifies the type of task included in the request, and sends a query to the database to extract the necessary information (e.g., contract number), thereby obtaining the data necessary for the cancellation procedure.

[0148] Input: Request data sent from the device (task details)

[0149] Output: Analyzed task type and extracted required information (e.g. contract number)

[0150] Specific behavior:

[0151] Received request.

[0152] Analyze the request content.

[0153] Identify the type of task.

[0154] Obtain necessary information (contract number, etc.) from the database.

[0155] Step 3:

[0156] The server calls the voice AI

[0157] Based on the analysis results, the server calls the voice AI to execute the appropriate task. The server passes the necessary information (e.g., contract number) to the voice AI and sends an instruction to execute. For example, the server issues an instruction such as "Contact customer support to cancel the subscription for contract number 123456."

[0158] Input: Analysis results and necessary information (contract number, etc.)

[0159] Output: Instructions and task information passed to the voice AI

[0160] Specific behavior:

[0161] Call up the voice AI.

[0162] Pass the necessary information to the voice AI.

[0163] Send instructions to perform tasks.

[0164] Step 4:

[0165] Voice AI initiates and manages calls

[0166] The voice AI will call the specified phone number and start the call. Once the call begins, the voice AI will use natural language processing technology to proceed with the conversation. For example, when a customer support operator answers, they will say, "Hello, I would like to cancel my subscription for contract number 123456," and proceed with the process.

[0167] Input: Execution instructions and task information passed from the server

[0168] Output: Dialogue outcome and progress during the call

[0169] Specific behavior:

[0170] Dial the specified phone number.

[0171] Customer support responds.

[0172] Say, "Hello, I would like to cancel my subscription for contract number 123456."

[0173] Ask questions and provide information as necessary depending on the procedure.

[0174] Step 5:

[0175] The server records the results and sends feedback

[0176] After the voice AI completes the task, it sends the results of the call to the server. The server analyzes the call results and records the necessary information. Finally, the server notifies the user device that the task has been completed. For example, it may send a notification to the user saying, "Your subscription has been canceled."

[0177] Input: Call result sent from voice AI

[0178] Output: Parsed results and user notification

[0179] Specific behavior:

[0180] Receive call results from voice AI.

[0181] Analyze the results and record the necessary information.

[0182] Sends notifications to the user's device.

[0183] A message will be sent stating that your subscription has been cancelled.

[0184] (Application example 1)

[0185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0186] Online shopping and food delivery services are rapidly becoming more popular, but there is a problem in that it takes time and effort for users to make changes to their orders or make inquiries. In particular, inquiries and changes made over the phone are prone to waiting times and communication problems. It also increases the burden on customer support, making it difficult to provide efficient service. To solve these issues, there is a need for a system that allows users to easily request changes to their orders and automates the entire process.

[0187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0188] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the server to receive and analyze the request from the user; a means for the voice AI to autonomously execute the task and make a call; a means for the server to record the result of the call and notify the user; a means for the user to request an order change and specify the new order content; and a means for the server to contact the ordering party via the voice AI and request the order change. This allows the user to easily request an order change, and the entire process is automated, eliminating problems such as waiting times and communication errors, enabling efficient service provision.

[0189] "Means for users to input task details and necessary information" refers to an interface through which users input information necessary to perform a specific task, such as a smartphone application or a web interface.

[0190] "Means for the server to receive and analyze a request from the user" refers to a function that enables the server to receive request data sent from the user, analyze it, and identify the necessary processing.

[0191] "Means for voice AI to autonomously execute tasks and make calls" refers to a function that enables voice AI to autonomously execute designated tasks on behalf of the user via telephone.

[0192] "Means for the server to record the results of the call and notify the user" refers to a function that enables the server to record the results of the call made by voice AI in a database and notify the user of the results.

[0193] The "means by which a user requests an order change and specifies the new order details" is an interface by which a user can change the current order details and specify the new order details after the changes.

[0194] "Means for the server to contact the ordering party via voice AI and request a change to the order" is a function that allows the server to use voice AI to call the designated ordering party and request a change to the user's order.

[0195] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by the use of voice AI to automate food delivery order changes and inquiries. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0196] System configuration

[0197] 1. User Device

[0198] A user terminal is a device, such as a smartphone or computer, that a user uses to enter change order details and submit requests. A dedicated application or web interface is provided through which the user interacts with the system.

[0199] 2. Server

[0200] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user. This process uses software such as natural language processing (NLP) libraries (e.g., Google® Cloud Text-to-Speech, Amazon Lex) and communication libraries (e.g., Twilio).

[0201] 3. Voice AI

[0202] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and engage in natural conversations with the other party (e.g., restaurant staff). Voice AI calls use NLP technologies such as Google Cloud Text-to-Speech and Amazon Lex.

[0203] What the program does

[0204] User submits a request

[0205] The user uses a dedicated application or web interface to input the details of the order change (e.g., change to two pizzas). For example, if the user wants to "change the order," they enter the current order number and the new order details and press the submit button. The device then sends this request to the server.

[0206] The server parses the request

[0207] The server analyzes the received request and understands the task. It checks the necessary information (e.g., the current order number and the new order details) and prepares the voice AI. For example, the server determines that the user wants to change the order and prepares the necessary data for the change.

[0208] Voice AI performs tasks

[0209] The voice AI connected to the server dials a real phone number on behalf of the user. It dials the specified restaurant's phone number, and after the caller answers, the voice AI uses natural language processing technology to carry out the conversation. For example, the voice AI might say, "Hello, I'd like to change the order for order number 123456. The new order is for two pizzas." As the conversation progresses, the AI ​​provides necessary information to complete the task.

[0210] Call outcome feedback

[0211] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification saying "Order change completed" will be sent to the user's device. This allows the user to complete the necessary tasks without having to spend time on their own.

[0212] Specific examples

[0213] Example 1: Order change procedure

[0214] The user opens the application and selects "Change Order." They enter the current order number "123456" and the new order details "2 pizzas," and submit. The server analyzes the request, and the voice AI calls the restaurant. The voice AI responds by saying, "Please change order number 123456 to 2 pizzas," and continues the conversation. Once the change procedure is complete, the server notifies the user, "The order change has been completed."

[0215] Prompt Sentence Examples

[0216] "Imagine a food delivery application where a user requests a change to their order. When the user inputs that they would like to change the order for order number 123456 to two pizzas, write a program in which the voice AI automatically calls the restaurant and requests the order change. If the call is successful, please also include a function to notify the user that "the order change has been completed."

[0217] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0218] Step 1:

[0219] The user opens the smartphone application and accesses the screen for requesting an order change. The user enters the current order number and new order details (e.g., two pizzas) and presses the "Submit" button. The data entered is the order number and new order details. The device sends this request to the server as JSON-formatted data.

[0220] Step 2:

[0221] The server receives the request sent from the terminal. The received data includes the order number and new order details. The server analyzes the data and recognizes that the user wants to change the order. The server first retrieves the relevant order information from the database and locally checks the updated status of the order data based on the new order details. The output is the necessary information for preparing the voice AI.

[0222] Step 3:

[0223] The server provides the voice AI with details of the order change and contact information for the person on the other end of the call (e.g., restaurant staff). The input is the detailed order change information obtained in the previous step. The server uses a generative AI model (e.g., Google Cloud Text-to-Speech or Amazon Lex) to generate a dialogue script for the voice AI. The script includes a greeting at the start of the call, providing the order number, and explaining the new order details. The generated dialogue script is output.

[0224] Step 4:

[0225] The voice AI automatically calls the specified restaurant's phone number based on the dialogue script provided by the server. The input is the dialogue script and phone number. Natural language processing (NLP) is used to analyze and generate dialogue content, and ask the restaurant staff to change the order. Specifically, it says, "Hello, I'd like to change order number 123456. The new order is for two pizzas." The necessary information is presented appropriately during the call, and the dialogue continues until the order change is completed. The output is the call result.

[0226] Step 5:

[0227] Once the voice AI completes the call, it sends the results back to the server, including confirmation of the call's success or failure and any specific changes that were made. The input is the call result data. The server uses this information to update the database and record the changes. The output is the updated order information.

[0228] Step 6:

[0229] The server creates a notification for the user based on the call result. The input is the updated order information. It generates a message such as "Order changes completed" and sends it to the user's device. The output is a notification message.

[0230] Step 7:

[0231] The user terminal receives the notification message from the server and displays it on the screen. The user confirms the message "Order change completed." Through this process, the user can complete the order change without any hassle.

[0232] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0233] This invention is a system that allows users to efficiently perform specific tasks. It is characterized by utilizing voice AI to automate operations via voice calls, and by combining it with an emotion engine, it recognizes the user's emotions and adaptively changes the response. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine. Each element works together to autonomously process the user's tasks.

[0234] System configuration

[0235] 1. User Device

[0236] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[0237] 2. Server

[0238] The server receives and analyzes user requests, retrieves necessary information from a database, and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user.

[0239] 3. Voice AI

[0240] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[0241] 4. Emotion Engine

[0242] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue, enabling optimal dialogue according to the user's emotional state. It also learns from the user's past emotional data to improve the quality of future dialogue.

[0243] What the program does

[0244] User submits a request

[0245] The user enters the details of the task using a dedicated application or web interface. For example, if they want to cancel a subscription, they enter their contract number and other necessary information and press the submit button. The device then sends this request to the server.

[0246] The server parses the request

[0247] The server analyzes the received request and understands the task. It checks the necessary information (e.g., contract information and reservation number) and prepares the voice AI and emotion engine. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[0248] Voice AI and emotion engines perform tasks

[0249] The voice AI and emotion engine connected to the server work together to dial the phone on behalf of the user. After dialing the specified phone number and the caller answers, the voice AI uses natural language processing technology to carry out the dialogue. The emotion engine recognizes emotions from the user's voice and changes the dialogue content as necessary. For example, when the voice AI says, "Hello, I'd like to cancel my subscription for contract number 123456," if the emotion engine recognizes the user's anxiety or irritation, the voice AI adds a phrase that provides reassurance, such as, "Don't worry about that. We'll deal with it right away."

[0250] Call outcome feedback

[0251] After the task is completed, the server records the call result and sends a notification to the user. For example, a notification saying "Your subscription has been canceled" will be sent to the user's device. The emotional data collected by the emotion engine can be used to improve the quality of future conversations.

[0252] Specific examples

[0253] Example 1: Canceling a subscription

[0254] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[0255] Example 2: Changing a restaurant reservation

[0256] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server notifies the user, "The reservation change has been completed."

[0257] In this way, by using voice AI and an emotion engine, this system can provide optimal dialogue based on the user's emotions and complete tasks efficiently.

[0258] The processing flow will be explained below.

[0259] Step 1:

[0260] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[0261] Step 2:

[0262] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[0263] Step 3:

[0264] The device sends a user request to the server, including the type of task and the required details.

[0265] Step 4:

[0266] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[0267] Step 5:

[0268] The server analyzes the request, recognizes what the task is (e.g. cancel a subscription), retrieves the necessary information from the database, and issues instructions to the voice AI and emotion engine.

[0269] Step 6:

[0270] The server generates a dialogue script for the voice AI to execute a task and prepares the emotion engine. For example, the server prepares a script that says, "I would like to cancel contract number 123456."

[0271] Step 7:

[0272] The voice AI will dial the specified phone number, wait for the call to connect, and begin the conversation once the other party answers.

[0273] Step 8:

[0274] The emotion engine analyzes the user's voice to determine their emotions and provides that information to the voice AI, which then adapts the dialogue accordingly. For example, if the user sounds anxious, the AI ​​might add a phrase like, "Don't worry, we'll get back to you shortly."

[0275] Step 9:

[0276] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[0277] Step 10:

[0278] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[0279] Step 11:

[0280] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[0281] Step 12:

[0282] The server analyzes the results of the call and notifies the user of the outcome. The user receives a notification on their device confirming that the task was completed as planned, for example, a message saying "Subscription cancellation successful."

[0283] Step 13:

[0284] The emotion engine learns from the emotional data collected during processing and uses it to improve the quality of future interactions, resulting in a continuously improved user experience.

[0285] Example 2

[0286] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0287] Conventional voice communication systems require users to perform detailed operations themselves, making it difficult to complete tasks efficiently. Furthermore, they only provide a uniform response without considering the user's feelings, resulting in low user satisfaction. This has led to a demand for a system that can reduce the user's time and effort while also providing optimal responses tailored to individual feelings.

[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0289] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the terminal to send the input information to the server; a means for the server to receive and analyze the request from the user; a means for the server to acquire data based on the received and analyzed information and issue preparation instructions to the voice AI and emotion engine; a means for the voice AI to autonomously execute the task and make a call; a means for the emotion engine to recognize emotions from the user's voice; and a means for the server to record the results of the call and notify the user. This allows the user to complete the task efficiently without detailed operations, and also enables optimal responses according to the user's emotions.

[0290] A "user" is someone who utilizes the system to perform a task.

[0291] A "terminal" is a device that a user uses to interact with the system, such as a smartphone or computer.

[0292] A "server" is a computer system that receives and analyzes requests from users.

[0293] "Voice AI" refers to artificial intelligence technology that generates voice and performs tasks autonomously.

[0294] The "emotion engine" is a technology that recognizes emotions from the user's voice and adaptively changes the content of the dialogue according to those emotions.

[0295] "Request" refers to a request including the content of a task that a user sends to a server via a terminal.

[0296] "Natural language processing technology" is a technology that allows computers to understand and generate human language.

[0297] A "call" is a means of communication through voice, and specifically refers to an exchange using a telephone.

[0298] "Data acquisition" refers to the process by which the server gathers the necessary information from internal databases and external sources.

[0299] "Notification" refers to a message sent by the server to inform the user of the outcome of a task.

[0300] This invention is a system that enables users to efficiently perform specific tasks, automating operations via voice calls by combining primarily voice AI and an emotion engine, and providing optimal responses according to the user's emotions. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine.

[0301] User Device

[0302] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system. For example, if a user wants to cancel a subscription, they open the dedicated application, enter their contract number and other required information, and press the submit button.

[0303] server

[0304] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user. The server analyzes the requests it receives, parses the request data, and extracts necessary information to understand the content of the task.

[0305] voice AI

[0306] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have a natural conversation with the other party (e.g., a customer support operator). For example, when the voice AI says, "Hello, I would like to cancel my subscription for contract number 123456," it uses natural language processing technology to smoothly guide the conversation.

[0307] Emotion Engine

[0308] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue. As a result, it is possible to have an optimal dialogue according to the user's emotional state. If the emotion engine detects the user's anxiety, the voice AI will add a reassuring phrase such as "Don't worry about that. We will respond immediately." It also learns from the user's past emotional data to improve the quality of future dialogue.

[0309] Specific examples

[0310] Example 1: Canceling a subscription

[0311] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server sends a notification to the user's device stating, "Cancellation completed."

[0312] Example 2: Changing a restaurant reservation

[0313] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server sends a notification to the user's device stating, "The reservation change has been completed."

[0314] Prompt Sentence Examples

[0315] Here are some example prompts to input to the generative AI model:

[0316] 1. Example of a user terminal prompt:

[0317] A user has requested to cancel their subscription using a dedicated application. What are their next steps?

[0318] 2. Server prompt example:

[0319] The server has received a user cancellation request, how can we parse the request and understand the task?

[0320] 3. Voice AI prompt example:

[0321] The voice AI needs to dial a specified phone number and proceed with canceling the subscription. How do you adapt the dialogue if the user is unsure?

[0322] 4. Emotion Engine Prompt Example:

[0323] The emotion engine detects anxiety in the user's voice. How can the voice AI change the dialogue to reassure the user?

[0324] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0325] Step 1:

[0326] The user enters the task.

[0327] The user enters task details such as cancellation or reservation change using a dedicated application or web interface. For example, they enter the contract number "123456," the reservation number "789012," and the desired change date and time "October 10th, 7:00 PM," and then press the send button. The contract number and reservation number are given as input, and JSON-formatted request data is generated as output.

[0328] Step 2:

[0329] The device sends a request to the server.

[0330] The terminal sends the information entered by the user to the server. Specifically, it sends the JSON formatted request data mentioned earlier to the server via an HTTP request. This request data is input and sent to the server as output.

[0331] Step 3:

[0332] The server receives and analyzes the request.

[0333] The server receives the request sent from the terminal and analyzes the request data. Specifically, it parses the JSON data and extracts necessary data such as the contract number and reservation information. The input to this analysis process is the request data, and the output is the analyzed task information.

[0334] Step 4:

[0335] The server retrieves the data.

[0336] The server retrieves the necessary contract or reservation information from an internal database. For example, retrieve the user's contract information associated with contract number "123456" from the database. The input to this database query is the contract or reservation number, and the output is any additional related information.

[0337] Step 5:

[0338] The server issues preparation instructions to the voice AI and emotion engine.

[0339] Based on the analysis results, the server instructs the voice AI and emotion engine to prepare for task execution. For example, it sends an instruction to "cancel contract number 123456." The input for this process is the analysis results and acquired data, and the output is a preparation instruction for the voice AI and emotion engine.

[0340] Step 6:

[0341] The server will have the voice AI start the call.

[0342] The server instructs the voice AI to dial the specified phone number. The voice AI receives this instruction and makes the call. The input of this process is the dial instruction and the phone number, and the output is the outgoing call status.

[0343] Step 7:

[0344] The voice AI begins the conversation.

[0345] When the other party responds, the voice AI begins a dialogue to accomplish a task. For example, say, "I would like to cancel the subscription for contract number 123456." The input for this process is the task content, and the output is the dialogue content.

[0346] Step 8:

[0347] The emotion engine recognizes emotions.

[0348] The emotion engine analyzes the user's voice and the voice of the other party to recognize emotions. For example, it can detect anxiety from the user's tone of voice. The input for this analysis is voice data, and the output is recognized emotional information.

[0349] Step 9:

[0350] The server records the call results.

[0351] Once the task is completed, the server receives feedback from the voice AI and records the call results. Specifically, the call content and the success or failure of the cancellation are stored in a database. The input of this process is the call feedback, and the output is the recorded data.

[0352] Step 10:

[0353] The server notifies the user of the results.

[0354] The server sends a notification to the user based on the recorded call result. For example, it sends a message to the user's device saying, "Subscription cancellation has been completed." The input of this notification is the recorded call result, and the output is a notification message to the user's device.

[0355] (Application example 2)

[0356] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0357] Modern security services face the challenge of requiring a lot of manual effort and time for users to efficiently complete specific tasks. Furthermore, they may not respond appropriately to the user's emotional state, which can increase anxiety and stress. This leads to a poor user experience and slows down the completion of security tasks, resulting in inefficiencies.

[0358] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0359] In this invention, the server includes a means for the user to input the task content and necessary information, a means for the server to receive and analyze the request from the user, a means for the voice AI to autonomously execute the task and make a call, a means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI, and a means for the server to record the results of the call and notify the user. This makes it possible to complete security tasks efficiently and quickly while taking the user's emotions into consideration.

[0360] "Means for users to input task details and required information" refers to an interface that allows users to input detailed information about the task they wish to perform at that time and required data. This includes smartphone apps and web interfaces.

[0361] The "means for the server to receive and analyze the request from the user" refers to a system that receives a request sent from a user terminal, analyzes it, and understands the specific content of the task and the required information, thereby preparing for the next processing step.

[0362] "Means for voice AI to autonomously execute tasks and make calls" refers to a function in which voice AI automatically communicates with designated parties on behalf of the user and carries out tasks. It uses natural language processing technology to carry out dialogue and carry out tasks.

[0363] "Means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI" refers to a function that analyzes the emotional state of the user from their voice and input content, and the voice AI generates and changes appropriate dialogue content according to that emotion. This makes it possible to respond in a way that takes the user's emotions into consideration.

[0364] The "means for the server to record the results of the call and notify the user" refers to a system in which the server records the results of the call made by the voice AI and notifies the user of that information, thereby letting the user know whether the task was performed correctly.

[0365] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language, including speech recognition, semantic analysis, and dialogue generation.

[0366] A "dialogue script" is a set of defined phrases and sentences that a voice AI uses during a call, allowing the voice AI to have a natural conversation.

[0367] A "task" is a detailed description of the specific action or task the user is trying to accomplish, such as locking a bank account, calling an emergency service, or controlling a home security system.

[0368] "Call outcome" refers to the final outcome or status of a call made by a voice AI, including whether the task was successful or failed and its detailed status.

[0369] "Feedback" is the process of providing users with information about the outcome of a call, the progress of a task, etc., so that they know the status of their request.

[0370] This invention is a system that allows users to efficiently and flexibly execute specific tasks, and in particular, by combining voice AI and an emotion engine, it realizes dialogue that takes into account the user's emotions. This system is composed of the following elements:

[0371] 1. User Device

[0372] A user terminal is a device such as a smartphone or computer that allows users to input the task details and necessary information and submit requests. A dedicated application or web interface is provided through which users interact with the system.

[0373] 2. Server

[0374] The server receives the user's request, analyzes it, and retrieves the necessary information from the database. The server performs the following operations:

[0375] Analyze the voice data received from the user and understand the content of the task.

[0376] It obtains the necessary information and issues instructions to the voice AI and emotion engine.

[0377] The voice AI records the results of the tasks it performs and notifies the user.

[0378] 3. Voice AI

[0379] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue, using natural language processing technologies such as Google Cloud Speech-to-Text and Dialogflow. Voice AI does the following:

[0380] Automatically call the specified person.

[0381] Analyzes call content and generates natural dialogue.

[0382] 4. Emotion Engine

[0383] The emotion engine recognizes emotions from the user's voice and adaptively changes the dialogue content of the voice AI. For this purpose, IBM Watson (registered trademark) Tone Analyzer is used. The emotion engine performs the following processes:

[0384] Analyzes the user's voice data and extracts their emotional state.

[0385] Depending on the emotion, instructions are sent to the voice AI to change the content of the dialogue.

[0386] Specific examples

[0387] Example 1: Locking a bank account

[0388] The user launches the application and commands it by voice, "Lock my bank account." The server receives the user's request and analyzes the voice data. Once the analysis is complete, the server instructs the voice AI to call the bank's customer support. If the emotion engine recognizes the user's nervousness, the voice AI will say, "Don't worry, we'll take care of it right away." Once the task is completed, the server notifies the user, "Your bank account has been successfully locked."

[0389] Prompt Sentence Examples

[0390] User input: "Lock my bank account"

[0391] Example prompt for a generative AI model: "Security-related task: User breathlessly commands, 'Lock my bank account.' The voice AI automatically calls a call center, and the emotion engine recognizes the tension and generates a reassuring response, 'Don't worry, we'll get back to you shortly.'"

[0392] In this way, the present invention enables a system that combines voice AI and an emotion engine to efficiently and flexibly execute tasks through dialogue that takes into consideration the user's emotions.

[0393] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0394] Step 1:

[0395] User Input

[0396] The user uses a smartphone or computer application to enter the task and required information by voice, for example, "lock my bank account."

[0397] Input: Voice data (e.g., "Lock my bank account")

[0398] Output: Audio data is sent to the system

[0399] How it works: The user launches a specific application and speaks a task into the microphone.

[0400] Step 2:

[0401] Receiving and analyzing requests (server)

[0402] The server receives the user's voice data, converts it into text using Google Cloud Speech-to-Text, and then analyzes the task content using Dialogflow.

[0403] Input: Voice data (e.g., "Lock my bank account")

[0404] Output: Text data (e.g. "Lock your bank account")

[0405] How it works: The server receives the voice data, converts it into text using speech recognition technology, and then analyzes the specific content of the task.

[0406] Step 3:

[0407] Emotion Recognition (Emotion Engine)

[0408] The server analyzes the user's emotional state from the voice data using IBM Watson Tone Analyzer.

[0409] Input: Audio data

[0410] Output: Emotional state data (e.g., tension, anxiety)

[0411] How it works: The server inputs the voice data into the emotion engine and analyzes the emotional state.

[0412] Step 4:

[0413] Task preparation (server)

[0414] Based on the analysis results, the server prepares the data necessary for the voice AI, such as obtaining the bank's customer support phone number and account information.

[0415] Input: Text data, emotional state data

[0416] Output: Data to be passed to the voice AI (e.g., phone number, account information)

[0417] Action: Prepares the data required to perform the task and generates a script for the voice AI to interact.

[0418] Step 5:

[0419] Task execution (voice AI)

[0420] The voice AI automatically calls the specified person and engages in a conversation. It uses Dialogflow to generate natural conversations and perform tasks. Based on data from the emotion engine, the voice AI changes the content of the conversation and speaks in a way that takes the user's emotions into consideration.

[0421] Input: Subject data, emotional state data

[0422] Output: Voice AI dialogue script, call results

[0423] How it works: The voice AI makes a call and proceeds with the conversation according to the specified content. For example, it generates and uses reassuring phrases such as "Don't worry, we'll respond right away."

[0424] Step 6:

[0425] Recording and notifying call results (server)

[0426] After the call is finished, the server records the result of the call and notifies the user, sending a message such as "Your bank account has been successfully locked."

[0427] Input: Call result data

[0428] Output: Record of call outcome, notification to user

[0429] How it works: The server stores the call results in a database and sends a notification to the user's device.

[0430] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0431] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0432] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0433] [Second embodiment]

[0434] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0435] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0436] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0437] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0438] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0440] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0441] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0442] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0443] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0444] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0445] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0446] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by its use of voice AI to automate operations via voice calls. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0447] System configuration

[0448] 1. User Device

[0449] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[0450] 2. Server

[0451] The server receives and analyzes requests from users, retrieves necessary information from a database, and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user.

[0452] 3. Voice AI

[0453] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[0454] What the program does

[0455] User submits a request

[0456] The user uses a dedicated application or web interface to input the details of the task (e.g., canceling a subscription or changing a reservation). For example, if the user wants to cancel a subscription, they enter the contract number and other necessary information and press the submit button. The device then sends this request to the server.

[0457] The server parses the request

[0458] The server analyzes the received request and understands the task. It checks the necessary information (e.g., subscription contract information and reservation number) and prepares the voice AI. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[0459] Voice AI performs tasks

[0460] The voice AI connected to the server dials a real phone number on behalf of the user. After dialing the specified phone number and receiving the answer, the voice AI uses natural language processing technology to carry out the dialogue. For example, the voice AI might say, "Hello, I would like to cancel my subscription for contract number 123456." As the dialogue progresses, the AI ​​provides necessary information to complete the task.

[0461] Call outcome feedback

[0462] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification such as "Subscription cancellation completed" will be sent to the user's device. This allows the user to complete the necessary task without having to spend time on their own.

[0463] Specific examples

[0464] Example 1: Canceling a subscription

[0465] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI calls customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[0466] Example 2: Changing a restaurant reservation

[0467] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI calls the restaurant and says, "Please change reservation number 789012." After confirming the change, the server notifies the user that "the reservation change has been completed."

[0468] In this way, this system uses voice AI to efficiently perform tasks on behalf of the user and provide feedback on the results.

[0469] The processing flow will be explained below.

[0470] Step 1:

[0471] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[0472] Step 2:

[0473] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[0474] Step 3:

[0475] The device sends a user request to the server, including the type of task and the required details.

[0476] Step 4:

[0477] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[0478] Step 5:

[0479] The server analyzes the request and identifies the type of task (e.g., canceling a subscription). It then retrieves the information needed to execute the task (such as contract and reservation information) from the database.

[0480] Step 6:

[0481] The server calls the voice AI module and prepares to execute the task. It generates a dialogue script according to the task. For example, if it is a cancellation procedure, it prepares a script that says, "I would like to cancel contract number 123456."

[0482] Step 7:

[0483] The voice AI automatically dials the specified phone number, waits for the call to connect, and starts the conversation when the other party answers.

[0484] Step 8:

[0485] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[0486] Step 9:

[0487] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[0488] Step 10:

[0489] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[0490] Step 11:

[0491] The server analyzes the results of the call and notifies the user of the outcome, including information about task completion and next steps.

[0492] Step 12:

[0493] The user receives a notification on their device confirming that the task was completed as expected, for example, a message saying "Your subscription has been successfully canceled."

[0494] This process allows users to complete tasks efficiently without spending time on their own.

[0495] Example 1

[0496] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0497] In today's society, tasks users perform over the phone (e.g., canceling subscriptions or changing reservations) are time-consuming and cumbersome, necessitating efficient solutions. Furthermore, negotiating directly over the phone can be stressful for users, so technology is needed to simplify and automate these processes. Furthermore, conventional voice recognition systems struggle to deliver natural dialogue and appropriate responses, so improvements are needed.

[0498] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0499] In this invention, the server includes a means for receiving and analyzing requests from users, a means for calling a voice AI and sending instructions including necessary information, and a means for recording the results of the call and notifying the user. This allows the user to efficiently automate tasks without having to make a call, reducing stress and enabling appropriate responses.

[0500] "User" means any person or entity that intends to use the System to perform a specific task.

[0501] A "task" is a specific purpose or action a user attempts to perform through the system (e.g., canceling a subscription or changing a reservation).

[0502] "Information" refers to data required to perform a task (e.g., contract number or reservation number).

[0503] "Means" refers to a method or combination of methods for achieving a specific function or operation.

[0504] "Server" refers to the central processing unit that receives and analyzes requests from users and controls the voice AI.

[0505] A "request" refers to a communication that describes the task or request sent by a user to a system.

[0506] "Analysis" refers to the operation by the server to interpret the contents of the request and determine the necessary processing.

[0507] "Voice AI" refers to artificial intelligence that makes voice calls on behalf of users and uses natural language processing technology to guide the conversation.

[0508] "Calling" refers to the operation in which the server sends instructions to the voice AI to perform a specific task.

[0509] "Call" refers to the voice communication that the voice AI has with customer support or other parties.

[0510] "Results" refers to information showing the status and results after the voice AI has performed a task.

[0511] "Notification" refers to communication from the server to the user informing them of task completion and progress.

[0512] "Natural language processing technology" refers to technology for understanding and generating human language.

[0513] A "dialogue script" refers to a scenario of a series of responses and questions that a voice AI uses during a call.

[0514] This invention is a system that allows users to efficiently perform specific tasks, and in particular automates operations via voice calls using voice AI. The main elements of this system are the user device, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0515] User Device

[0516] A user terminal is a device such as a smartphone or computer. The user interacts with the system through a dedicated application or a web interface. Using this interface, the user enters the details of the task (e.g., canceling a subscription, changing a reservation) and provides the required information. For example, if the user wants to cancel a subscription, they enter their contract number and other required information and press the submit button. The terminal then sends this request to the server.

[0517] server

[0518] The server receives and analyzes the request sent from the user's device. This analysis allows it to understand the task content and retrieve the necessary information from the database. For example, the server may determine that the user wishes to cancel a subscription and gather the necessary data for cancellation. The server also issues instructions to the voice AI to execute the task. It passes the necessary information (e.g., contract number 123456) to the voice AI and instructs it to execute a specific task.

[0519] voice AI

[0520] The voice AI calls a designated phone number on behalf of the user, and after the caller answers, it uses natural language processing technology to continue the conversation. For example, when a customer support operator answers, they say, "Hello, I'd like to cancel my subscription for contract number 123456," and then proceeds with the necessary procedures. They ask questions and provide information as the conversation progresses.

[0521] Call outcome feedback

[0522] After the task is completed, the voice AI sends the call results to the server. The server analyzes the call results and records the necessary information (e.g., cancellation completed, reservation change completed). The server then notifies the user's device that the task has been completed. For example, the user's device may receive a notification saying, "Subscription cancellation has been completed."

[0523] Specific examples

[0524] Example 1: Canceling a subscription

[0525] 1. The user opens the application and selects "Cancel Subscription."

[0526] 2. Enter the contract number "123456" and other required information and submit.

[0527] 3. The server analyzes the request and the voice AI calls customer support.

[0528] 4. The voice AI says, "Please cancel contract number 123456," and continues the conversation.

[0529] 5. Once the cancellation procedure is complete, the server will notify the user that "Cancellation has been completed."

[0530] Example 2: Changing a restaurant reservation

[0531] 1. The user selects "Change Restaurant Reservation" in the application.

[0532] 2. Enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit.

[0533] 3. The server analyzes the request and the voice AI calls the restaurant.

[0534] 4. The voice AI says, "Please change reservation number 789012."

[0535] 5. After confirming the change, the server notifies the user that the reservation change has been completed.

[0536] Prompt Sentence Examples

[0537] Example 1: Subscription cancellation prompt

[0538] Prompt: Please cancel your subscription. Contract number is 123456.

[0539] Example 2: Prompt for changing a restaurant reservation

[0540] Prompt: I would like to change my restaurant reservation. The reservation number is 789012, and the desired date and time is October 10th at 7:00 PM.

[0541] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0542] Step 1:

[0543] The user enters the task

[0544] The user opens a dedicated application or web interface on their smartphone or computer and enters the details of the task. For example, if they want to cancel a subscription, they enter the contract number "123456" and add any other necessary information. When the user presses the "Submit" button, the device collects this input data and sends it to the server.

[0545] Input: Task details and contract number entered by the user

[0546] Output: Request data sent to the server (detailed task information)

[0547] Specific behavior:

[0548] Open the application.

[0549] Select "Cancel Subscription."

[0550] Enter the contract number "123456".

[0551] Enter other necessary information.

[0552] Press the send button.

[0553] Step 2:

[0554] The server receives and analyzes the request

[0555] The server receives the request sent from the device, analyzes its contents, identifies the type of task included in the request, and sends a query to the database to extract the necessary information (e.g., contract number), thereby obtaining the data necessary for the cancellation procedure.

[0556] Input: Request data sent from the device (task details)

[0557] Output: Analyzed task type and extracted required information (e.g. contract number)

[0558] Specific behavior:

[0559] Received request.

[0560] Analyze the request content.

[0561] Identify the type of task.

[0562] Obtain necessary information (contract number, etc.) from the database.

[0563] Step 3:

[0564] The server calls the voice AI

[0565] Based on the analysis results, the server calls the voice AI to execute the appropriate task. The server passes the necessary information (e.g., contract number) to the voice AI and sends an instruction to execute. For example, the server issues an instruction such as "Contact customer support to cancel the subscription for contract number 123456."

[0566] Input: Analysis results and necessary information (contract number, etc.)

[0567] Output: Instructions and task information passed to the voice AI

[0568] Specific behavior:

[0569] Call up the voice AI.

[0570] Pass the necessary information to the voice AI.

[0571] Send instructions to perform tasks.

[0572] Step 4:

[0573] Voice AI initiates and manages calls

[0574] The voice AI will call the specified phone number and start the call. Once the call begins, the voice AI will use natural language processing technology to proceed with the conversation. For example, when a customer support operator answers, they will say, "Hello, I would like to cancel my subscription for contract number 123456," and proceed with the process.

[0575] Input: Execution instructions and task information passed from the server

[0576] Output: Dialogue outcome and progress during the call

[0577] Specific behavior:

[0578] Dial the specified phone number.

[0579] Customer support responds.

[0580] Say, "Hello, I would like to cancel my subscription for contract number 123456."

[0581] Ask questions and provide information as necessary depending on the procedure.

[0582] Step 5:

[0583] The server records the results and sends feedback

[0584] After the voice AI completes the task, it sends the results of the call to the server. The server analyzes the call results and records the necessary information. Finally, the server notifies the user device that the task has been completed. For example, it may send a notification to the user saying, "Your subscription has been canceled."

[0585] Input: Call result sent from voice AI

[0586] Output: Parsed results and user notification

[0587] Specific behavior:

[0588] Receive call results from voice AI.

[0589] Analyze the results and record the necessary information.

[0590] Sends notifications to the user's device.

[0591] A message will be sent stating that your subscription has been cancelled.

[0592] (Application example 1)

[0593] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0594] Online shopping and food delivery services are rapidly becoming more popular, but there is a problem in that it takes time and effort for users to make changes to their orders or make inquiries. In particular, inquiries and changes made over the phone are prone to waiting times and communication problems. It also increases the burden on customer support, making it difficult to provide efficient service. To solve these issues, there is a need for a system that allows users to easily request changes to their orders and automates the entire process.

[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0596] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the server to receive and analyze the request from the user; a means for the voice AI to autonomously execute the task and make a call; a means for the server to record the result of the call and notify the user; a means for the user to request an order change and specify the new order content; and a means for the server to contact the ordering party via the voice AI and request the order change. This allows the user to easily request an order change, and the entire process is automated, eliminating problems such as waiting times and communication errors, enabling efficient service provision.

[0597] "Means for users to input task details and necessary information" refers to an interface through which users input information necessary to perform a specific task, such as a smartphone application or a web interface.

[0598] "Means for the server to receive and analyze a request from the user" refers to a function that enables the server to receive request data sent from the user, analyze it, and identify the necessary processing.

[0599] "Means for voice AI to autonomously execute tasks and make calls" refers to a function that enables voice AI to autonomously execute designated tasks on behalf of the user via telephone.

[0600] "Means for the server to record the results of the call and notify the user" refers to a function that enables the server to record the results of the call made by voice AI in a database and notify the user of the results.

[0601] The "means by which a user requests an order change and specifies the new order details" is an interface by which a user can change the current order details and specify the new order details after the changes.

[0602] "Means for the server to contact the ordering party via voice AI and request a change to the order" is a function that allows the server to use voice AI to call the designated ordering party and request a change to the user's order.

[0603] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by the use of voice AI to automate food delivery order changes and inquiries. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0604] System configuration

[0605] 1. User Device

[0606] A user terminal is a device, such as a smartphone or computer, that a user uses to enter change order details and submit requests. A dedicated application or web interface is provided through which the user interacts with the system.

[0607] 2. Server

[0608] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user. This process uses software such as natural language processing (NLP) libraries (e.g., Google Cloud Text-to-Speech, Amazon Lex) and communication libraries (e.g., Twilio).

[0609] 3. Voice AI

[0610] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and engage in natural conversations with the other party (e.g., restaurant staff). Voice AI calls use NLP technologies such as Google Cloud Text-to-Speech and Amazon Lex.

[0611] What the program does

[0612] User submits a request

[0613] The user uses a dedicated application or web interface to input the details of the order change (e.g., change to two pizzas). For example, if the user wants to "change the order," they enter the current order number and the new order details and press the submit button. The device then sends this request to the server.

[0614] The server parses the request

[0615] The server analyzes the received request and understands the task. It checks the necessary information (e.g., the current order number and the new order details) and prepares the voice AI. For example, the server determines that the user wants to change the order and prepares the necessary data for the change.

[0616] Voice AI performs tasks

[0617] The voice AI connected to the server dials a real phone number on behalf of the user. It dials the specified restaurant's phone number, and after the caller answers, the voice AI uses natural language processing technology to carry out the conversation. For example, the voice AI might say, "Hello, I'd like to change the order for order number 123456. The new order is for two pizzas." As the conversation progresses, the AI ​​provides necessary information to complete the task.

[0618] Call outcome feedback

[0619] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification saying "Order change completed" will be sent to the user's device. This allows the user to complete the necessary tasks without having to spend time on their own.

[0620] Specific examples

[0621] Example 1: Order change procedure

[0622] The user opens the application and selects "Change Order." They enter the current order number "123456" and the new order details "2 pizzas," and submit. The server analyzes the request, and the voice AI calls the restaurant. The voice AI responds by saying, "Please change order number 123456 to 2 pizzas," and continues the conversation. Once the change procedure is complete, the server notifies the user, "The order change has been completed."

[0623] Prompt Sentence Examples

[0624] "Imagine a food delivery application where a user requests a change to their order. When the user inputs that they would like to change the order for order number 123456 to two pizzas, write a program in which the voice AI automatically calls the restaurant and requests the order change. If the call is successful, please also include a function to notify the user that "the order change has been completed."

[0625] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0626] Step 1:

[0627] The user opens the smartphone application and accesses the screen for requesting an order change. The user enters the current order number and new order details (e.g., two pizzas) and presses the "Submit" button. The data entered is the order number and new order details. The device sends this request to the server as JSON-formatted data.

[0628] Step 2:

[0629] The server receives the request sent from the terminal. The received data includes the order number and new order details. The server analyzes the data and recognizes that the user wants to change the order. The server first retrieves the relevant order information from the database and locally checks the updated status of the order data based on the new order details. The output is the necessary information for preparing the voice AI.

[0630] Step 3:

[0631] The server provides the voice AI with details of the order change and contact information for the person on the other end of the call (e.g., restaurant staff). The input is the detailed order change information obtained in the previous step. The server uses a generative AI model (e.g., Google Cloud Text-to-Speech or Amazon Lex) to generate a dialogue script for the voice AI. The script includes a greeting at the start of the call, providing the order number, and explaining the new order details. The generated dialogue script is output.

[0632] Step 4:

[0633] The voice AI automatically calls the specified restaurant's phone number based on the dialogue script provided by the server. The input is the dialogue script and phone number. Natural language processing (NLP) is used to analyze and generate dialogue content, and ask the restaurant staff to change the order. Specifically, it says, "Hello, I'd like to change order number 123456. The new order is for two pizzas." The necessary information is presented appropriately during the call, and the dialogue continues until the order change is completed. The output is the call result.

[0634] Step 5:

[0635] Once the voice AI completes the call, it sends the results back to the server, including confirmation of the call's success or failure and any specific changes that were made. The input is the call result data. The server uses this information to update the database and record the changes. The output is the updated order information.

[0636] Step 6:

[0637] The server creates a notification for the user based on the call result. The input is the updated order information. It generates a message such as "Order changes completed" and sends it to the user's device. The output is a notification message.

[0638] Step 7:

[0639] The user terminal receives the notification message from the server and displays it on the screen. The user confirms the message "Order change completed." Through this process, the user can complete the order change without any hassle.

[0640] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0641] This invention is a system that allows users to efficiently perform specific tasks. It is characterized by utilizing voice AI to automate operations via voice calls, and by combining it with an emotion engine, it recognizes the user's emotions and adaptively changes the response. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine. Each element works together to autonomously process the user's tasks.

[0642] System configuration

[0643] 1. User Device

[0644] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[0645] 2. Server

[0646] The server receives and analyzes user requests, retrieves necessary information from a database, and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user.

[0647] 3. Voice AI

[0648] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[0649] 4. Emotion Engine

[0650] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue, enabling optimal dialogue according to the user's emotional state. It also learns from the user's past emotional data to improve the quality of future dialogue.

[0651] What the program does

[0652] User submits a request

[0653] The user enters the details of the task using a dedicated application or web interface. For example, if they want to cancel a subscription, they enter their contract number and other necessary information and press the submit button. The device then sends this request to the server.

[0654] The server parses the request

[0655] The server analyzes the received request and understands the task. It checks the necessary information (e.g., contract information and reservation number) and prepares the voice AI and emotion engine. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[0656] Voice AI and emotion engines perform tasks

[0657] The voice AI and emotion engine connected to the server work together to dial the phone on behalf of the user. After dialing the specified phone number and the caller answers, the voice AI uses natural language processing technology to carry out the dialogue. The emotion engine recognizes emotions from the user's voice and changes the dialogue content as necessary. For example, when the voice AI says, "Hello, I'd like to cancel my subscription for contract number 123456," if the emotion engine recognizes the user's anxiety or irritation, the voice AI adds a phrase that provides reassurance, such as, "Don't worry about that. We'll deal with it right away."

[0658] Call outcome feedback

[0659] After the task is completed, the server records the call result and sends a notification to the user. For example, a notification saying "Your subscription has been canceled" will be sent to the user's device. The emotional data collected by the emotion engine can be used to improve the quality of future conversations.

[0660] Specific examples

[0661] Example 1: Canceling a subscription

[0662] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[0663] Example 2: Changing a restaurant reservation

[0664] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server notifies the user, "The reservation change has been completed."

[0665] In this way, by using voice AI and an emotion engine, this system can provide optimal dialogue based on the user's emotions and complete tasks efficiently.

[0666] The processing flow will be explained below.

[0667] Step 1:

[0668] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[0669] Step 2:

[0670] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[0671] Step 3:

[0672] The device sends a user request to the server, including the type of task and the required details.

[0673] Step 4:

[0674] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[0675] Step 5:

[0676] The server analyzes the request, recognizes what the task is (e.g. cancel a subscription), retrieves the necessary information from the database, and issues instructions to the voice AI and emotion engine.

[0677] Step 6:

[0678] The server generates a dialogue script for the voice AI to execute a task and prepares the emotion engine. For example, the server prepares a script that says, "I would like to cancel contract number 123456."

[0679] Step 7:

[0680] The voice AI will dial the specified phone number, wait for the call to connect, and begin the conversation once the other party answers.

[0681] Step 8:

[0682] The emotion engine analyzes the user's voice to determine their emotions and provides that information to the voice AI, which then adapts the dialogue accordingly. For example, if the user sounds anxious, the AI ​​might add a phrase like, "Don't worry, we'll get back to you shortly."

[0683] Step 9:

[0684] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[0685] Step 10:

[0686] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[0687] Step 11:

[0688] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[0689] Step 12:

[0690] The server analyzes the results of the call and notifies the user of the outcome. The user receives a notification on their device confirming that the task was completed as planned, for example, a message saying "Subscription cancellation successful."

[0691] Step 13:

[0692] The emotion engine learns from the emotional data collected during processing and uses it to improve the quality of future interactions, resulting in a continuously improved user experience.

[0693] Example 2

[0694] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0695] Conventional voice communication systems require users to perform detailed operations themselves, making it difficult to complete tasks efficiently. Furthermore, they only provide a uniform response without considering the user's feelings, resulting in low user satisfaction. This has led to a demand for a system that can reduce the user's time and effort while also providing optimal responses tailored to individual feelings.

[0696] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0697] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the terminal to send the input information to the server; a means for the server to receive and analyze the request from the user; a means for the server to acquire data based on the received and analyzed information and issue preparation instructions to the voice AI and emotion engine; a means for the voice AI to autonomously execute the task and make a call; a means for the emotion engine to recognize emotions from the user's voice; and a means for the server to record the results of the call and notify the user. This allows the user to complete the task efficiently without detailed operations, and also enables optimal responses according to the user's emotions.

[0698] A "user" is someone who utilizes the system to perform a task.

[0699] A "terminal" is a device that a user uses to interact with the system, such as a smartphone or computer.

[0700] A "server" is a computer system that receives and analyzes requests from users.

[0701] "Voice AI" refers to artificial intelligence technology that generates voice and performs tasks autonomously.

[0702] The "emotion engine" is a technology that recognizes emotions from the user's voice and adaptively changes the content of the dialogue according to those emotions.

[0703] "Request" refers to a request including the content of a task that a user sends to a server via a terminal.

[0704] "Natural language processing technology" is a technology that allows computers to understand and generate human language.

[0705] A "call" is a means of communication through voice, and specifically refers to an exchange using a telephone.

[0706] "Data acquisition" refers to the process by which the server gathers the necessary information from internal databases and external sources.

[0707] "Notification" refers to a message sent by the server to inform the user of the outcome of a task.

[0708] This invention is a system that enables users to efficiently perform specific tasks, automating operations via voice calls by combining primarily voice AI and an emotion engine, and providing optimal responses according to the user's emotions. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine.

[0709] User Device

[0710] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system. For example, if a user wants to cancel a subscription, they open the dedicated application, enter their contract number and other required information, and press the submit button.

[0711] server

[0712] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user. The server analyzes the requests it receives, parses the request data, and extracts necessary information to understand the content of the task.

[0713] voice AI

[0714] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have a natural conversation with the other party (e.g., a customer support operator). For example, when the voice AI says, "Hello, I would like to cancel my subscription for contract number 123456," it uses natural language processing technology to smoothly guide the conversation.

[0715] Emotion Engine

[0716] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue. As a result, it is possible to have an optimal dialogue according to the user's emotional state. If the emotion engine detects the user's anxiety, the voice AI will add a reassuring phrase such as "Don't worry about that. We will respond immediately." It also learns from the user's past emotional data to improve the quality of future dialogue.

[0717] Specific examples

[0718] Example 1: Canceling a subscription

[0719] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server sends a notification to the user's device stating, "Cancellation completed."

[0720] Example 2: Changing a restaurant reservation

[0721] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server sends a notification to the user's device stating, "The reservation change has been completed."

[0722] Prompt Sentence Examples

[0723] Here are some example prompts to input to the generative AI model:

[0724] 1. Example of a user terminal prompt:

[0725] A user has requested to cancel their subscription using a dedicated application. What are their next steps?

[0726] 2. Server prompt example:

[0727] The server has received a user cancellation request, how can we parse the request and understand the task?

[0728] 3. Voice AI prompt example:

[0729] The voice AI needs to dial a specified phone number and proceed with canceling the subscription. How do you adapt the dialogue if the user is unsure?

[0730] 4. Emotion Engine Prompt Example:

[0731] The emotion engine detects anxiety in the user's voice. How can the voice AI change the dialogue to reassure the user?

[0732] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0733] Step 1:

[0734] The user enters the task.

[0735] The user enters task details such as cancellation or reservation change using a dedicated application or web interface. For example, they enter the contract number "123456," the reservation number "789012," and the desired change date and time "October 10th, 7:00 PM," and then press the send button. The contract number and reservation number are given as input, and JSON-formatted request data is generated as output.

[0736] Step 2:

[0737] The device sends a request to the server.

[0738] The terminal sends the information entered by the user to the server. Specifically, it sends the JSON formatted request data mentioned earlier to the server via an HTTP request. This request data is input and sent to the server as output.

[0739] Step 3:

[0740] The server receives and analyzes the request.

[0741] The server receives the request sent from the terminal and analyzes the request data. Specifically, it parses the JSON data and extracts necessary data such as the contract number and reservation information. The input to this analysis process is the request data, and the output is the analyzed task information.

[0742] Step 4:

[0743] The server retrieves the data.

[0744] The server retrieves the necessary contract or reservation information from an internal database. For example, retrieve the user's contract information associated with contract number "123456" from the database. The input to this database query is the contract or reservation number, and the output is any additional related information.

[0745] Step 5:

[0746] The server issues preparation instructions to the voice AI and emotion engine.

[0747] Based on the analysis results, the server instructs the voice AI and emotion engine to prepare for task execution. For example, it sends an instruction to "cancel contract number 123456." The input for this process is the analysis results and acquired data, and the output is a preparation instruction for the voice AI and emotion engine.

[0748] Step 6:

[0749] The server will have the voice AI start the call.

[0750] The server instructs the voice AI to dial the specified phone number. The voice AI receives this instruction and makes the call. The input of this process is the dial instruction and the phone number, and the output is the outgoing call status.

[0751] Step 7:

[0752] The voice AI begins the conversation.

[0753] When the other party responds, the voice AI begins a dialogue to accomplish a task. For example, say, "I would like to cancel the subscription for contract number 123456." The input for this process is the task content, and the output is the dialogue content.

[0754] Step 8:

[0755] The emotion engine recognizes emotions.

[0756] The emotion engine analyzes the user's voice and the voice of the other party to recognize emotions. For example, it can detect anxiety from the user's tone of voice. The input for this analysis is voice data, and the output is recognized emotional information.

[0757] Step 9:

[0758] The server records the call results.

[0759] Once the task is completed, the server receives feedback from the voice AI and records the call results. Specifically, the call content and the success or failure of the cancellation are stored in a database. The input of this process is the call feedback, and the output is the recorded data.

[0760] Step 10:

[0761] The server notifies the user of the results.

[0762] The server sends a notification to the user based on the recorded call result. For example, it sends a message to the user's device saying, "Subscription cancellation has been completed." The input of this notification is the recorded call result, and the output is a notification message to the user's device.

[0763] (Application example 2)

[0764] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0765] Modern security services face the challenge of requiring a lot of manual effort and time for users to efficiently complete specific tasks. Furthermore, they may not respond appropriately to the user's emotional state, which can increase anxiety and stress. This leads to a poor user experience and slows down the completion of security tasks, resulting in inefficiencies.

[0766] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0767] In this invention, the server includes a means for the user to input the task content and necessary information, a means for the server to receive and analyze the request from the user, a means for the voice AI to autonomously execute the task and make a call, a means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI, and a means for the server to record the results of the call and notify the user. This makes it possible to complete security tasks efficiently and quickly while taking the user's emotions into consideration.

[0768] "Means for users to input task details and required information" refers to an interface that allows users to input detailed information about the task they wish to perform at that time and required data. This includes smartphone apps and web interfaces.

[0769] The "means for the server to receive and analyze the request from the user" refers to a system that receives a request sent from a user terminal, analyzes it, and understands the specific content of the task and the required information, thereby preparing for the next processing step.

[0770] "Means for voice AI to autonomously execute tasks and make calls" refers to a function in which voice AI automatically communicates with designated parties on behalf of the user and carries out tasks. It uses natural language processing technology to carry out dialogue and carry out tasks.

[0771] "Means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI" refers to a function that analyzes the emotional state of the user from their voice and input content, and the voice AI generates and changes appropriate dialogue content according to that emotion. This makes it possible to respond in a way that takes the user's emotions into consideration.

[0772] The "means for the server to record the results of the call and notify the user" refers to a system in which the server records the results of the call made by the voice AI and notifies the user of that information, thereby letting the user know whether the task was performed correctly.

[0773] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language, including speech recognition, semantic analysis, and dialogue generation.

[0774] A "dialogue script" is a set of defined phrases and sentences that a voice AI uses during a call, allowing the voice AI to have a natural conversation.

[0775] A "task" is a detailed description of the specific action or task the user is trying to accomplish, such as locking a bank account, calling an emergency service, or controlling a home security system.

[0776] "Call outcome" refers to the final outcome or status of a call made by a voice AI, including whether the task was successful or failed and its detailed status.

[0777] "Feedback" is the process of providing users with information about the outcome of a call, the progress of a task, etc., so that they know the status of their request.

[0778] This invention is a system that allows users to efficiently and flexibly execute specific tasks, and in particular, by combining voice AI and an emotion engine, it realizes dialogue that takes into account the user's emotions. This system is composed of the following elements:

[0779] 1. User Device

[0780] A user terminal is a device such as a smartphone or computer that allows users to input the task details and necessary information and submit requests. A dedicated application or web interface is provided through which users interact with the system.

[0781] 2. Server

[0782] The server receives the user's request, analyzes it, and retrieves the necessary information from the database. The server performs the following operations:

[0783] Analyze the voice data received from the user and understand the content of the task.

[0784] It obtains the necessary information and issues instructions to the voice AI and emotion engine.

[0785] The voice AI records the results of the tasks it performs and notifies the user.

[0786] 3. Voice AI

[0787] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue, using natural language processing technologies such as Google Cloud Speech-to-Text and Dialogflow. Voice AI does the following:

[0788] Automatically call the specified person.

[0789] Analyzes call content and generates natural dialogue.

[0790] 4. Emotion Engine

[0791] The emotion engine recognizes emotions from the user's voice and adaptively changes the dialogue content of the voice AI. For this purpose, IBM Watson Tone Analyzer is used. The emotion engine performs the following processes:

[0792] Analyzes the user's voice data and extracts their emotional state.

[0793] Depending on the emotion, instructions are sent to the voice AI to change the content of the dialogue.

[0794] Specific examples

[0795] Example 1: Locking a bank account

[0796] The user launches the application and commands it by voice, "Lock my bank account." The server receives the user's request and analyzes the voice data. Once the analysis is complete, the server instructs the voice AI to call the bank's customer support. If the emotion engine recognizes the user's nervousness, the voice AI will say, "Don't worry, we'll take care of it right away." Once the task is completed, the server notifies the user, "Your bank account has been successfully locked."

[0797] Prompt Sentence Examples

[0798] User input: "Lock my bank account"

[0799] Example prompt for a generative AI model: "Security-related task: User breathlessly commands, 'Lock my bank account.' The voice AI automatically calls a call center, and the emotion engine recognizes the tension and generates a reassuring response, 'Don't worry, we'll get back to you shortly.'"

[0800] In this way, the present invention enables a system that combines voice AI and an emotion engine to efficiently and flexibly execute tasks through dialogue that takes into consideration the user's emotions.

[0801] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0802] Step 1:

[0803] User Input

[0804] The user uses a smartphone or computer application to enter the task and required information by voice, for example, "lock my bank account."

[0805] Input: Voice data (e.g., "Lock my bank account")

[0806] Output: Audio data is sent to the system

[0807] How it works: The user launches a specific application and speaks a task into the microphone.

[0808] Step 2:

[0809] Receiving and analyzing requests (server)

[0810] The server receives the user's voice data, converts it into text using Google Cloud Speech-to-Text, and then analyzes the task content using Dialogflow.

[0811] Input: Voice data (e.g., "Lock my bank account")

[0812] Output: Text data (e.g. "Lock your bank account")

[0813] How it works: The server receives the voice data, converts it into text using speech recognition technology, and then analyzes the specific content of the task.

[0814] Step 3:

[0815] Emotion Recognition (Emotion Engine)

[0816] The server analyzes the user's emotional state from the voice data using IBM Watson Tone Analyzer.

[0817] Input: Audio data

[0818] Output: Emotional state data (e.g., tension, anxiety)

[0819] How it works: The server inputs the voice data into the emotion engine and analyzes the emotional state.

[0820] Step 4:

[0821] Task preparation (server)

[0822] Based on the analysis results, the server prepares the data necessary for the voice AI, such as obtaining the bank's customer support phone number and account information.

[0823] Input: Text data, emotional state data

[0824] Output: Data to be passed to the voice AI (e.g., phone number, account information)

[0825] Action: Prepares the data required to perform the task and generates a script for the voice AI to interact.

[0826] Step 5:

[0827] Task execution (voice AI)

[0828] The voice AI automatically calls the specified person and engages in a conversation. It uses Dialogflow to generate natural conversations and perform tasks. Based on data from the emotion engine, the voice AI changes the content of the conversation and speaks in a way that takes the user's emotions into consideration.

[0829] Input: Subject data, emotional state data

[0830] Output: Voice AI dialogue script, call results

[0831] How it works: The voice AI makes a call and proceeds with the conversation according to the specified content. For example, it generates and uses reassuring phrases such as "Don't worry, we'll respond right away."

[0832] Step 6:

[0833] Recording and notifying call results (server)

[0834] After the call is finished, the server records the result of the call and notifies the user, sending a message such as "Your bank account has been successfully locked."

[0835] Input: Call result data

[0836] Output: Record of call outcome, notification to user

[0837] How it works: The server stores the call results in a database and sends a notification to the user's device.

[0838] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0839] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0840] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0841] [Third embodiment]

[0842] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0843] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0844] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0845] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0846] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0847] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0848] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0849] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0850] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0851] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0852] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0853] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0854] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by its use of voice AI to automate operations via voice calls. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0855] System configuration

[0856] 1. User Device

[0857] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[0858] 2. Server

[0859] The server receives and analyzes requests from users, retrieves necessary information from a database, and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user.

[0860] 3. Voice AI

[0861] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[0862] What the program does

[0863] User submits a request

[0864] The user uses a dedicated application or web interface to input the details of the task (e.g., canceling a subscription or changing a reservation). For example, if the user wants to cancel a subscription, they enter the contract number and other necessary information and press the submit button. The device then sends this request to the server.

[0865] The server parses the request

[0866] The server analyzes the received request and understands the task. It checks the necessary information (e.g., subscription contract information and reservation number) and prepares the voice AI. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[0867] Voice AI performs tasks

[0868] The voice AI connected to the server dials a real phone number on behalf of the user. After dialing the specified phone number and receiving the answer, the voice AI uses natural language processing technology to carry out the dialogue. For example, the voice AI might say, "Hello, I would like to cancel my subscription for contract number 123456." As the dialogue progresses, the AI ​​provides necessary information to complete the task.

[0869] Call outcome feedback

[0870] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification such as "Subscription cancellation completed" will be sent to the user's device. This allows the user to complete the necessary task without having to spend time on their own.

[0871] Specific examples

[0872] Example 1: Canceling a subscription

[0873] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI calls customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[0874] Example 2: Changing a restaurant reservation

[0875] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI calls the restaurant and says, "Please change reservation number 789012." After confirming the change, the server notifies the user that "the reservation change has been completed."

[0876] In this way, this system uses voice AI to efficiently perform tasks on behalf of the user and provide feedback on the results.

[0877] The processing flow will be explained below.

[0878] Step 1:

[0879] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[0880] Step 2:

[0881] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[0882] Step 3:

[0883] The device sends a user request to the server, including the type of task and the required details.

[0884] Step 4:

[0885] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[0886] Step 5:

[0887] The server analyzes the request and identifies the type of task (e.g., canceling a subscription). It then retrieves the information needed to execute the task (such as contract and reservation information) from the database.

[0888] Step 6:

[0889] The server calls the voice AI module and prepares to execute the task. It generates a dialogue script according to the task. For example, if it is a cancellation procedure, it prepares a script that says, "I would like to cancel contract number 123456."

[0890] Step 7:

[0891] The voice AI automatically dials the specified phone number, waits for the call to connect, and starts the conversation when the other party answers.

[0892] Step 8:

[0893] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[0894] Step 9:

[0895] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[0896] Step 10:

[0897] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[0898] Step 11:

[0899] The server analyzes the results of the call and notifies the user of the outcome, including information about task completion and next steps.

[0900] Step 12:

[0901] The user receives a notification on their device confirming that the task was completed as expected, for example, a message saying "Your subscription has been successfully canceled."

[0902] This process allows users to complete tasks efficiently without spending time on their own.

[0903] Example 1

[0904] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0905] In today's society, tasks users perform over the phone (e.g., canceling subscriptions or changing reservations) are time-consuming and cumbersome, necessitating efficient solutions. Furthermore, negotiating directly over the phone can be stressful for users, so technology is needed to simplify and automate these processes. Furthermore, conventional voice recognition systems struggle to deliver natural dialogue and appropriate responses, so improvements are needed.

[0906] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0907] In this invention, the server includes a means for receiving and analyzing requests from users, a means for calling a voice AI and sending instructions including necessary information, and a means for recording the results of the call and notifying the user. This allows the user to efficiently automate tasks without having to make a call, reducing stress and enabling appropriate responses.

[0908] "User" means any person or entity that intends to use the System to perform a specific task.

[0909] A "task" is a specific purpose or action a user attempts to perform through the system (e.g., canceling a subscription or changing a reservation).

[0910] "Information" refers to data required to perform a task (e.g., contract number or reservation number).

[0911] "Means" refers to a method or combination of methods for achieving a specific function or operation.

[0912] "Server" refers to the central processing unit that receives and analyzes requests from users and controls the voice AI.

[0913] A "request" refers to a communication that describes the task or request sent by a user to a system.

[0914] "Analysis" refers to the operation by the server to interpret the contents of the request and determine the necessary processing.

[0915] "Voice AI" refers to artificial intelligence that makes voice calls on behalf of users and uses natural language processing technology to guide the conversation.

[0916] "Calling" refers to the operation in which the server sends instructions to the voice AI to perform a specific task.

[0917] "Call" refers to the voice communication that the voice AI has with customer support or other parties.

[0918] "Results" refers to information showing the status and results after the voice AI has performed a task.

[0919] "Notification" refers to communication from the server to the user informing them of task completion and progress.

[0920] "Natural language processing technology" refers to technology for understanding and generating human language.

[0921] A "dialogue script" refers to a scenario of a series of responses and questions that a voice AI uses during a call.

[0922] This invention is a system that allows users to efficiently perform specific tasks, and in particular automates operations via voice calls using voice AI. The main elements of this system are the user device, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[0923] User Device

[0924] A user terminal is a device such as a smartphone or computer. The user interacts with the system through a dedicated application or a web interface. Using this interface, the user enters the details of the task (e.g., canceling a subscription, changing a reservation) and provides the required information. For example, if the user wants to cancel a subscription, they enter their contract number and other required information and press the submit button. The terminal then sends this request to the server.

[0925] server

[0926] The server receives and analyzes the request sent from the user's device. This analysis allows it to understand the task content and retrieve the necessary information from the database. For example, the server may determine that the user wishes to cancel a subscription and gather the necessary data for cancellation. The server also issues instructions to the voice AI to execute the task. It passes the necessary information (e.g., contract number 123456) to the voice AI and instructs it to execute a specific task.

[0927] voice AI

[0928] The voice AI calls a designated phone number on behalf of the user, and after the caller answers, it uses natural language processing technology to continue the conversation. For example, when a customer support operator answers, they say, "Hello, I'd like to cancel my subscription for contract number 123456," and then proceeds with the necessary procedures. They ask questions and provide information as the conversation progresses.

[0929] Call outcome feedback

[0930] After the task is completed, the voice AI sends the call results to the server. The server analyzes the call results and records the necessary information (e.g., cancellation completed, reservation change completed). The server then notifies the user's device that the task has been completed. For example, the user's device may receive a notification saying, "Subscription cancellation has been completed."

[0931] Specific examples

[0932] Example 1: Canceling a subscription

[0933] 1. The user opens the application and selects "Cancel Subscription."

[0934] 2. Enter the contract number "123456" and other required information and submit.

[0935] 3. The server analyzes the request and the voice AI calls customer support.

[0936] 4. The voice AI says, "Please cancel contract number 123456," and continues the conversation.

[0937] 5. Once the cancellation procedure is complete, the server will notify the user that "Cancellation has been completed."

[0938] Example 2: Changing a restaurant reservation

[0939] 1. The user selects "Change Restaurant Reservation" in the application.

[0940] 2. Enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit.

[0941] 3. The server analyzes the request and the voice AI calls the restaurant.

[0942] 4. The voice AI says, "Please change reservation number 789012."

[0943] 5. After confirming the change, the server notifies the user that the reservation change has been completed.

[0944] Prompt Sentence Examples

[0945] Example 1: Subscription cancellation prompt

[0946] Prompt: Please cancel your subscription. Contract number is 123456.

[0947] Example 2: Prompt for changing a restaurant reservation

[0948] Prompt: I would like to change my restaurant reservation. The reservation number is 789012, and the desired date and time is October 10th at 7:00 PM.

[0949] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0950] Step 1:

[0951] The user enters the task

[0952] The user opens a dedicated application or web interface on their smartphone or computer and enters the details of the task. For example, if they want to cancel a subscription, they enter the contract number "123456" and add any other necessary information. When the user presses the "Submit" button, the device collects this input data and sends it to the server.

[0953] Input: Task details and contract number entered by the user

[0954] Output: Request data sent to the server (detailed task information)

[0955] Specific behavior:

[0956] Open the application.

[0957] Select "Cancel Subscription."

[0958] Enter the contract number "123456".

[0959] Enter other necessary information.

[0960] Press the send button.

[0961] Step 2:

[0962] The server receives and analyzes the request

[0963] The server receives the request sent from the device, analyzes its contents, identifies the type of task included in the request, and sends a query to the database to extract the necessary information (e.g., contract number), thereby obtaining the data necessary for the cancellation procedure.

[0964] Input: Request data sent from the device (task details)

[0965] Output: Analyzed task type and extracted required information (e.g. contract number)

[0966] Specific behavior:

[0967] Received request.

[0968] Analyze the request content.

[0969] Identify the type of task.

[0970] Obtain necessary information (contract number, etc.) from the database.

[0971] Step 3:

[0972] The server calls the voice AI

[0973] Based on the analysis results, the server calls the voice AI to execute the appropriate task. The server passes the necessary information (e.g., contract number) to the voice AI and sends an instruction to execute. For example, the server issues an instruction such as "Contact customer support to cancel the subscription for contract number 123456."

[0974] Input: Analysis results and necessary information (contract number, etc.)

[0975] Output: Instructions and task information passed to the voice AI

[0976] Specific behavior:

[0977] Call up the voice AI.

[0978] Pass the necessary information to the voice AI.

[0979] Send instructions to perform tasks.

[0980] Step 4:

[0981] Voice AI initiates and manages calls

[0982] The voice AI will call the specified phone number and start the call. Once the call begins, the voice AI will use natural language processing technology to proceed with the conversation. For example, when a customer support operator answers, they will say, "Hello, I would like to cancel my subscription for contract number 123456," and proceed with the process.

[0983] Input: Execution instructions and task information passed from the server

[0984] Output: Dialogue outcome and progress during the call

[0985] Specific behavior:

[0986] Dial the specified phone number.

[0987] Customer support responds.

[0988] Say, "Hello, I would like to cancel my subscription for contract number 123456."

[0989] Ask questions and provide information as necessary depending on the procedure.

[0990] Step 5:

[0991] The server records the results and sends feedback

[0992] After the voice AI completes the task, it sends the results of the call to the server. The server analyzes the call results and records the necessary information. Finally, the server notifies the user device that the task has been completed. For example, it may send a notification to the user saying, "Your subscription has been canceled."

[0993] Input: Call result sent from voice AI

[0994] Output: Parsed results and user notification

[0995] Specific behavior:

[0996] Receive call results from voice AI.

[0997] Analyze the results and record the necessary information.

[0998] Sends notifications to the user's device.

[0999] A message will be sent stating that your subscription has been cancelled.

[1000] (Application example 1)

[1001] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1002] Online shopping and food delivery services are rapidly becoming more popular, but there is a problem in that it takes time and effort for users to make changes to their orders or make inquiries. In particular, inquiries and changes made over the phone are prone to waiting times and communication problems. It also increases the burden on customer support, making it difficult to provide efficient service. To solve these issues, there is a need for a system that allows users to easily request changes to their orders and automates the entire process.

[1003] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1004] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the server to receive and analyze the request from the user; a means for the voice AI to autonomously execute the task and make a call; a means for the server to record the result of the call and notify the user; a means for the user to request an order change and specify the new order content; and a means for the server to contact the ordering party via the voice AI and request the order change. This allows the user to easily request an order change, and the entire process is automated, eliminating problems such as waiting times and communication errors, enabling efficient service provision.

[1005] "Means for users to input task details and necessary information" refers to an interface through which users input information necessary to perform a specific task, such as a smartphone application or a web interface.

[1006] "Means for the server to receive and analyze a request from the user" refers to a function that enables the server to receive request data sent from the user, analyze it, and identify the necessary processing.

[1007] "Means for voice AI to autonomously execute tasks and make calls" refers to a function that enables voice AI to autonomously execute designated tasks on behalf of the user via telephone.

[1008] "Means for the server to record the results of the call and notify the user" refers to a function that enables the server to record the results of the call made by voice AI in a database and notify the user of the results.

[1009] The "means by which a user requests an order change and specifies the new order details" is an interface by which a user can change the current order details and specify the new order details after the changes.

[1010] "Means for the server to contact the ordering party via voice AI and request a change to the order" is a function that allows the server to use voice AI to call the designated ordering party and request a change to the user's order.

[1011] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by the use of voice AI to automate food delivery order changes and inquiries. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[1012] System configuration

[1013] 1. User Device

[1014] A user terminal is a device, such as a smartphone or computer, that a user uses to enter change order details and submit requests. A dedicated application or web interface is provided through which the user interacts with the system.

[1015] 2. Server

[1016] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user. This process uses software such as natural language processing (NLP) libraries (e.g., Google Cloud Text-to-Speech, Amazon Lex) and communication libraries (e.g., Twilio).

[1017] 3. Voice AI

[1018] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and engage in natural conversations with the other party (e.g., restaurant staff). Voice AI calls use NLP technologies such as Google Cloud Text-to-Speech and Amazon Lex.

[1019] What the program does

[1020] User submits a request

[1021] The user uses a dedicated application or web interface to input the details of the order change (e.g., change to two pizzas). For example, if the user wants to "change the order," they enter the current order number and the new order details and press the submit button. The device then sends this request to the server.

[1022] The server parses the request

[1023] The server analyzes the received request and understands the task. It checks the necessary information (e.g., the current order number and the new order details) and prepares the voice AI. For example, the server determines that the user wants to change the order and prepares the necessary data for the change.

[1024] Voice AI performs tasks

[1025] The voice AI connected to the server dials a real phone number on behalf of the user. It dials the specified restaurant's phone number, and after the caller answers, the voice AI uses natural language processing technology to carry out the conversation. For example, the voice AI might say, "Hello, I'd like to change the order for order number 123456. The new order is for two pizzas." As the conversation progresses, the AI ​​provides necessary information to complete the task.

[1026] Call outcome feedback

[1027] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification saying "Order change completed" will be sent to the user's device. This allows the user to complete the necessary tasks without having to spend time on their own.

[1028] Specific examples

[1029] Example 1: Order change procedure

[1030] The user opens the application and selects "Change Order." They enter the current order number "123456" and the new order details "2 pizzas," and submit. The server analyzes the request, and the voice AI calls the restaurant. The voice AI responds by saying, "Please change order number 123456 to 2 pizzas," and continues the conversation. Once the change procedure is complete, the server notifies the user, "The order change has been completed."

[1031] Prompt Sentence Examples

[1032] "Imagine a food delivery application where a user requests a change to their order. When the user inputs that they would like to change the order for order number 123456 to two pizzas, write a program in which the voice AI automatically calls the restaurant and requests the order change. If the call is successful, please also include a function to notify the user that "the order change has been completed."

[1033] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1034] Step 1:

[1035] The user opens the smartphone application and accesses the screen for requesting an order change. The user enters the current order number and new order details (e.g., two pizzas) and presses the "Submit" button. The data entered is the order number and new order details. The device sends this request to the server as JSON-formatted data.

[1036] Step 2:

[1037] The server receives the request sent from the terminal. The received data includes the order number and new order details. The server analyzes the data and recognizes that the user wants to change the order. The server first retrieves the relevant order information from the database and locally checks the updated status of the order data based on the new order details. The output is the necessary information for preparing the voice AI.

[1038] Step 3:

[1039] The server provides the voice AI with details of the order change and contact information for the person on the other end of the call (e.g., restaurant staff). The input is the detailed order change information obtained in the previous step. The server uses a generative AI model (e.g., Google Cloud Text-to-Speech or Amazon Lex) to generate a dialogue script for the voice AI. The script includes a greeting at the start of the call, providing the order number, and explaining the new order details. The generated dialogue script is output.

[1040] Step 4:

[1041] The voice AI automatically calls the specified restaurant's phone number based on the dialogue script provided by the server. The input is the dialogue script and phone number. Natural language processing (NLP) is used to analyze and generate dialogue content, and ask the restaurant staff to change the order. Specifically, it says, "Hello, I'd like to change order number 123456. The new order is for two pizzas." The necessary information is presented appropriately during the call, and the dialogue continues until the order change is completed. The output is the call result.

[1042] Step 5:

[1043] Once the voice AI completes the call, it sends the results back to the server, including confirmation of the call's success or failure and any specific changes that were made. The input is the call result data. The server uses this information to update the database and record the changes. The output is the updated order information.

[1044] Step 6:

[1045] The server creates a notification for the user based on the call result. The input is the updated order information. It generates a message such as "Order changes completed" and sends it to the user's device. The output is a notification message.

[1046] Step 7:

[1047] The user terminal receives the notification message from the server and displays it on the screen. The user confirms the message "Order change completed." Through this process, the user can complete the order change without any hassle.

[1048] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1049] This invention is a system that allows users to efficiently perform specific tasks. It is characterized by utilizing voice AI to automate operations via voice calls, and by combining it with an emotion engine, it recognizes the user's emotions and adaptively changes the response. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine. Each element works together to autonomously process the user's tasks.

[1050] System configuration

[1051] 1. User Device

[1052] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[1053] 2. Server

[1054] The server receives and analyzes user requests, retrieves necessary information from a database, and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user.

[1055] 3. Voice AI

[1056] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[1057] 4. Emotion Engine

[1058] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue, enabling optimal dialogue according to the user's emotional state. It also learns from the user's past emotional data to improve the quality of future dialogue.

[1059] What the program does

[1060] User submits a request

[1061] The user enters the details of the task using a dedicated application or web interface. For example, if they want to cancel a subscription, they enter their contract number and other necessary information and press the submit button. The device then sends this request to the server.

[1062] The server parses the request

[1063] The server analyzes the received request and understands the task. It checks the necessary information (e.g., contract information and reservation number) and prepares the voice AI and emotion engine. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[1064] Voice AI and emotion engines perform tasks

[1065] The voice AI and emotion engine connected to the server work together to dial the phone on behalf of the user. After dialing the specified phone number and the caller answers, the voice AI uses natural language processing technology to carry out the dialogue. The emotion engine recognizes emotions from the user's voice and changes the dialogue content as necessary. For example, when the voice AI says, "Hello, I'd like to cancel my subscription for contract number 123456," if the emotion engine recognizes the user's anxiety or irritation, the voice AI adds a phrase that provides reassurance, such as, "Don't worry about that. We'll deal with it right away."

[1066] Call outcome feedback

[1067] After the task is completed, the server records the call result and sends a notification to the user. For example, a notification saying "Your subscription has been canceled" will be sent to the user's device. The emotional data collected by the emotion engine can be used to improve the quality of future conversations.

[1068] Specific examples

[1069] Example 1: Canceling a subscription

[1070] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[1071] Example 2: Changing a restaurant reservation

[1072] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server notifies the user, "The reservation change has been completed."

[1073] In this way, by using voice AI and an emotion engine, this system can provide optimal dialogue based on the user's emotions and complete tasks efficiently.

[1074] The processing flow will be explained below.

[1075] Step 1:

[1076] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[1077] Step 2:

[1078] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[1079] Step 3:

[1080] The device sends a user request to the server, including the type of task and the required details.

[1081] Step 4:

[1082] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[1083] Step 5:

[1084] The server analyzes the request, recognizes what the task is (e.g. cancel a subscription), retrieves the necessary information from the database, and issues instructions to the voice AI and emotion engine.

[1085] Step 6:

[1086] The server generates a dialogue script for the voice AI to execute a task and prepares the emotion engine. For example, the server prepares a script that says, "I would like to cancel contract number 123456."

[1087] Step 7:

[1088] The voice AI will dial the specified phone number, wait for the call to connect, and begin the conversation once the other party answers.

[1089] Step 8:

[1090] The emotion engine analyzes the user's voice to determine their emotions and provides that information to the voice AI, which then adapts the dialogue accordingly. For example, if the user sounds anxious, the AI ​​might add a phrase like, "Don't worry, we'll get back to you shortly."

[1091] Step 9:

[1092] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[1093] Step 10:

[1094] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[1095] Step 11:

[1096] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[1097] Step 12:

[1098] The server analyzes the results of the call and notifies the user of the outcome. The user receives a notification on their device confirming that the task was completed as planned, for example, a message saying "Subscription cancellation successful."

[1099] Step 13:

[1100] The emotion engine learns from the emotional data collected during processing and uses it to improve the quality of future interactions, resulting in a continuously improved user experience.

[1101] Example 2

[1102] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1103] Conventional voice communication systems require users to perform detailed operations themselves, making it difficult to complete tasks efficiently. Furthermore, they only provide a uniform response without considering the user's feelings, resulting in low user satisfaction. This has led to a demand for a system that can reduce the user's time and effort while also providing optimal responses tailored to individual feelings.

[1104] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1105] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the terminal to send the input information to the server; a means for the server to receive and analyze the request from the user; a means for the server to acquire data based on the received and analyzed information and issue preparation instructions to the voice AI and emotion engine; a means for the voice AI to autonomously execute the task and make a call; a means for the emotion engine to recognize emotions from the user's voice; and a means for the server to record the results of the call and notify the user. This allows the user to complete the task efficiently without detailed operations, and also enables optimal responses according to the user's emotions.

[1106] A "user" is someone who utilizes the system to perform a task.

[1107] A "terminal" is a device that a user uses to interact with the system, such as a smartphone or computer.

[1108] A "server" is a computer system that receives and analyzes requests from users.

[1109] "Voice AI" refers to artificial intelligence technology that generates voice and performs tasks autonomously.

[1110] The "emotion engine" is a technology that recognizes emotions from the user's voice and adaptively changes the content of the dialogue according to those emotions.

[1111] "Request" refers to a request including the content of a task that a user sends to a server via a terminal.

[1112] "Natural language processing technology" is a technology that allows computers to understand and generate human language.

[1113] A "call" is a means of communication through voice, and specifically refers to an exchange using a telephone.

[1114] "Data acquisition" refers to the process by which the server gathers the necessary information from internal databases and external sources.

[1115] "Notification" refers to a message sent by the server to inform the user of the outcome of a task.

[1116] This invention is a system that enables users to efficiently perform specific tasks, automating operations via voice calls by combining primarily voice AI and an emotion engine, and providing optimal responses according to the user's emotions. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine.

[1117] User Device

[1118] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system. For example, if a user wants to cancel a subscription, they open the dedicated application, enter their contract number and other required information, and press the submit button.

[1119] server

[1120] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user. The server analyzes the requests it receives, parses the request data, and extracts necessary information to understand the content of the task.

[1121] voice AI

[1122] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have a natural conversation with the other party (e.g., a customer support operator). For example, when the voice AI says, "Hello, I would like to cancel my subscription for contract number 123456," it uses natural language processing technology to smoothly guide the conversation.

[1123] Emotion Engine

[1124] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue. As a result, it is possible to have an optimal dialogue according to the user's emotional state. If the emotion engine detects the user's anxiety, the voice AI will add a reassuring phrase such as "Don't worry about that. We will respond immediately." It also learns from the user's past emotional data to improve the quality of future dialogue.

[1125] Specific examples

[1126] Example 1: Canceling a subscription

[1127] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server sends a notification to the user's device stating, "Cancellation completed."

[1128] Example 2: Changing a restaurant reservation

[1129] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server sends a notification to the user's device stating, "The reservation change has been completed."

[1130] Prompt Sentence Examples

[1131] Here are some example prompts to input to the generative AI model:

[1132] 1. Example of a user terminal prompt:

[1133] A user has requested to cancel their subscription using a dedicated application. What are their next steps?

[1134] 2. Server prompt example:

[1135] The server has received a user cancellation request, how can we parse the request and understand the task?

[1136] 3. Voice AI prompt example:

[1137] The voice AI needs to dial a specified phone number and proceed with canceling the subscription. How do you adapt the dialogue if the user is unsure?

[1138] 4. Emotion Engine Prompt Example:

[1139] The emotion engine detects anxiety in the user's voice. How can the voice AI change the dialogue to reassure the user?

[1140] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1141] Step 1:

[1142] The user enters the task.

[1143] The user enters task details such as cancellation or reservation change using a dedicated application or web interface. For example, they enter the contract number "123456," the reservation number "789012," and the desired change date and time "October 10th, 7:00 PM," and then press the send button. The contract number and reservation number are given as input, and JSON-formatted request data is generated as output.

[1144] Step 2:

[1145] The device sends a request to the server.

[1146] The terminal sends the information entered by the user to the server. Specifically, it sends the JSON formatted request data mentioned earlier to the server via an HTTP request. This request data is input and sent to the server as output.

[1147] Step 3:

[1148] The server receives and analyzes the request.

[1149] The server receives the request sent from the terminal and analyzes the request data. Specifically, it parses the JSON data and extracts necessary data such as the contract number and reservation information. The input to this analysis process is the request data, and the output is the analyzed task information.

[1150] Step 4:

[1151] The server retrieves the data.

[1152] The server retrieves the necessary contract or reservation information from an internal database. For example, retrieve the user's contract information associated with contract number "123456" from the database. The input to this database query is the contract or reservation number, and the output is any additional related information.

[1153] Step 5:

[1154] The server issues preparation instructions to the voice AI and emotion engine.

[1155] Based on the analysis results, the server instructs the voice AI and emotion engine to prepare for task execution. For example, it sends an instruction to "cancel contract number 123456." The input for this process is the analysis results and acquired data, and the output is a preparation instruction for the voice AI and emotion engine.

[1156] Step 6:

[1157] The server will have the voice AI start the call.

[1158] The server instructs the voice AI to dial the specified phone number. The voice AI receives this instruction and makes the call. The input of this process is the dial instruction and the phone number, and the output is the outgoing call status.

[1159] Step 7:

[1160] The voice AI begins the conversation.

[1161] When the other party responds, the voice AI begins a dialogue to accomplish a task. For example, say, "I would like to cancel the subscription for contract number 123456." The input for this process is the task content, and the output is the dialogue content.

[1162] Step 8:

[1163] The emotion engine recognizes emotions.

[1164] The emotion engine analyzes the user's voice and the voice of the other party to recognize emotions. For example, it can detect anxiety from the user's tone of voice. The input for this analysis is voice data, and the output is recognized emotional information.

[1165] Step 9:

[1166] The server records the call results.

[1167] Once the task is completed, the server receives feedback from the voice AI and records the call results. Specifically, the call content and the success or failure of the cancellation are stored in a database. The input of this process is the call feedback, and the output is the recorded data.

[1168] Step 10:

[1169] The server notifies the user of the results.

[1170] The server sends a notification to the user based on the recorded call result. For example, it sends a message to the user's device saying, "Subscription cancellation has been completed." The input of this notification is the recorded call result, and the output is a notification message to the user's device.

[1171] (Application example 2)

[1172] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1173] Modern security services face the challenge of requiring a lot of manual effort and time for users to efficiently complete specific tasks. Furthermore, they may not respond appropriately to the user's emotional state, which can increase anxiety and stress. This leads to a poor user experience and slows down the completion of security tasks, resulting in inefficiencies.

[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1175] In this invention, the server includes a means for the user to input the task content and necessary information, a means for the server to receive and analyze the request from the user, a means for the voice AI to autonomously execute the task and make a call, a means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI, and a means for the server to record the results of the call and notify the user. This makes it possible to complete security tasks efficiently and quickly while taking the user's emotions into consideration.

[1176] "Means for users to input task details and required information" refers to an interface that allows users to input detailed information about the task they wish to perform at that time and required data. This includes smartphone apps and web interfaces.

[1177] The "means for the server to receive and analyze the request from the user" refers to a system that receives a request sent from a user terminal, analyzes it, and understands the specific content of the task and the required information, thereby preparing for the next processing step.

[1178] "Means for voice AI to autonomously execute tasks and make calls" refers to a function in which voice AI automatically communicates with designated parties on behalf of the user and carries out tasks. It uses natural language processing technology to carry out dialogue and carry out tasks.

[1179] "Means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI" refers to a function that analyzes the emotional state of the user from their voice and input content, and the voice AI generates and changes appropriate dialogue content according to that emotion. This makes it possible to respond in a way that takes the user's emotions into consideration.

[1180] The "means for the server to record the results of the call and notify the user" refers to a system in which the server records the results of the call made by the voice AI and notifies the user of that information, thereby letting the user know whether the task was performed correctly.

[1181] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language, including speech recognition, semantic analysis, and dialogue generation.

[1182] A "dialogue script" is a set of defined phrases and sentences that a voice AI uses during a call, allowing the voice AI to have a natural conversation.

[1183] A "task" is a detailed description of the specific action or task the user is trying to accomplish, such as locking a bank account, calling an emergency service, or controlling a home security system.

[1184] "Call outcome" refers to the final outcome or status of a call made by a voice AI, including whether the task was successful or failed and its detailed status.

[1185] "Feedback" is the process of providing users with information about the outcome of a call, the progress of a task, etc., so that they know the status of their request.

[1186] This invention is a system that allows users to efficiently and flexibly execute specific tasks, and in particular, by combining voice AI and an emotion engine, it realizes dialogue that takes into account the user's emotions. This system is composed of the following elements:

[1187] 1. User Device

[1188] A user terminal is a device such as a smartphone or computer that allows users to input the task details and necessary information and submit requests. A dedicated application or web interface is provided through which users interact with the system.

[1189] 2. Server

[1190] The server receives the user's request, analyzes it, and retrieves the necessary information from the database. The server performs the following operations:

[1191] Analyze the voice data received from the user and understand the content of the task.

[1192] It obtains the necessary information and issues instructions to the voice AI and emotion engine.

[1193] The voice AI records the results of the tasks it performs and notifies the user.

[1194] 3. Voice AI

[1195] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue, using natural language processing technologies such as Google Cloud Speech-to-Text and Dialogflow. Voice AI does the following:

[1196] Automatically call the specified person.

[1197] Analyzes call content and generates natural dialogue.

[1198] 4. Emotion Engine

[1199] The emotion engine recognizes emotions from the user's voice and adaptively changes the dialogue content of the voice AI. For this purpose, IBM Watson Tone Analyzer is used. The emotion engine performs the following processes:

[1200] Analyzes the user's voice data and extracts their emotional state.

[1201] Depending on the emotion, instructions are sent to the voice AI to change the content of the dialogue.

[1202] Specific examples

[1203] Example 1: Locking a bank account

[1204] The user launches the application and commands it by voice, "Lock my bank account." The server receives the user's request and analyzes the voice data. Once the analysis is complete, the server instructs the voice AI to call the bank's customer support. If the emotion engine recognizes the user's nervousness, the voice AI will say, "Don't worry, we'll take care of it right away." Once the task is completed, the server notifies the user, "Your bank account has been successfully locked."

[1205] Prompt Sentence Examples

[1206] User input: "Lock my bank account"

[1207] Example prompt for a generative AI model: "Security-related task: User breathlessly commands, 'Lock my bank account.' The voice AI automatically calls a call center, and the emotion engine recognizes the tension and generates a reassuring response, 'Don't worry, we'll get back to you shortly.'"

[1208] In this way, the present invention enables a system that combines voice AI and an emotion engine to efficiently and flexibly execute tasks through dialogue that takes into consideration the user's emotions.

[1209] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1210] Step 1:

[1211] User Input

[1212] The user uses a smartphone or computer application to enter the task and required information by voice, for example, "lock my bank account."

[1213] Input: Voice data (e.g., "Lock my bank account")

[1214] Output: Audio data is sent to the system

[1215] How it works: The user launches a specific application and speaks a task into the microphone.

[1216] Step 2:

[1217] Receiving and analyzing requests (server)

[1218] The server receives the user's voice data, converts it into text using Google Cloud Speech-to-Text, and then analyzes the task content using Dialogflow.

[1219] Input: Voice data (e.g., "Lock my bank account")

[1220] Output: Text data (e.g. "Lock your bank account")

[1221] How it works: The server receives the voice data, converts it into text using speech recognition technology, and then analyzes the specific content of the task.

[1222] Step 3:

[1223] Emotion Recognition (Emotion Engine)

[1224] The server analyzes the user's emotional state from the voice data using IBM Watson Tone Analyzer.

[1225] Input: Audio data

[1226] Output: Emotional state data (e.g., tension, anxiety)

[1227] How it works: The server inputs the voice data into the emotion engine and analyzes the emotional state.

[1228] Step 4:

[1229] Task preparation (server)

[1230] Based on the analysis results, the server prepares the data necessary for the voice AI, such as obtaining the bank's customer support phone number and account information.

[1231] Input: Text data, emotional state data

[1232] Output: Data to be passed to the voice AI (e.g., phone number, account information)

[1233] Action: Prepares the data required to perform the task and generates a script for the voice AI to interact.

[1234] Step 5:

[1235] Task execution (voice AI)

[1236] The voice AI automatically calls the specified person and engages in a conversation. It uses Dialogflow to generate natural conversations and perform tasks. Based on data from the emotion engine, the voice AI changes the content of the conversation and speaks in a way that takes the user's emotions into consideration.

[1237] Input: Subject data, emotional state data

[1238] Output: Voice AI dialogue script, call results

[1239] How it works: The voice AI makes a call and proceeds with the conversation according to the specified content. For example, it generates and uses reassuring phrases such as "Don't worry, we'll respond right away."

[1240] Step 6:

[1241] Recording and notifying call results (server)

[1242] After the call is finished, the server records the result of the call and notifies the user, sending a message such as "Your bank account has been successfully locked."

[1243] Input: Call result data

[1244] Output: Record of call outcome, notification to user

[1245] How it works: The server stores the call results in a database and sends a notification to the user's device.

[1246] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1247] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1248] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1249] [Fourth embodiment]

[1250] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1251] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1252] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1253] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1254] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1255] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1256] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1257] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1258] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1259] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1260] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1261] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1262] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1263] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by its use of voice AI to automate operations via voice calls. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[1264] System configuration

[1265] 1. User Device

[1266] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[1267] 2. Server

[1268] The server receives and analyzes requests from users, retrieves necessary information from a database, and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user.

[1269] 3. Voice AI

[1270] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[1271] What the program does

[1272] User submits a request

[1273] The user uses a dedicated application or web interface to input the details of the task (e.g., canceling a subscription or changing a reservation). For example, if the user wants to cancel a subscription, they enter the contract number and other necessary information and press the submit button. The device then sends this request to the server.

[1274] The server parses the request

[1275] The server analyzes the received request and understands the task. It checks the necessary information (e.g., subscription contract information and reservation number) and prepares the voice AI. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[1276] Voice AI performs tasks

[1277] The voice AI connected to the server dials a real phone number on behalf of the user. After dialing the specified phone number and receiving the answer, the voice AI uses natural language processing technology to carry out the dialogue. For example, the voice AI might say, "Hello, I would like to cancel my subscription for contract number 123456." As the dialogue progresses, the AI ​​provides necessary information to complete the task.

[1278] Call outcome feedback

[1279] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification such as "Subscription cancellation completed" will be sent to the user's device. This allows the user to complete the necessary task without having to spend time on their own.

[1280] Specific examples

[1281] Example 1: Canceling a subscription

[1282] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI calls customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[1283] Example 2: Changing a restaurant reservation

[1284] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI calls the restaurant and says, "Please change reservation number 789012." After confirming the change, the server notifies the user that "the reservation change has been completed."

[1285] In this way, this system uses voice AI to efficiently perform tasks on behalf of the user and provide feedback on the results.

[1286] The processing flow will be explained below.

[1287] Step 1:

[1288] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[1289] Step 2:

[1290] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[1291] Step 3:

[1292] The device sends a user request to the server, including the type of task and the required details.

[1293] Step 4:

[1294] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[1295] Step 5:

[1296] The server analyzes the request and identifies the type of task (e.g., canceling a subscription). It then retrieves the information needed to execute the task (such as contract and reservation information) from the database.

[1297] Step 6:

[1298] The server calls the voice AI module and prepares to execute the task. It generates a dialogue script according to the task. For example, if it is a cancellation procedure, it prepares a script that says, "I would like to cancel contract number 123456."

[1299] Step 7:

[1300] The voice AI automatically dials the specified phone number, waits for the call to connect, and starts the conversation when the other party answers.

[1301] Step 8:

[1302] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[1303] Step 9:

[1304] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[1305] Step 10:

[1306] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[1307] Step 11:

[1308] The server analyzes the results of the call and notifies the user of the outcome, including information about task completion and next steps.

[1309] Step 12:

[1310] The user receives a notification on their device confirming that the task was completed as expected, for example, a message saying "Your subscription has been successfully canceled."

[1311] This process allows users to complete tasks efficiently without spending time on their own.

[1312] Example 1

[1313] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1314] In today's society, tasks users perform over the phone (e.g., canceling subscriptions or changing reservations) are time-consuming and cumbersome, necessitating efficient solutions. Furthermore, negotiating directly over the phone can be stressful for users, so technology is needed to simplify and automate these processes. Furthermore, conventional voice recognition systems struggle to deliver natural dialogue and appropriate responses, so improvements are needed.

[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1316] In this invention, the server includes a means for receiving and analyzing requests from users, a means for calling a voice AI and sending instructions including necessary information, and a means for recording the results of the call and notifying the user. This allows the user to efficiently automate tasks without having to make a call, reducing stress and enabling appropriate responses.

[1317] "User" means any person or entity that intends to use the System to perform a specific task.

[1318] A "task" is a specific purpose or action a user attempts to perform through the system (e.g., canceling a subscription or changing a reservation).

[1319] "Information" refers to data required to perform a task (e.g., contract number or reservation number).

[1320] "Means" refers to a method or combination of methods for achieving a specific function or operation.

[1321] "Server" refers to the central processing unit that receives and analyzes requests from users and controls the voice AI.

[1322] A "request" refers to a communication that describes the task or request sent by a user to a system.

[1323] "Analysis" refers to the operation by the server to interpret the contents of the request and determine the necessary processing.

[1324] "Voice AI" refers to artificial intelligence that makes voice calls on behalf of users and uses natural language processing technology to guide the conversation.

[1325] "Calling" refers to the operation in which the server sends instructions to the voice AI to perform a specific task.

[1326] "Call" refers to the voice communication that the voice AI has with customer support or other parties.

[1327] "Results" refers to information showing the status and results after the voice AI has performed a task.

[1328] "Notification" refers to communication from the server to the user informing them of task completion and progress.

[1329] "Natural language processing technology" refers to technology for understanding and generating human language.

[1330] A "dialogue script" refers to a scenario of a series of responses and questions that a voice AI uses during a call.

[1331] This invention is a system that allows users to efficiently perform specific tasks, and in particular automates operations via voice calls using voice AI. The main elements of this system are the user device, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[1332] User Device

[1333] A user terminal is a device such as a smartphone or computer. The user interacts with the system through a dedicated application or a web interface. Using this interface, the user enters the details of the task (e.g., canceling a subscription, changing a reservation) and provides the required information. For example, if the user wants to cancel a subscription, they enter their contract number and other required information and press the submit button. The terminal then sends this request to the server.

[1334] server

[1335] The server receives and analyzes the request sent from the user's device. This analysis allows it to understand the task content and retrieve the necessary information from the database. For example, the server may determine that the user wishes to cancel a subscription and gather the necessary data for cancellation. The server also issues instructions to the voice AI to execute the task. It passes the necessary information (e.g., contract number 123456) to the voice AI and instructs it to execute a specific task.

[1336] voice AI

[1337] The voice AI calls a designated phone number on behalf of the user, and after the caller answers, it uses natural language processing technology to continue the conversation. For example, when a customer support operator answers, they say, "Hello, I'd like to cancel my subscription for contract number 123456," and then proceeds with the necessary procedures. They ask questions and provide information as the conversation progresses.

[1338] Call outcome feedback

[1339] After the task is completed, the voice AI sends the call results to the server. The server analyzes the call results and records the necessary information (e.g., cancellation completed, reservation change completed). The server then notifies the user's device that the task has been completed. For example, the user's device may receive a notification saying, "Subscription cancellation has been completed."

[1340] Specific examples

[1341] Example 1: Canceling a subscription

[1342] 1. The user opens the application and selects "Cancel Subscription."

[1343] 2. Enter the contract number "123456" and other required information and submit.

[1344] 3. The server analyzes the request and the voice AI calls customer support.

[1345] 4. The voice AI says, "Please cancel contract number 123456," and continues the conversation.

[1346] 5. Once the cancellation procedure is complete, the server will notify the user that "Cancellation has been completed."

[1347] Example 2: Changing a restaurant reservation

[1348] 1. The user selects "Change Restaurant Reservation" in the application.

[1349] 2. Enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit.

[1350] 3. The server analyzes the request and the voice AI calls the restaurant.

[1351] 4. The voice AI says, "Please change reservation number 789012."

[1352] 5. After confirming the change, the server notifies the user that the reservation change has been completed.

[1353] Prompt Sentence Examples

[1354] Example 1: Subscription cancellation prompt

[1355] Prompt: Please cancel your subscription. Contract number is 123456.

[1356] Example 2: Prompt for changing a restaurant reservation

[1357] Prompt: I would like to change my restaurant reservation. The reservation number is 789012, and the desired date and time is October 10th at 7:00 PM.

[1358] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1359] Step 1:

[1360] The user enters the task

[1361] The user opens a dedicated application or web interface on their smartphone or computer and enters the details of the task. For example, if they want to cancel a subscription, they enter the contract number "123456" and add any other necessary information. When the user presses the "Submit" button, the device collects this input data and sends it to the server.

[1362] Input: Task details and contract number entered by the user

[1363] Output: Request data sent to the server (detailed task information)

[1364] Specific behavior:

[1365] Open the application.

[1366] Select "Cancel Subscription."

[1367] Enter the contract number "123456".

[1368] Enter other necessary information.

[1369] Press the send button.

[1370] Step 2:

[1371] The server receives and analyzes the request

[1372] The server receives the request sent from the device, analyzes its contents, identifies the type of task included in the request, and sends a query to the database to extract the necessary information (e.g., contract number), thereby obtaining the data necessary for the cancellation procedure.

[1373] Input: Request data sent from the device (task details)

[1374] Output: Analyzed task type and extracted required information (e.g. contract number)

[1375] Specific behavior:

[1376] Received request.

[1377] Analyze the request content.

[1378] Identify the type of task.

[1379] Obtain necessary information (contract number, etc.) from the database.

[1380] Step 3:

[1381] The server calls the voice AI

[1382] Based on the analysis results, the server calls the voice AI to execute the appropriate task. The server passes the necessary information (e.g., contract number) to the voice AI and sends an instruction to execute. For example, the server issues an instruction such as "Contact customer support to cancel the subscription for contract number 123456."

[1383] Input: Analysis results and necessary information (contract number, etc.)

[1384] Output: Instructions and task information passed to the voice AI

[1385] Specific behavior:

[1386] Call up the voice AI.

[1387] Pass the necessary information to the voice AI.

[1388] Send instructions to perform tasks.

[1389] Step 4:

[1390] Voice AI initiates and manages calls

[1391] The voice AI will call the specified phone number and start the call. Once the call begins, the voice AI will use natural language processing technology to proceed with the conversation. For example, when a customer support operator answers, they will say, "Hello, I would like to cancel my subscription for contract number 123456," and proceed with the process.

[1392] Input: Execution instructions and task information passed from the server

[1393] Output: Dialogue outcome and progress during the call

[1394] Specific behavior:

[1395] Dial the specified phone number.

[1396] Customer support responds.

[1397] Say, "Hello, I would like to cancel my subscription for contract number 123456."

[1398] Ask questions and provide information as necessary depending on the procedure.

[1399] Step 5:

[1400] The server records the results and sends feedback

[1401] After the voice AI completes the task, it sends the results of the call to the server. The server analyzes the call results and records the necessary information. Finally, the server notifies the user device that the task has been completed. For example, it may send a notification to the user saying, "Your subscription has been canceled."

[1402] Input: Call result sent from voice AI

[1403] Output: Parsed results and user notification

[1404] Specific behavior:

[1405] Receive call results from voice AI.

[1406] Analyze the results and record the necessary information.

[1407] Sends notifications to the user's device.

[1408] A message will be sent stating that your subscription has been cancelled.

[1409] (Application example 1)

[1410] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1411] Online shopping and food delivery services are rapidly becoming more popular, but there is a problem in that it takes time and effort for users to make changes to their orders or make inquiries. In particular, inquiries and changes made over the phone are prone to waiting times and communication problems. It also increases the burden on customer support, making it difficult to provide efficient service. To solve these issues, there is a need for a system that allows users to easily request changes to their orders and automates the entire process.

[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1413] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the server to receive and analyze the request from the user; a means for the voice AI to autonomously execute the task and make a call; a means for the server to record the result of the call and notify the user; a means for the user to request an order change and specify the new order content; and a means for the server to contact the ordering party via the voice AI and request the order change. This allows the user to easily request an order change, and the entire process is automated, eliminating problems such as waiting times and communication errors, enabling efficient service provision.

[1414] "Means for users to input task details and necessary information" refers to an interface through which users input information necessary to perform a specific task, such as a smartphone application or a web interface.

[1415] "Means for the server to receive and analyze a request from the user" refers to a function that enables the server to receive request data sent from the user, analyze it, and identify the necessary processing.

[1416] "Means for voice AI to autonomously execute tasks and make calls" refers to a function that enables voice AI to autonomously execute designated tasks on behalf of the user via telephone.

[1417] "Means for the server to record the results of the call and notify the user" refers to a function that enables the server to record the results of the call made by voice AI in a database and notify the user of the results.

[1418] The "means by which a user requests an order change and specifies the new order details" is an interface by which a user can change the current order details and specify the new order details after the changes.

[1419] "Means for the server to contact the ordering party via voice AI and request a change to the order" is a function that allows the server to use voice AI to call the designated ordering party and request a change to the user's order.

[1420] This invention is a system that allows users to efficiently perform specific tasks, and is characterized by the use of voice AI to automate food delivery order changes and inquiries. The main elements that make up this system are the user terminal, the server, and the voice AI. Each element works together to autonomously process the user's tasks.

[1421] System configuration

[1422] 1. User Device

[1423] A user terminal is a device, such as a smartphone or computer, that a user uses to enter change order details and submit requests. A dedicated application or web interface is provided through which the user interacts with the system.

[1424] 2. Server

[1425] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI. The server also records the results of the tasks performed by the voice AI and notifies the user. This process uses software such as natural language processing (NLP) libraries (e.g., Google Cloud Text-to-Speech, Amazon Lex) and communication libraries (e.g., Twilio).

[1426] 3. Voice AI

[1427] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and engage in natural conversations with the other party (e.g., restaurant staff). Voice AI calls use NLP technologies such as Google Cloud Text-to-Speech and Amazon Lex.

[1428] What the program does

[1429] User submits a request

[1430] The user uses a dedicated application or web interface to input the details of the order change (e.g., change to two pizzas). For example, if the user wants to "change the order," they enter the current order number and the new order details and press the submit button. The device then sends this request to the server.

[1431] The server parses the request

[1432] The server analyzes the received request and understands the task. It checks the necessary information (e.g., the current order number and the new order details) and prepares the voice AI. For example, the server determines that the user wants to change the order and prepares the necessary data for the change.

[1433] Voice AI performs tasks

[1434] The voice AI connected to the server dials a real phone number on behalf of the user. It dials the specified restaurant's phone number, and after the caller answers, the voice AI uses natural language processing technology to carry out the conversation. For example, the voice AI might say, "Hello, I'd like to change the order for order number 123456. The new order is for two pizzas." As the conversation progresses, the AI ​​provides necessary information to complete the task.

[1435] Call outcome feedback

[1436] After the task is completed, the server records the results of the call and sends a notification to the user. For example, a notification saying "Order change completed" will be sent to the user's device. This allows the user to complete the necessary tasks without having to spend time on their own.

[1437] Specific examples

[1438] Example 1: Order change procedure

[1439] The user opens the application and selects "Change Order." They enter the current order number "123456" and the new order details "2 pizzas," and submit. The server analyzes the request, and the voice AI calls the restaurant. The voice AI responds by saying, "Please change order number 123456 to 2 pizzas," and continues the conversation. Once the change procedure is complete, the server notifies the user, "The order change has been completed."

[1440] Prompt Sentence Examples

[1441] "Imagine a food delivery application where a user requests a change to their order. When the user inputs that they would like to change the order for order number 123456 to two pizzas, write a program in which the voice AI automatically calls the restaurant and requests the order change. If the call is successful, please also include a function to notify the user that "the order change has been completed."

[1442] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1443] Step 1:

[1444] The user opens the smartphone application and accesses the screen for requesting an order change. The user enters the current order number and new order details (e.g., two pizzas) and presses the "Submit" button. The data entered is the order number and new order details. The device sends this request to the server as JSON-formatted data.

[1445] Step 2:

[1446] The server receives the request sent from the terminal. The received data includes the order number and new order details. The server analyzes the data and recognizes that the user wants to change the order. The server first retrieves the relevant order information from the database and locally checks the updated status of the order data based on the new order details. The output is the necessary information for preparing the voice AI.

[1447] Step 3:

[1448] The server provides the voice AI with details of the order change and contact information for the person on the other end of the call (e.g., restaurant staff). The input is the detailed order change information obtained in the previous step. The server uses a generative AI model (e.g., Google Cloud Text-to-Speech or Amazon Lex) to generate a dialogue script for the voice AI. The script includes a greeting at the start of the call, providing the order number, and explaining the new order details. The generated dialogue script is output.

[1449] Step 4:

[1450] The voice AI automatically calls the specified restaurant's phone number based on the dialogue script provided by the server. The input is the dialogue script and phone number. Natural language processing (NLP) is used to analyze and generate dialogue content, and ask the restaurant staff to change the order. Specifically, it says, "Hello, I'd like to change order number 123456. The new order is for two pizzas." The necessary information is presented appropriately during the call, and the dialogue continues until the order change is completed. The output is the call result.

[1451] Step 5:

[1452] Once the voice AI completes the call, it sends the results back to the server, including confirmation of the call's success or failure and any specific changes that were made. The input is the call result data. The server uses this information to update the database and record the changes. The output is the updated order information.

[1453] Step 6:

[1454] The server creates a notification for the user based on the call result. The input is the updated order information. It generates a message such as "Order changes completed" and sends it to the user's device. The output is a notification message.

[1455] Step 7:

[1456] The user terminal receives the notification message from the server and displays it on the screen. The user confirms the message "Order change completed." Through this process, the user can complete the order change without any hassle.

[1457] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1458] This invention is a system that allows users to efficiently perform specific tasks. It is characterized by utilizing voice AI to automate operations via voice calls, and by combining it with an emotion engine, it recognizes the user's emotions and adaptively changes the response. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine. Each element works together to autonomously process the user's tasks.

[1459] System configuration

[1460] 1. User Device

[1461] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system.

[1462] 2. Server

[1463] The server receives and analyzes user requests, retrieves necessary information from a database, and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user.

[1464] 3. Voice AI

[1465] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have natural conversations with the other party (e.g., a call center operator).

[1466] 4. Emotion Engine

[1467] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue, enabling optimal dialogue according to the user's emotional state. It also learns from the user's past emotional data to improve the quality of future dialogue.

[1468] What the program does

[1469] User submits a request

[1470] The user enters the details of the task using a dedicated application or web interface. For example, if they want to cancel a subscription, they enter their contract number and other necessary information and press the submit button. The device then sends this request to the server.

[1471] The server parses the request

[1472] The server analyzes the received request and understands the task. It checks the necessary information (e.g., contract information and reservation number) and prepares the voice AI and emotion engine. For example, the server determines that the user wants to cancel a subscription and prepares the data necessary for cancellation.

[1473] Voice AI and emotion engines perform tasks

[1474] The voice AI and emotion engine connected to the server work together to dial the phone on behalf of the user. After dialing the specified phone number and the caller answers, the voice AI uses natural language processing technology to carry out the dialogue. The emotion engine recognizes emotions from the user's voice and changes the dialogue content as necessary. For example, when the voice AI says, "Hello, I'd like to cancel my subscription for contract number 123456," if the emotion engine recognizes the user's anxiety or irritation, the voice AI adds a phrase that provides reassurance, such as, "Don't worry about that. We'll deal with it right away."

[1475] Call outcome feedback

[1476] After the task is completed, the server records the call result and sends a notification to the user. For example, a notification saying "Your subscription has been canceled" will be sent to the user's device. The emotional data collected by the emotion engine can be used to improve the quality of future conversations.

[1477] Specific examples

[1478] Example 1: Canceling a subscription

[1479] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and submit. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server notifies the user, "Cancellation has been completed."

[1480] Example 2: Changing a restaurant reservation

[1481] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and submit. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server notifies the user, "The reservation change has been completed."

[1482] In this way, by using voice AI and an emotion engine, this system can provide optimal dialogue based on the user's emotions and complete tasks efficiently.

[1483] The processing flow will be explained below.

[1484] Step 1:

[1485] The user accesses a dedicated application or web interface and selects a task type, such as "cancel a subscription" or "change a restaurant reservation."

[1486] Step 2:

[1487] The user enters the details of the task, for example, the contract number and subscriber name if canceling a subscription. The user then clicks the "Submit" button to submit the request.

[1488] Step 3:

[1489] The device sends a user request to the server, including the type of task and the required details.

[1490] Step 4:

[1491] The server receives and parses the request from the user, first checking that the request is in the correct format and verifying that there is no missing information.

[1492] Step 5:

[1493] The server analyzes the request, recognizes what the task is (e.g. cancel a subscription), retrieves the necessary information from the database, and issues instructions to the voice AI and emotion engine.

[1494] Step 6:

[1495] The server generates a dialogue script for the voice AI to execute a task and prepares the emotion engine. For example, the server prepares a script that says, "I would like to cancel contract number 123456."

[1496] Step 7:

[1497] The voice AI will dial the specified phone number, wait for the call to connect, and begin the conversation once the other party answers.

[1498] Step 8:

[1499] The emotion engine analyzes the user's voice to determine their emotions and provides that information to the voice AI, which then adapts the dialogue accordingly. For example, if the user sounds anxious, the AI ​​might add a phrase like, "Don't worry, we'll get back to you shortly."

[1500] Step 9:

[1501] The voice AI requests a task from the other party based on the dialogue script, for example, "Hello, I would like to cancel the subscription for contract number 123456."

[1502] Step 10:

[1503] The voice AI analyzes the other person's response and continues the conversation according to the necessary confirmations, for example, by providing appropriate answers to questions from the other person.

[1504] Step 11:

[1505] After the task is completed, the voice AI sends the result of the call to the server, for example, recording the result as "Subscription cancellation completed."

[1506] Step 12:

[1507] The server analyzes the results of the call and notifies the user of the outcome. The user receives a notification on their device confirming that the task was completed as planned, for example, a message saying "Subscription cancellation successful."

[1508] Step 13:

[1509] The emotion engine learns from the emotional data collected during processing and uses it to improve the quality of future interactions, resulting in a continuously improved user experience.

[1510] Example 2

[1511] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1512] Conventional voice communication systems require users to perform detailed operations themselves, making it difficult to complete tasks efficiently. Furthermore, they only provide a uniform response without considering the user's feelings, resulting in low user satisfaction. This has led to a demand for a system that can reduce the user's time and effort while also providing optimal responses tailored to individual feelings.

[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1514] In this invention, the server includes: a means for the user to input the content of the task and necessary information; a means for the terminal to send the input information to the server; a means for the server to receive and analyze the request from the user; a means for the server to acquire data based on the received and analyzed information and issue preparation instructions to the voice AI and emotion engine; a means for the voice AI to autonomously execute the task and make a call; a means for the emotion engine to recognize emotions from the user's voice; and a means for the server to record the results of the call and notify the user. This allows the user to complete the task efficiently without detailed operations, and also enables optimal responses according to the user's emotions.

[1515] A "user" is someone who utilizes the system to perform a task.

[1516] A "terminal" is a device that a user uses to interact with the system, such as a smartphone or computer.

[1517] A "server" is a computer system that receives and analyzes requests from users.

[1518] "Voice AI" refers to artificial intelligence technology that generates voice and performs tasks autonomously.

[1519] The "emotion engine" is a technology that recognizes emotions from the user's voice and adaptively changes the content of the dialogue according to those emotions.

[1520] "Request" refers to a request including the content of a task that a user sends to a server via a terminal.

[1521] "Natural language processing technology" is a technology that allows computers to understand and generate human language.

[1522] A "call" is a means of communication through voice, and specifically refers to an exchange using a telephone.

[1523] "Data acquisition" refers to the process by which the server gathers the necessary information from internal databases and external sources.

[1524] "Notification" refers to a message sent by the server to inform the user of the outcome of a task.

[1525] This invention is a system that enables users to efficiently perform specific tasks, automating operations via voice calls by combining primarily voice AI and an emotion engine, and providing optimal responses according to the user's emotions. The main elements that make up this system are the user terminal, server, voice AI, and emotion engine.

[1526] User Device

[1527] A user terminal is a device, such as a smartphone or computer, that a user uses to enter task details and submit a request. A dedicated application or web interface is provided through which the user interacts with the system. For example, if a user wants to cancel a subscription, they open the dedicated application, enter their contract number and other required information, and press the submit button.

[1528] server

[1529] The server is responsible for receiving and analyzing requests from users. It also retrieves necessary information from a database and issues instructions to the voice AI and emotion engine. The server also records the results of tasks performed by the voice AI and notifies the user. The server analyzes the requests it receives, parses the request data, and extracts necessary information to understand the content of the task.

[1530] voice AI

[1531] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue. It uses natural language processing technology to analyze and generate dialogue content and have a natural conversation with the other party (e.g., a customer support operator). For example, when the voice AI says, "Hello, I would like to cancel my subscription for contract number 123456," it uses natural language processing technology to smoothly guide the conversation.

[1532] Emotion Engine

[1533] The emotion engine recognizes emotions from the user's voice and adaptively changes the content of the voice AI dialogue. As a result, it is possible to have an optimal dialogue according to the user's emotional state. If the emotion engine detects the user's anxiety, the voice AI will add a reassuring phrase such as "Don't worry about that. We will respond immediately." It also learns from the user's past emotional data to improve the quality of future dialogue.

[1534] Specific examples

[1535] Example 1: Canceling a subscription

[1536] The user opens the application and selects "Cancel subscription." They enter the contract number "123456" and other required information and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call customer support. The voice AI says, "Please cancel contract number 123456," and continues the conversation. If the emotion engine recognizes the user's anxiety, the voice AI adds, "Don't worry, the procedure is simple." Once the cancellation procedure is complete, the server sends a notification to the user's device stating, "Cancellation completed."

[1537] Example 2: Changing a restaurant reservation

[1538] The user selects "Change restaurant reservation" in the application. They enter the reservation number "789012" and the desired change date and time "October 10th, 7:00 PM" and press the send button. The device sends this request to the server. The server analyzes the request, and the voice AI and emotion engine call the restaurant and say, "Please change reservation number 789012." If the emotion engine recognizes the user's excitement, the voice AI adds a phrase such as, "Is that a special day? I'm looking forward to it." After confirming the change, the server sends a notification to the user's device stating, "The reservation change has been completed."

[1539] Prompt Sentence Examples

[1540] Here are some example prompts to input to the generative AI model:

[1541] 1. Example of a user terminal prompt:

[1542] A user has requested to cancel their subscription using a dedicated application. What are their next steps?

[1543] 2. Server prompt example:

[1544] The server has received a user cancellation request, how can we parse the request and understand the task?

[1545] 3. Voice AI prompt example:

[1546] The voice AI needs to dial a specified phone number and proceed with canceling the subscription. How do you adapt the dialogue if the user is unsure?

[1547] 4. Emotion Engine Prompt Example:

[1548] The emotion engine detects anxiety in the user's voice. How can the voice AI change the dialogue to reassure the user?

[1549] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1550] Step 1:

[1551] The user enters the task.

[1552] The user enters task details such as cancellation or reservation change using a dedicated application or web interface. For example, they enter the contract number "123456," the reservation number "789012," and the desired change date and time "October 10th, 7:00 PM," and then press the send button. The contract number and reservation number are given as input, and JSON-formatted request data is generated as output.

[1553] Step 2:

[1554] The device sends a request to the server.

[1555] The terminal sends the information entered by the user to the server. Specifically, it sends the JSON formatted request data mentioned earlier to the server via an HTTP request. This request data is input and sent to the server as output.

[1556] Step 3:

[1557] The server receives and analyzes the request.

[1558] The server receives the request sent from the terminal and analyzes the request data. Specifically, it parses the JSON data and extracts necessary data such as the contract number and reservation information. The input to this analysis process is the request data, and the output is the analyzed task information.

[1559] Step 4:

[1560] The server retrieves the data.

[1561] The server retrieves the necessary contract or reservation information from an internal database. For example, retrieve the user's contract information associated with contract number "123456" from the database. The input to this database query is the contract or reservation number, and the output is any additional related information.

[1562] Step 5:

[1563] The server issues preparation instructions to the voice AI and emotion engine.

[1564] Based on the analysis results, the server instructs the voice AI and emotion engine to prepare for task execution. For example, it sends an instruction to "cancel contract number 123456." The input for this process is the analysis results and acquired data, and the output is a preparation instruction for the voice AI and emotion engine.

[1565] Step 6:

[1566] The server will have the voice AI start the call.

[1567] The server instructs the voice AI to dial the specified phone number. The voice AI receives this instruction and makes the call. The input of this process is the dial instruction and the phone number, and the output is the outgoing call status.

[1568] Step 7:

[1569] The voice AI begins the conversation.

[1570] When the other party responds, the voice AI begins a dialogue to accomplish a task. For example, say, "I would like to cancel the subscription for contract number 123456." The input for this process is the task content, and the output is the dialogue content.

[1571] Step 8:

[1572] The emotion engine recognizes emotions.

[1573] The emotion engine analyzes the user's voice and the voice of the other party to recognize emotions. For example, it can detect anxiety from the user's tone of voice. The input for this analysis is voice data, and the output is recognized emotional information.

[1574] Step 9:

[1575] The server records the call results.

[1576] Once the task is completed, the server receives feedback from the voice AI and records the call results. Specifically, the call content and the success or failure of the cancellation are stored in a database. The input of this process is the call feedback, and the output is the recorded data.

[1577] Step 10:

[1578] The server notifies the user of the results.

[1579] The server sends a notification to the user based on the recorded call result. For example, it sends a message to the user's device saying, "Subscription cancellation has been completed." The input of this notification is the recorded call result, and the output is a notification message to the user's device.

[1580] (Application example 2)

[1581] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1582] Modern security services face the challenge of requiring a lot of manual effort and time for users to efficiently complete specific tasks. Furthermore, they may not respond appropriately to the user's emotional state, which can increase anxiety and stress. This leads to a poor user experience and slows down the completion of security tasks, resulting in inefficiencies.

[1583] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1584] In this invention, the server includes a means for the user to input the task content and necessary information, a means for the server to receive and analyze the request from the user, a means for the voice AI to autonomously execute the task and make a call, a means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI, and a means for the server to record the results of the call and notify the user. This makes it possible to complete security tasks efficiently and quickly while taking the user's emotions into consideration.

[1585] "Means for users to input task details and required information" refers to an interface that allows users to input detailed information about the task they wish to perform at that time and required data. This includes smartphone apps and web interfaces.

[1586] The "means for the server to receive and analyze the request from the user" refers to a system that receives a request sent from a user terminal, analyzes it, and understands the specific content of the task and the required information, thereby preparing for the next processing step.

[1587] "Means for voice AI to autonomously execute tasks and make calls" refers to a function in which voice AI automatically communicates with designated parties on behalf of the user and carries out tasks. It uses natural language processing technology to carry out dialogue and carry out tasks.

[1588] "Means for the emotion engine to recognize the user's emotions and adaptively change the dialogue content of the voice AI" refers to a function that analyzes the emotional state of the user from their voice and input content, and the voice AI generates and changes appropriate dialogue content according to that emotion. This makes it possible to respond in a way that takes the user's emotions into consideration.

[1589] The "means for the server to record the results of the call and notify the user" refers to a system in which the server records the results of the call made by the voice AI and notifies the user of that information, thereby letting the user know whether the task was performed correctly.

[1590] "Natural language processing technology" refers to a set of technologies that enable computers to understand, analyze, and generate human language, including speech recognition, semantic analysis, and dialogue generation.

[1591] A "dialogue script" is a set of defined phrases and sentences that a voice AI uses during a call, allowing the voice AI to have a natural conversation.

[1592] A "task" is a detailed description of the specific action or task the user is trying to accomplish, such as locking a bank account, calling an emergency service, or controlling a home security system.

[1593] "Call outcome" refers to the final outcome or status of a call made by a voice AI, including whether the task was successful or failed and its detailed status.

[1594] "Feedback" is the process of providing users with information about the outcome of a call, the progress of a task, etc., so that they know the status of their request.

[1595] This invention is a system that allows users to efficiently and flexibly execute specific tasks, and in particular, by combining voice AI and an emotion engine, it realizes dialogue that takes into account the user's emotions. This system is composed of the following elements:

[1596] 1. User Device

[1597] A user terminal is a device such as a smartphone or computer that allows users to input the task details and necessary information and submit requests. A dedicated application or web interface is provided through which users interact with the system.

[1598] 2. Server

[1599] The server receives the user's request, analyzes it, and retrieves the necessary information from the database. The server performs the following operations:

[1600] Analyze the voice data received from the user and understand the content of the task.

[1601] It obtains the necessary information and issues instructions to the voice AI and emotion engine.

[1602] The voice AI records the results of the tasks it performs and notifies the user.

[1603] 3. Voice AI

[1604] Voice AI makes actual voice calls on behalf of users and performs tasks through dialogue, using natural language processing technologies such as Google Cloud Speech-to-Text and Dialogflow. Voice AI does the following:

[1605] Automatically call the specified person.

[1606] Analyzes call content and generates natural dialogue.

[1607] 4. Emotion Engine

[1608] The emotion engine recognizes emotions from the user's voice and adaptively changes the dialogue content of the voice AI. For this purpose, IBM Watson Tone Analyzer is used. The emotion engine performs the following processes:

[1609] Analyzes the user's voice data and extracts their emotional state.

[1610] Depending on the emotion, instructions are sent to the voice AI to change the content of the dialogue.

[1611] Specific examples

[1612] Example 1: Locking a bank account

[1613] The user launches the application and commands it by voice, "Lock my bank account." The server receives the user's request and analyzes the voice data. Once the analysis is complete, the server instructs the voice AI to call the bank's customer support. If the emotion engine recognizes the user's nervousness, the voice AI will say, "Don't worry, we'll take care of it right away." Once the task is completed, the server notifies the user, "Your bank account has been successfully locked."

[1614] Prompt Sentence Examples

[1615] User input: "Lock my bank account"

[1616] Example prompt for a generative AI model: "Security-related task: User breathlessly commands, 'Lock my bank account.' The voice AI automatically calls a call center, and the emotion engine recognizes the tension and generates a reassuring response, 'Don't worry, we'll get back to you shortly.'"

[1617] In this way, the present invention enables a system that combines voice AI and an emotion engine to efficiently and flexibly execute tasks through dialogue that takes into consideration the user's emotions.

[1618] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1619] Step 1:

[1620] User Input

[1621] The user uses a smartphone or computer application to enter the task and required information by voice, for example, "lock my bank account."

[1622] Input: Voice data (e.g., "Lock my bank account")

[1623] Output: Audio data is sent to the system

[1624] How it works: The user launches a specific application and speaks a task into the microphone.

[1625] Step 2:

[1626] Receiving and analyzing requests (server)

[1627] The server receives the user's voice data, converts it into text using Google Cloud Speech-to-Text, and then analyzes the task content using Dialogflow.

[1628] Input: Voice data (e.g., "Lock my bank account")

[1629] Output: Text data (e.g. "Lock your bank account")

[1630] How it works: The server receives the voice data, converts it into text using speech recognition technology, and then analyzes the specific content of the task.

[1631] Step 3:

[1632] Emotion Recognition (Emotion Engine)

[1633] The server analyzes the user's emotional state from the voice data using IBM Watson Tone Analyzer.

[1634] Input: Audio data

[1635] Output: Emotional state data (e.g., tension, anxiety)

[1636] How it works: The server inputs the voice data into the emotion engine and analyzes the emotional state.

[1637] Step 4:

[1638] Task preparation (server)

[1639] Based on the analysis results, the server prepares the data necessary for the voice AI, such as obtaining the bank's customer support phone number and account information.

[1640] Input: Text data, emotional state data

[1641] Output: Data to be passed to the voice AI (e.g., phone number, account information)

[1642] Action: Prepares the data required to perform the task and generates a script for the voice AI to interact.

[1643] Step 5:

[1644] Task execution (voice AI)

[1645] The voice AI automatically calls the specified person and engages in a conversation. It uses Dialogflow to generate natural conversations and perform tasks. Based on data from the emotion engine, the voice AI changes the content of the conversation and speaks in a way that takes the user's emotions into consideration.

[1646] Input: Subject data, emotional state data

[1647] Output: Voice AI dialogue script, call results

[1648] How it works: The voice AI makes a call and proceeds with the conversation according to the specified content. For example, it generates and uses reassuring phrases such as "Don't worry, we'll respond right away."

[1649] Step 6:

[1650] Recording and notifying call results (server)

[1651] After the call is finished, the server records the result of the call and notifies the user, sending a message such as "Your bank account has been successfully locked."

[1652] Input: Call result data

[1653] Output: Record of call outcome, notification to user

[1654] How it works: The server stores the call results in a database and sends a notification to the user's device.

[1655] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1656] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1657] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1658] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1659] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1660] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1661] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1662] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1663] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1664] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1665] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1666] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1667] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1668] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1669] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1670] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1671] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1672] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1673] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1674] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1675] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1676] The following is further disclosed regarding the above embodiment.

[1677] (Claim 1)

[1678] A means for the user to input the content of the task and necessary information;

[1679] means for the server to receive and analyze requests from said users;

[1680] A means for voice AI to autonomously perform tasks and make calls;

[1681] means for the server to record the result of the call and notify the user;

[1682] A system including:

[1683] (Claim 2)

[1684] The system of claim 1, wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

[1685] (Claim 3)

[1686] The system according to claim 1, wherein the server includes means for generating a dialogue script for the voice AI according to the task content.

[1687] "Example 1"

[1688] (Claim 1)

[1689] A means for the user to input the content of the task and necessary information;

[1690] means for the server to receive and analyze requests from said users;

[1691] A means for the server to call the voice AI and send instructions including necessary information;

[1692] A means for voice AI to autonomously perform tasks and make calls;

[1693] means for the server to record the result of the call and notify the user;

[1694] A system including:

[1695] (Claim 2)

[1696] The system of claim 1, wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

[1697] (Claim 3)

[1698] The system according to claim 1, wherein the server includes means for generating a dialogue script for the voice AI according to the task content.

[1699] "Application Example 1"

[1700] (Claim 1)

[1701] A means for the user to input the content of the task and necessary information;

[1702] means for the server to receive and analyze requests from said users;

[1703] A means for voice AI to autonomously perform tasks and make calls;

[1704] means for the server to record the result of the call and notify the user;

[1705] A means for users to request order changes and specify new order details;

[1706] The server contacts the customer via voice AI to request a change to the order.

[1707] A system including:

[1708] (Claim 2)

[1709] The system of claim 1, wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

[1710] (Claim 3)

[1711] The system according to claim 1, wherein the server generates a dialogue script for a voice AI according to the task content and includes means for requesting order changes.

[1712] "Example 2: Combining Emotion Engines"

[1713] (Claim 1)

[1714] A means for the user to input the content of the task and necessary information;

[1715] means for transmitting the input information to a server by the terminal;

[1716] means for the server to receive and analyze requests from said users;

[1717] A means for acquiring data based on the information received and analyzed by the server and issuing preparation instructions to the voice AI and emotion engine;

[1718] A means for voice AI to autonomously perform tasks and make calls;

[1719] a means for the emotion engine to recognize emotions from the user's voice;

[1720] means for the server to record the result of the call and notify the user;

[1721] A system including:

[1722] (Claim 2)

[1723] The system of claim 1, wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

[1724] (Claim 3)

[1725] 2. The system of claim 1, wherein the emotion engine includes means for modifying dialogue content in response to a user's emotion.

[1726] "Application example 2 when combining emotion engines"

[1727] (Claim 1)

[1728] A means for the user to input the content of the task and necessary information;

[1729] means for the server to receive and analyze requests from said users;

[1730] A means for voice AI to autonomously perform tasks and make calls;

[1731] The emotion engine recognizes the user's emotions and adaptively changes the dialogue content of the voice AI.

[1732] means for the server to record the result of the call and notify the user;

[1733] A system including:

[1734] (Claim 2)

[1735] The system of claim 1, wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

[1736] (Claim 3)

[1737] The system of claim 1, wherein the server generates a dialogue script for the voice AI according to the task content, and the emotion engine includes means for adaptively changing the dialogue content based on the user's emotions. [Explanation of symbols]

[1738] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for the user to input the content of the task and necessary information; means for the server to receive and analyze requests from said users; A means for voice AI to autonomously perform tasks and make calls; means for the server to record the result of the call and notify the user; A system including:

2. The system of claim 1 , wherein the voice AI includes means for conducting dialogue using natural language processing techniques.

3. 2. The system according to claim 1, wherein the server includes means for generating a dialogue script for a voice AI according to task content.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A