system

The integration of IVR with voice-based generative AI in a centralized system addresses the limitations of conventional IVR systems by providing flexible and immediate responses to user inquiries, enhancing user experience and service integration.

JP7808661B2Active Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024163732
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-09-20
Filing Date
2024-09-20
Publication Date
2026-01-29
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

Conventional IVR systems lack flexibility in responding to user questions and requests, require multiple systems for various services, and fail to integrate voice data with generative AI for efficient responses.

Method used

A system integrating IVR with voice-based generative AI, providing comprehensive services such as number provision, voice guidance, and generative AI integration, including restaurant reservations, directions, and phone orders, using a centralized system for voice responses, SMS notifications, and text-based customer interactions.

Benefits of technology

Enables flexible and immediate responses to user inquiries, improves user experience by addressing multiple service needs through a unified system, and enhances interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808661000001
    Figure 0007808661000001
  • Figure 0007808661000002
    Figure 0007808661000002
  • Figure 0007808661000003
    Figure 0007808661000003
Patent Text Reader

Abstract

To provide a system.SOLUTION: The system includes means for receiving a telephone call from a user, means for converting audio data of the user into a text, means for sending the converted text to a generative AI model as a prompt sentence, means for having the generative AI model generate a response, means for converting the generated response text into audio, and means for providing the audio to the user as voice guidance.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional interactive voice response (IVR) systems operate based on fixed scripts, making it difficult to flexibly respond to user questions and requests. In addition, since there is no centralized system for providing various services and collecting information, users must use multiple systems, which is inconvenient. [Means for solving the problem]

[0005] To solve this problem, we offer a system that links an IVR with a voice-based generative AI. This system is offered for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and generative AI integration. In particular, generative AI can handle restaurant reservations, directions to the restaurant, business hours confirmation, and phone orders. By linking a pre-prepared response database with the voice-based generative AI, it is possible to provide voice responses, SMS notifications, and questions for confirmation, and store customer responses as text. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 2 is a sequence diagram showing a flow of processing in the data processing system according to the first embodiment of the first form example. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Embodiment 1. [Figure 13] FIG. 10 is a sequence diagram showing a processing flow of a data processing system in a second embodiment of the second form example. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Embodiment Example 2. [Figure 15] FIG. 10 is a sequence diagram showing the flow of processing in a data processing system according to a third embodiment of the third embodiment. [Figure 16] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Embodiment 3. [Figure 17] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the first embodiment of the first form example when an emotion engine is combined. [Figure 18] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the second embodiment of the second form example when an emotion engine is combined. [Figure 20] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the third embodiment of the third form example when an emotion engine is combined. [Figure 22] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0007] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0008] First, the terms used in the following description will be explained.

[0009] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)).

[0010] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0011] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0012] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0014] [First embodiment]

[0015] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0019] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0022] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0023] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0026] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0027] "Example 1"

[0028] One embodiment of the present invention is a system that integrates an interactive voice response (IVR) with a voice-based AI generator. This system is offered for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and AI generator integration functions. Specifically, when a user calls the system, the system responds to the user's questions and requests using the voice-based AI generator. For example, if a user asks, "What are your business hours?", the system responds using the voice-based AI generator, "Business hours are from 9:00 AM to 5:00 PM."

[0029] "Example 2"

[0030] Another embodiment of the present invention is a system that can use generative AI to solve restaurant reservations, directions to the restaurant, confirmation of business hours, and phone orders. Specifically, when a user requests to "make a restaurant reservation," the system uses a voice-based generative AI to ask, "How many people would you like to make a reservation for, and what time would you like to make a reservation for?" and makes the reservation based on the user's response.

[0031] "Example 3"

[0032] In yet another embodiment of the present invention, a system is provided that links a prepared answer database with a voice-based generation AI to provide voice responses, SMS notifications, and ask questions for confirmation, and then saves the customer's responses as text. Specifically, when a user requests to "place an order," the system uses the voice-based generation AI to ask "What would you like to order?", accepts the order based on the user's response, and saves the contents as text.

[0033] The processing flow of each embodiment will be described below.

[0034] "Example 1"

[0035] Step 1: A user calls the system.

[0036] Step 2: The system responds to the user's questions and requests using a voice-generative AI.

[0037] Step 3: The user asks, "What are your business hours?"

[0038] Step 4: The system uses a voice-generating AI to respond, "Business hours are from 9:00 AM to 5:00 PM."

[0039] "Example 2"

[0040] Step 1: The user requests to make a reservation.

[0041] Step 2: The system uses a voice-generating AI to ask, "How many people would you like to make a reservation for, and what time would you like to start?"

[0042] Step 3: The system makes the reservation based on the user's answers.

[0043] "Example 3"

[0044] Step 1: The user requests to place an order.

[0045] Step 2: The system uses a voice-generative AI to ask, "What would you like to order?" Step 3: The system accepts the order based on the user's answer.

[0046] Step 4: The system saves the order details as text.

[0047] Example 1

[0048] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0049] Conventional automated voice response systems have had difficulty responding appropriately and quickly to a variety of user questions and requests. Furthermore, they lacked the ability to convert voice data into text and integrate it with generative AI models, resulting in a poor user experience. Furthermore, certain industries, such as restaurants, lacked systems that could address specific needs, such as making reservations or checking opening hours.

[0050] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0051] In this invention, the server includes means for receiving a call from a user, means for providing voice guidance, means for converting the user's voice data into text, means for transmitting the converted text to the generative AI model as a prompt sentence, means for the generative AI model to generate an answer, means for converting the generated text into speech, and means for transmitting the converted speech data to the user, thereby enabling appropriate and prompt responses to a variety of user questions and requests.

[0052] The "means for receiving a call from a user" is a function that allows the system to receive a call made by a user and start a call session.

[0053] The "means for providing voice guidance" is a function for reproducing a voice message to the user and prompting the user to perform the next operation or input.

[0054] The "means for converting user's voice data into text" is a function for converting the voice uttered by the user into character data.

[0055] "Means for sending the converted text to the generative AI model as a prompt sentence" refers to a function that inputs text data converted from speech into the generative AI model and gives instructions for generating an appropriate response.

[0056] The "means by which the generative AI model generates an answer" is the function by which the generative AI model generates an appropriate answer based on the prompt sentence.

[0057] "Means for converting generated text into speech" refers to a function for converting text data generated by a generative AI model into speech data.

[0058] The "means for transmitting converted voice data to the user" is a function for transmitting data converted into voice to the user so that the user can listen to the voice.

[0059] This invention is a system that combines an interactive voice response (IVR) and a generative AI model, and aims to respond appropriately and quickly to calls from users. A specific embodiment of this system is described below.

[0060] System configuration

[0061] The server system is configured using the following hardware and software.

[0062] Hardware: Servers, communication equipment, audio input / output devices

[0063] Software: Twilio API, Google® Cloud Speech-to-Text API, Google Cloud Text-to-Speech API, generative AI models (e.g., OpenAI®'s GPT-3®)

[0064] Program processing

[0065] The server receives a call from the user using the Twilio API. When the call connects, the server provides a voice prompt, such as "How can I help you?"

[0066] When a user speaks a question or request, the server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then sent to the generative AI model as a prompt.

[0067] The generative AI model generates appropriate answers to user questions. For example, if a user asks, "What are your business hours?", the generative AI model will respond with, "Business hours are from 9:00 AM to 5:00 PM."

[0068] The generated text response is then converted back into audio. For this, the server uses the Google Cloud Text-to-Speech API. The converted audio data is then sent to the user via the Twilio API.

[0069] Specific examples

[0070] For example, the following shows the processing when a user asks, "What are your business hours?"

[0071] 1. A user calls the system.

[0072] 2. The server receives the call using the Twilio API and provides a voice prompt saying, "Please tell us your business."

[0073] 3. A user asks, "What are your business hours?"

[0074] 4. The server converts the speech to text using the Google Cloud Speech-to-Text API.

[0075] 5. The converted text, "What are your business hours?", is input as a prompt to the generative AI model.

[0076] 6. The generative AI model generates the answer, "Business hours are 9:00 AM to 5:00 PM."

[0077] 7. The server converts the generated text response into audio using the Google Cloud Text-to-Speech API.

[0078] 8. The server sends the converted voice data to the user via the Twilio API.

[0079] Prompt Sentence Examples

[0080] "A user is asking about business hours. Generate an appropriate answer based on the following text: 'What are your business hours?'"

[0081] In this way, the system operates in cooperation with the server, terminals, and users.

[0082] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0083] Step 1:

[0084] User makes a call

[0085] The user calls the provided phone number. The input is the user making the call, and the output is the call connecting to the server.

[0086] Step 2:

[0087] The server receives the call

[0088] The server receives a call from the user using a communication API. The input is a call connection from the user, and the output is the start of a call session. The server then prepares to manage the call session.

[0089] Step 3:

[0090] The server provides audio guidance

[0091] The server provides voice guidance to the user through a communication API. The input is the start of a call session, and the output is the playback of a voice guidance message, such as "Please tell us your business."

[0092] Step 4:

[0093] Users speak their questions or requests

[0094] The user follows the voice guidance to input questions or requests by voice. The input is the user's voice, and the output is the transmission of voice data to the server. For example, a question might be, "What are your business hours?"

[0095] Step 5:

[0096] The server converts the audio data into text

[0097] The server uses a speech recognition API to convert the user's voice data into text. The input is the user's voice data, and the output is the converted text data. For example, the generated text is "What are your business hours?"

[0098] Step 6:

[0099] The server sends a prompt to the generative AI model

[0100] The server sends the converted text to the generative AI model as a prompt. The input is the converted text data, and the output is the prompt sent to the generative AI model. The prompt is, "The user is asking about business hours. Please generate an appropriate answer based on the following text: 'What are your business hours?'"

[0101] Step 7:

[0102] Generative AI models generate answers

[0103] A generative AI model generates an appropriate answer based on a prompt. The input is the prompt, and the output is a generated text answer. For example, the generated answer might be, "Business hours are from 9:00 AM to 5:00 PM."

[0104] Step 8:

[0105] The server converts the generated text to speech

[0106] The server converts the generated text response into speech using a speech synthesis API. The input is the generated text response, and the output is the converted speech data. For example, the speech generated is "Business hours are from 9:00 AM to 5:00 PM."

[0107] Step 9:

[0108] The server sends the audio data to the user

[0109] The server sends the converted voice data to the user through a communication API. The input is the converted voice data, and the output is a voice transmission to the user. The user can hear the answer over the phone.

[0110] (Application example 1)

[0111] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0112] Conventional automated voice response systems could only provide pre-set, fixed responses, making it difficult to flexibly respond to a variety of user questions and requests. Furthermore, they were unable to respond immediately to security-related questions and requests, leaving users uneasy. This resulted in a poor user experience and limited the system's value.

[0113] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0114] In this invention, the server is a system that links an interactive voice response (IVR) with a voice-based generative AI, and is provided for a fixed fee. It includes a means for providing a number, voice guidance, generative AI linkage functions, etc. in an integrated manner, a means for speech recognition, a means for generating responses to user questions using a generative AI model, and a means for speech synthesis of the generated responses. This makes it possible to respond flexibly and immediately to a variety of user questions and requests, and in particular to quickly respond to questions and requests related to security, thereby alleviating user anxiety and improving the user experience.

[0115] An "Interactive Voice Response (IVR)" is a system that automatically responds with voice when a user calls and provides appropriate information based on the user's input.

[0116] "Voice generative AI" is an artificial intelligence technology that analyzes the user's voice input and generates an appropriate response.

[0117] "Providing a number" means providing a telephone number for a user to access.

[0118] "Voice guidance" is a function that provides voice guidance and instructions to the user.

[0119] The "generative AI integration function" is a function that links the voice-based generative AI with other systems and databases.

[0120] "Speech recognition" is a technology that converts a user's voice into text.

[0121] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on user input.

[0122] "Speech synthesis" is a technology that converts text into speech.

[0123] A "security question or request" is a question or request where a user requests information regarding the status or configuration of a security system.

[0124] The "answer database" is a database that stores prepared questions and their answers.

[0125] "Message notification" is a function that notifies users of information via text message.

[0126] "Questions for confirmation" is a function that asks the user questions about matters that need to be confirmed.

[0127] "Text save" is a function that saves the user's answers in text format.

[0128] As an embodiment of the present invention, a security assistant system will be described as an example. In this system, when a user inputs security-related questions or requests by voice, a voice-based AI system responds immediately.

[0129] Hardware and software used

[0130] Hardware: Smartphone (microphone, speaker)

[0131] software:

[0132] speech_recognition library: Used to perform speech recognition.

[0133] pyttsx3 library: Used to perform speech synthesis.

[0134] openai library: Uses generative AI models using the GPT-3 API.

[0135] System Operation

[0136] 1. Voice recognition: A user speaks a security question or request into the smartphone microphone, for example, "What is the status of my home security system?"

[0137] 2. Speech to text conversion: Uses the speech_recognition library to convert the user's speech into text.

[0138] 3. Response generation using a generative AI model: The converted text is sent to GPT-3 as a prompt to generate an appropriate response. An example of a prompt is "Security system question: What is the status of my home security system?"

[0139] 4. Speech synthesis: The generated response is converted into speech using the pyttsx3 library and transmitted back to the user through the smartphone speaker.

[0140] Specific examples

[0141] When a user asks, "What's the status of my home security system?" the system works as follows:

[0142] 1. The smartphone microphone captures the user's voice.

[0143] 2. The speech_recognition library converts the speech to text, generating the text "What is the status of my home security system?"

[0144] 3. The generated text is sent as a prompt to GPT-3, which generates an appropriate response, such as "All sensors are currently working properly."

[0145] 4. The pyttsx3 library converts the generated response into speech and delivers it back to the user through the smartphone speaker.

[0146] In this way, the security assistant system can respond immediately to users' security questions and requests, alleviating their concerns and improving their experience.

[0147] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0148] Step 1:

[0149] The user speaks their security question or request into the smartphone's microphone.

[0150] Input: User's voice

[0151] Output: Audio data captured by a microphone

[0152] Specific Action: A user says, "What's the status of my home security system?"

[0153] Step 2:

[0154] The device uses the speech_recognition library to convert the captured audio data into text.

[0155] Input: Audio data

[0156] Output: Text data

[0157] What it does: The speech recognition engine analyzes the voice data and generates text such as "What is the status of my home security system?"

[0158] Step 3:

[0159] The device sends the generated text as a prompt to GPT-3, which generates an appropriate response.

[0160] Input: Text data (prompt sentence)

[0161] Output: Response text

[0162] What it does: Sends the prompt "Security System Question: What's the status of my home security system?" to the GPT-3 API and receives the response "Currently, all sensors are working properly."

[0163] Step 4:

[0164] The device converts the generated response text into speech using the pyttsx3 library.

[0165] Input: Response text

[0166] Output: Audio data

[0167] What happens: The speech synthesis engine converts the text "All sensors are currently working properly" into speech data.

[0168] Step 5:

[0169] The device responds to the user with voice data via the smartphone speaker.

[0170] Input: Audio data

[0171] Output: The audio the user hears

[0172] What happens: The smartphone speaker will play a voice message saying "All sensors are currently working properly."

[0173] In this way, the security assistant system can respond immediately to the user's security questions and requests.

[0174] Example 2

[0175] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] Conventional automated voice response systems have difficulty responding flexibly and quickly to user requests, and efficient responses are required, especially for complex requests such as making restaurant reservations or confirming opening hours. Furthermore, there is a lack of technology to accurately understand users' voice requests and generate appropriate responses, making it difficult to improve the user experience.

[0177] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting a user's voice request, means for converting voice to text, means for generating and transmitting a prompt sentence to the generative AI model, means for receiving a response from the generative AI model, means for converting the text response to voice data, means for transmitting the voice data to the user, means for accepting the user's answer, means for registering reservation information, means for generating a reservation confirmation message and converting it into voice data, and means for transmitting the voice data to the user. This makes it possible to respond quickly and accurately to user requests and efficiently handle complex requests such as making a reservation at a restaurant or confirming business hours.

[0178] The "means for accepting a user's voice request" refers to a device or software that recognizes the voice uttered by the user and inputs the content of that voice into the system.

[0179] A "speech-to-text converter" is a device or software that uses speech recognition technology to convert a user's speech into textual information.

[0180] A "means for generating and sending prompts to a generative AI model" is a device or software for generating appropriate questions or instructions based on a user's request and sending them to a generative AI model.

[0181] A "means for receiving a response from a generative AI model" is a device or software for receiving a response returned from a generative AI model.

[0182] A "means for converting text responses into audio data" is a device or software for converting text responses from a generative AI model into audio data.

[0183] "Means for transmitting voice data to a user" refers to a device or software for transmitting the generated voice data to a user's terminal and conveying it to the user.

[0184] "Means for accepting user responses" refers to a device or software that recognizes additional voice information provided by the user and incorporates that information into the system.

[0185] The "means for registering reservation information" refers to a device or software for registering reservation information received from a user in a database or the like.

[0186] The "means for generating a reservation confirmation message and converting it into voice data" refers to a device or software that generates a message to notify the user that the reservation has been confirmed and converts it into voice data.

[0187] The "means for transmitting voice data to the user" refers to a device or software for transmitting the generated voice data to the user's terminal and conveying the contents of the reservation confirmation to the user.

[0188] This invention is a system that uses a generative AI model to solve restaurant reservations, restaurant directions, checking business hours, and phone orders. This system accepts a user's voice request, converts it into text, generates and sends a prompt to the generative AI model, and then generates an appropriate response to provide to the user.

[0189] Hardware and software used

[0190] Hardware: Servers, user devices (smartphones, tablets, PCs, etc.)

[0191] Software: Generative AI models (e.g., large-scale language models), speech recognition software (e.g., speech recognition engines), and speech synthesis software (e.g., speech synthesis engines).

[0192] Specific operation of the system

[0193] 1. Acceptance of user requests

[0194] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[0195] Device: Uses speech recognition software to convert your speech into text.

[0196] Terminal: Sends the converted text to the server.

[0197] 2. Prompt generation of generative AI models

[0198] Server: Parses the received text and generates prompts to send to the generative AI model.

[0199] Server: Generate an example prompt: "A user would like to make a reservation. How many people would you like to book and what time would you like to make the reservation for?"

[0200] 3. Response generation using generative AI models

[0201] Server: Sends the generated prompts to the generative AI model.

[0202] Generative AI model: Generates appropriate responses based on prompts.

[0203] Generative AI model: Generates an example response: "How many people would you like to make a reservation for, and what time?"

[0204] 4. Voice synthesis of responses

[0205] Server: Passes the generated text response to speech synthesis software to generate audio data.

[0206] 5. Sending a response to the user

[0207] Server: Sends the generated voice data to the user device.

[0208] Terminal: Plays the audio data and communicates the response to the user.

[0209] 6. Acceptance of user responses

[0210] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[0211] Device: Uses speech recognition software to convert your speech into text.

[0212] Terminal: Sends the converted text to the server.

[0213] 7. Confirmation of reservation

[0214] Server: Receives the user's response and registers the information in the reservation system.

[0215] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[0216] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We are waiting for you with XX guests on XX / XX / XX at XX time."

[0217] 8. Sending a confirmation message

[0218] Server: The confirmation message is converted into voice using text-to-speech software.

[0219] Text-to-speech software: converts text into audio data.

[0220] Server: Sends the generated voice data to the user device.

[0221] Terminal: Plays audio data to inform the user that the reservation is complete.

[0222] Specific examples

[0223] User request: "I want to make a reservation to visit the store."

[0224] Prompt to generative AI model: "A user wants to make a reservation. How many people would like to come and what time would they like to make the reservation?"

[0225] The generative AI model responds: "How many people would you like to book and what time would you like to book?"

[0226] User's answer: "Two people, starting tomorrow at 7pm."

[0227] Reservation confirmation message: "Your reservation is complete. We look forward to seeing you tomorrow at 7pm for two people."

[0228] In this way, the system uses a generative AI model to generate an appropriate response to a user's request and confirm the reservation.

[0229] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0230] Step 1: Accepting a user request

[0231] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[0232] Input: User's voice request

[0233] Device: Uses speech recognition software to convert your speech into text.

[0234] Data processing: Converting voice data into text data

[0235] Output: "I'd like to make an appointment"

[0236] Terminal: Sends the converted text to the server.

[0237] Step 2: Prompt generation for generative AI models

[0238] Server: Parses the received text and generates prompts to send to the generative AI model.

[0239] Input: "I'd like to make an appointment"

[0240] Data Calculation: Text Analysis and Prompt Generation

[0241] Output: A prompt saying "A user wants to make an appointment. How many people would like to book and what time would you like to book?"

[0242] Server: Sends prompts to the generative AI model.

[0243] Step 3: Generate a response using a generative AI model

[0244] Server: Sends the generated prompts to the generative AI model.

[0245] Input: The prompt "A user wants to make a reservation. How many people and what time would you like to book?"

[0246] Generative AI model: Generates appropriate responses based on prompts.

[0247] Data Calculation: Prompt-Based Response Generation

[0248] Output: Response "How many people would you like to book and what time?"

[0249] Generative AI model: Sends the response to the server.

[0250] Step 4: Speech synthesis to vocalize the response

[0251] Server: Passes the generated text response to speech synthesis software to generate audio data.

[0252] Input: Text response "How many people would you like to book and what time?"

[0253] Text-to-speech software: converts text into audio data.

[0254] Data processing: Convert text data into audio data

[0255] Output: Audio data

[0256] Server: Sends the generated voice data to the user device.

[0257] Step 5: Send a response to the user

[0258] Server: Sends the generated voice data to the user device.

[0259] Input: Audio data

[0260] Terminal: Plays the audio data and communicates the response to the user.

[0261] Output: A voice response saying "How many people would you like to book and what time would you like to book?"

[0262] Step 6: Accept user responses

[0263] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[0264] Input: User's spoken response

[0265] Device: Uses speech recognition software to convert your speech into text.

[0266] Data processing: Converting voice data into text data

[0267] Output: Text "Two people, starting tomorrow at 7pm"

[0268] Terminal: Sends the converted text to the server.

[0269] Step 7: Confirm your booking

[0270] Server: Receives the user's response and registers the information in the reservation system.

[0271] Input: Text "Two people, please meet tomorrow from 7pm"

[0272] Data calculation: Reservation information registration

[0273] Output: Reservation information registration completed

[0274] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[0275] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We look forward to seeing you tomorrow at 7pm for two guests."

[0276] Output: Confirmation message

[0277] Step 8: Send a confirmation message

[0278] Server: The confirmation message is converted into voice using text-to-speech software.

[0279] Input: Confirmation message text

[0280] Text-to-speech software: converts text into audio data.

[0281] Data processing: Convert text data into audio data

[0282] Output: Audio data

[0283] Server: Sends the generated voice data to the user device.

[0284] Terminal: Plays audio data to inform the user that the reservation is complete.

[0285] Output: Voice response "Your reservation is complete. We will be waiting for you tomorrow at 7pm for two people."

[0286] (Application example 2)

[0287] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0288] In the past, restaurant operations such as making reservations, getting directions to the restaurant, checking business hours, and calling to order were often done manually, which was time-consuming and labor-intensive. In addition, users had to use multiple methods to obtain this information, which was inconvenient.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an automatic voice response means, a voice-based generative AI means, a voice recognition means, a response generation means using a generative AI model, a means for making reservations based on a user's voice input, a means for providing directions, a means for checking business hours, and a means for automatically making an order call. This allows users to efficiently perform tasks such as making store reservations, getting directions to stores, checking business hours, and making an order call through a single system.

[0290] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[0291] A "voice generation AI means" is a system that uses artificial intelligence technology to analyze voice input and generate an appropriate voice response.

[0292] "Speech recognition means" refers to a system that has the technology to convert a user's voice into text data.

[0293] A "response generation means using a generative AI model" is a system that has the function of using a generative AI model to generate an appropriate response based on user input.

[0294] The "means for making a reservation based on a user's voice input" is a system that has the function of processing a reservation request made by a user through voice and confirming the reservation.

[0295] A "means for providing directions" is a system that has the function of providing users with directions to their destination via voice or text.

[0296] The "means for checking business hours" is a system that has the function of checking the business hours of a store specified by the user and providing the information to the user.

[0297] "Means for automatically making an order call" refers to a system that has the function of automatically transmitting the specified order details over the phone based on the user's voice input.

[0298] The "answer database" is a database that stores answers to questions prepared in advance.

[0299] "Short Message Service Notification" means a system that has the function of notifying users via short message service.

[0300] The "means for asking questions to confirm matters" is a system that has the function of asking the user necessary questions to confirm matters and obtaining the answers.

[0301] "Means that allow for text saving of user responses" refers to a system that has the function of saving the content of a user's voice responses as text data.

[0302] The following system configuration will be described as an embodiment of the present invention.

[0303] System Configuration

[0304] This system includes a server, a user terminal, and a voice recognition device. The server includes a response generation means using a generative AI model, a voice-based generative AI means, a voice recognition means, and a database. The user terminal includes a microphone and a speaker and receives voice input from the user.

[0305] Hardware and software used

[0306] Hardware:

[0307] Microphone: Accepts user voice input.

[0308] Speaker: Provides audio responses from the server to the user.

[0309] Server: Provides the computational resources to run generative AI models.

[0310] software:

[0311] speech_recognition library: Performs speech recognition.

[0312] The transformers library: Runs generative AI models.

[0313] Data processing and calculation

[0314] 1. Speech Recognition:

[0315] A microphone on the user terminal accepts the user's voice input.

[0316] A speech recognition means converts the speech input into text data.

[0317] 2. Response generation:

[0318] A response generation means using the server's generative AI model analyzes the text data obtained from the speech recognition means and generates an appropriate response.

[0319] A voice generation AI means converts the generated response into voice data.

[0320] 3. Booking Process:

[0321] Based on the user's voice input, the means for making reservations stores the reservation information in a database.

[0322] 4. Directions and business hours:

[0323] The server provides the necessary information in response to a user's request, using means for providing directions and means for checking opening hours.

[0324] 5. Order by phone:

[0325] The server uses a means for automatically placing an order call based on the user's voice input to convey the specified order details over the phone.

[0326] Specific examples

[0327] When a user says, "I would like to make a reservation," the system operates as follows.

[0328] 1. The microphone on the user device accepts voice input.

[0329] 2. A speech recognition tool converts the speech into text.

[0330] 3. The server's generative AI model generates a response generation method that asks the question, "How many people would you like to make a reservation for, and what time would you like to start?"

[0331] 4. The voice generation AI means converts the question into voice data and provides it to the user through the speaker on the user's device.

[0332] 5. When the user answers "Two people, starting at 7pm", the speech recognition means again converts the speech into text and saves the reservation information in the database.

[0333] Prompt Sentence Examples

[0334] A user says, "I'd like to schedule an appointment." What's your next question?

[0335] In this way, the system can efficiently perform tasks such as making reservations, getting directions, checking business hours, and placing orders over the phone, based on the user's voice input.

[0336] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0337] Step 1:

[0338] The user inputs "I would like to make a reservation" by voice. The microphone on the user terminal accepts this voice input. The input is the user's voice data, and the output is the voice data itself.

[0339] Step 2:

[0340] The voice recognition means of the user terminal converts the received voice data into text data. The input is voice data, and the output is text data such as "I would like to make a reservation to visit the restaurant."

[0341] Step 3:

[0342] The server's generative AI model is used to generate a response, which analyzes the text data and generates an appropriate question. The input is the text "I would like to make a reservation," and the output is the text "How many people would you like to make a reservation for, and what time would you like to make a reservation for?"

[0343] Step 4:

[0344] The server's voice generation AI means converts the generated text data of the question into voice data. The input is text data such as "How many people would you like to make a reservation for, and what time would you like to start?", and the output is voice data.

[0345] Step 5:

[0346] The speaker on the user's device plays the audio data sent from the server and presents the question to the user. The input is the audio data, and the output is the audio heard by the user.

[0347] Step 6:

[0348] The user responds by voice, "Two people, starting at 7 PM." The microphone on the user's device accepts this voice input. The input is the user's voice data, and the output is the voice data itself.

[0349] Step 7:

[0350] The voice recognition means of the user terminal converts the received voice data into text data. The input is voice data, and the output is text data such as "Two people, starting at 7pm."

[0351] Step 8:

[0352] The server's reservation processing means analyzes the text data and saves the reservation information in a database. The input is text data such as "2 people, starting at 7 PM," and the output is a database containing the reservation information.

[0353] Step 9:

[0354] The server generates voice data to notify the user that the reservation has been completed, and converts it into voice data using a voice generation AI means. The input is text data saying "Reservation completed," and the output is voice data.

[0355] Step 10:

[0356] The speaker on the user's terminal plays the audio data sent from the server and notifies the user that the reservation has been completed. The input is audio data, and the output is the audio heard by the user.

[0357] In this way, the system can efficiently make reservations based on the user's voice input.

[0358] Example 3

[0359] Next, a third embodiment of the third embodiment will be described. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0360] Conventional automated voice response systems struggled to generate appropriate voice questions in response to user requests and efficiently store user responses as text data. They also lacked the functionality to automatically send SMS notifications to confirm user responses or ask additional questions for confirmation. This resulted in a poor user experience and impaired system efficiency.

[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0362] In this invention, the server includes a means for receiving a request from a user, a means for generating a voice question using a voice generation AI, a means for receiving a user's response, a means for converting the received response into text data and saving it, a means for sending an SMS notification, and a means for generating and sending a confirmation question. This makes it possible to generate an appropriate voice question in response to a user's request and efficiently save the user's response as text data. Furthermore, by automatically sending an SMS notification or asking additional confirmation questions, it is possible to improve the user experience and increase the efficiency of the system.

[0363] The "means for receiving requests from a user" is a device or software for receiving requests made by a user to the system in voice or text format.

[0364] "Voice generation AI" is an artificial intelligence technology for converting text data into voice data, and is a system that can ask questions and provide guidance to users in a natural voice.

[0365] A "means for generating voice questions" is a device or software that uses a voice generation AI to generate appropriate questions in voice format based on the user's request.

[0366] A "means for receiving a user's response" is a device or software for receiving a response given by a user in voice or text form.

[0367] The "means for converting received responses into text data and saving the same" refers to a device or software for converting the user's voice responses into text data and saving the text data in a storage device such as a database.

[0368] A "means for sending SMS notifications" is a device or software for sending notifications to a user using Short Message Service (SMS).

[0369] A "means for generating and transmitting confirmation questions" is a device or software for generating additional questions to confirm the user's answers and transmitting them to the user in voice or text format.

[0370] MODE FOR CARRYING OUT THE INVENTION

[0371] This invention is a system that receives requests from users, generates voice questions using a voice generation AI, receives user responses, converts them into text data, and saves them. It also includes the functionality to send SMS notifications and generate and send confirmation questions.

[0372] Hardware and software used

[0373] Hardware: Servers, user devices (smartphones, tablets, PCs, etc.)

[0374] Software: Voice-generative AI (e.g., Google Cloud Text-to-Speech API), speech recognition technology (e.g., Google Cloud Speech-to-Text API), SMS notification systems (e.g., Twilio), database management systems (e.g., MySQL®)

[0375] Specific operation of the system

[0376] The server receives requests from users. For example, when a user uses a smartphone to request, "I would like to place an order," the request is sent to the server. The server uses a voice generation AI to generate a voice question such as, "What would you like to order?" and sends the voice data to the user's device.

[0377] The terminal plays the voice question sent from the server to the user. When the user answers "I'd like to order pizza," the terminal receives the voice data and sends it to the server.

[0378] The server analyzes the user's voice response received from the terminal and converts it into text data using voice recognition technology. For example, the generated text data is "I would like to order pizza." This text data is then stored in a database management system.

[0379] The server uses an SMS notification system to send a confirmation message to the user to confirm the user's order, for example, "Your pizza order has been received. Please reply to confirm."

[0380] The server receives the user's reply and asks additional questions if necessary. For example, a question like "What kind of pizza do you want?" is generated using a voice-based AI and sent to the device. The device then plays this question back to the user and sends the user's answer back to the server.

[0381] Specific examples

[0382] Example: Restaurant ordering system

[0383] 1. The user uses their smartphone to request an order.

[0384] 2. The server uses a voice generation AI to generate a voice question such as "What would you like to order?" and sends it to the terminal.

[0385] 3. The user responds, "I would like to order pizza," and the voice data is sent to the server via the device.

[0386] 4. The server converts the voice data into text and stores the message "I would like to order pizza" in the database.

[0387] 5. The server sends an SMS to the user saying, "Your pizza order has been accepted. Please reply to confirm."

[0388] 6. When the user replies, the server generates a follow-up question, "What kind of pizza do you want?" and sends it to the device.

[0389] Example prompts for generative AI models

[0390] When a user requests to "order," please describe a system that uses a voice-generative AI to ask "What would you like to order?" and saves the user's response as text.

[0391] The flow of the identification process in the third embodiment will be described with reference to FIG.

[0392] Program processing flow

[0393] Step 1: Receiving a user request

[0394] Subject: Terminal

[0395] The terminal receives a request from the user saying, "I would like to place an order." When the user speaks "I would like to place an order" into the microphone of their smartphone, the voice data is input into the terminal. The terminal then sends this voice data to the server.

[0396] Input: User's voice request

[0397] Output: Sending audio data to the server

[0398] Step 2: Generate a voice question

[0399] Subject: Server

[0400] The server analyzes the user's request received from the device and determines the next action to take. The server uses a voice generation AI to generate a voice question such as "What would you like to order?" This voice data is sent from the server to the device.

[0401] Input: User's voice request data

[0402] Output: Generated voice question data

[0403] Step 3: Receiving the user's response

[0404] Subject: Terminal

[0405] The terminal plays the voice question sent from the server to the user. When the user answers "I would like to order pizza," the terminal receives the voice data. The terminal then transmits this voice data to the server.

[0406] Input: Voice question data from the server, user's voice response

[0407] Output: Sending voice response data to the server

[0408] Step 4: Save the answer as text

[0409] Subject: Server

[0410] The server analyzes the user's voice response received from the terminal and converts it into text data using voice recognition technology. For example, the generated text data is "I would like to order pizza." This text data is then stored in a database management system.

[0411] Input: User's voice response data

[0412] Output: Save the converted text data

[0413] Step 5: Sending SMS notifications

[0414] Subject: Server

[0415] The server uses an SMS notification system to send a confirmation message to the user to confirm the user's order, for example, "Your pizza order has been received. Please reply to confirm."

[0416] Input: Text data (order details)

[0417] Output: SMS notification to the user

[0418] Step 6: Verification Questions

[0419] Subject: Server

[0420] The server receives the user's reply and asks additional questions if necessary. For example, a question like "What kind of pizza do you want?" is generated using a voice-based AI and sent to the device. The device then plays this question back to the user and sends the user's answer back to the server.

[0421] Input: User reply data

[0422] Output: Send generated additional question data

[0423] (Application example 3)

[0424] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0425] Conventional food delivery systems require users to manually input their orders, making operation cumbersome. It is also difficult to confirm or change order details, which can lead to a poor user experience. Furthermore, order details are not confirmed or notified in real time, making it difficult for users to understand the status of their order.

[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes an automatic voice response means, a voice version generation system AI means, a means for converting voice input into text, a means for saving the text, a means for sending SMS notifications, and a means for asking confirmation questions. This allows the user to easily place an order by voice, and the order details can be confirmed or changed in real time, improving the user experience. In addition, the order details are saved as text and notified via SMS, making it easier for the user to understand the order status.

[0427] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[0428] "Voice generation AI means" is a system that uses artificial intelligence technology to analyze voice data and generate appropriate voice responses.

[0429] A "means for converting voice input into text" is a system that has the ability to recognize a user's voice and convert the content into text data.

[0430] The "means for storing text" is a system that has the function of storing the converted text data in a database or file system.

[0431] "Means for sending SMS notifications" refers to a system that has the function of sending notifications to users via short message service (SMS) based on stored text data.

[0432] The "means for asking confirmation questions" is a system that has the function of asking additional confirmation questions by voice based on the user's input and receiving responses from the user.

[0433] A system for implementing this invention has the following configuration: The server includes an automatic voice response means, a voice version generation system AI means, a means for converting voice input into text, a means for saving the text, a means for sending SMS notifications, and a means for asking confirmation questions.

[0434] Hardware and software used

[0435] Hardware: Smartphone (microphone, speaker)

[0436] Software: Python, Flask (web framework), SpeechRecognition (voice recognition library), pyttsx3 (speech synthesis library), smtplib (email sending library)

[0437] Processing flow

[0438] 1. User voice input:

[0439] A user launches a smartphone app and places an order by voice. For example, the user says, "I'd like to order a pizza."

[0440] 2. Speech Recognition:

[0441] The smartphone's microphone captures the user's voice and converts it into text using a speech recognition library (SpeechRecognition). For example, the speech "I want to order a pizza" is converted into text "I want to order a pizza."

[0442] 3. Response by voice-generative AI:

[0443] A speech-based AI solution generates appropriate responses based on the user's voice input and asks the user questions through the smartphone speaker, such as "What would you like to order?"

[0444] 4. User Answer:

[0445] The user responds, "One Margherita pizza." Again, the speech recognition library is used to convert this speech to text.

[0446] 5. Save text:

[0447] The converted text data is stored in a database or file system on the server. For example, the text "One Margherita pizza" is stored.

[0448] 6. SMS notification:

[0449] Based on the saved text data, the user is notified via short message service (SMS). For example, an SMS message with the content "Order: 1 Margherita pizza" is sent to the user.

[0450] Specific examples

[0451] When a user says, "I'd like to order a pizza," the system asks aloud, "What would you like to order?", and the user replies, "One Margherita pizza." The system saves this as text and sends an SMS message saying, "Order: One Margherita pizza."

[0452] Prompt Sentence Examples

[0453] User: "I want to order a pizza."

[0454] System: "What would you like to order?"

[0455] User: "One Margherita pizza."

[0456] System: "Your order has been accepted. Order: 1 Margherita pizza."

[0457] In this way, a food delivery application can be realized that allows users to easily place orders by voice.

[0458] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0459] Step 1:

[0460] A user launches a smartphone app and places an order by voice. The user says, "I'd like to order a pizza." The smartphone's microphone captures this voice. The input is the user's voice data, and the output is the captured voice data.

[0461] Step 2:

[0462] The smartphone sends the captured voice data to a speech recognition library (SpeechRecognition), which converts the voice into text. The input is the captured voice data, and the output is the text data "I would like to order a pizza." The server receives this text data and proceeds to the next step.

[0463] Step 3:

[0464] The server uses a voice generation AI method to generate an appropriate response based on the user's text input. For example, it generates a voice response such as "What would you like to order?" The input is text data such as "I would like to order pizza," and the output is voice data such as "What would you like to order?" The server then sends this voice data to the smartphone.

[0465] Step 4:

[0466] The smartphone plays the voice data received from the server and asks the user, "What would you like to order?" The input is the voice data sent from the server, and the output is the voice question to the user. The user answers, "One Margherita pizza."

[0467] Step 5:

[0468] The smartphone captures the user's answer again and converts it into text using a speech recognition library. The input is the user's voice data, and the output is the text data "One Margherita pizza." The server receives this text data and proceeds to the next step.

[0469] Step 6:

[0470] The server saves the converted text data in a database or file system. The input is the text data "one Margherita pizza" and the output is the saved text data. The saved data can be referenced later.

[0471] Step 7:

[0472] The server notifies the user via short message service (SMS) based on the saved text data. The input is the saved text data, and the output is an SMS with the content "Order: 1 Margherita pizza." The server sends this SMS to the user's smartphone.

[0473] In this way, a system is realized that allows users to easily place orders by voice and receive real-time confirmation and notification of order details.

[0474] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0475] "Example 1"

[0476] One embodiment of the present invention is a system that incorporates an emotion engine. This system recognizes emotions from the user's tone of voice and language and responds accordingly. For example, if the system determines that the user is feeling angry or frustrated, it responds with more polite language and a softer tone. On the other hand, if the system determines that the user is feeling happy or excited, it responds with a more lively tone. This allows for optimal responses according to the user's emotions.

[0477] "Example 2"

[0478] Another embodiment of the present invention is a system that optimizes restaurant service according to the user's emotions. This system recognizes the user's emotions and optimizes the restaurant's service accordingly. For example, if the system determines that the user is feeling angry or frustrated, it conveys this information to the restaurant staff, encouraging them to provide more courteous service. If the system determines that the user is feeling happy or excited, it conveys this information to the restaurant staff, encouraging them to provide more lively service. This makes it possible to provide optimal service according to the user's emotions.

[0479] "Example 3"

[0480] Furthermore, another embodiment of the present invention is a system that provides a voice response or SMS notification according to a user's emotions. This system recognizes the user's emotions and provides a voice response or SMS notification accordingly. For example, if the system determines that the user is feeling angry or frustrated, it provides a voice response using more polite language and a softer tone. On the other hand, if the system determines that the user is feeling happy or excited, it provides a voice response in a more lively tone. This enables optimal voice responses and SMS notifications according to the user's emotions.

[0481] The processing flow of each embodiment will be described below.

[0482] "Example 1"

[0483] Step 1: Receive speech input from the user.

[0484] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[0485] Step 3: Adjust the tone and wording of your voice response depending on the emotion you recognize.

[0486] "Example 2"

[0487] Step 1: Receive speech input from the user.

[0488] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[0489] Step 3: Communicate the recognized emotions to restaurant staff to encourage them to optimize their service.

[0490] "Example 3"

[0491] Step 1: Receive speech input from the user.

[0492] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[0493] Step 3: Tailor the content of voice responses and SMS notifications depending on the recognized emotion.

[0494] Example 1

[0495] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0496] Conventional automated voice response systems could only provide fixed responses to user questions and requests, making it difficult to respond flexibly to the user's emotions. Furthermore, the process of converting the user's voice data into text and generating appropriate responses using a generative AI model was complex, creating a need for an efficient system. Furthermore, certain industries, such as restaurants, required systems that could respond to specific needs, such as making reservations and checking opening hours.

[0497] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0498] In this invention, the server includes means for linking an automated voice response with a voice generation AI, means for providing the service for a fixed fee, means for providing a number, means for providing voice guidance, means for linking a voice generation AI function, means for converting user voice data into text, means for transmitting a prompt to the AI ​​generation model, means for the AI ​​generation model to generate an answer, means for converting the generated answer into voice, means for providing the answer to the user in voice, means for analyzing the user's emotions, means for instructing the AI ​​generation model to respond according to the emotion, means for generating an answer according to the emotion, means for converting the answer according to the emotion into voice, and means for providing an answer in voice according to the emotion. This makes it possible to provide appropriate answers to user questions and requests and to respond flexibly according to the user's emotions.

[0499] An "automatic voice response" is a system that automatically responds to phone calls and voice inputs from users.

[0500] "Speech generation AI" is an AI technology that generates natural-sounding speech based on user text input and prompts.

[0501] A "fee-based means" is a means that indicates that a service or system is available for a set fee.

[0502] "Means for providing numbers" refers to means for providing telephone numbers or identification numbers for users to access the system.

[0503] The "audio guidance means" is a means for providing guidance and instructions to the user by voice.

[0504] "Generative AI linking function means" is a means for linking speech-generating AI with other systems and functions.

[0505] The "means for converting user's voice data into text" refers to a means for converting voice data input by a user into text format.

[0506] The "means for sending a prompt sentence to the generative artificial intelligence model" is a means for sending a prompt sentence including an instruction or question to the generative artificial intelligence model.

[0507] "Means by which a generative artificial intelligence model generates an answer" refers to means by which a generative artificial intelligence model generates an appropriate answer based on a prompt sentence.

[0508] The "means for converting the generated answer into speech" refers to a means for converting the text-format answer generated by the generative artificial intelligence model into speech format.

[0509] The "means for providing a user with an answer in the form of voice" refers to a means for providing a user with an answer that has been converted into voice format.

[0510] The "means for analyzing user emotions" refers to a means for analyzing emotions from user voice data and text data.

[0511] The "means for instructing the generative artificial intelligence model to respond according to the emotion" is a means for instructing the generative artificial intelligence model to respond appropriately based on the emotion of the user.

[0512] The "means for generating an answer according to emotions" is a means for a generative artificial intelligence model to generate an appropriate answer according to the user's emotions.

[0513] The "means for converting a response according to emotion into voice" is a means for converting a response in text format generated according to emotion into voice format.

[0514] The "means for providing an answer in a voice corresponding to an emotion" is a means for providing a user with an answer in the form of a voice corresponding to an emotion.

[0515] This invention is a system that combines an automatic voice response system with a speech generation artificial intelligence (AI) to provide appropriate answers to user questions and requests, and also enables flexible responses according to the user's emotions. This system is implemented using the following hardware and software.

[0516] Hardware and software used

[0517] Server: A central processing unit that manages the entire system and calls various APIs.

[0518] Twilio API: A communications API for managing incoming and outgoing phone calls.

[0519] Google Cloud Speech-to-Text API: A speech recognition API for converting user voice data into text.

[0520] Generative AI model (GPT-4 (registered trademark)): An artificial intelligence model for generating appropriate answers to user questions and requests.

[0521] Google Cloud Text-to-Speech API: A speech synthesis API for converting generated text responses into audio.

[0522] Emotion recognition engine (IBM Watson(R) Tone Analyzer): An engine for analyzing emotions from user voice and text data.

[0523] Specific operation of the system

[0524] 1. The user makes a call

[0525] The user calls the provided phone number, for example, a customer support phone number.

[0526] 2. The server receives the call

[0527] The server receives a call from the user using the Twilio API. If the call connects, the server proceeds to the next step.

[0528] 3. The server provides audio guidance

[0529] The server provides voice prompts to the user through the Twilio API, for example, playing a message such as "Please tell us your business."

[0530] 4. The user speaks their question or request

[0531] The user follows the voice prompts to input a question or request by voice, for example, "Please tell me how to return an item."

[0532] 5. The server converts the audio data into text

[0533] The server uses the Google Cloud Speech-to-Text API to convert the user's voice data into text, which becomes "How do I return this item?"

[0534] 6. The server sends a prompt to the generative AI model

[0535] The server sends a prompt to the generative AI model (GPT-4). For example, it sends a prompt such as, "The user is asking how to return a product. Please tell me how to return it."

[0536] 7. Generative AI models generate answers

[0537] The generative AI model generates an appropriate answer based on the prompt, for example, "To return the product, bring it to the store with your receipt within 30 days of purchase."

[0538] 8. The server converts the answer into audio

[0539] The server converts the generated answer into audio using the Google Cloud Text-to-Speech API.

[0540] 9. The server provides the user with a spoken response

[0541] The server provides the user with a response converted into voice via the Twilio API. For example, it might say, "To return the product, please bring it to the store with your receipt within 30 days of purchase."

[0542] 10. The server uses an emotion recognition engine to analyze the user's emotions.

[0543] The server sends the user's voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to analyze the user's emotions.

[0544] 11. The server instructs the generative AI model on how to respond based on the emotion.

[0545] The server sends instructions to the generative AI model based on the results of the emotion recognition engine, for example, sending a prompt such as "The user is angry. Please respond in a polite manner."

[0546] 12. Generative AI models generate emotionally relevant answers

[0547] The generative AI model generates appropriate responses based on the customer's sentiment, such as, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase."

[0548] 13. The server converts the response into speech based on the emotion.

[0549] The server uses the Google Cloud Text-to-Speech API to convert the emotional response into speech.

[0550] 14. The server provides the user with a response in voice based on their emotions.

[0551] The server provides the user with a response based on the emotion through the Twilio API. For example, it might say, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase."

[0552] Prompt Sentence Examples

[0553] "Users are asking about business hours. What are your business hours?"

[0554] "A user is asking how to return an item. How do I return it?"

[0555] "User is angry. Please respond politely."

[0556] In this way, the server provides appropriate answers to the user's questions and requests, realizing a system that responds according to emotions.

[0557] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0558] Step 1:

[0559] User makes a call

[0560] The user calls a provided phone number, for example a customer support phone number. The input is the user's phone number and the call start time, and the output is the call connection to the server.

[0561] Step 2:

[0562] The server receives the call

[0563] The server receives a call from the user using the Twilio API. Once the call connects, the server proceeds to the next step. The input is the call signal from the user, and the output is the establishment of the call connection.

[0564] Step 3:

[0565] The server provides audio guidance

[0566] The server provides a voice prompt to the user through the Twilio API. For example, it plays a message such as "Please tell us your business." The input is the establishment of a call connection, and the output is the playback of the voice prompt.

[0567] Step 4:

[0568] Users speak their questions or requests

[0569] The user inputs a question or request by voice, following the voice guidance. For example, say, "Please tell me how to return a product." The input is the voice guidance, and the output is the user's voice data.

[0570] Step 5:

[0571] The server converts the audio data into text

[0572] The server converts the user's voice data into text using the Google Cloud Speech-to-Text API. The converted text is "How do I return the item?" The input is the user's voice data, and the output is text data.

[0573] Step 6:

[0574] The server sends a prompt to the generative AI model

[0575] The server sends a prompt to the generative AI model (GPT-4). For example, it sends a prompt such as, "The user is asking how to return a product. Please tell me how to return it." The input is text data, and the output is the generation and transmission of a prompt.

[0576] Step 7:

[0577] Generative AI models generate answers

[0578] The generative AI model generates an appropriate answer based on the prompt sentence. For example, it generates an answer such as, "To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is the prompt sentence, and the output is the generated answer text.

[0579] Step 8:

[0580] The server converts the answer into audio

[0581] The server converts the generated answer into speech using the Google Cloud Text-to-Speech API, where the input is the generated answer text and the output is the audio data.

[0582] Step 9:

[0583] The server provides the user with a spoken response

[0584] The server provides the user with a response converted into voice via the Twilio API. For example, it may respond with a voice message such as, "To find out how to return the product, please bring it to the store with your receipt within 30 days of purchase." The input is voice data, and the output is a voice response to the user.

[0585] Step 10:

[0586] The server uses an emotion recognition engine to analyze the user's emotions.

[0587] The server sends the user's voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to analyze the user's emotions. The input is the user's voice data, and the output is the emotion analysis result.

[0588] Step 11:

[0589] The server instructs the generative AI model to respond according to the emotion.

[0590] The server sends instructions to the generative AI model based on the results of the emotion recognition engine. For example, it sends a prompt such as, "The user is feeling angry. Please respond in a polite manner." The input is the emotion analysis results, and the output is the generation and transmission of a prompt.

[0591] Step 12:

[0592] Generative AI models generate emotionally relevant answers

[0593] The generative AI model generates an appropriate response based on the sentiment. For example, it generates an answer like, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is the prompt sentence, and the output is the generated answer text.

[0594] Step 13:

[0595] The server converts the response into speech based on the emotion.

[0596] The server converts the sentiment-based response into speech using the Google Cloud Text-to-Speech API. The input is the generated response text, and the output is the speech data.

[0597] Step 14:

[0598] The server provides the user with a response in voice that corresponds to their emotions.

[0599] The server provides the user with a voice response based on the emotion through the Twilio API. For example, it might say, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is voice data, and the output is a voice response to the user.

[0600] (Application example 1)

[0601] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0602] Conventional automated voice response systems have had problems in that they are unable to provide appropriate answers to user questions and respond in a way that reflects the user's emotions. In particular, solving these problems is important for customer service in brick-and-mortar stores, where quick and appropriate responses are required.

[0603] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a means for linking an automatic voice response (IVR) with a voice-based generation AI, a means provided for a fixed fee, a number providing means, a voice guidance means, a generation AI linkage function means, a means for converting a user's voice into text, a means for generating answers to user questions using a generation AI model, a means for converting the generated answers into voice, and a means for recognizing the user's emotions and responding in an appropriate tone. This makes it possible to provide quick and appropriate answers to user questions and respond in accordance with the user's emotions.

[0604] An "Interactive Voice Response (IVR)" is a system that automatically responds via voice over the telephone, providing appropriate information based on user input.

[0605] "Voice generative AI" is an artificial intelligence technology that analyzes voice data and generates appropriate answers to users' questions and requests.

[0606] A "fixed fee means" is a system in which services are provided at a fixed fee.

[0607] "Number provisioning means" is a function that provides a telephone number for users to access the system.

[0608] The "audio guidance means" is a function that provides guidance and instructions to the user by voice.

[0609] "Generative AI collaboration function means" refers to technology that allows generative AI to collaborate with other systems and functions.

[0610] "Means for converting user speech into text" refers to technology that converts the user's speech into text information.

[0611] "Means for generating answers to user questions using a generative AI model" refers to technology that uses a generative AI model to create appropriate answers to user questions.

[0612] The "means for converting the generated answer into speech" is a technique for converting the generated text-format answer into speech format.

[0613] "Means to recognize the user's emotions and respond in an appropriate tone" refers to technology that analyzes the user's emotions from their tone of voice and choice of words, and responds in an appropriate tone accordingly.

[0614] "Means for improving customer service in physical stores" refers to technologies and methods for streamlining customer service in physical stores and improving customer satisfaction.

[0615] The "prepared answer database" is a database of questions and their answers that have been prepared in advance.

[0616] "Short message notification" is a function that sends short text messages to users.

[0617] "Means for asking questions to confirm and saving the customer's answers as text" refers to a technology that asks the user questions to confirm and saves the answers in text format.

[0618] As an embodiment of the present invention, a smartphone application for improving customer service in a brick-and-mortar store will be described as an example.

[0619] System Program

[0620] This system operates using the following hardware and software:

[0621] Hardware:

[0622] Smartphone (including microphone and speaker)

[0623] software:

[0624] Python

[0625] speech_recognition library: used to convert speech to text

[0626] gTTS library: used to convert text to speech

[0627] playsound library: used to play sounds

[0628] OpenAI API: Uses generative AI models to generate answers to user questions

[0629] Processing flow

[0630] The server first records the user's voice using the smartphone's microphone, then converts the recorded voice into text using the speech_recognition library, which is then sent to the OpenAI API to generate an appropriate answer using a generative AI model, which is then converted into speech using the gTTS library and played over the smartphone's speaker.

[0631] In addition, to recognize the user's emotions, the system analyzes voice data and determines the user's emotions from the tone of their voice and the way they speak. Based on this information, the system responds with an appropriate tone.

[0632] Specific examples

[0633] For example, if a user opens a smartphone app and asks, "What are your business hours?", the app converts the question into text and sends it to the OpenAI API. The generative AI model generates the answer, "Business hours are from 9:00 AM to 5:00 PM," and converts this answer into audio and plays it back to the user.

[0634] Example prompt sentence:

[0635] User Question: What are your opening hours?

[0636] answer:

[0637] In this way, it is possible to provide a quick and appropriate answer to the user's question and also to respond in accordance with the user's feelings.

[0638] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0639] Step 1:

[0640] A user launches a smartphone app and speaks a question into the microphone. The input is the user's voice data, and the output is a recorded audio file. Specifically, the smartphone's microphone captures the user's voice and saves it as an audio file.

[0641] Step 2:

[0642] The device converts the recorded audio file into text using the speech_recognition library. The input is an audio file and the output is text data. Specifically, the speech_recognition library analyzes the audio file and converts the audio into text information.

[0643] Step 3:

[0644] The device sends the generated text data to the OpenAI API and uses the generative AI model to generate an appropriate answer. The input is the text data, and the output is the generated answer text. Specifically, the text data is sent to the OpenAI API as a prompt, and the generative AI model generates an answer.

[0645] Step 4:

[0646] The device converts the generated response text into speech using the gTTS library. The input is the response text, and the output is an audio file. Specifically, the gTTS library converts the text data into speech data and saves it as an audio file.

[0647] Step 5:

[0648] The device plays the generated audio file on the smartphone's speaker. The input is the audio file, and the output is the audio that the user can hear. Specifically, it uses the playsound library to play the audio file and outputs the audio from the speaker.

[0649] Step 6:

[0650] The server analyzes the user's voice data and uses an emotion engine to recognize the user's emotions. The input is the user's voice data and the output is emotional information. Specifically, the server analyzes the voice data and determines the user's emotions from the tone of voice and the use of words.

[0651] Step 7:

[0652] The server generates a response in an appropriate tone based on the recognized emotional information and converts it into speech. The input is the emotional information and the response text, and the output is a speech file with a tone corresponding to the emotion. Specifically, the response text is adjusted taking into account the emotional information and converted into speech using the gTTS library.

[0653] Step 8:

[0654] The device plays an audio file with a tone corresponding to the emotion and provides it to the user. The input is an audio file with a tone corresponding to the emotion, and the output is a sound that the user can hear. Specifically, it uses the playsound library to play the audio file and outputs the sound from the speaker.

[0655] Example 2

[0656] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0657] Conventional automated voice response systems have difficulty responding flexibly to user requests, making it difficult to increase user satisfaction, especially when it comes to restaurant reservations and optimizing services. Furthermore, there were no systems that recognized user emotions and responded accordingly, making it impossible to improve the quality of service.

[0658] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice input, means for converting voice data into text, means for generating appropriate questions and answers using a generative AI model, means for converting the generated questions and answers into voice and conveying them to the user, means for recognizing the user's emotions and generating a response according to the emotions, and means for notifying staff of the generated response. This enables flexible response to user requests and optimization of restaurant reservations and services. Furthermore, by responding according to the user's emotions, the quality of service can be improved and user satisfaction can be increased.

[0659] "Means for receiving user voice input" refers to devices or techniques for capturing user-uttered voice and inputting it into the system.

[0660] A "voice-to-text converter" is a process or device that uses voice recognition technology to convert captured voice data into written information.

[0661] "Means for generating appropriate questions and answers using generative AI models" refers to systems or algorithms that utilize artificial intelligence technology to automatically generate questions and answers in response to user requests.

[0662] The "means for converting the generated questions and answers into voice and communicating them to the user" refers to a technology or device for converting text information into voice and providing the information to the user by voice.

[0663] "Means for recognizing a user's emotions and generating a response that corresponds to those emotions" refers to a system or technology that analyzes the user's emotions from their voice and facial expressions, and automatically generates an appropriate response based on the results.

[0664] The "means for notifying staff of the generated response" refers to a device or technology for notifying restaurant staff of the generated response information.

[0665] MODE FOR CARRYING OUT THE INVENTION

[0666] This invention is a system that uses a generative AI model to solve services such as restaurant reservations, restaurant directions, checking business hours, and phone orders. It also includes a function to recognize user emotions and optimize services accordingly.

[0667] Hardware and software used

[0668] server

[0669] The server receives voice input from the user and uses speech recognition technology to convert the voice data into text. Specifically, it uses the Google Cloud Speech-to-Text API. It then uses a generative AI model (e.g., OpenAI's GPT-4) to generate appropriate questions and answers based on the user's request. It uses speech synthesis technology (e.g., Amazon Polly) to convert the generated text into speech. It also uses emotion recognition technology (e.g., Microsoft® Azure®'s Emotion API) to recognize the user's emotions.

[0670] Terminal

[0671] The device captures the user's voice input and sends it to the server. It also uses speech synthesis technology to return responses from the server to the user as voice. It also has a camera and microphone to capture the user's facial expressions and send data for emotion recognition to the server.

[0672] User

[0673] The user inputs a request by voice into the terminal. For example, if the user says, "I'd like to make a reservation," the system generates an appropriate question and returns it to the user by voice. When the user inputs the answer, the system makes the reservation based on that information.

[0674] Specific examples

[0675] When a user says, "I'd like to make a reservation," the server converts the voice data into text using the Google Cloud Speech-to-Text API. Next, it uses a generative AI model (OpenAI's GPT-4) to generate the question, "How many people would you like to make a reservation for, and what time would you like to start?" The server converts this question into speech using Amazon Polly and relays it to the user via the device. When the user replies, "We'd like to make a reservation for three people, starting at 7 p.m.", the device sends the voice data to the server, which again converts it into text using the Google Cloud Speech-to-Text API. The server processes the reservation information, generates a confirmation message stating, "Your reservation for three people starting at 7 p.m.", and relays it to the user via Amazon Polly.

[0676] Prompt Sentence Examples

[0677] "A user wants to make a reservation. How many people would you like to book and what time would you like to book?"

[0678] If the user is dissatisfied with the restaurant, the server will analyze their emotions using Microsoft Azure's Emotion API and generate a response such as, "The user is dissatisfied. Please respond more politely." The device will then notify the staff of this information and prompt them to take appropriate action.

[0679] Prompt Sentence Examples

[0680] "Users are angry and frustrated. Please encourage your staff to be more polite."

[0681] In this way, restaurants can flexibly respond to user requests and optimize restaurant reservations and services. Furthermore, by responding to user emotions, the quality of service can be improved and user satisfaction can be increased.

[0682] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0683] Program processing flow

[0684] A system for making reservations, providing directions to the store, checking business hours, and answering phone orders

[0685] Step 1:

[0686] The user inputs voice data. The user speaks to the terminal, saying, "I'd like to make a reservation." The input is the user's voice data.

[0687] Step 2:

[0688] The device sends voice data to the server. The device captures the user's voice and sends the data to the server. The input is the user's voice data, and the output is the voice data sent to the server.

[0689] Step 3:

[0690] The server converts the voice data into text. The server converts the voice data into text using the Google Cloud Speech-to-Text API. The input is voice data and the output is text data.

[0691] Step 4:

[0692] The server uses a generative AI model to generate an appropriate question. The server uses OpenAI's GPT-4 to generate the question, "How many people would you like to book, and what time would you like to start?" The input is text data, and the output is the generated question text.

[0693] Step 5:

[0694] The server sends the generated question to the terminal. The server sends the generated question to the terminal. The input is the text of the generated question, and the output is the text of the question sent to the terminal.

[0695] Step 6:

[0696] The device speaks the question to the user. The device uses Amazon Polly to convert the question into speech and speaks it to the user. The input is the text data of the question, and the output is speech data.

[0697] Step 7:

[0698] The user inputs their answer by voice. The user answers, "I would like to make a reservation for three people starting at 7 PM." The input is the user's voice data.

[0699] Step 8:

[0700] The device sends the voice data of the answer to the server. The device captures the user's answer and sends the data to the server. The input is the user's voice data, and the output is the voice data sent to the server.

[0701] Step 9:

[0702] The server converts the voice data into text. The server again converts the voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data.

[0703] Step 10:

[0704] The server processes the reservation information and generates a confirmation message. The server processes the reservation information and generates a confirmation message saying, "We have received a reservation for 3 people starting at 7 PM." The input is text data, and the output is the text of the generated confirmation message.

[0705] Step 11:

[0706] The server sends a confirmation message to the terminal. The server sends a confirmation message to the terminal. The input is the generated confirmation message text, and the output is the confirmation message text sent to the terminal.

[0707] Step 12:

[0708] The device will speak a confirmation message to the user. The device will convert the confirmation message into speech using Amazon Polly and speak it to the user. The input is the text data of the confirmation message, and the output is the speech data.

[0709] A system that optimizes services based on emotions

[0710] Step 1:

[0711] A user receives a service at a restaurant. A user enjoys a meal at a restaurant. The input is the user's behavior.

[0712] Step 2:

[0713] The device captures the user's voice and facial expressions. The device captures the user's voice and facial expressions using a camera and microphone. The input is the user's voice data and facial expression data.

[0714] Step 3:

[0715] The device sends the captured data to the server. The device sends the captured data to the server. The input is voice data and facial expression data, and the output is the data sent to the server.

[0716] Step 4:

[0717] The server analyzes the user's emotions using emotion recognition technology. The server analyzes the user's emotions using Microsoft Azure's Emotion API. The input is voice data and facial expression data, and the output is the emotion analysis results.

[0718] Step 5:

[0719] The server generates an appropriate response based on the analysis results. The server uses OpenAI's GPT-4 to generate a response such as "The user is dissatisfied. Please respond more politely." The input is the sentiment analysis result, and the output is the generated response text.

[0720] Step 6:

[0721] The server sends the generated response to the terminal. The server sends the generated response to the terminal. The input is the text of the generated response, and the output is the text of the response sent to the terminal.

[0722] Step 7:

[0723] The terminal notifies the staff of the response. The terminal notifies the staff of the response using the display and voice output function. The input is the text data of the response, and the output is the information notified to the staff.

[0724] (Application example 2)

[0725] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0726] Conventional restaurant services such as making reservations, getting directions, checking business hours, and calling to order are cumbersome for users, and they have the problem of being difficult to respond to efficiently. In addition, services are not optimized according to user emotions, making it difficult to improve customer satisfaction. To solve these problems, a system utilizing voice recognition technology, generative AI, and emotion recognition technology is needed.

[0727] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0728] In this invention, the server includes an automatic voice response means, a voice-based AI means, an emotion recognition means, a location information providing means, a voice guidance means, and an AI generative function means. This allows users to easily make reservations, get directions, check business hours, place orders, and more by voice. Furthermore, by recognizing the user's emotions and providing optimal services in response to them, customer satisfaction can be improved.

[0729] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[0730] A "voice generation AI means" is a system that uses artificial intelligence technology to analyze voice input and generate an appropriate voice response.

[0731] An "emotion recognition means" is a system that has the technology to analyze emotions from the user's voice and facial expressions and recognize those emotions.

[0732] "Location information providing means" is a system that has the function of providing location information of the user's current location and destination.

[0733] "Audio guidance means" refers to a system that has the function of providing guidance and instructions to the user by voice.

[0734] A "generative AI collaboration function means" is a system that has the functionality to work in collaboration with generative AI and other systems and databases.

[0735] The "restaurant reservation means" is a system that allows users to make reservations for restaurants by voice.

[0736] The "restaurant navigation means" is a system that has the function of providing users with directions to restaurants.

[0737] The "means for checking business hours" is a system that has a function that allows users to check the business hours of restaurants by voice.

[0738] The "telephone ordering means" is a system that allows users to place orders with restaurants by telephone.

[0739] "Means for optimizing services according to emotions" is a system that has the function of recognizing the user's emotions and providing the optimal service according to those emotions.

[0740] The "prepared response database means" is a system having a function of using a database that stores prepared responses.

[0741] "Voice response means" refers to a system that has the function of providing answers to users by voice.

[0742] The "short message notification means" is a system that has the function of notifying the user by short message.

[0743] The "means for asking questions to confirm matters" is a system that has the function of asking the user questions to confirm matters.

[0744] "Means for saving customer responses in text" refers to a system that has the function of saving user responses in text format.

[0745] The system for implementing this invention uses the following hardware and software: a smartphone, microphone, and GPS are required for the hardware, and OpenAI API, SpeechRecognition, Geopy, and EmotionRecognizer are used for the software.

[0746] The server receives user voice input via a microphone and converts it to text using the SpeechRecognition library. It then uses the OpenAI API to generate an appropriate response based on the user's input. For example, if a user says, "I'd like to make a reservation," the generative AI will ask, "How many people would you like to make a reservation for, and what time would you like to make a reservation?"

[0747] The server also recognizes the user's emotions using EmotionRecognizer and notifies the staff of the appropriate response based on that emotion. For example, if the user is angry, the system will notify the staff, "The user is angry. Please respond politely."

[0748] Furthermore, the server uses location information provision means to provide directions from the user's current location to the restaurant. It uses the Geopy library to obtain the user's current location and calculate the optimal route to the destination.

[0749] As a specific example, if a user says, "Tell me where the store is," the generative AI will respond, "I will provide directions from your current location to the store," and begin providing directions using location information provision means.

[0750] An example of a prompt is as follows:

[0751] "A user wants to make a reservation. How many people would you like to book and what time would you like to book?"

[0752] "User wants to know where your store is. Please provide directions from their current location to the store."

[0753] "Users want to know your hours. What are your hours?"

[0754] In this way, users can easily make reservations, get directions, check business hours, place orders, etc. by voice. Furthermore, the system can recognize the user's emotions and provide optimal service accordingly, which can improve customer satisfaction.

[0755] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0756] Step 1:

[0757] The user inputs voice using the microphone on their smartphone. For example, the user's voice input might be, "I'd like to make a reservation to visit the restaurant." The input is voice data, and the output is the voice data itself.

[0758] Step 2:

[0759] The device converts voice input to text using the SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device analyzes the voice data and generates corresponding text.

[0760] Step 3:

[0761] The server receives the text data and generates an appropriate response using the OpenAI API. The input is text data, and the output is the generated response text. For example, in response to the input "I would like to make a reservation," the server generates the response "How many people would you like to make a reservation for, and what time would you like to make a reservation?"

[0762] Step 4:

[0763] The server converts the generated response text into speech and provides audio guidance to the user. The input is the generated response text, and the output is audio data. Specifically, the text is converted into audio data using speech synthesis technology.

[0764] Step 5:

[0765] The user again speaks to provide reservation details, for example, "I'd like to make a reservation for two people starting at 7 PM." The input is the voice data, and the output is the voice data itself.

[0766] Step 6:

[0767] The device again converts the voice input into text using the SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device analyzes the voice data and generates the corresponding text.

[0768] Step 7:

[0769] The server receives the text data and stores the reservation information in a database. The input is the text data, and the output is the result of saving it to the database. Specifically, the server analyzes the text data and stores it in the database as reservation information.

[0770] Step 8:

[0771] The server recognizes the user's emotions using EmotionRecognizer and notifies the staff of the appropriate response based on that emotion. The input is the user's voice data, and the output is the emotion recognition results and the notification content. Specifically, the server analyzes the voice data, recognizes the emotion, and generates the corresponding notification.

[0772] Step 9:

[0773] The server uses location information providing means to provide directions from the user's current location to the restaurant. The input is the user's current location data, and the output is route guidance information. Specifically, the server uses the Geopy library to obtain the user's current location and calculates the optimal route to the destination.

[0774] Step 10:

[0775] The server converts the route guidance information into voice and provides the user with voice guidance. The input is the route guidance information and the output is voice data. Specifically, the text is converted into voice data using voice synthesis technology.

[0776] Example 3

[0777] Next, a third embodiment of the third embodiment will be described. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0778] Conventional automated voice response systems have difficulty generating appropriate voice responses to user requests and are unable to respond according to the user's emotions. Furthermore, they lack the functionality to save user responses as text or send SMS notifications, limiting the user experience. To solve these issues, a system is needed that integrates the generation of voice responses according to user requests, responses according to emotions, text saving, and SMS notifications.

[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0780] In this invention, the server includes means for accepting user requests by voice, means for generating questions using a voice generation system AI, means for accepting user responses by voice and converting the voice data into text and saving it, means for recognizing the user's emotions and generating voice responses in response to the emotions, and means for sending SMS notifications. This makes it possible to generate appropriate voice responses in response to user requests, respond according to emotions, save the user's responses as text, and send SMS notifications.

[0781] "User" means any person or entity making a request using the System.

[0782] "Voice-based generative AI" is an artificial intelligence technology that converts text data into voice data.

[0783] "Audio data" is data that represents audio in digital form.

[0784] "Text data" is data that represents character information in digital form.

[0785] "Emotion recognition" is a technology that analyzes a user's emotions from voice data and text data.

[0786] "Voice response" refers to a voice response to a user's request or question.

[0787] "SMS Notification" means sending information to a user using short message service.

[0788] A "database" is a system for efficiently storing, managing, and retrieving data.

[0789] A "server" is a computer system that processes data and provides services over a network.

[0790] A "terminal" is a device that is directly operated by a user and that inputs and outputs audio.

[0791] This invention is a system that integrates the generation of voice responses according to user requests, responses according to emotions, text storage of user responses, and SMS notifications. This system operates in cooperation with the server, the terminal, and the user.

[0792] The server includes a means for accepting user requests by voice, a means for generating questions using a voice generation AI system, a means for accepting user responses by voice, converting the voice data into text and saving it, a means for recognizing the user's emotions and generating corresponding voice responses, and a means for sending SMS notifications.

[0793] Specifically, the server operates as follows: First, the device accepts a request from the user via voice and sends that voice data to the server. The server then uses a voice-based AI (e.g., Google Cloud Text-to-Speech API) to generate a question for the user as voice data and sends it to the device. The device then plays back this voice data and conveys the question to the user.

[0794] Next, the user speaks a response into the device, which then sends the voice data to the server. The server uses voice recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text and store it in a database. The server then uses emotion recognition software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions and generate a voice response with a tone appropriate to the emotion. The generated voice data is then sent to the device and played back to the user.

[0795] The server also uses an SMS notification service (e.g., Twilio API) to notify the user of confirmations and order details, allowing the user to receive confirmation of their request via SMS.

[0796] As a concrete example, if a user requests to "place an order," the system operates as follows: First, the device accepts the user's request and sends it to the server. The server uses a voice generation AI to generate the question "What would you like to order?" and sends it to the device. The device plays the question and conveys it to the user. When the user replies "I'd like to order pizza," the device sends the voice data to the server. The server converts the voice data into text and stores it in a database. The server then uses emotion recognition software to analyze the user's emotions and generates a voice response in an appropriate tone. Finally, the server uses an SMS notification service to send a confirmation of the order to the user.

[0797] An example of a prompt sentence is "How does the system behave when a user requests to place an order?" By inputting this prompt sentence into the generative AI model, a sentence explaining the specific operating procedure of the system can be generated. The flow of the specific processing in Example 3 will be explained using FIG. 21.

[0798] Step 1:

[0799] The user speaks a request into the terminal.

[0800] Input: User's spoken request (e.g., "I would like to place an order")

[0801] How it works: The device receives the user's voice using a microphone and generates voice data.

[0802] Output: Generated audio data

[0803] Step 2:

[0804] The terminal transmits the generated voice data to the server.

[0805] Input: Device-generated audio data

[0806] Operation: The device sends audio data to the server over the network.

[0807] Output: Audio data sent to the server

[0808] Step 3:

[0809] The server generates questions using a voice-based AI.

[0810] Input: Audio data sent to the server

[0811] How it works: The server analyzes the voice data and recognizes the user's request. Using a voice-generating AI (e.g., Google Cloud Text-to-Speech API), it generates the question "What would you like to order?" as voice data.

[0812] Output: Audio data of the generated question

[0813] Step 4:

[0814] The server sends the generated voice data of the question to the terminal.

[0815] Input: Audio data of the generated question

[0816] Operation: The server sends the voice data of the question to the terminal via the network.

[0817] Output: Audio data of the question sent to the device

[0818] Step 5:

[0819] The terminal reproduces the audio data of the question and conveys the question to the user.

[0820] Input: Voice data of the question sent to the terminal

[0821] Operation: The device uses the speaker to play back the audio data of the question and convey it to the user.

[0822] Output: The question conveyed to the user

[0823] Step 6:

[0824] The user speaks their answer into the terminal.

[0825] Input: User's spoken response (e.g., "I'd like to order a pizza")

[0826] How it works: The device receives the user's voice using a microphone and generates voice data.

[0827] Output: Generated audio data

[0828] Step 7:

[0829] The terminal transmits the generated voice data to the server.

[0830] Input: Device-generated audio data

[0831] Operation: The device sends audio data to the server over the network.

[0832] Output: Audio data sent to the server

[0833] Step 8:

[0834] The server converts the voice data into text and stores it in a database.

[0835] Input: Audio data sent to the server

[0836] How it works: The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text, which is then stored in a database.

[0837] Output: Text data stored in a database

[0838] Step 9:

[0839] The server recognizes the user's emotions and generates a voice response accordingly.

[0840] Input: Audio data sent to the server

[0841] How it works: The server uses emotion recognition software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. Based on the analysis, it generates a voice response with an appropriate tone.

[0842] Output: Generated voice response data

[0843] Step 10:

[0844] The server transmits the generated voice response data to the terminal.

[0845] Input: Generated voice response data

[0846] Operation: The server sends voice response data to the terminal via the network.

[0847] Output: Voice response data sent to the device

[0848] Step 11:

[0849] The terminal reproduces the voice response data and conveys the response to the user.

[0850] Input: Voice response data sent to the terminal

[0851] Operation: The device uses a speaker to play back the voice response data and convey it to the user.

[0852] Output: The spoken response given to the user

[0853] Step 12:

[0854] The server sends an SMS notification.

[0855] Input: Text data stored in a database

[0856] What it does: The server uses an SMS notification service (e.g., Twilio API) to notify the user of confirmations and order details.

[0857] Output: SMS notification sent to the user

[0858] (Application example 3)

[0859] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0860] Conventional automated voice response systems were unable to respond to users' emotions, resulting in a poor user experience. They also had insufficient responses to voice orders and inquiries, making it difficult to efficiently process orders and manage orders. Furthermore, the accuracy of voice and emotion recognition was low, making it difficult to accurately understand the user's intentions.

[0861] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server is a system that links an automatic voice response (IVR) and a voice-based generation AI, and is provided for a fixed fee. It includes: a means for providing a number, voice guidance, and a generation AI linking function in an integrated manner; a means for linking a pre-prepared answer database with the voice-based generation AI, providing voice answers, SMS notifications, and asking confirmation questions, and storing the user's answers as text; a means for recognizing the user's emotions and providing voice answers and SMS notifications accordingly; and a means for allowing the user to place an order by voice, saving the order details as text, and managing the order history. This enables optimal responses according to the user's emotions and enables efficient voice order processing and history management.

[0862] An "Interactive Voice Response (IVR)" is a system that automatically responds to phone calls and voice input from users.

[0863] "Voice generative AI" is an artificial intelligence technology that generates voice data and responds to users via voice.

[0864] The "answer database" is a database that stores prepared questions and answers.

[0865] "SMS Notification" is a function that sends text messages to users using short message service.

[0866] "Emotion recognition" is a technology that analyzes emotions from a user's voice or text and identifies those emotions.

[0867] "Voice guidance" is a function that provides guidance and instructions to the user using voice.

[0868] "Save order details as text" is a function that saves the order details entered by the user via voice as text data.

[0869] "History management" is a function that stores data on past orders and inquiries so that it can be referenced as needed.

[0870] The system for implementing this invention has the following configuration. The server is a system that links an IVR (Interactive Voice Response) and a voice-based AI generator. It is provided for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and AI generator linkage functions. It also links a pre-prepared response database with the voice-based AI generator, allowing for voice responses, SMS notifications, and confirmation questions, and the text of user responses can be saved. Furthermore, it has the function of recognizing the user's emotions and providing voice responses and SMS notifications accordingly. Users can place orders by voice, and the order details can be saved as text and managed as a history.

[0871] Hardware and software used

[0872] Hardware: Smartphone (microphone, speaker)

[0873] Software: Python, speech_recognition library (voice recognition), textblob library (sentiment analysis), gTTS library (voice generation), smtplib library (SMS sending), sqlite3 library (database storage)

[0874] Processing flow

[0875] The server first uses the speech_recognition library to convert the user's speech into text for speech recognition. Next, it uses the textblob library to analyze the sentiment of the text and recognize the user's emotion. It then uses the gTTS library to generate a voice response based on the user's sentiment and plays it back from the speaker. It then uses the smtplib library to send the user's order details via SMS. Finally, it uses the sqlite3 library to save the order details in a database and manage them as a history.

[0876] Specific examples

[0877] If a user says, "I want to order a pizza," the server will ask "What would you like to order?" and if the user answers, "One Margherita pizza," the server will save the answer as text and respond with a voice response according to the user's sentiment. The server will also send the order details via SMS and store them in a database.

[0878] Prompt Sentence Examples

[0879] User: "I want to order a pizza"

[0880] Server: "What would you like to order?"

[0881] User: "One Margherita pizza."

[0882] Server: "Thank you for your order!"

[0883] In this way, a system can be implemented that efficiently processes food delivery orders based on user voice input.

[0884] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0885] Step 1:

[0886] To perform speech recognition, the server uses the speech_recognition library to convert the user's speech into text. Specifically, when the user speaks into the microphone of their smartphone, "I would like to order a pizza," the speech data is sent to the server. The server receives this speech data as input and converts it into text data using the speech_recognition library. The text "I would like to order a pizza" is obtained as output.

[0887] Step 2:

[0888] The server uses the textblob library to analyze the text data and recognize the user's emotions. Specifically, it receives the text data "I want to order pizza" obtained in step 1 as input and performs sentiment analysis using the textblob library. As a result of the sentiment analysis, it recognizes the user's sentiment as "neutral." The output is the sentiment data "neutral."

[0889] Step 3:

[0890] The server uses the gTTS library to generate a voice response that corresponds to the user's emotion. Specifically, it receives the emotion data "Neutral" obtained in step 2 and the user's order "I would like to order pizza" as input, and generates voice data using the gTTS library. The generated voice data is "What would you like to order?" and is played back from the smartphone speaker. The voice data is obtained as output.

[0891] Step 4:

[0892] The user speaks into the microphone of their smartphone, "One Margherita pizza please." The server converts this voice data into text data again using the speech_recognition library. Specifically, the server receives the user's voice data as input and converts it into text data using the speech_recognition library. The output is the text "One Margherita pizza please."

[0893] Step 5:

[0894] The server uses the smtplib library to send the user's order details via SMS. Specifically, it receives the text data "One Margherita Pizza" obtained in step 4 as input, uses the smtplib library to generate an SMS message, and sends it to the user's registered phone number. The SMS message is sent as output.

[0895] Step 6:

[0896] The server uses the sqlite3 library to save the user's order details to the database. Specifically, it receives the text data "One Margherita Pizza" obtained in step 4 as input and saves it to the database using the sqlite3 library. As output, the order details are saved to the database and managed as a history.

[0897] In this way, a system can be implemented that efficiently processes food delivery orders based on user voice input.

[0898] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0899] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0900] Another example of generative AI is Gemini (registered trademark) (Internet search engine). <url: https: gemini.google.com ?hl="ja">) are mentioned.

[0901] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0902] [Second embodiment]

[0903] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0904] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0905] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0906] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0907] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0908] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0909] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0910] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0911] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0912] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0913] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0914] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0915] "Example 1"

[0916] One embodiment of the present invention is a system that integrates an interactive voice response (IVR) with a voice-based AI generator. This system is offered for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and AI generator integration functions. Specifically, when a user calls the system, the system responds to the user's questions and requests using the voice-based AI generator. For example, if a user asks, "What are your business hours?", the system responds using the voice-based AI generator, "Business hours are from 9:00 AM to 5:00 PM."

[0917] "Example 2"

[0918] Another embodiment of the present invention is a system that can use generative AI to solve restaurant reservations, directions to the restaurant, confirmation of business hours, and phone orders. Specifically, when a user requests to "make a restaurant reservation," the system uses a voice-based generative AI to ask, "How many people would you like to make a reservation for, and what time would you like to make a reservation for?" and makes the reservation based on the user's response.

[0919] "Example 3"

[0920] In yet another embodiment of the present invention, a system is provided that links a prepared answer database with a voice-based generation AI to provide voice responses, SMS notifications, and ask questions for confirmation, and then saves the customer's responses as text. Specifically, when a user requests to "place an order," the system uses the voice-based generation AI to ask "What would you like to order?", accepts the order based on the user's response, and saves the contents as text.

[0921] The processing flow of each embodiment will be described below.

[0922] "Example 1"

[0923] Step 1: A user calls the system.

[0924] Step 2: The system responds to the user's questions and requests using a voice-generative AI.

[0925] Step 3: The user asks, "What are your business hours?"

[0926] Step 4: The system uses a voice-generating AI to respond, "Business hours are from 9:00 AM to 5:00 PM."

[0927] "Example 2"

[0928] Step 1: The user requests to make a reservation.

[0929] Step 2: The system uses a voice-generating AI to ask, "How many people would you like to make a reservation for, and what time would you like to start?"

[0930] Step 3: The system makes the reservation based on the user's answers.

[0931] "Example 3"

[0932] Step 1: The user requests to place an order.

[0933] Step 2: The system uses a voice-generative AI to ask, "What would you like to order?" Step 3: The system accepts the order based on the user's answer.

[0934] Step 4: The system saves the order details as text.

[0935] Example 1

[0936] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0937] Conventional automated voice response systems have had difficulty responding appropriately and quickly to a variety of user questions and requests. Furthermore, they lacked the ability to convert voice data into text and integrate it with generative AI models, resulting in a poor user experience. Furthermore, certain industries, such as restaurants, lacked systems that could address specific needs, such as making reservations or checking opening hours.

[0938] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0939] In this invention, the server includes means for receiving a call from a user, means for providing voice guidance, means for converting the user's voice data into text, means for transmitting the converted text to the generative AI model as a prompt sentence, means for the generative AI model to generate an answer, means for converting the generated text into speech, and means for transmitting the converted speech data to the user, thereby enabling appropriate and prompt responses to a variety of user questions and requests.

[0940] The "means for receiving a call from a user" is a function that allows the system to receive a call made by a user and start a call session.

[0941] The "means for providing voice guidance" is a function for reproducing a voice message to the user and prompting the user to perform the next operation or input.

[0942] The "means for converting user's voice data into text" is a function for converting the voice uttered by the user into character data.

[0943] "Means for sending the converted text to the generative AI model as a prompt sentence" refers to a function that inputs text data converted from speech into the generative AI model and gives instructions for generating an appropriate response.

[0944] The "means by which the generative AI model generates an answer" is the function by which the generative AI model generates an appropriate answer based on the prompt sentence.

[0945] "Means for converting generated text into speech" refers to a function for converting text data generated by a generative AI model into speech data.

[0946] The "means for transmitting converted voice data to the user" is a function for transmitting data converted into voice to the user so that the user can listen to the voice.

[0947] This invention is a system that combines an interactive voice response (IVR) and a generative AI model, and aims to respond appropriately and quickly to calls from users. A specific embodiment of this system is described below.

[0948] System configuration

[0949] The server system is configured using the following hardware and software.

[0950] Hardware: Servers, communication equipment, audio input / output devices

[0951] Software: Twilio API, Google Cloud Speech-to-Text API, Google Cloud Text-to-Speech API, generative AI models (e.g., OpenAI's GPT-3)

[0952] Program processing

[0953] The server receives a call from the user using the Twilio API. When the call connects, the server provides a voice prompt, such as "How can I help you?"

[0954] When a user speaks a question or request, the server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then sent to the generative AI model as a prompt.

[0955] The generative AI model generates appropriate answers to user questions. For example, if a user asks, "What are your business hours?", the generative AI model will respond with, "Business hours are from 9:00 AM to 5:00 PM."

[0956] The generated text response is then converted back into audio. For this, the server uses the Google Cloud Text-to-Speech API. The converted audio data is then sent to the user via the Twilio API.

[0957] Specific examples

[0958] For example, the following shows the processing when a user asks, "What are your business hours?"

[0959] 1. A user calls the system.

[0960] 2. The server receives the call using the Twilio API and provides a voice prompt saying, "Please tell us your business."

[0961] 3. A user asks, "What are your business hours?"

[0962] 4. The server converts the speech to text using the Google Cloud Speech-to-Text API.

[0963] 5. The converted text, "What are your business hours?", is input as a prompt to the generative AI model.

[0964] 6. The generative AI model generates the answer, "Business hours are 9:00 AM to 5:00 PM."

[0965] 7. The server converts the generated text response into audio using the Google Cloud Text-to-Speech API.

[0966] 8. The server sends the converted voice data to the user via the Twilio API.

[0967] Prompt Sentence Examples

[0968] "A user is asking about business hours. Generate an appropriate answer based on the following text: 'What are your business hours?'"

[0969] In this way, the system operates in cooperation with the server, terminals, and users.

[0970] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0971] Step 1:

[0972] User makes a call

[0973] The user calls the provided phone number. The input is the user making the call, and the output is the call connecting to the server.

[0974] Step 2:

[0975] The server receives the call

[0976] The server receives a call from the user using a communication API. The input is a call connection from the user, and the output is the start of a call session. The server then prepares to manage the call session.

[0977] Step 3:

[0978] The server provides audio guidance

[0979] The server provides voice guidance to the user through a communication API. The input is the start of a call session, and the output is the playback of a voice guidance message, such as "Please tell us your business."

[0980] Step 4:

[0981] Users speak their questions or requests

[0982] The user follows the voice guidance to input questions or requests by voice. The input is the user's voice, and the output is the transmission of voice data to the server. For example, a question might be, "What are your business hours?"

[0983] Step 5:

[0984] The server converts the audio data into text

[0985] The server uses a speech recognition API to convert the user's voice data into text. The input is the user's voice data, and the output is the converted text data. For example, the generated text is "What are your business hours?"

[0986] Step 6:

[0987] The server sends a prompt to the generative AI model

[0988] The server sends the converted text to the generative AI model as a prompt. The input is the converted text data, and the output is the prompt sent to the generative AI model. The prompt is, "The user is asking about business hours. Please generate an appropriate answer based on the following text: 'What are your business hours?'"

[0989] Step 7:

[0990] Generative AI models generate answers

[0991] A generative AI model generates an appropriate answer based on a prompt. The input is the prompt, and the output is a generated text answer. For example, the generated answer might be, "Business hours are from 9:00 AM to 5:00 PM."

[0992] Step 8:

[0993] The server converts the generated text to speech

[0994] The server converts the generated text response into speech using a speech synthesis API. The input is the generated text response, and the output is the converted speech data. For example, the speech generated is "Business hours are from 9:00 AM to 5:00 PM."

[0995] Step 9:

[0996] The server sends the audio data to the user

[0997] The server sends the converted voice data to the user through a communication API. The input is the converted voice data, and the output is a voice transmission to the user. The user can hear the answer over the phone.

[0998] (Application example 1)

[0999] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1000] Conventional automated voice response systems could only provide pre-set, fixed responses, making it difficult to flexibly respond to a variety of user questions and requests. Furthermore, they were unable to respond immediately to security-related questions and requests, leaving users uneasy. This resulted in a poor user experience and limited the system's value.

[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1002] In this invention, the server is a system that links an interactive voice response (IVR) with a voice-based generative AI, and is provided for a fixed fee. It includes a means for providing a number, voice guidance, generative AI linkage functions, etc. in an integrated manner, a means for speech recognition, a means for generating responses to user questions using a generative AI model, and a means for speech synthesis of the generated responses. This makes it possible to respond flexibly and immediately to a variety of user questions and requests, and in particular to quickly respond to questions and requests related to security, thereby alleviating user anxiety and improving the user experience.

[1003] An "Interactive Voice Response (IVR)" is a system that automatically responds with voice when a user calls and provides appropriate information based on the user's input.

[1004] "Voice generative AI" is an artificial intelligence technology that analyzes the user's voice input and generates an appropriate response.

[1005] "Providing a number" means providing a telephone number for a user to access.

[1006] "Voice guidance" is a function that provides voice guidance and instructions to the user.

[1007] The "generative AI integration function" is a function that links the voice-based generative AI with other systems and databases.

[1008] "Speech recognition" is a technology that converts a user's voice into text.

[1009] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on user input.

[1010] "Speech synthesis" is a technology that converts text into speech.

[1011] A "security question or request" is a question or request where a user requests information regarding the status or configuration of a security system.

[1012] The "answer database" is a database that stores prepared questions and their answers.

[1013] "Message notification" is a function that notifies users of information via text message.

[1014] "Questions for confirmation" is a function that asks the user questions about matters that need to be confirmed.

[1015] "Text save" is a function that saves the user's answers in text format.

[1016] As an embodiment of the present invention, a security assistant system will be described as an example. In this system, when a user inputs security-related questions or requests by voice, a voice-based AI system responds immediately.

[1017] Hardware and software used

[1018] Hardware: Smartphone (microphone, speaker)

[1019] software:

[1020] speech_recognition library: Used to perform speech recognition.

[1021] pyttsx3 library: Used to perform speech synthesis.

[1022] openai library: Uses generative AI models using the GPT-3 API.

[1023] System Operation

[1024] 1. Voice recognition: A user speaks a security question or request into the smartphone microphone, for example, "What is the status of my home security system?"

[1025] 2. Speech to text conversion: Uses the speech_recognition library to convert the user's speech into text.

[1026] 3. Response generation using a generative AI model: The converted text is sent to GPT-3 as a prompt to generate an appropriate response. An example of a prompt is "Security system question: What is the status of my home security system?"

[1027] 4. Speech synthesis: The generated response is converted into speech using the pyttsx3 library and transmitted back to the user through the smartphone speaker.

[1028] Specific examples

[1029] When a user asks, "What's the status of my home security system?" the system works as follows:

[1030] 1. The smartphone microphone captures the user's voice.

[1031] 2. The speech_recognition library converts the speech to text, generating the text "What is the status of my home security system?"

[1032] 3. The generated text is sent as a prompt to GPT-3, which generates an appropriate response, such as "All sensors are currently working properly."

[1033] 4. The pyttsx3 library converts the generated response into speech and delivers it back to the user through the smartphone speaker.

[1034] In this way, the security assistant system can respond immediately to users' security questions and requests, alleviating their concerns and improving their experience.

[1035] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1036] Step 1:

[1037] The user speaks their security question or request into the smartphone's microphone.

[1038] Input: User's voice

[1039] Output: Audio data captured by a microphone

[1040] Specific Action: A user says, "What's the status of my home security system?"

[1041] Step 2:

[1042] The device uses the speech_recognition library to convert the captured audio data into text.

[1043] Input: Audio data

[1044] Output: Text data

[1045] What it does: The speech recognition engine analyzes the voice data and generates text such as "What is the status of my home security system?"

[1046] Step 3:

[1047] The device sends the generated text as a prompt to GPT-3, which generates an appropriate response.

[1048] Input: Text data (prompt sentence)

[1049] Output: Response text

[1050] What it does: Sends the prompt "Security System Question: What's the status of my home security system?" to the GPT-3 API and receives the response "Currently, all sensors are working properly."

[1051] Step 4:

[1052] The device converts the generated response text into speech using the pyttsx3 library.

[1053] Input: Response text

[1054] Output: Audio data

[1055] What happens: The speech synthesis engine converts the text "All sensors are currently working properly" into speech data.

[1056] Step 5:

[1057] The device responds to the user with voice data via the smartphone speaker.

[1058] Input: Audio data

[1059] Output: The audio the user hears

[1060] What happens: The smartphone speaker will play a voice message saying "All sensors are currently working properly."

[1061] In this way, the security assistant system can respond immediately to the user's security questions and requests.

[1062] Example 2

[1063] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1064] Conventional automated voice response systems have difficulty responding flexibly and quickly to user requests, and efficient responses are required, especially for complex requests such as making restaurant reservations or confirming opening hours. Furthermore, there is a lack of technology to accurately understand users' voice requests and generate appropriate responses, making it difficult to improve the user experience.

[1065] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting a user's voice request, means for converting voice to text, means for generating and transmitting a prompt sentence to the generative AI model, means for receiving a response from the generative AI model, means for converting the text response to voice data, means for transmitting the voice data to the user, means for accepting the user's answer, means for registering reservation information, means for generating a reservation confirmation message and converting it into voice data, and means for transmitting the voice data to the user. This makes it possible to respond quickly and accurately to user requests and efficiently handle complex requests such as making a reservation at a restaurant or confirming business hours.

[1066] The "means for accepting a user's voice request" refers to a device or software that recognizes the voice uttered by the user and inputs the content of that voice into the system.

[1067] A "speech-to-text converter" is a device or software that uses speech recognition technology to convert a user's speech into textual information.

[1068] A "means for generating and sending prompts to a generative AI model" is a device or software for generating appropriate questions or instructions based on a user's request and sending them to a generative AI model.

[1069] A "means for receiving a response from a generative AI model" is a device or software for receiving a response returned from a generative AI model.

[1070] A "means for converting text responses into audio data" is a device or software for converting text responses from a generative AI model into audio data.

[1071] "Means for transmitting voice data to a user" refers to a device or software for transmitting the generated voice data to a user's terminal and conveying it to the user.

[1072] "Means for accepting user responses" refers to a device or software that recognizes additional voice information provided by the user and incorporates that information into the system.

[1073] The "means for registering reservation information" refers to a device or software for registering reservation information received from a user in a database or the like.

[1074] The "means for generating a reservation confirmation message and converting it into voice data" refers to a device or software that generates a message to notify the user that the reservation has been confirmed and converts it into voice data.

[1075] The "means for transmitting voice data to the user" refers to a device or software for transmitting the generated voice data to the user's terminal and conveying the contents of the reservation confirmation to the user.

[1076] This invention is a system that uses a generative AI model to solve restaurant reservations, restaurant directions, checking business hours, and phone orders. This system accepts a user's voice request, converts it into text, generates and sends a prompt to the generative AI model, and then generates an appropriate response to provide to the user.

[1077] Hardware and software used

[1078] Hardware: Servers, user devices (smartphones, tablets, PCs, etc.)

[1079] Software: Generative AI models (e.g., large-scale language models), speech recognition software (e.g., speech recognition engines), and speech synthesis software (e.g., speech synthesis engines).

[1080] Specific operation of the system

[1081] 1. Acceptance of user requests

[1082] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[1083] Device: Uses speech recognition software to convert your speech into text.

[1084] Terminal: Sends the converted text to the server.

[1085] 2. Prompt generation of generative AI models

[1086] Server: Parses the received text and generates prompts to send to the generative AI model.

[1087] Server: Generate an example prompt: "A user would like to make a reservation. How many people would you like to book and what time would you like to make the reservation for?"

[1088] 3. Response generation using generative AI models

[1089] Server: Sends the generated prompts to the generative AI model.

[1090] Generative AI model: Generates appropriate responses based on prompts.

[1091] Generative AI model: Generates an example response: "How many people would you like to make a reservation for, and what time?"

[1092] 4. Voice synthesis of responses

[1093] Server: Passes the generated text response to speech synthesis software to generate audio data.

[1094] 5. Sending a response to the user

[1095] Server: Sends the generated voice data to the user device.

[1096] Terminal: Plays the audio data and communicates the response to the user.

[1097] 6. Acceptance of user responses

[1098] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[1099] Device: Uses speech recognition software to convert your speech into text.

[1100] Terminal: Sends the converted text to the server.

[1101] 7. Confirmation of reservation

[1102] Server: Receives the user's response and registers the information in the reservation system.

[1103] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[1104] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We are waiting for you with XX guests on XX / XX / XX at XX time."

[1105] 8. Sending a confirmation message

[1106] Server: The confirmation message is converted into voice using text-to-speech software.

[1107] Text-to-speech software: converts text into audio data.

[1108] Server: Sends the generated voice data to the user device.

[1109] Terminal: Plays audio data to inform the user that the reservation is complete.

[1110] Specific examples

[1111] User request: "I want to make a reservation to visit the store."

[1112] Prompt to generative AI model: "A user wants to make a reservation. How many people would like to come and what time would they like to make the reservation?"

[1113] The generative AI model responds: "How many people would you like to book and what time would you like to book?"

[1114] User's answer: "Two people, starting tomorrow at 7pm."

[1115] Reservation confirmation message: "Your reservation is complete. We look forward to seeing you tomorrow at 7pm for two people."

[1116] In this way, the system uses a generative AI model to generate an appropriate response to a user's request and confirm the reservation.

[1117] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1118] Step 1: Accepting a user request

[1119] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[1120] Input: User's voice request

[1121] Device: Uses speech recognition software to convert your speech into text.

[1122] Data processing: Converting voice data into text data

[1123] Output: "I'd like to make an appointment"

[1124] Terminal: Sends the converted text to the server.

[1125] Step 2: Prompt generation for generative AI models

[1126] Server: Parses the received text and generates prompts to send to the generative AI model.

[1127] Input: "I'd like to make an appointment"

[1128] Data Calculation: Text Analysis and Prompt Generation

[1129] Output: A prompt saying "A user wants to make an appointment. How many people would like to book and what time would you like to book?"

[1130] Server: Sends prompts to the generative AI model.

[1131] Step 3: Generate a response using a generative AI model

[1132] Server: Sends the generated prompts to the generative AI model.

[1133] Input: The prompt "A user wants to make a reservation. How many people and what time would you like to book?"

[1134] Generative AI model: Generates appropriate responses based on prompts.

[1135] Data Calculation: Prompt-Based Response Generation

[1136] Output: Response "How many people would you like to book and what time?"

[1137] Generative AI model: Sends the response to the server.

[1138] Step 4: Speech synthesis to vocalize the response

[1139] Server: Passes the generated text response to speech synthesis software to generate audio data.

[1140] Input: Text response "How many people would you like to book and what time?"

[1141] Text-to-speech software: converts text into audio data.

[1142] Data processing: Convert text data into audio data

[1143] Output: Audio data

[1144] Server: Sends the generated voice data to the user device.

[1145] Step 5: Send a response to the user

[1146] Server: Sends the generated voice data to the user device.

[1147] Input: Audio data

[1148] Terminal: Plays the audio data and communicates the response to the user.

[1149] Output: A voice response saying "How many people would you like to book and what time would you like to book?"

[1150] Step 6: Accept user responses

[1151] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[1152] Input: User's spoken response

[1153] Device: Uses speech recognition software to convert your speech into text.

[1154] Data processing: Converting voice data into text data

[1155] Output: Text "Two people, starting tomorrow at 7pm"

[1156] Terminal: Sends the converted text to the server.

[1157] Step 7: Confirm your booking

[1158] Server: Receives the user's response and registers the information in the reservation system.

[1159] Input: Text "Two people, please meet tomorrow from 7pm"

[1160] Data calculation: Reservation information registration

[1161] Output: Reservation information registration completed

[1162] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[1163] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We look forward to seeing you tomorrow at 7pm for two guests."

[1164] Output: Confirmation message

[1165] Step 8: Send a confirmation message

[1166] Server: The confirmation message is converted into voice using text-to-speech software.

[1167] Input: Confirmation message text

[1168] Text-to-speech software: converts text into audio data.

[1169] Data processing: Convert text data into audio data

[1170] Output: Audio data

[1171] Server: Sends the generated voice data to the user device.

[1172] Terminal: Plays audio data to inform the user that the reservation is complete.

[1173] Output: Voice response "Your reservation is complete. We will be waiting for you tomorrow at 7pm for two people."

[1174] (Application example 2)

[1175] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1176] In the past, restaurant operations such as making reservations, getting directions to the restaurant, checking business hours, and calling to order were often done manually, which was time-consuming and labor-intensive. In addition, users had to use multiple methods to obtain this information, which was inconvenient.

[1177] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an automatic voice response means, a voice-based generative AI means, a voice recognition means, a response generation means using a generative AI model, a means for making reservations based on a user's voice input, a means for providing directions, a means for checking business hours, and a means for automatically making an order call. This allows users to efficiently perform tasks such as making store reservations, getting directions to stores, checking business hours, and making an order call through a single system.

[1178] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[1179] A "voice generation AI means" is a system that uses artificial intelligence technology to analyze voice input and generate an appropriate voice response.

[1180] "Speech recognition means" refers to a system that has the technology to convert a user's voice into text data.

[1181] A "response generation means using a generative AI model" is a system that has the function of using a generative AI model to generate an appropriate response based on user input.

[1182] The "means for making a reservation based on a user's voice input" is a system that has the function of processing a reservation request made by a user through voice and confirming the reservation.

[1183] A "means for providing directions" is a system that has the function of providing users with directions to their destination via voice or text.

[1184] The "means for checking business hours" is a system that has the function of checking the business hours of a store specified by the user and providing the information to the user.

[1185] "Means for automatically making an order call" refers to a system that has the function of automatically transmitting the specified order details over the phone based on the user's voice input.

[1186] The "answer database" is a database that stores answers to questions prepared in advance.

[1187] "Short Message Service Notification" means a system that has the function of notifying users via short message service.

[1188] The "means for asking questions to confirm matters" is a system that has the function of asking the user necessary questions to confirm matters and obtaining the answers.

[1189] "Means that allow for text saving of user responses" refers to a system that has the function of saving the content of a user's voice responses as text data.

[1190] The following system configuration will be described as an embodiment of the present invention.

[1191] System Configuration

[1192] This system includes a server, a user terminal, and a voice recognition device. The server includes a response generation means using a generative AI model, a voice-based generative AI means, a voice recognition means, and a database. The user terminal includes a microphone and a speaker and receives voice input from the user.

[1193] Hardware and software used

[1194] Hardware:

[1195] Microphone: Accepts user voice input.

[1196] Speaker: Provides audio responses from the server to the user.

[1197] Server: Provides the computational resources to run generative AI models.

[1198] software:

[1199] speech_recognition library: Performs speech recognition.

[1200] The transformers library: Runs generative AI models.

[1201] Data processing and calculation

[1202] 1. Speech Recognition:

[1203] A microphone on the user terminal accepts the user's voice input.

[1204] A speech recognition means converts the speech input into text data.

[1205] 2. Response generation:

[1206] A response generation means using the server's generative AI model analyzes the text data obtained from the speech recognition means and generates an appropriate response.

[1207] A voice generation AI means converts the generated response into voice data.

[1208] 3. Booking Process:

[1209] Based on the user's voice input, the means for making reservations stores the reservation information in a database.

[1210] 4. Directions and business hours:

[1211] The server provides the necessary information in response to a user's request, using means for providing directions and means for checking opening hours.

[1212] 5. Order by phone:

[1213] The server uses a means for automatically placing an order call based on the user's voice input to convey the specified order details over the phone.

[1214] Specific examples

[1215] When a user says, "I would like to make a reservation," the system operates as follows.

[1216] 1. The microphone on the user device accepts voice input.

[1217] 2. A speech recognition tool converts the speech into text.

[1218] 3. The server's generative AI model generates a response generation method that asks the question, "How many people would you like to make a reservation for, and what time would you like to start?"

[1219] 4. The voice generation AI means converts the question into voice data and provides it to the user through the speaker on the user's device.

[1220] 5. When the user answers "Two people, starting at 7pm", the speech recognition means again converts the speech into text and saves the reservation information in the database.

[1221] Prompt Sentence Examples

[1222] A user says, "I'd like to schedule an appointment." What's your next question?

[1223] In this way, the system can efficiently perform tasks such as making reservations, getting directions, checking business hours, and placing orders over the phone, based on the user's voice input.

[1224] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1225] Step 1:

[1226] The user inputs "I would like to make a reservation" by voice. The microphone on the user terminal accepts this voice input. The input is the user's voice data, and the output is the voice data itself.

[1227] Step 2:

[1228] The voice recognition means of the user terminal converts the received voice data into text data. The input is voice data, and the output is text data such as "I would like to make a reservation to visit the restaurant."

[1229] Step 3:

[1230] The server's generative AI model is used to generate a response, which analyzes the text data and generates an appropriate question. The input is the text "I would like to make a reservation," and the output is the text "How many people would you like to make a reservation for, and what time would you like to make a reservation for?"

[1231] Step 4:

[1232] The server's voice generation AI means converts the generated text data of the question into voice data. The input is text data such as "How many people would you like to make a reservation for, and what time would you like to start?", and the output is voice data.

[1233] Step 5:

[1234] The speaker on the user's device plays the audio data sent from the server and presents the question to the user. The input is the audio data, and the output is the audio heard by the user.

[1235] Step 6:

[1236] The user responds by voice, "Two people, starting at 7 PM." The microphone on the user's device accepts this voice input. The input is the user's voice data, and the output is the voice data itself.

[1237] Step 7:

[1238] The voice recognition means of the user terminal converts the received voice data into text data. The input is voice data, and the output is text data such as "Two people, starting at 7pm."

[1239] Step 8:

[1240] The server's reservation processing means analyzes the text data and saves the reservation information in a database. The input is text data such as "2 people, starting at 7 PM," and the output is a database containing the reservation information.

[1241] Step 9:

[1242] The server generates voice data to notify the user that the reservation has been completed, and converts it into voice data using a voice generation AI means. The input is text data saying "Reservation completed," and the output is voice data.

[1243] Step 10:

[1244] The speaker on the user's terminal plays the audio data sent from the server and notifies the user that the reservation has been completed. The input is audio data, and the output is the audio heard by the user.

[1245] In this way, the system can efficiently make reservations based on the user's voice input.

[1246] Example 3

[1247] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1248] Conventional automated voice response systems struggled to generate appropriate voice questions in response to user requests and efficiently store user responses as text data. They also lacked the functionality to automatically send SMS notifications to confirm user responses or ask additional questions for confirmation. This resulted in a poor user experience and impaired system efficiency.

[1249] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1250] In this invention, the server includes a means for receiving a request from a user, a means for generating a voice question using a voice generation AI, a means for receiving a user's response, a means for converting the received response into text data and saving it, a means for sending an SMS notification, and a means for generating and sending a confirmation question. This makes it possible to generate an appropriate voice question in response to a user's request and efficiently save the user's response as text data. Furthermore, by automatically sending an SMS notification or asking additional confirmation questions, it is possible to improve the user experience and increase the efficiency of the system.

[1251] The "means for receiving requests from a user" is a device or software for receiving requests made by a user to the system in voice or text format.

[1252] "Voice generation AI" is an artificial intelligence technology for converting text data into voice data, and is a system that can ask questions and provide guidance to users in a natural voice.

[1253] A "means for generating voice questions" is a device or software that uses a voice generation AI to generate appropriate questions in voice format based on the user's request.

[1254] A "means for receiving a user's response" is a device or software for receiving a response given by a user in voice or text form.

[1255] The "means for converting received responses into text data and saving the same" refers to a device or software for converting the user's voice responses into text data and saving the text data in a storage device such as a database.

[1256] A "means for sending SMS notifications" is a device or software for sending notifications to a user using Short Message Service (SMS).

[1257] A "means for generating and transmitting confirmation questions" is a device or software for generating additional questions to confirm the user's answers and transmitting them to the user in voice or text format.

[1258] MODE FOR CARRYING OUT THE INVENTION

[1259] This invention is a system that receives requests from users, generates voice questions using a voice generation AI, receives user responses, converts them into text data, and saves them. It also includes the functionality to send SMS notifications and generate and send confirmation questions.

[1260] Hardware and software used

[1261] Hardware: Servers, user devices (smartphones, tablets, PCs, etc.)

[1262] Software: Generative voice AI (e.g., Google Cloud Text-to-Speech API), speech recognition technology (e.g., Google Cloud Speech-to-Text API), SMS notification systems (e.g., Twilio), database management systems (e.g., MySQL)

[1263] Specific operation of the system

[1264] The server receives requests from users. For example, when a user uses a smartphone to request, "I would like to place an order," the request is sent to the server. The server uses a voice generation AI to generate a voice question such as, "What would you like to order?" and sends the voice data to the user's device.

[1265] The terminal plays the voice question sent from the server to the user. When the user answers "I'd like to order pizza," the terminal receives the voice data and sends it to the server.

[1266] The server analyzes the user's voice response received from the terminal and converts it into text data using voice recognition technology. For example, the generated text data is "I would like to order pizza." This text data is then stored in a database management system.

[1267] The server uses an SMS notification system to send a confirmation message to the user to confirm the user's order, for example, "Your pizza order has been received. Please reply to confirm."

[1268] The server receives the user's reply and asks additional questions if necessary. For example, a question like "What kind of pizza do you want?" is generated using a voice-based AI and sent to the device. The device then plays this question back to the user and sends the user's answer back to the server.

[1269] Specific examples

[1270] Example: Restaurant ordering system

[1271] 1. The user uses their smartphone to request an order.

[1272] 2. The server uses a voice generation AI to generate a voice question such as "What would you like to order?" and sends it to the terminal.

[1273] 3. The user responds, "I would like to order pizza," and the voice data is sent to the server via the device.

[1274] 4. The server converts the voice data into text and stores the message "I would like to order pizza" in the database.

[1275] 5. The server sends an SMS to the user saying, "Your pizza order has been accepted. Please reply to confirm."

[1276] 6. When the user replies, the server generates a follow-up question, "What kind of pizza do you want?" and sends it to the device.

[1277] Example prompts for generative AI models

[1278] When a user requests to "order," please describe a system that uses a voice-generative AI to ask "What would you like to order?" and saves the user's response as text.

[1279] The flow of the identification process in the third embodiment will be described with reference to FIG.

[1280] Program processing flow

[1281] Step 1: Receiving a user request

[1282] Subject: Terminal

[1283] The terminal receives a request from the user saying, "I would like to place an order." When the user speaks "I would like to place an order" into the microphone of their smartphone, the voice data is input into the terminal. The terminal then sends this voice data to the server.

[1284] Input: User's voice request

[1285] Output: Sending audio data to the server

[1286] Step 2: Generate a voice question

[1287] Subject: Server

[1288] The server analyzes the user's request received from the device and determines the next action to take. The server uses a voice generation AI to generate a voice question such as "What would you like to order?" This voice data is sent from the server to the device.

[1289] Input: User's voice request data

[1290] Output: Generated voice question data

[1291] Step 3: Receiving the user's response

[1292] Subject: Terminal

[1293] The terminal plays the voice question sent from the server to the user. When the user answers "I would like to order pizza," the terminal receives the voice data. The terminal then transmits this voice data to the server.

[1294] Input: Voice question data from the server, user's voice response

[1295] Output: Sending voice response data to the server

[1296] Step 4: Save the answer as text

[1297] Subject: Server

[1298] The server analyzes the user's voice response received from the terminal and converts it into text data using voice recognition technology. For example, the generated text data is "I would like to order pizza." This text data is then stored in a database management system.

[1299] Input: User's voice response data

[1300] Output: Save the converted text data

[1301] Step 5: Sending SMS notifications

[1302] Subject: Server

[1303] The server uses an SMS notification system to send a confirmation message to the user to confirm the user's order, for example, "Your pizza order has been received. Please reply to confirm."

[1304] Input: Text data (order details)

[1305] Output: SMS notification to the user

[1306] Step 6: Verification Questions

[1307] Subject: Server

[1308] The server receives the user's reply and asks additional questions if necessary. For example, a question like "What kind of pizza do you want?" is generated using a voice-based AI and sent to the device. The device then plays this question back to the user and sends the user's answer back to the server.

[1309] Input: User reply data

[1310] Output: Send generated additional question data

[1311] (Application example 3)

[1312] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1313] Conventional food delivery systems require users to manually input their orders, making operation cumbersome. It is also difficult to confirm or change order details, which can lead to a poor user experience. Furthermore, order details are not confirmed or notified in real time, making it difficult for users to understand the status of their order.

[1314] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes an automatic voice response means, a voice version generation system AI means, a means for converting voice input into text, a means for saving the text, a means for sending SMS notifications, and a means for asking confirmation questions. This allows the user to easily place an order by voice, and the order details can be confirmed or changed in real time, improving the user experience. In addition, the order details are saved as text and notified via SMS, making it easier for the user to understand the order status.

[1315] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[1316] "Voice generation AI means" is a system that uses artificial intelligence technology to analyze voice data and generate appropriate voice responses.

[1317] A "means for converting voice input into text" is a system that has the ability to recognize a user's voice and convert the content into text data.

[1318] The "means for storing text" is a system that has the function of storing the converted text data in a database or file system.

[1319] "Means for sending SMS notifications" refers to a system that has the function of sending notifications to users via short message service (SMS) based on stored text data.

[1320] The "means for asking confirmation questions" is a system that has the function of asking additional confirmation questions by voice based on the user's input and receiving responses from the user.

[1321] A system for implementing this invention has the following configuration: The server includes an automatic voice response means, a voice version generation system AI means, a means for converting voice input into text, a means for saving the text, a means for sending SMS notifications, and a means for asking confirmation questions.

[1322] Hardware and software used

[1323] Hardware: Smartphone (microphone, speaker)

[1324] Software: Python, Flask (web framework), SpeechRecognition (voice recognition library), pyttsx3 (speech synthesis library), smtplib (email sending library)

[1325] Processing flow

[1326] 1. User voice input:

[1327] A user launches a smartphone app and places an order by voice. For example, the user says, "I'd like to order a pizza."

[1328] 2. Speech Recognition:

[1329] The smartphone's microphone captures the user's voice and converts it into text using a speech recognition library (SpeechRecognition). For example, the speech "I want to order a pizza" is converted into text "I want to order a pizza."

[1330] 3. Response by voice-generative AI:

[1331] A speech-based AI solution generates appropriate responses based on the user's voice input and asks the user questions through the smartphone speaker, such as "What would you like to order?"

[1332] 4. User Answer:

[1333] The user responds, "One Margherita pizza." Again, the speech recognition library is used to convert this speech to text.

[1334] 5. Save text:

[1335] The converted text data is stored in a database or file system on the server. For example, the text "One Margherita pizza" is stored.

[1336] 6. SMS notification:

[1337] Based on the saved text data, the user is notified via short message service (SMS). For example, an SMS message with the content "Order: 1 Margherita pizza" is sent to the user.

[1338] Specific examples

[1339] When a user says, "I'd like to order a pizza," the system asks aloud, "What would you like to order?", and the user replies, "One Margherita pizza." The system saves this as text and sends an SMS message saying, "Order: One Margherita pizza."

[1340] Prompt Sentence Examples

[1341] User: "I want to order a pizza."

[1342] System: "What would you like to order?"

[1343] User: "One Margherita pizza."

[1344] System: "Your order has been accepted. Order: 1 Margherita pizza."

[1345] In this way, a food delivery application can be realized that allows users to easily place orders by voice.

[1346] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1347] Step 1:

[1348] A user launches a smartphone app and places an order by voice. The user says, "I'd like to order a pizza." The smartphone's microphone captures this voice. The input is the user's voice data, and the output is the captured voice data.

[1349] Step 2:

[1350] The smartphone sends the captured voice data to a speech recognition library (SpeechRecognition), which converts the voice into text. The input is the captured voice data, and the output is the text data "I would like to order a pizza." The server receives this text data and proceeds to the next step.

[1351] Step 3:

[1352] The server uses a voice generation AI method to generate an appropriate response based on the user's text input. For example, it generates a voice response such as "What would you like to order?" The input is text data such as "I would like to order pizza," and the output is voice data such as "What would you like to order?" The server then sends this voice data to the smartphone.

[1353] Step 4:

[1354] The smartphone plays the voice data received from the server and asks the user, "What would you like to order?" The input is the voice data sent from the server, and the output is the voice question to the user. The user answers, "One Margherita pizza."

[1355] Step 5:

[1356] The smartphone captures the user's answer again and converts it into text using a speech recognition library. The input is the user's voice data, and the output is the text data "One Margherita pizza." The server receives this text data and proceeds to the next step.

[1357] Step 6:

[1358] The server saves the converted text data in a database or file system. The input is the text data "one Margherita pizza" and the output is the saved text data. The saved data can be referenced later.

[1359] Step 7:

[1360] The server notifies the user via short message service (SMS) based on the saved text data. The input is the saved text data, and the output is an SMS with the content "Order: 1 Margherita pizza." The server sends this SMS to the user's smartphone.

[1361] In this way, a system is realized that allows users to easily place orders by voice and receive real-time confirmation and notification of order details.

[1362] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1363] "Example 1"

[1364] One embodiment of the present invention is a system that incorporates an emotion engine. This system recognizes emotions from the user's tone of voice and language and responds accordingly. For example, if the system determines that the user is feeling angry or frustrated, it responds with more polite language and a softer tone. On the other hand, if the system determines that the user is feeling happy or excited, it responds with a more lively tone. This allows for optimal responses according to the user's emotions.

[1365] "Example 2"

[1366] Another embodiment of the present invention is a system that optimizes restaurant service according to the user's emotions. This system recognizes the user's emotions and optimizes the restaurant's service accordingly. For example, if the system determines that the user is feeling angry or frustrated, it conveys this information to the restaurant staff, encouraging them to provide more courteous service. If the system determines that the user is feeling happy or excited, it conveys this information to the restaurant staff, encouraging them to provide more lively service. This makes it possible to provide optimal service according to the user's emotions.

[1367] "Example 3"

[1368] Furthermore, another embodiment of the present invention is a system that provides a voice response or SMS notification according to a user's emotions. This system recognizes the user's emotions and provides a voice response or SMS notification accordingly. For example, if the system determines that the user is feeling angry or frustrated, it provides a voice response using more polite language and a softer tone. On the other hand, if the system determines that the user is feeling happy or excited, it provides a voice response in a more lively tone. This enables optimal voice responses and SMS notifications according to the user's emotions.

[1369] The processing flow of each embodiment will be described below.

[1370] "Example 1"

[1371] Step 1: Receive speech input from the user.

[1372] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[1373] Step 3: Adjust the tone and wording of your voice response depending on the emotion you recognize.

[1374] "Example 2"

[1375] Step 1: Receive speech input from the user.

[1376] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[1377] Step 3: Communicate the recognized emotions to restaurant staff to encourage them to optimize their service.

[1378] "Example 3"

[1379] Step 1: Receive speech input from the user.

[1380] Step 2: Use the emotion engine to recognize the user's emotion from the voice input.

[1381] Step 3: Tailor the content of voice responses and SMS notifications depending on the recognized emotion.

[1382] Example 1

[1383] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1384] Conventional automated voice response systems could only provide fixed responses to user questions and requests, making it difficult to respond flexibly to the user's emotions. Furthermore, the process of converting the user's voice data into text and generating appropriate responses using a generative AI model was complex, creating a need for an efficient system. Furthermore, certain industries, such as restaurants, required systems that could respond to specific needs, such as making reservations and checking opening hours.

[1385] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1386] In this invention, the server includes means for linking an automated voice response with a voice generation AI, means for providing the service for a fixed fee, means for providing a number, means for providing voice guidance, means for linking a voice generation AI function, means for converting user voice data into text, means for transmitting a prompt to the AI ​​generation model, means for the AI ​​generation model to generate an answer, means for converting the generated answer into voice, means for providing the answer to the user in voice, means for analyzing the user's emotions, means for instructing the AI ​​generation model to respond according to the emotion, means for generating an answer according to the emotion, means for converting the answer according to the emotion into voice, and means for providing an answer in voice according to the emotion. This makes it possible to provide appropriate answers to user questions and requests and to respond flexibly according to the user's emotions.

[1387] An "automatic voice response" is a system that automatically responds to phone calls and voice inputs from users.

[1388] "Speech generation AI" is an AI technology that generates natural-sounding speech based on user text input and prompts.

[1389] A "fee-based means" is a means that indicates that a service or system is available for a set fee.

[1390] "Means for providing numbers" refers to means for providing telephone numbers or identification numbers for users to access the system.

[1391] The "audio guidance means" is a means for providing guidance and instructions to the user by voice.

[1392] "Generative AI linking function means" is a means for linking speech-generating AI with other systems and functions.

[1393] The "means for converting user's voice data into text" refers to a means for converting voice data input by a user into text format.

[1394] The "means for sending a prompt sentence to the generative artificial intelligence model" is a means for sending a prompt sentence including an instruction or question to the generative artificial intelligence model.

[1395] "Means by which a generative artificial intelligence model generates an answer" refers to means by which a generative artificial intelligence model generates an appropriate answer based on a prompt sentence.

[1396] The "means for converting the generated answer into speech" refers to a means for converting the text-format answer generated by the generative artificial intelligence model into speech format.

[1397] The "means for providing a user with an answer in the form of voice" refers to a means for providing a user with an answer that has been converted into voice format.

[1398] The "means for analyzing user emotions" refers to a means for analyzing emotions from user voice data and text data.

[1399] The "means for instructing the generative artificial intelligence model to respond according to the emotion" is a means for instructing the generative artificial intelligence model to respond appropriately based on the emotion of the user.

[1400] The "means for generating an answer according to emotions" is a means for a generative artificial intelligence model to generate an appropriate answer according to the user's emotions.

[1401] The "means for converting a response according to emotion into voice" is a means for converting a response in text format generated according to emotion into voice format.

[1402] The "means for providing an answer in a voice corresponding to an emotion" is a means for providing a user with an answer in the form of a voice corresponding to an emotion.

[1403] This invention is a system that combines an automatic voice response system with a speech generation artificial intelligence (AI) to provide appropriate answers to user questions and requests, and also enables flexible responses according to the user's emotions. This system is implemented using the following hardware and software.

[1404] Hardware and software used

[1405] Server: A central processing unit that manages the entire system and calls various APIs.

[1406] Twilio API: A communications API for managing incoming and outgoing phone calls.

[1407] Google Cloud Speech-to-Text API: A speech recognition API for converting user voice data into text.

[1408] Generative AI model (GPT-4): An artificial intelligence model for generating appropriate answers to user questions and requests.

[1409] Google Cloud Text-to-Speech API: A speech synthesis API for converting generated text responses into audio.

[1410] Emotion recognition engine (IBM Watson Tone Analyzer): An engine for analyzing emotions from user voice and text data.

[1411] Specific operation of the system

[1412] 1. The user makes a call

[1413] The user calls the provided phone number, for example, a customer support phone number.

[1414] 2. The server receives the call

[1415] The server receives a call from the user using the Twilio API. If the call connects, the server proceeds to the next step.

[1416] 3. The server provides audio guidance

[1417] The server provides voice prompts to the user through the Twilio API, for example, playing a message such as "Please tell us your business."

[1418] 4. The user speaks their question or request

[1419] The user follows the voice prompts to input a question or request by voice, for example, "Please tell me how to return an item."

[1420] 5. The server converts the audio data into text

[1421] The server uses the Google Cloud Speech-to-Text API to convert the user's voice data into text, which becomes "How do I return this item?"

[1422] 6. The server sends a prompt to the generative AI model

[1423] The server sends a prompt to the generative AI model (GPT-4). For example, it sends a prompt such as, "The user is asking how to return a product. Please tell me how to return it."

[1424] 7. Generative AI models generate answers

[1425] The generative AI model generates an appropriate answer based on the prompt, for example, "To return the product, bring it to the store with your receipt within 30 days of purchase."

[1426] 8. The server converts the answer into audio

[1427] The server converts the generated answer into audio using the Google Cloud Text-to-Speech API.

[1428] 9. The server provides the user with a spoken response

[1429] The server provides the user with a response converted into voice via the Twilio API. For example, it might say, "To return the product, please bring it to the store with your receipt within 30 days of purchase."

[1430] 10. The server uses an emotion recognition engine to analyze the user's emotions.

[1431] The server sends the user's voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to analyze the user's emotions.

[1432] 11. The server instructs the generative AI model on how to respond based on the emotion.

[1433] The server sends instructions to the generative AI model based on the results of the emotion recognition engine, for example, sending a prompt such as "The user is angry. Please respond in a polite manner."

[1434] 12. Generative AI models generate emotionally relevant answers

[1435] The generative AI model generates appropriate responses based on the customer's sentiment, such as, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase."

[1436] 13. The server converts the response into speech based on the emotion.

[1437] The server uses the Google Cloud Text-to-Speech API to convert the emotional response into speech.

[1438] 14. The server provides the user with a response in voice based on their emotions.

[1439] The server provides the user with a response based on the emotion through the Twilio API. For example, it might say, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase."

[1440] Prompt Sentence Examples

[1441] "Users are asking about business hours. What are your business hours?"

[1442] "A user is asking how to return an item. How do I return it?"

[1443] "User is angry. Please respond politely."

[1444] In this way, the server provides appropriate answers to the user's questions and requests, realizing a system that responds according to emotions.

[1445] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1446] Step 1:

[1447] User makes a call

[1448] The user calls a provided phone number, for example a customer support phone number. The input is the user's phone number and the call start time, and the output is the call connection to the server.

[1449] Step 2:

[1450] The server receives the call

[1451] The server receives a call from the user using the Twilio API. Once the call connects, the server proceeds to the next step. The input is the call signal from the user, and the output is the establishment of the call connection.

[1452] Step 3:

[1453] The server provides audio guidance

[1454] The server provides a voice prompt to the user through the Twilio API. For example, it plays a message such as "Please tell us your business." The input is the establishment of a call connection, and the output is the playback of the voice prompt.

[1455] Step 4:

[1456] Users speak their questions or requests

[1457] The user inputs a question or request by voice, following the voice guidance. For example, say, "Please tell me how to return a product." The input is the voice guidance, and the output is the user's voice data.

[1458] Step 5:

[1459] The server converts the audio data into text

[1460] The server converts the user's voice data into text using the Google Cloud Speech-to-Text API. The converted text is "How do I return the item?" The input is the user's voice data, and the output is text data.

[1461] Step 6:

[1462] The server sends a prompt to the generative AI model

[1463] The server sends a prompt to the generative AI model (GPT-4). For example, it sends a prompt such as, "The user is asking how to return a product. Please tell me how to return it." The input is text data, and the output is the generation and transmission of a prompt.

[1464] Step 7:

[1465] Generative AI models generate answers

[1466] The generative AI model generates an appropriate answer based on the prompt sentence. For example, it generates an answer such as, "To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is the prompt sentence, and the output is the generated answer text.

[1467] Step 8:

[1468] The server converts the answer into audio

[1469] The server converts the generated answer into speech using the Google Cloud Text-to-Speech API, where the input is the generated answer text and the output is the audio data.

[1470] Step 9:

[1471] The server provides the user with a spoken response

[1472] The server provides the user with a response converted into voice via the Twilio API. For example, it may respond with a voice message such as, "To find out how to return the product, please bring it to the store with your receipt within 30 days of purchase." The input is voice data, and the output is a voice response to the user.

[1473] Step 10:

[1474] The server uses an emotion recognition engine to analyze the user's emotions.

[1475] The server sends the user's voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to analyze the user's emotions. The input is the user's voice data, and the output is the emotion analysis result.

[1476] Step 11:

[1477] The server instructs the generative AI model to respond according to the emotion.

[1478] The server sends instructions to the generative AI model based on the results of the emotion recognition engine. For example, it sends a prompt such as, "The user is feeling angry. Please respond in a polite manner." The input is the emotion analysis results, and the output is the generation and transmission of a prompt.

[1479] Step 12:

[1480] Generative AI models generate emotionally relevant answers

[1481] The generative AI model generates an appropriate response based on the sentiment. For example, it generates an answer like, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is the prompt sentence, and the output is the generated answer text.

[1482] Step 13:

[1483] The server converts the response into speech based on the emotion.

[1484] The server converts the sentiment-based response into speech using the Google Cloud Text-to-Speech API. The input is the generated response text, and the output is the speech data.

[1485] Step 14:

[1486] The server provides the user with a response in voice that corresponds to their emotions.

[1487] The server provides the user with a voice response based on the emotion through the Twilio API. For example, it might say, "We're sorry. To return the product, please bring it to the store with your receipt within 30 days of purchase." The input is voice data, and the output is a voice response to the user.

[1488] (Application example 1)

[1489] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1490] Conventional automated voice response systems have had problems in that they are unable to provide appropriate answers to user questions and respond in a way that reflects the user's emotions. In particular, solving these problems is important for customer service in brick-and-mortar stores, where quick and appropriate responses are required.

[1491] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a means for linking an automatic voice response (IVR) with a voice-based generation AI, a means provided for a fixed fee, a number providing means, a voice guidance means, a generation AI linkage function means, a means for converting a user's voice into text, a means for generating answers to user questions using a generation AI model, a means for converting the generated answers into voice, and a means for recognizing the user's emotions and responding in an appropriate tone. This makes it possible to provide quick and appropriate answers to user questions and respond in accordance with the user's emotions.

[1492] An "Interactive Voice Response (IVR)" is a system that automatically responds via voice over the telephone, providing appropriate information based on user input.

[1493] "Voice generative AI" is an artificial intelligence technology that analyzes voice data and generates appropriate answers to users' questions and requests.

[1494] A "fixed fee means" is a system in which services are provided at a fixed fee.

[1495] "Number provisioning means" is a function that provides a telephone number for users to access the system.

[1496] The "audio guidance means" is a function that provides guidance and instructions to the user by voice.

[1497] "Generative AI collaboration function means" refers to technology that allows generative AI to collaborate with other systems and functions.

[1498] "Means for converting user speech into text" refers to technology that converts the user's speech into text information.

[1499] "Means for generating answers to user questions using a generative AI model" refers to technology that uses a generative AI model to create appropriate answers to user questions.

[1500] The "means for converting the generated answer into speech" is a technique for converting the generated text-format answer into speech format.

[1501] "Means to recognize the user's emotions and respond in an appropriate tone" refers to technology that analyzes the user's emotions from their tone of voice and choice of words, and responds in an appropriate tone accordingly.

[1502] "Means for improving customer service in physical stores" refers to technologies and methods for streamlining customer service in physical stores and improving customer satisfaction.

[1503] The "prepared answer database" is a database of questions and their answers that have been prepared in advance.

[1504] "Short message notification" is a function that sends short text messages to users.

[1505] "Means for asking questions to confirm and saving the customer's answers as text" refers to a technology that asks the user questions to confirm and saves the answers in text format.

[1506] As an embodiment of the present invention, a smartphone application for improving customer service in a brick-and-mortar store will be described as an example.

[1507] System Program

[1508] This system operates using the following hardware and software:

[1509] Hardware:

[1510] Smartphone (including microphone and speaker)

[1511] software:

[1512] Python

[1513] speech_recognition library: used to convert speech to text

[1514] gTTS library: used to convert text to speech

[1515] playsound library: used to play sounds

[1516] OpenAI API: Uses generative AI models to generate answers to user questions

[1517] Processing flow

[1518] The server first records the user's voice using the smartphone's microphone, then converts the recorded voice into text using the speech_recognition library, which is then sent to the OpenAI API to generate an appropriate answer using a generative AI model, which is then converted into speech using the gTTS library and played over the smartphone's speaker.

[1519] In addition, to recognize the user's emotions, the system analyzes voice data and determines the user's emotions from the tone of their voice and the way they speak. Based on this information, the system responds with an appropriate tone.

[1520] Specific examples

[1521] For example, if a user opens a smartphone app and asks, "What are your business hours?", the app converts the question into text and sends it to the OpenAI API. The generative AI model generates the answer, "Business hours are from 9:00 AM to 5:00 PM," and converts this answer into audio and plays it back to the user.

[1522] Example prompt sentence:

[1523] User Question: What are your opening hours?

[1524] answer:

[1525] In this way, it is possible to provide a quick and appropriate answer to the user's question and also to respond in accordance with the user's feelings.

[1526] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1527] Step 1:

[1528] A user launches a smartphone app and speaks a question into the microphone. The input is the user's voice data, and the output is a recorded audio file. Specifically, the smartphone's microphone captures the user's voice and saves it as an audio file.

[1529] Step 2:

[1530] The device converts the recorded audio file into text using the speech_recognition library. The input is an audio file and the output is text data. Specifically, the speech_recognition library analyzes the audio file and converts the audio into text information.

[1531] Step 3:

[1532] The device sends the generated text data to the OpenAI API and uses the generative AI model to generate an appropriate answer. The input is the text data, and the output is the generated answer text. Specifically, the text data is sent to the OpenAI API as a prompt, and the generative AI model generates an answer.

[1533] Step 4:

[1534] The device converts the generated response text into speech using the gTTS library. The input is the response text, and the output is an audio file. Specifically, the gTTS library converts the text data into speech data and saves it as an audio file.

[1535] Step 5:

[1536] The device plays the generated audio file on the smartphone's speaker. The input is the audio file, and the output is the audio that the user can hear. Specifically, it uses the playsound library to play the audio file and outputs the audio from the speaker.

[1537] Step 6:

[1538] The server analyzes the user's voice data and uses an emotion engine to recognize the user's emotions. The input is the user's voice data and the output is emotional information. Specifically, the server analyzes the voice data and determines the user's emotions from the tone of voice and the use of words.

[1539] Step 7:

[1540] The server generates a response in an appropriate tone based on the recognized emotional information and converts it into speech. The input is the emotional information and the response text, and the output is a speech file with a tone corresponding to the emotion. Specifically, the response text is adjusted taking into account the emotional information and converted into speech using the gTTS library.

[1541] Step 8:

[1542] The device plays an audio file with a tone corresponding to the emotion and provides it to the user. The input is an audio file with a tone corresponding to the emotion, and the output is a sound that the user can hear. Specifically, it uses the playsound library to play the audio file and outputs the sound from the speaker.

[1543] Example 2

[1544] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1545] Conventional automated voice response systems have difficulty responding flexibly to user requests, making it difficult to increase user satisfaction, especially when it comes to restaurant reservations and optimizing services. Furthermore, there were no systems that recognized user emotions and responded accordingly, making it impossible to improve the quality of service.

[1546] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice input, means for converting voice data into text, means for generating appropriate questions and answers using a generative AI model, means for converting the generated questions and answers into voice and conveying them to the user, means for recognizing the user's emotions and generating a response according to the emotions, and means for notifying staff of the generated response. This enables flexible response to user requests and optimization of restaurant reservations and services. Furthermore, by responding according to the user's emotions, the quality of service can be improved and user satisfaction can be increased.

[1547] "Means for receiving user voice input" refers to devices or techniques for capturing user-uttered voice and inputting it into the system.

[1548] A "voice-to-text converter" is a process or device that uses voice recognition technology to convert captured voice data into written information.

[1549] "Means for generating appropriate questions and answers using generative AI models" refers to systems or algorithms that utilize artificial intelligence technology to automatically generate questions and answers in response to user requests.

[1550] The "means for converting the generated questions and answers into voice and communicating them to the user" refers to a technology or device for converting text information into voice and providing the information to the user by voice.

[1551] "Means for recognizing a user's emotions and generating a response that corresponds to those emotions" refers to a system or technology that analyzes the user's emotions from their voice and facial expressions, and automatically generates an appropriate response based on the results.

[1552] The "means for notifying staff of the generated response" refers to a device or technology for notifying restaurant staff of the generated response information.

[1553] MODE FOR CARRYING OUT THE INVENTION

[1554] This invention is a system that uses a generative AI model to solve services such as restaurant reservations, restaurant directions, checking business hours, and phone orders. It also includes a function to recognize user emotions and optimize services accordingly.

[1555] Hardware and software used

[1556] server

[1557] The server receives voice input from the user and uses speech recognition technology to convert the voice data into text. Specifically, it uses the Google Cloud Speech-to-Text API. Next, it uses a generative AI model (e.g., OpenAI's GPT-4) to generate appropriate questions and answers based on the user's request. It uses speech synthesis technology (e.g., Amazon Polly) to convert the generated text into speech. It also uses emotion recognition technology (e.g., Microsoft Azure's Emotion API) to recognize the user's emotions.

[1558] Terminal

[1559] The device captures the user's voice input and sends it to the server. It also uses speech synthesis technology to return responses from the server to the user as voice. It also has a camera and microphone to capture the user's facial expressions and send data for emotion recognition to the server.

[1560] User

[1561] The user inputs a request by voice into the terminal. For example, if the user says, "I'd like to make a reservation," the system generates an appropriate question and returns it to the user by voice. When the user inputs the answer, the system makes the reservation based on that information.

[1562] Specific examples

[1563] When a user says, "I'd like to make a reservation," the server converts the voice data into text using the Google Cloud Speech-to-Text API. Next, it uses a generative AI model (OpenAI's GPT-4) to generate the question, "How many people would you like to make a reservation for, and what time would you like to start?" The server converts this question into speech using Amazon Polly and relays it to the user via the device. When the user replies, "We'd like to make a reservation for three people, starting at 7 p.m.", the device sends the voice data to the server, which again converts it into text using the Google Cloud Speech-to-Text API. The server processes the reservation information, generates a confirmation message stating, "Your reservation for three people starting at 7 p.m.", and relays it to the user via Amazon Polly.

[1564] Prompt Sentence Examples

[1565] "A user wants to make a reservation. How many people would you like to book and what time would you like to book?"

[1566] If the user is dissatisfied with the restaurant, the server will analyze their emotions using Microsoft Azure's Emotion API and generate a response such as, "The user is dissatisfied. Please respond more politely." The device will then notify the staff of this information and prompt them to take appropriate action.

[1567] Prompt Sentence Examples

[1568] "Users are angry and frustrated. Please encourage your staff to be more polite."

[1569] In this way, restaurants can flexibly respond to user requests and optimize restaurant reservations and services. Furthermore, by responding to user emotions, the quality of service can be improved and user satisfaction can be increased.

[1570] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1571] Program processing flow

[1572] A system for making reservations, providing directions to the store, checking business hours, and answering phone orders

[1573] Step 1:

[1574] The user inputs voice data. The user speaks to the terminal, saying, "I'd like to make a reservation." The input is the user's voice data.

[1575] Step 2:

[1576] The device sends voice data to the server. The device captures the user's voice and sends the data to the server. The input is the user's voice data, and the output is the voice data sent to the server.

[1577] Step 3:

[1578] The server converts the voice data into text. The server converts the voice data into text using the Google Cloud Speech-to-Text API. The input is voice data and the output is text data.

[1579] Step 4:

[1580] The server uses a generative AI model to generate an appropriate question. The server uses OpenAI's GPT-4 to generate the question, "How many people would you like to book, and what time would you like to start?" The input is text data, and the output is the generated question text.

[1581] Step 5:

[1582] The server sends the generated question to the terminal. The server sends the generated question to the terminal. The input is the text of the generated question, and the output is the text of the question sent to the terminal.

[1583] Step 6:

[1584] The device speaks the question to the user. The device uses Amazon Polly to convert the question into speech and speaks it to the user. The input is the text data of the question, and the output is speech data.

[1585] Step 7:

[1586] The user inputs their answer by voice. The user answers, "I would like to make a reservation for three people starting at 7 PM." The input is the user's voice data.

[1587] Step 8:

[1588] The device sends the voice data of the answer to the server. The device captures the user's answer and sends the data to the server. The input is the user's voice data, and the output is the voice data sent to the server.

[1589] Step 9:

[1590] The server converts the voice data into text. The server again converts the voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data.

[1591] Step 10:

[1592] The server processes the reservation information and generates a confirmation message. The server processes the reservation information and generates a confirmation message saying, "We have received a reservation for 3 people starting at 7 PM." The input is text data, and the output is the text of the generated confirmation message.

[1593] Step 11:

[1594] The server sends a confirmation message to the terminal. The server sends a confirmation message to the terminal. The input is the generated confirmation message text, and the output is the confirmation message text sent to the terminal.

[1595] Step 12:

[1596] The device will speak a confirmation message to the user. The device will convert the confirmation message into speech using Amazon Polly and speak it to the user. The input is the text data of the confirmation message, and the output is the speech data.

[1597] A system that optimizes services based on emotions

[1598] Step 1:

[1599] A user receives a service at a restaurant. A user enjoys a meal at a restaurant. The input is the user's behavior.

[1600] Step 2:

[1601] The device captures the user's voice and facial expressions. The device captures the user's voice and facial expressions using a camera and microphone. The input is the user's voice data and facial expression data.

[1602] Step 3:

[1603] The device sends the captured data to the server. The device sends the captured data to the server. The input is voice data and facial expression data, and the output is the data sent to the server.

[1604] Step 4:

[1605] The server analyzes the user's emotions using emotion recognition technology. The server analyzes the user's emotions using Microsoft Azure's Emotion API. The input is voice data and facial expression data, and the output is the emotion analysis results.

[1606] Step 5:

[1607] The server generates an appropriate response based on the analysis results. The server uses OpenAI's GPT-4 to generate a response such as "The user is dissatisfied. Please respond more politely." The input is the sentiment analysis result, and the output is the generated response text.

[1608] Step 6:

[1609] The server sends the generated response to the terminal. The server sends the generated response to the terminal. The input is the text of the generated response, and the output is the text of the response sent to the terminal.

[1610] Step 7:

[1611] The terminal notifies the staff of the response. The terminal notifies the staff of the response using the display and voice output function. The input is the text data of the response, and the output is the information notified to the staff.

[1612] (Application example 2)

[1613] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1614] Conventional restaurant services such as making reservations, getting directions, checking business hours, and calling to order are cumbersome for users, and they have the problem of being difficult to respond to efficiently. In addition, services are not optimized according to user emotions, making it difficult to improve customer satisfaction. To solve these problems, a system utilizing voice recognition technology, generative AI, and emotion recognition technology is needed.

[1615] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1616] In this invention, the server includes an automatic voice response means, a voice-based AI means, an emotion recognition means, a location information providing means, a voice guidance means, and an AI generative function means. This allows users to easily make reservations, get directions, check business hours, place orders, and more by voice. Furthermore, by recognizing the user's emotions and providing optimal services in response to them, customer satisfaction can be improved.

[1617] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[1618] A "voice generation AI means" is a system that uses artificial intelligence technology to analyze voice input and generate an appropriate voice response.

[1619] An "emotion recognition means" is a system that has the technology to analyze emotions from the user's voice and facial expressions and recognize those emotions.

[1620] "Location information providing means" is a system that has the function of providing location information of the user's current location and destination.

[1621] "Audio guidance means" refers to a system that has the function of providing guidance and instructions to the user by voice.

[1622] A "generative AI collaboration function means" is a system that has the functionality to work in collaboration with generative AI and other systems and databases.

[1623] The "restaurant reservation means" is a system that allows users to make reservations for restaurants by voice.

[1624] The "restaurant navigation means" is a system that has the function of providing users with directions to restaurants.

[1625] The "means for checking business hours" is a system that has a function that allows users to check the business hours of restaurants by voice.

[1626] The "telephone ordering means" is a system that allows users to place orders with restaurants by telephone.

[1627] "Means for optimizing services according to emotions" is a system that has the function of recognizing the user's emotions and providing the optimal service according to those emotions.

[1628] The "prepared response database means" is a system having a function of using a database that stores prepared responses.

[1629] "Voice response means" refers to a system that has the function of providing answers to users by voice.

[1630] The "short message notification means" is a system that has the function of notifying the user by short message.

[1631] The "means for asking questions to confirm matters" is a system that has the function of asking the user questions to confirm matters.

[1632] "Means for saving customer responses in text" refers to a system that has the function of saving user responses in text format.

[1633] The system for implementing this invention uses the following hardware and software: a smartphone, microphone, and GPS are required for the hardware, and OpenAI API, SpeechRecognition, Geopy, and EmotionRecognizer are used for the software.

[1634] The server receives user voice input via a microphone and converts it to text using the SpeechRecognition library. It then uses the OpenAI API to generate an appropriate response based on the user's input. For example, if a user says, "I'd like to make a reservation," the generative AI will ask, "How many people would you like to make a reservation for, and what time would you like to make a reservation?"

[1635] The server also recognizes the user's emotions using EmotionRecognizer and notifies the staff of the appropriate response based on that emotion. For example, if the user is angry, the system will notify the staff, "The user is angry. Please respond politely."

[1636] Furthermore, the server uses location information provision means to provide directions from the user's current location to the restaurant. It uses the Geopy library to obtain the user's current location and calculate the optimal route to the destination.

[1637] As a specific example, if a user says, "Tell me where the store is," the generative AI will respond, "I will provide directions from your current location to the store," and begin providing directions using location information provision means.

[1638] An example of a prompt is as follows:

[1639] "A user wants to make a reservation. How many people would you like to book and what time would you like to book?"

[1640] "User wants to know where your store is. Please provide directions from their current location to the store."

[1641] "Users want to know your hours. What are your hours?"

[1642] In this way, users can easily make reservations, get directions, check business hours, place orders, etc. by voice. Furthermore, the system can recognize the user's emotions and provide optimal service accordingly, which can improve customer satisfaction.

[1643] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1644] Step 1:

[1645] The user inputs voice using the microphone on their smartphone. For example, the user's voice input might be, "I'd like to make a reservation to visit the restaurant." The input is voice data, and the output is the voice data itself.

[1646] Step 2:

[1647] The device converts voice input to text using the SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device analyzes the voice data and generates corresponding text.

[1648] Step 3:

[1649] The server receives the text data and generates an appropriate response using the OpenAI API. The input is text data, and the output is the generated response text. For example, in response to the input "I would like to make a reservation," the server generates the response "How many people would you like to make a reservation for, and what time would you like to make a reservation?"

[1650] Step 4:

[1651] The server converts the generated response text into speech and provides audio guidance to the user. The input is the generated response text, and the output is audio data. Specifically, the text is converted into audio data using speech synthesis technology.

[1652] Step 5:

[1653] The user again speaks to provide reservation details, for example, "I'd like to make a reservation for two people starting at 7 PM." The input is the voice data, and the output is the voice data itself.

[1654] Step 6:

[1655] The device again converts the voice input into text using the SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device analyzes the voice data and generates the corresponding text.

[1656] Step 7:

[1657] The server receives the text data and stores the reservation information in a database. The input is the text data, and the output is the result of saving it to the database. Specifically, the server analyzes the text data and stores it in the database as reservation information.

[1658] Step 8:

[1659] The server recognizes the user's emotions using EmotionRecognizer and notifies the staff of the appropriate response based on that emotion. The input is the user's voice data, and the output is the emotion recognition results and the notification content. Specifically, the server analyzes the voice data, recognizes the emotion, and generates the corresponding notification.

[1660] Step 9:

[1661] The server uses location information providing means to provide directions from the user's current location to the restaurant. The input is the user's current location data, and the output is route guidance information. Specifically, the server uses the Geopy library to obtain the user's current location and calculates the optimal route to the destination.

[1662] Step 10:

[1663] The server converts the route guidance information into voice and provides the user with voice guidance. The input is the route guidance information and the output is voice data. Specifically, the text is converted into voice data using voice synthesis technology.

[1664] Example 3

[1665] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1666] Conventional automated voice response systems have difficulty generating appropriate voice responses to user requests and are unable to respond according to the user's emotions. Furthermore, they lack the functionality to save user responses as text or send SMS notifications, limiting the user experience. To solve these issues, a system is needed that integrates the generation of voice responses according to user requests, responses according to emotions, text saving, and SMS notifications.

[1667] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1668] In this invention, the server includes means for accepting user requests by voice, means for generating questions using a voice generation system AI, means for accepting user responses by voice and converting the voice data into text and saving it, means for recognizing the user's emotions and generating voice responses in response to the emotions, and means for sending SMS notifications. This makes it possible to generate appropriate voice responses in response to user requests, respond according to emotions, save the user's responses as text, and send SMS notifications.

[1669] "User" means any person or entity making a request using the System.

[1670] "Voice-based generative AI" is an artificial intelligence technology that converts text data into voice data.

[1671] "Audio data" is data that represents audio in digital form.

[1672] "Text data" is data that represents character information in digital form.

[1673] "Emotion recognition" is a technology that analyzes a user's emotions from voice data and text data.

[1674] "Voice response" refers to a voice response to a user's request or question.

[1675] "SMS Notification" means sending information to a user using short message service.

[1676] A "database" is a system for efficiently storing, managing, and retrieving data.

[1677] A "server" is a computer system that processes data and provides services over a network.

[1678] A "terminal" is a device that is directly operated by a user and that inputs and outputs audio.

[1679] This invention is a system that integrates the generation of voice responses according to user requests, responses according to emotions, text storage of user responses, and SMS notifications. This system operates in cooperation with the server, the terminal, and the user.

[1680] The server includes a means for accepting user requests by voice, a means for generating questions using a voice generation AI system, a means for accepting user responses by voice, converting the voice data into text and saving it, a means for recognizing the user's emotions and generating corresponding voice responses, and a means for sending SMS notifications.

[1681] Specifically, the server operates as follows: First, the device accepts a request from the user via voice and sends that voice data to the server. The server then uses a voice-based AI (e.g., Google Cloud Text-to-Speech API) to generate a question for the user as voice data and sends it to the device. The device then plays back this voice data and conveys the question to the user.

[1682] Next, the user speaks a response into the device, which then sends the voice data to the server. The server uses voice recognition software (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text and store it in a database. The server then uses emotion recognition software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions and generate a voice response with a tone appropriate to the emotion. The generated voice data is then sent to the device and played back to the user.

[1683] The server also uses an SMS notification service (e.g., Twilio API) to notify the user of confirmations and order details, allowing the user to receive confirmation of their request via SMS.

[1684] As a concrete example, if a user requests to "place an order," the system operates as follows: First, the device accepts the user's request and sends it to the server. The server uses a voice generation AI to generate the question "What would you like to order?" and sends it to the device. The device plays the question and conveys it to the user. When the user replies "I'd like to order pizza," the device sends the voice data to the server. The server converts the voice data into text and stores it in a database. The server then uses emotion recognition software to analyze the user's emotions and generates a voice response in an appropriate tone. Finally, the server uses an SMS notification service to send a confirmation of the order to the user.

[1685] An example of a prompt sentence is "How does the system behave when a user requests to place an order?" By inputting this prompt sentence into the generative AI model, a sentence explaining the specific operating procedure of the system can be generated. The flow of the specific processing in Example 3 will be explained using FIG. 21.

[1686] Step 1:

[1687] The user speaks a request into the terminal.

[1688] Input: User's spoken request (e.g., "I would like to place an order")

[1689] How it works: The device receives the user's voice using a microphone and generates voice data.

[1690] Output: Generated audio data

[1691] Step 2:

[1692] The terminal transmits the generated voice data to the server.

[1693] Input: Device-generated audio data

[1694] Operation: The device sends audio data to the server over the network.

[1695] Output: Audio data sent to the server

[1696] Step 3:

[1697] The server generates questions using a voice-based AI.

[1698] Input: Audio data sent to the server

[1699] How it works: The server analyzes the voice data and recognizes the user's request. Using a voice-generating AI (e.g., Google Cloud Text-to-Speech API), it generates the question "What would you like to order?" as voice data.

[1700] Output: Audio data of the generated question

[1701] Step 4:

[1702] The server sends the generated voice data of the question to the terminal.

[1703] Input: Audio data of the generated question

[1704] Operation: The server sends the voice data of the question to the terminal via the network.

[1705] Output: Audio data of the question sent to the device

[1706] Step 5:

[1707] The terminal reproduces the audio data of the question and conveys the question to the user.

[1708] Input: Voice data of the question sent to the terminal

[1709] Operation: The device uses the speaker to play back the audio data of the question and convey it to the user.

[1710] Output: The question conveyed to the user

[1711] Step 6:

[1712] The user speaks their answer into the terminal.

[1713] Input: User's spoken response (e.g., "I'd like to order a pizza")

[1714] How it works: The device receives the user's voice using a microphone and generates voice data.

[1715] Output: Generated audio data

[1716] Step 7:

[1717] The terminal transmits the generated voice data to the server.

[1718] Input: Device-generated audio data

[1719] Operation: The device sends audio data to the server over the network.

[1720] Output: Audio data sent to the server

[1721] Step 8:

[1722] The server converts the voice data into text and stores it in a database.

[1723] Input: Audio data sent to the server

[1724] How it works: The server uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text, which is then stored in a database.

[1725] Output: Text data stored in a database

[1726] Step 9:

[1727] The server recognizes the user's emotions and generates a voice response accordingly.

[1728] Input: Audio data sent to the server

[1729] How it works: The server uses emotion recognition software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. Based on the analysis, it generates a voice response with an appropriate tone.

[1730] Output: Generated voice response data

[1731] Step 10:

[1732] The server transmits the generated voice response data to the terminal.

[1733] Input: Generated voice response data

[1734] Operation: The server sends voice response data to the terminal via the network.

[1735] Output: Voice response data sent to the device

[1736] Step 11:

[1737] The terminal reproduces the voice response data and conveys the response to the user.

[1738] Input: Voice response data sent to the terminal

[1739] Operation: The device uses a speaker to play back the voice response data and convey it to the user.

[1740] Output: The spoken response given to the user

[1741] Step 12:

[1742] The server sends an SMS notification.

[1743] Input: Text data stored in a database

[1744] What it does: The server uses an SMS notification service (e.g., Twilio API) to notify the user of confirmations and order details.

[1745] Output: SMS notification sent to the user

[1746] (Application example 3)

[1747] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1748] Conventional automated voice response systems were unable to respond to users' emotions, resulting in a poor user experience. They also had insufficient responses to voice orders and inquiries, making it difficult to efficiently process orders and manage orders. Furthermore, the accuracy of voice and emotion recognition was low, making it difficult to accurately understand the user's intentions.

[1749] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server is a system that links an automatic voice response (IVR) and a voice-based generation AI, and is provided for a fixed fee. It includes: a means for providing a number, voice guidance, and a generation AI linking function in an integrated manner; a means for linking a pre-prepared answer database with the voice-based generation AI, providing voice answers, SMS notifications, and asking confirmation questions, and storing the user's answers as text; a means for recognizing the user's emotions and providing voice answers and SMS notifications accordingly; and a means for allowing the user to place an order by voice, saving the order details as text, and managing the order history. This enables optimal responses according to the user's emotions and enables efficient voice order processing and history management.

[1750] An "Interactive Voice Response (IVR)" is a system that automatically responds to phone calls and voice input from users.

[1751] "Voice generative AI" is an artificial intelligence technology that generates voice data and responds to users via voice.

[1752] The "answer database" is a database that stores prepared questions and answers.

[1753] "SMS Notification" is a function that sends text messages to users using short message service.

[1754] "Emotion recognition" is a technology that analyzes emotions from a user's voice or text and identifies those emotions.

[1755] "Voice guidance" is a function that provides guidance and instructions to the user using voice.

[1756] "Save order details as text" is a function that saves the order details entered by the user via voice as text data.

[1757] "History management" is a function that stores data on past orders and inquiries so that it can be referenced as needed.

[1758] The system for implementing this invention has the following configuration. The server is a system that links an IVR (Interactive Voice Response) and a voice-based AI generator. It is provided for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and AI generator linkage functions. It also links a pre-prepared response database with the voice-based AI generator, allowing for voice responses, SMS notifications, and confirmation questions, and the text of user responses can be saved. Furthermore, it has the function of recognizing the user's emotions and providing voice responses and SMS notifications accordingly. Users can place orders by voice, and the order details can be saved as text and managed as a history.

[1759] Hardware and software used

[1760] Hardware: Smartphone (microphone, speaker)

[1761] Software: Python, speech_recognition library (voice recognition), textblob library (sentiment analysis), gTTS library (voice generation), smtplib library (SMS sending), sqlite3 library (database storage)

[1762] Processing flow

[1763] The server first uses the speech_recognition library to convert the user's speech into text for speech recognition. Next, it uses the textblob library to analyze the sentiment of the text and recognize the user's emotion. It then uses the gTTS library to generate a voice response based on the user's sentiment and plays it back from the speaker. It then uses the smtplib library to send the user's order details via SMS. Finally, it uses the sqlite3 library to save the order details in a database and manage them as a history.

[1764] Specific examples

[1765] If a user says, "I want to order a pizza," the server will ask "What would you like to order?" and if the user answers, "One Margherita pizza," the server will save the answer as text and respond with a voice response according to the user's sentiment. The server will also send the order details via SMS and store them in a database.

[1766] Prompt Sentence Examples

[1767] User: "I want to order a pizza"

[1768] Server: "What would you like to order?"

[1769] User: "One Margherita pizza."

[1770] Server: "Thank you for your order!"

[1771] In this way, a system can be implemented that efficiently processes food delivery orders based on user voice input.

[1772] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1773] Step 1:

[1774] To perform speech recognition, the server uses the speech_recognition library to convert the user's speech into text. Specifically, when the user speaks into the microphone of their smartphone, "I would like to order a pizza," the speech data is sent to the server. The server receives this speech data as input and converts it into text data using the speech_recognition library. The text "I would like to order a pizza" is obtained as output.

[1775] Step 2:

[1776] The server uses the textblob library to analyze the text data and recognize the user's emotions. Specifically, it receives the text data "I want to order pizza" obtained in step 1 as input and performs sentiment analysis using the textblob library. As a result of the sentiment analysis, it recognizes the user's sentiment as "neutral." The output is the sentiment data "neutral."

[1777] Step 3:

[1778] The server uses the gTTS library to generate a voice response that corresponds to the user's emotion. Specifically, it receives the emotion data "Neutral" obtained in step 2 and the user's order "I would like to order pizza" as input, and generates voice data using the gTTS library. The generated voice data is "What would you like to order?" and is played back from the smartphone speaker. The voice data is obtained as output.

[1779] Step 4:

[1780] The user speaks into the microphone of their smartphone, "One Margherita pizza please." The server converts this voice data into text data again using the speech_recognition library. Specifically, the server receives the user's voice data as input and converts it into text data using the speech_recognition library. The output is the text "One Margherita pizza please."

[1781] Step 5:

[1782] The server uses the smtplib library to send the user's order details via SMS. Specifically, it receives the text data "One Margherita Pizza" obtained in step 4 as input, uses the smtplib library to generate an SMS message, and sends it to the user's registered phone number. The SMS message is sent as output.

[1783] Step 6:

[1784] The server uses the sqlite3 library to save the user's order details to the database. Specifically, it receives the text data "One Margherita Pizza" obtained in step 4 as input and saves it to the database using the sqlite3 library. As output, the order details are saved to the database and managed as a history.

[1785] In this way, a system can be implemented that efficiently processes food delivery orders based on user voice input.

[1786] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1787] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1788] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are mentioned.

[1789] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1790] [Third embodiment]

[1791] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1792] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1793] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1794] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1795] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1796] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1797] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1798] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1799] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1800] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1801] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1802] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[1803] "Example 1"

[1804] One embodiment of the present invention is a system that integrates an interactive voice response (IVR) with a voice-based AI generator. This system is offered for a fixed fee and provides a comprehensive range of services, including number provision, voice guidance, and AI generator integration functions. Specifically, when a user calls the system, the system responds to the user's questions and requests using the voice-based AI generator. For example, if a user asks, "What are your business hours?", the system responds using the voice-based AI generator, "Business hours are from 9:00 AM to 5:00 PM."

[1805] "Example 2"

[1806] Another embodiment of the present invention is a system that can use generative AI to solve restaurant reservations, directions to the restaurant, confirmation of business hours, and phone orders. Specifically, when a user requests to "make a restaurant reservation," the system uses a voice-based generative AI to ask, "How many people would you like to make a reservation for, and what time would you like to make a reservation for?" and makes the reservation based on the user's response.

[1807] "Example 3"

[1808] In yet another embodiment of the present invention, a system is provided that links a prepared answer database with a voice-based generation AI to provide voice responses, SMS notifications, and ask questions for confirmation, and then saves the customer's responses as text. Specifically, when a user requests to "place an order," the system uses the voice-based generation AI to ask "What would you like to order?", accepts the order based on the user's response, and saves the contents as text.

[1809] The processing flow of each embodiment will be described below.

[1810] "Example 1"

[1811] Step 1: A user calls the system.

[1812] Step 2: The system responds to the user's questions and requests using a voice-generative AI.

[1813] Step 3: The user asks, "What are your business hours?"

[1814] Step 4: The system uses a voice-generating AI to respond, "Business hours are from 9:00 AM to 5:00 PM."

[1815] "Example 2"

[1816] Step 1: The user requests to make a reservation.

[1817] Step 2: The system uses a voice-generating AI to ask, "How many people would you like to make a reservation for, and what time would you like to start?"

[1818] Step 3: The system makes the reservation based on the user's answers.

[1819] "Example 3"

[1820] Step 1: The user requests to place an order.

[1821] Step 2: The system uses a voice-generative AI to ask, "What would you like to order?" Step 3: The system accepts the order based on the user's answer.

[1822] Step 4: The system saves the order details as text.

[1823] Example 1

[1824] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1825] Conventional automated voice response systems have had difficulty responding appropriately and quickly to a variety of user questions and requests. Furthermore, they lacked the ability to convert voice data into text and integrate it with generative AI models, resulting in a poor user experience. Furthermore, certain industries, such as restaurants, lacked systems that could address specific needs, such as making reservations or checking opening hours.

[1826] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1827] In this invention, the server includes means for receiving a call from a user, means for providing voice guidance, means for converting the user's voice data into text, means for transmitting the converted text to the generative AI model as a prompt sentence, means for the generative AI model to generate an answer, means for converting the generated text into speech, and means for transmitting the converted speech data to the user, thereby enabling appropriate and prompt responses to a variety of user questions and requests.

[1828] The "means for receiving a call from a user" is a function that allows the system to receive a call made by a user and start a call session.

[1829] The "means for providing voice guidance" is a function for reproducing a voice message to the user and prompting the user to perform the next operation or input.

[1830] The "means for converting user's voice data into text" is a function for converting the voice uttered by the user into character data.

[1831] "Means for sending the converted text to the generative AI model as a prompt sentence" refers to a function that inputs text data converted from speech into the generative AI model and gives instructions for generating an appropriate response.

[1832] The "means by which the generative AI model generates an answer" is the function by which the generative AI model generates an appropriate answer based on the prompt sentence.

[1833] "Means for converting generated text into speech" refers to a function for converting text data generated by a generative AI model into speech data.

[1834] The "means for transmitting converted voice data to the user" is a function for transmitting data converted into voice to the user so that the user can listen to the voice.

[1835] This invention is a system that combines an interactive voice response (IVR) and a generative AI model, and aims to respond appropriately and quickly to calls from users. A specific embodiment of this system is described below.

[1836] System configuration

[1837] The server system is configured using the following hardware and software.

[1838] Hardware: Servers, communication equipment, audio input / output devices

[1839] Software: Twilio API, Google Cloud Speech-to-Text API, Google Cloud Text-to-Speech API, generative AI models (e.g., OpenAI's GPT-3)

[1840] Program processing

[1841] The server receives a call from the user using the Twilio API. When the call connects, the server provides a voice prompt, such as "How can I help you?"

[1842] When a user speaks a question or request, the server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then sent to the generative AI model as a prompt.

[1843] The generative AI model generates appropriate answers to user questions. For example, if a user asks, "What are your business hours?", the generative AI model will respond with, "Business hours are from 9:00 AM to 5:00 PM."

[1844] The generated text response is then converted back into audio. For this, the server uses the Google Cloud Text-to-Speech API. The converted audio data is then sent to the user via the Twilio API.

[1845] Specific examples

[1846] For example, the following shows the processing when a user asks, "What are your business hours?"

[1847] 1. A user calls the system.

[1848] 2. The server receives the call using the Twilio API and provides a voice prompt saying, "Please tell us your business."

[1849] 3. A user asks, "What are your business hours?"

[1850] 4. The server converts the speech to text using the Google Cloud Speech-to-Text API.

[1851] 5. The converted text, "What are your business hours?", is input as a prompt to the generative AI model.

[1852] 6. The generative AI model generates the answer, "Business hours are 9:00 AM to 5:00 PM."

[1853] 7. The server converts the generated text response into audio using the Google Cloud Text-to-Speech API.

[1854] 8. The server sends the converted voice data to the user via the Twilio API.

[1855] Prompt Sentence Examples

[1856] "A user is asking about business hours. Generate an appropriate answer based on the following text: 'What are your business hours?'"

[1857] In this way, the system operates in cooperation with the server, terminals, and users.

[1858] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1859] Step 1:

[1860] User makes a call

[1861] The user calls the provided phone number. The input is the user making the call, and the output is the call connecting to the server.

[1862] Step 2:

[1863] The server receives the call

[1864] The server receives a call from the user using a communication API. The input is a call connection from the user, and the output is the start of a call session. The server then prepares to manage the call session.

[1865] Step 3:

[1866] The server provides audio guidance

[1867] The server provides voice guidance to the user through a communication API. The input is the start of a call session, and the output is the playback of a voice guidance message, such as "Please tell us your business."

[1868] Step 4:

[1869] Users speak their questions or requests

[1870] The user follows the voice guidance to input questions or requests by voice. The input is the user's voice, and the output is the transmission of voice data to the server. For example, a question might be, "What are your business hours?"

[1871] Step 5:

[1872] The server converts the audio data into text

[1873] The server uses a speech recognition API to convert the user's voice data into text. The input is the user's voice data, and the output is the converted text data. For example, the generated text is "What are your business hours?"

[1874] Step 6:

[1875] The server sends a prompt to the generative AI model

[1876] The server sends the converted text to the generative AI model as a prompt. The input is the converted text data, and the output is the prompt sent to the generative AI model. The prompt is, "The user is asking about business hours. Please generate an appropriate answer based on the following text: 'What are your business hours?'"

[1877] Step 7:

[1878] Generative AI models generate answers

[1879] A generative AI model generates an appropriate answer based on a prompt. The input is the prompt, and the output is a generated text answer. For example, the generated answer might be, "Business hours are from 9:00 AM to 5:00 PM."

[1880] Step 8:

[1881] The server converts the generated text to speech

[1882] The server converts the generated text response into speech using a speech synthesis API. The input is the generated text response, and the output is the converted speech data. For example, the speech generated is "Business hours are from 9:00 AM to 5:00 PM."

[1883] Step 9:

[1884] The server sends the audio data to the user

[1885] The server sends the converted voice data to the user through a communication API. The input is the converted voice data, and the output is a voice transmission to the user. The user can hear the answer over the phone.

[1886] (Application example 1)

[1887] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1888] Conventional automated voice response systems could only provide pre-set, fixed responses, making it difficult to flexibly respond to a variety of user questions and requests. Furthermore, they were unable to respond immediately to security-related questions and requests, leaving users uneasy. This resulted in a poor user experience and limited the system's value.

[1889] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1890] In this invention, the server is a system that links an interactive voice response (IVR) with a voice-based generative AI, and is provided for a fixed fee. It includes a means for providing a number, voice guidance, generative AI linkage functions, etc. in an integrated manner, a means for speech recognition, a means for generating responses to user questions using a generative AI model, and a means for speech synthesis of the generated responses. This makes it possible to respond flexibly and immediately to a variety of user questions and requests, and in particular to quickly respond to questions and requests related to security, thereby alleviating user anxiety and improving the user experience.

[1891] An "Interactive Voice Response (IVR)" is a system that automatically responds with voice when a user calls and provides appropriate information based on the user's input.

[1892] "Voice generative AI" is an artificial intelligence technology that analyzes the user's voice input and generates an appropriate response.

[1893] "Providing a number" means providing a telephone number for a user to access.

[1894] "Voice guidance" is a function that provides voice guidance and instructions to the user.

[1895] The "generative AI integration function" is a function that links the voice-based generative AI with other systems and databases.

[1896] "Speech recognition" is a technology that converts a user's voice into text.

[1897] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on user input.

[1898] "Speech synthesis" is a technology that converts text into speech.

[1899] A "security question or request" is a question or request where a user requests information regarding the status or configuration of a security system.

[1900] The "answer database" is a database that stores prepared questions and their answers.

[1901] "Message notification" is a function that notifies users of information via text message.

[1902] "Questions for confirmation" is a function that asks the user questions about matters that need to be confirmed.

[1903] "Text save" is a function that saves the user's answers in text format.

[1904] As an embodiment of the present invention, a security assistant system will be described as an example. In this system, when a user inputs security-related questions or requests by voice, a voice-based AI system responds immediately.

[1905] Hardware and software used

[1906] Hardware: Smartphone (microphone, speaker)

[1907] software:

[1908] speech_recognition library: Used to perform speech recognition.

[1909] pyttsx3 library: Used to perform speech synthesis.

[1910] openai library: Uses generative AI models using the GPT-3 API.

[1911] System Operation

[1912] 1. Voice recognition: A user speaks a security question or request into the smartphone microphone, for example, "What is the status of my home security system?"

[1913] 2. Speech to text conversion: Uses the speech_recognition library to convert the user's speech into text.

[1914] 3. Response generation using a generative AI model: The converted text is sent to GPT-3 as a prompt to generate an appropriate response. An example of a prompt is "Security system question: What is the status of my home security system?"

[1915] 4. Speech synthesis: The generated response is converted into speech using the pyttsx3 library and transmitted back to the user through the smartphone speaker.

[1916] Specific examples

[1917] When a user asks, "What's the status of my home security system?" the system works as follows:

[1918] 1. The smartphone microphone captures the user's voice.

[1919] 2. The speech_recognition library converts the speech to text, generating the text "What is the status of my home security system?"

[1920] 3. The generated text is sent as a prompt to GPT-3, which generates an appropriate response, such as "All sensors are currently working properly."

[1921] 4. The pyttsx3 library converts the generated response into speech and delivers it back to the user through the smartphone speaker.

[1922] In this way, the security assistant system can respond immediately to users' security questions and requests, alleviating their concerns and improving their experience.

[1923] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1924] Step 1:

[1925] The user speaks their security question or request into the smartphone's microphone.

[1926] Input: User's voice

[1927] Output: Audio data captured by a microphone

[1928] Specific Action: A user says, "What's the status of my home security system?"

[1929] Step 2:

[1930] The device uses the speech_recognition library to convert the captured audio data into text.

[1931] Input: Audio data

[1932] Output: Text data

[1933] What it does: The speech recognition engine analyzes the voice data and generates text such as "What is the status of my home security system?"

[1934] Step 3:

[1935] The device sends the generated text as a prompt to GPT-3, which generates an appropriate response.

[1936] Input: Text data (prompt sentence)

[1937] Output: Response text

[1938] What it does: Sends the prompt "Security System Question: What's the status of my home security system?" to the GPT-3 API and receives the response "Currently, all sensors are working properly."

[1939] Step 4:

[1940] The device converts the generated response text into speech using the pyttsx3 library.

[1941] Input: Response text

[1942] Output: Audio data

[1943] What happens: The speech synthesis engine converts the text "All sensors are currently working properly" into speech data.

[1944] Step 5:

[1945] The device responds to the user with voice data via the smartphone speaker.

[1946] Input: Audio data

[1947] Output: The audio the user hears

[1948] What happens: The smartphone speaker will play a voice message saying "All sensors are currently working properly."

[1949] In this way, the security assistant system can respond immediately to the user's security questions and requests.

[1950] Example 2

[1951] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1952] Conventional automated voice response systems have difficulty responding flexibly and quickly to user requests, and efficient responses are required, especially for complex requests such as making restaurant reservations or confirming opening hours. Furthermore, there is a lack of technology to accurately understand users' voice requests and generate appropriate responses, making it difficult to improve the user experience.

[1953] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for accepting a user's voice request, means for converting voice to text, means for generating and transmitting a prompt sentence to the generative AI model, means for receiving a response from the generative AI model, means for converting the text response to voice data, means for transmitting the voice data to the user, means for accepting the user's answer, means for registering reservation information, means for generating a reservation confirmation message and converting it into voice data, and means for transmitting the voice data to the user. This makes it possible to respond quickly and accurately to user requests and efficiently handle complex requests such as making a reservation at a restaurant or confirming business hours.

[1954] The "means for accepting a user's voice request" refers to a device or software that recognizes the voice uttered by the user and inputs the content of that voice into the system.

[1955] A "speech-to-text converter" is a device or software that uses speech recognition technology to convert a user's speech into textual information.

[1956] A "means for generating and sending prompts to a generative AI model" is a device or software for generating appropriate questions or instructions based on a user's request and sending them to a generative AI model.

[1957] A "means for receiving a response from a generative AI model" is a device or software for receiving a response returned from a generative AI model.

[1958] A "means for converting text responses into audio data" is a device or software for converting text responses from a generative AI model into audio data.

[1959] "Means for transmitting voice data to a user" refers to a device or software for transmitting the generated voice data to a user's terminal and conveying it to the user.

[1960] "Means for accepting user responses" refers to a device or software that recognizes additional voice information provided by the user and incorporates that information into the system.

[1961] The "means for registering reservation information" refers to a device or software for registering reservation information received from a user in a database or the like.

[1962] The "means for generating a reservation confirmation message and converting it into voice data" refers to a device or software that generates a message to notify the user that the reservation has been confirmed and converts it into voice data.

[1963] The "means for transmitting voice data to the user" refers to a device or software for transmitting the generated voice data to the user's terminal and conveying the contents of the reservation confirmation to the user.

[1964] This invention is a system that uses a generative AI model to solve restaurant reservations, restaurant directions, checking business hours, and phone orders. This system accepts a user's voice request, converts it into text, generates and sends a prompt to the generative AI model, and then generates an appropriate response to provide to the user.

[1965] Hardware and software used

[1966] Hardware: Servers, user devices (smartphones, tablets, PCs, etc.)

[1967] Software: Generative AI models (e.g., large-scale language models), speech recognition software (e.g., speech recognition engines), and speech synthesis software (e.g., speech synthesis engines).

[1968] Specific operation of the system

[1969] 1. Acceptance of user requests

[1970] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[1971] Device: Uses speech recognition software to convert your speech into text.

[1972] Terminal: Sends the converted text to the server.

[1973] 2. Prompt generation of generative AI models

[1974] Server: Parses the received text and generates prompts to send to the generative AI model.

[1975] Server: Generate an example prompt: "A user would like to make a reservation. How many people would you like to book and what time would you like to make the reservation for?"

[1976] 3. Response generation using generative AI models

[1977] Server: Sends the generated prompts to the generative AI model.

[1978] Generative AI model: Generates appropriate responses based on prompts.

[1979] Generative AI model: Generates an example response: "How many people would you like to make a reservation for, and what time?"

[1980] 4. Voice synthesis of responses

[1981] Server: Passes the generated text response to speech synthesis software to generate audio data.

[1982] 5. Sending a response to the user

[1983] Server: Sends the generated voice data to the user device.

[1984] Terminal: Plays the audio data and communicates the response to the user.

[1985] 6. Acceptance of user responses

[1986] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[1987] Device: Uses speech recognition software to convert your speech into text.

[1988] Terminal: Sends the converted text to the server.

[1989] 7. Confirmation of reservation

[1990] Server: Receives the user's response and registers the information in the reservation system.

[1991] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[1992] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We are waiting for you with XX guests on XX / XX / XX at XX time."

[1993] 8. Sending a confirmation message

[1994] Server: The confirmation message is converted into voice using text-to-speech software.

[1995] Text-to-speech software: converts text into audio data.

[1996] Server: Sends the generated voice data to the user device.

[1997] Terminal: Plays audio data to inform the user that the reservation is complete.

[1998] Specific examples

[1999] User request: "I want to make a reservation to visit the store."

[2000] Prompt to generative AI model: "A user wants to make a reservation. How many people would like to come and what time would they like to make the reservation?"

[2001] The generative AI model responds: "How many people would you like to book and what time would you like to book?"

[2002] User's answer: "Two people, starting tomorrow at 7pm."

[2003] Reservation confirmation message: "Your reservation is complete. We look forward to seeing you tomorrow at 7pm for two people."

[2004] In this way, the system uses a generative AI model to generate an appropriate response to a user's request and confirm the reservation.

[2005] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2006] Step 1: Accepting a user request

[2007] User: Speaks into the terminal, saying, "I'd like to make a reservation."

[2008] Input: User's voice request

[2009] Device: Uses speech recognition software to convert your speech into text.

[2010] Data processing: Converting voice data into text data

[2011] Output: "I'd like to make an appointment"

[2012] Terminal: Sends the converted text to the server.

[2013] Step 2: Prompt generation for generative AI models

[2014] Server: Parses the received text and generates prompts to send to the generative AI model.

[2015] Input: "I'd like to make an appointment"

[2016] Data Calculation: Text Analysis and Prompt Generation

[2017] Output: A prompt saying "A user wants to make an appointment. How many people would like to book and what time would you like to book?"

[2018] Server: Sends prompts to the generative AI model.

[2019] Step 3: Generate a response using a generative AI model

[2020] Server: Sends the generated prompts to the generative AI model.

[2021] Input: The prompt "A user wants to make a reservation. How many people and what time would you like to book?"

[2022] Generative AI model: Generates appropriate responses based on prompts.

[2023] Data Calculation: Prompt-Based Response Generation

[2024] Output: Response "How many people would you like to book and what time?"

[2025] Generative AI model: Sends the response to the server.

[2026] Step 4: Speech synthesis to vocalize the response

[2027] Server: Passes the generated text response to speech synthesis software to generate audio data.

[2028] Input: Text response "How many people would you like to book and what time?"

[2029] Text-to-speech software: converts text into audio data.

[2030] Data processing: Convert text data into audio data

[2031] Output: Audio data

[2032] Server: Sends the generated voice data to the user device.

[2033] Step 5: Send a response to the user

[2034] Server: Sends the generated voice data to the user device.

[2035] Input: Audio data

[2036] Terminal: Plays the audio data and communicates the response to the user.

[2037] Output: A voice response saying "How many people would you like to book and what time would you like to book?"

[2038] Step 6: Accept user responses

[2039] User: Answers by voice regarding reservation information (number of people, date and time, etc.).

[2040] Input: User's spoken response

[2041] Device: Uses speech recognition software to convert your speech into text.

[2042] Data processing: Converting voice data into text data

[2043] Output: Text "Two people, starting tomorrow at 7pm"

[2044] Terminal: Sends the converted text to the server.

[2045] Step 7: Confirm your booking

[2046] Server: Receives the user's response and registers the information in the reservation system.

[2047] Input: Text "Two people, please meet tomorrow from 7pm"

[2048] Data calculation: Reservation information registration

[2049] Output: Reservation information registration completed

[2050] Server: Informs the generative AI model that the reservation has been confirmed and generates a confirmation message.

[2051] Generative AI model: An example of a confirmation message would be "Your reservation has been completed. We look forward to seeing you tomorrow at 7pm for two guests."

[2052] Output: Confirmation message

[2053] Step 8: Send a confirmation message

[2054] Server: The confirmation message is converted into voice using text-to-speech software.

[2055] Input: Confirmation message text

[2056] Text-to-speech software: converts text into audio data.

[2057] Data processing: Convert text data into audio data

[2058] Output: Audio data

[2059] Server: Sends the generated voice data to the user device.

[2060] Terminal: Plays audio data to inform the user that the reservation is complete.

[2061] Output: Voice response "Your reservation is complete. We will be waiting for you tomorrow at 7pm for two people."

[2062] (Application example 2)

[2063] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2064] In the past, restaurant operations such as making reservations, getting directions to the restaurant, checking business hours, and calling to order were often done manually, which was time-consuming and labor-intensive. In addition, users had to use multiple methods to obtain this information, which was inconvenient.

[2065] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an automatic voice response means, a voice-based generative AI means, a voice recognition means, a response generation means using a generative AI model, a means for making reservations based on a user's voice input, a means for providing directions, a means for checking business hours, and a means for automatically making an order call. This allows users to efficiently perform tasks such as making store reservations, getting directions to stores, checking business hours, and making an order call through a single system.

[2066] An "automatic voice response means" is a system that has the function of automatically responding to voice input from a user.

[2067] A "voice generation AI means" is a system that uses artificial intelligence technology to analyze voice input and generate an appropriate voice response.

[2068] "Speech recognition means" refers to a system that has the technology to convert a user's voice into text data.

[2069] A "response generation means using a generative AI model" is a system that has the function of using a generative AI model to generate an appropriate response based on user input.

[2070] The "means for making a reservation based on a user's voice input" is a system that has the function of processing a reservation request made by a user through voice and confirming the reservation.

[2071] A "means for providing directions" is a system that has the function of providing users with directions to their destination via voice or text.

[2072] The "means for checking business hours" is a system that has the function of checking the business hours of a store specified by the user and providing the information to the user.

[2073] "Means for automatically making an order call" refers to a system that has the function of automatically transmitting the specified order details over the phone based on the user's voice input.

[2074] The "answer database" is a database that stores answers to questions prepared in advance.

[2075] "Short Message Service Notification" means a system that has the function of notifying users via short message service.

[2076] The "means for asking questions to confirm matters" is a system that has the function of asking the user necessary questions to confirm matters and obtaining the answers.

[2077] "Means that allow for text saving of user responses" refers to a system that has the function of saving the content of a user's voice responses as text data.

[2078] The following system configuration will be described as an embodiment of the present invention.

[2079] System Configuration

[2080] This system includes a server, a user terminal, and a voice recognition device. The server includes a response generation means using a generative AI model, a voice-based generative AI means, a voice recognition means, and a database. The user terminal includes a microphone and a speaker and receives voice input from the user.

[2081] Hardware and software used

[2082] Hardware:

[2083] Microphone: Accepts user voice input.

[2084] Speaker: Provides audio responses from the server to the user.

[2085] Server: Provides the computational resources to run generative AI models.

[2086] software:

[2087] speech_recognition library: Performs speech recognition.

[2088] The transformers library: Runs generative AI models.

[2089] Data processing and calculation

[2090] 1. Speech Recognition:

[2091] A microphone on the user terminal accepts the user's voice input.

[2092] A speech recognition means converts the speech input in...

Claims

1. means for receiving a call from a user; means for converting said user's voice data into text; means for transmitting the converted text to a generative AI model as a prompt sentence; means for generating an answer using the generative AI model; means for converting the generated text of the answer into speech; means for providing the voice to the user as voice guidance; emotion identification means for determining an emotion value indicating the emotion of the user according to an emotion map; and a control means for controlling at least one of the generation of an answer by the generation AI model and the vocalization of the answer according to the emotion value, thereby adapting the content or tone of the voice guidance; The prompt sentence is dynamically generated based on the text converted by the means for converting the user's voice data into text and the emotion of the user identified by the emotion identification means. system.

2. The generative AI model can solve restaurant reservations, directions to restaurants, confirmation of business hours, and phone orders by generating answers based on the prompt sentence. The system of claim 1 .

3. By linking the prepared response database with the above-mentioned AI model, voice responses, message notifications, and confirmation questions can be sent, and customer responses can be saved as text. The system of claim 1 .

Citation Information

Patent Citations

  • Intelligent customer service multi-round dialogue optimization method based on GPT

    CN116610786A

  • Response system, response method, and computer program

    JP2015211403A

  • Artificial intelligence-based automatic response method and system

    JP2021022928A

  • Persona chatbot control method and system

    JP2022180282A

  • Speech Agent

    JP2022523504A