system
A system combining speech recognition, natural language processing, and generative AI addresses the complexity and security issues of current systems, enabling safe and easy purchasing and dining activities for the elderly.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Current systems for purchasing and dining activities are complex and pose security concerns, making them difficult for the elderly to use safely and conveniently.
A system utilizing speech recognition, natural language processing, and generative AI to convert voice inputs into text, analyze user intent, generate responses, and confirm orders through external service APIs, ensuring ease and safety.
Provides a safe and convenient system for the elderly to perform purchasing and dining activities using voice-based operations.
Smart Images

Figure 2026062185000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The senior generation, including the elderly, is required to safely and conveniently carry out purchasing and dining activities in daily life. Since the current system is complex to operate and has security concerns, many elderly people are hesitant to use it. Therefore, it is necessary to provide a system that can be intuitively operated by the elderly and has high safety.
Means for Solving the Problems
[0005] This invention comprises a speech recognition means for receiving user voice input and a natural language processing means for converting the voice into natural language text and analyzing the intent. Furthermore, it uses a generative AI means to generate responses based on the intent, thereby providing the user with appropriate instructions and suggestions. This generative AI means cooperates with an external service API to obtain necessary information and generate a response again. Finally, by using a means to request confirmation from the user, the invention provides a system that enables safe and convenient purchasing and dining activities.
[0006] "Voice recognition means" refers to a device or software that receives voice input from a user and converts it into text data.
[0007] "Natural language processing means" refers to a device or software that analyzes text converted by speech recognition means and identifies intents.
[0008] A "generative AI system" is an artificial intelligence system that generates appropriate responses based on intents identified by natural language processing systems.
[0009] "Presentation means" refers to a device or software that displays or reads aloud a response generated by a generative AI means to the user.
[0010] "Information acquisition means" refers to a device or software that acquires necessary information from an external service API in accordance with the user's intent.
[0011] "External service integration means" refers to a device or software that makes requests to external purchasing systems or delivery services based on the generated response.
[0012] A "confirmation device" is a device or software used to request final confirmation from the user. [Brief explanation of the drawing]
[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] System Overview
[0035] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, and generative AI technologies. Specifically, when a user gives a voice command, a speech recognition means converts this command into text data, and a natural language processing means analyzes it. Then, a generative AI means generates an appropriate response and presents the response to the user while obtaining necessary information in cooperation with an external service API. Finally, after user confirmation, the request to the external service is completed.
[0036] Explanation of the program's processing
[0037] 1. Speech recognition
[0038] Terminal: The user says by voice, "I want to order a pizza for dinner tonight."
[0039] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[0040] Terminal: Sends text data to the server.
[0041] 2. Natural Language Processing
[0042] Server: Receives text data sent from the terminal.
[0043] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0044] 3. Response generation
[0045] Server: A generative AI (e.g., GPT-3®) generates responses such as, "What kind of pizza would you like to order?".
[0046] Server: Query the delivery service API to retrieve information on available pizza menus.
[0047] Server: Based on this, the generative AI generates a response that presents the user with the options "Margherita pizza, pepperoni pizza, vegetarian pizza".
[0048] 4. Presentation and User Selection
[0049] Server: Sends the generated response to the terminal.
[0050] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[0051] User: For example, respond with, "I'd like to order a Margherita pizza."
[0052] 5. Re-speech recognition and response generation
[0053] Terminal: It uses speech recognition to convert the user's response into text and sends it to the server.
[0054] Server: Upon receiving the response, the generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[0055] Server: Sends a confirmation message to the terminal.
[0056] 6. Final confirmation and order completion
[0057] Terminal: Displays a confirmation message to the user.
[0058] User: For example, respond with, "Yes, please place the order."
[0059] Terminal: Converts the response to text and sends it to the server.
[0060] Server: Sends the final request to the delivery service API to confirm the order.
[0061] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0062] Server: Sends a completion message to the terminal.
[0063] Terminal: Displays a completion message to the user.
[0064] Specific example
[0065] As an example, consider a scenario where a user says, "I want to order a pizza for dinner tonight." A speech recognition system converts this into text, and a natural language processing system analyzes the intent. A generative AI system generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. The user then responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[0066] This will result in a safe and easy-to-use purchasing and dining support system that is suitable even for the elderly.
[0067] The following describes the processing flow.
[0068] Step 1:
[0069] User: The user says, "I want to order a pizza for dinner tonight."
[0070] Step 2:
[0071] Terminal: The terminal receives voice input.
[0072] Step 3:
[0073] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[0074] Step 4:
[0075] Terminal: Sends the converted text data to the server.
[0076] Step 5:
[0077] Server: The server receives text data received from the terminal.
[0078] Step 6:
[0079] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0080] Step 7:
[0081] Server: The generative AI generates the response, "What kind of pizza would you like to order?"
[0082] Step 8:
[0083] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0084] Step 9:
[0085] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[0086] Step 10:
[0087] Server: Sends the generated response to the terminal.
[0088] Step 11:
[0089] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[0090] Step 12:
[0091] User: The user replies, "I want to order a Margherita pizza."
[0092] Step 13:
[0093] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0094] Step 14:
[0095] Terminal: Sends the converted text data to the server.
[0096] Step 15:
[0097] Server: Receives the converted text data.
[0098] Step 16:
[0099] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that alright?".
[0100] Step 17:
[0101] Server: Sends the generated confirmation message to the terminal.
[0102] Step 18:
[0103] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[0104] Step 19:
[0105] User: The user replies, "Yes, please place the order."
[0106] Step 20:
[0107] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0108] Step 21:
[0109] Terminal: Sends the converted text data to the server.
[0110] Step 22:
[0111] Server: Receives the converted text data.
[0112] Step 23:
[0113] Server: Sends the final order request to the delivery service API to confirm the order.
[0114] Step 24:
[0115] Server: Receives order confirmations from the delivery service API.
[0116] Step 25:
[0117] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0118] Step 26:
[0119] Server: Sends the generated completion message to the terminal.
[0120] Step 27:
[0121] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[0122] This completes the user's pizza ordering process.
[0123] (Example 1)
[0124] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0125] In modern society, it is difficult for the elderly to smoothly make online purchases or order meals. In particular, technology that enables natural dialogue is necessary for the process of extracting appropriate information from voice input and accurately confirming orders. Current systems often suffer from low accuracy in voice recognition and natural language processing, or insufficient integration with external services, making them difficult for the elderly to use. There is a need to solve these problems and provide a purchasing and dining activity support system that can be easily used by the elderly.
[0126] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0127] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; speech recognition means for re-recognizing the user's voice input and converting it into text data; and means for presenting a confirmation message to the user that is generated again based on the response from the external service API. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[0128] "Speech recognition means" refers to a device or program that receives a user's speech as audio data and converts it into text data.
[0129] "Natural language processing means" refers to a device or program that analyzes natural language text converted by speech recognition means and understands the user's intent.
[0130] A "generative AI means" is a device or program that utilizes artificial intelligence technology to generate an appropriate response based on intents analyzed by natural language processing means.
[0131] "Presentation means" refers to a device or program that presents a response generated by a generative AI means to the user visually or audibly.
[0132] "Information acquisition means" refers to a device or program for acquiring information corresponding to a user's intent from an external service API.
[0133] "External service integration means" refers to a device or program for making requests to external services in accordance with the generated response.
[0134] A "confirmation means" is a device or program that presents a generated response to the user and performs the user's final confirmation.
[0135] Modes for carrying out the invention
[0136] This invention is a system that supports elderly people in purchasing and dining activities through simple voice-based operation. Specifically, it combines voice recognition, natural language processing, and generative AI technology so that when a user gives a voice command, the system performs the appropriate processing and ultimately confirms the order corresponding to the user's instructions. A detailed explanation is as follows.
[0137] System Configuration
[0138] The system consists mainly of the following components:
[0139] 1. Speech recognition method: This is a method for converting speech input into text data. Specifically, Google's Cloud Speech-to-Text API is used.
[0140] 2. Natural Language Processing (NLP) Methods: These are methods for analyzing text data and extracting the user's intent. Specifically, natural language processing modules such as spaCy are used.
[0141] 3. Generative AI methods: These are methods for generating appropriate responses based on intents analyzed by natural language processing methods. Specifically, OpenAI® GPT-3 and similar tools are used.
[0142] 4. Presentation means: These are means of presenting the generated response to the user. Specifically, this includes a speaker for reading the response aloud and a display for showing the text.
[0143] 5. Information Acquisition Method: This involves using external service APIs to obtain the necessary information. For example, the Uber Eats API is used.
[0144] 6. External service integration means: This is a means for making requests to external services according to the generated response.
[0145] 7. Verification Method: This is a means for obtaining final confirmation from the user. The speech recognition method is used again to present a response to the user and obtain final confirmation.
[0146] Example of operation
[0147] The system works as follows:
[0148] 1. Voice input:
[0149] The user says, "I want to order a pizza for dinner tonight." This audio data is captured by the device.
[0150] 2. Speech recognition:
[0151] The device uses speech recognition (Google Cloud Speech-to-Text API) to convert the voice data into text data: "I want to order pizza for dinner tonight."
[0152] 3. Natural Language Processing:
[0153] The server uses natural language processing (spaCy) to analyze the text data and identify the user's intent, "Pizza Order".
[0154] 4. Response generation:
[0155] The server uses generative AI tools (OpenAI GPT-3) to generate the response, "What kind of pizza would you like to order?"
[0156] 5. Information acquisition:
[0157] The server uses an information retrieval method (Uber Eats API) to obtain information about available pizza menus.
[0158] 6. Response presentation:
[0159] Based on the menu information acquired by the generative AI, the server generates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[0160] 7. User selection and voice recognition again:
[0161] The terminal presents a response to the user, who replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to convert this response into text.
[0162] 8. Final confirmation and order completion:
[0163] The server uses a generative AI system to generate a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[0164] The terminal displays a confirmation message to the user, who responds, "Yes, please place the order." The terminal then uses speech recognition to convert this response back into text and sends it to the server.
[0165] The server sends the final request to an external service API to confirm the order. Upon receiving the order confirmation from the Uber Eats API, the server uses generative AI tools to generate a completion message that reads, "Your order is complete. Your pizza is expected to arrive in 30 minutes."
[0166] The device displays a completion message to the user.
[0167] Example of a prompt
[0168] An example of a prompt to input into a generative AI would be: "The user says they want to order a pizza. Ask the user what kinds of pizza are available."
[0169] The above is an embodiment of the present invention. This system provides the convenience of allowing even elderly people to easily purchase items and order meals using only their voice.
[0170] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0171] Step 1: Capture voice input
[0172] User: The user says, "I want to order a pizza for dinner tonight." This voice input becomes the system's initial input.
[0173] Step 2: Speech Recognition
[0174] Terminal: The terminal receives the user's voice input and sends it to a speech recognition system (e.g., Google Cloud Speech-to-Text API). The voice data is processed by the API, and the text data "I want to order pizza for dinner tonight" is generated.
[0175] Input: User's voice data
[0176] Output: Text data "I want to order a pizza for dinner tonight"
[0177] Specific operation: The device sends voice data to the API and receives the converted text data.
[0178] Step 3: Send text data
[0179] Terminal: Sends the generated text data to the server.
[0180] Input: Text data "I want to order a pizza for dinner tonight"
[0181] Output: Sending text data to the server
[0182] Specific action: The terminal sends text data to the server as an HTTP request.
[0183] Step 4: Natural Language Processing
[0184] Server: The server uses a natural language processing module (e.g., spaCy) to parse the text data and identify that the user's intent is "pizza order".
[0185] Input: Text data "I want to order a pizza for dinner tonight"
[0186] Output: Analysis result "Pizza order"
[0187] Specific operation: The server receives text data as input, parses it using a natural language processing module, and extracts intents.
[0188] Step 5: Create a prompt for response generation
[0189] Server: Creates and sends a prompt to a generative AI (e.g., OpenAI GPT-3) that says, "The user wants to order a pizza. Ask them what kind of pizza they want to order."
[0190] Input: Intent "Pizza Order"
[0191] Output: Prompt "The user says they want to order a pizza. Ask them what kind of pizza they would like to order."
[0192] Specific operation: The server creates a prompt to send to the generative AI based on the intent.
[0193] Step 6: Generate initial response
[0194] Server: The generative AI generates a response based on the prompt, asking, "What kind of pizza would you like to order?"
[0195] Input: Prompt
[0196] Output: Response "What kind of pizza would you like to order?"
[0197] Specific operation: The generative AI analyzes the prompt and generates an appropriate response.
[0198] Step 7: Obtaining menu information
[0199] Server: The server queries an information retrieval tool (e.g., Uber Eats API) to obtain information on available pizza menus.
[0200] Input: Intent "Pizza Order"
[0201] Output: Pizza menu information
[0202] Specific operation: The server sends a request to the API and receives menu information.
[0203] Step 8: Enhancing the response content
[0204] Server: Based on the acquired menu information, the generative AI regenerates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[0205] Input: Pizza menu information
[0206] Output: Detailed response "We have Margherita pizza, pepperoni pizza, and vegetarian pizza. Which type would you like?"
[0207] Specific operation: The generative AI generates a new response based on the menu information.
[0208] Step 9: Sending and presenting a response
[0209] Server: Sends the generated detailed response to the terminal.
[0210] Terminal: The terminal either reads the response aloud to the user or displays it as text.
[0211] Input: Detailed response
[0212] Output: Presenting a response to the user
[0213] Specific operation: The server sends a detailed response to the terminal in text format, and the terminal presents it to the user in an appropriate format.
[0214] Step 10: User selection and voice recognition again
[0215] User: The user replies, "I would like to order a Margherita pizza."
[0216] Terminal: Re-recognizes the user's voice, converts it into text data, and sends it to the server.
[0217] Input: User's voice response
[0218] Output: Text data "I want to order a Margherita pizza"
[0219] Specific operation: The speech recognition system converts the user's response into text and sends it to the server.
[0220] Step 11: Generating the final confirmation message
[0221] Server: The generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that alright?"
[0222] Input: Text data "I want to order a Margherita pizza"
[0223] Output: Confirmation message
[0224] Specific operation: The generative AI generates a confirmation message based on the text it receives.
[0225] Step 12: Presentation of confirmation message
[0226] Server: Sends a confirmation message to the terminal.
[0227] Terminal: The terminal displays a confirmation message to the user.
[0228] Input: Confirmation message
[0229] Output: Presentation of a confirmation message to the user.
[0230] Specific operation: The server sends a confirmation message to the terminal, and the terminal presents it to the user via voice or text.
[0231] Step 13: User confirmation and speech recognition
[0232] User: The user replies, "Yes, please place the order."
[0233] Terminal: The speech recognition system converts the user's response back into text and sends it to the server.
[0234] Input: User's final confirmation voice
[0235] Output: Text data "Yes, please place your order"
[0236] Specific operation: The device converts the audio to text and sends it to the server.
[0237] Step 14: Confirm the order and generate the completion message.
[0238] Server: Sends the final request to an external service API (e.g., Uber Eats API) to confirm the order.
[0239] Information retrieval: Order confirmation via API
[0240] Output: Order confirmation and completion messages
[0241] Specific operation: The server sends a request to the API, and after receiving confirmation of the order, it generates the message, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0242] Step 15: Send and present completion message
[0243] Server: Sends a completion message to the terminal.
[0244] Terminal: The terminal displays a completion message to the user.
[0245] Input: Completion message
[0246] Output: Presentation of a completion message to the user.
[0247] Specific operation: The server sends a completion message to the terminal, and the terminal presents the message to the user via voice or text.
[0248] (Application Example 1)
[0249] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0250] For the elderly and those unfamiliar with technology, ordering food delivery online using smartphones or PCs presents complex, time-consuming, and stressful challenges. Furthermore, long lists and numerous options can lead to confusion when confirming specific order details and making final confirmations. This invention aims to solve these problems and provide a system that allows the elderly to easily and safely order food delivery using voice-based technology.
[0251] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0252] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; and means for ordering food and beverages based on the user's voice input, thereby presenting available menu options and confirming the order after final confirmation. This makes it possible for elderly people to easily and safely order food delivery using only their voice.
[0253] "Voice recognition means" refers to technology that converts a user's voice input into text data.
[0254] "Natural language processing means" refers to technologies that analyze converted natural language text and understand the user's intent.
[0255] "Generative AI methods" refer to artificial intelligence technologies that generate appropriate responses based on analyzed intents.
[0256] "Presentation means" refers to a technology that presents responses generated by generative AI means to the user visually or audibly.
[0257] "Information acquisition means" refers to the technology of obtaining information corresponding to a user's intent from an external service API.
[0258] "External service integration means" refers to a technology that makes requests to external services according to the generated response.
[0259] A "confirmation method" is a technology that presents the generated response to the user for final confirmation.
[0260] "Means for ordering food and beverages" refers to technology that places orders for food and beverages based on the user's voice input and confirms the order details.
[0261] This invention provides a system that allows elderly people to easily order food delivery, and is provided by combining speech recognition technology, natural language processing technology, and generative AI technology. The specific implementation of this system includes the following elements.
[0262] hardware
[0263] Smartphone: A device used by users for voice input. Generally, iOS or Android (registered trademark) smartphones are used.
[0264] Server: A central system that performs speech recognition, natural language processing, and generative AI processing.
[0265] software
[0266] Speech recognition software:
[0267] This process converts voice input acquired from the device into text data. Specifically, it uses the Google Cloud Speech-to-Text API.
[0268] Natural language processing software:
[0269] The converted text data is analyzed to understand the user's intent. This is done using the Google Cloud Natural Language API.
[0270] Generative AI software:
[0271] The system generates a response based on the analyzed data. OpenAI's GPT-3 protocol is used.
[0272] External service API:
[0273] An API for integrating with food delivery services. For example, using the Uber Eats API or other food delivery service APIs.
[0274] Processing flow
[0275] 1. Speech recognition:
[0276] The user says "I want to order a pizza" into their smartphone. The device captures the audio and uses the Google Cloud Speech-to-Text API to convert the speech to text.
[0277] 2. Natural Language Processing:
[0278] The converted text data is sent to the server, where the Google Cloud Natural Language API is used to analyze the user's intent. This analysis identifies the user's intent to "order a pizza."
[0279] 3. Response generation and presentation:
[0280] Based on the analysis results, a generative AI (OpenAI GPT-3) generates a response asking "What kind of pizza would you like to order?", and the server sends this to the user's terminal.
[0281] 4. Integration with external services and order confirmation:
[0282] Retrieve the pizza menu available from the food delivery service API (e.g., UberEats API) and present it to the user as options. If the user selects "want to order a Margherita pizza" and sends it to the server again. The server uses generative AI to generate a final confirmation message (e.g., "Do you want to order one Margherita pizza?") and presents it to the user. Once the user's final confirmation is obtained, send the order to the food delivery service API to confirm the order.
[0283] Specific example
[0284] Consider a scenario where the user says "want to order pizza for dinner tonight". The speech recognition software converts this into text, and the natural language processing software analyzes the intent. The generative AI software generates a response "Which type of pizza do you want to order?" and retrieves menu information from an external service API. Then, "Margherita pizza" is selected and the order is confirmed through the confirmation means.
[0285] Example of prompt text
[0286] User's instruction: "Want to order pizza"
[0287] Generative AI's response: "Which type of pizza do you want to order?"
[0288] [[ID=#23]] User's instruction: "Margherita pizza"
[0289] Generative AI's final confirmation response: "Do you want to order one Margherita pizza?"
[0290] With these technical elements and processing flows, the elderly can easily place food delivery orders using voice.
[0291] The flow of specific processing in Application Example 1 will be described using Figure 12.
[0292] Step 1:
[0293] Voice input capture and conversion
[0294] The user provides voice input (e.g., "I want to order a pizza for dinner tonight"). The device captures this audio and sends it to the Google Cloud Speech-to-Text API to convert the audio data into text.
[0295] Input: User voice: "I want to order a pizza for dinner tonight."
[0296] Data processing: Audio data → Text data "I want to order a pizza for dinner tonight."
[0297] Output: Text data "I want to order a pizza for dinner tonight"
[0298] Step 2:
[0299] Text transmission and natural language processing
[0300] The device sends the converted text data to the server. The server uses the Google Cloud Natural Language API to analyze this text and identify the user's intent (e.g., ordering a pizza).
[0301] Input: Text data "I want to order a pizza for dinner tonight"
[0302] Data processing: Natural language processing → Intent "Pizza order"
[0303] Output: Intent "Pizza Order"
[0304] Step 3:
[0305] Response generation
[0306] Based on the analysis results, the server generates an appropriate response using generative AI (OpenAI GPT-3). Here, "What kind of pizza would you like to order?" is generated.
[0307] Input: Intent "Pizza order"
[0308] Data processing: Intent → Response generation
[0309] Output: Response "What kind of pizza would you like to order?"
[0310] Step 4:
[0311] Presentation of the response
[0312] The server sends the generated response to the terminal. The terminal presents the response to the user in an appropriate way. For example, text display or voice output.
[0313] [[ID=z7]] Input: Response "What kind of pizza would you like to order?"
[0314] Data processing: None
[0315] Output: Displayed response, voice output
[0316] Step 5:
[0317] User selection and re - voice recognition
[0318] The user makes a selection for the response (e.g., "I want to order a Margherita pizza"). The terminal captures this new voice input and sends it again to the Google Cloud Speech - to - Text API to convert it into text data.
[0319] Input: User voice "I want to order a Margherita pizza"
[0320] Data processing: Voice data → Text data "I want to order a Margherita pizza"
[0321] Output: Text data "I want to order a Margherita pizza"
[0322] Step 6:
[0323] Sending text and re-processing natural language
[0324] The device sends the converted text data back to the server. The server then uses the Google Cloud Natural Language API again to parse the user's intent (e.g., selecting a specific pizza type).
[0325] Input: Text data "I want to order a Margherita pizza"
[0326] Data processing: Natural language processing → Intent "Order for Margherita pizza"
[0327] Output: Intent "Order a Margherita pizza"
[0328] Step 7:
[0329] Generation of final acknowledgment
[0330] The server uses a generative AI (OpenAI GPT-3) based on the analysis results to generate a final confirmation response. In this case, the response "I would like to order one Margherita pizza. Is that alright?" is generated.
[0331] Input: Intent "Order a Margherita pizza"
[0332] Data processing: Intent → Final confirmation response generation
[0333] Output: Final confirmation response: "I would like to order one Margherita pizza. Is that correct?"
[0334] Step 8:
[0335] Final confirmation presentation and confirmation response transmission
[0336] The server sends the generated final confirmation to the terminal, which then presents it to the user. The user replies, "Yes, please place the order." The terminal captures this reply and sends it back to the Google Cloud Speech-to-Text API to convert it back into text data.
[0337] Input: Final confirmation response "I would like to order one Margherita pizza. Is that correct?"
[0338] Data processing: None
[0339] Output: User confirmation "Yes, please place the order"
[0340] Step 9:
[0341] Final order confirmation
[0342] The terminal sends final confirmation text data to the server. The server sends the final order information to an external service API (e.g., a food delivery API) to confirm the order.
[0343] Input: User confirmation "Yes, please place the order"
[0344] Data processing: Sending final order information
[0345] Output: Confirmed order "One Margherita Pizza"
[0346] Step 10:
[0347] Generating and displaying completion messages
[0348] The server receives an order confirmation from the food delivery API, generates a completion message (e.g., "Your order is complete. Your pizza is expected to arrive in 30 minutes") using OpenAI GPT-3, and sends it to the terminal. The terminal then displays this message to the user.
[0349] Input: Confirmed order information
[0350] Data processing: Order confirmation → Completion message generation
[0351] Output: Completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0352] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0353] System Overview
[0354] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, generative AI, and emotion recognition technology. When a user gives a voice command, a speech recognition means converts the voice into text, and a natural language processing means analyzes the text to identify the user's intent. A generative AI generates an appropriate response and obtains necessary information in cooperation with an external service API. Subsequently, an emotion engine analyzes the user's emotions and adjusts the response based on the results. Finally, the user is asked for confirmation and the order is completed.
[0355] Explanation of the program's processing
[0356] 1. Speech recognition
[0357] Terminal: The user says, "I want to order a pizza for dinner tonight."
[0358] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[0359] Terminal: Sends text data to the server.
[0360] 2. Natural Language Processing
[0361] Server: Receives text data sent from the terminal.
[0362] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0363] 3. Emotion recognition
[0364] Server: The emotion engine analyzes the user's emotions from text data.
[0365] Server: The emotion engine analyzes the emotion data and sends it to the generative AI.
[0366] 4. Response generation
[0367] Server: A generative AI (e.g., GPT-3) generates a response such as, "What kind of pizza would you like to order?"
[0368] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0369] Server: The generative AI generates responses that present the user with options such as "Margherita pizza, pepperoni pizza, and vegetarian pizza" based on menu information, and adjusts them according to the user's mood.
[0370] 5. Presentation and User Selection
[0371] Server: Sends the generated response to the terminal.
[0372] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[0373] User: For example, respond with, "I'd like to order a Margherita pizza."
[0374] 6. Re-speech recognition and response generation
[0375] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0376] Terminal: Sends the converted text data to the server.
[0377] Server: Upon receiving the response, the generative AI generates a confirmation message, "I would like to order one Margherita pizza. Is that alright?", and makes adjustments based on the user's emotions.
[0378] Server: Sends a confirmation message to the terminal.
[0379] 7. Final confirmation and order completion
[0380] Terminal: Displays a confirmation message to the user.
[0381] User: For example, respond with, "Yes, please place the order."
[0382] Terminal: Converts the response to text and sends it to the server.
[0383] Server: Sends the final request to the delivery service API to confirm the order.
[0384] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0385] Server: Sends a completion message to the terminal.
[0386] Terminal: Displays a completion message to the user.
[0387] Specific example
[0388] For example, if a user says, "I want to order a pizza for dinner tonight," a speech recognition system converts this into text. A natural language processing system analyzes the intent, and an emotion engine recognizes the user's emotions. A generative AI system generates a response such as, "What kind of pizza would you like to order?" and retrieves menu information from an external service API. The user responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[0389] The emotion engine reduces the anxiety and questions users experience during the ordering process, enabling the provision of reassuring and accurate support. This will result in a purchasing and dining activity support system that is easy for even the elderly to use.
[0390] The following describes the processing flow.
[0391] Step 1:
[0392] User: The user says, "I want to order a pizza for dinner tonight."
[0393] Step 2:
[0394] Terminal: The terminal receives voice input.
[0395] Step 3:
[0396] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[0397] Step 4:
[0398] Terminal: Sends the converted text data to the server.
[0399] Step 5:
[0400] Server: The server receives text data sent from the terminal.
[0401] Step 6:
[0402] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0403] Step 7:
[0404] Server: The emotion engine analyzes the user's emotions from text data.
[0405] Step 8:
[0406] Server: The emotion engine sends the recognized emotion data to the generative AI.
[0407] Step 9:
[0408] Server: The generative AI generates a response to the question "What kind of pizza would you like to order?", adjusting it based on the user's emotions.
[0409] Step 10:
[0410] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0411] Step 11:
[0412] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[0413] Step 12:
[0414] Server: Sends the generated response to the terminal.
[0415] Step 13:
[0416] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[0417] Step 14:
[0418] User: The user replies, "I want to order a Margherita pizza."
[0419] Step 15:
[0420] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0421] Step 16:
[0422] Terminal: Sends the converted text data to the server.
[0423] Step 17:
[0424] Server: Receives the converted text data.
[0425] Step 18:
[0426] Server: The emotion engine analyzes the user's emotions again from the text data.
[0427] Step 19:
[0428] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that correct?".
[0429] Step 20:
[0430] Server: Sends the generated confirmation message to the terminal.
[0431] Step 20:
[0432] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[0433] Step 21:
[0434] User: The user replies, "Yes, please place the order."
[0435] Step 22:
[0436] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0437] Step 23:
[0438] Terminal: Sends the converted text data to the server.
[0439] Step 24:
[0440] Server: Receives the converted text data.
[0441] Step 25:
[0442] Server: Sends the final order request to the delivery service API to confirm the order.
[0443] Step 26:
[0444] Server: Receives order confirmations from the delivery service API.
[0445] Step 27:
[0446] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0447] Step 28:
[0448] Server: Sends the generated completion message to the terminal.
[0449] Step 29:
[0450] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[0451] This allows the user's pizza ordering process to be completed in a way that includes emotion recognition.
[0452] (Example 2)
[0453] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0454] Today, many elderly people find online shopping and dining difficult due to the complex interfaces and operation methods. Furthermore, despite advancements in speech recognition and natural language processing technologies, comprehensive support systems, including emotion recognition, are still lacking. As a result, users often struggle to resolve anxieties and questions they encounter during the ordering process.
[0455] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0456] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making requests to external services according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; emotion recognition means for analyzing the user's emotions from the natural language text; and means for adjusting the response based on the emotion data analyzed by the emotion recognition means. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[0457] "Voice recognition means" refers to a technology or device that receives voice input from a user and converts it into text data.
[0458] "Natural language processing means" refers to a technology or device that analyzes natural language text converted by speech recognition means and understands its intent and meaning.
[0459] "Generative AI means" refers to artificial intelligence technology or systems that generate appropriate responses based on intents analyzed by natural language processing means.
[0460] "Presentation means" refers to a technology or device for showing a response generated by a generative AI means to a user.
[0461] "Information acquisition means" refers to a technology or system that obtains information corresponding to a user's intent from an external service API.
[0462] "External service integration means" refers to a technology or system that makes requests to external services according to the generated response.
[0463] A "confirmation method" is a technology or system that presents a response to the user and performs final confirmation.
[0464] "Emotion recognition means" refers to a technology or system that analyzes a user's emotions from natural language text.
[0465] "Means for adjusting responses" refers to technologies or systems for appropriately modifying responses based on emotional data analyzed by emotion recognition means.
[0466] This invention is a system that assists elderly people in easily engaging in purchasing and dining activities through voice communication. Specifically, it is realized through a system that integrates voice recognition, natural language processing, generative AI, and emotion recognition technology.
[0467] Hardware and software to be used
[0468] This system consists of the following main components:
[0469] 1. Speech recognition means:
[0470] The hardware uses a microphone, and the software uses a speech recognition system such as the Google Speech-to-Text API.
[0471] 2. Natural language processing methods:
[0472] The software used will be a natural language processing library (such as spaCy or BERT).
[0473] 3. Generative AI means:
[0474] For example, generative AI models such as OpenAI's GPT-3 can be used.
[0475] 4. Emotion recognition means:
[0476] The software used includes emotion recognition engines such as IBM Watson® Tone Analyzer.
[0477] 5. External Service APIs:
[0478] Use an API (such as the Uber Eats API) to retrieve and send user order information.
[0479] Operating principle
[0480] The operation of this system begins with the user's voice input and goes through several processing steps before finally sending the order to an external service.
[0481] Voice input:
[0482] The user speaks into the microphone and says, "I'd like to order a pizza for dinner tonight."
[0483] Speech recognition:
[0484] The device receives the audio data via its microphone and uses the Google Speech-to-Text API to generate the text data "I want to order a pizza for dinner tonight." This text data is then sent to the server.
[0485] Natural language processing:
[0486] The server analyzes the received text data using a natural language processing library to identify that the user's intent is "pizza order".
[0487] Emotion recognition:
[0488] The server uses an emotion recognition engine to analyze the user's emotions from text data. Based on the results, it generates data to adjust its response.
[0489] Response generation:
[0490] Generative AI (for example, OpenAI's GPT-3) generates a response such as, "What kind of pizza would you like to order?". In doing so, it queries a delivery service API to obtain menu information on available pizzas (for example, "Margherita pizza, pepperoni pizza, vegetarian pizza") and adjusts the response based on sentiment data.
[0491] Presentation and selection:
[0492] The server sends the generated response to the terminal, which then presents it to the user via voice or text. The user responds, for example, "I would like to order a Margherita pizza."
[0493] Speech recognition and response generation again:
[0494] The device accepts voice input again, performs speech recognition to convert it into text data, and sends it to the server. The server understands the response, and a generative AI generates a confirmation message (for example, "I'd like to order one Margherita pizza. Is that alright?") and makes adjustments based on emotion.
[0495] Final confirmation and order completion:
[0496] The server generates a final confirmation message which is sent to the terminal and displayed to the user. The user replies, "Yes, please place the order," and sends it to the server. The server sends the final order request to the delivery service API, and the order is confirmed. Finally, the generative AI generates a completion message, "Your order is complete. Your pizza is expected to arrive within 30 minutes," and displays this to the user.
[0497] Specific example
[0498] For example, a user's voice input, such as "I want to order a pizza for dinner tonight," is received by the microphone and converted into text data using the Google Speech-to-Text API. Next, this text data is analyzed by a natural language processing library to identify the intent "order pizza." An emotion recognition engine analyzes the user's emotions, and a generative AI generates and presents the response, "What kind of pizza would you like to order?" The user replies, "I would like to order a Margherita pizza," and after confirmation, the order is finalized.
[0499] Example of a prompt:
[0500] User: I want to order a pizza for dinner tonight.
[0501] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0502] Step 1:
[0503] The user gives a voice command saying, "I want to order a pizza for dinner tonight." The device receives this voice data via its microphone. The received voice data is sent to a speech recognition system (Google Speech-to-Text API), which analyzes it and generates the text data "I want to order a pizza for dinner tonight." The device then sends the generated text data to the server.
[0504] Step 2:
[0505] The server receives the text data "I want to order pizza for dinner tonight" from the terminal. Using natural language processing tools (e.g., spaCy, BERT), the server analyzes the received text data and identifies that the user's intent is "order pizza". The result of the intent identification is stored in an internal database.
[0506] Step 3:
[0507] The server uses emotion recognition tools (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., joy, anxiety) from text data. The emotion data is stored in an internal database for transmission to generative AI tools.
[0508] Step 4:
[0509] The server invokes a generative AI (such as OpenAI's GPT-3) to generate an appropriate response, "What kind of pizza would you like to order?". During this process, the server queries a delivery service API (e.g., the Uber Eats API) to obtain information on available pizza menus. Based on this menu information, the generative AI generates a response, which is then refined based on sentiment data. The server then sends the generated response message to the device.
[0510] Step 5:
[0511] The terminal receives a response message from the server: "What kind of pizza would you like to order?" It either reads this message aloud to the user or displays it as text. The user replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to process the user's response again, converts it back into text data, "I would like to order a Margherita pizza," and sends it to the server.
[0512] Step 6:
[0513] The server receives the text data "I want to order a Margherita pizza" sent again from the terminal. The generative AI generates a confirmation message "I would like to order one Margherita pizza. Is that alright?". The generated confirmation message is then adjusted based on sentiment data and sent to the terminal.
[0514] Step 7:
[0515] The terminal displays a confirmation message received from the server, "You would like to order one Margherita pizza. Is that alright?" (either read aloud or displayed as text). The user replies, "Yes, please place the order." The terminal uses a speech recognition system to convert the user's reply into text data and sends it to the server.
[0516] Step 8:
[0517] The server receives the final text data "Yes, please place the order" and sends the final request to the delivery service API. Once the order is confirmed, it receives confirmation from the delivery service API, and the generative AI generates a completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes." The completion message is sent to the device, which then displays it to the user (either read aloud or displayed as text).
[0518] (Application Example 2)
[0519] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0520] A system is needed that allows users, including the elderly, to easily support their daily purchasing and dining activities using voice commands. In particular, a system is required that analyzes the user's emotions and provides appropriate responses based on those emotions, allowing users to use the service without feeling anxious or confused. Furthermore, a mechanism is needed that allows the elderly to complete orders and inquiries without performing complex operations.
[0521] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a speech recognition means, a natural language processing means, a response generation means, an emotion recognition means, a generative AI means, a presentation means, an information acquisition means, an external linkage means, and a confirmation means. This enables the provision of appropriate responses based on the intent and emotion of the user when they easily place an order or make an inquiry using voice input, thereby supporting purchasing and dining activities that can be used safely even by the elderly.
[0522] "Voice recognition means" refers to devices or technologies that convert a user's voice input into text data.
[0523] "Natural language processing means" refers to technology that analyzes natural language text converted by speech recognition means and understands the user's intent.
[0524] "Response generation means" refers to a technology that generates an appropriate response based on intents analyzed by natural language processing means.
[0525] "Emotion recognition means" refers to a technology that analyzes a user's emotions based on the response generated by a response generation means.
[0526] "Generative AI methods" refer to artificial intelligence technologies that adjust and generate responses based on emotional data analyzed by emotion recognition methods.
[0527] "Presentation means" refers to technologies or devices that present a user with a refined response generated by a generative AI means.
[0528] "Information acquisition means" refers to technology that acquires information corresponding to the user's intent from an external information provider.
[0529] "External collaboration means" refers to a technology that makes requests to external information providers according to the generated response.
[0530] A "confirmation method" refers to a technology or device that presents a response to the user and performs final confirmation.
[0531] In a mode for carrying out the invention, a food delivery support system is realized that allows elderly people to easily order meals by voice by using a system that combines speech recognition, natural language processing, generative AI, and emotion recognition technology.
[0532] The server includes the following hardware and software: First, it uses the speech_recognition library for speech recognition. Second, it employs the GPT-3 model from the transformers library for natural language processing. Furthermore, it uses the Hugging Face sentiment analysis model for emotion recognition. A program to integrate these elements and manage the entire system is implemented in Python.
[0533] Specifically, when a user gives a voice command such as "I want to order curry for dinner tonight," a voice recognition system converts the voice into text and sends it to the server. The server uses a natural language processing system to analyze this text data and recognize the user's intent. Next, an emotion recognition system analyzes the user's emotions and sends the result to a generative AI system. The generative AI system generates an appropriate response based on the user's emotions and issues a prompt such as "What kind of curry would you like to order?"
[0534] The server, as an external means of providing information, connects with the API of a food delivery service to retrieve selectable curry menus. Based on the retrieved menu information, a generative AI generates a list including "butter chicken curry, beef curry, vegetable curry," etc., and presents it to the user through a presentation system.
[0535] For example, if a user replies, "I'd like to order butter chicken curry," the speech recognition system converts this speech back into text and sends it to the server. The server performs natural language processing and sentiment recognition as before, and then a generative AI system generates a confirmation message, "Is butter chicken curry alright?" Once the user confirms, the order is finalized using an external integration system, and a final confirmation message is presented to the user.
[0536] Example of a prompt:
[0537] "The user's intention is to order curry, and their emotion is calmness. The available menu items are butter chicken curry, beef curry, and vegetable curry. Which curry would you like to choose?"
[0538] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0539] Step 1:
[0540] The user gives instructions for their order by voice. For example, the user might say, "I'd like to order curry for dinner tonight." The voice data becomes the input.
[0541] Step 2:
[0542] The device uses speech recognition to convert the user's voice into text data. The text output is "I want to order curry for dinner tonight." The speech recognition library is used.
[0543] Step 3:
[0544] The terminal sends the generated text data to the server. The server receives this text data.
[0545] Step 4:
[0546] The server uses natural language processing to analyze text data and recognize the user's intent. In this example, the intent is "order curry". The GPT-3 model from the transformers library is used for this process. The intent is output as the analysis result.
[0547] Step 5:
[0548] The server uses emotion recognition to analyze the user's emotions from text data. In this example, the emotion is determined to be "calm." The Hugging Face emotion analysis model is used as the emotion recognition tool. Emotion data is output as the analysis result.
[0549] Step 6:
[0550] The server uses generative AI tools to generate an appropriate response based on the user's intent and emotions. In this example, the prompt "What kind of curry would you like to order?" is generated. A generative AI model (e.g., GPT-3) is used for this process. The generated response is then output.
[0551] Step 7:
[0552] The server connects with an external information source (a food delivery service API) to retrieve information on available menu items. The retrieved information includes "butter chicken curry, beef curry, and vegetable curry." The menu information retrieved from the API is then output.
[0553] Step 8:
[0554] The server generates a more specific response based on the retrieved menu information. This response creates a prompt containing a list of options: "Butter Chicken Curry, Beef Curry, Vegetable Curry." The generated specific response is then output.
[0555] Step 9:
[0556] The server presents the response to the user through a presentation mechanism. The user responds, "I would like to order butter chicken curry." This audio data becomes the input.
[0557] Step 10:
[0558] The device uses speech recognition again to convert the user's response into text data. The text output is "I would like to order butter chicken curry."
[0559] Step 11:
[0560] The terminal sends the generated text data to the server. The server receives this text data.
[0561] Step 12:
[0562] The server again uses natural language processing and emotion recognition to analyze the user's intent and emotion. It recognizes the intent as "order butter chicken curry" and the emotion as "calm." The intent and emotion are output as analysis results.
[0563] Step 13:
[0564] The server uses a generative AI system to generate a message for final confirmation. The response "Is butter chicken curry alright?" is generated. The generated confirmation response is output.
[0565] Step 14:
[0566] The server presents an acknowledgment to the user through a display mechanism. The user responds with "Yes, please place the order." The voice data becomes the input.
[0567] Step 15:
[0568] The terminal uses voice recognition again to convert the user's confirmation response into text data. The text "Yes, please place your order" is output.
[0569] Step 16:
[0570] The server receives the generated text data and uses a generative AI to finally generate a confirmation message. A completion message such as "Your order is complete. Your pizza is expected to arrive within 30 minutes" is generated. The generated completion message is then output.
[0571] Step 17:
[0572] The server sends the user's order information to the food delivery service's API via an external connection method, and confirms the order. A confirmation message is output from the API.
[0573] Step 18:
[0574] The server displays a completion message to the user through a designated means. The user confirms that the order has been completed.
[0575] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0576] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0577] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0578] [Second Embodiment]
[0579] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0580] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0581] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0582] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0583] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0584] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0585] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0586] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0587] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0588] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0589] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0590] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0591] System Overview
[0592] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, and generative AI technologies. Specifically, when a user gives a voice command, a speech recognition means converts this command into text data, and a natural language processing means analyzes it. Then, a generative AI means generates an appropriate response and presents the response to the user while obtaining necessary information in cooperation with an external service API. Finally, after user confirmation, the request to the external service is completed.
[0593] Explanation of the program's processing
[0594] 1. Speech recognition
[0595] Terminal: The user says by voice, "I want to order a pizza for dinner tonight."
[0596] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[0597] Terminal: Sends text data to the server.
[0598] 2. Natural Language Processing
[0599] Server: Receives text data sent from the terminal.
[0600] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0601] 3. Response generation
[0602] Server: A generative AI (e.g., GPT-3) generates responses such as, "What kind of pizza would you like to order?".
[0603] Server: Query the delivery service API to retrieve information on available pizza menus.
[0604] Server: Based on this, the generative AI generates a response that presents the user with the options "Margherita pizza, pepperoni pizza, vegetarian pizza".
[0605] 4. Presentation and User Selection
[0606] Server: Sends the generated response to the terminal.
[0607] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[0608] User: For example, respond with, "I'd like to order a Margherita pizza."
[0609] 5. Re-speech recognition and response generation
[0610] Terminal: It uses speech recognition to convert the user's response into text and sends it to the server.
[0611] Server: Upon receiving the response, the generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[0612] Server: Sends a confirmation message to the terminal.
[0613] 6. Final confirmation and order completion
[0614] Terminal: Displays a confirmation message to the user.
[0615] User: For example, respond with, "Yes, please place the order."
[0616] Terminal: Converts the response to text and sends it to the server.
[0617] Server: Sends the final request to the delivery service API to confirm the order.
[0618] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0619] Server: Sends a completion message to the terminal.
[0620] Terminal: Displays a completion message to the user.
[0621] Specific example
[0622] As an example, consider a scenario where a user says, "I want to order a pizza for dinner tonight." A speech recognition system converts this into text, and a natural language processing system analyzes the intent. A generative AI system generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. The user then responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[0623] This will result in a safe and easy-to-use purchasing and dining support system that is suitable even for the elderly.
[0624] The following describes the processing flow.
[0625] Step 1:
[0626] User: The user says, "I want to order a pizza for dinner tonight."
[0627] Step 2:
[0628] Terminal: The terminal receives voice input.
[0629] Step 3:
[0630] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[0631] Step 4:
[0632] Terminal: Sends the converted text data to the server.
[0633] Step 5:
[0634] Server: The server receives text data received from the terminal.
[0635] Step 6:
[0636] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0637] Step 7:
[0638] Server: The generative AI generates the response, "What kind of pizza would you like to order?"
[0639] Step 8:
[0640] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0641] Step 9:
[0642] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[0643] Step 10:
[0644] Server: Sends the generated response to the terminal.
[0645] Step 11:
[0646] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[0647] Step 12:
[0648] User: The user replies, "I want to order a Margherita pizza."
[0649] Step 13:
[0650] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0651] Step 14:
[0652] Terminal: Sends the converted text data to the server.
[0653] Step 15:
[0654] Server: Receives the converted text data.
[0655] Step 16:
[0656] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that alright?".
[0657] Step 17:
[0658] Server: Sends the generated confirmation message to the terminal.
[0659] Step 18:
[0660] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[0661] Step 19:
[0662] User: The user replies, "Yes, please place the order."
[0663] Step 20:
[0664] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0665] Step 21:
[0666] Terminal: Sends the converted text data to the server.
[0667] Step 22:
[0668] Server: Receives the converted text data.
[0669] Step 23:
[0670] Server: Sends the final order request to the delivery service API to confirm the order.
[0671] Step 24:
[0672] Server: Receives order confirmations from the delivery service API.
[0673] Step 25:
[0674] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0675] Step 26:
[0676] Server: Sends the generated completion message to the terminal.
[0677] Step 27:
[0678] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[0679] This completes the user's pizza ordering process.
[0680] (Example 1)
[0681] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0682] In modern society, it is difficult for the elderly to smoothly make online purchases or order meals. In particular, technology that enables natural dialogue is necessary for the process of extracting appropriate information from voice input and accurately confirming orders. Current systems often suffer from low accuracy in voice recognition and natural language processing, or insufficient integration with external services, making them difficult for the elderly to use. There is a need to solve these problems and provide a purchasing and dining activity support system that can be easily used by the elderly.
[0683] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0684] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; speech recognition means for re-recognizing the user's voice input and converting it into text data; and means for presenting a confirmation message to the user that is generated again based on the response from the external service API. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[0685] "Speech recognition means" refers to a device or program that receives a user's speech as audio data and converts it into text data.
[0686] "Natural language processing means" refers to a device or program that analyzes natural language text converted by speech recognition means and understands the user's intent.
[0687] A "generative AI means" is a device or program that utilizes artificial intelligence technology to generate an appropriate response based on intents analyzed by natural language processing means.
[0688] "Presentation means" refers to a device or program that presents a response generated by a generative AI means to the user visually or audibly.
[0689] "Information acquisition means" refers to a device or program for acquiring information corresponding to a user's intent from an external service API.
[0690] "External service integration means" refers to a device or program for making requests to external services in accordance with the generated response.
[0691] A "confirmation means" is a device or program that presents a generated response to the user and performs the user's final confirmation.
[0692] Modes for carrying out the invention
[0693] This invention is a system that supports elderly people in purchasing and dining activities through simple voice-based operation. Specifically, it combines voice recognition, natural language processing, and generative AI technology so that when a user gives a voice command, the system performs the appropriate processing and ultimately confirms the order corresponding to the user's instructions. A detailed explanation is as follows.
[0694] System Configuration
[0695] The system consists mainly of the following components:
[0696] 1. Speech recognition method: This is a method for converting speech input into text data. Specifically, the Google Cloud Speech-to-Text API is used.
[0697] 2. Natural Language Processing (NLP) Methods: These are methods for analyzing text data and extracting the user's intent. Specifically, natural language processing modules such as spaCy are used.
[0698] 3. Generative AI methods: These are methods for generating appropriate responses based on intents analyzed by natural language processing tools. Specifically, OpenAI GPT-3 and similar tools are used.
[0699] 4. Presentation means: These are means of presenting the generated response to the user. Specifically, this includes a speaker for reading the response aloud and a display for showing the text.
[0700] 5. Information Acquisition Method: This involves using external service APIs to obtain the necessary information. For example, the Uber Eats API is used.
[0701] 6. External service integration means: This is a means for making requests to external services according to the generated response.
[0702] 7. Verification Method: This is a means for obtaining final confirmation from the user. The speech recognition method is used again to present a response to the user and obtain final confirmation.
[0703] Example of operation
[0704] The system works as follows:
[0705] 1. Voice input:
[0706] The user says, "I want to order a pizza for dinner tonight." This audio data is captured by the device.
[0707] 2. Speech recognition:
[0708] The device uses speech recognition (Google Cloud Speech-to-Text API) to convert the voice data into text data: "I want to order pizza for dinner tonight."
[0709] 3. Natural Language Processing:
[0710] The server uses natural language processing (spaCy) to analyze the text data and identify the user's intent, "Pizza Order".
[0711] 4. Response generation:
[0712] The server uses generative AI tools (OpenAI GPT-3) to generate the response, "What kind of pizza would you like to order?"
[0713] 5. Information acquisition:
[0714] The server uses an information retrieval method (Uber Eats API) to obtain information about available pizza menus.
[0715] 6. Response presentation:
[0716] Based on the menu information acquired by the generative AI, the server generates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[0717] 7. User selection and voice recognition again:
[0718] The terminal presents a response to the user, who replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to convert this response into text.
[0719] 8. Final confirmation and order completion:
[0720] The server uses a generative AI system to generate a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[0721] The terminal displays a confirmation message to the user, who responds, "Yes, please place the order." The terminal then uses speech recognition to convert this response back into text and sends it to the server.
[0722] The server sends the final request to an external service API to confirm the order. Upon receiving the order confirmation from the Uber Eats API, the server uses generative AI tools to generate a completion message that reads, "Your order is complete. Your pizza is expected to arrive in 30 minutes."
[0723] The device displays a completion message to the user.
[0724] Example of a prompt
[0725] An example of a prompt to input into a generative AI would be: "The user says they want to order a pizza. Ask the user what kinds of pizza are available."
[0726] The above is an embodiment of the present invention. This system provides the convenience of allowing even elderly people to easily purchase items and order meals using only their voice.
[0727] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0728] Step 1: Capture voice input
[0729] User: The user says, "I want to order a pizza for dinner tonight." This voice input becomes the system's initial input.
[0730] Step 2: Speech Recognition
[0731] Terminal: The terminal receives the user's voice input and sends it to a speech recognition system (e.g., Google Cloud Speech-to-Text API). The voice data is processed by the API, and the text data "I want to order pizza for dinner tonight" is generated.
[0732] Input: User's voice data
[0733] Output: Text data "I want to order a pizza for dinner tonight"
[0734] Specific operation: The device sends voice data to the API and receives the converted text data.
[0735] Step 3: Send text data
[0736] Terminal: Sends the generated text data to the server.
[0737] Input: Text data "I want to order a pizza for dinner tonight"
[0738] Output: Sending text data to the server
[0739] Specific action: The terminal sends text data to the server as an HTTP request.
[0740] Step 4: Natural Language Processing
[0741] Server: The server uses a natural language processing module (e.g., spaCy) to parse the text data and identify that the user's intent is "pizza order".
[0742] Input: Text data "I want to order a pizza for dinner tonight"
[0743] Output: Analysis result "Pizza order"
[0744] Specific operation: The server receives text data as input, parses it using a natural language processing module, and extracts intents.
[0745] Step 5: Create a prompt for response generation
[0746] Server: Creates and sends a prompt to a generative AI (e.g., OpenAI GPT-3) that says, "The user wants to order a pizza. Ask them what kind of pizza they want to order."
[0747] Input: Intent "Pizza Order"
[0748] Output: Prompt "The user says they want to order a pizza. Ask them what kind of pizza they would like to order."
[0749] Specific operation: The server creates a prompt to send to the generative AI based on the intent.
[0750] Step 6: Generate initial response
[0751] Server: The generative AI generates a response based on the prompt, asking, "What kind of pizza would you like to order?"
[0752] Input: Prompt
[0753] Output: Response "What kind of pizza would you like to order?"
[0754] Specific operation: The generative AI analyzes the prompt and generates an appropriate response.
[0755] Step 7: Obtaining menu information
[0756] Server: The server queries an information retrieval tool (e.g., Uber Eats API) to obtain information on available pizza menus.
[0757] Input: Intent "Pizza Order"
[0758] Output: Pizza menu information
[0759] Specific operation: The server sends a request to the API and receives menu information.
[0760] Step 8: Enhancing the response content
[0761] Server: Based on the acquired menu information, the generative AI regenerates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[0762] Input: Pizza menu information
[0763] Output: Detailed response "We have Margherita pizza, pepperoni pizza, and vegetarian pizza. Which type would you like?"
[0764] Specific operation: The generative AI generates a new response based on the menu information.
[0765] Step 9: Sending and presenting a response
[0766] Server: Sends the generated detailed response to the terminal.
[0767] Terminal: The terminal either reads the response aloud to the user or displays it as text.
[0768] Input: Detailed response
[0769] Output: Presenting a response to the user
[0770] Specific operation: The server sends a detailed response to the terminal in text format, and the terminal presents it to the user in an appropriate format.
[0771] Step 10: User selection and voice recognition again
[0772] User: The user replies, "I would like to order a Margherita pizza."
[0773] Terminal: Re-recognizes the user's voice, converts it into text data, and sends it to the server.
[0774] Input: User's voice response
[0775] Output: Text data "I want to order a Margherita pizza"
[0776] Specific operation: The speech recognition system converts the user's response into text and sends it to the server.
[0777] Step 11: Generating the final confirmation message
[0778] Server: The generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that alright?"
[0779] Input: Text data "I want to order a Margherita pizza"
[0780] Output: Confirmation message
[0781] Specific operation: The generative AI generates a confirmation message based on the text it receives.
[0782] Step 12: Presentation of confirmation message
[0783] Server: Sends a confirmation message to the terminal.
[0784] Terminal: The terminal displays a confirmation message to the user.
[0785] Input: Confirmation message
[0786] Output: Presentation of a confirmation message to the user.
[0787] Specific operation: The server sends a confirmation message to the terminal, and the terminal presents it to the user via voice or text.
[0788] Step 13: User confirmation and speech recognition
[0789] User: The user replies, "Yes, please place the order."
[0790] Terminal: The speech recognition system converts the user's response back into text and sends it to the server.
[0791] Input: User's final confirmation voice
[0792] Output: Text data "Yes, please place your order"
[0793] Specific operation: The device converts the audio to text and sends it to the server.
[0794] Step 14: Confirm the order and generate the completion message.
[0795] Server: Sends the final request to an external service API (e.g., Uber Eats API) to confirm the order.
[0796] Information retrieval: Order confirmation via API
[0797] Output: Order confirmation and completion messages
[0798] Specific operation: The server sends a request to the API, and after receiving confirmation of the order, it generates the message, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0799] Step 15: Send and present completion message
[0800] Server: Sends a completion message to the terminal.
[0801] Terminal: The terminal displays a completion message to the user.
[0802] Input: Completion message
[0803] Output: Presentation of a completion message to the user.
[0804] Specific operation: The server sends a completion message to the terminal, and the terminal presents the message to the user via voice or text.
[0805] (Application Example 1)
[0806] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0807] For the elderly and those unfamiliar with technology, ordering food delivery online using smartphones or PCs presents complex, time-consuming, and stressful challenges. Furthermore, long lists and numerous options can lead to confusion when confirming specific order details and making final confirmations. This invention aims to solve these problems and provide a system that allows the elderly to easily and safely order food delivery using voice-based technology.
[0808] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0809] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; and means for ordering food and beverages based on the user's voice input, thereby presenting available menu options and confirming the order after final confirmation. This makes it possible for elderly people to easily and safely order food delivery using only their voice.
[0810] "Voice recognition means" refers to technology that converts a user's voice input into text data.
[0811] "Natural language processing means" refers to technologies that analyze converted natural language text and understand the user's intent.
[0812] "Generative AI methods" refer to artificial intelligence technologies that generate appropriate responses based on analyzed intents.
[0813] "Presentation means" refers to a technology that presents responses generated by generative AI means to the user visually or audibly.
[0814] "Information acquisition means" refers to the technology of obtaining information corresponding to a user's intent from an external service API.
[0815] "External service integration means" refers to a technology that makes requests to external services according to the generated response.
[0816] A "confirmation method" is a technology that presents the generated response to the user for final confirmation.
[0817] "Means for ordering food and beverages" refers to technology that places orders for food and beverages based on the user's voice input and confirms the order details.
[0818] This invention provides a system that allows elderly people to easily order food delivery, and is provided by combining speech recognition technology, natural language processing technology, and generative AI technology. The specific implementation of this system includes the following elements.
[0819] hardware
[0820] Smartphone: A device used by users for voice input. Typically, iOS or Android smartphones are used.
[0821] Server: A central system that performs speech recognition, natural language processing, and generative AI processing.
[0822] software
[0823] Speech recognition software:
[0824] This process converts voice input acquired from the device into text data. Specifically, it uses the Google Cloud Speech-to-Text API.
[0825] Natural language processing software:
[0826] The converted text data is analyzed to understand the user's intent. This is done using the Google Cloud Natural Language API.
[0827] Generative AI software:
[0828] The system generates a response based on the analyzed data. OpenAI's GPT-3 protocol is used.
[0829] External service API:
[0830] An API for integrating with food delivery services. For example, using the Uber Eats API or other food delivery service APIs.
[0831] Processing flow
[0832] 1. Speech recognition:
[0833] The user says "I want to order a pizza" into their smartphone. The device captures the audio and uses the Google Cloud Speech-to-Text API to convert the speech to text.
[0834] 2. Natural Language Processing:
[0835] The converted text data is sent to the server, where the Google Cloud Natural Language API is used to analyze the user's intent. This analysis identifies the user's intent to "order a pizza."
[0836] 3. Response generation and presentation:
[0837] Based on the analysis results, a generative AI (OpenAI GPT-3) generates a response asking "What kind of pizza would you like to order?", and the server sends this to the user's terminal.
[0838] 4. Integration with external services and order confirmation:
[0839] The system retrieves available pizza menus from a food delivery service API (e.g., UberEats API) and presents them to the user as options. The user selects "I want to order a Margherita pizza" and sends the selection back to the server. The server uses generative AI to generate a final confirmation message (e.g., "I would like to order one Margherita pizza. Is that correct?") and presents it to the user. Once the user confirms, the order is sent to the food delivery service API to finalize the order.
[0840] Specific example
[0841] Consider a scenario where a user says, "I want to order a pizza for dinner tonight." Speech recognition software converts this to text, and natural language processing software analyzes the intent. Generative AI software generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. Then, "Margherita pizza" is selected, and the order is confirmed after a confirmation process.
[0842] Example of a prompt
[0843] User instruction: "I want to order a pizza."
[0844] Generative AI response: "What kind of pizza would you like to order?"
[0845] User instructions: "Margherita pizza"
[0846] Final confirmation response from a generative AI: "I'd like to order one Margherita pizza. Is that alright?"
[0847] These technological elements and processing flows allow elderly people to easily order food delivery using their voice.
[0848] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0849] Step 1:
[0850] Voice input capture and conversion
[0851] The user provides voice input (e.g., "I want to order a pizza for dinner tonight"). The device captures this audio and sends it to the Google Cloud Speech-to-Text API to convert the audio data into text.
[0852] Input: User voice: "I want to order a pizza for dinner tonight."
[0853] Data processing: Audio data → Text data "I want to order a pizza for dinner tonight."
[0854] Output: Text data "I want to order a pizza for dinner tonight"
[0855] Step 2:
[0856] Text transmission and natural language processing
[0857] The device sends the converted text data to the server. The server uses the Google Cloud Natural Language API to analyze this text and identify the user's intent (e.g., ordering a pizza).
[0858] Input: Text data "I want to order a pizza for dinner tonight"
[0859] Data processing: Natural language processing → Intent "Pizza order"
[0860] Output: Intent "Pizza Order"
[0861] Step 3:
[0862] Response generation
[0863] Based on the analysis results, the server generates an appropriate response using a generative AI (OpenAI GPT-3). In this case, the response "What kind of pizza would you like to order?" is generated.
[0864] Input: Intent "Pizza Order"
[0865] Data processing: Intent → Response generation
[0866] Output: Response "What kind of pizza would you like to order?"
[0867] Step 4:
[0868] Presentation of response
[0869] The server sends the generated response to the terminal. The terminal presents the response to the user in an appropriate manner, such as text display or audio output.
[0870] Input: Response "What kind of pizza would you like to order?"
[0871] Data processing: None
[0872] Output: Displayed response, audio output
[0873] Step 5:
[0874] User selection and repeated voice recognition
[0875] The user makes a selection in response (e.g., "I want to order a Margherita pizza"). The device captures this new voice input and sends it back to the Google Cloud Speech-to-Text API, where it is converted into text data.
[0876] Input: User voice: "I want to order a Margherita pizza."
[0877] Data processing: Audio data → Text data "I want to order a Margherita pizza"
[0878] Output: Text data "I want to order a Margherita pizza"
[0879] Step 6:
[0880] Sending text and re-processing natural language
[0881] The device sends the converted text data back to the server. The server then uses the Google Cloud Natural Language API again to parse the user's intent (e.g., selecting a specific pizza type).
[0882] Input: Text data "I want to order a Margherita pizza"
[0883] Data processing: Natural language processing → Intent "Order for Margherita pizza"
[0884] Output: Intent "Order a Margherita pizza"
[0885] Step 7:
[0886] Generation of final acknowledgment
[0887] The server uses a generative AI (OpenAI GPT-3) based on the analysis results to generate a final confirmation response. In this case, the response "I would like to order one Margherita pizza. Is that alright?" is generated.
[0888] Input: Intent "Order a Margherita pizza"
[0889] Data processing: Intent → Final confirmation response generation
[0890] Output: Final confirmation response: "I would like to order one Margherita pizza. Is that correct?"
[0891] Step 8:
[0892] Final confirmation presentation and confirmation response transmission
[0893] The server sends the generated final confirmation to the terminal, which then presents it to the user. The user replies, "Yes, please place the order." The terminal captures this reply and sends it back to the Google Cloud Speech-to-Text API to convert it back into text data.
[0894] Input: Final confirmation response "I would like to order one Margherita pizza. Is that correct?"
[0895] Data processing: None
[0896] Output: User confirmation "Yes, please place the order"
[0897] Step 9:
[0898] Final order confirmation
[0899] The terminal sends final confirmation text data to the server. The server sends the final order information to an external service API (e.g., a food delivery API) to confirm the order.
[0900] Input: User confirmation "Yes, please place the order"
[0901] Data processing: Sending final order information
[0902] Output: Confirmed order "One Margherita Pizza"
[0903] Step 10:
[0904] Generating and displaying completion messages
[0905] The server receives an order confirmation from the food delivery API, generates a completion message (e.g., "Your order is complete. Your pizza is expected to arrive in 30 minutes") using OpenAI GPT-3, and sends it to the terminal. The terminal then displays this message to the user.
[0906] Input: Confirmed order information
[0907] Data processing: Order confirmation → Completion message generation
[0908] Output: Completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0909] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0910] System Overview
[0911] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, generative AI, and emotion recognition technology. When a user gives a voice command, a speech recognition means converts the voice into text, and a natural language processing means analyzes the text to identify the user's intent. A generative AI generates an appropriate response and obtains necessary information in cooperation with an external service API. Subsequently, an emotion engine analyzes the user's emotions and adjusts the response based on the results. Finally, the user is asked for confirmation and the order is completed.
[0912] Explanation of the program's processing
[0913] 1. Speech recognition
[0914] Terminal: The user says, "I want to order a pizza for dinner tonight."
[0915] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[0916] Terminal: Sends text data to the server.
[0917] 2. Natural Language Processing
[0918] Server: Receives text data sent from the terminal.
[0919] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0920] 3. Emotion recognition
[0921] Server: The emotion engine analyzes the user's emotions from text data.
[0922] Server: The emotion engine analyzes the emotion data and sends it to the generative AI.
[0923] 4. Response generation
[0924] Server: A generative AI (e.g., GPT-3) generates a response such as, "What kind of pizza would you like to order?"
[0925] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0926] Server: The generative AI generates responses that present the user with options such as "Margherita pizza, pepperoni pizza, and vegetarian pizza" based on menu information, and adjusts them according to the user's mood.
[0927] 5. Presentation and User Selection
[0928] Server: Sends the generated response to the terminal.
[0929] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[0930] User: For example, respond with, "I'd like to order a Margherita pizza."
[0931] 6. Re-speech recognition and response generation
[0932] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0933] Terminal: Sends the converted text data to the server.
[0934] Server: Upon receiving the response, the generative AI generates a confirmation message, "I would like to order one Margherita pizza. Is that alright?", and makes adjustments based on the user's emotions.
[0935] Server: Sends a confirmation message to the terminal.
[0936] 7. Final confirmation and order completion
[0937] Terminal: Displays a confirmation message to the user.
[0938] User: For example, respond with, "Yes, please place the order."
[0939] Terminal: Converts the response to text and sends it to the server.
[0940] Server: Sends the final request to the delivery service API to confirm the order.
[0941] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[0942] Server: Sends a completion message to the terminal.
[0943] Terminal: Displays a completion message to the user.
[0944] Specific example
[0945] For example, if a user says, "I want to order a pizza for dinner tonight," a speech recognition system converts this into text. A natural language processing system analyzes the intent, and an emotion engine recognizes the user's emotions. A generative AI system generates a response such as, "What kind of pizza would you like to order?" and retrieves menu information from an external service API. The user responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[0946] The emotion engine reduces the anxiety and questions users experience during the ordering process, enabling the provision of reassuring and accurate support. This will result in a purchasing and dining activity support system that is easy for even the elderly to use.
[0947] The following describes the processing flow.
[0948] Step 1:
[0949] User: The user says, "I want to order a pizza for dinner tonight."
[0950] Step 2:
[0951] Terminal: The terminal receives voice input.
[0952] Step 3:
[0953] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[0954] Step 4:
[0955] Terminal: Sends the converted text data to the server.
[0956] Step 5:
[0957] Server: The server receives text data sent from the terminal.
[0958] Step 6:
[0959] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[0960] Step 7:
[0961] Server: The emotion engine analyzes the user's emotions from text data.
[0962] Step 8:
[0963] Server: The emotion engine sends the recognized emotion data to the generative AI.
[0964] Step 9:
[0965] Server: The generative AI generates a response to the question "What kind of pizza would you like to order?", adjusting it based on the user's emotions.
[0966] Step 10:
[0967] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[0968] Step 11:
[0969] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[0970] Step 12:
[0971] Server: Sends the generated response to the terminal.
[0972] Step 13:
[0973] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[0974] Step 14:
[0975] User: The user replies, "I want to order a Margherita pizza."
[0976] Step 15:
[0977] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0978] Step 16:
[0979] Terminal: Sends the converted text data to the server.
[0980] Step 17:
[0981] Server: Receives the converted text data.
[0982] Step 18:
[0983] Server: The emotion engine analyzes the user's emotions again from the text data.
[0984] Step 19:
[0985] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that correct?".
[0986] Step 20:
[0987] Server: Sends the generated confirmation message to the terminal.
[0988] Step 20:
[0989] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[0990] Step 21:
[0991] User: The user replies, "Yes, please place the order."
[0992] Step 22:
[0993] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[0994] Step 23:
[0995] Terminal: Sends the converted text data to the server.
[0996] Step 24:
[0997] Server: Receives the converted text data.
[0998] Step 25:
[0999] Server: Sends the final order request to the delivery service API to confirm the order.
[1000] Step 26:
[1001] Server: Receives order confirmations from the delivery service API.
[1002] Step 27:
[1003] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1004] Step 28:
[1005] Server: Sends the generated completion message to the terminal.
[1006] Step 29:
[1007] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[1008] This allows the user's pizza ordering process to be completed in a way that includes emotion recognition.
[1009] (Example 2)
[1010] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1011] Today, many elderly people find online shopping and dining difficult due to the complex interfaces and operation methods. Furthermore, despite advancements in speech recognition and natural language processing technologies, comprehensive support systems, including emotion recognition, are still lacking. As a result, users often struggle to resolve anxieties and questions they encounter during the ordering process.
[1012] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1013] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making requests to external services according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; emotion recognition means for analyzing the user's emotions from the natural language text; and means for adjusting the response based on the emotion data analyzed by the emotion recognition means. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[1014] "Voice recognition means" refers to a technology or device that receives voice input from a user and converts it into text data.
[1015] "Natural language processing means" refers to a technology or device that analyzes natural language text converted by speech recognition means and understands its intent and meaning.
[1016] "Generative AI means" refers to artificial intelligence technology or systems that generate appropriate responses based on intents analyzed by natural language processing means.
[1017] "Presentation means" refers to a technology or device for showing a response generated by a generative AI means to a user.
[1018] "Information acquisition means" refers to a technology or system that obtains information corresponding to a user's intent from an external service API.
[1019] "External service integration means" refers to a technology or system that makes requests to external services according to the generated response.
[1020] A "confirmation method" is a technology or system that presents a response to the user and performs final confirmation.
[1021] "Emotion recognition means" refers to a technology or system that analyzes a user's emotions from natural language text.
[1022] "Means for adjusting responses" refers to technologies or systems for appropriately modifying responses based on emotional data analyzed by emotion recognition means.
[1023] This invention is a system that assists elderly people in easily engaging in purchasing and dining activities through voice communication. Specifically, it is realized through a system that integrates voice recognition, natural language processing, generative AI, and emotion recognition technology.
[1024] Hardware and software to be used
[1025] This system consists of the following main components:
[1026] 1. Speech recognition means:
[1027] The hardware uses a microphone, and the software uses a speech recognition system such as the Google Speech-to-Text API.
[1028] 2. Natural language processing methods:
[1029] The software used will be a natural language processing library (such as spaCy or BERT).
[1030] 3. Generative AI means:
[1031] For example, generative AI models such as OpenAI's GPT-3 can be used.
[1032] 4. Emotion recognition means:
[1033] The software used will be an emotion recognition engine such as IBM Watson Tone Analyzer.
[1034] 5. External Service APIs:
[1035] Use an API (such as the Uber Eats API) to retrieve and send user order information.
[1036] Operating principle
[1037] The operation of this system begins with the user's voice input and goes through several processing steps before finally sending the order to an external service.
[1038] Voice input:
[1039] The user speaks into the microphone and says, "I'd like to order a pizza for dinner tonight."
[1040] Speech recognition:
[1041] The device receives the audio data via its microphone and uses the Google Speech-to-Text API to generate the text data "I want to order a pizza for dinner tonight." This text data is then sent to the server.
[1042] Natural language processing:
[1043] The server analyzes the received text data using a natural language processing library to identify that the user's intent is "pizza order".
[1044] Emotion recognition:
[1045] The server uses an emotion recognition engine to analyze the user's emotions from text data. Based on the results, it generates data to adjust its response.
[1046] Response generation:
[1047] Generative AI (for example, OpenAI's GPT-3) generates a response such as, "What kind of pizza would you like to order?". In doing so, it queries a delivery service API to obtain menu information on available pizzas (for example, "Margherita pizza, pepperoni pizza, vegetarian pizza") and adjusts the response based on sentiment data.
[1048] Presentation and selection:
[1049] The server sends the generated response to the terminal, which then presents it to the user via voice or text. The user responds with something like, "I'd like to order a Margherita pizza."
[1050] Speech recognition and response generation again:
[1051] The device accepts voice input again, performs speech recognition to convert it into text data, and sends it to the server. The server understands the response, and a generative AI generates a confirmation message (for example, "I'd like to order one Margherita pizza. Is that alright?") and makes adjustments based on emotion.
[1052] Final confirmation and order completion:
[1053] The server generates a final confirmation message which is sent to the terminal and displayed to the user. The user replies, "Yes, please place the order," and sends it to the server. The server sends the final order request to the delivery service API, and the order is confirmed. Finally, the generative AI generates a completion message, "Your order is complete. Your pizza is expected to arrive within 30 minutes," and displays this to the user.
[1054] Specific example
[1055] For example, a user's voice input, such as "I want to order a pizza for dinner tonight," is received by the microphone and converted into text data using the Google Speech-to-Text API. Next, this text data is analyzed by a natural language processing library to identify the intent "order pizza." An emotion recognition engine analyzes the user's emotions, and a generative AI generates and presents the response, "What kind of pizza would you like to order?" The user replies, "I would like to order a Margherita pizza," and after confirmation, the order is finalized.
[1056] Example of a prompt:
[1057] User: I want to order a pizza for dinner tonight.
[1058] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1059] Step 1:
[1060] The user gives a voice command saying, "I want to order a pizza for dinner tonight." The device receives this voice data via its microphone. The received voice data is sent to a speech recognition system (Google Speech-to-Text API), which analyzes it and generates the text data "I want to order a pizza for dinner tonight." The device then sends the generated text data to the server.
[1061] Step 2:
[1062] The server receives the text data "I want to order pizza for dinner tonight" from the terminal. Using natural language processing tools (e.g., spaCy, BERT), the server analyzes the received text data and identifies that the user's intent is "order pizza". The result of the intent identification is stored in an internal database.
[1063] Step 3:
[1064] The server uses emotion recognition tools (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., joy, anxiety) from text data. The emotion data is stored in an internal database for transmission to generative AI tools.
[1065] Step 4:
[1066] The server invokes a generative AI (such as OpenAI's GPT-3) to generate an appropriate response, "What kind of pizza would you like to order?". During this process, the server queries a delivery service API (e.g., the Uber Eats API) to obtain information on available pizza menus. Based on this menu information, the generative AI generates a response, which is then refined based on sentiment data. The server then sends the generated response message to the device.
[1067] Step 5:
[1068] The terminal receives a response message from the server: "What kind of pizza would you like to order?" It either reads this message aloud to the user or displays it as text. The user replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to process the user's response again, converts it back into text data, "I would like to order a Margherita pizza," and sends it to the server.
[1069] Step 6:
[1070] The server receives the text data "I want to order a Margherita pizza" sent again from the terminal. The generative AI generates a confirmation message "I would like to order one Margherita pizza. Is that alright?". The generated confirmation message is then adjusted based on sentiment data and sent to the terminal.
[1071] Step 7:
[1072] The terminal displays a confirmation message received from the server, "You would like to order one Margherita pizza. Is that alright?" (either read aloud or displayed as text). The user replies, "Yes, please place the order." The terminal uses a speech recognition system to convert the user's reply into text data and sends it to the server.
[1073] Step 8:
[1074] The server receives the final text data "Yes, please place the order" and sends the final request to the delivery service API. Once the order is confirmed, it receives confirmation from the delivery service API, and the generative AI generates a completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes." The completion message is sent to the device, which then displays it to the user (either read aloud or displayed as text).
[1075] (Application Example 2)
[1076] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1077] A system is needed that allows users, including the elderly, to easily support their daily purchasing and dining activities using voice commands. In particular, a system is required that analyzes the user's emotions and provides appropriate responses based on those emotions, allowing users to use the service without feeling anxious or confused. Furthermore, a mechanism is needed that allows the elderly to complete orders and inquiries without performing complex operations.
[1078] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a speech recognition means, a natural language processing means, a response generation means, an emotion recognition means, a generative AI means, a presentation means, an information acquisition means, an external linkage means, and a confirmation means. This enables the provision of appropriate responses based on the intent and emotion of the user when they easily place an order or make an inquiry using voice input, thereby supporting purchasing and dining activities that can be used safely even by the elderly.
[1079] "Voice recognition means" refers to devices or technologies that convert a user's voice input into text data.
[1080] "Natural language processing means" refers to technology that analyzes natural language text converted by speech recognition means and understands the user's intent.
[1081] "Response generation means" refers to a technology that generates an appropriate response based on intents analyzed by natural language processing means.
[1082] "Emotion recognition means" refers to a technology that analyzes a user's emotions based on the response generated by a response generation means.
[1083] "Generative AI methods" refer to artificial intelligence technologies that adjust and generate responses based on emotional data analyzed by emotion recognition methods.
[1084] "Presentation means" refers to technologies or devices that present a user with a refined response generated by a generative AI means.
[1085] "Information acquisition means" refers to technology that acquires information corresponding to the user's intent from an external information provider.
[1086] "External collaboration means" refers to a technology that makes requests to external information providers according to the generated response.
[1087] A "confirmation method" refers to a technology or device that presents a response to the user and performs final confirmation.
[1088] In a mode for carrying out the invention, a food delivery support system is realized that allows elderly people to easily order meals by voice by using a system that combines speech recognition, natural language processing, generative AI, and emotion recognition technology.
[1089] The server includes the following hardware and software: First, it uses the speech_recognition library for speech recognition. Second, it employs the GPT-3 model from the transformers library for natural language processing. Furthermore, it uses the Hugging Face sentiment analysis model for emotion recognition. A program to integrate these elements and manage the entire system is implemented in Python.
[1090] Specifically, when a user gives a voice command such as "I want to order curry for dinner tonight," a voice recognition system converts the voice into text and sends it to the server. The server uses a natural language processing system to analyze this text data and recognize the user's intent. Next, an emotion recognition system analyzes the user's emotions and sends the result to a generative AI system. The generative AI system generates an appropriate response based on the user's emotions and issues a prompt such as "What kind of curry would you like to order?"
[1091] The server, as an external means of providing information, connects with the API of a food delivery service to retrieve selectable curry menus. Based on the retrieved menu information, a generative AI generates a list including "butter chicken curry, beef curry, vegetable curry," etc., and presents it to the user through a presentation system.
[1092] For example, if a user replies, "I'd like to order butter chicken curry," the speech recognition system converts this speech back into text and sends it to the server. The server performs natural language processing and sentiment recognition as before, and then a generative AI system generates a confirmation message, "Is butter chicken curry alright?" Once the user confirms, the order is finalized using an external integration system, and a final confirmation message is presented to the user.
[1093] Example of a prompt:
[1094] "The user's intention is to order curry, and their emotion is calmness. The available menu items are butter chicken curry, beef curry, and vegetable curry. Which curry would you like to choose?"
[1095] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1096] Step 1:
[1097] The user gives instructions for their order by voice. For example, the user might say, "I'd like to order curry for dinner tonight." The voice data becomes the input.
[1098] Step 2:
[1099] The device uses speech recognition to convert the user's voice into text data. The text output is "I want to order curry for dinner tonight." The speech recognition library is used.
[1100] Step 3:
[1101] The terminal sends the generated text data to the server. The server receives this text data.
[1102] Step 4:
[1103] The server uses natural language processing to analyze text data and recognize the user's intent. In this example, the intent is "order curry". The GPT-3 model from the transformers library is used for this process. The intent is output as the analysis result.
[1104] Step 5:
[1105] The server uses emotion recognition to analyze the user's emotions from text data. In this example, the emotion is determined to be "calm." The Hugging Face emotion analysis model is used as the emotion recognition tool. Emotion data is output as the analysis result.
[1106] Step 6:
[1107] The server uses generative AI tools to generate an appropriate response based on the user's intent and emotions. In this example, the prompt "What kind of curry would you like to order?" is generated. A generative AI model (e.g., GPT-3) is used for this process. The generated response is then output.
[1108] Step 7:
[1109] The server connects with an external information source (a food delivery service API) to retrieve information on available menu items. The retrieved information includes "butter chicken curry, beef curry, and vegetable curry." The menu information retrieved from the API is then output.
[1110] Step 8:
[1111] The server generates a more specific response based on the retrieved menu information. This response creates a prompt containing a list of options: "Butter Chicken Curry, Beef Curry, Vegetable Curry." The generated specific response is then output.
[1112] Step 9:
[1113] The server presents the response to the user through a presentation mechanism. The user responds, "I would like to order butter chicken curry." This audio data becomes the input.
[1114] Step 10:
[1115] The device uses speech recognition again to convert the user's response into text data. The text output is "I would like to order butter chicken curry."
[1116] Step 11:
[1117] The terminal sends the generated text data to the server. The server receives this text data.
[1118] Step 12:
[1119] The server again uses natural language processing and emotion recognition to analyze the user's intent and emotion. It recognizes the intent as "order butter chicken curry" and the emotion as "calm." The intent and emotion are output as analysis results.
[1120] Step 13:
[1121] The server uses a generative AI system to generate a message for final confirmation. The response "Is butter chicken curry alright?" is generated. The generated confirmation response is output.
[1122] Step 14:
[1123] The server presents an acknowledgment to the user through a display mechanism. The user responds with "Yes, please place the order." The voice data becomes the input.
[1124] Step 15:
[1125] The terminal uses voice recognition again to convert the user's confirmation response into text data. The text "Yes, please place your order" is output.
[1126] Step 16:
[1127] The server receives the generated text data and uses a generative AI to finally generate a confirmation message. A completion message such as "Your order is complete. Your pizza is expected to arrive within 30 minutes" is generated. The generated completion message is then output.
[1128] Step 17:
[1129] The server sends the user's order information to the food delivery service's API via an external connection method, and confirms the order. A confirmation message is output from the API.
[1130] Step 18:
[1131] The server displays a completion message to the user through a designated means. The user confirms that the order has been completed.
[1132] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1133] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1134] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1135] [Third Embodiment]
[1136] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1137] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1138] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1139] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1140] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1141] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1142] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1143] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1144] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1145] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1146] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1147] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1148] System Overview
[1149] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, and generative AI technologies. Specifically, when a user gives a voice command, a speech recognition means converts this command into text data, and a natural language processing means analyzes it. Then, a generative AI means generates an appropriate response and presents the response to the user while obtaining necessary information in cooperation with an external service API. Finally, after user confirmation, the request to the external service is completed.
[1150] Explanation of the program's processing
[1151] 1. Speech recognition
[1152] Terminal: The user says by voice, "I want to order a pizza for dinner tonight."
[1153] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[1154] Terminal: Sends text data to the server.
[1155] 2. Natural Language Processing
[1156] Server: Receives text data sent from the terminal.
[1157] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1158] 3. Response generation
[1159] Server: A generative AI (e.g., GPT-3) generates responses such as, "What kind of pizza would you like to order?".
[1160] Server: Query the delivery service API to retrieve information on available pizza menus.
[1161] Server: Based on this, the generative AI generates a response that presents the user with the options "Margherita pizza, pepperoni pizza, vegetarian pizza".
[1162] 4. Presentation and User Selection
[1163] Server: Sends the generated response to the terminal.
[1164] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[1165] User: For example, respond with, "I'd like to order a Margherita pizza."
[1166] 5. Re-speech recognition and response generation
[1167] Terminal: It uses speech recognition to convert the user's response into text and sends it to the server.
[1168] Server: Upon receiving the response, the generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[1169] Server: Sends a confirmation message to the terminal.
[1170] 6. Final confirmation and order completion
[1171] Terminal: Displays a confirmation message to the user.
[1172] User: For example, respond with, "Yes, please place the order."
[1173] Terminal: Converts the response to text and sends it to the server.
[1174] Server: Sends the final request to the delivery service API to confirm the order.
[1175] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1176] Server: Sends a completion message to the terminal.
[1177] Terminal: Displays a completion message to the user.
[1178] Specific example
[1179] As an example, consider a scenario where a user says, "I want to order a pizza for dinner tonight." A speech recognition system converts this into text, and a natural language processing system analyzes the intent. A generative AI system generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. The user then responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[1180] This will result in a safe and easy-to-use purchasing and dining support system that is suitable even for the elderly.
[1181] The following describes the processing flow.
[1182] Step 1:
[1183] User: The user says, "I want to order a pizza for dinner tonight."
[1184] Step 2:
[1185] Terminal: The terminal receives voice input.
[1186] Step 3:
[1187] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[1188] Step 4:
[1189] Terminal: Sends the converted text data to the server.
[1190] Step 5:
[1191] Server: The server receives text data received from the terminal.
[1192] Step 6:
[1193] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1194] Step 7:
[1195] Server: The generative AI generates the response, "What kind of pizza would you like to order?"
[1196] Step 8:
[1197] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[1198] Step 9:
[1199] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[1200] Step 10:
[1201] Server: Sends the generated response to the terminal.
[1202] Step 11:
[1203] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[1204] Step 12:
[1205] User: The user replies, "I want to order a Margherita pizza."
[1206] Step 13:
[1207] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1208] Step 14:
[1209] Terminal: Sends the converted text data to the server.
[1210] Step 15:
[1211] Server: Receives the converted text data.
[1212] Step 16:
[1213] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that alright?".
[1214] Step 17:
[1215] Server: Sends the generated confirmation message to the terminal.
[1216] Step 18:
[1217] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[1218] Step 19:
[1219] User: The user replies, "Yes, please place the order."
[1220] Step 20:
[1221] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1222] Step 21:
[1223] Terminal: Sends the converted text data to the server.
[1224] Step 22:
[1225] Server: Receives the converted text data.
[1226] Step 23:
[1227] Server: Sends the final order request to the delivery service API to confirm the order.
[1228] Step 24:
[1229] Server: Receives order confirmations from the delivery service API.
[1230] Step 25:
[1231] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1232] Step 26:
[1233] Server: Sends the generated completion message to the terminal.
[1234] Step 27:
[1235] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[1236] This completes the user's pizza ordering process.
[1237] (Example 1)
[1238] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1239] In modern society, it is difficult for the elderly to smoothly make online purchases or order meals. In particular, technology that enables natural dialogue is necessary for the process of extracting appropriate information from voice input and accurately confirming orders. Current systems often suffer from low accuracy in voice recognition and natural language processing, or insufficient integration with external services, making them difficult for the elderly to use. There is a need to solve these problems and provide a purchasing and dining activity support system that can be easily used by the elderly.
[1240] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1241] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; speech recognition means for re-recognizing the user's voice input and converting it into text data; and means for presenting a confirmation message to the user that is generated again based on the response from the external service API. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[1242] "Speech recognition means" refers to a device or program that receives a user's speech as audio data and converts it into text data.
[1243] "Natural language processing means" refers to a device or program that analyzes natural language text converted by speech recognition means and understands the user's intent.
[1244] A "generative AI means" is a device or program that utilizes artificial intelligence technology to generate an appropriate response based on intents analyzed by natural language processing means.
[1245] "Presentation means" refers to a device or program that presents a response generated by a generative AI means to the user visually or audibly.
[1246] "Information acquisition means" refers to a device or program for acquiring information corresponding to a user's intent from an external service API.
[1247] "External service integration means" refers to a device or program for making requests to external services in accordance with the generated response.
[1248] A "confirmation means" is a device or program that presents a generated response to the user and performs the user's final confirmation.
[1249] Modes for carrying out the invention
[1250] This invention is a system that supports elderly people in purchasing and dining activities through simple voice-based operation. Specifically, it combines voice recognition, natural language processing, and generative AI technology so that when a user gives a voice command, the system performs the appropriate processing and ultimately confirms the order corresponding to the user's instructions. A detailed explanation is as follows.
[1251] System Configuration
[1252] The system consists mainly of the following components:
[1253] 1. Speech recognition method: This is a method for converting speech input into text data. Specifically, the Google Cloud Speech-to-Text API is used.
[1254] 2. Natural Language Processing (NLP) Methods: These are methods for analyzing text data and extracting the user's intent. Specifically, natural language processing modules such as spaCy are used.
[1255] 3. Generative AI methods: These are methods for generating appropriate responses based on intents analyzed by natural language processing tools. Specifically, OpenAI GPT-3 and similar tools are used.
[1256] 4. Presentation means: These are means of presenting the generated response to the user. Specifically, this includes a speaker for reading the response aloud and a display for showing the text.
[1257] 5. Information Acquisition Method: This involves using external service APIs to obtain the necessary information. For example, the Uber Eats API is used.
[1258] 6. External service integration means: This is a means for making requests to external services according to the generated response.
[1259] 7. Verification Method: This is a means for obtaining final confirmation from the user. The speech recognition method is used again to present a response to the user and obtain final confirmation.
[1260] Example of operation
[1261] The system works as follows:
[1262] 1. Voice input:
[1263] The user says, "I want to order a pizza for dinner tonight." This audio data is captured by the device.
[1264] 2. Speech recognition:
[1265] The device uses speech recognition (Google Cloud Speech-to-Text API) to convert the voice data into text data: "I want to order pizza for dinner tonight."
[1266] 3. Natural Language Processing:
[1267] The server uses natural language processing (spaCy) to analyze the text data and identify the user's intent, "Pizza Order".
[1268] 4. Response generation:
[1269] The server uses generative AI tools (OpenAI GPT-3) to generate the response, "What kind of pizza would you like to order?"
[1270] 5. Information acquisition:
[1271] The server uses an information retrieval method (Uber Eats API) to obtain information about available pizza menus.
[1272] 6. Response presentation:
[1273] Based on the menu information acquired by the generative AI, the server generates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[1274] 7. User selection and voice recognition again:
[1275] The terminal presents a response to the user, who replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to convert this response into text.
[1276] 8. Final confirmation and order completion:
[1277] The server uses a generative AI system to generate a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[1278] The terminal displays a confirmation message to the user, who responds, "Yes, please place the order." The terminal then uses speech recognition to convert this response back into text and sends it to the server.
[1279] The server sends the final request to an external service API to confirm the order. Upon receiving the order confirmation from the Uber Eats API, the server uses generative AI tools to generate a completion message that reads, "Your order is complete. Your pizza is expected to arrive in 30 minutes."
[1280] The device displays a completion message to the user.
[1281] Example of a prompt
[1282] An example of a prompt to input into a generative AI would be: "The user says they want to order a pizza. Ask the user what kinds of pizza are available."
[1283] The above is an embodiment of the present invention. This system provides the convenience of allowing even elderly people to easily purchase items and order meals using only their voice.
[1284] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1285] Step 1: Capture voice input
[1286] User: The user says, "I want to order a pizza for dinner tonight." This voice input becomes the system's initial input.
[1287] Step 2: Speech Recognition
[1288] Terminal: The terminal receives the user's voice input and sends it to a speech recognition system (e.g., Google Cloud Speech-to-Text API). The voice data is processed by the API, and the text data "I want to order pizza for dinner tonight" is generated.
[1289] Input: User's voice data
[1290] Output: Text data "I want to order a pizza for dinner tonight"
[1291] Specific operation: The device sends voice data to the API and receives the converted text data.
[1292] Step 3: Send text data
[1293] Terminal: Sends the generated text data to the server.
[1294] Input: Text data "I want to order a pizza for dinner tonight"
[1295] Output: Sending text data to the server
[1296] Specific action: The terminal sends text data to the server as an HTTP request.
[1297] Step 4: Natural Language Processing
[1298] Server: The server uses a natural language processing module (e.g., spaCy) to parse the text data and identify that the user's intent is "pizza order".
[1299] Input: Text data "I want to order a pizza for dinner tonight"
[1300] Output: Analysis result "Pizza order"
[1301] Specific operation: The server receives text data as input, parses it using a natural language processing module, and extracts intents.
[1302] Step 5: Create a prompt for response generation
[1303] Server: Creates and sends a prompt to a generative AI (e.g., OpenAI GPT-3) that says, "The user wants to order a pizza. Ask them what kind of pizza they want to order."
[1304] Input: Intent "Pizza Order"
[1305] Output: Prompt "The user says they want to order a pizza. Ask them what kind of pizza they would like to order."
[1306] Specific operation: The server creates a prompt to send to the generative AI based on the intent.
[1307] Step 6: Generate initial response
[1308] Server: The generative AI generates a response based on the prompt, asking, "What kind of pizza would you like to order?"
[1309] Input: Prompt
[1310] Output: Response "What kind of pizza would you like to order?"
[1311] Specific operation: The generative AI analyzes the prompt and generates an appropriate response.
[1312] Step 7: Obtaining menu information
[1313] Server: The server queries an information retrieval tool (e.g., Uber Eats API) to obtain information on available pizza menus.
[1314] Input: Intent "Pizza Order"
[1315] Output: Pizza menu information
[1316] Specific operation: The server sends a request to the API and receives menu information.
[1317] Step 8: Enhancing the response content
[1318] Server: Based on the acquired menu information, the generative AI regenerates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[1319] Input: Pizza menu information
[1320] Output: Detailed response "We have Margherita pizza, pepperoni pizza, and vegetarian pizza. Which type would you like?"
[1321] Specific operation: The generative AI generates a new response based on the menu information.
[1322] Step 9: Sending and presenting a response
[1323] Server: Sends the generated detailed response to the terminal.
[1324] Terminal: The terminal either reads the response aloud to the user or displays it as text.
[1325] Input: Detailed response
[1326] Output: Presenting a response to the user
[1327] Specific operation: The server sends a detailed response to the terminal in text format, and the terminal presents it to the user in an appropriate format.
[1328] Step 10: User selection and voice recognition again
[1329] User: The user replies, "I would like to order a Margherita pizza."
[1330] Terminal: Re-recognizes the user's voice, converts it into text data, and sends it to the server.
[1331] Input: User's voice response
[1332] Output: Text data "I want to order a Margherita pizza"
[1333] Specific operation: The speech recognition system converts the user's response into text and sends it to the server.
[1334] Step 11: Generating the final confirmation message
[1335] Server: The generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that alright?"
[1336] Input: Text data "I want to order a Margherita pizza"
[1337] Output: Confirmation message
[1338] Specific operation: The generative AI generates a confirmation message based on the text it receives.
[1339] Step 12: Presentation of confirmation message
[1340] Server: Sends a confirmation message to the terminal.
[1341] Terminal: The terminal displays a confirmation message to the user.
[1342] Input: Confirmation message
[1343] Output: Presentation of a confirmation message to the user.
[1344] Specific operation: The server sends a confirmation message to the terminal, and the terminal presents it to the user via voice or text.
[1345] Step 13: User confirmation and speech recognition
[1346] User: The user replies, "Yes, please place the order."
[1347] Terminal: The speech recognition system converts the user's response back into text and sends it to the server.
[1348] Input: User's final confirmation voice
[1349] Output: Text data "Yes, please place your order"
[1350] Specific operation: The device converts the audio to text and sends it to the server.
[1351] Step 14: Confirm the order and generate the completion message.
[1352] Server: Sends the final request to an external service API (e.g., Uber Eats API) to confirm the order.
[1353] Information retrieval: Order confirmation via API
[1354] Output: Order confirmation and completion messages
[1355] Specific operation: The server sends a request to the API, and after receiving confirmation of the order, it generates the message, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1356] Step 15: Send and present completion message
[1357] Server: Sends a completion message to the terminal.
[1358] Terminal: The terminal displays a completion message to the user.
[1359] Input: Completion message
[1360] Output: Presentation of a completion message to the user.
[1361] Specific operation: The server sends a completion message to the terminal, and the terminal presents the message to the user via voice or text.
[1362] (Application Example 1)
[1363] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1364] For the elderly and those unfamiliar with technology, ordering food delivery online using smartphones or PCs presents complex, time-consuming, and stressful challenges. Furthermore, long lists and numerous options can lead to confusion when confirming specific order details and making final confirmations. This invention aims to solve these problems and provide a system that allows the elderly to easily and safely order food delivery using voice-based technology.
[1365] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1366] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; and means for ordering food and beverages based on the user's voice input, thereby presenting available menu options and confirming the order after final confirmation. This makes it possible for elderly people to easily and safely order food delivery using only their voice.
[1367] "Voice recognition means" refers to technology that converts a user's voice input into text data.
[1368] "Natural language processing means" refers to technologies that analyze converted natural language text and understand the user's intent.
[1369] "Generative AI methods" refer to artificial intelligence technologies that generate appropriate responses based on analyzed intents.
[1370] "Presentation means" refers to a technology that presents responses generated by generative AI means to the user visually or audibly.
[1371] "Information acquisition means" refers to the technology of obtaining information corresponding to a user's intent from an external service API.
[1372] "External service integration means" refers to a technology that makes requests to external services according to the generated response.
[1373] A "confirmation method" is a technology that presents the generated response to the user for final confirmation.
[1374] "Means for ordering food and beverages" refers to technology that places orders for food and beverages based on the user's voice input and confirms the order details.
[1375] This invention provides a system that allows elderly people to easily order food delivery, and is provided by combining speech recognition technology, natural language processing technology, and generative AI technology. The specific implementation of this system includes the following elements.
[1376] hardware
[1377] Smartphone: A device used by users for voice input. Typically, iOS or Android smartphones are used.
[1378] Server: A central system that performs speech recognition, natural language processing, and generative AI processing.
[1379] software
[1380] Speech recognition software:
[1381] This process converts voice input acquired from the device into text data. Specifically, it uses the Google Cloud Speech-to-Text API.
[1382] Natural language processing software:
[1383] The converted text data is analyzed to understand the user's intent. This is done using the Google Cloud Natural Language API.
[1384] Generative AI software:
[1385] The system generates a response based on the analyzed data. OpenAI's GPT-3 protocol is used.
[1386] External service API:
[1387] An API for integrating with food delivery services. For example, using the Uber Eats API or other food delivery service APIs.
[1388] Processing flow
[1389] 1. Speech recognition:
[1390] The user says "I want to order a pizza" into their smartphone. The device captures the audio and uses the Google Cloud Speech-to-Text API to convert the speech to text.
[1391] 2. Natural Language Processing:
[1392] The converted text data is sent to the server, where the Google Cloud Natural Language API is used to analyze the user's intent. This analysis identifies the user's intent to "order a pizza."
[1393] 3. Response generation and presentation:
[1394] Based on the analysis results, a generative AI (OpenAI GPT-3) generates a response asking "What kind of pizza would you like to order?", and the server sends this to the user's terminal.
[1395] 4. Integration with external services and order confirmation:
[1396] The system retrieves available pizza menus from a food delivery service API (e.g., UberEats API) and presents them to the user as options. The user selects "I want to order a Margherita pizza" and sends the selection back to the server. The server uses generative AI to generate a final confirmation message (e.g., "I would like to order one Margherita pizza. Is that correct?") and presents it to the user. Once the user confirms, the order is sent to the food delivery service API to finalize the order.
[1397] Specific example
[1398] Consider a scenario where a user says, "I want to order a pizza for dinner tonight." Speech recognition software converts this to text, and natural language processing software analyzes the intent. Generative AI software generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. Then, "Margherita pizza" is selected, and the order is confirmed after a confirmation process.
[1399] Example of a prompt
[1400] User instruction: "I want to order a pizza."
[1401] Generative AI response: "What kind of pizza would you like to order?"
[1402] User instructions: "Margherita pizza"
[1403] Final confirmation response from a generative AI: "I'd like to order one Margherita pizza. Is that alright?"
[1404] These technological elements and processing flows allow elderly people to easily order food delivery using their voice.
[1405] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1406] Step 1:
[1407] Voice input capture and conversion
[1408] The user provides voice input (e.g., "I want to order a pizza for dinner tonight"). The device captures this audio and sends it to the Google Cloud Speech-to-Text API to convert the audio data into text.
[1409] Input: User voice: "I want to order a pizza for dinner tonight."
[1410] Data processing: Audio data → Text data "I want to order a pizza for dinner tonight."
[1411] Output: Text data "I want to order a pizza for dinner tonight"
[1412] Step 2:
[1413] Text transmission and natural language processing
[1414] The device sends the converted text data to the server. The server uses the Google Cloud Natural Language API to analyze this text and identify the user's intent (e.g., ordering a pizza).
[1415] Input: Text data "I want to order a pizza for dinner tonight"
[1416] Data processing: Natural language processing → Intent "Pizza order"
[1417] Output: Intent "Pizza Order"
[1418] Step 3:
[1419] Response generation
[1420] Based on the analysis results, the server generates an appropriate response using a generative AI (OpenAI GPT-3). In this case, the response "What kind of pizza would you like to order?" is generated.
[1421] Input: Intent "Pizza Order"
[1422] Data processing: Intent → Response generation
[1423] Output: Response "What kind of pizza would you like to order?"
[1424] Step 4:
[1425] Presentation of response
[1426] The server sends the generated response to the terminal. The terminal presents the response to the user in an appropriate manner, such as text display or audio output.
[1427] Input: Response "What kind of pizza would you like to order?"
[1428] Data processing: None
[1429] Output: Displayed response, audio output
[1430] Step 5:
[1431] User selection and repeated voice recognition
[1432] The user makes a selection in response (e.g., "I want to order a Margherita pizza"). The device captures this new voice input and sends it back to the Google Cloud Speech-to-Text API, where it is converted into text data.
[1433] Input: User voice: "I want to order a Margherita pizza."
[1434] Data processing: Audio data → Text data "I want to order a Margherita pizza"
[1435] Output: Text data "I want to order a Margherita pizza"
[1436] Step 6:
[1437] Sending text and re-processing natural language
[1438] The device sends the converted text data back to the server. The server then uses the Google Cloud Natural Language API again to parse the user's intent (e.g., selecting a specific pizza type).
[1439] Input: Text data "I want to order a Margherita pizza"
[1440] Data processing: Natural language processing → Intent "Order for Margherita pizza"
[1441] Output: Intent "Order a Margherita pizza"
[1442] Step 7:
[1443] Generation of final acknowledgment
[1444] The server uses a generative AI (OpenAI GPT-3) based on the analysis results to generate a final confirmation response. In this case, the response "I would like to order one Margherita pizza. Is that alright?" is generated.
[1445] Input: Intent "Order a Margherita pizza"
[1446] Data processing: Intent → Final confirmation response generation
[1447] Output: Final confirmation response: "I would like to order one Margherita pizza. Is that correct?"
[1448] Step 8:
[1449] Final confirmation presentation and confirmation response transmission
[1450] The server sends the generated final confirmation to the terminal, which then presents it to the user. The user replies, "Yes, please place the order." The terminal captures this reply and sends it back to the Google Cloud Speech-to-Text API to convert it back into text data.
[1451] Input: Final confirmation response "I would like to order one Margherita pizza. Is that correct?"
[1452] Data processing: None
[1453] Output: User confirmation "Yes, please place the order"
[1454] Step 9:
[1455] Final order confirmation
[1456] The terminal sends final confirmation text data to the server. The server sends the final order information to an external service API (e.g., a food delivery API) to confirm the order.
[1457] Input: User confirmation "Yes, please place the order"
[1458] Data processing: Sending final order information
[1459] Output: Confirmed order "One Margherita Pizza"
[1460] Step 10:
[1461] Generating and displaying completion messages
[1462] The server receives an order confirmation from the food delivery API, generates a completion message (e.g., "Your order is complete. Your pizza is expected to arrive in 30 minutes") using OpenAI GPT-3, and sends it to the terminal. The terminal then displays this message to the user.
[1463] Input: Confirmed order information
[1464] Data processing: Order confirmation → Completion message generation
[1465] Output: Completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1466] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1467] System Overview
[1468] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, generative AI, and emotion recognition technology. When a user gives a voice command, a speech recognition means converts the voice into text, and a natural language processing means analyzes the text to identify the user's intent. A generative AI generates an appropriate response and obtains necessary information in cooperation with an external service API. Subsequently, an emotion engine analyzes the user's emotions and adjusts the response based on the results. Finally, the user is asked for confirmation and the order is completed.
[1469] Explanation of the program's processing
[1470] 1. Speech recognition
[1471] Terminal: The user says, "I want to order a pizza for dinner tonight."
[1472] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[1473] Terminal: Sends text data to the server.
[1474] 2. Natural Language Processing
[1475] Server: Receives text data sent from the terminal.
[1476] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1477] 3. Emotion recognition
[1478] Server: The emotion engine analyzes the user's emotions from text data.
[1479] Server: The emotion engine analyzes the emotion data and sends it to the generative AI.
[1480] 4. Response generation
[1481] Server: A generative AI (e.g., GPT-3) generates a response such as, "What kind of pizza would you like to order?"
[1482] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[1483] Server: The generative AI generates responses that present the user with options such as "Margherita pizza, pepperoni pizza, and vegetarian pizza" based on menu information, and adjusts them according to the user's mood.
[1484] 5. Presentation and User Selection
[1485] Server: Sends the generated response to the terminal.
[1486] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[1487] User: For example, respond with, "I'd like to order a Margherita pizza."
[1488] 6. Re-speech recognition and response generation
[1489] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1490] Terminal: Sends the converted text data to the server.
[1491] Server: Upon receiving the response, the generative AI generates a confirmation message, "I would like to order one Margherita pizza. Is that alright?", and makes adjustments based on the user's emotions.
[1492] Server: Sends a confirmation message to the terminal.
[1493] 7. Final confirmation and order completion
[1494] Terminal: Displays a confirmation message to the user.
[1495] User: For example, respond with, "Yes, please place the order."
[1496] Terminal: Converts the response to text and sends it to the server.
[1497] Server: Sends the final request to the delivery service API to confirm the order.
[1498] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1499] Server: Sends a completion message to the terminal.
[1500] Terminal: Displays a completion message to the user.
[1501] Specific example
[1502] For example, if a user says, "I want to order a pizza for dinner tonight," a speech recognition system converts this into text. A natural language processing system analyzes the intent, and an emotion engine recognizes the user's emotions. A generative AI system generates a response such as, "What kind of pizza would you like to order?" and retrieves menu information from an external service API. The user responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[1503] The emotion engine reduces the anxiety and questions users experience during the ordering process, enabling the provision of reassuring and accurate support. This will result in a purchasing and dining activity support system that is easy for even the elderly to use.
[1504] The following describes the processing flow.
[1505] Step 1:
[1506] User: The user says, "I want to order a pizza for dinner tonight."
[1507] Step 2:
[1508] Terminal: The terminal receives voice input.
[1509] Step 3:
[1510] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[1511] Step 4:
[1512] Terminal: Sends the converted text data to the server.
[1513] Step 5:
[1514] Server: The server receives text data sent from the terminal.
[1515] Step 6:
[1516] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1517] Step 7:
[1518] Server: The emotion engine analyzes the user's emotions from text data.
[1519] Step 8:
[1520] Server: The emotion engine sends the recognized emotion data to the generative AI.
[1521] Step 9:
[1522] Server: The generative AI generates a response to the question "What kind of pizza would you like to order?", adjusting it based on the user's emotions.
[1523] Step 10:
[1524] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[1525] Step 11:
[1526] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[1527] Step 12:
[1528] Server: Sends the generated response to the terminal.
[1529] Step 13:
[1530] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[1531] Step 14:
[1532] User: The user replies, "I want to order a Margherita pizza."
[1533] Step 15:
[1534] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1535] Step 16:
[1536] Terminal: Sends the converted text data to the server.
[1537] Step 17:
[1538] Server: Receives the converted text data.
[1539] Step 18:
[1540] Server: The emotion engine analyzes the user's emotions again from the text data.
[1541] Step 19:
[1542] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that correct?".
[1543] Step 20:
[1544] Server: Sends the generated confirmation message to the terminal.
[1545] Step 20:
[1546] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[1547] Step 21:
[1548] User: The user replies, "Yes, please place the order."
[1549] Step 22:
[1550] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1551] Step 23:
[1552] Terminal: Sends the converted text data to the server.
[1553] Step 24:
[1554] Server: Receives the converted text data.
[1555] Step 25:
[1556] Server: Sends the final order request to the delivery service API to confirm the order.
[1557] Step 26:
[1558] Server: Receives order confirmations from the delivery service API.
[1559] Step 27:
[1560] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1561] Step 28:
[1562] Server: Sends the generated completion message to the terminal.
[1563] Step 29:
[1564] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[1565] This allows the user's pizza ordering process to be completed in a way that includes emotion recognition.
[1566] (Example 2)
[1567] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1568] Today, many elderly people find online shopping and dining difficult due to the complex interfaces and operation methods. Furthermore, despite advancements in speech recognition and natural language processing technologies, comprehensive support systems, including emotion recognition, are still lacking. As a result, users often struggle to resolve anxieties and questions they encounter during the ordering process.
[1569] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1570] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making requests to external services according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; emotion recognition means for analyzing the user's emotions from the natural language text; and means for adjusting the response based on the emotion data analyzed by the emotion recognition means. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[1571] "Voice recognition means" refers to a technology or device that receives voice input from a user and converts it into text data.
[1572] "Natural language processing means" refers to a technology or device that analyzes natural language text converted by speech recognition means and understands its intent and meaning.
[1573] "Generative AI means" refers to artificial intelligence technology or systems that generate appropriate responses based on intents analyzed by natural language processing means.
[1574] "Presentation means" refers to a technology or device for showing a response generated by a generative AI means to a user.
[1575] "Information acquisition means" refers to a technology or system that obtains information corresponding to a user's intent from an external service API.
[1576] "External service integration means" refers to a technology or system that makes requests to external services according to the generated response.
[1577] A "confirmation method" is a technology or system that presents a response to the user and performs final confirmation.
[1578] "Emotion recognition means" refers to a technology or system that analyzes a user's emotions from natural language text.
[1579] "Means for adjusting responses" refers to technologies or systems for appropriately modifying responses based on emotional data analyzed by emotion recognition means.
[1580] This invention is a system that assists elderly people in easily engaging in purchasing and dining activities through voice communication. Specifically, it is realized through a system that integrates voice recognition, natural language processing, generative AI, and emotion recognition technology.
[1581] Hardware and software to be used
[1582] This system consists of the following main components:
[1583] 1. Speech recognition means:
[1584] The hardware uses a microphone, and the software uses a speech recognition system such as the Google Speech-to-Text API.
[1585] 2. Natural language processing methods:
[1586] The software used will be a natural language processing library (such as spaCy or BERT).
[1587] 3. Generative AI means:
[1588] For example, generative AI models such as OpenAI's GPT-3 can be used.
[1589] 4. Emotion recognition means:
[1590] The software used will be an emotion recognition engine such as IBM Watson Tone Analyzer.
[1591] 5. External Service APIs:
[1592] Use an API (such as the Uber Eats API) to retrieve and send user order information.
[1593] Operating principle
[1594] The operation of this system begins with the user's voice input and goes through several processing steps before finally sending the order to an external service.
[1595] Voice input:
[1596] The user speaks into the microphone and says, "I'd like to order a pizza for dinner tonight."
[1597] Speech recognition:
[1598] The device receives the audio data via its microphone and uses the Google Speech-to-Text API to generate the text data "I want to order a pizza for dinner tonight." This text data is then sent to the server.
[1599] Natural language processing:
[1600] The server analyzes the received text data using a natural language processing library to identify that the user's intent is "pizza order".
[1601] Emotion recognition:
[1602] The server uses an emotion recognition engine to analyze the user's emotions from text data. Based on the results, it generates data to adjust its response.
[1603] Response generation:
[1604] Generative AI (for example, OpenAI's GPT-3) generates a response such as, "What kind of pizza would you like to order?". In doing so, it queries a delivery service API to obtain menu information on available pizzas (for example, "Margherita pizza, pepperoni pizza, vegetarian pizza") and adjusts the response based on sentiment data.
[1605] Presentation and selection:
[1606] The server sends the generated response to the terminal, which then presents it to the user via voice or text. The user responds with something like, "I'd like to order a Margherita pizza."
[1607] Speech recognition and response generation again:
[1608] The device accepts voice input again, performs speech recognition to convert it into text data, and sends it to the server. The server understands the response, and a generative AI generates a confirmation message (for example, "I'd like to order one Margherita pizza. Is that alright?") and makes adjustments based on emotion.
[1609] Final confirmation and order completion:
[1610] The server generates a final confirmation message which is sent to the terminal and displayed to the user. The user replies, "Yes, please place the order," and sends it to the server. The server sends the final order request to the delivery service API, and the order is confirmed. Finally, the generative AI generates a completion message, "Your order is complete. Your pizza is expected to arrive within 30 minutes," and displays this to the user.
[1611] Specific example
[1612] For example, a user's voice input, such as "I want to order a pizza for dinner tonight," is received by the microphone and converted into text data using the Google Speech-to-Text API. Next, this text data is analyzed by a natural language processing library to identify the intent "order pizza." An emotion recognition engine analyzes the user's emotions, and a generative AI generates and presents the response, "What kind of pizza would you like to order?" The user replies, "I would like to order a Margherita pizza," and after confirmation, the order is finalized.
[1613] Example of a prompt:
[1614] User: I want to order a pizza for dinner tonight.
[1615] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1616] Step 1:
[1617] The user gives a voice command saying, "I want to order a pizza for dinner tonight." The device receives this voice data via its microphone. The received voice data is sent to a speech recognition system (Google Speech-to-Text API), which analyzes it and generates the text data "I want to order a pizza for dinner tonight." The device then sends the generated text data to the server.
[1618] Step 2:
[1619] The server receives the text data "I want to order pizza for dinner tonight" from the terminal. Using natural language processing tools (e.g., spaCy, BERT), the server analyzes the received text data and identifies that the user's intent is "order pizza". The result of the intent identification is stored in an internal database.
[1620] Step 3:
[1621] The server uses emotion recognition tools (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., joy, anxiety) from text data. The emotion data is stored in an internal database for transmission to generative AI tools.
[1622] Step 4:
[1623] The server invokes a generative AI (such as OpenAI's GPT-3) to generate an appropriate response, "What kind of pizza would you like to order?". During this process, the server queries a delivery service API (e.g., the Uber Eats API) to obtain information on available pizza menus. Based on this menu information, the generative AI generates a response, which is then refined based on sentiment data. The server then sends the generated response message to the device.
[1624] Step 5:
[1625] The terminal receives a response message from the server: "What kind of pizza would you like to order?" It either reads this message aloud to the user or displays it as text. The user replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to process the user's response again, converts it back into text data, "I would like to order a Margherita pizza," and sends it to the server.
[1626] Step 6:
[1627] The server receives the text data "I want to order a Margherita pizza" sent again from the terminal. The generative AI generates a confirmation message "I would like to order one Margherita pizza. Is that alright?". The generated confirmation message is then adjusted based on sentiment data and sent to the terminal.
[1628] Step 7:
[1629] The terminal displays a confirmation message received from the server, "You would like to order one Margherita pizza. Is that alright?" (either read aloud or displayed as text). The user replies, "Yes, please place the order." The terminal uses a speech recognition system to convert the user's reply into text data and sends it to the server.
[1630] Step 8:
[1631] The server receives the final text data "Yes, please place the order" and sends the final request to the delivery service API. Once the order is confirmed, it receives confirmation from the delivery service API, and the generative AI generates a completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes." The completion message is sent to the device, which then displays it to the user (either read aloud or displayed as text).
[1632] (Application Example 2)
[1633] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1634] A system is needed that allows users, including the elderly, to easily support their daily purchasing and dining activities using voice commands. In particular, a system is required that analyzes the user's emotions and provides appropriate responses based on those emotions, allowing users to use the service without feeling anxious or confused. Furthermore, a mechanism is needed that allows the elderly to complete orders and inquiries without performing complex operations.
[1635] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a speech recognition means, a natural language processing means, a response generation means, an emotion recognition means, a generative AI means, a presentation means, an information acquisition means, an external linkage means, and a confirmation means. This enables the provision of appropriate responses based on the intent and emotion of the user when they easily place an order or make an inquiry using voice input, thereby supporting purchasing and dining activities that can be used safely even by the elderly.
[1636] "Voice recognition means" refers to devices or technologies that convert a user's voice input into text data.
[1637] "Natural language processing means" refers to technology that analyzes natural language text converted by speech recognition means and understands the user's intent.
[1638] "Response generation means" refers to a technology that generates an appropriate response based on intents analyzed by natural language processing means.
[1639] "Emotion recognition means" refers to a technology that analyzes a user's emotions based on the response generated by a response generation means.
[1640] "Generative AI methods" refer to artificial intelligence technologies that adjust and generate responses based on emotional data analyzed by emotion recognition methods.
[1641] "Presentation means" refers to technologies or devices that present a user with a refined response generated by a generative AI means.
[1642] "Information acquisition means" refers to technology that acquires information corresponding to the user's intent from an external information provider.
[1643] "External collaboration means" refers to a technology that makes requests to external information providers according to the generated response.
[1644] A "confirmation method" refers to a technology or device that presents a response to the user and performs final confirmation.
[1645] In a mode for carrying out the invention, a food delivery support system is realized that allows elderly people to easily order meals by voice by using a system that combines speech recognition, natural language processing, generative AI, and emotion recognition technology.
[1646] The server includes the following hardware and software: First, it uses the speech_recognition library for speech recognition. Second, it employs the GPT-3 model from the transformers library for natural language processing. Furthermore, it uses the Hugging Face sentiment analysis model for emotion recognition. A program to integrate these elements and manage the entire system is implemented in Python.
[1647] Specifically, when a user gives a voice command such as "I want to order curry for dinner tonight," a voice recognition system converts the voice into text and sends it to the server. The server uses a natural language processing system to analyze this text data and recognize the user's intent. Next, an emotion recognition system analyzes the user's emotions and sends the result to a generative AI system. The generative AI system generates an appropriate response based on the user's emotions and issues a prompt such as "What kind of curry would you like to order?"
[1648] The server, as an external means of providing information, connects with the API of a food delivery service to retrieve selectable curry menus. Based on the retrieved menu information, a generative AI generates a list including "butter chicken curry, beef curry, vegetable curry," etc., and presents it to the user through a presentation system.
[1649] For example, if a user replies, "I'd like to order butter chicken curry," the speech recognition system converts this speech back into text and sends it to the server. The server performs natural language processing and sentiment recognition as before, and then a generative AI system generates a confirmation message, "Is butter chicken curry alright?" Once the user confirms, the order is finalized using an external integration system, and a final confirmation message is presented to the user.
[1650] Example of a prompt:
[1651] "The user's intention is to order curry, and their emotion is calmness. The available menu items are butter chicken curry, beef curry, and vegetable curry. Which curry would you like to choose?"
[1652] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1653] Step 1:
[1654] The user gives instructions for their order by voice. For example, the user might say, "I'd like to order curry for dinner tonight." The voice data becomes the input.
[1655] Step 2:
[1656] The device uses speech recognition to convert the user's voice into text data. The text output is "I want to order curry for dinner tonight." The speech recognition library is used.
[1657] Step 3:
[1658] The terminal sends the generated text data to the server. The server receives this text data.
[1659] Step 4:
[1660] The server uses natural language processing to analyze text data and recognize the user's intent. In this example, the intent is "order curry". The GPT-3 model from the transformers library is used for this process. The intent is output as the analysis result.
[1661] Step 5:
[1662] The server uses emotion recognition to analyze the user's emotions from text data. In this example, the emotion is determined to be "calm." The Hugging Face emotion analysis model is used as the emotion recognition tool. Emotion data is output as the analysis result.
[1663] Step 6:
[1664] The server uses generative AI tools to generate an appropriate response based on the user's intent and emotions. In this example, the prompt "What kind of curry would you like to order?" is generated. A generative AI model (e.g., GPT-3) is used for this process. The generated response is then output.
[1665] Step 7:
[1666] The server connects with an external information source (a food delivery service API) to retrieve information on available menu items. The retrieved information includes "butter chicken curry, beef curry, and vegetable curry." The menu information retrieved from the API is then output.
[1667] Step 8:
[1668] The server generates a more specific response based on the retrieved menu information. This response creates a prompt containing a list of options: "Butter Chicken Curry, Beef Curry, Vegetable Curry." The generated specific response is then output.
[1669] Step 9:
[1670] The server presents the response to the user through a presentation mechanism. The user responds, "I would like to order butter chicken curry." This audio data becomes the input.
[1671] Step 10:
[1672] The device uses speech recognition again to convert the user's response into text data. The text output is "I would like to order butter chicken curry."
[1673] Step 11:
[1674] The terminal sends the generated text data to the server. The server receives this text data.
[1675] Step 12:
[1676] The server again uses natural language processing and emotion recognition to analyze the user's intent and emotion. It recognizes the intent as "order butter chicken curry" and the emotion as "calm." The intent and emotion are output as analysis results.
[1677] Step 13:
[1678] The server uses a generative AI system to generate a message for final confirmation. The response "Is butter chicken curry alright?" is generated. The generated confirmation response is output.
[1679] Step 14:
[1680] The server presents an acknowledgment to the user through a display mechanism. The user responds with "Yes, please place the order." The voice data becomes the input.
[1681] Step 15:
[1682] The terminal uses voice recognition again to convert the user's confirmation response into text data. The text "Yes, please place your order" is output.
[1683] Step 16:
[1684] The server receives the generated text data and uses a generative AI to finally generate a confirmation message. A completion message such as "Your order is complete. Your pizza is expected to arrive within 30 minutes" is generated. The generated completion message is then output.
[1685] Step 17:
[1686] The server sends the user's order information to the food delivery service's API via an external connection method, and confirms the order. A confirmation message is output from the API.
[1687] Step 18:
[1688] The server displays a completion message to the user through a designated means. The user confirms that the order has been completed.
[1689] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1690] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1691] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1692] [Fourth Embodiment]
[1693] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1694] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1695] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1696] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1697] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1698] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1699] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1700] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1701] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1702] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1703] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1704] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1705] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1706] System Overview
[1707] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, and generative AI technologies. Specifically, when a user gives a voice command, a speech recognition means converts this command into text data, and a natural language processing means analyzes it. Then, a generative AI means generates an appropriate response and presents the response to the user while obtaining necessary information in cooperation with an external service API. Finally, after user confirmation, the request to the external service is completed.
[1708] Explanation of the program's processing
[1709] 1. Speech recognition
[1710] Terminal: The user says by voice, "I want to order a pizza for dinner tonight."
[1711] Terminal: The voice input is sent to the speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[1712] Terminal: Sends text data to the server.
[1713] 2. Natural Language Processing
[1714] Server: Receives text data sent from the terminal.
[1715] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1716] 3. Response generation
[1717] Server: A generative AI (e.g., GPT-3) generates responses such as, "What kind of pizza would you like to order?".
[1718] Server: Query the delivery service API to retrieve information on available pizza menus.
[1719] Server: Based on this, the generative AI generates a response that presents the user with the options "Margherita pizza, pepperoni pizza, vegetarian pizza".
[1720] 4. Presentation and User Selection
[1721] Server: Sends the generated response to the terminal.
[1722] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[1723] User: For example, respond with, "I'd like to order a Margherita pizza."
[1724] 5. Re-speech recognition and response generation
[1725] Terminal: It uses speech recognition to convert the user's response into text and sends it to the server.
[1726] Server: Upon receiving the response, the generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[1727] Server: Sends a confirmation message to the terminal.
[1728] 6. Final confirmation and order completion
[1729] Terminal: Displays a confirmation message to the user.
[1730] User: For example, respond with, "Yes, please place the order."
[1731] Terminal: Converts the response to text and sends it to the server.
[1732] Server: Sends the final request to the delivery service API to confirm the order.
[1733] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1734] Server: Sends a completion message to the terminal.
[1735] Terminal: Displays a completion message to the user.
[1736] Specific example
[1737] As an example, consider a scenario where a user says, "I want to order a pizza for dinner tonight." A speech recognition system converts this into text, and a natural language processing system analyzes the intent. A generative AI system generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. The user then responds, "I want to order a Margherita pizza," and the order is finally confirmed after going through a confirmation process.
[1738] This will result in a safe and easy-to-use purchasing and dining support system that is suitable even for the elderly.
[1739] The following describes the processing flow.
[1740] Step 1:
[1741] User: The user says, "I want to order a pizza for dinner tonight."
[1742] Step 2:
[1743] Terminal: The terminal receives voice input.
[1744] Step 3:
[1745] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[1746] Step 4:
[1747] Terminal: Sends the converted text data to the server.
[1748] Step 5:
[1749] Server: The server receives text data received from the terminal.
[1750] Step 6:
[1751] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[1752] Step 7:
[1753] Server: The generative AI generates the response, "What kind of pizza would you like to order?"
[1754] Step 8:
[1755] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[1756] Step 9:
[1757] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[1758] Step 10:
[1759] Server: Sends the generated response to the terminal.
[1760] Step 11:
[1761] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[1762] Step 12:
[1763] User: The user replies, "I want to order a Margherita pizza."
[1764] Step 13:
[1765] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1766] Step 14:
[1767] Terminal: Sends the converted text data to the server.
[1768] Step 15:
[1769] Server: Receives the converted text data.
[1770] Step 16:
[1771] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that alright?".
[1772] Step 17:
[1773] Server: Sends the generated confirmation message to the terminal.
[1774] Step 18:
[1775] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[1776] Step 19:
[1777] User: The user replies, "Yes, please place the order."
[1778] Step 20:
[1779] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[1780] Step 21:
[1781] Terminal: Sends the converted text data to the server.
[1782] Step 22:
[1783] Server: Receives the converted text data.
[1784] Step 23:
[1785] Server: Sends the final order request to the delivery service API to confirm the order.
[1786] Step 24:
[1787] Server: Receives order confirmations from the delivery service API.
[1788] Step 25:
[1789] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1790] Step 26:
[1791] Server: Sends the generated completion message to the terminal.
[1792] Step 27:
[1793] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[1794] This completes the user's pizza ordering process.
[1795] (Example 1)
[1796] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1797] In modern society, it is difficult for the elderly to smoothly make online purchases or order meals. In particular, technology that enables natural dialogue is necessary for the process of extracting appropriate information from voice input and accurately confirming orders. Current systems often suffer from low accuracy in voice recognition and natural language processing, or insufficient integration with external services, making them difficult for the elderly to use. There is a need to solve these problems and provide a purchasing and dining activity support system that can be easily used by the elderly.
[1798] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1799] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; speech recognition means for re-recognizing the user's voice input and converting it into text data; and means for presenting a confirmation message to the user that is generated again based on the response from the external service API. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[1800] "Speech recognition means" refers to a device or program that receives a user's speech as audio data and converts it into text data.
[1801] "Natural language processing means" refers to a device or program that analyzes natural language text converted by speech recognition means and understands the user's intent.
[1802] A "generative AI means" is a device or program that utilizes artificial intelligence technology to generate an appropriate response based on intents analyzed by natural language processing means.
[1803] "Presentation means" refers to a device or program that presents a response generated by a generative AI means to the user visually or audibly.
[1804] "Information acquisition means" refers to a device or program for acquiring information corresponding to a user's intent from an external service API.
[1805] "External service integration means" refers to a device or program for making requests to external services in accordance with the generated response.
[1806] A "confirmation means" is a device or program that presents a generated response to the user and performs the user's final confirmation.
[1807] Modes for carrying out the invention
[1808] This invention is a system that supports elderly people in purchasing and dining activities through simple voice-based operation. Specifically, it combines voice recognition, natural language processing, and generative AI technology so that when a user gives a voice command, the system performs the appropriate processing and ultimately confirms the order corresponding to the user's instructions. A detailed explanation is as follows.
[1809] System Configuration
[1810] The system consists mainly of the following components:
[1811] 1. Speech recognition method: This is a method for converting speech input into text data. Specifically, the Google Cloud Speech-to-Text API is used.
[1812] 2. Natural Language Processing (NLP) Methods: These are methods for analyzing text data and extracting the user's intent. Specifically, natural language processing modules such as spaCy are used.
[1813] 3. Generative AI methods: These are methods for generating appropriate responses based on intents analyzed by natural language processing tools. Specifically, OpenAI GPT-3 and similar tools are used.
[1814] 4. Presentation means: These are means of presenting the generated response to the user. Specifically, this includes a speaker for reading the response aloud and a display for showing the text.
[1815] 5. Information Acquisition Method: This involves using external service APIs to obtain the necessary information. For example, the Uber Eats API is used.
[1816] 6. External service integration means: This is a means for making requests to external services according to the generated response.
[1817] 7. Verification Method: This is a means for obtaining final confirmation from the user. The speech recognition method is used again to present a response to the user and obtain final confirmation.
[1818] Example of operation
[1819] The system works as follows:
[1820] 1. Voice input:
[1821] The user says, "I want to order a pizza for dinner tonight." This audio data is captured by the device.
[1822] 2. Speech recognition:
[1823] The device uses speech recognition (Google Cloud Speech-to-Text API) to convert the voice data into text data: "I want to order pizza for dinner tonight."
[1824] 3. Natural Language Processing:
[1825] The server uses natural language processing (spaCy) to analyze the text data and identify the user's intent, "Pizza Order".
[1826] 4. Response generation:
[1827] The server uses generative AI tools (OpenAI GPT-3) to generate the response, "What kind of pizza would you like to order?"
[1828] 5. Information acquisition:
[1829] The server uses an information retrieval method (Uber Eats API) to obtain information about available pizza menus.
[1830] 6. Response presentation:
[1831] Based on the menu information acquired by the generative AI, the server generates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[1832] 7. User selection and voice recognition again:
[1833] The terminal presents a response to the user, who replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to convert this response into text.
[1834] 8. Final confirmation and order completion:
[1835] The server uses a generative AI system to generate a confirmation message: "I would like to order one Margherita pizza. Is that correct?"
[1836] The terminal displays a confirmation message to the user, who responds, "Yes, please place the order." The terminal then uses speech recognition to convert this response back into text and sends it to the server.
[1837] The server sends the final request to an external service API to confirm the order. Upon receiving the order confirmation from the Uber Eats API, the server uses generative AI tools to generate a completion message that reads, "Your order is complete. Your pizza is expected to arrive in 30 minutes."
[1838] The device displays a completion message to the user.
[1839] Example of a prompt
[1840] An example of a prompt to input into a generative AI would be: "The user says they want to order a pizza. Ask the user what kinds of pizza are available."
[1841] The above is an embodiment of the present invention. This system provides the convenience of allowing even elderly people to easily purchase items and order meals using only their voice.
[1842] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1843] Step 1: Capture voice input
[1844] User: The user says, "I want to order a pizza for dinner tonight." This voice input becomes the system's initial input.
[1845] Step 2: Speech Recognition
[1846] Terminal: The terminal receives the user's voice input and sends it to a speech recognition system (e.g., Google Cloud Speech-to-Text API). The voice data is processed by the API, and the text data "I want to order pizza for dinner tonight" is generated.
[1847] Input: User's voice data
[1848] Output: Text data "I want to order a pizza for dinner tonight"
[1849] Specific operation: The device sends voice data to the API and receives the converted text data.
[1850] Step 3: Send text data
[1851] Terminal: Sends the generated text data to the server.
[1852] Input: Text data "I want to order a pizza for dinner tonight"
[1853] Output: Sending text data to the server
[1854] Specific action: The terminal sends text data to the server as an HTTP request.
[1855] Step 4: Natural Language Processing
[1856] Server: The server uses a natural language processing module (e.g., spaCy) to parse the text data and identify that the user's intent is "pizza order".
[1857] Input: Text data "I want to order a pizza for dinner tonight"
[1858] Output: Analysis result "Pizza order"
[1859] Specific operation: The server receives text data as input, parses it using a natural language processing module, and extracts intents.
[1860] Step 5: Create a prompt for response generation
[1861] Server: Creates and sends a prompt to a generative AI (e.g., OpenAI GPT-3) that says, "The user wants to order a pizza. Ask them what kind of pizza they want to order."
[1862] Input: Intent "Pizza Order"
[1863] Output: Prompt "The user says they want to order a pizza. Ask them what kind of pizza they would like to order."
[1864] Specific operation: The server creates a prompt to send to the generative AI based on the intent.
[1865] Step 6: Generate initial response
[1866] Server: The generative AI generates a response based on the prompt, asking, "What kind of pizza would you like to order?"
[1867] Input: Prompt
[1868] Output: Response "What kind of pizza would you like to order?"
[1869] Specific operation: The generative AI analyzes the prompt and generates an appropriate response.
[1870] Step 7: Obtaining menu information
[1871] Server: The server queries an information retrieval tool (e.g., Uber Eats API) to obtain information on available pizza menus.
[1872] Input: Intent "Pizza Order"
[1873] Output: Pizza menu information
[1874] Specific operation: The server sends a request to the API and receives menu information.
[1875] Step 8: Enhancing the response content
[1876] Server: Based on the acquired menu information, the generative AI regenerates a response that includes options such as "Margherita pizza, pepperoni pizza, vegetarian pizza."
[1877] Input: Pizza menu information
[1878] Output: Detailed response "We have Margherita pizza, pepperoni pizza, and vegetarian pizza. Which type would you like?"
[1879] Specific operation: The generative AI generates a new response based on the menu information.
[1880] Step 9: Sending and presenting a response
[1881] Server: Sends the generated detailed response to the terminal.
[1882] Terminal: The terminal either reads the response aloud to the user or displays it as text.
[1883] Input: Detailed response
[1884] Output: Presenting a response to the user
[1885] Specific operation: The server sends a detailed response to the terminal in text format, and the terminal presents it to the user in an appropriate format.
[1886] Step 10: User selection and voice recognition again
[1887] User: The user replies, "I would like to order a Margherita pizza."
[1888] Terminal: Re-recognizes the user's voice, converts it into text data, and sends it to the server.
[1889] Input: User's voice response
[1890] Output: Text data "I want to order a Margherita pizza"
[1891] Specific operation: The speech recognition system converts the user's response into text and sends it to the server.
[1892] Step 11: Generating the final confirmation message
[1893] Server: The generation AI generates a confirmation message: "I would like to order one Margherita pizza. Is that alright?"
[1894] Input: Text data "I want to order a Margherita pizza"
[1895] Output: Confirmation message
[1896] Specific operation: The generative AI generates a confirmation message based on the text it receives.
[1897] Step 12: Presentation of confirmation message
[1898] Server: Sends a confirmation message to the terminal.
[1899] Terminal: The terminal displays a confirmation message to the user.
[1900] Input: Confirmation message
[1901] Output: Presentation of a confirmation message to the user.
[1902] Specific operation: The server sends a confirmation message to the terminal, and the terminal presents it to the user via voice or text.
[1903] Step 13: User confirmation and speech recognition
[1904] User: The user replies, "Yes, please place the order."
[1905] Terminal: The speech recognition system converts the user's response back into text and sends it to the server.
[1906] Input: User's final confirmation voice
[1907] Output: Text data "Yes, please place your order"
[1908] Specific operation: The device converts the audio to text and sends it to the server.
[1909] Step 14: Confirm the order and generate the completion message.
[1910] Server: Sends the final request to an external service API (e.g., Uber Eats API) to confirm the order.
[1911] Information retrieval: Order confirmation via API
[1912] Output: Order confirmation and completion messages
[1913] Specific operation: The server sends a request to the API, and after receiving confirmation of the order, it generates the message, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[1914] Step 15: Send and present completion message
[1915] Server: Sends a completion message to the terminal.
[1916] Terminal: The terminal displays a completion message to the user.
[1917] Input: Completion message
[1918] Output: Presentation of a completion message to the user.
[1919] Specific operation: The server sends a completion message to the terminal, and the terminal presents the message to the user via voice or text.
[1920] (Application Example 1)
[1921] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1922] For the elderly and those unfamiliar with technology, ordering food delivery online using smartphones or PCs presents complex, time-consuming, and stressful challenges. Furthermore, long lists and numerous options can lead to confusion when confirming specific order details and making final confirmations. This invention aims to solve these problems and provide a system that allows the elderly to easily and safely order food delivery using voice-based technology.
[1923] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1924] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making a request to an external service according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; and means for ordering food and beverages based on the user's voice input, thereby presenting available menu options and confirming the order after final confirmation. This makes it possible for elderly people to easily and safely order food delivery using only their voice.
[1925] "Voice recognition means" refers to technology that converts a user's voice input into text data.
[1926] "Natural language processing means" refers to technologies that analyze converted natural language text and understand the user's intent.
[1927] "Generative AI methods" refer to artificial intelligence technologies that generate appropriate responses based on analyzed intents.
[1928] "Presentation means" refers to a technology that presents responses generated by generative AI means to the user visually or audibly.
[1929] "Information acquisition means" refers to the technology of obtaining information corresponding to a user's intent from an external service API.
[1930] "External service integration means" refers to a technology that makes requests to external services according to the generated response.
[1931] A "confirmation method" is a technology that presents the generated response to the user for final confirmation.
[1932] "Means for ordering food and beverages" refers to technology that places orders for food and beverages based on the user's voice input and confirms the order details.
[1933] This invention provides a system that allows elderly people to easily order food delivery, and is provided by combining speech recognition technology, natural language processing technology, and generative AI technology. The specific implementation of this system includes the following elements.
[1934] hardware
[1935] Smartphone: A device used by users for voice input. Typically, iOS or Android smartphones are used.
[1936] Server: A central system that performs speech recognition, natural language processing, and generative AI processing.
[1937] software
[1938] Speech recognition software:
[1939] This process converts voice input acquired from the device into text data. Specifically, it uses the Google Cloud Speech-to-Text API.
[1940] Natural language processing software:
[1941] The converted text data is analyzed to understand the user's intent. This is done using the Google Cloud Natural Language API.
[1942] Generative AI software:
[1943] The system generates a response based on the analyzed data. OpenAI's GPT-3 protocol is used.
[1944] External service API:
[1945] An API for integrating with food delivery services. For example, using the Uber Eats API or other food delivery service APIs.
[1946] Processing flow
[1947] 1. Speech recognition:
[1948] The user says "I want to order a pizza" into their smartphone. The device captures the audio and uses the Google Cloud Speech-to-Text API to convert the speech to text.
[1949] 2. Natural Language Processing:
[1950] The converted text data is sent to the server, where the Google Cloud Natural Language API is used to analyze the user's intent. This analysis identifies the user's intent to "order a pizza."
[1951] 3. Response generation and presentation:
[1952] Based on the analysis results, a generative AI (OpenAI GPT-3) generates a response asking "What kind of pizza would you like to order?", and the server sends this to the user's terminal.
[1953] 4. Integration with external services and order confirmation:
[1954] The system retrieves available pizza menus from a food delivery service API (e.g., UberEats API) and presents them to the user as options. The user selects "I want to order a Margherita pizza" and sends the selection back to the server. The server uses generative AI to generate a final confirmation message (e.g., "I would like to order one Margherita pizza. Is that correct?") and presents it to the user. Once the user confirms, the order is sent to the food delivery service API to finalize the order.
[1955] Specific example
[1956] Consider a scenario where a user says, "I want to order a pizza for dinner tonight." Speech recognition software converts this to text, and natural language processing software analyzes the intent. Generative AI software generates a response, "What kind of pizza would you like to order?", and retrieves menu information from an external service API. Then, "Margherita pizza" is selected, and the order is confirmed after a confirmation process.
[1957] Example of a prompt
[1958] User instruction: "I want to order a pizza."
[1959] Generative AI response: "What kind of pizza would you like to order?"
[1960] User instructions: "Margherita pizza"
[1961] Final confirmation response from a generative AI: "I'd like to order one Margherita pizza. Is that alright?"
[1962] These technological elements and processing flows allow elderly people to easily order food delivery using their voice.
[1963] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1964] Step 1:
[1965] Voice input capture and conversion
[1966] The user provides voice input (e.g., "I want to order a pizza for dinner tonight"). The device captures this audio and sends it to the Google Cloud Speech-to-Text API to convert the audio data into text.
[1967] Input: User voice: "I want to order a pizza for dinner tonight."
[1968] Data processing: Audio data → Text data "I want to order a pizza for dinner tonight."
[1969] Output: Text data "I want to order a pizza for dinner tonight"
[1970] Step 2:
[1971] Text transmission and natural language processing
[1972] The device sends the converted text data to the server. The server uses the Google Cloud Natural Language API to analyze this text and identify the user's intent (e.g., ordering a pizza).
[1973] Input: Text data "I want to order a pizza for dinner tonight"
[1974] Data processing: Natural language processing → Intent "Pizza order"
[1975] Output: Intent "Pizza Order"
[1976] Step 3:
[1977] Response generation
[1978] Based on the analysis results, the server generates an appropriate response using a generative AI (OpenAI GPT-3). In this case, the response "What kind of pizza would you like to order?" is generated.
[1979] Input: Intent "Pizza Order"
[1980] Data processing: Intent → Response generation
[1981] Output: Response "What kind of pizza would you like to order?"
[1982] Step 4:
[1983] Presentation of response
[1984] The server sends the generated response to the terminal. The terminal presents the response to the user in an appropriate manner, such as text display or audio output.
[1985] Input: Response "What kind of pizza would you like to order?"
[1986] Data processing: None
[1987] Output: Displayed response, audio output
[1988] Step 5:
[1989] User selection and repeated voice recognition
[1990] The user makes a selection in response (e.g., "I want to order a Margherita pizza"). The device captures this new voice input and sends it back to the Google Cloud Speech-to-Text API, where it is converted into text data.
[1991] Input: User voice: "I want to order a Margherita pizza."
[1992] Data processing: Audio data → Text data "I want to order a Margherita pizza"
[1993] Output: Text data "I want to order a Margherita pizza"
[1994] Step 6:
[1995] Sending text and re-processing natural language
[1996] The device sends the converted text data back to the server. The server then uses the Google Cloud Natural Language API again to parse the user's intent (e.g., selecting a specific pizza type).
[1997] Input: Text data "I want to order a Margherita pizza"
[1998] Data processing: Natural language processing → Intent "Order for Margherita pizza"
[1999] Output: Intent "Order a Margherita pizza"
[2000] Step 7:
[2001] Generation of final acknowledgment
[2002] The server uses a generative AI (OpenAI GPT-3) based on the analysis results to generate a final confirmation response. In this case, the response "I would like to order one Margherita pizza. Is that alright?" is generated.
[2003] Input: Intent "Order a Margherita pizza"
[2004] Data processing: Intent → Final confirmation response generation
[2005] Output: Final confirmation response: "I would like to order one Margherita pizza. Is that correct?"
[2006] Step 8:
[2007] Final confirmation presentation and confirmation response transmission
[2008] The server sends the generated final confirmation to the terminal, which then presents it to the user. The user replies, "Yes, please place the order." The terminal captures this reply and sends it back to the Google Cloud Speech-to-Text API to convert it back into text data.
[2009] Input: Final confirmation response "I would like to order one Margherita pizza. Is that correct?"
[2010] Data processing: None
[2011] Output: User confirmation "Yes, please place the order"
[2012] Step 9:
[2013] Final order confirmation
[2014] The terminal sends final confirmation text data to the server. The server sends the final order information to an external service API (e.g., a food delivery API) to confirm the order.
[2015] Input: User confirmation "Yes, please place the order"
[2016] Data processing: Sending final order information
[2017] Output: Confirmed order "One Margherita Pizza"
[2018] Step 10:
[2019] Generating and displaying completion messages
[2020] The server receives an order confirmation from the food delivery API, generates a completion message (e.g., "Your order is complete. Your pizza is expected to arrive in 30 minutes") using OpenAI GPT-3, and sends it to the terminal. The terminal then displays this message to the user.
[2021] Input: Confirmed order information
[2022] Data processing: Order confirmation → Completion message generation
[2023] Output: Completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[2024] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[2025] System Overview
[2026] This invention provides a purchasing and dining activity support system that is easy for the elderly to use by combining speech recognition, natural language processing, generative AI, and emotion recognition technology. When a user gives a voice command, a speech recognition means converts the voice into text, and a natural language processing means analyzes the text to identify the user's intent. A generative AI generates an appropriate response and obtains necessary information in cooperation with an external service API. Subsequently, an emotion engine analyzes the user's emotions and adjusts the response based on the results. Finally, the user is asked for confirmation and the order is completed.
[2027] Explanation of the program's processing
[2028] 1. Speech recognition
[2029] Terminal: The user says, "I want to order a pizza for dinner tonight."
[2030] Terminal: The voice input is sent to a speech recognition system, which analyzes it and generates text data: "I want to order pizza for dinner tonight."
[2031] Terminal: Sends text data to the server.
[2032] 2. Natural Language Processing
[2033] Server: Receives text data sent from the terminal.
[2034] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[2035] 3. Emotion recognition
[2036] Server: The emotion engine analyzes the user's emotions from text data.
[2037] Server: The emotion engine analyzes the emotion data and sends it to the generative AI.
[2038] 4. Response generation
[2039] Server: A generative AI (e.g., GPT-3) generates a response such as, "What kind of pizza would you like to order?"
[2040] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[2041] Server: The generative AI generates responses that present the user with options such as "Margherita pizza, pepperoni pizza, and vegetarian pizza" based on menu information, and adjusts them according to the user's mood.
[2042] 5. Presentation and User Selection
[2043] Server: Sends the generated response to the terminal.
[2044] Terminal: Reads aloud or displays as text the question, "What kind of pizza would you like to order?" to the user.
[2045] User: For example, respond with, "I'd like to order a Margherita pizza."
[2046] 6. Re-speech recognition and response generation
[2047] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[2048] Terminal: Sends the converted text data to the server.
[2049] Server: Upon receiving the response, the generative AI generates a confirmation message, "I would like to order one Margherita pizza. Is that alright?", and makes adjustments based on the user's emotions.
[2050] Server: Sends a confirmation message to the terminal.
[2051] 7. Final confirmation and order completion
[2052] Terminal: Displays a confirmation message to the user.
[2053] User: For example, respond with, "Yes, please place the order."
[2054] Terminal: Converts the response to text and sends it to the server.
[2055] Server: Sends the final request to the delivery service API to confirm the order.
[2056] Server: Upon receiving confirmation from the delivery service API, the generative AI generates a completion message stating, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[2057] Server: Sends a completion message to the terminal.
[2058] Terminal: Displays a completion message to the user.
[2059] Specific example
[2060] For example, if a user says, "I want to order a pizza for dinner tonight," a speech recognition system converts this into text. A natural language processing system analyzes the intent, and an emotion engine recognizes the user's emotions. A generative AI system generates a response such as, "What kind of pizza would you like to order?" and retrieves menu information from an external service API. The user responds, "I want to order a Margherita pizza," and after confirmation, the order is finally finalized.
[2061] The emotion engine reduces the anxiety and questions users experience during the ordering process, enabling the provision of reassuring and accurate support. This will result in a purchasing and dining activity support system that is easy for even the elderly to use.
[2062] The following describes the processing flow.
[2063] Step 1:
[2064] User: The user says, "I want to order a pizza for dinner tonight."
[2065] Step 2:
[2066] Terminal: The terminal receives voice input.
[2067] Step 3:
[2068] Terminal: The voice recognition system converts the voice input into text data: "I want to order pizza for dinner tonight."
[2069] Step 4:
[2070] Terminal: Sends the converted text data to the server.
[2071] Step 5:
[2072] Server: The server receives text data sent from the terminal.
[2073] Step 6:
[2074] Server: The natural language processing module analyzes the text data and identifies that the user's intent is "pizza order".
[2075] Step 7:
[2076] Server: The emotion engine analyzes the user's emotions from text data.
[2077] Step 8:
[2078] Server: The emotion engine sends the recognized emotion data to the generative AI.
[2079] Step 9:
[2080] Server: The generative AI generates a response to the question "What kind of pizza would you like to order?", adjusting it based on the user's emotions.
[2081] Step 10:
[2082] Server: Query the delivery service API to retrieve menu information for pizzas that can be ordered.
[2083] Step 11:
[2084] Server: The generation AI generates a response based on the menu information, saying, "Please choose from Margherita pizza, pepperoni pizza, or vegetarian pizza."
[2085] Step 12:
[2086] Server: Sends the generated response to the terminal.
[2087] Step 13:
[2088] Terminal: Reads suggestions received from the server aloud to the user or displays them as text.
[2089] Step 14:
[2090] User: The user replies, "I want to order a Margherita pizza."
[2091] Step 15:
[2092] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[2093] Step 16:
[2094] Terminal: Sends the converted text data to the server.
[2095] Step 17:
[2096] Server: Receives the converted text data.
[2097] Step 18:
[2098] Server: The emotion engine analyzes the user's emotions again from the text data.
[2099] Step 19:
[2100] Server: The generation AI generates the confirmation message "I would like to order one Margherita pizza. Is that correct?".
[2101] Step 20:
[2102] Server: Sends the generated confirmation message to the terminal.
[2103] Step 20:
[2104] Terminal: Reads the confirmation message received from the server aloud to the user, or displays it as text.
[2105] Step 21:
[2106] User: The user replies, "Yes, please place the order."
[2107] Step 22:
[2108] Terminal: Receives voice input, and the speech recognition system converts the speech into text.
[2109] Step 23:
[2110] Terminal: Sends the converted text data to the server.
[2111] Step 24:
[2112] Server: Receives the converted text data.
[2113] Step 25:
[2114] Server: Sends the final order request to the delivery service API to confirm the order.
[2115] Step 26:
[2116] Server: Receives order confirmations from the delivery service API.
[2117] Step 27:
[2118] Server: The generation AI generates a completion message saying, "Your order is complete. Your pizza is expected to arrive within 30 minutes."
[2119] Step 28:
[2120] Server: Sends the generated completion message to the terminal.
[2121] Step 29:
[2122] Terminal: Reads the completion message received from the server aloud to the user, or displays it as text.
[2123] This allows the user's pizza ordering process to be completed in a way that includes emotion recognition.
[2124] (Example 2)
[2125] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2126] Today, many elderly people find online shopping and dining difficult due to the complex interfaces and operation methods. Furthermore, despite advancements in speech recognition and natural language processing technologies, comprehensive support systems, including emotion recognition, are still lacking. As a result, users often struggle to resolve anxieties and questions they encounter during the ordering process.
[2127] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[2128] In this invention, the server includes: speech recognition means for receiving voice input from a user; natural language processing means for analyzing natural language text converted by the speech recognition means; generative AI means for generating an appropriate response based on the intent analyzed by the natural language processing means; presentation means for presenting the response generated by the generative AI means to the user; information acquisition means for obtaining information corresponding to the user's intent from an external service API; generative AI means for generating another response based on the information acquired by the information acquisition means; external service linkage means for making requests to external services according to the generated response; confirmation means for presenting the response to the user and performing final confirmation; emotion recognition means for analyzing the user's emotions from the natural language text; and means for adjusting the response based on the emotion data analyzed by the emotion recognition means. This makes it possible to support purchasing and dining activities that can be easily used even by the elderly.
[2129] "Voice recognition means" refers to a technology or device that receives voice input from a user and converts it into text data.
[2130] "Natural language processing means" refers to a technology or device that analyzes natural language text converted by speech recognition means and understands its intent and meaning.
[2131] "Generative AI means" refers to artificial intelligence technology or systems that generate appropriate responses based on intents analyzed by natural language processing means.
[2132] "Presentation means" refers to a technology or device for showing a response generated by a generative AI means to a user.
[2133] "Information acquisition means" refers to a technology or system that obtains information corresponding to a user's intent from an external service API.
[2134] "External service integration means" refers to a technology or system that makes requests to external services according to the generated response.
[2135] A "confirmation method" is a technology or system that presents a response to the user and performs final confirmation.
[2136] "Emotion recognition means" refers to a technology or system that analyzes a user's emotions from natural language text.
[2137] "Means for adjusting responses" refers to technologies or systems for appropriately modifying responses based on emotional data analyzed by emotion recognition means.
[2138] This invention is a system that assists elderly people in easily engaging in purchasing and dining activities through voice communication. Specifically, it is realized through a system that integrates voice recognition, natural language processing, generative AI, and emotion recognition technology.
[2139] Hardware and software to be used
[2140] This system consists of the following main components:
[2141] 1. Speech recognition means:
[2142] The hardware uses a microphone, and the software uses a speech recognition system such as the Google Speech-to-Text API.
[2143] 2. Natural language processing methods:
[2144] The software used will be a natural language processing library (such as spaCy or BERT).
[2145] 3. Generative AI means:
[2146] For example, generative AI models such as OpenAI's GPT-3 can be used.
[2147] 4. Emotion recognition means:
[2148] The software used will be an emotion recognition engine such as IBM Watson Tone Analyzer.
[2149] 5. External Service APIs:
[2150] Use an API (such as the Uber Eats API) to retrieve and send user order information.
[2151] Operating principle
[2152] The operation of this system begins with the user's voice input and goes through several processing steps before finally sending the order to an external service.
[2153] Voice input:
[2154] The user speaks into the microphone and says, "I'd like to order a pizza for dinner tonight."
[2155] Speech recognition:
[2156] The device receives the audio data via its microphone and uses the Google Speech-to-Text API to generate the text data "I want to order a pizza for dinner tonight." This text data is then sent to the server.
[2157] Natural language processing:
[2158] The server analyzes the received text data using a natural language processing library to identify that the user's intent is "pizza order".
[2159] Emotion recognition:
[2160] The server uses an emotion recognition engine to analyze the user's emotions from text data. Based on the results, it generates data to adjust its response.
[2161] Response generation:
[2162] Generative AI (for example, OpenAI's GPT-3) generates a response such as, "What kind of pizza would you like to order?". In doing so, it queries a delivery service API to obtain menu information on available pizzas (for example, "Margherita pizza, pepperoni pizza, vegetarian pizza") and adjusts the response based on sentiment data.
[2163] Presentation and selection:
[2164] The server sends the generated response to the terminal, which then presents it to the user via voice or text. The user responds with something like, "I'd like to order a Margherita pizza."
[2165] Speech recognition and response generation again:
[2166] The device accepts voice input again, performs speech recognition to convert it into text data, and sends it to the server. The server understands the response, and a generative AI generates a confirmation message (for example, "I'd like to order one Margherita pizza. Is that alright?") and makes adjustments based on emotion.
[2167] Final confirmation and order completion:
[2168] The server generates a final confirmation message which is sent to the terminal and displayed to the user. The user replies, "Yes, please place the order," and sends it to the server. The server sends the final order request to the delivery service API, and the order is confirmed. Finally, the generative AI generates a completion message, "Your order is complete. Your pizza is expected to arrive within 30 minutes," and displays this to the user.
[2169] Specific example
[2170] For example, a user's voice input, such as "I want to order a pizza for dinner tonight," is received by the microphone and converted into text data using the Google Speech-to-Text API. Next, this text data is analyzed by a natural language processing library to identify the intent "order pizza." An emotion recognition engine analyzes the user's emotions, and a generative AI generates and presents the response, "What kind of pizza would you like to order?" The user replies, "I would like to order a Margherita pizza," and after confirmation, the order is finalized.
[2171] Example of a prompt:
[2172] User: I want to order a pizza for dinner tonight.
[2173] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2174] Step 1:
[2175] The user gives a voice command saying, "I want to order a pizza for dinner tonight." The device receives this voice data via its microphone. The received voice data is sent to a speech recognition system (Google Speech-to-Text API), which analyzes it and generates the text data "I want to order a pizza for dinner tonight." The device then sends the generated text data to the server.
[2176] Step 2:
[2177] The server receives the text data "I want to order pizza for dinner tonight" from the terminal. Using natural language processing tools (e.g., spaCy, BERT), the server analyzes the received text data and identifies that the user's intent is "order pizza". The result of the intent identification is stored in an internal database.
[2178] Step 3:
[2179] The server uses emotion recognition tools (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., joy, anxiety) from text data. The emotion data is stored in an internal database for transmission to generative AI tools.
[2180] Step 4:
[2181] The server invokes a generative AI (such as OpenAI's GPT-3) to generate an appropriate response, "What kind of pizza would you like to order?". During this process, the server queries a delivery service API (e.g., the Uber Eats API) to obtain information on available pizza menus. Based on this menu information, the generative AI generates a response, which is then refined based on sentiment data. The server then sends the generated response message to the device.
[2182] Step 5:
[2183] The terminal receives a response message from the server: "What kind of pizza would you like to order?" It either reads this message aloud to the user or displays it as text. The user replies, "I would like to order a Margherita pizza." The terminal then uses speech recognition to process the user's response again, converts it back into text data, "I would like to order a Margherita pizza," and sends it to the server.
[2184] Step 6:
[2185] The server receives the text data "I want to order a Margherita pizza" sent again from the terminal. The generative AI generates a confirmation message "I would like to order one Margherita pizza. Is that alright?". The generated confirmation message is then adjusted based on sentiment data and sent to the terminal.
[2186] Step 7:
[2187] The terminal displays a confirmation message received from the server, "You would like to order one Margherita pizza. Is that alright?" (either read aloud or displayed as text). The user replies, "Yes, please place the order." The terminal uses a speech recognition system to convert the user's reply into text data and sends it to the server.
[2188] Step 8:
[2189] The server receives the final text data "Yes, please place the order" and sends the final request to the delivery service API. Once the order is confirmed, it receives confirmation from the delivery service API, and the generative AI generates a completion message "Your order is complete. Your pizza is expected to arrive within 30 minutes." The completion message is sent to the device, which then displays it to the user (either read aloud or displayed as text).
[2190] (Application Example 2)
[2191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2192] A system is needed that allows users, including the elderly, to easily support their daily purchasing and dining activities using voice commands. In particular, a system is required that analyzes the user's emotions and provides appropriate responses based on those emotions, allowing users to use the service without feeling anxious or confused. Furthermore, a mechanism is needed that allows the elderly to complete orders and inquiries without performing complex operations.
[2193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a speech recognition means, a natural language processing means, a response generation means, an emotion recognition means, a generative AI means, a presentation means, an information acquisition means, an external linkage means, and a confirmation means. This enables the provision of appropriate responses based on the intent and emotion of the user when they easily place an order or make an inquiry using voice input, thereby supporting purchasing and dining activities that can be used safely even by the elderly.
[2194] "Voice recognition means" refers to devices or technologies that convert a user's voice input into text data.
[2195] "Natural language processing means" refers to technology that analyzes natural language text converted by speech recognition means and understands the user's intent.
[2196] "Response generation means" refers to a technology that generates an appropriate response based on intents analyzed by natural language processing means.
[2197] "Emotion recognition means" refers to a technology that analyzes a user's emotions based on the response generated by a response generation means.
[2198] "Generative AI methods" refer to artificial intelligence technologies that adjust and generate responses based on emotional data analyzed by emotion recognition methods.
[2199] "Presentation means" refers to technologies or devices that present a user with a refined response generated by a generative AI means.
[2200] "Information acquisition means" refers to technology that acquires information corresponding to the user's intent from an external information provider.
[2201] "External collaboration means" refers to a technology that makes requests to external information providers according to the generated response.
[2202] A "confirmation method" refers to a technology or device that presents a response to the user and performs final confirmation.
[2203] In a mode for carrying out the invention, a food delivery support system is realized that allows elderly people to easily order meals by voice by using a system that combines speech recognition, natural language processing, generative AI, and emotion recognition technology.
[2204] The server includes the following hardware and software: First, it uses the speech_recognition library for speech recognition. Second, it employs the GPT-3 model from the transformers library for natural language processing. Furthermore, it uses the Hugging Face sentiment analysis model for emotion recognition. A program to integrate these elements and manage the entire system is implemented in Python.
[2205] Specifically, when a user gives a voice command such as "I want to order curry for dinner tonight," a voice recognition system converts the voice into text and sends it to the server. The server uses a natural language processing system to analyze this text data and recognize the user's intent. Next, an emotion recognition system analyzes the user's emotions and sends the result to a generative AI system. The generative AI system generates an appropriate response based on the user's emotions and issues a prompt such as "What kind of curry would you like to order?"
[2206] The server, as an external means of providing information, connects with the API of a food delivery service to retrieve selectable curry menus. Based on the retrieved menu information, a generative AI generates a list including "butter chicken curry, beef curry, vegetable curry," etc., and presents it to the user through a presentation system.
[2207] For example, if a user replies, "I'd like to order butter chicken curry," the speech recognition system converts this speech back into text and sends it to the server. The server performs natural language processing and sentiment recognition as before, and then a generative AI system generates a confirmation message, "Is butter chicken curry alright?" Once the user confirms, the order is finalized using an external integration system, and a final confirmation message is presented to the user.
[2208] Example of a prompt:
[2209] "The user's intention is to order curry, and their emotion is calmness. The available menu items are butter chicken curry, beef curry, and vegetable curry. Which curry would you like to choose?"
[2210] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2211] Step 1:
[2212] The user gives instructions for their order by voice. For example, the user might say, "I'd like to order curry for dinner tonight." The voice data becomes the input.
[2213] Step 2:
[2214] The device uses speech recognition to convert the user's voice into text data. The text output is "I want to order curry for dinner tonight." The speech recognition library is used.
[2215] Step 3:
[2216] The terminal sends the generated text data to the server. The server receives this text data.
[2217] Step 4:
[2218] The server uses natural language processing to analyze text data and recognize the user's intent. In this example, the intent is "order curry". The GPT-3 model from the transformers library is used for this process. The intent is output as the analysis result.
[2219] Step 5:
[2220] The server uses emotion recognition to analyze the user's emotions from text data. In this example, the emotion is determined to be "calm." The Hugging Face emotion analysis model is used as the emotion recognition tool. Emotion data is output as the analysis result.
[2221] Step 6:
[2222] The server uses generative AI tools to generate an appropriate response based on the user's intent and emotions. In this example, the prompt "What kind of curry would you like to order?" is generated. A generative AI model (e.g., GPT-3) is used for this process. The generated response is then output.
[2223] Step 7:
[2224] The server connects with an external information source (a food delivery service API) to retrieve information on available menu items. The retrieved information includes "butter chicken curry, beef curry, and vegetable curry." The menu information retrieved from the API is then output.
[2225] Step 8:
[2226] The server generates a more specific response based on the retrieved menu information. This response creates a prompt containing a list of options: "Butter Chicken Curry, Beef Curry, Vegetable Curry." The generated specific response is then output.
[2227] Step 9:
[2228] The server presents the response to the user through a presentation mechanism. The user responds, "I would like to order butter chicken curry." This audio data becomes the input.
[2229] Step 10:
[2230] The device uses speech recognition again to convert the user's response into text data. The text output is "I would like to order butter chicken curry."
[2231] Step 11:
[2232] The terminal sends the generated text data to the server. The server receives this text data.
[2233] Step 12:
[2234] The server again uses natural language processing and emotion recognition to analyze the user's intent and emotion. It recognizes the intent as "order butter chicken curry" and the emotion as "calm." The intent and emotion are output as analysis results.
[2235] Step 13:
[2236] The server uses a generative AI system to generate a message for final confirmation. The response "Is butter chicken curry alright?" is generated. The generated confirmation response is output.
[2237] Step 14:
[2238] The server presents an acknowledgment to the user through a display mechanism. The user responds with "Yes, please place the order." The voice data becomes the input.
[2239] Step 15:
[2240] The terminal uses voice recognition again to convert the user's confirmation response into text data. The text "Yes, please place your order" is output.
[2241] Step 16:
[2242] The server receives the generated text data and uses a generative AI to finally generate a confirmation message. A completion message such as "Your order is complete. Your pizza is expected to arrive within 30 minutes" is generated. The generated completion message is then output.
[2243] Step 17:
[2244] The server sends the user's order information to the food delivery service's API via an external connection method, and confirms the order. A confirmation message is output from the API.
[2245] Step 18:
[2246] The server displays a completion message to the user through a designated means. The user confirms that the order has been completed.
[2247] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2248] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2249] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2250] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2251] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2252] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2253] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2254] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2255] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2256] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2257] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2258] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2259] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[2260] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2261] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[2262] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[2263] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[2264] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[2265] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[2266] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[2267] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[2268] The following is further disclosed regarding the embodiments described above.
[2269] (Claim 1)
[2270] A speech recognition means that receives voice input from the user,
[2271] A natural language processing means for analyzing the natural language text converted by the speech recognition means,
[2272] A generative AI means that generates an appropriate response based on the intent analyzed by the aforementioned natural language processing means,
[2273] A presentation means for presenting the response generated by the aforementioned generation AI means to the user,
[2274] Information acquisition means for obtaining information corresponding to the user's intent from an external servic...
Claims
1. A speech recognition means that receives voice input from the user, A natural language processing means for analyzing the natural language text converted by the speech recognition means, A generative AI means that generates an appropriate response based on the intent analyzed by the aforementioned natural language processing means, A presentation means for presenting the response generated by the aforementioned generation AI means to the user, Information acquisition means for obtaining information corresponding to the user's intent from an external service API, A generation AI means that generates a response again based on the information acquired by the aforementioned information acquisition means, An external service integration means that makes a request to an external service in accordance with the generated response, A confirmation means that presents the aforementioned response to the user and performs final confirmation, A system that includes this.
2. The system according to claim 1, wherein the speech recognition means converts the user's voice into text data and transmits the text data to a server.
3. The system according to claim 1, wherein the external service linkage means includes means for transmitting user order information to an external purchasing system or delivery service.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A