System

The system addresses inefficiencies in restaurant robot systems by using voice input, text conversion, and generative AI for flexible order handling and service optimization, improving user experience.

JP2026035247APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138090
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing robot systems in restaurants lack flexibility and natural communication skills to respond to diverse customer needs and uncertain situations, leading to inefficient processes such as order taking, serving food, and handling additional orders, which reduces user experience.

Method used

A system that includes voice input, text conversion, order analysis, server communication, robot control, and voice response using generative artificial intelligence to enable flexible and efficient order acceptance, serving, and additional order processing.

Benefits of technology

The system provides efficient and flexible services by automating user recognition, order taking, serving, and payment processing, optimizing restaurant operations and enhancing user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035247000001_ABST
    Figure 2026035247000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for inputting user's voice, a means for converting the voice into text, a means for analyzing the text to specify order contents, a means for transmitting order information to a server, a means for controlling a robot by receiving an instruction from the server, and a means for responding to the user by voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Existing robot systems lack the flexibility and natural communication skills to quickly respond to diverse customer needs and uncertain situations in restaurants. Furthermore, processes such as order taking, serving food, and handling additional orders are inefficient, preventing an improved user experience. The present invention aims to solve these problems and provide a system that comprehensively optimizes service in restaurants. [Means for solving the problem]

[0005] The present invention is a system that includes a means for inputting a user's voice, a means for converting the voice into text, a means for analyzing the text to identify the order contents, a means for sending the order information to a server, a means for controlling a robot by receiving instructions from the server, and a means for responding to the user by voice. This allows for flexible response to diverse user needs and realizes efficient order acceptance, serving, and processing of additional orders. The system also uses generative artificial intelligence to enable flexible responses to additional orders and questions from users. This improves the user experience and achieves the overall efficiency and optimization of services at restaurants.

[0006] "User" refers to the customer or client who uses the system, and is the entity that interacts with the system through speech and actions.

[0007] "Voice input means" refers to devices such as microphones and sensors that recognize the user's speech and input it into the system, as well as accompanying software.

[0008] A "voice recognition module" refers to an algorithm or program that analyzes input voice and converts it into text data.

[0009] "Generative AI" is an artificial intelligence system that analyzes text data converted by a voice recognition module and generates appropriate responses and processing.

[0010] The "order identification means" refers to an analysis algorithm or program for extracting specific order details from the user's speech and executing the related procedures.

[0011] A "server" is a central control device that manages and controls data and processing within a system via a network, and is a device that communicates with other terminals to exchange necessary information.

[0012] "Robot control means" refers to a software and hardware system that receives instructions from the server, controls the robot's operations, and provides services to users.

[0013] "Voice response means" refers to a speaker or other audio output device used to convey the voice response created by the generation AI to the user.

[0014] "Serving" refers to the task of bringing cooked food to the user's table.

[0015] "Flexible response" refers to the ability to respond quickly and appropriately to diverse user requests and situations. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a system for providing efficient and flexible service using robots in restaurants. Specific embodiments for carrying out the invention and the processing of the program therefor will be described below.

[0038] User recognition and conversation initiation

[0039] User entry recognition

[0040] The device uses a camera and microphone to recognize the user. When a user enters a store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech. Through facial recognition algorithms and voice input, it is possible to quickly respond to users even when meeting them for the first time.

[0041] Initial voice recognition and greeting

[0042] When a user says "hello," the device sends the speech to a speech recognition module, which converts the speech to text. The generative AI analyzes the text and generates an appropriate greeting. The device responds through the speaker, saying, "Hello, welcome. We'll show you to your seat."

[0043] Receiving orders

[0044] Voice recognition for ordering

[0045] If a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which the generation AI then analyzes to determine the order.

[0046] Communicating with the Server

[0047] The terminal sends the specified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin promptly.

[0048] Acknowledgement and Response

[0049] The AI ​​generates a message to confirm the order, and the device responds aloud, saying, "Understood. A hamburger and orange juice, please."

[0050] Cooking status management and serving

[0051] Cooking completion notification

[0052] The server receives a notification from the kitchen that the food is ready and sends that information to the terminal, allowing the robot to begin serving the food.

[0053] Serving instructions and actions

[0054] The server sends serving instructions to the terminal, which then follows those instructions to move to the kitchen and place the food on a tray. The robot then safely carries the food to the table and serves it to the user, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[0055] Continuing the conversation and being flexible

[0056] Accepting additional orders

[0057] If the user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. The device responds, "Okay, I'll order some extra fries," and sends the additional order information to the server.

[0058] Inquiry response

[0059] When a user asks, "What desserts do you have?", the device converts the voice into text, and the AI ​​analyzes the query and generates an appropriate answer. The device responds aloud, "Today's desserts include cake and ice cream."

[0060] settlement

[0061] Acceptance of settlement requests

[0062] When a user says, "I'd like to pay, please," the device converts the speech into text using a voice recognition module, and the generation AI identifies the payment request and notifies the server.

[0063] Settlement process

[0064] The server compiles all the order details, calculates the payment amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and says, "Please check on this device."

[0065] Check and complete the payment

[0066] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed, and the settlement process is complete.

[0067] As described above, the present invention provides a system that efficiently and flexibly performs user recognition, order taking, serving, conversation, and payment processing, thereby optimizing service in restaurants.

[0068] The processing flow will be explained below.

[0069] Step 1:

[0070] The user enters the store.

[0071] Step 2:

[0072] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[0073] Step 3:

[0074] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[0075] Step 4:

[0076] The user says "Hello."

[0077] Step 5:

[0078] The terminal (voice recognition module) converts the user's speech into text.

[0079] Step 6:

[0080] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[0081] Step 7:

[0082] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[0083] Step 8:

[0084] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[0085] Step 9:

[0086] The terminal (voice recognition module) converts the user's speech into text.

[0087] Step 10:

[0088] The terminal (generative AI) analyzes the text and identifies the order contents.

[0089] Step 11:

[0090] The terminal sends the order details to the server.

[0091] Step 12:

[0092] The server receives the order information and passes it on to the kitchen and bar counter.

[0093] Step 13:

[0094] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[0095] Step 14:

[0096] The terminal will audibly convey a confirmation message to the user.

[0097] Step 15:

[0098] The server receives a notification from the kitchen that the food is ready.

[0099] Step 16:

[0100] The server sends a cooking completion notification to the terminal.

[0101] Step 17:

[0102] The device receives a notification, goes to the kitchen and places the food on a tray.

[0103] Step 18:

[0104] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0105] Step 19:

[0106] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[0107] Step 20:

[0108] The terminal (voice recognition module) converts the user's speech into text.

[0109] Step 21:

[0110] The terminal (generative AI) analyzes the text and identifies the additional order details.

[0111] Step 22:

[0112] The terminal transmits the additional order details to the server.

[0113] Step 23:

[0114] The terminal responds, "Okay, we'll add fries."

[0115] Step 24:

[0116] The user asks about dessert. Say, "What desserts do you have?"

[0117] Step 25:

[0118] The terminal (voice recognition module) converts the user's speech into text.

[0119] Step 26:

[0120] The terminal (generative AI) analyzes the text and generates a dessert menu.

[0121] Step 27:

[0122] The device will respond aloud, "Today's desserts include cake and ice cream."

[0123] Step 28:

[0124] The user wishes to settle the bill and says, "Please pay."

[0125] Step 29:

[0126] The terminal (voice recognition module) converts the user's speech into text.

[0127] Step 30:

[0128] The terminal (generator AI) identifies the settlement request and notifies the server.

[0129] Step 31:

[0130] The server compiles all the order details and calculates the total amount.

[0131] Step 32:

[0132] The server sends the settlement information to the terminal.

[0133] Step 33:

[0134] The device brings a tablet device to the user's table and says, "Please check this device."

[0135] Step 34:

[0136] The user checks the amount on the tablet device and makes the payment.

[0137] Step 35:

[0138] The terminal notifies the server that the payment has been completed.

[0139] Example 1

[0140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] Providing efficient and flexible service is a key issue for modern restaurants. In particular, there is a demand for systems that automate and seamlessly perform all processes, from recognizing users as they enter the restaurant to taking their orders, serving food, handling additional orders, and processing payments. Current systems often only provide individual functions and are insufficient for providing comprehensive services. Furthermore, delays in the recognition of specific users and the timing of voice responses can lead to problems that reduce user satisfaction.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0143] In this invention, the server includes means for inputting user voice, means for converting voice to text, means for analyzing the text to identify the order details, means for transmitting the order information to the server, means for receiving instructions from the server and controlling the robot, means for responding to the user by voice, means for detecting the user using facial recognition technology, and means for receiving notification of cooking completion. This enables the consistent automation of everything from recognizing the user's entry to receiving the order, serving the food, handling additional orders, and processing the payment, enabling the provision of efficient and flexible services.

[0144] The "means for inputting user's voice" refers to a device that acquires the user's speech using a voice input device such as a microphone.

[0145] A "means for converting speech to text" is a device or software that converts speech data acquired using speech recognition technology into text data.

[0146] The "means for analyzing text to identify order details" refers to software or an algorithm for analyzing the acquired character data and identifying the order details from the content of that data.

[0147] The "means for transmitting order information to the server" refers to a device or software that has the function of transmitting the specified order details to the server via a communication means.

[0148] The "means for receiving instructions from the server and controlling the robot" refers to a device or software for receiving instructions from the server and controlling the operation of the robot based on those instructions.

[0149] The "means for responding to the user by voice" refers to a device or software for providing the generated response message to the user by voice through a voice synthesizer.

[0150] A "means for detecting a user using facial recognition technology" is a device or software that uses a camera and a facial recognition algorithm to detect and identify a user's face.

[0151] The "means for receiving notification of cooking completion" refers to a device or software for receiving notification of cooking completion from the kitchen and sharing that information within the system.

[0152] The present invention provides a system for providing efficient and flexible service using robots in restaurants. Specific processing steps for carrying out the present invention will be described in natural language.

[0153] This system uses a microphone to input the user's voice and speech recognition technology to convert the voice into text. This technology uses the Google (registered trademark) Speech-to-Text API, among others. It also uses generative artificial intelligence (generative AI model) to analyze the text and identify the order contents. This analysis uses OpenAI (registered trademark)'s GPT-4 (registered trademark), among others.

[0154] A device with communication capabilities is used to send order information to the server. This sends the order details to the server (e.g., Dell PowerEdge T30). The robot operates based on a communication protocol to receive instructions from the server and control the robot. A typical service robot (e.g., SoftBank's Pepper) is used as the robot.

[0155] To respond to the user via voice, a generated response message is provided to the user using a speech synthesizer (e.g., Amazon Polly).Furthermore, to detect the user using facial recognition technology, a camera (e.g., Logitech HD Pro Webcam C920) and a facial recognition algorithm (e.g., OpenCV library) are used.To receive a notification that cooking is complete, a communication function is used to receive notifications from the kitchen and share that information with the device.

[0156] As a concrete example, let's consider a scenario where a user enters a restaurant. When the user enters the restaurant, the camera recognizes their face and the microphone picks up the utterance "Hello." The device converts the utterance into text through a voice recognition module, and the generation AI analyzes it and responds with "Hello, welcome. We will show you to your seat."

[0157] Here's an example of ordering: When a user says, "I'd like a hamburger and orange juice, please," the device converts the speech into text, and the generation AI identifies the order. This information is sent to the server, which then relays the order to the kitchen or bar counter. After confirming the order, the device responds, "I understand. A hamburger and orange juice, please."

[0158] If the user wishes to place an additional order, for example by saying, "I'd like some extra fries, please," the device analyzes the voice and uses a generative AI to identify the additional order. The additional order information is then sent to the server. After the cooking is complete, the robot brings the food from the kitchen and serves it to the user. It responds, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0159] Some examples of prompts are:

[0160] "How do you greet users when they walk in?"

[0161] "When a user places an order, how do you verify the order and provide cooking instructions?"

[0162] "How do you transport the food once it's done cooking?"

[0163] "How do you handle additional orders or inquiries from users?"

[0164] "How do you guide users through the process at checkout?"

[0165] In this way, the system is able to provide consistently efficient and flexible service, optimizing service in restaurants.

[0166] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0167] Specific explanation of processing steps

[0168] Step 1: User entry recognition

[0169] Specifically: The device uses a camera and microphone to recognize the user. The camera (Logitech HD Pro Webcam C920) detects the user's face, and the microphone (Shure MV5) picks up the user's speech.

[0170] Input: Camera image and audio data

[0171] Output: Face detection data and audio clips

[0172] Data processing / calculation: Detect faces from camera images (using OpenCV library) and save audio clips

[0173] Specific operation: When a user enters the store, the camera recognizes their face and confirms what they are saying through voice input.

[0174] Step 2: First voice recognition and greeting

[0175] Specifically, the device sends the user's speech to a speech recognition module (Google Speech-to-Text API) and converts it into text data. A generative AI (OpenAI's GPT-4) analyzes the text and generates an appropriate greeting message.

[0176] Input: Audio clip (user speech)

[0177] Output: Character data and response message

[0178] Data processing / calculation: Converting voice data to text data (Speech-to-Text), character analysis and response generation (generative AI)

[0179] Specific operation: When the user says "Hello," the device responds with "Hello, welcome. We will show you to your seat."

[0180] Step 3: Voice recognition of orders

[0181] Specific explanation: When a user speaks an order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the order contents.

[0182] Input: Audio clip (order details)

[0183] Output: Text data (order details) and order information

[0184] Data processing / calculation: Converting voice data into text data and identifying order details through text analysis

[0185] Specific operation: The user says, "I'd like a hamburger and orange juice, please," and this is identified as the order information.

[0186] Step 4: Communicating with the Server

[0187] Specific explanation: The terminal sends the identified order details to the server, which then transmits the order information to the kitchen or bar counter.

[0188] Input: Order Information

[0189] Output: Instructions to the kitchen and bar counter

[0190] Data processing / calculation: Sends order information to the server and transfers cooking instructions to each department

[0191] Specific operation: The terminal sends an order for "hamburger and orange juice" to the server, which then distributes it to the kitchen and bar counter.

[0192] Step 5: Order confirmation and response

[0193] Specific explanation: The generation AI generates a message to confirm the order details, and the terminal responds with a voice message saying, "Okay, a hamburger and orange juice, right?"

[0194] Input: Order Information

[0195] Output: Confirmation message

[0196] Data processing / calculation: Checking order information and generating messages

[0197] Specific operation: The generation AI creates a confirmation message, which the device then audibly conveys to the user.

[0198] Step 6: Receive notification that your food is ready

[0199] Specific explanation: The server receives a cooking completion notification from the kitchen and sends that information to the terminal.

[0200] Input: Cooking complete notification

[0201] Output: Cooking completion information

[0202] Data processing / calculation: Receiving and transmitting cooking completion notification

[0203] Specific operation: The kitchen notifies the server that the food is ready, and the server sends that information to the device.

[0204] Step 7: Food serving instructions and actions

[0205] Details: The server sends serving instructions to the terminal, which then moves to the kitchen according to the instructions and places the food on a tray. The robot then carries it to the table and serves it to the user.

[0206] Input: Cooking completion information

[0207] Output: Food served

[0208] Data processing / calculation: Generation of serving instructions and robot motion control

[0209] What it does: The device moves to the kitchen and places the food on a tray, which the robot then carries to the table.

[0210] Step 8: Accepting additional orders

[0211] Specific explanation: When a user speaks an additional order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the additional order.

[0212] Input: Audio Clip (reorder)

[0213] Output: Text data (reorder) and reorder information

[0214] Data processing / calculation: Converting voice data into text data and identifying additional orders

[0215] Specific operation: When the user says, "I'd like some extra fries, please," the device analyzes it and sends it to the server.

[0216] Step 9: Handling inquiries

[0217] Specific explanation: When a user speaks a question, the device converts the voice into text data, which is then analyzed by the generation AI to generate an appropriate answer.

[0218] Input: Audio clip (user question)

[0219] Output: Text data (question content) and answer message

[0220] Data processing / calculation: Converting voice data into text data, analyzing questions, and generating answers

[0221] Specific behavior: When the user asks, "What desserts do you have?", the device responds, "Today's desserts include cake and ice cream."

[0222] Step 10: Accepting the settlement request

[0223] Specific explanation: When a user says, "I'd like to pay the bill, please," the device converts the voice into text data, and the generation AI identifies the payment request and notifies the server.

[0224] Input: Audio clip (payment request)

[0225] Output: Text data (settlement request) and settlement notice

[0226] Data processing / calculation: Converting voice data into text data, identifying and notifying payment requests

[0227] Specific operation: When the user says "I'd like to pay the bill," the terminal notifies the server.

[0228] Step 11: Checkout

[0229] Specifically: The server compiles all the order details, calculates the total amount, and sends it to the terminal, which then brings a tablet to the user's table and displays the total amount.

[0230] Input: Order details

[0231] Output: Settlement amount

[0232] Data processing / calculation: Aggregating order details and calculating payment amounts

[0233] Specific operation: The server performs the calculation and the terminal displays the amount on the tablet.

[0234] Step 12: Check and complete the payment

[0235] Specific explanation: The user checks the amount on the tablet and makes the payment. The terminal notifies the server that the payment has been completed, and the settlement process is complete.

[0236] Input: Payment Information

[0237] Output: Payment completion notification

[0238] Data processing / calculation: Confirmation of payment information and notification of settlement completion

[0239] Specific operation: The user checks the amount on the tablet, makes the payment, and the device notifies the server.

[0240] In this way, the system can efficiently automate restaurant operations through a series of processes and provide users with high quality service.

[0241] (Application example 1)

[0242] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0243] Conventional food delivery systems have the problem that the entire process from ordering to delivery and payment is divided into manual processes and multiple different systems, resulting in a lack of efficiency and flexibility. The present invention aims to solve these problems, improve the efficiency of services for users, and provide a system that unifies and smoothly processes each stage of ordering, delivery, and payment.

[0244] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0245] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order details;] [means for sending order information to the server;] [means for controlling the robot upon receiving instructions from the server;] [means for responding to the user by voice;] [means for recognizing the user using facial recognition technology;] [means for sending delivery instructions to the delivery robot after accepting the order;] [means for responding to payment requests and calculating the payment amount; and [means for completing the user's payment.] This makes it possible to carry out the entire process from ordering to delivery and payment in a unified and efficient manner.

[0246] The "means for inputting user's voice" refers to a means for acquiring voice uttered by the user through a device.

[0247] The "means for converting voice to text" is a means for analyzing acquired voice data and converting it into text data.

[0248] The "means for analyzing text to identify order details" refers to a means for analyzing text data and accurately identifying the user's intentions and order details.

[0249] The "means for transmitting order information to the server" is a means for transmitting the analyzed order details to the server via a network.

[0250] The "means for receiving instructions from the server and controlling the robot" refers to means for receiving instruction data from the server and controlling the robot's operation based on that data.

[0251] The "means for responding to the user by voice" refers to a means for generating voice in response to an input from the user and responding to the user.

[0252] "Means for recognizing a user using facial recognition technology" refers to technology that uses a camera or the like to identify and specify the user's face.

[0253] The "means for sending delivery instructions to a delivery robot after receiving an order" refers to a means for transmitting the order details to a delivery robot and issuing delivery instructions after receiving a user's order.

[0254] The "means for responding to a payment request and calculating the payment amount" is a means for receiving a payment request from a user and calculating the payment amount based on all order details.

[0255] The "means for completing the user's settlement" is the means by which the user confirms the amount presented and completes the payment procedure.

[0256] This invention is a system that automates efficient and flexible robot-based food delivery services. This system automates a series of processes, from user voice input, speech-to-text conversion, order analysis, order information transmission to a server, robot control, voice response, user identification using facial recognition technology, delivery instructions to the delivery robot, payment request handling, and payment completion.

[0257] Hardware and Software Configuration

[0258] This system mainly consists of the following hardware and software:

[0259] Hardware: smartphones, delivery robots, servers

[0260] Software: Speech recognition module (e.g., Google Speech-to-Text API), face recognition module (e.g., OpenCV), generative AI model (e.g., OpenAI GPT-4), Python, Django framework, Celery

[0261] Program processing

[0262] The main processing of the entire system will be explained.

[0263] 1. Enter the user's voice:

[0264] The server and the terminal use a microphone to input the user's voice, which is then converted into text data by a voice recognition module.

[0265] 2. Convert speech to text:

[0266] The server uses a speech recognition module to convert the voice data into text.

[0267] 3. Parse the text to identify the order:

[0268] The terminal uses a generative AI model (e.g., GPT-4) to analyze the converted text and identify the order details.

[0269] 4. Send the order information to the server:

[0270] The terminal transmits the parsed order details to the server.

[0271] 5. Control the robot by receiving instructions from the server:

[0272] The server sends the order details to the delivery robot and controls the robot to pick up and deliver the order.

[0273] 6. Respond to the user verbally:

[0274] The device uses the generative AI model to generate a response message and conveys it to the user via voice output.

[0275] 7. Recognize users using facial recognition technology:

[0276] The device uses a camera to perform facial recognition and identify the user.

[0277] 8. After receiving the order, send delivery instructions to the delivery robot:

[0278] After receiving the order, the server issues delivery instructions to the delivery robot.

[0279] 9. Respond to settlement requests and calculate settlement amounts:

[0280] The server receives the user's payment request, calculates all the order details, and calculates the payment amount.

[0281] 10. Complete the user's checkout:

[0282] The terminal presents the payment amount to the user and completes the payment.

[0283] Specific examples

[0284] For example, when a user says "Pizza and Coke please" on a smartphone, the following process is executed:

[0285] 1. Speech recognition: The speech is converted into text, and the text data "Pizza and Coke, please" is generated.

[0286] 2. Order Analysis: A generative AI model analyzes the text data to identify an order for a pizza and a cola.

[0287] 3. Robot control: The server sends instructions to the delivery robot to pick up the pizza and cola from the store and deliver it to the user's address.

[0288] 4. Voice response: "Thank you for your order. Your pizza and Coke will be on their way soon," the device informs the user.

[0289] 5. Settlement: When the user says, "Please pay the bill," the server calculates the amount and the payment is completed on the settlement screen displayed on the smartphone.

[0290] As an example of input to the generative AI model, we will use the following prompt:

[0291] "A user asks, 'What's on the menu today?' What are your menu recommendations?"

[0292] By inputting the above prompt sentences into the generative AI model, an appropriate response can be obtained. This process enables efficient and flexible food delivery services.

[0293] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0294] Step 1:

[0295] The user inputs voice into the smartphone.

[0296] The user speaks into the smartphone microphone, "Pizza and Coke please." The input data is the user's voice data. The device acquires this voice data and passes it on to the next step.

[0297] Step 2:

[0298] The device converts the voice data into text data.

[0299] The device uses a speech recognition module (e.g., Google Speech-to-Text API) to convert the voice data into text data. Here, the input is voice data, and the output is text data such as "Pizza and Coke, please."

[0300] Step 3:

[0301] The terminal analyzes the text data and identifies the order details

[0302] The device uses a generative AI model (e.g., GPT-4) to analyze the text data. Specifically, it extracts the order details of "pizza" and "cola." The input is text data, and the output is the identified order details.

[0303] Step 4:

[0304] The terminal sends the order information to the server

[0305] The terminal sends the specified order details to the server, where the input is the order details data and the output is the order information sent to the server.

[0306] Step 5:

[0307] The server receives the order information and sends delivery instructions to the delivery robot.

[0308] Based on the received order information, the server sends an instruction to the delivery robot to "deliver pizza and cola to the specified address." The input is the order information, and the output is the delivery instruction data.

[0309] Step 6:

[0310] Delivery robots will pick up orders from stores and deliver them to designated addresses.

[0311] The delivery robot receives instructions from the server, goes to the store, picks up the pizza and cola, and delivers it to the specified user's address. The input is delivery instruction data, and the output is a delivery completion notification.

[0312] Step 7:

[0313] The device notifies the user of the delivery progress

[0314] The device uses a generative AI model to generate a message about the delivery progress, such as "Your pizza and coke will arrive soon," and notify the user via voice. The input is the delivery progress data, and the output is the voice message.

[0315] Step 8:

[0316] The user makes a payment request

[0317] The user verbally requests the terminal, "Please pay the bill." The input is voice data, which the terminal recognizes and proceeds to the next step.

[0318] Step 9:

[0319] The terminal converts the voice data into text data and sends a payment request to the server.

[0320] The terminal converts the voice data into text data using a voice recognition module and sends this text data to the server. The input is the voice data, and the output is the text data and payment request data sent to the server.

[0321] Step 10:

[0322] The server calculates the settlement amount and sends it to the terminal.

[0323] The server calculates the total amount based on all the order details and sends the total amount data to the terminal. The input is the order information and the output is the total amount data.

[0324] Step 11:

[0325] The terminal presents the payment amount to the user and completes the payment procedure.

[0326] The terminal uses the generative AI model to generate a message saying, "The settlement amount is XX yen," and presents it to the user via voice. The settlement procedure is then completed after the user performs the payment operation. The input is the settlement amount data, and the output is a payment completion notification.

[0327] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0328] The present invention is a customer service system using a robot in a restaurant, which is particularly equipped with the function of recognizing the user's emotions and responding flexibly based on them. This system has the following main functions:

[0329] User recognition and conversation initiation

[0330] User entry recognition

[0331] The device uses a camera and microphone to recognize the user. When a user enters the store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech, allowing the device to prepare for the initial interaction.

[0332] Initial voice recognition and greeting

[0333] When a user says "hello," the device sends the speech to a speech recognition module. The speech is converted into text data, and a generative AI analyzes the text to generate an appropriate greeting. The message "Hello, welcome. We will show you to your seat" is played over the device's speaker.

[0334] Receiving orders

[0335] Voice recognition for ordering

[0336] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which is then analyzed by a generation AI to determine the order.

[0337] Communicating with the Server

[0338] The terminal sends the identified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin quickly.

[0339] Acknowledgement and Response

[0340] The AI ​​generates a message confirming the order details, and the device responds to the user verbally, saying, "I understand. A hamburger and orange juice, right?"

[0341] Cooking status management and serving

[0342] Cooking completion notification

[0343] The server receives a cooking completion notification from the kitchen, then sends the cooking completion notification to the terminal, allowing the robot to begin serving the food.

[0344] Serving instructions and actions

[0345] The terminal receives the serving instructions, moves to the kitchen, and places the food on a tray. The robot then safely delivers the food to the table, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy."

[0346] Continuing the conversation and being flexible

[0347] Reorder acceptance and emotion recognition

[0348] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion engine recognizes the user's emotions from their voice and reflects them in the response. For example, if the emotion engine detects fatigue in the user's voice, the device will respond in a gentle tone, saying, "I understand. I'll bring the fries right away, so please wait a moment."

[0349] Inquiry response

[0350] When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion engine analyzes the user's emotions, and the device responds to the user in a gentle tone, saying, "Today's desserts include cake and ice cream. What would you like?"

[0351] settlement

[0352] Acceptance of settlement requests

[0353] When a user says, "Please give me the bill," the device converts the speech into text using a speech recognition module. The AI ​​generator identifies the payment request and notifies the server.

[0354] Settlement process

[0355] The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and tells them to "please check on this device."

[0356] Check and complete the payment

[0357] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed.

[0358] The system, which includes the above functions, can recognize users' emotions and provide flexible and efficient services based on those emotions, significantly improving the customer experience in restaurants.

[0359] The processing flow will be explained below.

[0360] Step 1:

[0361] The user enters the store.

[0362] Step 2:

[0363] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[0364] Step 3:

[0365] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[0366] Step 4:

[0367] The user says "Hello."

[0368] Step 5:

[0369] The terminal (voice recognition module) converts the user's speech into text.

[0370] Step 6:

[0371] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[0372] Step 7:

[0373] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[0374] Step 8:

[0375] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[0376] Step 9:

[0377] The terminal (voice recognition module) converts the user's speech into text.

[0378] Step 10:

[0379] The terminal (generative AI) analyzes the text and identifies the order contents.

[0380] Step 11:

[0381] The terminal transmits the specified order details to the server.

[0382] Step 12:

[0383] The server receives the order information and passes it on to the kitchen and bar counter.

[0384] Step 13:

[0385] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[0386] Step 14:

[0387] The terminal will audibly convey a confirmation message to the user.

[0388] Step 15:

[0389] The server receives a notification from the kitchen that the food is ready.

[0390] Step 16:

[0391] The server sends a cooking completion notification to the terminal.

[0392] Step 17:

[0393] The device receives a notification, goes to the kitchen and places the food on a tray.

[0394] Step 18:

[0395] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0396] Step 19:

[0397] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[0398] Step 20:

[0399] The terminal (voice recognition module) converts the user's speech into text.

[0400] Step 21:

[0401] The terminal (generative AI) analyzes the text and identifies the additional order details.

[0402] Step 22:

[0403] The terminal transmits the additional order details to the server.

[0404] Step 23:

[0405] The terminal responds, "Okay, we'll add fries."

[0406] Step 24:

[0407] The user asks about dessert. Say, "What desserts do you have?"

[0408] Step 25:

[0409] The terminal (voice recognition module) converts the user's speech into text.

[0410] Step 26:

[0411] The terminal (generative AI) analyzes the text and generates a dessert menu.

[0412] Step 27:

[0413] The device will respond aloud, "Today's desserts include cake and ice cream."

[0414] Step 28:

[0415] The user wishes to settle the bill and says, "Please pay."

[0416] Step 29:

[0417] The terminal (voice recognition module) converts the user's speech into text.

[0418] Step 30:

[0419] The terminal (generator AI) identifies the settlement request and notifies the server.

[0420] Step 31:

[0421] The server compiles all the order details and calculates the total amount.

[0422] Step 32:

[0423] The server sends the settlement information to the terminal.

[0424] Step 33:

[0425] The device brings a tablet device to the user's table and says, "Please check this device."

[0426] Step 34:

[0427] The user checks the amount on the tablet device and makes the payment.

[0428] Step 35:

[0429] The terminal notifies the server that the payment has been completed.

[0430] Added emotion engine processing

[0431] Step 36:

[0432] The device (emotion engine) analyzes the user's facial expressions and voice to recognize their emotional state.

[0433] Step 37:

[0434] The device (generative AI) generates a flexible response based on the recognized emotion.

[0435] Step 38:

[0436] The device will respond to the user based on their emotions. For example, if it recognizes that the user is tired, it will respond in a gentle tone, saying, "I understand. I'll bring you the fries right away, so please wait a moment."

[0437] Step 39:

[0438] The terminal (emotion engine) analyzes the accumulated emotional data and generates feedback to improve the quality of the service.

[0439] Step 40:

[0440] The server receives the feedback and uses it to improve services across the system.

[0441] Example 2

[0442] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0443] In modern restaurants, improving the efficiency and flexibility of customer service is important. However, conventional systems have difficulty recognizing customer emotions and responding flexibly. Accurate order acceptance and rapid processing are also required, but there is a high possibility of human error. To solve these issues, a system is needed that automatically recognizes customer voices and provides appropriate responses and order processing.

[0444] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for analyzing the text data to identify the order contents, means for transmitting the order information to the server, means for controlling the robot based on instructions received from the server, means for responding to the user by voice, and means for recognizing the user's emotions and generating a response based on the same. This enables flexible responses that take into account the customer's emotions and accurate and prompt order processing.

[0445] The "means for acquiring the user's voice" is a combination of hardware and software for acquiring the voice uttered by the user.

[0446] The "means for converting voice into text data" refers to a voice recognition technology for converting acquired voice into digital text data.

[0447] The "means for analyzing text data to identify order details" refers to algorithms and software for analyzing text data and accurately identifying the order details intended by the user.

[0448] The "means for transmitting order information to the server" refers to a communication means and protocol for transmitting the specified order details to the server via a network.

[0449] The "means for controlling the robot based on instructions received from the server" refers to hardware and software for receiving instructions sent from the server and controlling the robot in accordance with those instructions.

[0450] "Means for providing audio responses to the user" refers to a speaker and voice generating software for playing audio messages to the user.

[0451] The "means for recognizing the user's emotions and generating responses based on them" refers to an emotion analysis engine and generative AI model that analyzes emotions from the user's speech and actions and generates responses that are adapted to those emotions.

[0452] This invention is a customer service system using a robot in a restaurant, and is equipped with a function to recognize the user's emotions and respond flexibly based on them. This system is designed to acquire and analyze the user's voice to identify the order, and further recognize the user's emotions and respond accordingly.

[0453] Hardware and Software Configuration

[0454] Terminal

[0455] Camera: A device that uses facial recognition technology to detect users entering the store.

[0456] Microphone: A device used to capture the user's voice.

[0457] Speaker: A device that provides audio responses to the user.

[0458] Speech recognition module: Software for converting speech into text data using the Google Cloud Speech-to-Text API, etc.

[0459] Generative AI: Software that uses tools such as GPT-4 to analyze text data and generate appropriate responses.

[0460] Emotion analysis engine: Software for recognizing user emotions using Microsoft® Azure® Emotion API, etc.

[0461] server

[0462] Communications module: A device that receives order information sent from the terminal and forwards it to the kitchen or bar counter.

[0463] Information management system: Software for managing order information and cooking status in real time.

[0464] Specific operation of the system

[0465] User recognition and conversation initiation

[0466] When a user enters a restaurant, the device's camera detects the user using facial recognition technology (e.g., OpenCV), and the microphone captures the user's speech. This allows the device to prepare for the initial interaction. When the user says "Hello," the speech recognition module converts the speech into text data. The converted text data is sent to a generative AI, which generates an appropriate greeting. The device's speaker plays the message, "Hello, welcome. We will show you to your seat."

[0467] Receiving and processing orders

[0468] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text data. The generation AI analyzes the text and determines the order. The order is then sent to a server, which forwards the information to the kitchen or bar counter and requests cooking and drink preparation. The generation AI generates a message confirming the order, and the device responds with a voice saying, "I understand. A hamburger and orange juice, please."

[0469] Cooking status management and serving

[0470] The server receives a notification from the kitchen that the food is ready and notifies the terminal. The terminal receives the serving instructions, moves to the kitchen, places the food on a tray, and delivers it to the table. The robot delivers the food and responds with a voice message saying, "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[0471] Reorder acceptance and emotion recognition

[0472] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[0473] Inquiry response and settlement

[0474] When a user asks, "What desserts do you have?", the device converts the voice into text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. Would you like that?" When the user says, "Please pay the bill, please," the device converts the voice into text, and the generation AI identifies the payment request and notifies the server. The server compiles all the order details, calculates the payment amount, and sends it to the device. The device then brings a tablet device to the user's table and says, "Please check on this device." The user checks the amount on the tablet device and makes the payment, and the device notifies the server that the payment is complete.

[0475] Prompt Sentence Examples

[0476] "What's the scenario when a user walks into a store?"

[0477] "Please explain how emotions are recognized when ordering more."

[0478] "Please tell me in detail what steps a user should take to request a checkout."

[0479] The above is a specific operation of the system according to the embodiment of the present invention. This system is expected to make customer service in restaurants more efficient and flexible, thereby improving customer satisfaction.

[0480] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0481] Step 1:

[0482] A user enters the store

[0483] Input: User entry, video data from camera, audio data from microphone

[0484] Operation: The device recognizes the user's face using a camera installed at the entrance, processes the video data to detect the user's presence, and captures audio data using a microphone.

[0485] Output: Recognition result that the user has entered the store

[0486] Step 2:

[0487] Recognize the user's first utterance

[0488] Input: User's speech

[0489] How it works: When a user says "hello," the device captures this audio through the microphone. The audio data is sent to the Google Cloud Speech-to-Text API and converted to text.

[0490] Output: Text data (e.g. "Hello")

[0491] Step 3:

[0492] Generative AI creates greeting messages

[0493] Input: Text data (e.g. "Hello")

[0494] How it works: A generative AI (e.g., GPT-4) analyzes text data and generates an appropriate greeting.

[0495] Output: Greeting message (e.g. "Hello, welcome. We will show you to your seat.")

[0496] Step 4:

[0497] Voice output of greetings

[0498] Input: Greeting message (e.g. "Hello, welcome. We'll show you to your seat.")

[0499] Behavior: The generated greeting message is played aloud through the device's speaker.

[0500] Output: A voice response to the user

[0501] Step 5:

[0502] Get the user's orders

[0503] Input: User's voice order

[0504] How it works: When a user says, "I'd like a hamburger and orange juice, please," the device captures this speech through the microphone. The speech data is then sent to the Google Cloud Speech-to-Text API, where it is converted into text data.

[0505] Output: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[0506] Step 6:

[0507] Order analysis

[0508] Input: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[0509] How it works: Generative AI analyzes text data and identifies the order details.

[0510] Output: Identification of the order (e.g. "Hamburger", "Orange juice")

[0511] Step 7:

[0512] Send the order to the server

[0513] Input: Identification of order details (e.g. "Hamburger", "Orange juice")

[0514] Operation: The terminal sends the order information to the server.

[0515] Output: Order information received by the server

[0516] Step 8:

[0517] Transferring order information to the kitchen

[0518] Input: Order information received by the server

[0519] How it works: The server forwards the order information to the kitchen or bar counter and instructs cooking and preparation.

[0520] Output: Instructions received by the kitchen or bar counter

[0521] Step 9:

[0522] Order confirmation

[0523] Input: The order details identified by the generation AI (e.g., "hamburger," "orange juice")

[0524] How it works: The AI ​​generates a confirmation message and responds audibly through the device's speaker: "Okay, a hamburger and orange juice, right?"

[0525] Output: Audible acknowledgment to the user

[0526] Step 10:

[0527] Receive notifications when food is cooked

[0528] Input: Notification from the kitchen that cooking is complete

[0529] Operation: The server receives a notification from the kitchen that the food is ready and notifies the terminal.

[0530] Output: Cooking completion notification to the device

[0531] Step 11:

[0532] Receiving and acting on serving instructions

[0533] Input: Delivery instructions from the server

[0534] How it works: The terminal receives serving instructions, and the robot moves to the kitchen, places the food on a tray, and delivers it to the table.

[0535] Output: Food served

[0536] Step 12:

[0537] Reorder acceptance and emotion recognition

[0538] Input: User's voice for reorder

[0539] How it works: When a user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[0540] Output: Response to the user about the reorder

[0541] Step 13:

[0542] Inquiry response

[0543] Input: User's spoken query

[0544] How it works: When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. What would you like?"

[0545] Output: Audio response to user queries

[0546] Step 14:

[0547] Acceptance of settlement requests

[0548] Input: User's voice request for payment

[0549] How it works: When a user says, "I'd like to pay," the device converts the speech into text, and the generation AI identifies the payment request and notifies the server.

[0550] Output: Settlement request notification to the server

[0551] Step 15:

[0552] Settlement process

[0553] Input: Settlement request notification to the server

[0554] How it works: The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet to the user's table and says, "Please check on this terminal."

[0555] Output: The settlement amount displayed to the user

[0556] Step 16:

[0557] Check and complete the payment

[0558] Input: User payment confirmation

[0559] How it works: When the user checks the amount on the tablet and makes the payment, the device notifies the server that the payment is complete.

[0560] Output: Notification of successful payment

[0561] The above is the flow of processing of the program of this system.

[0562] (Application example 2)

[0563] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0564] Conventional restaurant and food delivery systems lack the flexibility to consider user emotions, resulting in a uniform quality of customer experience and difficulty in improving satisfaction. Furthermore, appropriate responses that consider user emotions are required when notifying users of delivery status and collecting feedback after delivery. Therefore, a system that can provide attentive service based on emotion recognition, in addition to voice recognition, is required, but no such system currently exists.

[0565] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0566] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order contents;] [means for sending order information to the server;] [means for controlling the robot by receiving instructions from the server;] [means for responding to the user by voice;] [means for using generative artificial intelligence to analyze the user's emotions and generate a response based on those emotions;] [means for notifying the user of the delivery status in real time; and [means for collecting post-delivery evaluations based on the user's emotions.] This enables flexible and efficient service provision based on the user's emotions and a consistent improvement in the customer experience during and after delivery.

[0567] "Means for inputting user voice" refers to technology for inputting user speech using a microphone or similar device in restaurants or food delivery systems.

[0568] The "means for converting speech to text" is a speech recognition technology that converts a user's speech data into character string data.

[0569] The "means for analyzing text to identify order details" is a technology for analyzing text data that has been speech-recognized to identify and specify the user's order details.

[0570] The "means for transmitting order information to a server" is a technique for transmitting user order information to a central server via a communication network.

[0571] "Means for controlling a robot by receiving instructions from a server" refers to a control technology that receives instructions from a server, operates the robot, and executes the specified task.

[0572] "Means for responding to the user by voice" refers to a technology for generating an appropriate voice response to the user's speech and transmitting it to the user through a speaker.

[0573] "Means using generative artificial intelligence to analyze a user's emotions and generate responses based on those emotions" refers to artificial intelligence technology that recognizes emotions from the user's voice and automatically generates natural responses based on those emotions.

[0574] "Means for notifying the user of the status during delivery in real time" refers to technology that notifies the user of the current status, estimated arrival time, etc. in real time during food delivery.

[0575] The "means for collecting post-delivery evaluations based on user emotions" is a technology for collecting delivery evaluation feedback based on user emotions after delivery is completed.

[0576] The present invention provides a system for restaurants and food delivery systems that provides flexible responses that take into account user emotions. This system has the following main functions:

[0577] Basic configuration

[0578] The system includes a means for inputting the user's voice, a means for converting the voice into text, a means for analyzing the text to identify the order contents, a means for sending the order information to a server, a means for receiving instructions from the server to control the robot, and a means for responding to the user by voice.

[0579] Additionally, it also includes generative artificial intelligence that analyzes the user's emotions and generates responses based on those emotions, a means of notifying the user of the delivery status in real time, and a means of collecting post-delivery evaluations based on the user's emotions.

[0580] Hardware and Software

[0581] Hardware: Uses the computer's built-in camera, microphone, robot, and speaker. These devices are used to input the user's voice and image, and to output response voices and delivery information.

[0582] software:

[0583] Speech Recognition Module: Uses the speech_recognition library, which converts the user's speech into text data.

[0584] Facial Recognition Software: Recognizes the user's face using the OpenCV library.

[0585] Generative AI: We use emotion recognition and text generation models from the transformers library, specifically the 'michellejieli / emotion_text' model and the 'GPT-3(R).5-turbo' model.

[0586] User voice recognition and emotion analysis

[0587] The device uses a camera and microphone to recognize the user, detects the user through facial recognition, and converts the user's speech into text using a speech recognition module. This text data is then input into an emotion recognition model to identify the user's emotions.

[0588] Response Generation

[0589] The generative AI generates an appropriate response based on the user's emotions. For example, if the user's voice is recognized as saying "hello" and the voice contains the emotion of "welcome," the generative AI model will generate a response such as "Welcome, please place your order." This response is then conveyed to the user through the speaker.

[0590] Examples and prompts

[0591] Examples:

[0592] Scenario: A user launches your app and says, "Good evening, I'd like a pizza and a Coke, please."

[0593] Emotion recognition: The emotion of "friendliness" is recognized from the user's voice.

[0594] Output Response: "Thank you for your order. Pizza and Coke, please. Now, please tell me your delivery address."

[0595] Example prompt sentence:

[0596] "Respond in a friendly tone: Welcome, your order is welcome."

[0597] "Please proceed to the next step to confirm your delivery address."

[0598] This will enable flexible and efficient service provision based on user emotions and improve the consistent customer experience during and after delivery.

[0599] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0600] Step 1:

[0601] The device recognizes the user using a camera and microphone. The camera detects the user's face and uses facial recognition software (OpenCV) to confirm the user's presence. At the same time, the microphone collects the user's speech and inputs it as audio data. The output is the user's facial image data and audio data.

[0602] Step 2:

[0603] The device sends the collected voice data to a voice recognition module (speech_recognition library), which converts the voice data into text data. The input is the user's voice data, and the output is text data. Specifically, the voice recognition module analyzes the voice waveform and generates a corresponding string of characters.

[0604] Step 3:

[0605] The server receives the text data and uses generative artificial intelligence (the transformers library) to analyze the spoken text and identify the user's emotions. The input is text data generated by speech recognition, and the output is the recognized emotion label. Specifically, the emotion recognition model analyzes the text and determines emotions such as "welcoming" or "friendliness."

[0606] Step 4:

[0607] The server uses generative artificial intelligence to generate an appropriate response based on the identified emotion. The input is text data containing the user's emotion label and order details, and the output is a response text. Specifically, the response generation model forms a natural-sounding answer while taking the emotion label into account.

[0608] Step 5:

[0609] The server sends the generated response text to the terminal, and the terminal responds to the user audibly through the speaker. The input is the response text sent from the server, and the output is the audio response to the user. Specifically, the text is converted into audio and played back so that the user can hear it.

[0610] Step 6:

[0611] The terminal sends the user's order information to the server, which then transmits the order information to the cooking station or delivery station. The input is the order information written in text, and the output is an execution instruction sent to the cooking station or delivery station. Specifically, the order information is sent to each station via the network.

[0612] Step 7:

[0613] The server receives delivery status information from the delivery station in real time and notifies the user of the delivery status. The input is status report data from the delivery station, and the output is delivery progress information displayed to the user. Specifically, the status data is notified to the user's smartphone or other device.

[0614] Step 8:

[0615] When the delivery is completed, the terminal collects the user's evaluation based on their emotions and sends it to the server. The input is the user's voice feedback, and the output is the evaluation data sent to the server. Specifically, the voice feedback is converted into text and saved along with the user's emotional evaluation.

[0616] Step 9:

[0617] The server analyzes the collected evaluation data to improve service quality. The input is the collected evaluation data, and the output is reports and statistical information for improvement. Specifically, the evaluation data in the database is analyzed to identify areas for improvement.

[0618] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0619] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0620] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0621] [Second embodiment]

[0622] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0623] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0624] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0625] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0626] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0627] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0628] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0629] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0630] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0631] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0632] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0633] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0634] The present invention is a system for providing efficient and flexible service using robots in restaurants. Specific embodiments for carrying out the invention and the processing of the program therefor will be described below.

[0635] User recognition and conversation initiation

[0636] User entry recognition

[0637] The device uses a camera and microphone to recognize the user. When a user enters a store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech. Through facial recognition algorithms and voice input, it is possible to quickly respond to users even when meeting them for the first time.

[0638] Initial voice recognition and greeting

[0639] When a user says "hello," the device sends the speech to a speech recognition module, which converts the speech to text. The generative AI analyzes the text and generates an appropriate greeting. The device responds through the speaker, saying, "Hello, welcome. We'll show you to your seat."

[0640] Receiving orders

[0641] Voice recognition for ordering

[0642] If a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which the generation AI then analyzes to determine the order.

[0643] Communicating with the Server

[0644] The terminal sends the specified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin promptly.

[0645] Acknowledgement and Response

[0646] The AI ​​generates a message to confirm the order, and the device responds aloud, saying, "Understood. A hamburger and orange juice, please."

[0647] Cooking status management and serving

[0648] Cooking completion notification

[0649] The server receives a notification from the kitchen that the food is ready and sends that information to the terminal, allowing the robot to begin serving the food.

[0650] Serving instructions and actions

[0651] The server sends serving instructions to the terminal, which then follows those instructions to move to the kitchen and place the food on a tray. The robot then safely carries the food to the table and serves it to the user, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[0652] Continuing the conversation and being flexible

[0653] Accepting additional orders

[0654] If the user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. The device responds, "Okay, I'll order some extra fries," and sends the additional order information to the server.

[0655] Inquiry response

[0656] When a user asks, "What desserts do you have?", the device converts the voice into text, and the AI ​​analyzes the query and generates an appropriate answer. The device responds aloud, "Today's desserts include cake and ice cream."

[0657] settlement

[0658] Acceptance of settlement requests

[0659] When a user says, "I'd like to pay, please," the device converts the speech into text using a voice recognition module, and the generation AI identifies the payment request and notifies the server.

[0660] Settlement process

[0661] The server compiles all the order details, calculates the payment amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and says, "Please check on this device."

[0662] Check and complete the payment

[0663] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed, and the settlement process is complete.

[0664] As described above, the present invention provides a system that efficiently and flexibly performs user recognition, order taking, serving, conversation, and payment processing, thereby optimizing service in restaurants.

[0665] The processing flow will be explained below.

[0666] Step 1:

[0667] The user enters the store.

[0668] Step 2:

[0669] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[0670] Step 3:

[0671] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[0672] Step 4:

[0673] The user says "Hello."

[0674] Step 5:

[0675] The terminal (voice recognition module) converts the user's speech into text.

[0676] Step 6:

[0677] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[0678] Step 7:

[0679] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[0680] Step 8:

[0681] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[0682] Step 9:

[0683] The terminal (voice recognition module) converts the user's speech into text.

[0684] Step 10:

[0685] The terminal (generative AI) analyzes the text and identifies the order contents.

[0686] Step 11:

[0687] The terminal sends the order details to the server.

[0688] Step 12:

[0689] The server receives the order information and passes it on to the kitchen and bar counter.

[0690] Step 13:

[0691] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[0692] Step 14:

[0693] The terminal will audibly convey a confirmation message to the user.

[0694] Step 15:

[0695] The server receives a notification from the kitchen that the food is ready.

[0696] Step 16:

[0697] The server sends a cooking completion notification to the terminal.

[0698] Step 17:

[0699] The device receives a notification, goes to the kitchen and places the food on a tray.

[0700] Step 18:

[0701] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0702] Step 19:

[0703] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[0704] Step 20:

[0705] The terminal (voice recognition module) converts the user's speech into text.

[0706] Step 21:

[0707] The terminal (generative AI) analyzes the text and identifies the additional order details.

[0708] Step 22:

[0709] The terminal transmits the additional order details to the server.

[0710] Step 23:

[0711] The terminal responds, "Okay, we'll add fries."

[0712] Step 24:

[0713] The user asks about dessert. Say, "What desserts do you have?"

[0714] Step 25:

[0715] The terminal (voice recognition module) converts the user's speech into text.

[0716] Step 26:

[0717] The terminal (generative AI) analyzes the text and generates a dessert menu.

[0718] Step 27:

[0719] The device will respond aloud, "Today's desserts include cake and ice cream."

[0720] Step 28:

[0721] The user wishes to settle the bill and says, "Please pay."

[0722] Step 29:

[0723] The terminal (voice recognition module) converts the user's speech into text.

[0724] Step 30:

[0725] The terminal (generator AI) identifies the settlement request and notifies the server.

[0726] Step 31:

[0727] The server compiles all the order details and calculates the total amount.

[0728] Step 32:

[0729] The server sends the settlement information to the terminal.

[0730] Step 33:

[0731] The device brings a tablet device to the user's table and says, "Please check this device."

[0732] Step 34:

[0733] The user checks the amount on the tablet device and makes the payment.

[0734] Step 35:

[0735] The terminal notifies the server that the payment has been completed.

[0736] Example 1

[0737] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0738] Providing efficient and flexible service is a key issue for modern restaurants. In particular, there is a demand for systems that automate and seamlessly perform all processes, from recognizing users as they enter the restaurant to taking their orders, serving food, handling additional orders, and processing payments. Current systems often only provide individual functions and are insufficient for providing comprehensive services. Furthermore, delays in the recognition of specific users and the timing of voice responses can lead to problems that reduce user satisfaction.

[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0740] In this invention, the server includes means for inputting user voice, means for converting voice to text, means for analyzing the text to identify the order details, means for transmitting the order information to the server, means for receiving instructions from the server and controlling the robot, means for responding to the user by voice, means for detecting the user using facial recognition technology, and means for receiving notification of cooking completion. This enables the consistent automation of everything from recognizing the user's entry to receiving the order, serving the food, handling additional orders, and processing the payment, enabling the provision of efficient and flexible services.

[0741] The "means for inputting user's voice" refers to a device that acquires the user's speech using a voice input device such as a microphone.

[0742] A "means for converting speech to text" is a device or software that converts speech data acquired using speech recognition technology into text data.

[0743] The "means for analyzing text to identify order details" refers to software or an algorithm for analyzing the acquired character data and identifying the order details from the content of that data.

[0744] The "means for transmitting order information to the server" refers to a device or software that has the function of transmitting the specified order details to the server via a communication means.

[0745] The "means for receiving instructions from the server and controlling the robot" refers to a device or software for receiving instructions from the server and controlling the operation of the robot based on those instructions.

[0746] The "means for responding to the user by voice" refers to a device or software for providing the generated response message to the user by voice through a voice synthesizer.

[0747] A "means for detecting a user using facial recognition technology" is a device or software that uses a camera and a facial recognition algorithm to detect and identify a user's face.

[0748] The "means for receiving notification of cooking completion" refers to a device or software for receiving notification of cooking completion from the kitchen and sharing that information within the system.

[0749] The present invention provides a system for providing efficient and flexible service using robots in restaurants. Specific processing steps for carrying out the present invention will be described in natural language.

[0750] This system uses a microphone to input the user's voice and speech recognition technology to convert the voice into text. This is done using the Google Speech-to-Text API, among others. It also uses generative artificial intelligence (generative AI model) to analyze the text and identify the order contents. This analysis utilizes OpenAI's GPT-4, among others.

[0751] A device with communication capabilities is used to send order information to the server. This sends the order details to the server (e.g., Dell PowerEdge T30). The robot operates based on a communication protocol to receive instructions from the server and control the robot. A typical service robot (e.g., SoftBank's Pepper) is used as the robot.

[0752] To respond to the user via voice, a generated response message is provided to the user using a speech synthesizer (e.g., Amazon Polly).Furthermore, to detect the user using facial recognition technology, a camera (e.g., Logitech HD Pro Webcam C920) and a facial recognition algorithm (e.g., OpenCV library) are used.To receive a notification that cooking is complete, a communication function is used to receive notifications from the kitchen and share that information with the device.

[0753] As a concrete example, let's consider a scenario where a user enters a restaurant. When the user enters the restaurant, the camera recognizes their face and the microphone picks up the utterance "Hello." The device converts the utterance into text through a voice recognition module, and the generation AI analyzes it and responds with "Hello, welcome. We will show you to your seat."

[0754] Here's an example of ordering: When a user says, "I'd like a hamburger and orange juice, please," the device converts the speech into text, and the generation AI identifies the order. This information is sent to the server, which then relays the order to the kitchen or bar counter. After confirming the order, the device responds, "I understand. A hamburger and orange juice, please."

[0755] If the user wishes to place an additional order, for example by saying, "I'd like some extra fries, please," the device analyzes the voice and uses a generative AI to identify the additional order. The additional order information is then sent to the server. After the cooking is complete, the robot brings the food from the kitchen and serves it to the user. It responds, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0756] Some examples of prompts are:

[0757] "How do you greet users when they walk in?"

[0758] "When a user places an order, how do you verify the order and provide cooking instructions?"

[0759] "How do you transport the food once it's done cooking?"

[0760] "How do you handle additional orders or inquiries from users?"

[0761] "How do you guide users through the process at checkout?"

[0762] In this way, the system is able to provide consistently efficient and flexible service, optimizing service in restaurants.

[0763] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0764] Specific explanation of processing steps

[0765] Step 1: User entry recognition

[0766] Specifically: The device uses a camera and microphone to recognize the user. The camera (Logitech HD Pro Webcam C920) detects the user's face, and the microphone (Shure MV5) picks up the user's speech.

[0767] Input: Camera image and audio data

[0768] Output: Face detection data and audio clips

[0769] Data processing / calculation: Detect faces from camera images (using OpenCV library) and save audio clips

[0770] Specific operation: When a user enters the store, the camera recognizes their face and confirms what they are saying through voice input.

[0771] Step 2: First voice recognition and greeting

[0772] Specifically, the device sends the user's speech to a speech recognition module (Google Speech-to-Text API) and converts it into text data. A generative AI (OpenAI's GPT-4) analyzes the text and generates an appropriate greeting message.

[0773] Input: Audio clip (user speech)

[0774] Output: Character data and response message

[0775] Data processing / calculation: Converting voice data to text data (Speech-to-Text), character analysis and response generation (generative AI)

[0776] Specific operation: When the user says "Hello," the device responds with "Hello, welcome. We will show you to your seat."

[0777] Step 3: Voice recognition of orders

[0778] Specific explanation: When a user speaks an order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the order contents.

[0779] Input: Audio clip (order details)

[0780] Output: Text data (order details) and order information

[0781] Data processing / calculation: Converting voice data into text data and identifying order details through text analysis

[0782] Specific operation: The user says, "I'd like a hamburger and orange juice, please," and this is identified as the order information.

[0783] Step 4: Communicating with the Server

[0784] Specific explanation: The terminal sends the identified order details to the server, which then transmits the order information to the kitchen or bar counter.

[0785] Input: Order Information

[0786] Output: Instructions to the kitchen and bar counter

[0787] Data processing / calculation: Sends order information to the server and transfers cooking instructions to each department

[0788] Specific operation: The terminal sends an order for "hamburger and orange juice" to the server, which then distributes it to the kitchen and bar counter.

[0789] Step 5: Order confirmation and response

[0790] Specific explanation: The generation AI generates a message to confirm the order details, and the terminal responds with a voice message saying, "Okay, a hamburger and orange juice, right?"

[0791] Input: Order Information

[0792] Output: Confirmation message

[0793] Data processing / calculation: Checking order information and generating messages

[0794] Specific operation: The generation AI creates a confirmation message, which the device then audibly conveys to the user.

[0795] Step 6: Receive notification that your food is ready

[0796] Specific explanation: The server receives a cooking completion notification from the kitchen and sends that information to the terminal.

[0797] Input: Cooking complete notification

[0798] Output: Cooking completion information

[0799] Data processing / calculation: Receiving and transmitting cooking completion notification

[0800] Specific operation: The kitchen notifies the server that the food is ready, and the server sends that information to the device.

[0801] Step 7: Food serving instructions and actions

[0802] Details: The server sends serving instructions to the terminal, which then moves to the kitchen according to the instructions and places the food on a tray. The robot then carries it to the table and serves it to the user.

[0803] Input: Cooking completion information

[0804] Output: Food served

[0805] Data processing / calculation: Generation of serving instructions and robot motion control

[0806] What it does: The device moves to the kitchen and places the food on a tray, which the robot then carries to the table.

[0807] Step 8: Accepting additional orders

[0808] Specific explanation: When a user speaks an additional order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the additional order.

[0809] Input: Audio Clip (reorder)

[0810] Output: Text data (reorder) and reorder information

[0811] Data processing / calculation: Converting voice data into text data and identifying additional orders

[0812] Specific operation: When the user says, "I'd like some extra fries, please," the device analyzes it and sends it to the server.

[0813] Step 9: Handling inquiries

[0814] Specific explanation: When a user speaks a question, the device converts the voice into text data, which is then analyzed by the generation AI to generate an appropriate answer.

[0815] Input: Audio clip (user question)

[0816] Output: Text data (question content) and answer message

[0817] Data processing / calculation: Converting voice data into text data, analyzing questions, and generating answers

[0818] Specific behavior: When the user asks, "What desserts do you have?", the device responds, "Today's desserts include cake and ice cream."

[0819] Step 10: Accepting the settlement request

[0820] Specific explanation: When a user says, "I'd like to pay the bill, please," the device converts the voice into text data, and the generation AI identifies the payment request and notifies the server.

[0821] Input: Audio clip (payment request)

[0822] Output: Text data (settlement request) and settlement notice

[0823] Data processing / calculation: Converting voice data into text data, identifying and notifying payment requests

[0824] Specific operation: When the user says "I'd like to pay the bill," the terminal notifies the server.

[0825] Step 11: Checkout

[0826] Specifically: The server compiles all the order details, calculates the total amount, and sends it to the terminal, which then brings a tablet to the user's table and displays the total amount.

[0827] Input: Order details

[0828] Output: Settlement amount

[0829] Data processing / calculation: Aggregating order details and calculating payment amounts

[0830] Specific operation: The server performs the calculation and the terminal displays the amount on the tablet.

[0831] Step 12: Check and complete the payment

[0832] Specific explanation: The user checks the amount on the tablet and makes the payment. The terminal notifies the server that the payment has been completed, and the settlement process is complete.

[0833] Input: Payment Information

[0834] Output: Payment completion notification

[0835] Data processing / calculation: Confirmation of payment information and notification of settlement completion

[0836] Specific operation: The user checks the amount on the tablet, makes the payment, and the device notifies the server.

[0837] In this way, the system can efficiently automate restaurant operations through a series of processes and provide users with high quality service.

[0838] (Application example 1)

[0839] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0840] Conventional food delivery systems have the problem that the entire process from ordering to delivery and payment is divided into manual processes and multiple different systems, resulting in a lack of efficiency and flexibility. The present invention aims to solve these problems, improve the efficiency of services for users, and provide a system that unifies and smoothly processes each stage of ordering, delivery, and payment.

[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0842] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order details;] [means for sending order information to the server;] [means for controlling the robot upon receiving instructions from the server;] [means for responding to the user by voice;] [means for recognizing the user using facial recognition technology;] [means for sending delivery instructions to the delivery robot after accepting the order;] [means for responding to payment requests and calculating the payment amount; and [means for completing the user's payment.] This makes it possible to carry out the entire process from ordering to delivery and payment in a unified and efficient manner.

[0843] The "means for inputting user's voice" refers to a means for acquiring voice uttered by the user through a device.

[0844] The "means for converting voice to text" is a means for analyzing acquired voice data and converting it into text data.

[0845] The "means for analyzing text to identify order details" refers to a means for analyzing text data and accurately identifying the user's intentions and order details.

[0846] The "means for transmitting order information to the server" is a means for transmitting the analyzed order details to the server via a network.

[0847] The "means for receiving instructions from the server and controlling the robot" refers to means for receiving instruction data from the server and controlling the robot's operation based on that data.

[0848] The "means for responding to the user by voice" refers to a means for generating voice in response to an input from the user and responding to the user.

[0849] "Means for recognizing a user using facial recognition technology" refers to technology that uses a camera or the like to identify and specify the user's face.

[0850] The "means for sending delivery instructions to a delivery robot after receiving an order" refers to a means for transmitting the order details to a delivery robot and issuing delivery instructions after receiving a user's order.

[0851] The "means for responding to a payment request and calculating the payment amount" is a means for receiving a payment request from a user and calculating the payment amount based on all order details.

[0852] The "means for completing the user's settlement" is the means by which the user confirms the amount presented and completes the payment procedure.

[0853] This invention is a system that automates efficient and flexible robot-based food delivery services. This system automates a series of processes, from user voice input, speech-to-text conversion, order analysis, order information transmission to a server, robot control, voice response, user identification using facial recognition technology, delivery instructions to the delivery robot, payment request handling, and payment completion.

[0854] Hardware and Software Configuration

[0855] This system mainly consists of the following hardware and software:

[0856] Hardware: smartphones, delivery robots, servers

[0857] Software: Speech recognition module (e.g., Google Speech-to-Text API), face recognition module (e.g., OpenCV), generative AI model (e.g., OpenAI GPT-4), Python, Django framework, Celery

[0858] Program processing

[0859] The main processing of the entire system will be explained.

[0860] 1. Enter the user's voice:

[0861] The server and the terminal use a microphone to input the user's voice, which is then converted into text data by a voice recognition module.

[0862] 2. Convert speech to text:

[0863] The server uses a speech recognition module to convert the voice data into text.

[0864] 3. Parse the text to identify the order:

[0865] The terminal uses a generative AI model (e.g., GPT-4) to analyze the converted text and identify the order details.

[0866] 4. Send the order information to the server:

[0867] The terminal transmits the parsed order details to the server.

[0868] 5. Control the robot by receiving instructions from the server:

[0869] The server sends the order details to the delivery robot and controls the robot to pick up and deliver the order.

[0870] 6. Respond to the user verbally:

[0871] The device uses the generative AI model to generate a response message and conveys it to the user via voice output.

[0872] 7. Recognize users using facial recognition technology:

[0873] The device uses a camera to perform facial recognition and identify the user.

[0874] 8. After receiving the order, send delivery instructions to the delivery robot:

[0875] After receiving the order, the server issues delivery instructions to the delivery robot.

[0876] 9. Respond to settlement requests and calculate settlement amounts:

[0877] The server receives the user's payment request, calculates all the order details, and calculates the payment amount.

[0878] 10. Complete the user's checkout:

[0879] The terminal presents the payment amount to the user and completes the payment.

[0880] Specific examples

[0881] For example, when a user says "Pizza and Coke please" on a smartphone, the following process is executed:

[0882] 1. Speech recognition: The speech is converted into text, and the text data "Pizza and Coke, please" is generated.

[0883] 2. Order Analysis: A generative AI model analyzes the text data to identify an order for a pizza and a cola.

[0884] 3. Robot control: The server sends instructions to the delivery robot to pick up the pizza and cola from the store and deliver it to the user's address.

[0885] 4. Voice response: "Thank you for your order. Your pizza and Coke will be on their way soon," the device informs the user.

[0886] 5. Settlement: When the user says, "Please pay the bill," the server calculates the amount and the payment is completed on the settlement screen displayed on the smartphone.

[0887] As an example of input to the generative AI model, we will use the following prompt:

[0888] "A user asks, 'What's on the menu today?' What are your menu recommendations?"

[0889] By inputting the above prompt sentences into the generative AI model, an appropriate response can be obtained. This process enables efficient and flexible food delivery services.

[0890] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0891] Step 1:

[0892] The user inputs voice into the smartphone.

[0893] The user speaks into the smartphone microphone, "Pizza and Coke please." The input data is the user's voice data. The device acquires this voice data and passes it on to the next step.

[0894] Step 2:

[0895] The device converts the voice data into text data.

[0896] The device uses a speech recognition module (e.g., Google Speech-to-Text API) to convert the voice data into text data. Here, the input is voice data, and the output is text data such as "Pizza and Coke, please."

[0897] Step 3:

[0898] The terminal analyzes the text data and identifies the order details

[0899] The device uses a generative AI model (e.g., GPT-4) to analyze the text data. Specifically, it extracts the order details of "pizza" and "cola." The input is text data, and the output is the identified order details.

[0900] Step 4:

[0901] The terminal sends the order information to the server

[0902] The terminal sends the specified order details to the server, where the input is the order details data and the output is the order information sent to the server.

[0903] Step 5:

[0904] The server receives the order information and sends delivery instructions to the delivery robot.

[0905] Based on the received order information, the server sends an instruction to the delivery robot to "deliver pizza and cola to the specified address." The input is the order information, and the output is the delivery instruction data.

[0906] Step 6:

[0907] Delivery robots will pick up orders from stores and deliver them to designated addresses.

[0908] The delivery robot receives instructions from the server, goes to the store, picks up the pizza and cola, and delivers it to the specified user's address. The input is delivery instruction data, and the output is a delivery completion notification.

[0909] Step 7:

[0910] The device notifies the user of the delivery progress

[0911] The device uses a generative AI model to generate a message about the delivery progress, such as "Your pizza and coke will arrive soon," and notify the user via voice. The input is the delivery progress data, and the output is the voice message.

[0912] Step 8:

[0913] The user makes a payment request

[0914] The user verbally requests the terminal, "Please pay the bill." The input is voice data, which the terminal recognizes and proceeds to the next step.

[0915] Step 9:

[0916] The terminal converts the voice data into text data and sends a payment request to the server.

[0917] The terminal converts the voice data into text data using a voice recognition module and sends this text data to the server. The input is the voice data, and the output is the text data and payment request data sent to the server.

[0918] Step 10:

[0919] The server calculates the settlement amount and sends it to the terminal.

[0920] The server calculates the total amount based on all the order details and sends the total amount data to the terminal. The input is the order information and the output is the total amount data.

[0921] Step 11:

[0922] The terminal presents the payment amount to the user and completes the payment procedure.

[0923] The terminal uses the generative AI model to generate a message saying, "The settlement amount is XX yen," and presents it to the user via voice. The settlement procedure is then completed after the user performs the payment operation. The input is the settlement amount data, and the output is a payment completion notification.

[0924] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0925] The present invention is a customer service system using a robot in a restaurant, which is particularly equipped with the function of recognizing the user's emotions and responding flexibly based on them. This system has the following main functions:

[0926] User recognition and conversation initiation

[0927] User entry recognition

[0928] The device uses a camera and microphone to recognize the user. When a user enters the store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech, allowing the device to prepare for the initial interaction.

[0929] Initial voice recognition and greeting

[0930] When a user says "hello," the device sends the speech to a speech recognition module. The speech is converted into text data, and a generative AI analyzes the text to generate an appropriate greeting. The message "Hello, welcome. We will show you to your seat" is played over the device's speaker.

[0931] Receiving orders

[0932] Voice recognition for ordering

[0933] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which is then analyzed by a generation AI to determine the order.

[0934] Communicating with the Server

[0935] The terminal sends the identified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin quickly.

[0936] Acknowledgement and Response

[0937] The AI ​​generates a message confirming the order details, and the device responds to the user verbally, saying, "I understand. A hamburger and orange juice, right?"

[0938] Cooking status management and serving

[0939] Cooking completion notification

[0940] The server receives a cooking completion notification from the kitchen, then sends the cooking completion notification to the terminal, allowing the robot to begin serving the food.

[0941] Serving instructions and actions

[0942] The terminal receives the serving instructions, moves to the kitchen, and places the food on a tray. The robot then safely delivers the food to the table, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy."

[0943] Continuing the conversation and being flexible

[0944] Reorder acceptance and emotion recognition

[0945] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion engine recognizes the user's emotions from their voice and reflects them in the response. For example, if the emotion engine detects fatigue in the user's voice, the device will respond in a gentle tone, saying, "I understand. I'll bring the fries right away, so please wait a moment."

[0946] Inquiry response

[0947] When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion engine analyzes the user's emotions, and the device responds to the user in a gentle tone, saying, "Today's desserts include cake and ice cream. What would you like?"

[0948] settlement

[0949] Acceptance of settlement requests

[0950] When a user says, "Please give me the bill," the device converts the speech into text using a speech recognition module. The AI ​​generator identifies the payment request and notifies the server.

[0951] Settlement process

[0952] The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and tells them to "please check on this device."

[0953] Check and complete the payment

[0954] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed.

[0955] The system, which includes the above functions, can recognize users' emotions and provide flexible and efficient services based on those emotions, significantly improving the customer experience in restaurants.

[0956] The processing flow will be explained below.

[0957] Step 1:

[0958] The user enters the store.

[0959] Step 2:

[0960] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[0961] Step 3:

[0962] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[0963] Step 4:

[0964] The user says "Hello."

[0965] Step 5:

[0966] The terminal (voice recognition module) converts the user's speech into text.

[0967] Step 6:

[0968] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[0969] Step 7:

[0970] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[0971] Step 8:

[0972] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[0973] Step 9:

[0974] The terminal (voice recognition module) converts the user's speech into text.

[0975] Step 10:

[0976] The terminal (generative AI) analyzes the text and identifies the order contents.

[0977] Step 11:

[0978] The terminal transmits the specified order details to the server.

[0979] Step 12:

[0980] The server receives the order information and passes it on to the kitchen and bar counter.

[0981] Step 13:

[0982] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[0983] Step 14:

[0984] The terminal will audibly convey a confirmation message to the user.

[0985] Step 15:

[0986] The server receives a notification from the kitchen that the food is ready.

[0987] Step 16:

[0988] The server sends a cooking completion notification to the terminal.

[0989] Step 17:

[0990] The device receives a notification, goes to the kitchen and places the food on a tray.

[0991] Step 18:

[0992] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[0993] Step 19:

[0994] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[0995] Step 20:

[0996] The terminal (voice recognition module) converts the user's speech into text.

[0997] Step 21:

[0998] The terminal (generative AI) analyzes the text and identifies the additional order details.

[0999] Step 22:

[1000] The terminal transmits the additional order details to the server.

[1001] Step 23:

[1002] The terminal responds, "Okay, we'll add fries."

[1003] Step 24:

[1004] The user asks about dessert. Say, "What desserts do you have?"

[1005] Step 25:

[1006] The terminal (voice recognition module) converts the user's speech into text.

[1007] Step 26:

[1008] The terminal (generative AI) analyzes the text and generates a dessert menu.

[1009] Step 27:

[1010] The device will respond aloud, "Today's desserts include cake and ice cream."

[1011] Step 28:

[1012] The user wishes to settle the bill and says, "Please pay."

[1013] Step 29:

[1014] The terminal (voice recognition module) converts the user's speech into text.

[1015] Step 30:

[1016] The terminal (generator AI) identifies the settlement request and notifies the server.

[1017] Step 31:

[1018] The server compiles all the order details and calculates the total amount.

[1019] Step 32:

[1020] The server sends the settlement information to the terminal.

[1021] Step 33:

[1022] The device brings a tablet device to the user's table and says, "Please check this device."

[1023] Step 34:

[1024] The user checks the amount on the tablet device and makes the payment.

[1025] Step 35:

[1026] The terminal notifies the server that the payment has been completed.

[1027] Added emotion engine processing

[1028] Step 36:

[1029] The device (emotion engine) analyzes the user's facial expressions and voice to recognize their emotional state.

[1030] Step 37:

[1031] The device (generative AI) generates a flexible response based on the recognized emotion.

[1032] Step 38:

[1033] The device will respond to the user based on their emotions. For example, if it recognizes that the user is tired, it will respond in a gentle tone, saying, "I understand. I'll bring you the fries right away, so please wait a moment."

[1034] Step 39:

[1035] The terminal (emotion engine) analyzes the accumulated emotional data and generates feedback to improve the quality of the service.

[1036] Step 40:

[1037] The server receives the feedback and uses it to improve services across the system.

[1038] Example 2

[1039] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1040] In modern restaurants, improving the efficiency and flexibility of customer service is important. However, conventional systems have difficulty recognizing customer emotions and responding flexibly. Accurate order acceptance and rapid processing are also required, but there is a high possibility of human error. To solve these issues, a system is needed that automatically recognizes customer voices and provides appropriate responses and order processing.

[1041] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for analyzing the text data to identify the order contents, means for transmitting the order information to the server, means for controlling the robot based on instructions received from the server, means for responding to the user by voice, and means for recognizing the user's emotions and generating a response based on the same. This enables flexible responses that take into account the customer's emotions and accurate and prompt order processing.

[1042] The "means for acquiring the user's voice" is a combination of hardware and software for acquiring the voice uttered by the user.

[1043] The "means for converting voice into text data" refers to a voice recognition technology for converting acquired voice into digital text data.

[1044] The "means for analyzing text data to identify order details" refers to algorithms and software for analyzing text data and accurately identifying the order details intended by the user.

[1045] The "means for transmitting order information to the server" refers to a communication means and protocol for transmitting the specified order details to the server via a network.

[1046] The "means for controlling the robot based on instructions received from the server" refers to hardware and software for receiving instructions sent from the server and controlling the robot in accordance with those instructions.

[1047] "Means for providing audio responses to the user" refers to a speaker and voice generating software for playing audio messages to the user.

[1048] The "means for recognizing the user's emotions and generating responses based on them" refers to an emotion analysis engine and generative AI model that analyzes emotions from the user's speech and actions and generates responses that are adapted to those emotions.

[1049] This invention is a customer service system using a robot in a restaurant, and is equipped with a function to recognize the user's emotions and respond flexibly based on them. This system is designed to acquire and analyze the user's voice to identify the order, and further recognize the user's emotions and respond accordingly.

[1050] Hardware and Software Configuration

[1051] Terminal

[1052] Camera: A device that uses facial recognition technology to detect users entering the store.

[1053] Microphone: A device used to capture the user's voice.

[1054] Speaker: A device that provides audio responses to the user.

[1055] Speech recognition module: Software for converting speech into text data using the Google Cloud Speech-to-Text API, etc.

[1056] Generative AI: Software that uses tools such as GPT-4 to analyze text data and generate appropriate responses.

[1057] Emotion analysis engine: Software for recognizing user emotions using Microsoft Azure Emotion API, etc.

[1058] server

[1059] Communications module: A device that receives order information sent from the terminal and forwards it to the kitchen or bar counter.

[1060] Information management system: Software for managing order information and cooking status in real time.

[1061] Specific operation of the system

[1062] User recognition and conversation initiation

[1063] When a user enters a restaurant, the device's camera detects the user using facial recognition technology (e.g., OpenCV), and the microphone captures the user's speech. This allows the device to prepare for the initial interaction. When the user says "Hello," the speech recognition module converts the speech into text data. The converted text data is sent to a generative AI, which generates an appropriate greeting. The device's speaker plays the message, "Hello, welcome. We will show you to your seat."

[1064] Receiving and processing orders

[1065] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text data. The generation AI analyzes the text and determines the order. The order is then sent to a server, which forwards the information to the kitchen or bar counter and requests cooking and drink preparation. The generation AI generates a message confirming the order, and the device responds with a voice saying, "I understand. A hamburger and orange juice, please."

[1066] Cooking status management and serving

[1067] The server receives a notification from the kitchen that the food is ready and notifies the terminal. The terminal receives the serving instructions, moves to the kitchen, places the food on a tray, and delivers it to the table. The robot delivers the food and responds with a voice message saying, "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[1068] Reorder acceptance and emotion recognition

[1069] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[1070] Inquiry response and settlement

[1071] When a user asks, "What desserts do you have?", the device converts the voice into text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. Would you like that?" When the user says, "Please pay the bill, please," the device converts the voice into text, and the generation AI identifies the payment request and notifies the server. The server compiles all the order details, calculates the payment amount, and sends it to the device. The device then brings a tablet device to the user's table and says, "Please check on this device." The user checks the amount on the tablet device and makes the payment, and the device notifies the server that the payment is complete.

[1072] Prompt Sentence Examples

[1073] "What's the scenario when a user walks into a store?"

[1074] "Please explain how emotions are recognized when ordering more."

[1075] "Please tell me in detail what steps a user should take to request a checkout."

[1076] The above is a specific operation of the system according to the embodiment of the present invention. This system is expected to make customer service in restaurants more efficient and flexible, thereby improving customer satisfaction.

[1077] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1078] Step 1:

[1079] A user enters the store

[1080] Input: User entry, video data from camera, audio data from microphone

[1081] Operation: The device recognizes the user's face using a camera installed at the entrance, processes the video data to detect the user's presence, and captures audio data using a microphone.

[1082] Output: Recognition result that the user has entered the store

[1083] Step 2:

[1084] Recognize the user's first utterance

[1085] Input: User's speech

[1086] How it works: When a user says "hello," the device captures this audio through the microphone. The audio data is sent to the Google Cloud Speech-to-Text API and converted to text.

[1087] Output: Text data (e.g. "Hello")

[1088] Step 3:

[1089] Generative AI creates greeting messages

[1090] Input: Text data (e.g. "Hello")

[1091] How it works: A generative AI (e.g., GPT-4) analyzes text data and generates an appropriate greeting.

[1092] Output: Greeting message (e.g. "Hello, welcome. We will show you to your seat.")

[1093] Step 4:

[1094] Voice output of greetings

[1095] Input: Greeting message (e.g. "Hello, welcome. We'll show you to your seat.")

[1096] Behavior: The generated greeting message is played aloud through the device's speaker.

[1097] Output: A voice response to the user

[1098] Step 5:

[1099] Get the user's orders

[1100] Input: User's voice order

[1101] How it works: When a user says, "I'd like a hamburger and orange juice, please," the device captures this speech through the microphone. The speech data is then sent to the Google Cloud Speech-to-Text API, where it is converted into text data.

[1102] Output: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[1103] Step 6:

[1104] Order analysis

[1105] Input: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[1106] How it works: Generative AI analyzes text data and identifies the order details.

[1107] Output: Identification of the order (e.g. "Hamburger", "Orange juice")

[1108] Step 7:

[1109] Send the order to the server

[1110] Input: Identification of order details (e.g. "Hamburger", "Orange juice")

[1111] Operation: The terminal sends the order information to the server.

[1112] Output: Order information received by the server

[1113] Step 8:

[1114] Transferring order information to the kitchen

[1115] Input: Order information received by the server

[1116] How it works: The server forwards the order information to the kitchen or bar counter and instructs cooking and preparation.

[1117] Output: Instructions received by the kitchen or bar counter

[1118] Step 9:

[1119] Order confirmation

[1120] Input: The order details identified by the generation AI (e.g., "hamburger," "orange juice")

[1121] How it works: The AI ​​generates a confirmation message and responds audibly through the device's speaker: "Okay, a hamburger and orange juice, right?"

[1122] Output: Audible acknowledgment to the user

[1123] Step 10:

[1124] Receive notifications when food is cooked

[1125] Input: Notification from the kitchen that cooking is complete

[1126] Operation: The server receives a notification from the kitchen that the food is ready and notifies the terminal.

[1127] Output: Cooking completion notification to the device

[1128] Step 11:

[1129] Receiving and acting on serving instructions

[1130] Input: Delivery instructions from the server

[1131] How it works: The terminal receives serving instructions, and the robot moves to the kitchen, places the food on a tray, and delivers it to the table.

[1132] Output: Food served

[1133] Step 12:

[1134] Reorder acceptance and emotion recognition

[1135] Input: User's voice for reorder

[1136] How it works: When a user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[1137] Output: Response to the user about the reorder

[1138] Step 13:

[1139] Inquiry response

[1140] Input: User's spoken query

[1141] How it works: When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. What would you like?"

[1142] Output: Audio response to user queries

[1143] Step 14:

[1144] Acceptance of settlement requests

[1145] Input: User's voice request for payment

[1146] How it works: When a user says, "I'd like to pay," the device converts the speech into text, and the generation AI identifies the payment request and notifies the server.

[1147] Output: Settlement request notification to the server

[1148] Step 15:

[1149] Settlement process

[1150] Input: Settlement request notification to the server

[1151] How it works: The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet to the user's table and says, "Please check on this terminal."

[1152] Output: The settlement amount displayed to the user

[1153] Step 16:

[1154] Check and complete the payment

[1155] Input: User payment confirmation

[1156] How it works: When the user checks the amount on the tablet and makes the payment, the device notifies the server that the payment is complete.

[1157] Output: Notification of successful payment

[1158] The above is the flow of processing of the program of this system.

[1159] (Application example 2)

[1160] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1161] Conventional restaurant and food delivery systems lack the flexibility to consider user emotions, resulting in a uniform quality of customer experience and difficulty in improving satisfaction. Furthermore, appropriate responses that consider user emotions are required when notifying users of delivery status and collecting feedback after delivery. Therefore, a system that can provide attentive service based on emotion recognition, in addition to voice recognition, is required, but no such system currently exists.

[1162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1163] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order contents;] [means for sending order information to the server;] [means for controlling the robot by receiving instructions from the server;] [means for responding to the user by voice;] [means for using generative artificial intelligence to analyze the user's emotions and generate a response based on those emotions;] [means for notifying the user of the delivery status in real time; and [means for collecting post-delivery evaluations based on the user's emotions.] This enables flexible and efficient service provision based on the user's emotions and a consistent improvement in the customer experience during and after delivery.

[1164] "Means for inputting user voice" refers to technology for inputting user speech using a microphone or similar device in restaurants or food delivery systems.

[1165] The "means for converting speech to text" is a speech recognition technology that converts a user's speech data into character string data.

[1166] The "means for analyzing text to identify order details" is a technology for analyzing text data that has been speech-recognized to identify and specify the user's order details.

[1167] The "means for transmitting order information to a server" is a technique for transmitting user order information to a central server via a communication network.

[1168] "Means for controlling a robot by receiving instructions from a server" refers to a control technology that receives instructions from a server, operates the robot, and executes the specified task.

[1169] "Means for responding to the user by voice" refers to a technology for generating an appropriate voice response to the user's speech and transmitting it to the user through a speaker.

[1170] "Means using generative artificial intelligence to analyze a user's emotions and generate responses based on those emotions" refers to artificial intelligence technology that recognizes emotions from the user's voice and automatically generates natural responses based on those emotions.

[1171] "Means for notifying the user of the status during delivery in real time" refers to technology that notifies the user of the current status, estimated arrival time, etc. in real time during food delivery.

[1172] The "means for collecting post-delivery evaluations based on user emotions" is a technology for collecting delivery evaluation feedback based on user emotions after delivery is completed.

[1173] The present invention provides a system for restaurants and food delivery systems that provides flexible responses that take into account user emotions. This system has the following main functions:

[1174] Basic configuration

[1175] The system includes a means for inputting the user's voice, a means for converting the voice into text, a means for analyzing the text to identify the order contents, a means for sending the order information to a server, a means for receiving instructions from the server to control the robot, and a means for responding to the user by voice.

[1176] Additionally, it also includes generative artificial intelligence that analyzes the user's emotions and generates responses based on those emotions, a means of notifying the user of the delivery status in real time, and a means of collecting post-delivery evaluations based on the user's emotions.

[1177] Hardware and Software

[1178] Hardware: Uses the computer's built-in camera, microphone, robot, and speaker. These devices are used to input the user's voice and image, and to output response voices and delivery information.

[1179] software:

[1180] Speech Recognition Module: Uses the speech_recognition library, which converts the user's speech into text data.

[1181] Facial Recognition Software: Recognizes the user's face using the OpenCV library.

[1182] Generative AI: We use emotion recognition and text generation models from the transformers library, specifically the 'michellejieli / emotion_text' and 'gpt-3.5-turbo' models.

[1183] User voice recognition and emotion analysis

[1184] The device uses a camera and microphone to recognize the user, detects the user through facial recognition, and converts the user's speech into text using a speech recognition module. This text data is then input into an emotion recognition model to identify the user's emotions.

[1185] Response Generation

[1186] The generative AI generates an appropriate response based on the user's emotions. For example, if the user's voice is recognized as saying "hello" and the voice contains the emotion of "welcome," the generative AI model will generate a response such as "Welcome, please place your order." This response is then conveyed to the user through the speaker.

[1187] Examples and prompts

[1188] Examples:

[1189] Scenario: A user launches your app and says, "Good evening, I'd like a pizza and a Coke, please."

[1190] Emotion recognition: The emotion of "friendliness" is recognized from the user's voice.

[1191] Output Response: "Thank you for your order. Pizza and Coke, please. Now, please tell me your delivery address."

[1192] Example prompt sentence:

[1193] "Respond in a friendly tone: Welcome, your order is welcome."

[1194] "Please proceed to the next step to confirm your delivery address."

[1195] This will enable flexible and efficient service provision based on user emotions and improve the consistent customer experience during and after delivery.

[1196] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1197] Step 1:

[1198] The device recognizes the user using a camera and microphone. The camera detects the user's face and uses facial recognition software (OpenCV) to confirm the user's presence. At the same time, the microphone collects the user's speech and inputs it as audio data. The output is the user's facial image data and audio data.

[1199] Step 2:

[1200] The device sends the collected voice data to a voice recognition module (speech_recognition library), which converts the voice data into text data. The input is the user's voice data, and the output is text data. Specifically, the voice recognition module analyzes the voice waveform and generates a corresponding string of characters.

[1201] Step 3:

[1202] The server receives the text data and uses generative artificial intelligence (the transformers library) to analyze the spoken text and identify the user's emotions. The input is text data generated by speech recognition, and the output is the recognized emotion label. Specifically, the emotion recognition model analyzes the text and determines emotions such as "welcoming" or "friendliness."

[1203] Step 4:

[1204] The server uses generative artificial intelligence to generate an appropriate response based on the identified emotion. The input is text data containing the user's emotion label and order details, and the output is a response text. Specifically, the response generation model forms a natural-sounding answer while taking the emotion label into account.

[1205] Step 5:

[1206] The server sends the generated response text to the terminal, and the terminal responds to the user audibly through the speaker. The input is the response text sent from the server, and the output is the audio response to the user. Specifically, the text is converted into audio and played back so that the user can hear it.

[1207] Step 6:

[1208] The terminal sends the user's order information to the server, which then transmits the order information to the cooking station or delivery station. The input is the order information written in text, and the output is an execution instruction sent to the cooking station or delivery station. Specifically, the order information is sent to each station via the network.

[1209] Step 7:

[1210] The server receives delivery status information from the delivery station in real time and notifies the user of the delivery status. The input is status report data from the delivery station, and the output is delivery progress information displayed to the user. Specifically, the status data is notified to the user's smartphone or other device.

[1211] Step 8:

[1212] When the delivery is completed, the terminal collects the user's evaluation based on their emotions and sends it to the server. The input is the user's voice feedback, and the output is the evaluation data sent to the server. Specifically, the voice feedback is converted into text and saved along with the user's emotional evaluation.

[1213] Step 9:

[1214] The server analyzes the collected evaluation data to improve service quality. The input is the collected evaluation data, and the output is reports and statistical information for improvement. Specifically, the evaluation data in the database is analyzed to identify areas for improvement.

[1215] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1216] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1217] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1218] [Third embodiment]

[1219] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1220] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1221] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1222] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1223] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1225] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1226] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1227] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1228] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1229] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1230] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1231] The present invention is a system for providing efficient and flexible service using robots in restaurants. Specific embodiments for carrying out the invention and the processing of the program therefor will be described below.

[1232] User recognition and conversation initiation

[1233] User entry recognition

[1234] The device uses a camera and microphone to recognize the user. When a user enters a store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech. Through facial recognition algorithms and voice input, it is possible to quickly respond to users even when meeting them for the first time.

[1235] Initial voice recognition and greeting

[1236] When a user says "hello," the device sends the speech to a speech recognition module, which converts the speech to text. The generative AI analyzes the text and generates an appropriate greeting. The device responds through the speaker, saying, "Hello, welcome. We'll show you to your seat."

[1237] Receiving orders

[1238] Voice recognition for ordering

[1239] If a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which the generation AI then analyzes to determine the order.

[1240] Communicating with the Server

[1241] The terminal sends the specified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin promptly.

[1242] Acknowledgement and Response

[1243] The AI ​​generates a message to confirm the order, and the device responds aloud, saying, "Understood. A hamburger and orange juice, please."

[1244] Cooking status management and serving

[1245] Cooking completion notification

[1246] The server receives a notification from the kitchen that the food is ready and sends that information to the terminal, allowing the robot to begin serving the food.

[1247] Serving instructions and actions

[1248] The server sends serving instructions to the terminal, which then follows those instructions to move to the kitchen and place the food on a tray. The robot then safely carries the food to the table and serves it to the user, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[1249] Continuing the conversation and being flexible

[1250] Accepting additional orders

[1251] If the user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. The device responds, "Okay, I'll order some extra fries," and sends the additional order information to the server.

[1252] Inquiry response

[1253] When a user asks, "What desserts do you have?", the device converts the voice into text, and the AI ​​analyzes the query and generates an appropriate answer. The device responds aloud, "Today's desserts include cake and ice cream."

[1254] settlement

[1255] Acceptance of settlement requests

[1256] When a user says, "I'd like to pay, please," the device converts the speech into text using a voice recognition module, and the generation AI identifies the payment request and notifies the server.

[1257] Settlement process

[1258] The server compiles all the order details, calculates the payment amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and says, "Please check on this device."

[1259] Check and complete the payment

[1260] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed, and the settlement process is complete.

[1261] As described above, the present invention provides a system that efficiently and flexibly performs user recognition, order taking, serving, conversation, and payment processing, thereby optimizing service in restaurants.

[1262] The processing flow will be explained below.

[1263] Step 1:

[1264] The user enters the store.

[1265] Step 2:

[1266] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[1267] Step 3:

[1268] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[1269] Step 4:

[1270] The user says "Hello."

[1271] Step 5:

[1272] The terminal (voice recognition module) converts the user's speech into text.

[1273] Step 6:

[1274] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[1275] Step 7:

[1276] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[1277] Step 8:

[1278] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[1279] Step 9:

[1280] The terminal (voice recognition module) converts the user's speech into text.

[1281] Step 10:

[1282] The terminal (generative AI) analyzes the text and identifies the order contents.

[1283] Step 11:

[1284] The terminal sends the order details to the server.

[1285] Step 12:

[1286] The server receives the order information and passes it on to the kitchen and bar counter.

[1287] Step 13:

[1288] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[1289] Step 14:

[1290] The terminal will audibly convey a confirmation message to the user.

[1291] Step 15:

[1292] The server receives a notification from the kitchen that the food is ready.

[1293] Step 16:

[1294] The server sends a cooking completion notification to the terminal.

[1295] Step 17:

[1296] The device receives a notification, goes to the kitchen and places the food on a tray.

[1297] Step 18:

[1298] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[1299] Step 19:

[1300] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[1301] Step 20:

[1302] The terminal (voice recognition module) converts the user's speech into text.

[1303] Step 21:

[1304] The terminal (generative AI) analyzes the text and identifies the additional order details.

[1305] Step 22:

[1306] The terminal transmits the additional order details to the server.

[1307] Step 23:

[1308] The terminal responds, "Okay, we'll add fries."

[1309] Step 24:

[1310] The user asks about dessert. Say, "What desserts do you have?"

[1311] Step 25:

[1312] The terminal (voice recognition module) converts the user's speech into text.

[1313] Step 26:

[1314] The terminal (generative AI) analyzes the text and generates a dessert menu.

[1315] Step 27:

[1316] The device will respond aloud, "Today's desserts include cake and ice cream."

[1317] Step 28:

[1318] The user wishes to settle the bill and says, "Please pay."

[1319] Step 29:

[1320] The terminal (voice recognition module) converts the user's speech into text.

[1321] Step 30:

[1322] The terminal (generator AI) identifies the settlement request and notifies the server.

[1323] Step 31:

[1324] The server compiles all the order details and calculates the total amount.

[1325] Step 32:

[1326] The server sends the settlement information to the terminal.

[1327] Step 33:

[1328] The device brings a tablet device to the user's table and says, "Please check this device."

[1329] Step 34:

[1330] The user checks the amount on the tablet device and makes the payment.

[1331] Step 35:

[1332] The terminal notifies the server that the payment has been completed.

[1333] Example 1

[1334] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1335] Providing efficient and flexible service is a key issue for modern restaurants. In particular, there is a demand for systems that automate and seamlessly perform all processes, from recognizing users as they enter the restaurant to taking their orders, serving food, handling additional orders, and processing payments. Current systems often only provide individual functions and are insufficient for providing comprehensive services. Furthermore, delays in the recognition of specific users and the timing of voice responses can lead to problems that reduce user satisfaction.

[1336] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1337] In this invention, the server includes means for inputting user voice, means for converting voice to text, means for analyzing the text to identify the order details, means for transmitting the order information to the server, means for receiving instructions from the server and controlling the robot, means for responding to the user by voice, means for detecting the user using facial recognition technology, and means for receiving notification of cooking completion. This enables the consistent automation of everything from recognizing the user's entry to receiving the order, serving the food, handling additional orders, and processing the payment, enabling the provision of efficient and flexible services.

[1338] The "means for inputting user's voice" refers to a device that acquires the user's speech using a voice input device such as a microphone.

[1339] A "means for converting speech to text" is a device or software that converts speech data acquired using speech recognition technology into text data.

[1340] The "means for analyzing text to identify order details" refers to software or an algorithm for analyzing the acquired character data and identifying the order details from the content of that data.

[1341] The "means for transmitting order information to the server" refers to a device or software that has the function of transmitting the specified order details to the server via a communication means.

[1342] The "means for receiving instructions from the server and controlling the robot" refers to a device or software for receiving instructions from the server and controlling the operation of the robot based on those instructions.

[1343] The "means for responding to the user by voice" refers to a device or software for providing the generated response message to the user by voice through a voice synthesizer.

[1344] A "means for detecting a user using facial recognition technology" is a device or software that uses a camera and a facial recognition algorithm to detect and identify a user's face.

[1345] The "means for receiving notification of cooking completion" refers to a device or software for receiving notification of cooking completion from the kitchen and sharing that information within the system.

[1346] The present invention provides a system for providing efficient and flexible service using robots in restaurants. Specific processing steps for carrying out the present invention will be described in natural language.

[1347] This system uses a microphone to input the user's voice and speech recognition technology to convert the voice into text. This is done using the Google Speech-to-Text API, among others. It also uses generative artificial intelligence (generative AI model) to analyze the text and identify the order contents. This analysis utilizes OpenAI's GPT-4, among others.

[1348] A device with communication capabilities is used to send order information to the server. This sends the order details to the server (e.g., Dell PowerEdge T30). The robot operates based on a communication protocol to receive instructions from the server and control the robot. A typical service robot (e.g., SoftBank's Pepper) is used as the robot.

[1349] To respond to the user via voice, a generated response message is provided to the user using a speech synthesizer (e.g., Amazon Polly).Furthermore, to detect the user using facial recognition technology, a camera (e.g., Logitech HD Pro Webcam C920) and a facial recognition algorithm (e.g., OpenCV library) are used.To receive a notification that cooking is complete, a communication function is used to receive notifications from the kitchen and share that information with the device.

[1350] As a concrete example, let's consider a scenario where a user enters a restaurant. When the user enters the restaurant, the camera recognizes their face and the microphone picks up the utterance "Hello." The device converts the utterance into text through a voice recognition module, and the generation AI analyzes it and responds with "Hello, welcome. We will show you to your seat."

[1351] Here's an example of ordering: When a user says, "I'd like a hamburger and orange juice, please," the device converts the speech into text, and the generation AI identifies the order. This information is sent to the server, which then relays the order to the kitchen or bar counter. After confirming the order, the device responds, "I understand. A hamburger and orange juice, please."

[1352] If the user wishes to place an additional order, for example by saying, "I'd like some extra fries, please," the device analyzes the voice and uses a generative AI to identify the additional order. The additional order information is then sent to the server. After the cooking is complete, the robot brings the food from the kitchen and serves it to the user. It responds, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[1353] Some examples of prompts are:

[1354] "How do you greet users when they walk in?"

[1355] "When a user places an order, how do you verify the order and provide cooking instructions?"

[1356] "How do you transport the food once it's done cooking?"

[1357] "How do you handle additional orders or inquiries from users?"

[1358] "How do you guide users through the process at checkout?"

[1359] In this way, the system is able to provide consistently efficient and flexible service, optimizing service in restaurants.

[1360] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1361] Specific explanation of processing steps

[1362] Step 1: User entry recognition

[1363] Specifically: The device uses a camera and microphone to recognize the user. The camera (Logitech HD Pro Webcam C920) detects the user's face, and the microphone (Shure MV5) picks up the user's speech.

[1364] Input: Camera image and audio data

[1365] Output: Face detection data and audio clips

[1366] Data processing / calculation: Detect faces from camera images (using OpenCV library) and save audio clips

[1367] Specific operation: When a user enters the store, the camera recognizes their face and confirms what they are saying through voice input.

[1368] Step 2: First voice recognition and greeting

[1369] Specifically, the device sends the user's speech to a speech recognition module (Google Speech-to-Text API) and converts it into text data. A generative AI (OpenAI's GPT-4) analyzes the text and generates an appropriate greeting message.

[1370] Input: Audio clip (user speech)

[1371] Output: Character data and response message

[1372] Data processing / calculation: Converting voice data to text data (Speech-to-Text), character analysis and response generation (generative AI)

[1373] Specific operation: When the user says "Hello," the device responds with "Hello, welcome. We will show you to your seat."

[1374] Step 3: Voice recognition of orders

[1375] Specific explanation: When a user speaks an order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the order contents.

[1376] Input: Audio clip (order details)

[1377] Output: Text data (order details) and order information

[1378] Data processing / calculation: Converting voice data into text data and identifying order details through text analysis

[1379] Specific operation: The user says, "I'd like a hamburger and orange juice, please," and this is identified as the order information.

[1380] Step 4: Communicating with the Server

[1381] Specific explanation: The terminal sends the identified order details to the server, which then transmits the order information to the kitchen or bar counter.

[1382] Input: Order Information

[1383] Output: Instructions to the kitchen and bar counter

[1384] Data processing / calculation: Sends order information to the server and transfers cooking instructions to each department

[1385] Specific operation: The terminal sends an order for "hamburger and orange juice" to the server, which then distributes it to the kitchen and bar counter.

[1386] Step 5: Order confirmation and response

[1387] Specific explanation: The generation AI generates a message to confirm the order details, and the terminal responds with a voice message saying, "Okay, a hamburger and orange juice, right?"

[1388] Input: Order Information

[1389] Output: Confirmation message

[1390] Data processing / calculation: Checking order information and generating messages

[1391] Specific operation: The generation AI creates a confirmation message, which the device then audibly conveys to the user.

[1392] Step 6: Receive notification that your food is ready

[1393] Specific explanation: The server receives a cooking completion notification from the kitchen and sends that information to the terminal.

[1394] Input: Cooking complete notification

[1395] Output: Cooking completion information

[1396] Data processing / calculation: Receiving and transmitting cooking completion notification

[1397] Specific operation: The kitchen notifies the server that the food is ready, and the server sends that information to the device.

[1398] Step 7: Food serving instructions and actions

[1399] Details: The server sends serving instructions to the terminal, which then moves to the kitchen according to the instructions and places the food on a tray. The robot then carries it to the table and serves it to the user.

[1400] Input: Cooking completion information

[1401] Output: Food served

[1402] Data processing / calculation: Generation of serving instructions and robot motion control

[1403] What it does: The device moves to the kitchen and places the food on a tray, which the robot then carries to the table.

[1404] Step 8: Accepting additional orders

[1405] Specific explanation: When a user speaks an additional order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the additional order.

[1406] Input: Audio Clip (reorder)

[1407] Output: Text data (reorder) and reorder information

[1408] Data processing / calculation: Converting voice data into text data and identifying additional orders

[1409] Specific operation: When the user says, "I'd like some extra fries, please," the device analyzes it and sends it to the server.

[1410] Step 9: Handling inquiries

[1411] Specific explanation: When a user speaks a question, the device converts the voice into text data, which is then analyzed by the generation AI to generate an appropriate answer.

[1412] Input: Audio clip (user question)

[1413] Output: Text data (question content) and answer message

[1414] Data processing / calculation: Converting voice data into text data, analyzing questions, and generating answers

[1415] Specific behavior: When the user asks, "What desserts do you have?", the device responds, "Today's desserts include cake and ice cream."

[1416] Step 10: Accepting the settlement request

[1417] Specific explanation: When a user says, "I'd like to pay the bill, please," the device converts the voice into text data, and the generation AI identifies the payment request and notifies the server.

[1418] Input: Audio clip (payment request)

[1419] Output: Text data (settlement request) and settlement notice

[1420] Data processing / calculation: Converting voice data into text data, identifying and notifying payment requests

[1421] Specific operation: When the user says "I'd like to pay the bill," the terminal notifies the server.

[1422] Step 11: Checkout

[1423] Specifically: The server compiles all the order details, calculates the total amount, and sends it to the terminal, which then brings a tablet to the user's table and displays the total amount.

[1424] Input: Order details

[1425] Output: Settlement amount

[1426] Data processing / calculation: Aggregating order details and calculating payment amounts

[1427] Specific operation: The server performs the calculation and the terminal displays the amount on the tablet.

[1428] Step 12: Check and complete the payment

[1429] Specific explanation: The user checks the amount on the tablet and makes the payment. The terminal notifies the server that the payment has been completed, and the settlement process is complete.

[1430] Input: Payment Information

[1431] Output: Payment completion notification

[1432] Data processing / calculation: Confirmation of payment information and notification of settlement completion

[1433] Specific operation: The user checks the amount on the tablet, makes the payment, and the device notifies the server.

[1434] In this way, the system can efficiently automate restaurant operations through a series of processes and provide users with high quality service.

[1435] (Application example 1)

[1436] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1437] Conventional food delivery systems have the problem that the entire process from ordering to delivery and payment is divided into manual processes and multiple different systems, resulting in a lack of efficiency and flexibility. The present invention aims to solve these problems, improve the efficiency of services for users, and provide a system that unifies and smoothly processes each stage of ordering, delivery, and payment.

[1438] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1439] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order details;] [means for sending order information to the server;] [means for controlling the robot upon receiving instructions from the server;] [means for responding to the user by voice;] [means for recognizing the user using facial recognition technology;] [means for sending delivery instructions to the delivery robot after accepting the order;] [means for responding to payment requests and calculating the payment amount; and [means for completing the user's payment.] This makes it possible to carry out the entire process from ordering to delivery and payment in a unified and efficient manner.

[1440] The "means for inputting user's voice" refers to a means for acquiring voice uttered by the user through a device.

[1441] The "means for converting voice to text" is a means for analyzing acquired voice data and converting it into text data.

[1442] The "means for analyzing text to identify order details" refers to a means for analyzing text data and accurately identifying the user's intentions and order details.

[1443] The "means for transmitting order information to the server" is a means for transmitting the analyzed order details to the server via a network.

[1444] The "means for receiving instructions from the server and controlling the robot" refers to means for receiving instruction data from the server and controlling the robot's operation based on that data.

[1445] The "means for responding to the user by voice" refers to a means for generating voice in response to an input from the user and responding to the user.

[1446] "Means for recognizing a user using facial recognition technology" refers to technology that uses a camera or the like to identify and specify the user's face.

[1447] The "means for sending delivery instructions to a delivery robot after receiving an order" refers to a means for transmitting the order details to a delivery robot and issuing delivery instructions after receiving a user's order.

[1448] The "means for responding to a payment request and calculating the payment amount" is a means for receiving a payment request from a user and calculating the payment amount based on all order details.

[1449] The "means for completing the user's settlement" is the means by which the user confirms the amount presented and completes the payment procedure.

[1450] This invention is a system that automates efficient and flexible robot-based food delivery services. This system automates a series of processes, from user voice input, speech-to-text conversion, order analysis, order information transmission to a server, robot control, voice response, user identification using facial recognition technology, delivery instructions to the delivery robot, payment request handling, and payment completion.

[1451] Hardware and Software Configuration

[1452] This system mainly consists of the following hardware and software:

[1453] Hardware: smartphones, delivery robots, servers

[1454] Software: Speech recognition module (e.g., Google Speech-to-Text API), face recognition module (e.g., OpenCV), generative AI model (e.g., OpenAI GPT-4), Python, Django framework, Celery

[1455] Program processing

[1456] The main processing of the entire system will be explained.

[1457] 1. Enter the user's voice:

[1458] The server and the terminal use a microphone to input the user's voice, which is then converted into text data by a voice recognition module.

[1459] 2. Convert speech to text:

[1460] The server uses a speech recognition module to convert the voice data into text.

[1461] 3. Parse the text to identify the order:

[1462] The terminal uses a generative AI model (e.g., GPT-4) to analyze the converted text and identify the order details.

[1463] 4. Send the order information to the server:

[1464] The terminal transmits the parsed order details to the server.

[1465] 5. Control the robot by receiving instructions from the server:

[1466] The server sends the order details to the delivery robot and controls the robot to pick up and deliver the order.

[1467] 6. Respond to the user verbally:

[1468] The device uses the generative AI model to generate a response message and conveys it to the user via voice output.

[1469] 7. Recognize users using facial recognition technology:

[1470] The device uses a camera to perform facial recognition and identify the user.

[1471] 8. After receiving the order, send delivery instructions to the delivery robot:

[1472] After receiving the order, the server issues delivery instructions to the delivery robot.

[1473] 9. Respond to settlement requests and calculate settlement amounts:

[1474] The server receives the user's payment request, calculates all the order details, and calculates the payment amount.

[1475] 10. Complete the user's checkout:

[1476] The terminal presents the payment amount to the user and completes the payment.

[1477] Specific examples

[1478] For example, when a user says "Pizza and Coke please" on a smartphone, the following process is executed:

[1479] 1. Speech recognition: The speech is converted into text, and the text data "Pizza and Coke, please" is generated.

[1480] 2. Order Analysis: A generative AI model analyzes the text data to identify an order for a pizza and a cola.

[1481] 3. Robot control: The server sends instructions to the delivery robot to pick up the pizza and cola from the store and deliver it to the user's address.

[1482] 4. Voice response: "Thank you for your order. Your pizza and Coke will be on their way soon," the device informs the user.

[1483] 5. Settlement: When the user says, "Please pay the bill," the server calculates the amount and the payment is completed on the settlement screen displayed on the smartphone.

[1484] As an example of input to the generative AI model, we will use the following prompt:

[1485] "A user asks, 'What's on the menu today?' What are your menu recommendations?"

[1486] By inputting the above prompt sentences into the generative AI model, an appropriate response can be obtained. This process enables efficient and flexible food delivery services.

[1487] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1488] Step 1:

[1489] The user inputs voice into the smartphone.

[1490] The user speaks into the smartphone microphone, "Pizza and Coke please." The input data is the user's voice data. The device acquires this voice data and passes it on to the next step.

[1491] Step 2:

[1492] The device converts the voice data into text data.

[1493] The device uses a speech recognition module (e.g., Google Speech-to-Text API) to convert the voice data into text data. Here, the input is voice data, and the output is text data such as "Pizza and Coke, please."

[1494] Step 3:

[1495] The terminal analyzes the text data and identifies the order details

[1496] The device uses a generative AI model (e.g., GPT-4) to analyze the text data. Specifically, it extracts the order details of "pizza" and "cola." The input is text data, and the output is the identified order details.

[1497] Step 4:

[1498] The terminal sends the order information to the server

[1499] The terminal sends the specified order details to the server, where the input is the order details data and the output is the order information sent to the server.

[1500] Step 5:

[1501] The server receives the order information and sends delivery instructions to the delivery robot.

[1502] Based on the received order information, the server sends an instruction to the delivery robot to "deliver pizza and cola to the specified address." The input is the order information, and the output is the delivery instruction data.

[1503] Step 6:

[1504] Delivery robots will pick up orders from stores and deliver them to designated addresses.

[1505] The delivery robot receives instructions from the server, goes to the store, picks up the pizza and cola, and delivers it to the specified user's address. The input is delivery instruction data, and the output is a delivery completion notification.

[1506] Step 7:

[1507] The device notifies the user of the delivery progress

[1508] The device uses a generative AI model to generate a message about the delivery progress, such as "Your pizza and coke will arrive soon," and notify the user via voice. The input is the delivery progress data, and the output is the voice message.

[1509] Step 8:

[1510] The user makes a payment request

[1511] The user verbally requests the terminal, "Please pay the bill." The input is voice data, which the terminal recognizes and proceeds to the next step.

[1512] Step 9:

[1513] The terminal converts the voice data into text data and sends a payment request to the server.

[1514] The terminal converts the voice data into text data using a voice recognition module and sends this text data to the server. The input is the voice data, and the output is the text data and payment request data sent to the server.

[1515] Step 10:

[1516] The server calculates the settlement amount and sends it to the terminal.

[1517] The server calculates the total amount based on all the order details and sends the total amount data to the terminal. The input is the order information and the output is the total amount data.

[1518] Step 11:

[1519] The terminal presents the payment amount to the user and completes the payment procedure.

[1520] The terminal uses the generative AI model to generate a message saying, "The settlement amount is XX yen," and presents it to the user via voice. The settlement procedure is then completed after the user performs the payment operation. The input is the settlement amount data, and the output is a payment completion notification.

[1521] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1522] The present invention is a customer service system using a robot in a restaurant, which is particularly equipped with the function of recognizing the user's emotions and responding flexibly based on them. This system has the following main functions:

[1523] User recognition and conversation initiation

[1524] User entry recognition

[1525] The device uses a camera and microphone to recognize the user. When a user enters the store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech, allowing the device to prepare for the initial interaction.

[1526] Initial voice recognition and greeting

[1527] When a user says "hello," the device sends the speech to a speech recognition module. The speech is converted into text data, and a generative AI analyzes the text to generate an appropriate greeting. The message "Hello, welcome. We will show you to your seat" is played over the device's speaker.

[1528] Receiving orders

[1529] Voice recognition for ordering

[1530] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which is then analyzed by a generation AI to determine the order.

[1531] Communicating with the Server

[1532] The terminal sends the identified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin quickly.

[1533] Acknowledgement and Response

[1534] The AI ​​generates a message confirming the order details, and the device responds to the user verbally, saying, "I understand. A hamburger and orange juice, right?"

[1535] Cooking status management and serving

[1536] Cooking completion notification

[1537] The server receives a cooking completion notification from the kitchen, then sends the cooking completion notification to the terminal, allowing the robot to begin serving the food.

[1538] Serving instructions and actions

[1539] The terminal receives the serving instructions, moves to the kitchen, and places the food on a tray. The robot then safely delivers the food to the table, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy."

[1540] Continuing the conversation and being flexible

[1541] Reorder acceptance and emotion recognition

[1542] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion engine recognizes the user's emotions from their voice and reflects them in the response. For example, if the emotion engine detects fatigue in the user's voice, the device will respond in a gentle tone, saying, "I understand. I'll bring the fries right away, so please wait a moment."

[1543] Inquiry response

[1544] When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion engine analyzes the user's emotions, and the device responds to the user in a gentle tone, saying, "Today's desserts include cake and ice cream. What would you like?"

[1545] settlement

[1546] Acceptance of settlement requests

[1547] When a user says, "Please give me the bill," the device converts the speech into text using a speech recognition module. The AI ​​generator identifies the payment request and notifies the server.

[1548] Settlement process

[1549] The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and tells them to "please check on this device."

[1550] Check and complete the payment

[1551] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed.

[1552] The system, which includes the above functions, can recognize users' emotions and provide flexible and efficient services based on those emotions, significantly improving the customer experience in restaurants.

[1553] The processing flow will be explained below.

[1554] Step 1:

[1555] The user enters the store.

[1556] Step 2:

[1557] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[1558] Step 3:

[1559] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[1560] Step 4:

[1561] The user says "Hello."

[1562] Step 5:

[1563] The terminal (voice recognition module) converts the user's speech into text.

[1564] Step 6:

[1565] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[1566] Step 7:

[1567] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[1568] Step 8:

[1569] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[1570] Step 9:

[1571] The terminal (voice recognition module) converts the user's speech into text.

[1572] Step 10:

[1573] The terminal (generative AI) analyzes the text and identifies the order contents.

[1574] Step 11:

[1575] The terminal transmits the specified order details to the server.

[1576] Step 12:

[1577] The server receives the order information and passes it on to the kitchen and bar counter.

[1578] Step 13:

[1579] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[1580] Step 14:

[1581] The terminal will audibly convey a confirmation message to the user.

[1582] Step 15:

[1583] The server receives a notification from the kitchen that the food is ready.

[1584] Step 16:

[1585] The server sends a cooking completion notification to the terminal.

[1586] Step 17:

[1587] The device receives a notification, goes to the kitchen and places the food on a tray.

[1588] Step 18:

[1589] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[1590] Step 19:

[1591] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[1592] Step 20:

[1593] The terminal (voice recognition module) converts the user's speech into text.

[1594] Step 21:

[1595] The terminal (generative AI) analyzes the text and identifies the additional order details.

[1596] Step 22:

[1597] The terminal transmits the additional order details to the server.

[1598] Step 23:

[1599] The terminal responds, "Okay, we'll add fries."

[1600] Step 24:

[1601] The user asks about dessert. Say, "What desserts do you have?"

[1602] Step 25:

[1603] The terminal (voice recognition module) converts the user's speech into text.

[1604] Step 26:

[1605] The terminal (generative AI) analyzes the text and generates a dessert menu.

[1606] Step 27:

[1607] The device will respond aloud, "Today's desserts include cake and ice cream."

[1608] Step 28:

[1609] The user wishes to settle the bill and says, "Please pay."

[1610] Step 29:

[1611] The terminal (voice recognition module) converts the user's speech into text.

[1612] Step 30:

[1613] The terminal (generator AI) identifies the settlement request and notifies the server.

[1614] Step 31:

[1615] The server compiles all the order details and calculates the total amount.

[1616] Step 32:

[1617] The server sends the settlement information to the terminal.

[1618] Step 33:

[1619] The device brings a tablet device to the user's table and says, "Please check this device."

[1620] Step 34:

[1621] The user checks the amount on the tablet device and makes the payment.

[1622] Step 35:

[1623] The terminal notifies the server that the payment has been completed.

[1624] Added emotion engine processing

[1625] Step 36:

[1626] The device (emotion engine) analyzes the user's facial expressions and voice to recognize their emotional state.

[1627] Step 37:

[1628] The device (generative AI) generates a flexible response based on the recognized emotion.

[1629] Step 38:

[1630] The device will respond to the user based on their emotions. For example, if it recognizes that the user is tired, it will respond in a gentle tone, saying, "I understand. I'll bring you the fries right away, so please wait a moment."

[1631] Step 39:

[1632] The terminal (emotion engine) analyzes the accumulated emotional data and generates feedback to improve the quality of the service.

[1633] Step 40:

[1634] The server receives the feedback and uses it to improve services across the system.

[1635] Example 2

[1636] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1637] In modern restaurants, improving the efficiency and flexibility of customer service is important. However, conventional systems have difficulty recognizing customer emotions and responding flexibly. Accurate order acceptance and rapid processing are also required, but there is a high possibility of human error. To solve these issues, a system is needed that automatically recognizes customer voices and provides appropriate responses and order processing.

[1638] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for analyzing the text data to identify the order contents, means for transmitting the order information to the server, means for controlling the robot based on instructions received from the server, means for responding to the user by voice, and means for recognizing the user's emotions and generating a response based on the same. This enables flexible responses that take into account the customer's emotions and accurate and prompt order processing.

[1639] The "means for acquiring the user's voice" is a combination of hardware and software for acquiring the voice uttered by the user.

[1640] The "means for converting voice into text data" refers to a voice recognition technology for converting acquired voice into digital text data.

[1641] The "means for analyzing text data to identify order details" refers to algorithms and software for analyzing text data and accurately identifying the order details intended by the user.

[1642] The "means for transmitting order information to the server" refers to a communication means and protocol for transmitting the specified order details to the server via a network.

[1643] The "means for controlling the robot based on instructions received from the server" refers to hardware and software for receiving instructions sent from the server and controlling the robot in accordance with those instructions.

[1644] "Means for providing audio responses to the user" refers to a speaker and voice generating software for playing audio messages to the user.

[1645] The "means for recognizing the user's emotions and generating responses based on them" refers to an emotion analysis engine and generative AI model that analyzes emotions from the user's speech and actions and generates responses that are adapted to those emotions.

[1646] This invention is a customer service system using a robot in a restaurant, and is equipped with a function to recognize the user's emotions and respond flexibly based on them. This system is designed to acquire and analyze the user's voice to identify the order, and further recognize the user's emotions and respond accordingly.

[1647] Hardware and Software Configuration

[1648] Terminal

[1649] Camera: A device that uses facial recognition technology to detect users entering the store.

[1650] Microphone: A device used to capture the user's voice.

[1651] Speaker: A device that provides audio responses to the user.

[1652] Speech recognition module: Software for converting speech into text data using the Google Cloud Speech-to-Text API, etc.

[1653] Generative AI: Software that uses tools such as GPT-4 to analyze text data and generate appropriate responses.

[1654] Emotion analysis engine: Software for recognizing user emotions using Microsoft Azure Emotion API, etc.

[1655] server

[1656] Communications module: A device that receives order information sent from the terminal and forwards it to the kitchen or bar counter.

[1657] Information management system: Software for managing order information and cooking status in real time.

[1658] Specific operation of the system

[1659] User recognition and conversation initiation

[1660] When a user enters a restaurant, the device's camera detects the user using facial recognition technology (e.g., OpenCV), and the microphone captures the user's speech. This allows the device to prepare for the initial interaction. When the user says "Hello," the speech recognition module converts the speech into text data. The converted text data is sent to a generative AI, which generates an appropriate greeting. The device's speaker plays the message, "Hello, welcome. We will show you to your seat."

[1661] Receiving and processing orders

[1662] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text data. The generation AI analyzes the text and determines the order. The order is then sent to a server, which forwards the information to the kitchen or bar counter and requests cooking and drink preparation. The generation AI generates a message confirming the order, and the device responds with a voice saying, "I understand. A hamburger and orange juice, please."

[1663] Cooking status management and serving

[1664] The server receives a notification from the kitchen that the food is ready and notifies the terminal. The terminal receives the serving instructions, moves to the kitchen, places the food on a tray, and delivers it to the table. The robot delivers the food and responds with a voice message saying, "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[1665] Reorder acceptance and emotion recognition

[1666] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[1667] Inquiry response and settlement

[1668] When a user asks, "What desserts do you have?", the device converts the voice into text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. Would you like that?" When the user says, "Please pay the bill, please," the device converts the voice into text, and the generation AI identifies the payment request and notifies the server. The server compiles all the order details, calculates the payment amount, and sends it to the device. The device then brings a tablet device to the user's table and says, "Please check on this device." The user checks the amount on the tablet device and makes the payment, and the device notifies the server that the payment is complete.

[1669] Prompt Sentence Examples

[1670] "What's the scenario when a user walks into a store?"

[1671] "Please explain how emotions are recognized when ordering more."

[1672] "Please tell me in detail what steps a user should take to request a checkout."

[1673] The above is a specific operation of the system according to the embodiment of the present invention. This system is expected to make customer service in restaurants more efficient and flexible, thereby improving customer satisfaction.

[1674] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1675] Step 1:

[1676] A user enters the store

[1677] Input: User entry, video data from camera, audio data from microphone

[1678] Operation: The device recognizes the user's face using a camera installed at the entrance, processes the video data to detect the user's presence, and captures audio data using a microphone.

[1679] Output: Recognition result that the user has entered the store

[1680] Step 2:

[1681] Recognize the user's first utterance

[1682] Input: User's speech

[1683] How it works: When a user says "hello," the device captures this audio through the microphone. The audio data is sent to the Google Cloud Speech-to-Text API and converted to text.

[1684] Output: Text data (e.g. "Hello")

[1685] Step 3:

[1686] Generative AI creates greeting messages

[1687] Input: Text data (e.g. "Hello")

[1688] How it works: A generative AI (e.g., GPT-4) analyzes text data and generates an appropriate greeting.

[1689] Output: Greeting message (e.g. "Hello, welcome. We will show you to your seat.")

[1690] Step 4:

[1691] Voice output of greetings

[1692] Input: Greeting message (e.g. "Hello, welcome. We'll show you to your seat.")

[1693] Behavior: The generated greeting message is played aloud through the device's speaker.

[1694] Output: A voice response to the user

[1695] Step 5:

[1696] Get the user's orders

[1697] Input: User's voice order

[1698] How it works: When a user says, "I'd like a hamburger and orange juice, please," the device captures this speech through the microphone. The speech data is then sent to the Google Cloud Speech-to-Text API, where it is converted into text data.

[1699] Output: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[1700] Step 6:

[1701] Order analysis

[1702] Input: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[1703] How it works: Generative AI analyzes text data and identifies the order details.

[1704] Output: Identification of the order (e.g. "Hamburger", "Orange juice")

[1705] Step 7:

[1706] Send the order to the server

[1707] Input: Identification of order details (e.g. "Hamburger", "Orange juice")

[1708] Operation: The terminal sends the order information to the server.

[1709] Output: Order information received by the server

[1710] Step 8:

[1711] Transferring order information to the kitchen

[1712] Input: Order information received by the server

[1713] How it works: The server forwards the order information to the kitchen or bar counter and instructs cooking and preparation.

[1714] Output: Instructions received by the kitchen or bar counter

[1715] Step 9:

[1716] Order confirmation

[1717] Input: The order details identified by the generation AI (e.g., "hamburger," "orange juice")

[1718] How it works: The AI ​​generates a confirmation message and responds audibly through the device's speaker: "Okay, a hamburger and orange juice, right?"

[1719] Output: Audible acknowledgment to the user

[1720] Step 10:

[1721] Receive notifications when food is cooked

[1722] Input: Notification from the kitchen that cooking is complete

[1723] Operation: The server receives a notification from the kitchen that the food is ready and notifies the terminal.

[1724] Output: Cooking completion notification to the device

[1725] Step 11:

[1726] Receiving and acting on serving instructions

[1727] Input: Delivery instructions from the server

[1728] How it works: The terminal receives serving instructions, and the robot moves to the kitchen, places the food on a tray, and delivers it to the table.

[1729] Output: Food served

[1730] Step 12:

[1731] Reorder acceptance and emotion recognition

[1732] Input: User's voice for reorder

[1733] How it works: When a user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[1734] Output: Response to the user about the reorder

[1735] Step 13:

[1736] Inquiry response

[1737] Input: User's spoken query

[1738] How it works: When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. What would you like?"

[1739] Output: Audio response to user queries

[1740] Step 14:

[1741] Acceptance of settlement requests

[1742] Input: User's voice request for payment

[1743] How it works: When a user says, "I'd like to pay," the device converts the speech into text, and the generation AI identifies the payment request and notifies the server.

[1744] Output: Settlement request notification to the server

[1745] Step 15:

[1746] Settlement process

[1747] Input: Settlement request notification to the server

[1748] How it works: The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet to the user's table and says, "Please check on this terminal."

[1749] Output: The settlement amount displayed to the user

[1750] Step 16:

[1751] Check and complete the payment

[1752] Input: User payment confirmation

[1753] How it works: When the user checks the amount on the tablet and makes the payment, the device notifies the server that the payment is complete.

[1754] Output: Notification of successful payment

[1755] The above is the flow of processing of the program of this system.

[1756] (Application example 2)

[1757] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1758] Conventional restaurant and food delivery systems lack the flexibility to consider user emotions, resulting in a uniform quality of customer experience and difficulty in improving satisfaction. Furthermore, appropriate responses that consider user emotions are required when notifying users of delivery status and collecting feedback after delivery. Therefore, a system that can provide attentive service based on emotion recognition, in addition to voice recognition, is required, but no such system currently exists.

[1759] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1760] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order contents;] [means for sending order information to the server;] [means for controlling the robot by receiving instructions from the server;] [means for responding to the user by voice;] [means for using generative artificial intelligence to analyze the user's emotions and generate a response based on those emotions;] [means for notifying the user of the delivery status in real time; and [means for collecting post-delivery evaluations based on the user's emotions.] This enables flexible and efficient service provision based on the user's emotions and a consistent improvement in the customer experience during and after delivery.

[1761] "Means for inputting user voice" refers to technology for inputting user speech using a microphone or similar device in restaurants or food delivery systems.

[1762] The "means for converting speech to text" is a speech recognition technology that converts a user's speech data into character string data.

[1763] The "means for analyzing text to identify order details" is a technology for analyzing text data that has been speech-recognized to identify and specify the user's order details.

[1764] The "means for transmitting order information to a server" is a technique for transmitting user order information to a central server via a communication network.

[1765] "Means for controlling a robot by receiving instructions from a server" refers to a control technology that receives instructions from a server, operates the robot, and executes the specified task.

[1766] "Means for responding to the user by voice" refers to a technology for generating an appropriate voice response to the user's speech and transmitting it to the user through a speaker.

[1767] "Means using generative artificial intelligence to analyze a user's emotions and generate responses based on those emotions" refers to artificial intelligence technology that recognizes emotions from the user's voice and automatically generates natural responses based on those emotions.

[1768] "Means for notifying the user of the status during delivery in real time" refers to technology that notifies the user of the current status, estimated arrival time, etc. in real time during food delivery.

[1769] The "means for collecting post-delivery evaluations based on user emotions" is a technology for collecting delivery evaluation feedback based on user emotions after delivery is completed.

[1770] The present invention provides a system for restaurants and food delivery systems that provides flexible responses that take into account user emotions. This system has the following main functions:

[1771] Basic configuration

[1772] The system includes a means for inputting the user's voice, a means for converting the voice into text, a means for analyzing the text to identify the order contents, a means for sending the order information to a server, a means for receiving instructions from the server to control the robot, and a means for responding to the user by voice.

[1773] Additionally, it also includes generative artificial intelligence that analyzes the user's emotions and generates responses based on those emotions, a means of notifying the user of the delivery status in real time, and a means of collecting post-delivery evaluations based on the user's emotions.

[1774] Hardware and Software

[1775] Hardware: Uses the computer's built-in camera, microphone, robot, and speaker. These devices are used to input the user's voice and image, and to output response voices and delivery information.

[1776] software:

[1777] Speech Recognition Module: Uses the speech_recognition library, which converts the user's speech into text data.

[1778] Facial Recognition Software: Recognizes the user's face using the OpenCV library.

[1779] Generative AI: We use emotion recognition and text generation models from the transformers library, specifically the 'michellejieli / emotion_text' and 'gpt-3.5-turbo' models.

[1780] User voice recognition and emotion analysis

[1781] The device uses a camera and microphone to recognize the user, detects the user through facial recognition, and converts the user's speech into text using a speech recognition module. This text data is then input into an emotion recognition model to identify the user's emotions.

[1782] Response Generation

[1783] The generative AI generates an appropriate response based on the user's emotions. For example, if the user's voice is recognized as saying "hello" and the voice contains the emotion of "welcome," the generative AI model will generate a response such as "Welcome, please place your order." This response is then conveyed to the user through the speaker.

[1784] Examples and prompts

[1785] Examples:

[1786] Scenario: A user launches your app and says, "Good evening, I'd like a pizza and a Coke, please."

[1787] Emotion recognition: The emotion of "friendliness" is recognized from the user's voice.

[1788] Output Response: "Thank you for your order. Pizza and Coke, please. Now, please tell me your delivery address."

[1789] Example prompt sentence:

[1790] "Respond in a friendly tone: Welcome, your order is welcome."

[1791] "Please proceed to the next step to confirm your delivery address."

[1792] This will enable flexible and efficient service provision based on user emotions and improve the consistent customer experience during and after delivery.

[1793] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1794] Step 1:

[1795] The device recognizes the user using a camera and microphone. The camera detects the user's face and uses facial recognition software (OpenCV) to confirm the user's presence. At the same time, the microphone collects the user's speech and inputs it as audio data. The output is the user's facial image data and audio data.

[1796] Step 2:

[1797] The device sends the collected voice data to a voice recognition module (speech_recognition library), which converts the voice data into text data. The input is the user's voice data, and the output is text data. Specifically, the voice recognition module analyzes the voice waveform and generates a corresponding string of characters.

[1798] Step 3:

[1799] The server receives the text data and uses generative artificial intelligence (the transformers library) to analyze the spoken text and identify the user's emotions. The input is text data generated by speech recognition, and the output is the recognized emotion label. Specifically, the emotion recognition model analyzes the text and determines emotions such as "welcoming" or "friendliness."

[1800] Step 4:

[1801] The server uses generative artificial intelligence to generate an appropriate response based on the identified emotion. The input is text data containing the user's emotion label and order details, and the output is a response text. Specifically, the response generation model forms a natural-sounding answer while taking the emotion label into account.

[1802] Step 5:

[1803] The server sends the generated response text to the terminal, and the terminal responds to the user audibly through the speaker. The input is the response text sent from the server, and the output is the audio response to the user. Specifically, the text is converted into audio and played back so that the user can hear it.

[1804] Step 6:

[1805] The terminal sends the user's order information to the server, which then transmits the order information to the cooking station or delivery station. The input is the order information written in text, and the output is an execution instruction sent to the cooking station or delivery station. Specifically, the order information is sent to each station via the network.

[1806] Step 7:

[1807] The server receives delivery status information from the delivery station in real time and notifies the user of the delivery status. The input is status report data from the delivery station, and the output is delivery progress information displayed to the user. Specifically, the status data is notified to the user's smartphone or other device.

[1808] Step 8:

[1809] When the delivery is completed, the terminal collects the user's evaluation based on their emotions and sends it to the server. The input is the user's voice feedback, and the output is the evaluation data sent to the server. Specifically, the voice feedback is converted into text and saved along with the user's emotional evaluation.

[1810] Step 9:

[1811] The server analyzes the collected evaluation data to improve service quality. The input is the collected evaluation data, and the output is reports and statistical information for improvement. Specifically, the evaluation data in the database is analyzed to identify areas for improvement.

[1812] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1813] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1814] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1815] [Fourth embodiment]

[1816] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1817] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1818] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1819] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1820] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1821] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1822] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1823] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1824] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1825] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1826] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1827] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1828] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1829] The present invention is a system for providing efficient and flexible service using robots in restaurants. Specific embodiments for carrying out the invention and the processing of the program therefor will be described below.

[1830] User recognition and conversation initiation

[1831] User entry recognition

[1832] The device uses a camera and microphone to recognize the user. When a user enters a store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech. Through facial recognition algorithms and voice input, it is possible to quickly respond to users even when meeting them for the first time.

[1833] Initial voice recognition and greeting

[1834] When a user says "hello," the device sends the speech to a speech recognition module, which converts the speech to text. The generative AI analyzes the text and generates an appropriate greeting. The device responds through the speaker, saying, "Hello, welcome. We'll show you to your seat."

[1835] Receiving orders

[1836] Voice recognition for ordering

[1837] If a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which the generation AI then analyzes to determine the order.

[1838] Communicating with the Server

[1839] The terminal sends the specified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin promptly.

[1840] Acknowledgement and Response

[1841] The AI ​​generates a message to confirm the order, and the device responds aloud, saying, "Understood. A hamburger and orange juice, please."

[1842] Cooking status management and serving

[1843] Cooking completion notification

[1844] The server receives a notification from the kitchen that the food is ready and sends that information to the terminal, allowing the robot to begin serving the food.

[1845] Serving instructions and actions

[1846] The server sends serving instructions to the terminal, which then follows those instructions to move to the kitchen and place the food on a tray. The robot then safely carries the food to the table and serves it to the user, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[1847] Continuing the conversation and being flexible

[1848] Accepting additional orders

[1849] If the user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. The device responds, "Okay, I'll order some extra fries," and sends the additional order information to the server.

[1850] Inquiry response

[1851] When a user asks, "What desserts do you have?", the device converts the voice into text, and the AI ​​analyzes the query and generates an appropriate answer. The device responds aloud, "Today's desserts include cake and ice cream."

[1852] settlement

[1853] Acceptance of settlement requests

[1854] When a user says, "I'd like to pay, please," the device converts the speech into text using a voice recognition module, and the generation AI identifies the payment request and notifies the server.

[1855] Settlement process

[1856] The server compiles all the order details, calculates the payment amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and says, "Please check on this device."

[1857] Check and complete the payment

[1858] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed, and the settlement process is complete.

[1859] As described above, the present invention provides a system that efficiently and flexibly performs user recognition, order taking, serving, conversation, and payment processing, thereby optimizing service in restaurants.

[1860] The processing flow will be explained below.

[1861] Step 1:

[1862] The user enters the store.

[1863] Step 2:

[1864] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[1865] Step 3:

[1866] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[1867] Step 4:

[1868] The user says "Hello."

[1869] Step 5:

[1870] The terminal (voice recognition module) converts the user's speech into text.

[1871] Step 6:

[1872] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[1873] Step 7:

[1874] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[1875] Step 8:

[1876] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[1877] Step 9:

[1878] The terminal (voice recognition module) converts the user's speech into text.

[1879] Step 10:

[1880] The terminal (generative AI) analyzes the text and identifies the order contents.

[1881] Step 11:

[1882] The terminal sends the order details to the server.

[1883] Step 12:

[1884] The server receives the order information and passes it on to the kitchen and bar counter.

[1885] Step 13:

[1886] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[1887] Step 14:

[1888] The terminal will audibly convey a confirmation message to the user.

[1889] Step 15:

[1890] The server receives a notification from the kitchen that the food is ready.

[1891] Step 16:

[1892] The server sends a cooking completion notification to the terminal.

[1893] Step 17:

[1894] The device receives a notification, goes to the kitchen and places the food on a tray.

[1895] Step 18:

[1896] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[1897] Step 19:

[1898] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[1899] Step 20:

[1900] The terminal (voice recognition module) converts the user's speech into text.

[1901] Step 21:

[1902] The terminal (generative AI) analyzes the text and identifies the additional order details.

[1903] Step 22:

[1904] The terminal transmits the additional order details to the server.

[1905] Step 23:

[1906] The terminal responds, "Okay, we'll add fries."

[1907] Step 24:

[1908] The user asks about dessert. Say, "What desserts do you have?"

[1909] Step 25:

[1910] The terminal (voice recognition module) converts the user's speech into text.

[1911] Step 26:

[1912] The terminal (generative AI) analyzes the text and generates a dessert menu.

[1913] Step 27:

[1914] The device will respond aloud, "Today's desserts include cake and ice cream."

[1915] Step 28:

[1916] The user wishes to settle the bill and says, "Please pay."

[1917] Step 29:

[1918] The terminal (voice recognition module) converts the user's speech into text.

[1919] Step 30:

[1920] The terminal (generator AI) identifies the settlement request and notifies the server.

[1921] Step 31:

[1922] The server compiles all the order details and calculates the total amount.

[1923] Step 32:

[1924] The server sends the settlement information to the terminal.

[1925] Step 33:

[1926] The device brings a tablet device to the user's table and says, "Please check this device."

[1927] Step 34:

[1928] The user checks the amount on the tablet device and makes the payment.

[1929] Step 35:

[1930] The terminal notifies the server that the payment has been completed.

[1931] Example 1

[1932] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1933] Providing efficient and flexible service is a key issue for modern restaurants. In particular, there is a demand for systems that automate and seamlessly perform all processes, from recognizing users as they enter the restaurant to taking their orders, serving food, handling additional orders, and processing payments. Current systems often only provide individual functions and are insufficient for providing comprehensive services. Furthermore, delays in the recognition of specific users and the timing of voice responses can lead to problems that reduce user satisfaction.

[1934] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1935] In this invention, the server includes means for inputting user voice, means for converting voice to text, means for analyzing the text to identify the order details, means for transmitting the order information to the server, means for receiving instructions from the server and controlling the robot, means for responding to the user by voice, means for detecting the user using facial recognition technology, and means for receiving notification of cooking completion. This enables the consistent automation of everything from recognizing the user's entry to receiving the order, serving the food, handling additional orders, and processing the payment, enabling the provision of efficient and flexible services.

[1936] The "means for inputting user's voice" refers to a device that acquires the user's speech using a voice input device such as a microphone.

[1937] A "means for converting speech to text" is a device or software that converts speech data acquired using speech recognition technology into text data.

[1938] The "means for analyzing text to identify order details" refers to software or an algorithm for analyzing the acquired character data and identifying the order details from the content of that data.

[1939] The "means for transmitting order information to the server" refers to a device or software that has the function of transmitting the specified order details to the server via a communication means.

[1940] The "means for receiving instructions from the server and controlling the robot" refers to a device or software for receiving instructions from the server and controlling the operation of the robot based on those instructions.

[1941] The "means for responding to the user by voice" refers to a device or software for providing the generated response message to the user by voice through a voice synthesizer.

[1942] A "means for detecting a user using facial recognition technology" is a device or software that uses a camera and a facial recognition algorithm to detect and identify a user's face.

[1943] The "means for receiving notification of cooking completion" refers to a device or software for receiving notification of cooking completion from the kitchen and sharing that information within the system.

[1944] The present invention provides a system for providing efficient and flexible service using robots in restaurants. Specific processing steps for carrying out the present invention will be described in natural language.

[1945] This system uses a microphone to input the user's voice and speech recognition technology to convert the voice into text. This is done using the Google Speech-to-Text API, among others. It also uses generative artificial intelligence (generative AI model) to analyze the text and identify the order contents. This analysis utilizes OpenAI's GPT-4, among others.

[1946] A device with communication capabilities is used to send order information to the server. This sends the order details to the server (e.g., Dell PowerEdge T30). The robot operates based on a communication protocol to receive instructions from the server and control the robot. A typical service robot (e.g., SoftBank's Pepper) is used as the robot.

[1947] To respond to the user via voice, a generated response message is provided to the user using a speech synthesizer (e.g., Amazon Polly).Furthermore, to detect the user using facial recognition technology, a camera (e.g., Logitech HD Pro Webcam C920) and a facial recognition algorithm (e.g., OpenCV library) are used.To receive a notification that cooking is complete, a communication function is used to receive notifications from the kitchen and share that information with the device.

[1948] As a concrete example, let's consider a scenario where a user enters a restaurant. When the user enters the restaurant, the camera recognizes their face and the microphone picks up the utterance "Hello." The device converts the utterance into text through a voice recognition module, and the generation AI analyzes it and responds with "Hello, welcome. We will show you to your seat."

[1949] Here's an example of ordering: When a user says, "I'd like a hamburger and orange juice, please," the device converts the speech into text, and the generation AI identifies the order. This information is sent to the server, which then relays the order to the kitchen or bar counter. After confirming the order, the device responds, "I understand. A hamburger and orange juice, please."

[1950] If the user wishes to place an additional order, for example by saying, "I'd like some extra fries, please," the device analyzes the voice and uses a generative AI to identify the additional order. The additional order information is then sent to the server. After the cooking is complete, the robot brings the food from the kitchen and serves it to the user. It responds, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[1951] Some examples of prompts are:

[1952] "How do you greet users when they walk in?"

[1953] "When a user places an order, how do you verify the order and provide cooking instructions?"

[1954] "How do you transport the food once it's done cooking?"

[1955] "How do you handle additional orders or inquiries from users?"

[1956] "How do you guide users through the process at checkout?"

[1957] In this way, the system is able to provide consistently efficient and flexible service, optimizing service in restaurants.

[1958] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1959] Specific explanation of processing steps

[1960] Step 1: User entry recognition

[1961] Specifically: The device uses a camera and microphone to recognize the user. The camera (Logitech HD Pro Webcam C920) detects the user's face, and the microphone (Shure MV5) picks up the user's speech.

[1962] Input: Camera image and audio data

[1963] Output: Face detection data and audio clips

[1964] Data processing / calculation: Detect faces from camera images (using OpenCV library) and save audio clips

[1965] Specific operation: When a user enters the store, the camera recognizes their face and confirms what they are saying through voice input.

[1966] Step 2: First voice recognition and greeting

[1967] Specifically, the device sends the user's speech to a speech recognition module (Google Speech-to-Text API) and converts it into text data. A generative AI (OpenAI's GPT-4) analyzes the text and generates an appropriate greeting message.

[1968] Input: Audio clip (user speech)

[1969] Output: Character data and response message

[1970] Data processing / calculation: Converting voice data to text data (Speech-to-Text), character analysis and response generation (generative AI)

[1971] Specific operation: When the user says "Hello," the device responds with "Hello, welcome. We will show you to your seat."

[1972] Step 3: Voice recognition of orders

[1973] Specific explanation: When a user speaks an order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the order contents.

[1974] Input: Audio clip (order details)

[1975] Output: Text data (order details) and order information

[1976] Data processing / calculation: Converting voice data into text data and identifying order details through text analysis

[1977] Specific operation: The user says, "I'd like a hamburger and orange juice, please," and this is identified as the order information.

[1978] Step 4: Communicating with the Server

[1979] Specific explanation: The terminal sends the identified order details to the server, which then transmits the order information to the kitchen or bar counter.

[1980] Input: Order Information

[1981] Output: Instructions to the kitchen and bar counter

[1982] Data processing / calculation: Sends order information to the server and transfers cooking instructions to each department

[1983] Specific operation: The terminal sends an order for "hamburger and orange juice" to the server, which then distributes it to the kitchen and bar counter.

[1984] Step 5: Order confirmation and response

[1985] Specific explanation: The generation AI generates a message to confirm the order details, and the terminal responds with a voice message saying, "Okay, a hamburger and orange juice, right?"

[1986] Input: Order Information

[1987] Output: Confirmation message

[1988] Data processing / calculation: Checking order information and generating messages

[1989] Specific operation: The generation AI creates a confirmation message, which the device then audibly conveys to the user.

[1990] Step 6: Receive notification that your food is ready

[1991] Specific explanation: The server receives a cooking completion notification from the kitchen and sends that information to the terminal.

[1992] Input: Cooking complete notification

[1993] Output: Cooking completion information

[1994] Data processing / calculation: Receiving and transmitting cooking completion notification

[1995] Specific operation: The kitchen notifies the server that the food is ready, and the server sends that information to the device.

[1996] Step 7: Food serving instructions and actions

[1997] Details: The server sends serving instructions to the terminal, which then moves to the kitchen according to the instructions and places the food on a tray. The robot then carries it to the table and serves it to the user.

[1998] Input: Cooking completion information

[1999] Output: Food served

[2000] Data processing / calculation: Generation of serving instructions and robot motion control

[2001] What it does: The device moves to the kitchen and places the food on a tray, which the robot then carries to the table.

[2002] Step 8: Accepting additional orders

[2003] Specific explanation: When a user speaks an additional order, the device converts the voice into text data using a voice recognition module, and the generation AI analyzes it to identify the additional order.

[2004] Input: Audio Clip (reorder)

[2005] Output: Text data (reorder) and reorder information

[2006] Data processing / calculation: Converting voice data into text data and identifying additional orders

[2007] Specific operation: When the user says, "I'd like some extra fries, please," the device analyzes it and sends it to the server.

[2008] Step 9: Handling inquiries

[2009] Specific explanation: When a user speaks a question, the device converts the voice into text data, which is then analyzed by the generation AI to generate an appropriate answer.

[2010] Input: Audio clip (user question)

[2011] Output: Text data (question content) and answer message

[2012] Data processing / calculation: Converting voice data into text data, analyzing questions, and generating answers

[2013] Specific behavior: When the user asks, "What desserts do you have?", the device responds, "Today's desserts include cake and ice cream."

[2014] Step 10: Accepting the settlement request

[2015] Specific explanation: When a user says, "I'd like to pay the bill, please," the device converts the voice into text data, and the generation AI identifies the payment request and notifies the server.

[2016] Input: Audio clip (payment request)

[2017] Output: Text data (settlement request) and settlement notice

[2018] Data processing / calculation: Converting voice data into text data, identifying and notifying payment requests

[2019] Specific operation: When the user says "I'd like to pay the bill," the terminal notifies the server.

[2020] Step 11: Checkout

[2021] Specifically: The server compiles all the order details, calculates the total amount, and sends it to the terminal, which then brings a tablet to the user's table and displays the total amount.

[2022] Input: Order details

[2023] Output: Settlement amount

[2024] Data processing / calculation: Aggregating order details and calculating payment amounts

[2025] Specific operation: The server performs the calculation and the terminal displays the amount on the tablet.

[2026] Step 12: Check and complete the payment

[2027] Specific explanation: The user checks the amount on the tablet and makes the payment. The terminal notifies the server that the payment has been completed, and the settlement process is complete.

[2028] Input: Payment Information

[2029] Output: Payment completion notification

[2030] Data processing / calculation: Confirmation of payment information and notification of settlement completion

[2031] Specific operation: The user checks the amount on the tablet, makes the payment, and the device notifies the server.

[2032] In this way, the system can efficiently automate restaurant operations through a series of processes and provide users with high quality service.

[2033] (Application example 1)

[2034] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2035] Conventional food delivery systems have the problem that the entire process from ordering to delivery and payment is divided into manual processes and multiple different systems, resulting in a lack of efficiency and flexibility. The present invention aims to solve these problems, improve the efficiency of services for users, and provide a system that unifies and smoothly processes each stage of ordering, delivery, and payment.

[2036] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2037] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order details;] [means for sending order information to the server;] [means for controlling the robot upon receiving instructions from the server;] [means for responding to the user by voice;] [means for recognizing the user using facial recognition technology;] [means for sending delivery instructions to the delivery robot after accepting the order;] [means for responding to payment requests and calculating the payment amount; and [means for completing the user's payment.] This makes it possible to carry out the entire process from ordering to delivery and payment in a unified and efficient manner.

[2038] The "means for inputting user's voice" refers to a means for acquiring voice uttered by the user through a device.

[2039] The "means for converting voice to text" is a means for analyzing acquired voice data and converting it into text data.

[2040] The "means for analyzing text to identify order details" refers to a means for analyzing text data and accurately identifying the user's intentions and order details.

[2041] The "means for transmitting order information to the server" is a means for transmitting the analyzed order details to the server via a network.

[2042] The "means for receiving instructions from the server and controlling the robot" refers to means for receiving instruction data from the server and controlling the robot's operation based on that data.

[2043] The "means for responding to the user by voice" refers to a means for generating voice in response to an input from the user and responding to the user.

[2044] "Means for recognizing a user using facial recognition technology" refers to technology that uses a camera or the like to identify and specify the user's face.

[2045] The "means for sending delivery instructions to a delivery robot after receiving an order" refers to a means for transmitting the order details to a delivery robot and issuing delivery instructions after receiving a user's order.

[2046] The "means for responding to a payment request and calculating the payment amount" is a means for receiving a payment request from a user and calculating the payment amount based on all order details.

[2047] The "means for completing the user's settlement" is the means by which the user confirms the amount presented and completes the payment procedure.

[2048] This invention is a system that automates efficient and flexible robot-based food delivery services. This system automates a series of processes, from user voice input, speech-to-text conversion, order analysis, order information transmission to a server, robot control, voice response, user identification using facial recognition technology, delivery instructions to the delivery robot, payment request handling, and payment completion.

[2049] Hardware and Software Configuration

[2050] This system mainly consists of the following hardware and software:

[2051] Hardware: smartphones, delivery robots, servers

[2052] Software: Speech recognition module (e.g., Google Speech-to-Text API), face recognition module (e.g., OpenCV), generative AI model (e.g., OpenAI GPT-4), Python, Django framework, Celery

[2053] Program processing

[2054] The main processing of the entire system will be explained.

[2055] 1. Enter the user's voice:

[2056] The server and the terminal use a microphone to input the user's voice, which is then converted into text data by a voice recognition module.

[2057] 2. Convert speech to text:

[2058] The server uses a speech recognition module to convert the voice data into text.

[2059] 3. Parse the text to identify the order:

[2060] The terminal uses a generative AI model (e.g., GPT-4) to analyze the converted text and identify the order details.

[2061] 4. Send the order information to the server:

[2062] The terminal transmits the parsed order details to the server.

[2063] 5. Control the robot by receiving instructions from the server:

[2064] The server sends the order details to the delivery robot and controls the robot to pick up and deliver the order.

[2065] 6. Respond to the user verbally:

[2066] The device uses the generative AI model to generate a response message and conveys it to the user via voice output.

[2067] 7. Recognize users using facial recognition technology:

[2068] The device uses a camera to perform facial recognition and identify the user.

[2069] 8. After receiving the order, send delivery instructions to the delivery robot:

[2070] After receiving the order, the server issues delivery instructions to the delivery robot.

[2071] 9. Respond to settlement requests and calculate settlement amounts:

[2072] The server receives the user's payment request, calculates all the order details, and calculates the payment amount.

[2073] 10. Complete the user's checkout:

[2074] The terminal presents the payment amount to the user and completes the payment.

[2075] Specific examples

[2076] For example, when a user says "Pizza and Coke please" on a smartphone, the following process is executed:

[2077] 1. Speech recognition: The speech is converted into text, and the text data "Pizza and Coke, please" is generated.

[2078] 2. Order Analysis: A generative AI model analyzes the text data to identify an order for a pizza and a cola.

[2079] 3. Robot control: The server sends instructions to the delivery robot to pick up the pizza and cola from the store and deliver it to the user's address.

[2080] 4. Voice response: "Thank you for your order. Your pizza and Coke will be on their way soon," the device informs the user.

[2081] 5. Settlement: When the user says, "Please pay the bill," the server calculates the amount and the payment is completed on the settlement screen displayed on the smartphone.

[2082] As an example of input to the generative AI model, we will use the following prompt:

[2083] "A user asks, 'What's on the menu today?' What are your menu recommendations?"

[2084] By inputting the above prompt sentences into the generative AI model, an appropriate response can be obtained. This process enables efficient and flexible food delivery services.

[2085] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2086] Step 1:

[2087] The user inputs voice into the smartphone.

[2088] The user speaks into the smartphone microphone, "Pizza and Coke please." The input data is the user's voice data. The device acquires this voice data and passes it on to the next step.

[2089] Step 2:

[2090] The device converts the voice data into text data.

[2091] The device uses a speech recognition module (e.g., Google Speech-to-Text API) to convert the voice data into text data. Here, the input is voice data, and the output is text data such as "Pizza and Coke, please."

[2092] Step 3:

[2093] The terminal analyzes the text data and identifies the order details

[2094] The device uses a generative AI model (e.g., GPT-4) to analyze the text data. Specifically, it extracts the order details of "pizza" and "cola." The input is text data, and the output is the identified order details.

[2095] Step 4:

[2096] The terminal sends the order information to the server

[2097] The terminal sends the specified order details to the server, where the input is the order details data and the output is the order information sent to the server.

[2098] Step 5:

[2099] The server receives the order information and sends delivery instructions to the delivery robot.

[2100] Based on the received order information, the server sends an instruction to the delivery robot to "deliver pizza and cola to the specified address." The input is the order information, and the output is the delivery instruction data.

[2101] Step 6:

[2102] Delivery robots will pick up orders from stores and deliver them to designated addresses.

[2103] The delivery robot receives instructions from the server, goes to the store, picks up the pizza and cola, and delivers it to the specified user's address. The input is delivery instruction data, and the output is a delivery completion notification.

[2104] Step 7:

[2105] The device notifies the user of the delivery progress

[2106] The device uses a generative AI model to generate a message about the delivery progress, such as "Your pizza and coke will arrive soon," and notify the user via voice. The input is the delivery progress data, and the output is the voice message.

[2107] Step 8:

[2108] The user makes a payment request

[2109] The user verbally requests the terminal, "Please pay the bill." The input is voice data, which the terminal recognizes and proceeds to the next step.

[2110] Step 9:

[2111] The terminal converts the voice data into text data and sends a payment request to the server.

[2112] The terminal converts the voice data into text data using a voice recognition module and sends this text data to the server. The input is the voice data, and the output is the text data and payment request data sent to the server.

[2113] Step 10:

[2114] The server calculates the settlement amount and sends it to the terminal.

[2115] The server calculates the total amount based on all the order details and sends the total amount data to the terminal. The input is the order information and the output is the total amount data.

[2116] Step 11:

[2117] The terminal presents the payment amount to the user and completes the payment procedure.

[2118] The terminal uses the generative AI model to generate a message saying, "The settlement amount is XX yen," and presents it to the user via voice. The settlement procedure is then completed after the user performs the payment operation. The input is the settlement amount data, and the output is a payment completion notification.

[2119] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2120] The present invention is a customer service system using a robot in a restaurant, which is particularly equipped with the function of recognizing the user's emotions and responding flexibly based on them. This system has the following main functions:

[2121] User recognition and conversation initiation

[2122] User entry recognition

[2123] The device uses a camera and microphone to recognize the user. When a user enters the store, the camera detects the user using facial recognition technology, and the microphone picks up the user's speech, allowing the device to prepare for the initial interaction.

[2124] Initial voice recognition and greeting

[2125] When a user says "hello," the device sends the speech to a speech recognition module. The speech is converted into text data, and a generative AI analyzes the text to generate an appropriate greeting. The message "Hello, welcome. We will show you to your seat" is played over the device's speaker.

[2126] Receiving orders

[2127] Voice recognition for ordering

[2128] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text, which is then analyzed by a generation AI to determine the order.

[2129] Communicating with the Server

[2130] The terminal sends the identified order details to the server, which then relays the order information to the kitchen or bar counter, allowing food and drink preparation to begin quickly.

[2131] Acknowledgement and Response

[2132] The AI ​​generates a message confirming the order details, and the device responds to the user verbally, saying, "I understand. A hamburger and orange juice, right?"

[2133] Cooking status management and serving

[2134] Cooking completion notification

[2135] The server receives a cooking completion notification from the kitchen, then sends the cooking completion notification to the terminal, allowing the robot to begin serving the food.

[2136] Serving instructions and actions

[2137] The terminal receives the serving instructions, moves to the kitchen, and places the food on a tray. The robot then safely delivers the food to the table, responding with a voice message: "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy."

[2138] Continuing the conversation and being flexible

[2139] Reorder acceptance and emotion recognition

[2140] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion engine recognizes the user's emotions from their voice and reflects them in the response. For example, if the emotion engine detects fatigue in the user's voice, the device will respond in a gentle tone, saying, "I understand. I'll bring the fries right away, so please wait a moment."

[2141] Inquiry response

[2142] When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion engine analyzes the user's emotions, and the device responds to the user in a gentle tone, saying, "Today's desserts include cake and ice cream. What would you like?"

[2143] settlement

[2144] Acceptance of settlement requests

[2145] When a user says, "Please give me the bill," the device converts the speech into text using a speech recognition module. The AI ​​generator identifies the payment request and notifies the server.

[2146] Settlement process

[2147] The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet device to the user's table and tells them to "please check on this device."

[2148] Check and complete the payment

[2149] The user checks the amount on the tablet device and makes the payment. The device then notifies the server that the payment has been completed.

[2150] The system, which includes the above functions, can recognize users' emotions and provide flexible and efficient services based on those emotions, significantly improving the customer experience in restaurants.

[2151] The processing flow will be explained below.

[2152] Step 1:

[2153] The user enters the store.

[2154] Step 2:

[2155] The device (robot camera) recognizes the user and detects the user's presence using facial recognition technology.

[2156] Step 3:

[2157] The terminal (robot microphone) waits for the user to speak. Voice input begins.

[2158] Step 4:

[2159] The user says "Hello."

[2160] Step 5:

[2161] The terminal (voice recognition module) converts the user's speech into text.

[2162] Step 6:

[2163] The device (generative AI) analyzes the text and generates an appropriate greeting message.

[2164] Step 7:

[2165] The device responds through the speaker, "Hello, welcome. We will show you to your seat."

[2166] Step 8:

[2167] The user begins to order by saying, "I'd like a hamburger and an orange juice, please."

[2168] Step 9:

[2169] The terminal (voice recognition module) converts the user's speech into text.

[2170] Step 10:

[2171] The terminal (generative AI) analyzes the text and identifies the order contents.

[2172] Step 11:

[2173] The terminal transmits the specified order details to the server.

[2174] Step 12:

[2175] The server receives the order information and passes it on to the kitchen and bar counter.

[2176] Step 13:

[2177] The terminal (generation AI) generates a confirmation message saying, "Okay, a hamburger and orange juice, right?"

[2178] Step 14:

[2179] The terminal will audibly convey a confirmation message to the user.

[2180] Step 15:

[2181] The server receives a notification from the kitchen that the food is ready.

[2182] Step 16:

[2183] The server sends a cooking completion notification to the terminal.

[2184] Step 17:

[2185] The device receives a notification, goes to the kitchen and places the food on a tray.

[2186] Step 18:

[2187] The device delivers the food to the user's table, responding, "Sorry to keep you waiting. Here's a hamburger and orange juice. Please enjoy."

[2188] Step 19:

[2189] The user wishes to order more food. He / she says, "I'd like some extra fries, please."

[2190] Step 20:

[2191] The terminal (voice recognition module) converts the user's speech into text.

[2192] Step 21:

[2193] The terminal (generative AI) analyzes the text and identifies the additional order details.

[2194] Step 22:

[2195] The terminal transmits the additional order details to the server.

[2196] Step 23:

[2197] The terminal responds, "Okay, we'll add fries."

[2198] Step 24:

[2199] The user asks about dessert. Say, "What desserts do you have?"

[2200] Step 25:

[2201] The terminal (voice recognition module) converts the user's speech into text.

[2202] Step 26:

[2203] The terminal (generative AI) analyzes the text and generates a dessert menu.

[2204] Step 27:

[2205] The device will respond aloud, "Today's desserts include cake and ice cream."

[2206] Step 28:

[2207] The user wishes to settle the bill and says, "Please pay."

[2208] Step 29:

[2209] The terminal (voice recognition module) converts the user's speech into text.

[2210] Step 30:

[2211] The terminal (generator AI) identifies the settlement request and notifies the server.

[2212] Step 31:

[2213] The server compiles all the order details and calculates the total amount.

[2214] Step 32:

[2215] The server sends the settlement information to the terminal.

[2216] Step 33:

[2217] The device brings a tablet device to the user's table and says, "Please check this device."

[2218] Step 34:

[2219] The user checks the amount on the tablet device and makes the payment.

[2220] Step 35:

[2221] The terminal notifies the server that the payment has been completed.

[2222] Added emotion engine processing

[2223] Step 36:

[2224] The device (emotion engine) analyzes the user's facial expressions and voice to recognize their emotional state.

[2225] Step 37:

[2226] The device (generative AI) generates a flexible response based on the recognized emotion.

[2227] Step 38:

[2228] The device will respond to the user based on their emotions. For example, if it recognizes that the user is tired, it will respond in a gentle tone, saying, "I understand. I'll bring you the fries right away, so please wait a moment."

[2229] Step 39:

[2230] The terminal (emotion engine) analyzes the accumulated emotional data and generates feedback to improve the quality of the service.

[2231] Step 40:

[2232] The server receives the feedback and uses it to improve services across the system.

[2233] Example 2

[2234] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2235] In modern restaurants, improving the efficiency and flexibility of customer service is important. However, conventional systems have difficulty recognizing customer emotions and responding flexibly. Accurate order acceptance and rapid processing are also required, but there is a high possibility of human error. To solve these issues, a system is needed that automatically recognizes customer voices and provides appropriate responses and order processing.

[2236] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for analyzing the text data to identify the order contents, means for transmitting the order information to the server, means for controlling the robot based on instructions received from the server, means for responding to the user by voice, and means for recognizing the user's emotions and generating a response based on the same. This enables flexible responses that take into account the customer's emotions and accurate and prompt order processing.

[2237] The "means for acquiring the user's voice" is a combination of hardware and software for acquiring the voice uttered by the user.

[2238] The "means for converting voice into text data" refers to a voice recognition technology for converting acquired voice into digital text data.

[2239] The "means for analyzing text data to identify order details" refers to algorithms and software for analyzing text data and accurately identifying the order details intended by the user.

[2240] The "means for transmitting order information to the server" refers to a communication means and protocol for transmitting the specified order details to the server via a network.

[2241] The "means for controlling the robot based on instructions received from the server" refers to hardware and software for receiving instructions sent from the server and controlling the robot in accordance with those instructions.

[2242] "Means for providing audio responses to the user" refers to a speaker and voice generating software for playing audio messages to the user.

[2243] The "means for recognizing the user's emotions and generating responses based on them" refers to an emotion analysis engine and generative AI model that analyzes emotions from the user's speech and actions and generates responses that are adapted to those emotions.

[2244] This invention is a customer service system using a robot in a restaurant, and is equipped with a function to recognize the user's emotions and respond flexibly based on them. This system is designed to acquire and analyze the user's voice to identify the order, and further recognize the user's emotions and respond accordingly.

[2245] Hardware and Software Configuration

[2246] Terminal

[2247] Camera: A device that uses facial recognition technology to detect users entering the store.

[2248] Microphone: A device used to capture the user's voice.

[2249] Speaker: A device that provides audio responses to the user.

[2250] Speech recognition module: Software for converting speech into text data using the Google Cloud Speech-to-Text API, etc.

[2251] Generative AI: Software that uses tools such as GPT-4 to analyze text data and generate appropriate responses.

[2252] Emotion analysis engine: Software for recognizing user emotions using Microsoft Azure Emotion API, etc.

[2253] server

[2254] Communications module: A device that receives order information sent from the terminal and forwards it to the kitchen or bar counter.

[2255] Information management system: Software for managing order information and cooking status in real time.

[2256] Specific operation of the system

[2257] User recognition and conversation initiation

[2258] When a user enters a restaurant, the device's camera detects the user using facial recognition technology (e.g., OpenCV), and the microphone captures the user's speech. This allows the device to prepare for the initial interaction. When the user says "Hello," the speech recognition module converts the speech into text data. The converted text data is sent to a generative AI, which generates an appropriate greeting. The device's speaker plays the message, "Hello, welcome. We will show you to your seat."

[2259] Receiving and processing orders

[2260] When a user says, "I'd like a hamburger and orange juice, please," the device uses a speech recognition module to convert the speech into text data. The generation AI analyzes the text and determines the order. The order is then sent to a server, which forwards the information to the kitchen or bar counter and requests cooking and drink preparation. The generation AI generates a message confirming the order, and the device responds with a voice saying, "I understand. A hamburger and orange juice, please."

[2261] Cooking status management and serving

[2262] The server receives a notification from the kitchen that the food is ready and notifies the terminal. The terminal receives the serving instructions, moves to the kitchen, places the food on a tray, and delivers it to the table. The robot delivers the food and responds with a voice message saying, "Sorry to keep you waiting. Here's your hamburger and orange juice. Please enjoy your meal."

[2263] Reorder acceptance and emotion recognition

[2264] If a user says, "I'd like some extra fries, please," the device's speech recognition module converts the speech into text, and the generation AI analyzes the text to determine if the additional order is needed. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[2265] Inquiry response and settlement

[2266] When a user asks, "What desserts do you have?", the device converts the voice into text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. Would you like that?" When the user says, "Please pay the bill, please," the device converts the voice into text, and the generation AI identifies the payment request and notifies the server. The server compiles all the order details, calculates the payment amount, and sends it to the device. The device then brings a tablet device to the user's table and says, "Please check on this device." The user checks the amount on the tablet device and makes the payment, and the device notifies the server that the payment is complete.

[2267] Prompt Sentence Examples

[2268] "What's the scenario when a user walks into a store?"

[2269] "Please explain how emotions are recognized when ordering more."

[2270] "Please tell me in detail what steps a user should take to request a checkout."

[2271] The above is a specific operation of the system according to the embodiment of the present invention. This system is expected to make customer service in restaurants more efficient and flexible, thereby improving customer satisfaction.

[2272] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2273] Step 1:

[2274] A user enters the store

[2275] Input: User entry, video data from camera, audio data from microphone

[2276] Operation: The device recognizes the user's face using a camera installed at the entrance, processes the video data to detect the user's presence, and captures audio data using a microphone.

[2277] Output: Recognition result that the user has entered the store

[2278] Step 2:

[2279] Recognize the user's first utterance

[2280] Input: User's speech

[2281] How it works: When a user says "hello," the device captures this audio through the microphone. The audio data is sent to the Google Cloud Speech-to-Text API and converted to text.

[2282] Output: Text data (e.g. "Hello")

[2283] Step 3:

[2284] Generative AI creates greeting messages

[2285] Input: Text data (e.g. "Hello")

[2286] How it works: A generative AI (e.g., GPT-4) analyzes text data and generates an appropriate greeting.

[2287] Output: Greeting message (e.g. "Hello, welcome. We will show you to your seat.")

[2288] Step 4:

[2289] Voice output of greetings

[2290] Input: Greeting message (e.g. "Hello, welcome. We'll show you to your seat.")

[2291] Behavior: The generated greeting message is played aloud through the device's speaker.

[2292] Output: A voice response to the user

[2293] Step 5:

[2294] Get the user's orders

[2295] Input: User's voice order

[2296] How it works: When a user says, "I'd like a hamburger and orange juice, please," the device captures this speech through the microphone. The speech data is then sent to the Google Cloud Speech-to-Text API, where it is converted into text data.

[2297] Output: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[2298] Step 6:

[2299] Order analysis

[2300] Input: Text data (e.g., "I'd like a hamburger and orange juice, please.")

[2301] How it works: Generative AI analyzes text data and identifies the order details.

[2302] Output: Identification of the order (e.g. "Hamburger", "Orange juice")

[2303] Step 7:

[2304] Send the order to the server

[2305] Input: Identification of order details (e.g. "Hamburger", "Orange juice")

[2306] Operation: The terminal sends the order information to the server.

[2307] Output: Order information received by the server

[2308] Step 8:

[2309] Transferring order information to the kitchen

[2310] Input: Order information received by the server

[2311] How it works: The server forwards the order information to the kitchen or bar counter and instructs cooking and preparation.

[2312] Output: Instructions received by the kitchen or bar counter

[2313] Step 9:

[2314] Order confirmation

[2315] Input: The order details identified by the generation AI (e.g., "hamburger," "orange juice")

[2316] How it works: The AI ​​generates a confirmation message and responds audibly through the device's speaker: "Okay, a hamburger and orange juice, right?"

[2317] Output: Audible acknowledgment to the user

[2318] Step 10:

[2319] Receive notifications when food is cooked

[2320] Input: Notification from the kitchen that cooking is complete

[2321] Operation: The server receives a notification from the kitchen that the food is ready and notifies the terminal.

[2322] Output: Cooking completion notification to the device

[2323] Step 11:

[2324] Receiving and acting on serving instructions

[2325] Input: Delivery instructions from the server

[2326] How it works: The terminal receives serving instructions, and the robot moves to the kitchen, places the food on a tray, and delivers it to the table.

[2327] Output: Food served

[2328] Step 12:

[2329] Reorder acceptance and emotion recognition

[2330] Input: User's voice for reorder

[2331] How it works: When a user says, "I'd like some extra fries, please," the device uses a speech recognition module to convert the speech into text, and the generation AI analyzes the text to identify the additional order. At the same time, the emotion analysis engine recognizes the emotion in the user's voice and reflects it in the response. For example, if the device senses fatigue, it might respond in a gentle tone, "I understand. I'll bring the fries right away, so please wait a moment."

[2332] Output: Response to the user about the reorder

[2333] Step 13:

[2334] Inquiry response

[2335] Input: User's spoken query

[2336] How it works: When a user asks, "What desserts do you have?", the device converts the voice to text, and the generation AI analyzes the inquiry and generates an appropriate dessert menu. At the same time, the emotion analysis engine analyzes the user's emotions, and the device responds in a gentle tone, "Today's desserts include cake and ice cream. What would you like?"

[2337] Output: Audio response to user queries

[2338] Step 14:

[2339] Acceptance of settlement requests

[2340] Input: User's voice request for payment

[2341] How it works: When a user says, "I'd like to pay," the device converts the speech into text, and the generation AI identifies the payment request and notifies the server.

[2342] Output: Settlement request notification to the server

[2343] Step 15:

[2344] Settlement process

[2345] Input: Settlement request notification to the server

[2346] How it works: The server compiles all the order details, calculates the total amount, and sends it to the terminal. The terminal then brings a tablet to the user's table and says, "Please check on this terminal."

[2347] Output: The settlement amount displayed to the user

[2348] Step 16:

[2349] Check and complete the payment

[2350] Input: User payment confirmation

[2351] How it works: When the user checks the amount on the tablet and makes the payment, the device notifies the server that the payment is complete.

[2352] Output: Notification of successful payment

[2353] The above is the flow of processing of the program of this system.

[2354] (Application example 2)

[2355] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2356] Conventional restaurant and food delivery systems lack the flexibility to consider user emotions, resulting in a uniform quality of customer experience and difficulty in improving satisfaction. Furthermore, appropriate responses that consider user emotions are required when notifying users of delivery status and collecting feedback after delivery. Therefore, a system that can provide attentive service based on emotion recognition, in addition to voice recognition, is required, but no such system currently exists.

[2357] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2358] In this invention, the server includes: [means for inputting user voice;] [means for converting voice into text;] [means for analyzing the text to identify the order contents;] [means for sending order information to the server;] [means for controlling the robot by receiving instructions from the server;] [means for responding to the user by voice;] [means for using generative artificial intelligence to analyze the user's emotions and generate a response based on those emotions;] [means for notifying the user of the delivery status in real time; and [means for collecting post-delivery evaluations based on the user's emotions.] This enables flexible and efficient service provision based on the user's emotions and a consistent improvement in the customer experience during and after delivery.

[2359] "Means for inputting user voice" refers to technology for inputting user speech using a microphone or similar device in restaurants or food delivery systems.

[2360] The "means for converting speech to text" is a speech recognition technology that converts a user's speech data into character string data.

[2361] The "means for analyzing text to identify order details" is a technology for analyzing text data that has been speech-recognized to identify and specify the user's order details.

[2362] The "means for transmitting order information to a server" is a technique for transmitting user order information to a central server via a communication network.

[2363] "Means for controlling a robot by receiving instructions from a server" refers to a control technology that receives instructions from a server, operates the robot, and executes the specified task.

[2364] "Means for responding to the user by voice" refers to a technology for generating an appropriate voice response to the user's speech and transmitting it to the user through a speaker.

[2365] "Means using generative artificial intelligence to analyze a user's emotions and generate responses based on those emotions" refers to artificial intelligence technology that recognizes emotions from the user's voice and automatically generates natural responses based on those emotions.

[2366] "Means for notifying the user of the status during delivery in real time" refers to technology that notifies the user of the current status, estimated arrival time, etc. in real time during food delivery.

[2367] The "means for collecting post-delivery evaluations based on user emotions" is a technology for collecting delivery evaluation feedback based on user emotions after delivery is completed.

[2368] The present invention provides a system for restaurants and food delivery systems that provides flexible responses that take into account user emotions. This system has the following main functions:

[2369] Basic configuration

[2370] The system includes a means for inputting the user's voice, a means for converting the voice into text, a means for analyzing the text to identify the order contents, a means for sending the order information to a server, a means for receiving instructions from the server to control the robot, and a means for responding to the user by voice.

[2371] Additionally, it also includes generative artificial intelligence that analyzes the user's emotions and generates responses based on those emotions, a means of notifying the user of the delivery status in real time, and a means of collecting post-delivery evaluations based on the user's emotions.

[2372] Hardware and Software

[2373] Hardware: Uses the computer's built-in camera, microphone, robot, and speaker. These devices are used to input the user's voice and image, and to output response voices and delivery information.

[2374] software:

[2375] Speech Recognition Module: Uses the speech_recognition library, which converts the user's speech into text data.

[2376] Facial Recognition Software: Recognizes the user's face using the OpenCV library.

[2377] Generative AI: We use emotion recognition and text generation models from the transformers library, specifically the 'michellejieli / emotion_text' and 'gpt-3.5-turbo' models.

[2378] User voice recognition and emotion analysis

[2379] The device uses a camera and microphone to recognize the user, detects the user through facial recognition, and converts the user's speech into text using a speech recognition module. This text data is then input into an emotion recognition model to identify the user's emotions.

[2380] Response Generation

[2381] The generative AI generates an appropriate response based on the user's emotions. For example, if the user's voice is recognized as saying "hello" and the voice contains the emotion of "welcome," the generative AI model will generate a response such as "Welcome, please place your order."...

Claims

1. a means for inputting a user's voice; a means for converting speech to text; a means for analyzing the text to identify the order; means for transmitting order information to a server; A means for receiving instructions from a server and controlling the robot; A means for responding to the user by voice; A system including:

2. 10. The system according to claim 1, wherein the system receives order information and manages cooking status in real time.

3. The system of claim 1, wherein generative artificial intelligence is used to flexibly respond to user requests and questions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A