System
A system using generative AI to analyze customer data and automate order tracking and billing on tablets addresses customer service challenges and staff shortages, enhancing restaurant efficiency and convenience.
Patent Information
- Application Number
- JP2024133507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
In restaurants, customers often face challenges such as employees being unavailable or too busy to respond to additional orders, difficulty in tracking orders, and the hassle of finding a second restaurant, while employees struggle with staff shortages and operational inefficiencies.
A system that collects customer audio and video data, analyzes it using generative AI models to assist and answer questions, recommend a second restaurant, and records orders and bills, utilizing tablets with cameras and microphones to enhance customer service and operational efficiency.
The system improves customer convenience by providing immediate responses and reducing staff shortages, enhancing operational efficiency by automating order tracking and billing, and improving the overall quality of service in restaurants.
Smart Images

Figure 2026030524000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In restaurants, even when customers request additional orders, there is a problem that an employee cannot be found or is too busy to respond. It is also desirable for customers to be able to answer questions immediately while enjoying conversation during their meal, but in many cases, employees are unable to respond adequately. Other issues include the hassle of finding a second restaurant and the difficulty of keeping track of who has ordered what. The present invention aims to improve customer convenience and alleviate the shortage of employees and improve operational efficiency in restaurants. [Means for solving the problem]
[0005] The present invention is a system that includes a means for collecting customer audio and video data, a means for transmitting the collected audio and video data to a server, a means for the server to analyze the audio and video data and provide the analysis results to a generative AI model, a means for the generative AI model to assist the customer in the conversation based on the analysis results, a means for the generative AI model to answer any additional questions from the customer, a means for the generative AI model to recommend and make reservations for a second restaurant, and a means for recording who ordered what and apportioning the bill individually. The system also includes a means for tracking the progress of the customer's meal in real time using image recognition and a means for suggesting additional orders based on the progress. Furthermore, the system includes a means for recognizing customer conversations in multiple languages and providing appropriate information and responses, making it possible to accommodate foreign customers.
[0006] "Customer" refers to a person who visits a restaurant and receives services.
[0007] "Audio and video data" is a form of digital data that includes the customer's speech, video, actions, facial expressions, etc.
[0008] "Collection" is the process of capturing audio and video data using equipment such as cameras and microphones.
[0009] A "server" is a computer system that receives, analyzes, and responds to data over a network.
[0010] A "generative AI model" is an algorithm that uses techniques such as machine learning and natural language processing to analyze input data and generate appropriate responses and suggestions.
[0011] "Analysis" is the process of analyzing collected audio and video data and extracting useful information.
[0012] "Conversation assistance" is the act of analyzing the content of a customer's conversation and providing appropriate responses or supplementary information.
[0013] A "follow-up question" is any further question that the customer asks related to the original question or the content of the conversation.
[0014] "Second restaurant" refers to the restaurant a customer plans to visit after the first one they visited.
[0015] "Introduction and reservation" is the process of providing information about the next store to visit and making a reservation if necessary.
[0016] "Recording" is the act of saving the orders and actions made by customers as data.
[0017] "Individual payment allocation" is a process of allocating the amount so that multiple customers pay for their respective orders.
[0018] The "meal progress" is a state that indicates how much food and drink the customer has consumed.
[0019] "Image recognition" is a technology that analyzes camera footage to automatically recognize specific objects and situations.
[0020] "Multilingual recognition" refers to the ability to parse and understand customer speech in multiple languages.
[0021] "Providing appropriate information and answers" refers to the act of generating and providing optimal information and answers to customer questions.
[0022] "Foreign customers" are customers who speak a language other than the language of the country in which they visit the restaurant. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] System Overview
[0045] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI to assist the customer's conversation, answering any follow-up questions, recommending a second restaurant, and making reservations. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily comprised of a server, a terminal (tablet), and a user (customer).
[0046] System Details
[0047] 1. Setup
[0048] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[0049] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with mechanisms for voice recognition, image recognition, and generative AI models.
[0050] 2. Real-time analysis of customer behavior and eating habits
[0051] Device: The tablet's camera and microphone are used to collect audio and video of the customer in real time. The sensors are periodically checked to ensure they are working properly.
[0052] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[0053] 3. Data analysis and response generation
[0054] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the customer's conversation, and extracts the necessary information.
[0055] Server: Feeds the analysis results to the generative AI model to generate appropriate answers to customer questions and suggestions for additional orders.
[0056] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[0057] Specific processing examples
[0058] Conversation assistance
[0059] 1. User: A customer asks, "What's your recommendation today?"
[0060] 2. Terminal: Collects the question voice and sends it to the server.
[0061] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[0062] 4. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[0063] Reservation for the second restaurant
[0064] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[0065] 2. Terminal: Sends the request to the server.
[0066] 3. Server: Based on the customer's location and desired conditions, the server searches for information on nearby bars and lists available reservations.
[0067] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[0068] 5. User: The customer selects a store and confirms the reservation.
[0069] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[0070] 7. Terminal: Display a confirmation message to the customer.
[0071] Accounting responsibilities
[0072] 1. Terminal: During the meal, you enter additional orders into the terminal and send the details to the server.
[0073] 2. Server: Automatically records who ordered what using image recognition technology.
[0074] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0075] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[0076] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[0077] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[0078] The processing flow will be explained below.
[0079] Collection and analysis of customer audio and video data
[0080] Step 1:
[0081] Device: The tablet activates the camera and microphone to capture the customer's voice and video in real time, capturing details of the customer's movements and conversations.
[0082] Step 2:
[0083] Terminal: The collected audio and video data is compressed and securely sent to the server. The data transfer protocol is encrypted to protect customer privacy.
[0084] Step 3:
[0085] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[0086] Conversation assistance and question handling
[0087] Step 1:
[0088] User: A customer asks, "What's your recommendation today?"
[0089] Step 2:
[0090] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[0091] Step 3:
[0092] Server: Analyzes the received text data, understands the content of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[0093] Step 4:
[0094] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[0095] Step 5:
[0096] Server: Sends the generated answer to the device.
[0097] Step 6:
[0098] Terminal: Receives the answer from the server and displays it to the customer. At the same time, it reads the answer aloud.
[0099] Introduction and reservation of the second restaurant
[0100] Step 1:
[0101] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[0102] Step 2:
[0103] Terminal: Sends the request contents to the server.
[0104] Step 3:
[0105] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[0106] Step 4:
[0107] Server: Sends the listed store information to the terminal.
[0108] Step 5:
[0109] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[0110] Step 6:
[0111] User: The customer selects the desired store and confirms the reservation.
[0112] Step 7:
[0113] Terminal: Sends the selected information to the server.
[0114] Step 8:
[0115] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[0116] Step 9:
[0117] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[0118] Accounting responsibilities
[0119] Step 1:
[0120] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[0121] Step 2:
[0122] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[0123] Step 3:
[0124] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0125] Step 4:
[0126] User: The customer checks the details of each order and enters them into the tablet device.
[0127] Step 5:
[0128] Terminal: Sends input to the server.
[0129] Step 6:
[0130] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[0131] Step 7:
[0132] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[0133] The above are the specific processing steps in the system of the present invention, which can reduce customer waiting times and significantly improve store operational efficiency.
[0134] Example 1
[0135] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0136] Conventional customer service systems in restaurants lack the technology to effectively utilize customer voice and video data to assist conversations, answer additional questions, recommend and reserve the next restaurant, record who ordered what, and allocate individual bills. Therefore, there is a need to improve the quality of customer service and solve the issues of employee shortages and work efficiency.
[0137] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0138] In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to a data processing device, means for the data processing device to analyze the audio and video data and supply the analysis results to a generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to introduce and make reservations for the next store to be visited, and means for recording who ordered what and dividing up the bill individually. This enables a system that improves the quality of customer service and solves issues such as employee shortages and work efficiency.
[0139] "Means for collecting customer audio and video data" refers to equipment installed to collect customer audio and video information in real time, such as devices including cameras and microphones.
[0140] "Means for transmitting collected audio and video data to a data processing device" refers to network communication functions and protocols for transmitting audio and video data collected from customers to a data processing device (server) in a secure manner.
[0141] "Means for analyzing audio and video data using a data processing device and providing the analysis results to a generative AI model" refers to a process and system that analyzes data using speech recognition and image recognition algorithms running on a server and provides the analysis results to a generative AI model.
[0142] "Means for the generative AI model to assist customer conversations based on the analysis results" refers to a function in which the generative AI model provides information to assist customer questions and conversations based on analyzed data.
[0143] "Means for the generative AI model to answer follow-up questions from customers" refers to a system in which the generative AI model generates and provides appropriate answers to further questions from customers.
[0144] "Means for the generative AI model to introduce and make reservations for the next store to visit" is a function in which the generative AI model suggests the next store to visit based on the customer's preferences and carries out the reservation process.
[0145] "A means of recording who ordered what and dividing the bill individually" is a system that records the menu items ordered by customers and calculates and divides the amount due for each customer at the time of payment.
[0146] System Overview
[0147] The system of the present invention uses terminals installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes a generative AI model to assist customer conversations, answer follow-up questions, and recommend and make reservations for the next restaurant to visit. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[0148] Detailed configuration and operation
[0149] set up
[0150] Terminal: A tablet is installed at each customer's seat and equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, enabling communication with the server.
[0151] Server: The server stores customer information, menu information, and information about participating restaurants in a database. It also sets up voice recognition software (e.g., Google Cloud Speech-to-Text), image recognition software (e.g., AWS Rekognition), and generative AI models (e.g., OpenAI GPT-4), and sets the necessary API keys.
[0152] Collection of customer audio and video data
[0153] Device: The device uses the tablet's camera and microphone to collect real-time audio and video data from customers. It periodically checks the operation of sensors and issues alerts if there are any problems.
[0154] Terminal: The collected data is encrypted and sent to the data processing device (server). Security protocols (e.g. SSL / TLS) are used to ensure the safety of the transfer.
[0155] Data Analysis and Response Generation
[0156] Server: The server uses a speech recognition algorithm to convert the received audio data into text, and an image recognition algorithm to analyze the video data.
[0157] Server: The analyzed data is fed into a generative AI model, which generates appropriate answers and suggestions based on the customer's questions and conversations.
[0158] Terminal: Displays generated responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[0159] Specific processing examples
[0160] Conversation assistance
[0161] 1. User: A customer asks, "What's your recommendation today?"
[0162] 2. Device: The device collects the audio of this question and sends it to the server.
[0163] 3. Server: The server analyzes the voice data and generates menu recommendation information using a generative AI model.
[0164] 4. Terminal: The terminal displays the recommended menu information received from the server to the customer and reads it out loud.
[0165] Prompt Sentence Examples
[0166] "What's today's recommended menu?"
[0167] "What's the special tonight?"
[0168] Introduction and reservation of next store to visit
[0169] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[0170] 2. Terminal: The terminal sends the request to the server.
[0171] 3. Server: The server searches for nearby bars based on the customer's location and desired conditions, and lists available bars.
[0172] 4. Terminal: The terminal displays the list of suggested stores to the customer and confirms the reservation.
[0173] 5. User: The customer selects a store and confirms the reservation.
[0174] 6. Server: The server completes the reservation procedure with the selected store and sends a confirmation message to the terminal.
[0175] 7. Terminal: The terminal displays a confirmation message to the customer.
[0176] Prompt Sentence Examples
[0177] "Find recommended bars near you."
[0178] "Can you help me book a bar?"
[0179] Accounting responsibilities
[0180] 1. Terminal: When an additional order is entered during the meal, the terminal sends the details to the server.
[0181] 2. Server: The server uses image recognition technology to automatically record who ordered what.
[0182] 3. Terminal: Once the meal is finished, the terminal provides a screen where customers can enter who ordered what.
[0183] 4. Server: The server calculates the bill and calculates the individual amounts to be shared.
[0184] 5. Terminal: The terminal displays the payment amount for each customer and provides a confirmation screen.
[0185] Example prompts for implementing the invention
[0186] Conversation assistance prompts:
[0187] "What's today's recommended menu?"
[0188] "What's the special tonight?"
[0189] Prompt for next store introduction and reservation:
[0190] "Find recommended bars near you."
[0191] "Can you help me book a bar?"
[0192] Prompt for accounting allocation:
[0193] "Calculate the amount each customer pays."
[0194] "Please split the bill based on your order."
[0195] In this way, the system of the present invention improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component works in cooperation with the others, significantly improving the quality of service in restaurants.
[0196] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0197] Step 1:
[0198] set up
[0199] Devices: Tablets are installed at each customer's seat, and the camera, microphone, and touch screen are set up. The dedicated application is installed and the network connection is checked.
[0200] Input: Initial tablet information
[0201] Output: Tablet configured and network connection confirmed
[0202] Step 2:
[0203] Collection of customer audio and video data
[0204] Device: The tablet's camera and microphone are used to collect audio and video from customers in real time. The collected data is encrypted and sent to a server.
[0205] Input: Customer audio and video
[0206] Output: Encrypted audio and video data
[0207] Step 3:
[0208] Sending data to the server
[0209] Terminal: Collected audio and video data is sent to the server using a security protocol (e.g., SSL / TLS).
[0210] Input: Encrypted audio and video data
[0211] Output: Audio and video data received by the server
[0212] Step 4:
[0213] Audio and video data analysis
[0214] Server: Uses voice recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data, and uses image recognition software (e.g., AWS Rekognition) to analyze the video data.
[0215] Input: Audio and video data received by the server
[0216] Output: Parsed character data and analysis results
[0217] Step 5:
[0218] Feed to generative AI models and generate responses
[0219] Server: Feeds the analysis results to a generative AI model (e.g., OpenAI GPT-4) to generate appropriate answers and suggestions for customer questions.
[0220] Input: Parsed text data and video information
[0221] Output: Generated answers and suggestions
[0222] Step 6:
[0223] Sending the response to the terminal
[0224] Server: Sends generated answers and suggestions to the device.
[0225] Input: Generated answers and suggestions
[0226] Output: Answers and suggestions sent to the terminal
[0227] Step 7:
[0228] Displayed to customers and read aloud
[0229] Terminal: Displays responses and suggestions received from the server to the customer and reads them aloud if necessary.
[0230] Input: Response and proposal received from the server
[0231] Output: Information displayed to the customer and spoken aloud
[0232] Step 8:
[0233] Introducing the next store to visit and making reservations
[0234] 1. User: A customer requests a reservation for their next store visit.
[0235] 2. Terminal: Sends the request to the server.
[0236] 3. Server: Searches for nearby store information based on the customer's location and desired conditions, and lists stores that can be reserved.
[0237] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[0238] 5. User: The customer selects a store and confirms the reservation.
[0239] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[0240] 7. Terminal: Display a confirmation message to the customer.
[0241] Input: Customer request, location, preferences
[0242] Output: Reservation confirmation message
[0243] Step 9:
[0244] Accounting responsibilities
[0245] 1. Terminal: Additional orders are entered during the meal and sent to the server.
[0246] 2. Server: Automatically records who ordered what using image recognition technology.
[0247] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0248] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[0249] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[0250] Input: Customer additional order information, image data
[0251] Output: Recorded order information, calculated billing amount
[0252] (Application example 1)
[0253] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0254] There is a need to solve issues such as improving work efficiency at logistics centers, responding immediately to worker questions, proposing optimal work spots, and managing individual work progress.It is also important to achieve smooth and efficient work by performing these tasks in real time.
[0255] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0256] In this invention, the server includes a means for collecting voice and video data of workers, a means for transmitting the collected voice and video data to the server, a means for the generative AI model to propose optimal work procedures in response to questions from workers, a means for the generative AI model to propose and reserve work spots, and a means for recording who performed which work and reporting individual progress. This makes it possible to improve work efficiency, respond to questions, propose optimal spots, and manage individual progress in real time.
[0257] "Audio and video data of workers" refers to the audio made by workers in a logistics center, as well as video footage of their appearance and actions.
[0258] A "generative AI model" refers to an artificial intelligence system that uses a generative approach to generate diverse outputs based on input data.
[0259] "Distribution center" refers to a facility that receives, stores, ships, and distributes goods and materials.
[0260] "Work procedures" refer to specific instructions and methods for effectively and efficiently carrying out various tasks within a logistics center.
[0261] "Work Spot" refers to a location or area within a distribution center where specific work is performed.
[0262] "Progress report" refers to a report by a worker on the progress or achievement of the work he or she has done.
[0263] System Overview
[0264] The system of the present invention collects voice and video data of workers at a logistics center and transmits it to a server in real time, and then utilizes a generative AI model to assist workers with their questions, suggest work spots, and report on the progress of individual tasks. The system is primarily composed of a server, terminals, and users (workers).
[0265] System Details
[0266] 1. Setup
[0267] Terminal: A robot is installed in each work area within the logistics facility. The robot is equipped with a camera, microphone, and display, and has a dedicated application installed. The robot is capable of communicating with the server.
[0268] Server: The server registers data related to workers, work spots, and work content in a database. It is equipped with mechanisms for running voice recognition, image recognition, and generative AI models.
[0269] 2. Real-time analysis of work status
[0270] Terminal: The robot's camera and microphone collect the worker's voice and video in real time. The sensors periodically check whether the robot is working properly.
[0271] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[0272] 3. Data analysis and response generation
[0273] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the worker's questions, and extracts the necessary information.
[0274] Server: Provides the analysis results to the generation AI model, generating appropriate answers to worker questions and proposing work procedures.
[0275] Terminal: Displays responses and suggestions sent from the server to the worker and reads them out loud if necessary.
[0276] Specific processing examples
[0277] Work procedure suggestions
[0278] 1. User: A worker asks, "Please tell me where to place the next package most efficiently."
[0279] 2. Terminal: Collects the question voice and sends it to the server.
[0280] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[0281] 4. Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[0282] Work Spot Reservation
[0283] 1. User: A worker requests that the next work spot be reserved.
[0284] 2. Terminal: Sends the request to the server.
[0285] 3. Server: Suggests the optimal location based on the work content and information on available work spots within the logistics center.
[0286] 4. Terminal: Shows the worker a list of suggested spots and confirms the reservation.
[0287] 5. User: The worker selects a spot and confirms the reservation.
[0288] 6. Server: Makes a reservation for the selected spot and sends a confirmation message to the terminal.
[0289] 7. Terminal: Display a confirmation message to the worker.
[0290] Progress reports for individual tasks
[0291] 1. Terminal: Enter additional work while working and send the content to the server.
[0292] 2. Server: Records who has done what work and manages progress.
[0293] 3. Terminal: Provides the worker with a progress report screen when the work is completed.
[0294] 4. Server: Organizes the progress of work and reports on individual tasks.
[0295] 5. Terminal: Displays the progress of each worker and provides a confirmation screen.
[0296] Hardware and software used
[0297] Camera and microphone: Installed on the robot to collect the worker's voice and video.
[0298] Robot: Moves around the distribution center and is positioned in each designated work area.
[0299] Server: Runs speech and image recognition and generative AI models (e.g., GPT-3.5-turbo).
[0300] Prompt Sentence Examples
[0301] "Please tell me where to place the next shipment for maximum efficiency."
[0302] "Please reserve the next work spot for me."
[0303] In this way, this system aims to improve work efficiency at logistics centers and provides real-time assistance to workers.
[0304] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0305] Step 1:
[0306] User: A worker asks, "Please tell me where to place the next package most efficiently."
[0307] Input: The worker's spoken question.
[0308] Output: Audio data is input to the robot's microphone.
[0309] Step 2:
[0310] Terminal: Collects voice questions and sends them to the server.
[0311] Input: Worker voice data.
[0312] Data processing: Audio data captured by the robot's microphone is converted into a digital signal and sent to a server via a secure communication protocol.
[0313] Output: Audio data transferred to the server.
[0314] Step 3:
[0315] Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[0316] Input: The transmitted audio data.
[0317] Data computation: Speech data is converted into text using a speech recognition algorithm, and then fed into a generative AI model (e.g., GPT-3.5-turbo) for analysis.
[0318] Output: Text data containing optimal work procedure information.
[0319] Step 4:
[0320] Server: Based on the analysis results, the AI model generates work procedures and sends them to the device.
[0321] Input: Text data containing optimal work procedure information.
[0322] Data calculation: The generative AI model generates appropriate work procedures based on the analysis results.
[0323] Output: Work procedure information sent to the robot's terminal.
[0324] Step 5:
[0325] Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[0326] Input: Work procedure information sent from the server.
[0327] Data processing: Work procedure information is displayed on the robot's display and read aloud using a speech synthesis algorithm.
[0328] Output: Visual and audio feedback to the worker.
[0329] Step 6:
[0330] User: A worker requests to reserve the next work spot.
[0331] Input: A voice request to reserve a work spot.
[0332] Output: Audio data is input to the robot's microphone.
[0333] Step 7:
[0334] Terminal: Sends the request to the server.
[0335] Input: Worker voice data.
[0336] Data processing: Converts voice data into digital signals and sends them to a server via a secure communication protocol.
[0337] Output: Audio data transferred to the server.
[0338] Step 8:
[0339] Server: Suggests the optimal location based on the work content and information on available work spots within the distribution center.
[0340] Input: Voice request data from workers and work spot data within the distribution center.
[0341] Data Computation: Uses speech recognition algorithms to convert voice data into text, search for available work spots, and use generative AI models to suggest the best locations.
[0342] Output: Text data containing the suggested work spot information.
[0343] Step 9:
[0344] Terminal: Displays the proposed spot list received from the server to the worker and confirms the reservation.
[0345] Input: Suggested spot information sent from the server.
[0346] Data processing: Display a list of suggested spots on the robot's display and provide an interface for reservation confirmation.
[0347] Output: A confirmation screen that is displayed to the worker.
[0348] Step 10:
[0349] User: The worker selects a spot and confirms the reservation.
[0350] Input: Work spot information selected by the worker.
[0351] Output: The input selection information is obtained through the robot interface.
[0352] Step 11:
[0353] Terminal: Sends the selected information to the server.
[0354] Input: Worker selection information.
[0355] Data processing: Selected data is converted into a digital signal and sent to the server via a secure communication protocol.
[0356] Output: The selected data that is transferred to the server.
[0357] Step 12:
[0358] Server: Makes a reservation at the selected spot and sends a confirmation message to the device.
[0359] Input: Worker selection information.
[0360] Data calculation: Executes the reservation procedure and generates a confirmation message.
[0361] Output: Sends a confirmation message to the robot's terminal.
[0362] Step 13:
[0363] Terminal: Display a confirmation message to the worker.
[0364] Input: The confirmation message sent by the server.
[0365] Data processing: Display the message on the robot's display.
[0366] Output: A confirmation message that is displayed to the worker.
[0367] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0368] System Overview
[0369] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI and an emotion engine to assist customer conversations, answer follow-up questions, and recommend and make reservations for second restaurants. It also records who ordered what and has the ability to allocate the bill individually, recognizing user emotions in real time to improve the quality of service. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[0370] System Details
[0371] 1. Setup
[0372] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[0373] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with a system that runs voice recognition, image recognition, generative AI models, and an emotion engine.
[0374] 2. Real-time analysis of customer behavior and eating habits
[0375] Device: The tablet's camera and microphone are used to collect the customer's voice and video in real time.
[0376] Terminal: Collected data is sent to the server. The data transfer protocol is encrypted and secure.
[0377] 3. Data analysis and response generation
[0378] Server: Analyzes the received audio and video data using voice and image recognition algorithms, analyzes the customer's conversation and the progress of their meal, and extracts the necessary information.
[0379] Server: Feeds the analysis results to the generative AI model and emotion engine to generate appropriate answers to customer questions and suggestions for additional orders.
[0380] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[0381] 4. Use of Emotion Engine
[0382] Server: The emotion engine analyzes the customer's facial expressions and tone of voice based on collected audio and video data to grasp their emotional state (e.g., joy, displeasure, interest, anger, etc.) in real time.
[0383] Server: Based on the analysis results of the emotion engine, the generative AI model adjusts the responses and assistance provided accordingly. For example, if the customer shows signs of discomfort, the server softens the response and refrains from making meal suggestions.
[0384] Specific processing examples
[0385] Conversation assistance
[0386] 1. User: A customer asks, "What's your recommendation today?"
[0387] 2. Terminal: Collects voice questions, converts them into text, and sends them to the server.
[0388] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[0389] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the generated response appropriately (e.g., adding more information if the customer is interested).
[0390] 5. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[0391] Reservation for the second restaurant
[0392] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[0393] 2. Terminal: Sends the request to the server.
[0394] 3. Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. It lists available establishments and retrieves detailed information from the database.
[0395] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations (e.g., if the customer indicates they want to relax, prioritize quiet bars).
[0396] 5. Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[0397] 6. User: The customer selects the desired store and confirms the reservation.
[0398] 7. Terminal: Sends the selection information to the server.
[0399] 8. Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[0400] 9. Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[0401] Accounting responsibilities
[0402] 1. Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[0403] 2. Server: Automatically records who ordered what. Using image recognition technology, it automatically tracks who placed what order.
[0404] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0405] 4. User: The customer checks the individual order details and enters them into the tablet device.
[0406] 5. Terminal: Sends the input to the server.
[0407] 6. Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user contributes based on the items ordered.
[0408] 7. Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each person's payment.
[0409] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[0410] The processing flow will be explained below.
[0411] Collection and analysis of customer audio and video data
[0412] Step 1:
[0413] Device: The tablet activates the camera and microphone to collect the customer's voice and video in real time, capturing the content of the conversation and the customer's behavior.
[0414] Step 2:
[0415] Terminal: Collected audio and video data is temporarily stored in internal memory and compressed, allowing for efficient data transfer.
[0416] Step 3:
[0417] Terminal: Compressed audio and video data is sent to the server. The data is encrypted to prevent information leakage during transmission.
[0418] Step 4:
[0419] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[0420] Conversation assistance and question handling
[0421] Step 1:
[0422] User: A customer asks, "What's your recommendation today?"
[0423] Step 2:
[0424] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[0425] Step 3:
[0426] Server: Analyzes the received text data, understands the intent of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[0427] Step 4:
[0428] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[0429] Step 5:
[0430] Server: Uses analyzed customer sentiment data to adjust the tone and content of the responses it generates, for example, providing more detailed information if the customer is interested.
[0431] Step 6:
[0432] Server: Sends the generated answer to the device.
[0433] Step 7:
[0434] Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[0435] Introduction and reservation of the second restaurant
[0436] Step 1:
[0437] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[0438] Step 2:
[0439] Terminal: Sends the request contents to the server.
[0440] Step 3:
[0441] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[0442] Step 4:
[0443] Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations accordingly. For example, if a customer indicates that they want to relax, it will prioritize quiet bars.
[0444] Step 5:
[0445] Server: Sends the listed store information to the terminal.
[0446] Step 6:
[0447] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[0448] Step 7:
[0449] User: The customer selects the desired store and confirms the reservation.
[0450] Step 8:
[0451] Terminal: Sends the selected information to the server.
[0452] Step 9:
[0453] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[0454] Step 10:
[0455] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[0456] Accounting responsibilities
[0457] Step 1:
[0458] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[0459] Step 2:
[0460] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[0461] Step 3:
[0462] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0463] Step 4:
[0464] User: The customer checks the details of each order and enters them into the tablet device.
[0465] Step 5:
[0466] Terminal: Sends input to the server.
[0467] Step 6:
[0468] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[0469] Step 7:
[0470] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[0471] Step 8:
[0472] Server: The emotion engine analyzes the customer's emotional state at checkout and adjusts payment methods and responses as needed.
[0473] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[0474] Example 2
[0475] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0476] Modern restaurants are required to improve customer convenience while simultaneously increasing operational efficiency. However, the number of employees is limited, and it is often difficult to provide adequate service, especially during busy periods. It is also difficult to grasp customers' emotional state in real time and provide service accordingly. Furthermore, it is time-consuming to accurately record who ordered what and assign individual tasks at the checkout. A system is needed to solve these issues and improve customer satisfaction.
[0477] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for encrypting the customer's voice and video data and transmitting it to the server, means for analyzing the data using voice recognition and image recognition algorithms and providing the analysis results to the generative AI model and emotion analysis engine, and means for the generative AI model to assist the customer in the conversation based on the analysis results. This enables secure transfer and analysis of the customer's voice and video data, immediate conversation support based on the analysis results, and adjustment of response content according to emotions. It also includes functions for introducing and making reservations for the next store and sharing individual billing responsibilities, greatly improving customer convenience.
[0478] "Audio and video data" refers to digital data that records the voice and video of a customer.
[0479] "Encryption" is a technology that converts data into a form that cannot be understood by third parties in order to transfer it securely.
[0480] A "server" is a computer system that handles back-end operations such as data management, analysis, and communication.
[0481] A "voice recognition algorithm" is a program for converting voice data into text data.
[0482] An "image recognition algorithm" is a program that analyzes video data to detect specific patterns or objects.
[0483] A "generative AI model" is an artificial intelligence model that generates natural language responses or suggestions based on input data.
[0484] An "emotion analysis engine" is a program for analyzing an individual's emotional state from audio and video data.
[0485] "Conversational assistance" means providing appropriate responses and suggestions to customer questions and requests.
[0486] The "means for answering additional questions" is a function for generating answers to additional questions from customers.
[0487] "Next store introduction and reservation" is a function that suggests the next store that the customer would like to visit and makes a reservation for it.
[0488] The "means for dividing the bill amount individually" is a function that records the order details for each customer and calculates the individual payment amount based on that.
[0489] The system of the present invention is designed to improve customer convenience in restaurants and is composed of multiple elements including terminals (tablets), servers, and users (customers). Each element will be described in detail below.
[0490] 1. Setup
[0491] Device:
[0492] The tablet is installed at the customer's seat and is equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, and it can communicate with the server. The tablet uses a Wi-Fi connection, and plugins and drivers are pre-configured. Normal operation is initiated by clicking the "Start" button on the main screen.
[0493] server:
[0494] The server registers customer information, menus, and data on affiliated restaurants in a database. The database uses MySQL, and an environment is set up in which voice recognition, image recognition, generative AI models, and an emotion analysis engine can operate. The necessary software and libraries are installed on the server in advance.
[0495] 2. Real-time analysis of customer behavior and eating habits
[0496] Device:
[0497] The tablet's built-in camera and microphone are used to collect customer audio and video data in real time. For example, when a customer speaks into the tablet, their audio and video are recorded.
[0498] Device:
[0499] The collected data is encrypted using the AES encryption algorithm and transmitted to the server via the HTTPS protocol, which ensures the security of the data.
[0500] 3. Data analysis and response generation
[0501] server:
[0502] The server uses a speech recognition algorithm (e.g., Reve) to convert the voice data into text. For example, the voice data "What's the recommendation today?" is converted into text data "What's the recommendation today?"
[0503] server:
[0504] Image recognition algorithms (e.g., OpenCV) are used to analyze video data and grasp customer facial expressions and movements in real time.
[0505] server:
[0506] A generative AI model (e.g., GPT-3) generates an appropriate response based on the analysis results. For example, in response to the text data "What's today's recommendation?", it generates a response such as "Today's recommendation is seafood pasta."
[0507] Device:
[0508] The response data from the server is sent to the tablet, displayed on the screen, and read aloud using a Text-to-Speech (TTS) engine.
[0509] 4. Use of sentiment analysis engines
[0510] server:
[0511] The emotion analysis engine analyzes the customer's emotional state from audio and video data. For example, if the customer is smiling, it is determined to be "happy."
[0512] server:
[0513] The generative AI model adjusts responses based on the results of sentiment analysis. For example, if a customer expresses displeasure, the next response will be generated in a softer tone. For example, a response such as, "May I suggest some other dishes so you can enjoy them?"
[0514] Examples of concrete examples and prompts
[0515] Conversation Assistance:
[0516] When a user asks, "What's the recommendation today?", the device collects the voice and sends it to the server, which analyzes the voice and generates a response such as, "Today's recommendation is seafood pasta."
[0517] 2nd restaurant reservation:
[0518] When a customer requests a reservation for the next bar they want to visit, the server will suggest a bar based on the customer's desired conditions and make the reservation.
[0519] Example prompt sentence:
[0520] "When a customer asks, 'What's the special today?' show them the current menu specials and read out the details."
[0521] With the above configuration, the present invention significantly improves customer convenience and improves the operational efficiency of restaurants. In particular, by combining it with an emotion analysis engine, it becomes possible to respond flexibly to customer emotions, thereby providing more personalized service.
[0522] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0523] Step 1:
[0524] Device:
[0525] A tablet is placed at the customer's seat, and the camera, microphone, and touchscreen are set up. A dedicated application is then launched. When the customer speaks into the tablet, the camera and microphone collect audio and video. The collected data is processed in real time, resulting in minimal latency.
[0526] Input: Customer audio and video data
[0527] Output: Encrypted audio and video data
[0528] For example, a customer asks, "What's the recommendation today?" Audio and video are captured and the application encrypts the data.
[0529] Step 2:
[0530] Device:
[0531] The collected data is encrypted using the AES encryption algorithm and sent to the server via the HTTPS protocol, ensuring data security and privacy.
[0532] Input: Unencrypted audio and video data
[0533] Output: Encrypted data sent to the server
[0534] As a specific example of operation, the encrypted data is sent to the server, and a message indicating that the transfer is complete is displayed.
[0535] Step 3:
[0536] server:
[0537] The voice data that arrives at the server is analyzed using a voice recognition algorithm and converted into text data. For example, the voice data "What's recommended today?" is converted into text.
[0538] Input: Encrypted audio data
[0539] Output: Text data
[0540] As a specific example of how it works, the voice recognition algorithm is activated and the result "What's recommended today?" is output.
[0541] Step 4:
[0542] server:
[0543] Image recognition algorithms are used to analyze the customer's facial expressions and movements from the incoming video data, thereby capturing the customer's real-time emotional state.
[0544] Input: Encrypted video data
[0545] Output: Emotional state data
[0546] As a specific example of operation, an image recognition algorithm is run to determine that "the customer is showing interest."
[0547] Step 5:
[0548] server:
[0549] A generative AI model is used to generate an appropriate response based on the analyzed text data and emotional state, for example, "Today's recommendation is seafood pasta."
[0550] Input: Text data and emotional state data
[0551] Output: Response message
[0552] As a specific example of how it works, the generative AI model generates the message "Today's recommendation is seafood pasta."
[0553] Step 6:
[0554] server:
[0555] The generated response message is sent to the terminal.
[0556] Input: Response message
[0557] Output: Message sent to the terminal
[0558] As a specific example of operation, a response message is sent from the server to the terminal.
[0559] Step 7:
[0560] Device:
[0561] The terminal receives the response message, displays it on the screen, and uses a Text-to-Speech (TTS) engine to read it aloud to the customer.
[0562] Input: Received response message
[0563] Output: Voice and text output to the customer
[0564] As a specific example of how it works, the message "Today's recommendation is seafood pasta" is displayed on the screen and simultaneously read aloud.
[0565] By having each step work in conjunction with each other, we aim to create a system that can respond to customers quickly and appropriately.
[0566] (Application example 2)
[0567] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0568] In recent years, restaurants have been required to improve the quality of customer service and streamline operations. However, staff shortages and inconsistent service quality have become problems. In particular, there are limitations to providing multilingual support and individualized service in real time, making it difficult to maintain customer satisfaction. It is also difficult to allocate individual amounts at the time of payment and provide service that responds to customer emotions.
[0569] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to the server, means for the server to analyze the audio and video data and provide the analysis results to the generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to recommend and reserve the next restaurant, means for recording who ordered what and individually apportioning the bill, means for analyzing the customer's emotional state and adjusting the quality of service, means for displaying real-time video of the customer on a smartphone or other mobile device, and means for encrypting the collected data and securely transferring it to the server. This improves the quality of customer service and enables efficient business operations. It is also possible to provide multilingual support, real-time personalized services, and flexible service provision based on individual apportionment of the bill and emotions.
[0570] "Customer audio and video data" refers to audio produced by the customer and video information such as the customer's actions and facial expressions.
[0571] "Collecting means" refers to devices and methods for capturing audio and video data.
[0572] "Transmission means" refers to the communications technologies and protocols used to transmit collected data to the server.
[0573] "Server" refers to the computer system that analyzes collected data and runs the generative AI model and emotion engine.
[0574] "Means for analysis" refers to the algorithms and software used to analyze audio and video data and extract the necessary information.
[0575] "Generative AI model" refers to an artificial intelligence algorithm or program that generates responses or suggestions based on data it receives.
[0576] "Means for assisting conversation" refers to the method by which the generative AI model assists customer conversations based on the analysis results.
[0577] "Means for answering additional questions" refers to a method for generating and providing appropriate answers to new questions from customers.
[0578] "Means for introducing and making reservations at the next store" refers to a method for suggesting another store to the customer and making a reservation at that store if necessary.
[0579] "Means of recording who ordered what and allocating the bill individually" refers to a method of recording the order details by linking them to a specific customer and ultimately calculating the individual payment amount.
[0580] "Means for analyzing emotional state and adjusting quality of service" refers to a method for analyzing the emotions of a customer and providing service accordingly.
[0581] "Means for displaying real-time video on a smartphone or other mobile device" refers to methods and technologies for capturing video of a customer in real time and displaying it on a smartphone or other mobile device.
[0582] "Means for encrypting and securely transferring data to the server" refers to the technology and methods for encrypting collected data and transmitting it securely to the server while preventing unauthorized access from outside.
[0583] System Overview
[0584] The system of the present invention collects customer voice and video data and analyzes it on a server to improve the quality of service. The main components of the system are a server, a terminal, and a user, and is realized using various devices and software.
[0585] Hardware and Software Configuration
[0586] The main components running on the server and terminals are:
[0587] Hardware
[0588] Device: A mobile device such as a tablet or smartphone equipped with a camera and microphone to collect customer audio and video data.
[0589] Server: A powerful computer system responsible for analyzing data and running generative AI models.
[0590] Software and APIs
[0591] Speech Recognition API: Google Cloud Speech-to-Text API
[0592] Image Recognition API: Amazon Rekognition
[0593] Generative AI model: GPT-4 (OpenAI)
[0594] Emotion engine: IBM Watson Tone Analyzer
[0595] Database: Firebase Realtime Database
[0596] Communication protocol: HTTPS (SSL / TLS encryption)
[0597] Data processing and analysis flow
[0598] 1. Data Collection:
[0599] The device's camera and microphone are used to collect customer audio and video data.
[0600] The collected data is temporarily stored on the device.
[0601] 2. Data Transfer:
[0602] The collected data is sent to the server in real time using the HTTPS protocol.
[0603] Data is transferred securely via SSL / TLS encryption.
[0604] 3. Data Analysis:
[0605] The server converts the received voice data into text using the Google Cloud Speech-to-Text API.
[0606] The video data is analyzed using Amazon Rekognition to recognize customer facial expressions and behavior.
[0607] The analysis results are fed into a generative AI model (GPT-4) to generate responses and suggestions to customer questions and requests.
[0608] The emotion engine (IBM Watson Tone Analyzer) analyzes the customer's emotional state and adjusts the response content and service delivery method.
[0609] Provision of services
[0610] 1. Conversation assistance:
[0611] The server uses a generative AI model to generate appropriate answers to customer questions, such as suggesting menu recommendations in response to the question, "What's your recommendation today?"
[0612] The generated answers are displayed on the device screen and, if necessary, read aloud.
[0613] 2. Introduction and reservation of the following stores:
[0614] Based on the customer's request, the server searches for and suggests the next restaurant. The server processes the reservation and sends a confirmation message to the terminal.
[0615] 3. Accounting responsibilities:
[0616] The terminal records who has ordered what, and the server calculates the individual payment amount, which is displayed on the terminal for each customer to see.
[0617] Specific examples
[0618] For example, if a customer asks, "What dish do you recommend?" the system will process it as follows:
[0619] Prompt Sentence Examples
[0620] A customer asks, "What dish do you recommend?" Use the following information to generate a suitable answer:
[0621] Today's recommended dishes are "Omurice," "Steak," and "Salad."
[0622] The customer's emotional state is "interested."
[0623] answer:
[0624] Using this prompt, the generative AI model generates a response such as, "Today's recommended dishes are omelet rice, steak, and salad. They're all delicious. Let me know if you need more information," and displays it on the device. It also decides whether to provide more detailed information depending on the customer's emotional state.
[0625] The above is a specific embodiment for carrying out the invention. This system improves the quality of customer service and enables efficient business operations.
[0626] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0627] Step 1:
[0628] The user turns on the device (smartphone or tablet) and launches the dedicated application. The user then enables the camera and microphone. The camera and microphone collect the customer's audio and video data in real time. The customer's audio and video are captured as input and converted into digital data within the application.
[0629] Step 2:
[0630] The collected audio and video data is sent from the device to the server using the SSL / TLS encrypted HTTPS protocol. This data is securely received by the server. The collected audio and video data is sent as input, reaches the server, is safely received, and is output.
[0631] Step 3:
[0632] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. Voice data is given as input, and it is analyzed and converted into text data. The output is text data.
[0633] Step 4:
[0634] The server analyzes the received video data using the Amazon Rekognition API to recognize the customer's facial expressions and behavior. The video data is given as input, and it is analyzed to output information about facial expressions and behavior.
[0635] Step 5:
[0636] The analyzed audio and video data is integrated on the server and supplied to a generative AI model (GPT-4). The generative AI model creates prompts to generate appropriate answers to customer questions and requests. The textual customer utterances and emotional state are given as input, and an appropriate answer text is generated as output.
[0637] Step 6:
[0638] The server uses IBM Watson Tone Analyzer to analyze the customer's emotional state based on their text data. The generative AI model adjusts the answers and suggestions it provides based on their emotional state. The customer's text data is given as input, and the emotional state is obtained as output.
[0639] Step 7:
[0640] The generated response or suggestion is then re-encrypted and sent to the device, where it is received and displayed on the application screen and optionally read aloud. The generated response text is given as input, and the output is the answer appropriate to the customer's request.
[0641] Step 8:
[0642] Based on a specific need, if a customer requests "I would like to reserve the next store," the server receives the request, searches for an appropriate store based on the location information and desired conditions, and performs the reservation procedure. The customer's request is given as input, and a reservation confirmation message is generated as output.
[0643] Step 9:
[0644] When a customer places an order, the terminal sends the details to the server, which automatically records who ordered what. After the meal is finished, the server displays the bill based on the recorded order data and allocates it to each customer. The order data is given as input, and the individual payment amounts are calculated as output.
[0645] Step 10:
[0646] Based on the customer's emotional state, the server adjusts the response content and quality of service to provide optimal service. The emotion analysis results are given as input, and the adjusted service content is provided as output.
[0647] In this way, the system can improve customer service in real time and also improve operational efficiency.
[0648] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0649] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0650] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0651] [Second embodiment]
[0652] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0653] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0654] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0655] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0656] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0657] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0658] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0659] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0660] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0661] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0662] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0663] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0664] System Overview
[0665] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI to assist the customer's conversation, answering any follow-up questions, recommending a second restaurant, and making reservations. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily comprised of a server, a terminal (tablet), and a user (customer).
[0666] System Details
[0667] 1. Setup
[0668] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[0669] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with mechanisms for voice recognition, image recognition, and generative AI models.
[0670] 2. Real-time analysis of customer behavior and eating habits
[0671] Device: The tablet's camera and microphone are used to collect audio and video of the customer in real time. The sensors are periodically checked to ensure they are working properly.
[0672] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[0673] 3. Data analysis and response generation
[0674] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the customer's conversation, and extracts the necessary information.
[0675] Server: Feeds the analysis results to the generative AI model to generate appropriate answers to customer questions and suggestions for additional orders.
[0676] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[0677] Specific processing examples
[0678] Conversation assistance
[0679] 1. User: A customer asks, "What's your recommendation today?"
[0680] 2. Terminal: Collects the question voice and sends it to the server.
[0681] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[0682] 4. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[0683] Reservation for the second restaurant
[0684] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[0685] 2. Terminal: Sends the request to the server.
[0686] 3. Server: Based on the customer's location and desired conditions, the server searches for information on nearby bars and lists available reservations.
[0687] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[0688] 5. User: The customer selects a store and confirms the reservation.
[0689] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[0690] 7. Terminal: Display a confirmation message to the customer.
[0691] Accounting responsibilities
[0692] 1. Terminal: During the meal, you enter additional orders into the terminal and send the details to the server.
[0693] 2. Server: Automatically records who ordered what using image recognition technology.
[0694] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0695] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[0696] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[0697] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[0698] The processing flow will be explained below.
[0699] Collection and analysis of customer audio and video data
[0700] Step 1:
[0701] Device: The tablet activates the camera and microphone to capture the customer's voice and video in real time, capturing details of the customer's movements and conversations.
[0702] Step 2:
[0703] Terminal: The collected audio and video data is compressed and securely sent to the server. The data transfer protocol is encrypted to protect customer privacy.
[0704] Step 3:
[0705] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[0706] Conversation assistance and question handling
[0707] Step 1:
[0708] User: A customer asks, "What's your recommendation today?"
[0709] Step 2:
[0710] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[0711] Step 3:
[0712] Server: Analyzes the received text data, understands the content of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[0713] Step 4:
[0714] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[0715] Step 5:
[0716] Server: Sends the generated answer to the device.
[0717] Step 6:
[0718] Terminal: Receives the answer from the server and displays it to the customer. At the same time, it reads the answer aloud.
[0719] Introduction and reservation of the second restaurant
[0720] Step 1:
[0721] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[0722] Step 2:
[0723] Terminal: Sends the request contents to the server.
[0724] Step 3:
[0725] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[0726] Step 4:
[0727] Server: Sends the listed store information to the terminal.
[0728] Step 5:
[0729] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[0730] Step 6:
[0731] User: The customer selects the desired store and confirms the reservation.
[0732] Step 7:
[0733] Terminal: Sends the selected information to the server.
[0734] Step 8:
[0735] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[0736] Step 9:
[0737] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[0738] Accounting responsibilities
[0739] Step 1:
[0740] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[0741] Step 2:
[0742] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[0743] Step 3:
[0744] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0745] Step 4:
[0746] User: The customer checks the details of each order and enters them into the tablet device.
[0747] Step 5:
[0748] Terminal: Sends input to the server.
[0749] Step 6:
[0750] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[0751] Step 7:
[0752] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[0753] The above are the specific processing steps in the system of the present invention, which can reduce customer waiting times and significantly improve store operational efficiency.
[0754] Example 1
[0755] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0756] Conventional customer service systems in restaurants lack the technology to effectively utilize customer voice and video data to assist conversations, answer additional questions, recommend and reserve the next restaurant, record who ordered what, and allocate individual bills. Therefore, there is a need to improve the quality of customer service and solve the issues of employee shortages and work efficiency.
[0757] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0758] In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to a data processing device, means for the data processing device to analyze the audio and video data and supply the analysis results to a generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to introduce and make reservations for the next store to be visited, and means for recording who ordered what and dividing up the bill individually. This enables a system that improves the quality of customer service and solves issues such as employee shortages and work efficiency.
[0759] "Means for collecting customer audio and video data" refers to equipment installed to collect customer audio and video information in real time, such as devices including cameras and microphones.
[0760] "Means for transmitting collected audio and video data to a data processing device" refers to network communication functions and protocols for transmitting audio and video data collected from customers to a data processing device (server) in a secure manner.
[0761] "Means for analyzing audio and video data using a data processing device and providing the analysis results to a generative AI model" refers to a process and system that analyzes data using speech recognition and image recognition algorithms running on a server and provides the analysis results to a generative AI model.
[0762] "Means for the generative AI model to assist customer conversations based on the analysis results" refers to a function in which the generative AI model provides information to assist customer questions and conversations based on analyzed data.
[0763] "Means for the generative AI model to answer follow-up questions from customers" refers to a system in which the generative AI model generates and provides appropriate answers to further questions from customers.
[0764] "Means for the generative AI model to introduce and make reservations for the next store to visit" is a function in which the generative AI model suggests the next store to visit based on the customer's preferences and carries out the reservation process.
[0765] "A means of recording who ordered what and dividing the bill individually" is a system that records the menu items ordered by customers and calculates and divides the amount due for each customer at the time of payment.
[0766] System Overview
[0767] The system of the present invention uses terminals installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes a generative AI model to assist customer conversations, answer follow-up questions, and recommend and make reservations for the next restaurant to visit. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[0768] Detailed configuration and operation
[0769] set up
[0770] Terminal: A tablet is installed at each customer's seat and equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, enabling communication with the server.
[0771] Server: The server stores customer information, menu information, and information about participating restaurants in a database. It also sets up voice recognition software (e.g., Google Cloud Speech-to-Text), image recognition software (e.g., AWS Rekognition), and generative AI models (e.g., OpenAI GPT-4), and sets the necessary API keys.
[0772] Collection of customer audio and video data
[0773] Device: The device uses the tablet's camera and microphone to collect real-time audio and video data from customers. It periodically checks the operation of sensors and issues alerts if there are any problems.
[0774] Terminal: The collected data is encrypted and sent to the data processing device (server). Security protocols (e.g. SSL / TLS) are used to ensure the safety of the transfer.
[0775] Data Analysis and Response Generation
[0776] Server: The server uses a speech recognition algorithm to convert the received audio data into text, and an image recognition algorithm to analyze the video data.
[0777] Server: The analyzed data is fed into a generative AI model, which generates appropriate answers and suggestions based on the customer's questions and conversations.
[0778] Terminal: Displays generated responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[0779] Specific processing examples
[0780] Conversation assistance
[0781] 1. User: A customer asks, "What's your recommendation today?"
[0782] 2. Device: The device collects the audio of this question and sends it to the server.
[0783] 3. Server: The server analyzes the voice data and generates menu recommendation information using a generative AI model.
[0784] 4. Terminal: The terminal displays the recommended menu information received from the server to the customer and reads it out loud.
[0785] Prompt Sentence Examples
[0786] "What's today's recommended menu?"
[0787] "What's the special tonight?"
[0788] Introduction and reservation of next store to visit
[0789] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[0790] 2. Terminal: The terminal sends the request to the server.
[0791] 3. Server: The server searches for nearby bars based on the customer's location and desired conditions, and lists available bars.
[0792] 4. Terminal: The terminal displays the list of suggested stores to the customer and confirms the reservation.
[0793] 5. User: The customer selects a store and confirms the reservation.
[0794] 6. Server: The server completes the reservation procedure with the selected store and sends a confirmation message to the terminal.
[0795] 7. Terminal: The terminal displays a confirmation message to the customer.
[0796] Prompt Sentence Examples
[0797] "Find recommended bars near you."
[0798] "Can you help me book a bar?"
[0799] Accounting responsibilities
[0800] 1. Terminal: When an additional order is entered during the meal, the terminal sends the details to the server.
[0801] 2. Server: The server uses image recognition technology to automatically record who ordered what.
[0802] 3. Terminal: Once the meal is finished, the terminal provides a screen where customers can enter who ordered what.
[0803] 4. Server: The server calculates the bill and calculates the individual amounts to be shared.
[0804] 5. Terminal: The terminal displays the payment amount for each customer and provides a confirmation screen.
[0805] Example prompts for implementing the invention
[0806] Conversation assistance prompts:
[0807] "What's today's recommended menu?"
[0808] "What's the special tonight?"
[0809] Prompt for next store introduction and reservation:
[0810] "Find recommended bars near you."
[0811] "Can you help me book a bar?"
[0812] Prompt for accounting allocation:
[0813] "Calculate the amount each customer pays."
[0814] "Please split the bill based on your order."
[0815] In this way, the system of the present invention improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component works in cooperation with the others, significantly improving the quality of service in restaurants.
[0816] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0817] Step 1:
[0818] set up
[0819] Devices: Tablets are installed at each customer's seat, and the camera, microphone, and touch screen are set up. The dedicated application is installed and the network connection is checked.
[0820] Input: Initial tablet information
[0821] Output: Tablet configured and network connection confirmed
[0822] Step 2:
[0823] Collection of customer audio and video data
[0824] Device: The tablet's camera and microphone are used to collect audio and video from customers in real time. The collected data is encrypted and sent to a server.
[0825] Input: Customer audio and video
[0826] Output: Encrypted audio and video data
[0827] Step 3:
[0828] Sending data to the server
[0829] Terminal: Collected audio and video data is sent to the server using a security protocol (e.g., SSL / TLS).
[0830] Input: Encrypted audio and video data
[0831] Output: Audio and video data received by the server
[0832] Step 4:
[0833] Audio and video data analysis
[0834] Server: Uses voice recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data, and uses image recognition software (e.g., AWS Rekognition) to analyze the video data.
[0835] Input: Audio and video data received by the server
[0836] Output: Parsed character data and analysis results
[0837] Step 5:
[0838] Feed to generative AI models and generate responses
[0839] Server: Feeds the analysis results to a generative AI model (e.g., OpenAI GPT-4) to generate appropriate answers and suggestions for customer questions.
[0840] Input: Parsed text data and video information
[0841] Output: Generated answers and suggestions
[0842] Step 6:
[0843] Sending the response to the terminal
[0844] Server: Sends generated answers and suggestions to the device.
[0845] Input: Generated answers and suggestions
[0846] Output: Answers and suggestions sent to the terminal
[0847] Step 7:
[0848] Displayed to customers and read aloud
[0849] Terminal: Displays responses and suggestions received from the server to the customer and reads them aloud if necessary.
[0850] Input: Response and proposal received from the server
[0851] Output: Information displayed to the customer and spoken aloud
[0852] Step 8:
[0853] Introducing the next store to visit and making reservations
[0854] 1. User: A customer requests a reservation for their next store visit.
[0855] 2. Terminal: Sends the request to the server.
[0856] 3. Server: Searches for nearby store information based on the customer's location and desired conditions, and lists stores that can be reserved.
[0857] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[0858] 5. User: The customer selects a store and confirms the reservation.
[0859] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[0860] 7. Terminal: Display a confirmation message to the customer.
[0861] Input: Customer request, location, preferences
[0862] Output: Reservation confirmation message
[0863] Step 9:
[0864] Accounting responsibilities
[0865] 1. Terminal: Additional orders are entered during the meal and sent to the server.
[0866] 2. Server: Automatically records who ordered what using image recognition technology.
[0867] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[0868] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[0869] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[0870] Input: Customer additional order information, image data
[0871] Output: Recorded order information, calculated billing amount
[0872] (Application example 1)
[0873] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0874] There is a need to solve issues such as improving work efficiency at logistics centers, responding immediately to worker questions, proposing optimal work spots, and managing individual work progress.It is also important to achieve smooth and efficient work by performing these tasks in real time.
[0875] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0876] In this invention, the server includes a means for collecting voice and video data of workers, a means for transmitting the collected voice and video data to the server, a means for the generative AI model to propose optimal work procedures in response to questions from workers, a means for the generative AI model to propose and reserve work spots, and a means for recording who performed which work and reporting individual progress. This makes it possible to improve work efficiency, respond to questions, propose optimal spots, and manage individual progress in real time.
[0877] "Audio and video data of workers" refers to the audio made by workers in a logistics center, as well as video footage of their appearance and actions.
[0878] A "generative AI model" refers to an artificial intelligence system that uses a generative approach to generate diverse outputs based on input data.
[0879] "Distribution center" refers to a facility that receives, stores, ships, and distributes goods and materials.
[0880] "Work procedures" refer to specific instructions and methods for effectively and efficiently carrying out various tasks within a logistics center.
[0881] "Work Spot" refers to a location or area within a distribution center where specific work is performed.
[0882] "Progress report" refers to a report by a worker on the progress or achievement of the work he or she has done.
[0883] System Overview
[0884] The system of the present invention collects voice and video data of workers at a logistics center and transmits it to a server in real time, and then utilizes a generative AI model to assist workers with their questions, suggest work spots, and report on the progress of individual tasks. The system is primarily composed of a server, terminals, and users (workers).
[0885] System Details
[0886] 1. Setup
[0887] Terminal: A robot is installed in each work area within the logistics facility. The robot is equipped with a camera, microphone, and display, and has a dedicated application installed. The robot is capable of communicating with the server.
[0888] Server: The server registers data related to workers, work spots, and work content in a database. It is equipped with mechanisms for running voice recognition, image recognition, and generative AI models.
[0889] 2. Real-time analysis of work status
[0890] Terminal: The robot's camera and microphone collect the worker's voice and video in real time. The sensors periodically check whether the robot is working properly.
[0891] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[0892] 3. Data analysis and response generation
[0893] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the worker's questions, and extracts the necessary information.
[0894] Server: Provides the analysis results to the generation AI model, generating appropriate answers to worker questions and proposing work procedures.
[0895] Terminal: Displays responses and suggestions sent from the server to the worker and reads them out loud if necessary.
[0896] Specific processing examples
[0897] Work procedure suggestions
[0898] 1. User: A worker asks, "Please tell me where to place the next package most efficiently."
[0899] 2. Terminal: Collects the question voice and sends it to the server.
[0900] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[0901] 4. Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[0902] Work Spot Reservation
[0903] 1. User: A worker requests that the next work spot be reserved.
[0904] 2. Terminal: Sends the request to the server.
[0905] 3. Server: Suggests the optimal location based on the work content and information on available work spots within the logistics center.
[0906] 4. Terminal: Shows the worker a list of suggested spots and confirms the reservation.
[0907] 5. User: The worker selects a spot and confirms the reservation.
[0908] 6. Server: Makes a reservation for the selected spot and sends a confirmation message to the terminal.
[0909] 7. Terminal: Display a confirmation message to the worker.
[0910] Progress reports for individual tasks
[0911] 1. Terminal: Enter additional work while working and send the content to the server.
[0912] 2. Server: Records who has done what work and manages progress.
[0913] 3. Terminal: Provides the worker with a progress report screen when the work is completed.
[0914] 4. Server: Organizes the progress of work and reports on individual tasks.
[0915] 5. Terminal: Displays the progress of each worker and provides a confirmation screen.
[0916] Hardware and software used
[0917] Camera and microphone: Installed on the robot to collect the worker's voice and video.
[0918] Robot: Moves around the distribution center and is positioned in each designated work area.
[0919] Server: Runs speech and image recognition and generative AI models (e.g., GPT-3.5-turbo).
[0920] Prompt Sentence Examples
[0921] "Please tell me where to place the next shipment for maximum efficiency."
[0922] "Please reserve the next work spot for me."
[0923] In this way, this system aims to improve work efficiency at logistics centers and provides real-time assistance to workers.
[0924] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0925] Step 1:
[0926] User: A worker asks, "Please tell me where to place the next package most efficiently."
[0927] Input: The worker's spoken question.
[0928] Output: Audio data is input to the robot's microphone.
[0929] Step 2:
[0930] Terminal: Collects voice questions and sends them to the server.
[0931] Input: Worker voice data.
[0932] Data processing: Audio data captured by the robot's microphone is converted into a digital signal and sent to a server via a secure communication protocol.
[0933] Output: Audio data transferred to the server.
[0934] Step 3:
[0935] Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[0936] Input: The transmitted audio data.
[0937] Data computation: Speech data is converted into text using a speech recognition algorithm, and then fed into a generative AI model (e.g., GPT-3.5-turbo) for analysis.
[0938] Output: Text data containing optimal work procedure information.
[0939] Step 4:
[0940] Server: Based on the analysis results, the AI model generates work procedures and sends them to the device.
[0941] Input: Text data containing optimal work procedure information.
[0942] Data calculation: The generative AI model generates appropriate work procedures based on the analysis results.
[0943] Output: Work procedure information sent to the robot's terminal.
[0944] Step 5:
[0945] Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[0946] Input: Work procedure information sent from the server.
[0947] Data processing: Work procedure information is displayed on the robot's display and read aloud using a speech synthesis algorithm.
[0948] Output: Visual and audio feedback to the worker.
[0949] Step 6:
[0950] User: A worker requests to reserve the next work spot.
[0951] Input: A voice request to reserve a work spot.
[0952] Output: Audio data is input to the robot's microphone.
[0953] Step 7:
[0954] Terminal: Sends the request to the server.
[0955] Input: Worker voice data.
[0956] Data processing: Converts voice data into digital signals and sends them to a server via a secure communication protocol.
[0957] Output: Audio data transferred to the server.
[0958] Step 8:
[0959] Server: Suggests the optimal location based on the work content and information on available work spots within the distribution center.
[0960] Input: Voice request data from workers and work spot data within the distribution center.
[0961] Data Computation: Uses speech recognition algorithms to convert voice data into text, search for available work spots, and use generative AI models to suggest the best locations.
[0962] Output: Text data containing the suggested work spot information.
[0963] Step 9:
[0964] Terminal: Displays the proposed spot list received from the server to the worker and confirms the reservation.
[0965] Input: Suggested spot information sent from the server.
[0966] Data processing: Display a list of suggested spots on the robot's display and provide an interface for reservation confirmation.
[0967] Output: A confirmation screen that is displayed to the worker.
[0968] Step 10:
[0969] User: The worker selects a spot and confirms the reservation.
[0970] Input: Work spot information selected by the worker.
[0971] Output: The input selection information is obtained through the robot interface.
[0972] Step 11:
[0973] Terminal: Sends the selected information to the server.
[0974] Input: Worker selection information.
[0975] Data processing: Selected data is converted into a digital signal and sent to the server via a secure communication protocol.
[0976] Output: The selected data that is transferred to the server.
[0977] Step 12:
[0978] Server: Makes a reservation at the selected spot and sends a confirmation message to the device.
[0979] Input: Worker selection information.
[0980] Data calculation: Executes the reservation procedure and generates a confirmation message.
[0981] Output: Sends a confirmation message to the robot's terminal.
[0982] Step 13:
[0983] Terminal: Display a confirmation message to the worker.
[0984] Input: The confirmation message sent by the server.
[0985] Data processing: Display the message on the robot's display.
[0986] Output: A confirmation message that is displayed to the worker.
[0987] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0988] System Overview
[0989] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI and an emotion engine to assist customer conversations, answer follow-up questions, and recommend and make reservations for second restaurants. It also records who ordered what and has the ability to allocate the bill individually, recognizing user emotions in real time to improve the quality of service. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[0990] System Details
[0991] 1. Setup
[0992] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[0993] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with a system that runs voice recognition, image recognition, generative AI models, and an emotion engine.
[0994] 2. Real-time analysis of customer behavior and eating habits
[0995] Device: The tablet's camera and microphone are used to collect the customer's voice and video in real time.
[0996] Terminal: Collected data is sent to the server. The data transfer protocol is encrypted and secure.
[0997] 3. Data analysis and response generation
[0998] Server: Analyzes the received audio and video data using voice and image recognition algorithms, analyzes the customer's conversation and the progress of their meal, and extracts the necessary information.
[0999] Server: Feeds the analysis results to the generative AI model and emotion engine to generate appropriate answers to customer questions and suggestions for additional orders.
[1000] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[1001] 4. Use of Emotion Engine
[1002] Server: The emotion engine analyzes the customer's facial expressions and tone of voice based on collected audio and video data to grasp their emotional state (e.g., joy, displeasure, interest, anger, etc.) in real time.
[1003] Server: Based on the analysis results of the emotion engine, the generative AI model adjusts the responses and assistance provided accordingly. For example, if the customer shows signs of discomfort, the server softens the response and refrains from making meal suggestions.
[1004] Specific processing examples
[1005] Conversation assistance
[1006] 1. User: A customer asks, "What's your recommendation today?"
[1007] 2. Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1008] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[1009] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the generated response appropriately (e.g., adding more information if the customer is interested).
[1010] 5. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1011] Reservation for the second restaurant
[1012] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[1013] 2. Terminal: Sends the request to the server.
[1014] 3. Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. It lists available establishments and retrieves detailed information from the database.
[1015] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations (e.g., if the customer indicates they want to relax, prioritize quiet bars).
[1016] 5. Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1017] 6. User: The customer selects the desired store and confirms the reservation.
[1018] 7. Terminal: Sends the selection information to the server.
[1019] 8. Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1020] 9. Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1021] Accounting responsibilities
[1022] 1. Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1023] 2. Server: Automatically records who ordered what. Using image recognition technology, it automatically tracks who placed what order.
[1024] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1025] 4. User: The customer checks the individual order details and enters them into the tablet device.
[1026] 5. Terminal: Sends the input to the server.
[1027] 6. Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user contributes based on the items ordered.
[1028] 7. Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each person's payment.
[1029] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1030] The processing flow will be explained below.
[1031] Collection and analysis of customer audio and video data
[1032] Step 1:
[1033] Device: The tablet activates the camera and microphone to collect the customer's voice and video in real time, capturing the content of the conversation and the customer's behavior.
[1034] Step 2:
[1035] Terminal: Collected audio and video data is temporarily stored in internal memory and compressed, allowing for efficient data transfer.
[1036] Step 3:
[1037] Terminal: Compressed audio and video data is sent to the server. The data is encrypted to prevent information leakage during transmission.
[1038] Step 4:
[1039] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[1040] Conversation assistance and question handling
[1041] Step 1:
[1042] User: A customer asks, "What's your recommendation today?"
[1043] Step 2:
[1044] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1045] Step 3:
[1046] Server: Analyzes the received text data, understands the intent of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[1047] Step 4:
[1048] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[1049] Step 5:
[1050] Server: Uses analyzed customer sentiment data to adjust the tone and content of the responses it generates, for example, providing more detailed information if the customer is interested.
[1051] Step 6:
[1052] Server: Sends the generated answer to the device.
[1053] Step 7:
[1054] Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1055] Introduction and reservation of the second restaurant
[1056] Step 1:
[1057] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[1058] Step 2:
[1059] Terminal: Sends the request contents to the server.
[1060] Step 3:
[1061] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[1062] Step 4:
[1063] Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations accordingly. For example, if a customer indicates that they want to relax, it will prioritize quiet bars.
[1064] Step 5:
[1065] Server: Sends the listed store information to the terminal.
[1066] Step 6:
[1067] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1068] Step 7:
[1069] User: The customer selects the desired store and confirms the reservation.
[1070] Step 8:
[1071] Terminal: Sends the selected information to the server.
[1072] Step 9:
[1073] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1074] Step 10:
[1075] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1076] Accounting responsibilities
[1077] Step 1:
[1078] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1079] Step 2:
[1080] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[1081] Step 3:
[1082] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1083] Step 4:
[1084] User: The customer checks the details of each order and enters them into the tablet device.
[1085] Step 5:
[1086] Terminal: Sends input to the server.
[1087] Step 6:
[1088] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[1089] Step 7:
[1090] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[1091] Step 8:
[1092] Server: The emotion engine analyzes the customer's emotional state at checkout and adjusts payment methods and responses as needed.
[1093] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1094] Example 2
[1095] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1096] Modern restaurants are required to improve customer convenience while simultaneously increasing operational efficiency. However, the number of employees is limited, and it is often difficult to provide adequate service, especially during busy periods. It is also difficult to grasp customers' emotional state in real time and provide service accordingly. Furthermore, it is time-consuming to accurately record who ordered what and assign individual tasks at the checkout. A system is needed to solve these issues and improve customer satisfaction.
[1097] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for encrypting the customer's voice and video data and transmitting it to the server, means for analyzing the data using voice recognition and image recognition algorithms and providing the analysis results to the generative AI model and emotion analysis engine, and means for the generative AI model to assist the customer in the conversation based on the analysis results. This enables secure transfer and analysis of the customer's voice and video data, immediate conversation support based on the analysis results, and adjustment of response content according to emotions. It also includes functions for introducing and making reservations for the next store and sharing individual billing responsibilities, greatly improving customer convenience.
[1098] "Audio and video data" refers to digital data that records the voice and video of a customer.
[1099] "Encryption" is a technology that converts data into a form that cannot be understood by third parties in order to transfer it securely.
[1100] A "server" is a computer system that handles back-end operations such as data management, analysis, and communication.
[1101] A "voice recognition algorithm" is a program for converting voice data into text data.
[1102] An "image recognition algorithm" is a program that analyzes video data to detect specific patterns or objects.
[1103] A "generative AI model" is an artificial intelligence model that generates natural language responses or suggestions based on input data.
[1104] An "emotion analysis engine" is a program for analyzing an individual's emotional state from audio and video data.
[1105] "Conversational assistance" means providing appropriate responses and suggestions to customer questions and requests.
[1106] The "means for answering additional questions" is a function for generating answers to additional questions from customers.
[1107] "Next store introduction and reservation" is a function that suggests the next store that the customer would like to visit and makes a reservation for it.
[1108] The "means for dividing the bill amount individually" is a function that records the order details for each customer and calculates the individual payment amount based on that.
[1109] The system of the present invention is designed to improve customer convenience in restaurants and is composed of multiple elements including terminals (tablets), servers, and users (customers). Each element will be described in detail below.
[1110] 1. Setup
[1111] Device:
[1112] The tablet is installed at the customer's seat and is equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, and it can communicate with the server. The tablet uses a Wi-Fi connection, and plugins and drivers are pre-configured. Normal operation is initiated by clicking the "Start" button on the main screen.
[1113] server:
[1114] The server registers customer information, menus, and data on affiliated restaurants in a database. The database uses MySQL, and an environment is set up in which voice recognition, image recognition, generative AI models, and an emotion analysis engine can operate. The necessary software and libraries are installed on the server in advance.
[1115] 2. Real-time analysis of customer behavior and eating habits
[1116] Device:
[1117] The tablet's built-in camera and microphone are used to collect customer audio and video data in real time. For example, when a customer speaks into the tablet, their audio and video are recorded.
[1118] Device:
[1119] The collected data is encrypted using the AES encryption algorithm and transmitted to the server via the HTTPS protocol, which ensures the security of the data.
[1120] 3. Data analysis and response generation
[1121] server:
[1122] The server uses a speech recognition algorithm (e.g., Reve) to convert the voice data into text. For example, the voice data "What's the recommendation today?" is converted into text data "What's the recommendation today?"
[1123] server:
[1124] Image recognition algorithms (e.g., OpenCV) are used to analyze video data and grasp customer facial expressions and movements in real time.
[1125] server:
[1126] A generative AI model (e.g., GPT-3) generates an appropriate response based on the analysis results. For example, in response to the text data "What's today's recommendation?", it generates a response such as "Today's recommendation is seafood pasta."
[1127] Device:
[1128] The response data from the server is sent to the tablet, displayed on the screen, and read aloud using a Text-to-Speech (TTS) engine.
[1129] 4. Use of sentiment analysis engines
[1130] server:
[1131] The emotion analysis engine analyzes the customer's emotional state from audio and video data. For example, if the customer is smiling, it is determined to be "happy."
[1132] server:
[1133] The generative AI model adjusts responses based on the results of sentiment analysis. For example, if a customer expresses displeasure, the next response will be generated in a softer tone. For example, a response such as, "May I suggest some other dishes so you can enjoy them?"
[1134] Examples of concrete examples and prompts
[1135] Conversation Assistance:
[1136] When a user asks, "What's the recommendation today?", the device collects the voice and sends it to the server, which analyzes the voice and generates a response such as, "Today's recommendation is seafood pasta."
[1137] 2nd restaurant reservation:
[1138] When a customer requests a reservation for the next bar they want to visit, the server will suggest a bar based on the customer's desired conditions and make the reservation.
[1139] Example prompt sentence:
[1140] "When a customer asks, 'What's the special today?' show them the current menu specials and read out the details."
[1141] With the above configuration, the present invention significantly improves customer convenience and improves the operational efficiency of restaurants. In particular, by combining it with an emotion analysis engine, it becomes possible to respond flexibly to customer emotions, thereby providing more personalized service.
[1142] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1143] Step 1:
[1144] Device:
[1145] A tablet is placed at the customer's seat, and the camera, microphone, and touchscreen are set up. A dedicated application is then launched. When the customer speaks into the tablet, the camera and microphone collect audio and video. The collected data is processed in real time, resulting in minimal latency.
[1146] Input: Customer audio and video data
[1147] Output: Encrypted audio and video data
[1148] For example, a customer asks, "What's the recommendation today?" Audio and video are captured and the application encrypts the data.
[1149] Step 2:
[1150] Device:
[1151] The collected data is encrypted using the AES encryption algorithm and sent to the server via the HTTPS protocol, ensuring data security and privacy.
[1152] Input: Unencrypted audio and video data
[1153] Output: Encrypted data sent to the server
[1154] As a specific example of operation, the encrypted data is sent to the server, and a message indicating that the transfer is complete is displayed.
[1155] Step 3:
[1156] server:
[1157] The voice data that arrives at the server is analyzed using a voice recognition algorithm and converted into text data. For example, the voice data "What's recommended today?" is converted into text.
[1158] Input: Encrypted audio data
[1159] Output: Text data
[1160] As a specific example of how it works, the voice recognition algorithm is activated and the result "What's recommended today?" is output.
[1161] Step 4:
[1162] server:
[1163] Image recognition algorithms are used to analyze the customer's facial expressions and movements from the incoming video data, thereby capturing the customer's real-time emotional state.
[1164] Input: Encrypted video data
[1165] Output: Emotional state data
[1166] As a specific example of operation, an image recognition algorithm is run to determine that "the customer is showing interest."
[1167] Step 5:
[1168] server:
[1169] A generative AI model is used to generate an appropriate response based on the analyzed text data and emotional state, for example, "Today's recommendation is seafood pasta."
[1170] Input: Text data and emotional state data
[1171] Output: Response message
[1172] As a specific example of how it works, the generative AI model generates the message "Today's recommendation is seafood pasta."
[1173] Step 6:
[1174] server:
[1175] The generated response message is sent to the terminal.
[1176] Input: Response message
[1177] Output: Message sent to the terminal
[1178] As a specific example of operation, a response message is sent from the server to the terminal.
[1179] Step 7:
[1180] Device:
[1181] The terminal receives the response message, displays it on the screen, and uses a Text-to-Speech (TTS) engine to read it aloud to the customer.
[1182] Input: Received response message
[1183] Output: Voice and text output to the customer
[1184] As a specific example of how it works, the message "Today's recommendation is seafood pasta" is displayed on the screen and simultaneously read aloud.
[1185] By having each step work in conjunction with each other, we aim to create a system that can respond to customers quickly and appropriately.
[1186] (Application example 2)
[1187] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1188] In recent years, restaurants have been required to improve the quality of customer service and streamline operations. However, staff shortages and inconsistent service quality have become problems. In particular, there are limitations to providing multilingual support and individualized service in real time, making it difficult to maintain customer satisfaction. It is also difficult to allocate individual amounts at the time of payment and provide service that responds to customer emotions.
[1189] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to the server, means for the server to analyze the audio and video data and provide the analysis results to the generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to recommend and reserve the next restaurant, means for recording who ordered what and individually apportioning the bill, means for analyzing the customer's emotional state and adjusting the quality of service, means for displaying real-time video of the customer on a smartphone or other mobile device, and means for encrypting the collected data and securely transferring it to the server. This improves the quality of customer service and enables efficient business operations. It is also possible to provide multilingual support, real-time personalized services, and flexible service provision based on individual apportionment of the bill and emotions.
[1190] "Customer audio and video data" refers to audio produced by the customer and video information such as the customer's actions and facial expressions.
[1191] "Collecting means" refers to devices and methods for capturing audio and video data.
[1192] "Transmission means" refers to the communications technologies and protocols used to transmit collected data to the server.
[1193] "Server" refers to the computer system that analyzes collected data and runs the generative AI model and emotion engine.
[1194] "Means for analysis" refers to the algorithms and software used to analyze audio and video data and extract the necessary information.
[1195] "Generative AI model" refers to an artificial intelligence algorithm or program that generates responses or suggestions based on data it receives.
[1196] "Means for assisting conversation" refers to the method by which the generative AI model assists customer conversations based on the analysis results.
[1197] "Means for answering additional questions" refers to a method for generating and providing appropriate answers to new questions from customers.
[1198] "Means for introducing and making reservations at the next store" refers to a method for suggesting another store to the customer and making a reservation at that store if necessary.
[1199] "Means of recording who ordered what and allocating the bill individually" refers to a method of recording the order details by linking them to a specific customer and ultimately calculating the individual payment amount.
[1200] "Means for analyzing emotional state and adjusting quality of service" refers to a method for analyzing the emotions of a customer and providing service accordingly.
[1201] "Means for displaying real-time video on a smartphone or other mobile device" refers to methods and technologies for capturing video of a customer in real time and displaying it on a smartphone or other mobile device.
[1202] "Means for encrypting and securely transferring data to the server" refers to the technology and methods for encrypting collected data and transmitting it securely to the server while preventing unauthorized access from outside.
[1203] System Overview
[1204] The system of the present invention collects customer voice and video data and analyzes it on a server to improve the quality of service. The main components of the system are a server, a terminal, and a user, and is realized using various devices and software.
[1205] Hardware and Software Configuration
[1206] The main components running on the server and terminals are:
[1207] Hardware
[1208] Device: A mobile device such as a tablet or smartphone equipped with a camera and microphone to collect customer audio and video data.
[1209] Server: A powerful computer system responsible for analyzing data and running generative AI models.
[1210] Software and APIs
[1211] Speech Recognition API: Google Cloud Speech-to-Text API
[1212] Image Recognition API: Amazon Rekognition
[1213] Generative AI model: GPT-4 (OpenAI)
[1214] Emotion engine: IBM Watson Tone Analyzer
[1215] Database: Firebase Realtime Database
[1216] Communication protocol: HTTPS (SSL / TLS encryption)
[1217] Data processing and analysis flow
[1218] 1. Data Collection:
[1219] The device's camera and microphone are used to collect customer audio and video data.
[1220] The collected data is temporarily stored on the device.
[1221] 2. Data Transfer:
[1222] The collected data is sent to the server in real time using the HTTPS protocol.
[1223] Data is transferred securely via SSL / TLS encryption.
[1224] 3. Data Analysis:
[1225] The server converts the received voice data into text using the Google Cloud Speech-to-Text API.
[1226] The video data is analyzed using Amazon Rekognition to recognize customer facial expressions and behavior.
[1227] The analysis results are fed into a generative AI model (GPT-4) to generate responses and suggestions to customer questions and requests.
[1228] The emotion engine (IBM Watson Tone Analyzer) analyzes the customer's emotional state and adjusts the response content and service delivery method.
[1229] Provision of services
[1230] 1. Conversation assistance:
[1231] The server uses a generative AI model to generate appropriate answers to customer questions, such as suggesting menu recommendations in response to the question, "What's your recommendation today?"
[1232] The generated answers are displayed on the device screen and, if necessary, read aloud.
[1233] 2. Introduction and reservation of the following stores:
[1234] Based on the customer's request, the server searches for and suggests the next restaurant. The server processes the reservation and sends a confirmation message to the terminal.
[1235] 3. Accounting responsibilities:
[1236] The terminal records who has ordered what, and the server calculates the individual payment amount, which is displayed on the terminal for each customer to see.
[1237] Specific examples
[1238] For example, if a customer asks, "What dish do you recommend?" the system will process it as follows:
[1239] Prompt Sentence Examples
[1240] A customer asks, "What dish do you recommend?" Use the following information to generate a suitable answer:
[1241] Today's recommended dishes are "Omurice," "Steak," and "Salad."
[1242] The customer's emotional state is "interested."
[1243] answer:
[1244] Using this prompt, the generative AI model generates a response such as, "Today's recommended dishes are omelet rice, steak, and salad. They're all delicious. Let me know if you need more information," and displays it on the device. It also decides whether to provide more detailed information depending on the customer's emotional state.
[1245] The above is a specific embodiment for carrying out the invention. This system improves the quality of customer service and enables efficient business operations.
[1246] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1247] Step 1:
[1248] The user turns on the device (smartphone or tablet) and launches the dedicated application. The user then enables the camera and microphone. The camera and microphone collect the customer's audio and video data in real time. The customer's audio and video are captured as input and converted into digital data within the application.
[1249] Step 2:
[1250] The collected audio and video data is sent from the device to the server using the SSL / TLS encrypted HTTPS protocol. This data is securely received by the server. The collected audio and video data is sent as input, reaches the server, is safely received, and is output.
[1251] Step 3:
[1252] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. Voice data is given as input, and it is analyzed and converted into text data. The output is text data.
[1253] Step 4:
[1254] The server analyzes the received video data using the Amazon Rekognition API to recognize the customer's facial expressions and behavior. The video data is given as input, and it is analyzed to output information about facial expressions and behavior.
[1255] Step 5:
[1256] The analyzed audio and video data is integrated on the server and supplied to a generative AI model (GPT-4). The generative AI model creates prompts to generate appropriate answers to customer questions and requests. The textual customer utterances and emotional state are given as input, and an appropriate answer text is generated as output.
[1257] Step 6:
[1258] The server uses IBM Watson Tone Analyzer to analyze the customer's emotional state based on their text data. The generative AI model adjusts the answers and suggestions it provides based on their emotional state. The customer's text data is given as input, and the emotional state is obtained as output.
[1259] Step 7:
[1260] The generated response or suggestion is then re-encrypted and sent to the device, where it is received and displayed on the application screen and optionally read aloud. The generated response text is given as input, and the output is the answer appropriate to the customer's request.
[1261] Step 8:
[1262] Based on a specific need, if a customer requests "I would like to reserve the next store," the server receives the request, searches for an appropriate store based on the location information and desired conditions, and performs the reservation procedure. The customer's request is given as input, and a reservation confirmation message is generated as output.
[1263] Step 9:
[1264] When a customer places an order, the terminal sends the details to the server, which automatically records who ordered what. After the meal is finished, the server displays the bill based on the recorded order data and allocates it to each customer. The order data is given as input, and the individual payment amounts are calculated as output.
[1265] Step 10:
[1266] Based on the customer's emotional state, the server adjusts the response content and quality of service to provide optimal service. The emotion analysis results are given as input, and the adjusted service content is provided as output.
[1267] In this way, the system can improve customer service in real time and also improve operational efficiency.
[1268] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1269] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1270] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1271] [Third embodiment]
[1272] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1273] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1274] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1275] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1276] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1277] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1278] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1279] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1280] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1281] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1282] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1283] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1284] System Overview
[1285] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI to assist the customer's conversation, answering any follow-up questions, recommending a second restaurant, and making reservations. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily comprised of a server, a terminal (tablet), and a user (customer).
[1286] System Details
[1287] 1. Setup
[1288] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[1289] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with mechanisms for voice recognition, image recognition, and generative AI models.
[1290] 2. Real-time analysis of customer behavior and eating habits
[1291] Device: The tablet's camera and microphone are used to collect audio and video of the customer in real time. The sensors are periodically checked to ensure they are working properly.
[1292] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[1293] 3. Data analysis and response generation
[1294] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the customer's conversation, and extracts the necessary information.
[1295] Server: Feeds the analysis results to the generative AI model to generate appropriate answers to customer questions and suggestions for additional orders.
[1296] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[1297] Specific processing examples
[1298] Conversation assistance
[1299] 1. User: A customer asks, "What's your recommendation today?"
[1300] 2. Terminal: Collects the question voice and sends it to the server.
[1301] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[1302] 4. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1303] Reservation for the second restaurant
[1304] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[1305] 2. Terminal: Sends the request to the server.
[1306] 3. Server: Based on the customer's location and desired conditions, the server searches for information on nearby bars and lists available reservations.
[1307] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[1308] 5. User: The customer selects a store and confirms the reservation.
[1309] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[1310] 7. Terminal: Display a confirmation message to the customer.
[1311] Accounting responsibilities
[1312] 1. Terminal: During the meal, you enter additional orders into the terminal and send the details to the server.
[1313] 2. Server: Automatically records who ordered what using image recognition technology.
[1314] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1315] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[1316] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[1317] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1318] The processing flow will be explained below.
[1319] Collection and analysis of customer audio and video data
[1320] Step 1:
[1321] Device: The tablet activates the camera and microphone to capture the customer's voice and video in real time, capturing details of the customer's movements and conversations.
[1322] Step 2:
[1323] Terminal: The collected audio and video data is compressed and securely sent to the server. The data transfer protocol is encrypted to protect customer privacy.
[1324] Step 3:
[1325] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[1326] Conversation assistance and question handling
[1327] Step 1:
[1328] User: A customer asks, "What's your recommendation today?"
[1329] Step 2:
[1330] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1331] Step 3:
[1332] Server: Analyzes the received text data, understands the content of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[1333] Step 4:
[1334] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[1335] Step 5:
[1336] Server: Sends the generated answer to the device.
[1337] Step 6:
[1338] Terminal: Receives the answer from the server and displays it to the customer. At the same time, it reads the answer aloud.
[1339] Introduction and reservation of the second restaurant
[1340] Step 1:
[1341] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[1342] Step 2:
[1343] Terminal: Sends the request contents to the server.
[1344] Step 3:
[1345] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[1346] Step 4:
[1347] Server: Sends the listed store information to the terminal.
[1348] Step 5:
[1349] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1350] Step 6:
[1351] User: The customer selects the desired store and confirms the reservation.
[1352] Step 7:
[1353] Terminal: Sends the selected information to the server.
[1354] Step 8:
[1355] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1356] Step 9:
[1357] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1358] Accounting responsibilities
[1359] Step 1:
[1360] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1361] Step 2:
[1362] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[1363] Step 3:
[1364] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1365] Step 4:
[1366] User: The customer checks the details of each order and enters them into the tablet device.
[1367] Step 5:
[1368] Terminal: Sends input to the server.
[1369] Step 6:
[1370] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[1371] Step 7:
[1372] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[1373] The above are the specific processing steps in the system of the present invention, which can reduce customer waiting times and significantly improve store operational efficiency.
[1374] Example 1
[1375] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1376] Conventional customer service systems in restaurants lack the technology to effectively utilize customer voice and video data to assist conversations, answer additional questions, recommend and reserve the next restaurant, record who ordered what, and allocate individual bills. Therefore, there is a need to improve the quality of customer service and solve the issues of employee shortages and work efficiency.
[1377] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1378] In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to a data processing device, means for the data processing device to analyze the audio and video data and supply the analysis results to a generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to introduce and make reservations for the next store to be visited, and means for recording who ordered what and dividing up the bill individually. This enables a system that improves the quality of customer service and solves issues such as employee shortages and work efficiency.
[1379] "Means for collecting customer audio and video data" refers to equipment installed to collect customer audio and video information in real time, such as devices including cameras and microphones.
[1380] "Means for transmitting collected audio and video data to a data processing device" refers to network communication functions and protocols for transmitting audio and video data collected from customers to a data processing device (server) in a secure manner.
[1381] "Means for analyzing audio and video data using a data processing device and providing the analysis results to a generative AI model" refers to a process and system that analyzes data using speech recognition and image recognition algorithms running on a server and provides the analysis results to a generative AI model.
[1382] "Means for the generative AI model to assist customer conversations based on the analysis results" refers to a function in which the generative AI model provides information to assist customer questions and conversations based on analyzed data.
[1383] "Means for the generative AI model to answer follow-up questions from customers" refers to a system in which the generative AI model generates and provides appropriate answers to further questions from customers.
[1384] "Means for the generative AI model to introduce and make reservations for the next store to visit" is a function in which the generative AI model suggests the next store to visit based on the customer's preferences and carries out the reservation process.
[1385] "A means of recording who ordered what and dividing the bill individually" is a system that records the menu items ordered by customers and calculates and divides the amount due for each customer at the time of payment.
[1386] System Overview
[1387] The system of the present invention uses terminals installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes a generative AI model to assist customer conversations, answer follow-up questions, and recommend and make reservations for the next restaurant to visit. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[1388] Detailed configuration and operation
[1389] set up
[1390] Terminal: A tablet is installed at each customer's seat and equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, enabling communication with the server.
[1391] Server: The server stores customer information, menu information, and information about participating restaurants in a database. It also sets up voice recognition software (e.g., Google Cloud Speech-to-Text), image recognition software (e.g., AWS Rekognition), and generative AI models (e.g., OpenAI GPT-4), and sets the necessary API keys.
[1392] Collection of customer audio and video data
[1393] Device: The device uses the tablet's camera and microphone to collect real-time audio and video data from customers. It periodically checks the operation of sensors and issues alerts if there are any problems.
[1394] Terminal: The collected data is encrypted and sent to the data processing device (server). Security protocols (e.g. SSL / TLS) are used to ensure the safety of the transfer.
[1395] Data Analysis and Response Generation
[1396] Server: The server uses a speech recognition algorithm to convert the received audio data into text, and an image recognition algorithm to analyze the video data.
[1397] Server: The analyzed data is fed into a generative AI model, which generates appropriate answers and suggestions based on the customer's questions and conversations.
[1398] Terminal: Displays generated responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[1399] Specific processing examples
[1400] Conversation assistance
[1401] 1. User: A customer asks, "What's your recommendation today?"
[1402] 2. Device: The device collects the audio of this question and sends it to the server.
[1403] 3. Server: The server analyzes the voice data and generates menu recommendation information using a generative AI model.
[1404] 4. Terminal: The terminal displays the recommended menu information received from the server to the customer and reads it out loud.
[1405] Prompt Sentence Examples
[1406] "What's today's recommended menu?"
[1407] "What's the special tonight?"
[1408] Introduction and reservation of next store to visit
[1409] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[1410] 2. Terminal: The terminal sends the request to the server.
[1411] 3. Server: The server searches for nearby bars based on the customer's location and desired conditions, and lists available bars.
[1412] 4. Terminal: The terminal displays the list of suggested stores to the customer and confirms the reservation.
[1413] 5. User: The customer selects a store and confirms the reservation.
[1414] 6. Server: The server completes the reservation procedure with the selected store and sends a confirmation message to the terminal.
[1415] 7. Terminal: The terminal displays a confirmation message to the customer.
[1416] Prompt Sentence Examples
[1417] "Find recommended bars near you."
[1418] "Can you help me book a bar?"
[1419] Accounting responsibilities
[1420] 1. Terminal: When an additional order is entered during the meal, the terminal sends the details to the server.
[1421] 2. Server: The server uses image recognition technology to automatically record who ordered what.
[1422] 3. Terminal: Once the meal is finished, the terminal provides a screen where customers can enter who ordered what.
[1423] 4. Server: The server calculates the bill and calculates the individual amounts to be shared.
[1424] 5. Terminal: The terminal displays the payment amount for each customer and provides a confirmation screen.
[1425] Example prompts for implementing the invention
[1426] Conversation assistance prompts:
[1427] "What's today's recommended menu?"
[1428] "What's the special tonight?"
[1429] Prompt for next store introduction and reservation:
[1430] "Find recommended bars near you."
[1431] "Can you help me book a bar?"
[1432] Prompt for accounting allocation:
[1433] "Calculate the amount each customer pays."
[1434] "Please split the bill based on your order."
[1435] In this way, the system of the present invention improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component works in cooperation with the others, significantly improving the quality of service in restaurants.
[1436] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1437] Step 1:
[1438] set up
[1439] Devices: Tablets are installed at each customer's seat, and the camera, microphone, and touch screen are set up. The dedicated application is installed and the network connection is checked.
[1440] Input: Initial tablet information
[1441] Output: Tablet configured and network connection confirmed
[1442] Step 2:
[1443] Collection of customer audio and video data
[1444] Device: The tablet's camera and microphone are used to collect audio and video from customers in real time. The collected data is encrypted and sent to a server.
[1445] Input: Customer audio and video
[1446] Output: Encrypted audio and video data
[1447] Step 3:
[1448] Sending data to the server
[1449] Terminal: Collected audio and video data is sent to the server using a security protocol (e.g., SSL / TLS).
[1450] Input: Encrypted audio and video data
[1451] Output: Audio and video data received by the server
[1452] Step 4:
[1453] Audio and video data analysis
[1454] Server: Uses voice recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data, and uses image recognition software (e.g., AWS Rekognition) to analyze the video data.
[1455] Input: Audio and video data received by the server
[1456] Output: Parsed character data and analysis results
[1457] Step 5:
[1458] Feed to generative AI models and generate responses
[1459] Server: Feeds the analysis results to a generative AI model (e.g., OpenAI GPT-4) to generate appropriate answers and suggestions for customer questions.
[1460] Input: Parsed text data and video information
[1461] Output: Generated answers and suggestions
[1462] Step 6:
[1463] Sending the response to the terminal
[1464] Server: Sends generated answers and suggestions to the device.
[1465] Input: Generated answers and suggestions
[1466] Output: Answers and suggestions sent to the terminal
[1467] Step 7:
[1468] Displayed to customers and read aloud
[1469] Terminal: Displays responses and suggestions received from the server to the customer and reads them aloud if necessary.
[1470] Input: Response and proposal received from the server
[1471] Output: Information displayed to the customer and spoken aloud
[1472] Step 8:
[1473] Introducing the next store to visit and making reservations
[1474] 1. User: A customer requests a reservation for their next store visit.
[1475] 2. Terminal: Sends the request to the server.
[1476] 3. Server: Searches for nearby store information based on the customer's location and desired conditions, and lists stores that can be reserved.
[1477] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[1478] 5. User: The customer selects a store and confirms the reservation.
[1479] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[1480] 7. Terminal: Display a confirmation message to the customer.
[1481] Input: Customer request, location, preferences
[1482] Output: Reservation confirmation message
[1483] Step 9:
[1484] Accounting responsibilities
[1485] 1. Terminal: Additional orders are entered during the meal and sent to the server.
[1486] 2. Server: Automatically records who ordered what using image recognition technology.
[1487] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1488] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[1489] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[1490] Input: Customer additional order information, image data
[1491] Output: Recorded order information, calculated billing amount
[1492] (Application example 1)
[1493] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1494] There is a need to solve issues such as improving work efficiency at logistics centers, responding immediately to worker questions, proposing optimal work spots, and managing individual work progress.It is also important to achieve smooth and efficient work by performing these tasks in real time.
[1495] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1496] In this invention, the server includes a means for collecting voice and video data of workers, a means for transmitting the collected voice and video data to the server, a means for the generative AI model to propose optimal work procedures in response to questions from workers, a means for the generative AI model to propose and reserve work spots, and a means for recording who performed which work and reporting individual progress. This makes it possible to improve work efficiency, respond to questions, propose optimal spots, and manage individual progress in real time.
[1497] "Audio and video data of workers" refers to the audio made by workers in a logistics center, as well as video footage of their appearance and actions.
[1498] A "generative AI model" refers to an artificial intelligence system that uses a generative approach to generate diverse outputs based on input data.
[1499] "Distribution center" refers to a facility that receives, stores, ships, and distributes goods and materials.
[1500] "Work procedures" refer to specific instructions and methods for effectively and efficiently carrying out various tasks within a logistics center.
[1501] "Work Spot" refers to a location or area within a distribution center where specific work is performed.
[1502] "Progress report" refers to a report by a worker on the progress or achievement of the work he or she has done.
[1503] System Overview
[1504] The system of the present invention collects voice and video data of workers at a logistics center and transmits it to a server in real time, and then utilizes a generative AI model to assist workers with their questions, suggest work spots, and report on the progress of individual tasks. The system is primarily composed of a server, terminals, and users (workers).
[1505] System Details
[1506] 1. Setup
[1507] Terminal: A robot is installed in each work area within the logistics facility. The robot is equipped with a camera, microphone, and display, and has a dedicated application installed. The robot is capable of communicating with the server.
[1508] Server: The server registers data related to workers, work spots, and work content in a database. It is equipped with mechanisms for running voice recognition, image recognition, and generative AI models.
[1509] 2. Real-time analysis of work status
[1510] Terminal: The robot's camera and microphone collect the worker's voice and video in real time. The sensors periodically check whether the robot is working properly.
[1511] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[1512] 3. Data analysis and response generation
[1513] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the worker's questions, and extracts the necessary information.
[1514] Server: Provides the analysis results to the generation AI model, generating appropriate answers to worker questions and proposing work procedures.
[1515] Terminal: Displays responses and suggestions sent from the server to the worker and reads them out loud if necessary.
[1516] Specific processing examples
[1517] Work procedure suggestions
[1518] 1. User: A worker asks, "Please tell me where to place the next package most efficiently."
[1519] 2. Terminal: Collects the question voice and sends it to the server.
[1520] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[1521] 4. Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[1522] Work Spot Reservation
[1523] 1. User: A worker requests that the next work spot be reserved.
[1524] 2. Terminal: Sends the request to the server.
[1525] 3. Server: Suggests the optimal location based on the work content and information on available work spots within the logistics center.
[1526] 4. Terminal: Shows the worker a list of suggested spots and confirms the reservation.
[1527] 5. User: The worker selects a spot and confirms the reservation.
[1528] 6. Server: Makes a reservation for the selected spot and sends a confirmation message to the terminal.
[1529] 7. Terminal: Display a confirmation message to the worker.
[1530] Progress reports for individual tasks
[1531] 1. Terminal: Enter additional work while working and send the content to the server.
[1532] 2. Server: Records who has done what work and manages progress.
[1533] 3. Terminal: Provides the worker with a progress report screen when the work is completed.
[1534] 4. Server: Organizes the progress of work and reports on individual tasks.
[1535] 5. Terminal: Displays the progress of each worker and provides a confirmation screen.
[1536] Hardware and software used
[1537] Camera and microphone: Installed on the robot to collect the worker's voice and video.
[1538] Robot: Moves around the distribution center and is positioned in each designated work area.
[1539] Server: Runs speech and image recognition and generative AI models (e.g., GPT-3.5-turbo).
[1540] Prompt Sentence Examples
[1541] "Please tell me where to place the next shipment for maximum efficiency."
[1542] "Please reserve the next work spot for me."
[1543] In this way, this system aims to improve work efficiency at logistics centers and provides real-time assistance to workers.
[1544] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1545] Step 1:
[1546] User: A worker asks, "Please tell me where to place the next package most efficiently."
[1547] Input: The worker's spoken question.
[1548] Output: Audio data is input to the robot's microphone.
[1549] Step 2:
[1550] Terminal: Collects voice questions and sends them to the server.
[1551] Input: Worker voice data.
[1552] Data processing: Audio data captured by the robot's microphone is converted into a digital signal and sent to a server via a secure communication protocol.
[1553] Output: Audio data transferred to the server.
[1554] Step 3:
[1555] Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[1556] Input: The transmitted audio data.
[1557] Data computation: Speech data is converted into text using a speech recognition algorithm, and then fed into a generative AI model (e.g., GPT-3.5-turbo) for analysis.
[1558] Output: Text data containing optimal work procedure information.
[1559] Step 4:
[1560] Server: Based on the analysis results, the AI model generates work procedures and sends them to the device.
[1561] Input: Text data containing optimal work procedure information.
[1562] Data calculation: The generative AI model generates appropriate work procedures based on the analysis results.
[1563] Output: Work procedure information sent to the robot's terminal.
[1564] Step 5:
[1565] Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[1566] Input: Work procedure information sent from the server.
[1567] Data processing: Work procedure information is displayed on the robot's display and read aloud using a speech synthesis algorithm.
[1568] Output: Visual and audio feedback to the worker.
[1569] Step 6:
[1570] User: A worker requests to reserve the next work spot.
[1571] Input: A voice request to reserve a work spot.
[1572] Output: Audio data is input to the robot's microphone.
[1573] Step 7:
[1574] Terminal: Sends the request to the server.
[1575] Input: Worker voice data.
[1576] Data processing: Converts voice data into digital signals and sends them to a server via a secure communication protocol.
[1577] Output: Audio data transferred to the server.
[1578] Step 8:
[1579] Server: Suggests the optimal location based on the work content and information on available work spots within the distribution center.
[1580] Input: Voice request data from workers and work spot data within the distribution center.
[1581] Data Computation: Uses speech recognition algorithms to convert voice data into text, search for available work spots, and use generative AI models to suggest the best locations.
[1582] Output: Text data containing the suggested work spot information.
[1583] Step 9:
[1584] Terminal: Displays the proposed spot list received from the server to the worker and confirms the reservation.
[1585] Input: Suggested spot information sent from the server.
[1586] Data processing: Display a list of suggested spots on the robot's display and provide an interface for reservation confirmation.
[1587] Output: A confirmation screen that is displayed to the worker.
[1588] Step 10:
[1589] User: The worker selects a spot and confirms the reservation.
[1590] Input: Work spot information selected by the worker.
[1591] Output: The input selection information is obtained through the robot interface.
[1592] Step 11:
[1593] Terminal: Sends the selected information to the server.
[1594] Input: Worker selection information.
[1595] Data processing: Selected data is converted into a digital signal and sent to the server via a secure communication protocol.
[1596] Output: The selected data that is transferred to the server.
[1597] Step 12:
[1598] Server: Makes a reservation at the selected spot and sends a confirmation message to the device.
[1599] Input: Worker selection information.
[1600] Data calculation: Executes the reservation procedure and generates a confirmation message.
[1601] Output: Sends a confirmation message to the robot's terminal.
[1602] Step 13:
[1603] Terminal: Display a confirmation message to the worker.
[1604] Input: The confirmation message sent by the server.
[1605] Data processing: Display the message on the robot's display.
[1606] Output: A confirmation message that is displayed to the worker.
[1607] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1608] System Overview
[1609] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI and an emotion engine to assist customer conversations, answer follow-up questions, and recommend and make reservations for second restaurants. It also records who ordered what and has the ability to allocate the bill individually, recognizing user emotions in real time to improve the quality of service. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[1610] System Details
[1611] 1. Setup
[1612] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[1613] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with a system that runs voice recognition, image recognition, generative AI models, and an emotion engine.
[1614] 2. Real-time analysis of customer behavior and eating habits
[1615] Device: The tablet's camera and microphone are used to collect the customer's voice and video in real time.
[1616] Terminal: Collected data is sent to the server. The data transfer protocol is encrypted and secure.
[1617] 3. Data analysis and response generation
[1618] Server: Analyzes the received audio and video data using voice and image recognition algorithms, analyzes the customer's conversation and the progress of their meal, and extracts the necessary information.
[1619] Server: Feeds the analysis results to the generative AI model and emotion engine to generate appropriate answers to customer questions and suggestions for additional orders.
[1620] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[1621] 4. Use of Emotion Engine
[1622] Server: The emotion engine analyzes the customer's facial expressions and tone of voice based on collected audio and video data to grasp their emotional state (e.g., joy, displeasure, interest, anger, etc.) in real time.
[1623] Server: Based on the analysis results of the emotion engine, the generative AI model adjusts the responses and assistance provided accordingly. For example, if the customer shows signs of discomfort, the server softens the response and refrains from making meal suggestions.
[1624] Specific processing examples
[1625] Conversation assistance
[1626] 1. User: A customer asks, "What's your recommendation today?"
[1627] 2. Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1628] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[1629] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the generated response appropriately (e.g., adding more information if the customer is interested).
[1630] 5. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1631] Reservation for the second restaurant
[1632] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[1633] 2. Terminal: Sends the request to the server.
[1634] 3. Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. It lists available establishments and retrieves detailed information from the database.
[1635] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations (e.g., if the customer indicates they want to relax, prioritize quiet bars).
[1636] 5. Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1637] 6. User: The customer selects the desired store and confirms the reservation.
[1638] 7. Terminal: Sends the selection information to the server.
[1639] 8. Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1640] 9. Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1641] Accounting responsibilities
[1642] 1. Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1643] 2. Server: Automatically records who ordered what. Using image recognition technology, it automatically tracks who placed what order.
[1644] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1645] 4. User: The customer checks the individual order details and enters them into the tablet device.
[1646] 5. Terminal: Sends the input to the server.
[1647] 6. Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user contributes based on the items ordered.
[1648] 7. Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each person's payment.
[1649] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1650] The processing flow will be explained below.
[1651] Collection and analysis of customer audio and video data
[1652] Step 1:
[1653] Device: The tablet activates the camera and microphone to collect the customer's voice and video in real time, capturing the content of the conversation and the customer's behavior.
[1654] Step 2:
[1655] Terminal: Collected audio and video data is temporarily stored in internal memory and compressed, allowing for efficient data transfer.
[1656] Step 3:
[1657] Terminal: Compressed audio and video data is sent to the server. The data is encrypted to prevent information leakage during transmission.
[1658] Step 4:
[1659] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[1660] Conversation assistance and question handling
[1661] Step 1:
[1662] User: A customer asks, "What's your recommendation today?"
[1663] Step 2:
[1664] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1665] Step 3:
[1666] Server: Analyzes the received text data, understands the intent of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[1667] Step 4:
[1668] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[1669] Step 5:
[1670] Server: Uses analyzed customer sentiment data to adjust the tone and content of the responses it generates, for example, providing more detailed information if the customer is interested.
[1671] Step 6:
[1672] Server: Sends the generated answer to the device.
[1673] Step 7:
[1674] Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1675] Introduction and reservation of the second restaurant
[1676] Step 1:
[1677] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[1678] Step 2:
[1679] Terminal: Sends the request contents to the server.
[1680] Step 3:
[1681] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[1682] Step 4:
[1683] Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations accordingly. For example, if a customer indicates that they want to relax, it will prioritize quiet bars.
[1684] Step 5:
[1685] Server: Sends the listed store information to the terminal.
[1686] Step 6:
[1687] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1688] Step 7:
[1689] User: The customer selects the desired store and confirms the reservation.
[1690] Step 8:
[1691] Terminal: Sends the selected information to the server.
[1692] Step 9:
[1693] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1694] Step 10:
[1695] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1696] Accounting responsibilities
[1697] Step 1:
[1698] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1699] Step 2:
[1700] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[1701] Step 3:
[1702] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1703] Step 4:
[1704] User: The customer checks the details of each order and enters them into the tablet device.
[1705] Step 5:
[1706] Terminal: Sends input to the server.
[1707] Step 6:
[1708] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[1709] Step 7:
[1710] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[1711] Step 8:
[1712] Server: The emotion engine analyzes the customer's emotional state at checkout and adjusts payment methods and responses as needed.
[1713] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1714] Example 2
[1715] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1716] Modern restaurants are required to improve customer convenience while simultaneously increasing operational efficiency. However, the number of employees is limited, and it is often difficult to provide adequate service, especially during busy periods. It is also difficult to grasp customers' emotional state in real time and provide service accordingly. Furthermore, it is time-consuming to accurately record who ordered what and assign individual tasks at the checkout. A system is needed to solve these issues and improve customer satisfaction.
[1717] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for encrypting the customer's voice and video data and transmitting it to the server, means for analyzing the data using voice recognition and image recognition algorithms and providing the analysis results to the generative AI model and emotion analysis engine, and means for the generative AI model to assist the customer in the conversation based on the analysis results. This enables secure transfer and analysis of the customer's voice and video data, immediate conversation support based on the analysis results, and adjustment of response content according to emotions. It also includes functions for introducing and making reservations for the next store and sharing individual billing responsibilities, greatly improving customer convenience.
[1718] "Audio and video data" refers to digital data that records the voice and video of a customer.
[1719] "Encryption" is a technology that converts data into a form that cannot be understood by third parties in order to transfer it securely.
[1720] A "server" is a computer system that handles back-end operations such as data management, analysis, and communication.
[1721] A "voice recognition algorithm" is a program for converting voice data into text data.
[1722] An "image recognition algorithm" is a program that analyzes video data to detect specific patterns or objects.
[1723] A "generative AI model" is an artificial intelligence model that generates natural language responses or suggestions based on input data.
[1724] An "emotion analysis engine" is a program for analyzing an individual's emotional state from audio and video data.
[1725] "Conversational assistance" means providing appropriate responses and suggestions to customer questions and requests.
[1726] The "means for answering additional questions" is a function for generating answers to additional questions from customers.
[1727] "Next store introduction and reservation" is a function that suggests the next store that the customer would like to visit and makes a reservation for it.
[1728] The "means for dividing the bill amount individually" is a function that records the order details for each customer and calculates the individual payment amount based on that.
[1729] The system of the present invention is designed to improve customer convenience in restaurants and is composed of multiple elements including terminals (tablets), servers, and users (customers). Each element will be described in detail below.
[1730] 1. Setup
[1731] Device:
[1732] The tablet is installed at the customer's seat and is equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, and it can communicate with the server. The tablet uses a Wi-Fi connection, and plugins and drivers are pre-configured. Normal operation is initiated by clicking the "Start" button on the main screen.
[1733] server:
[1734] The server registers customer information, menus, and data on affiliated restaurants in a database. The database uses MySQL, and an environment is set up in which voice recognition, image recognition, generative AI models, and an emotion analysis engine can operate. The necessary software and libraries are installed on the server in advance.
[1735] 2. Real-time analysis of customer behavior and eating habits
[1736] Device:
[1737] The tablet's built-in camera and microphone are used to collect customer audio and video data in real time. For example, when a customer speaks into the tablet, their audio and video are recorded.
[1738] Device:
[1739] The collected data is encrypted using the AES encryption algorithm and transmitted to the server via the HTTPS protocol, which ensures the security of the data.
[1740] 3. Data analysis and response generation
[1741] server:
[1742] The server uses a speech recognition algorithm (e.g., Reve) to convert the voice data into text. For example, the voice data "What's the recommendation today?" is converted into text data "What's the recommendation today?"
[1743] server:
[1744] Image recognition algorithms (e.g., OpenCV) are used to analyze video data and grasp customer facial expressions and movements in real time.
[1745] server:
[1746] A generative AI model (e.g., GPT-3) generates an appropriate response based on the analysis results. For example, in response to the text data "What's today's recommendation?", it generates a response such as "Today's recommendation is seafood pasta."
[1747] Device:
[1748] The response data from the server is sent to the tablet, displayed on the screen, and read aloud using a Text-to-Speech (TTS) engine.
[1749] 4. Use of sentiment analysis engines
[1750] server:
[1751] The emotion analysis engine analyzes the customer's emotional state from audio and video data. For example, if the customer is smiling, it is determined to be "happy."
[1752] server:
[1753] The generative AI model adjusts responses based on the results of sentiment analysis. For example, if a customer expresses displeasure, the next response will be generated in a softer tone. For example, a response such as, "May I suggest some other dishes so you can enjoy them?"
[1754] Examples of concrete examples and prompts
[1755] Conversation Assistance:
[1756] When a user asks, "What's the recommendation today?", the device collects the voice and sends it to the server, which analyzes the voice and generates a response such as, "Today's recommendation is seafood pasta."
[1757] 2nd restaurant reservation:
[1758] When a customer requests a reservation for the next bar they want to visit, the server will suggest a bar based on the customer's desired conditions and make the reservation.
[1759] Example prompt sentence:
[1760] "When a customer asks, 'What's the special today?' show them the current menu specials and read out the details."
[1761] With the above configuration, the present invention significantly improves customer convenience and improves the operational efficiency of restaurants. In particular, by combining it with an emotion analysis engine, it becomes possible to respond flexibly to customer emotions, thereby providing more personalized service.
[1762] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1763] Step 1:
[1764] Device:
[1765] A tablet is placed at the customer's seat, and the camera, microphone, and touchscreen are set up. A dedicated application is then launched. When the customer speaks into the tablet, the camera and microphone collect audio and video. The collected data is processed in real time, resulting in minimal latency.
[1766] Input: Customer audio and video data
[1767] Output: Encrypted audio and video data
[1768] For example, a customer asks, "What's the recommendation today?" Audio and video are captured and the application encrypts the data.
[1769] Step 2:
[1770] Device:
[1771] The collected data is encrypted using the AES encryption algorithm and sent to the server via the HTTPS protocol, ensuring data security and privacy.
[1772] Input: Unencrypted audio and video data
[1773] Output: Encrypted data sent to the server
[1774] As a specific example of operation, the encrypted data is sent to the server, and a message indicating that the transfer is complete is displayed.
[1775] Step 3:
[1776] server:
[1777] The voice data that arrives at the server is analyzed using a voice recognition algorithm and converted into text data. For example, the voice data "What's recommended today?" is converted into text.
[1778] Input: Encrypted audio data
[1779] Output: Text data
[1780] As a specific example of how it works, the voice recognition algorithm is activated and the result "What's recommended today?" is output.
[1781] Step 4:
[1782] server:
[1783] Image recognition algorithms are used to analyze the customer's facial expressions and movements from the incoming video data, thereby capturing the customer's real-time emotional state.
[1784] Input: Encrypted video data
[1785] Output: Emotional state data
[1786] As a specific example of operation, an image recognition algorithm is run to determine that "the customer is showing interest."
[1787] Step 5:
[1788] server:
[1789] A generative AI model is used to generate an appropriate response based on the analyzed text data and emotional state, for example, "Today's recommendation is seafood pasta."
[1790] Input: Text data and emotional state data
[1791] Output: Response message
[1792] As a specific example of how it works, the generative AI model generates the message "Today's recommendation is seafood pasta."
[1793] Step 6:
[1794] server:
[1795] The generated response message is sent to the terminal.
[1796] Input: Response message
[1797] Output: Message sent to the terminal
[1798] As a specific example of operation, a response message is sent from the server to the terminal.
[1799] Step 7:
[1800] Device:
[1801] The terminal receives the response message, displays it on the screen, and uses a Text-to-Speech (TTS) engine to read it aloud to the customer.
[1802] Input: Received response message
[1803] Output: Voice and text output to the customer
[1804] As a specific example of how it works, the message "Today's recommendation is seafood pasta" is displayed on the screen and simultaneously read aloud.
[1805] By having each step work in conjunction with each other, we aim to create a system that can respond to customers quickly and appropriately.
[1806] (Application example 2)
[1807] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1808] In recent years, restaurants have been required to improve the quality of customer service and streamline operations. However, staff shortages and inconsistent service quality have become problems. In particular, there are limitations to providing multilingual support and individualized service in real time, making it difficult to maintain customer satisfaction. It is also difficult to allocate individual amounts at the time of payment and provide service that responds to customer emotions.
[1809] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to the server, means for the server to analyze the audio and video data and provide the analysis results to the generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to recommend and reserve the next restaurant, means for recording who ordered what and individually apportioning the bill, means for analyzing the customer's emotional state and adjusting the quality of service, means for displaying real-time video of the customer on a smartphone or other mobile device, and means for encrypting the collected data and securely transferring it to the server. This improves the quality of customer service and enables efficient business operations. It is also possible to provide multilingual support, real-time personalized services, and flexible service provision based on individual apportionment of the bill and emotions.
[1810] "Customer audio and video data" refers to audio produced by the customer and video information such as the customer's actions and facial expressions.
[1811] "Collecting means" refers to devices and methods for capturing audio and video data.
[1812] "Transmission means" refers to the communications technologies and protocols used to transmit collected data to the server.
[1813] "Server" refers to the computer system that analyzes collected data and runs the generative AI model and emotion engine.
[1814] "Means for analysis" refers to the algorithms and software used to analyze audio and video data and extract the necessary information.
[1815] "Generative AI model" refers to an artificial intelligence algorithm or program that generates responses or suggestions based on data it receives.
[1816] "Means for assisting conversation" refers to the method by which the generative AI model assists customer conversations based on the analysis results.
[1817] "Means for answering additional questions" refers to a method for generating and providing appropriate answers to new questions from customers.
[1818] "Means for introducing and making reservations at the next store" refers to a method for suggesting another store to the customer and making a reservation at that store if necessary.
[1819] "Means of recording who ordered what and allocating the bill individually" refers to a method of recording the order details by linking them to a specific customer and ultimately calculating the individual payment amount.
[1820] "Means for analyzing emotional state and adjusting quality of service" refers to a method for analyzing the emotions of a customer and providing service accordingly.
[1821] "Means for displaying real-time video on a smartphone or other mobile device" refers to methods and technologies for capturing video of a customer in real time and displaying it on a smartphone or other mobile device.
[1822] "Means for encrypting and securely transferring data to the server" refers to the technology and methods for encrypting collected data and transmitting it securely to the server while preventing unauthorized access from outside.
[1823] System Overview
[1824] The system of the present invention collects customer voice and video data and analyzes it on a server to improve the quality of service. The main components of the system are a server, a terminal, and a user, and is realized using various devices and software.
[1825] Hardware and Software Configuration
[1826] The main components running on the server and terminals are:
[1827] Hardware
[1828] Device: A mobile device such as a tablet or smartphone equipped with a camera and microphone to collect customer audio and video data.
[1829] Server: A powerful computer system responsible for analyzing data and running generative AI models.
[1830] Software and APIs
[1831] Speech Recognition API: Google Cloud Speech-to-Text API
[1832] Image Recognition API: Amazon Rekognition
[1833] Generative AI model: GPT-4 (OpenAI)
[1834] Emotion engine: IBM Watson Tone Analyzer
[1835] Database: Firebase Realtime Database
[1836] Communication protocol: HTTPS (SSL / TLS encryption)
[1837] Data processing and analysis flow
[1838] 1. Data Collection:
[1839] The device's camera and microphone are used to collect customer audio and video data.
[1840] The collected data is temporarily stored on the device.
[1841] 2. Data Transfer:
[1842] The collected data is sent to the server in real time using the HTTPS protocol.
[1843] Data is transferred securely via SSL / TLS encryption.
[1844] 3. Data Analysis:
[1845] The server converts the received voice data into text using the Google Cloud Speech-to-Text API.
[1846] The video data is analyzed using Amazon Rekognition to recognize customer facial expressions and behavior.
[1847] The analysis results are fed into a generative AI model (GPT-4) to generate responses and suggestions to customer questions and requests.
[1848] The emotion engine (IBM Watson Tone Analyzer) analyzes the customer's emotional state and adjusts the response content and service delivery method.
[1849] Provision of services
[1850] 1. Conversation assistance:
[1851] The server uses a generative AI model to generate appropriate answers to customer questions, such as suggesting menu recommendations in response to the question, "What's your recommendation today?"
[1852] The generated answers are displayed on the device screen and, if necessary, read aloud.
[1853] 2. Introduction and reservation of the following stores:
[1854] Based on the customer's request, the server searches for and suggests the next restaurant. The server processes the reservation and sends a confirmation message to the terminal.
[1855] 3. Accounting responsibilities:
[1856] The terminal records who has ordered what, and the server calculates the individual payment amount, which is displayed on the terminal for each customer to see.
[1857] Specific examples
[1858] For example, if a customer asks, "What dish do you recommend?" the system will process it as follows:
[1859] Prompt Sentence Examples
[1860] A customer asks, "What dish do you recommend?" Use the following information to generate a suitable answer:
[1861] Today's recommended dishes are "Omurice," "Steak," and "Salad."
[1862] The customer's emotional state is "interested."
[1863] answer:
[1864] Using this prompt, the generative AI model generates a response such as, "Today's recommended dishes are omelet rice, steak, and salad. They're all delicious. Let me know if you need more information," and displays it on the device. It also decides whether to provide more detailed information depending on the customer's emotional state.
[1865] The above is a specific embodiment for carrying out the invention. This system improves the quality of customer service and enables efficient business operations.
[1866] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1867] Step 1:
[1868] The user turns on the device (smartphone or tablet) and launches the dedicated application. The user then enables the camera and microphone. The camera and microphone collect the customer's audio and video data in real time. The customer's audio and video are captured as input and converted into digital data within the application.
[1869] Step 2:
[1870] The collected audio and video data is sent from the device to the server using the SSL / TLS encrypted HTTPS protocol. This data is securely received by the server. The collected audio and video data is sent as input, reaches the server, is safely received, and is output.
[1871] Step 3:
[1872] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. Voice data is given as input, and it is analyzed and converted into text data. The output is text data.
[1873] Step 4:
[1874] The server analyzes the received video data using the Amazon Rekognition API to recognize the customer's facial expressions and behavior. The video data is given as input, and it is analyzed to output information about facial expressions and behavior.
[1875] Step 5:
[1876] The analyzed audio and video data is integrated on the server and supplied to a generative AI model (GPT-4). The generative AI model creates prompts to generate appropriate answers to customer questions and requests. The textual customer utterances and emotional state are given as input, and an appropriate answer text is generated as output.
[1877] Step 6:
[1878] The server uses IBM Watson Tone Analyzer to analyze the customer's emotional state based on their text data. The generative AI model adjusts the answers and suggestions it provides based on their emotional state. The customer's text data is given as input, and the emotional state is obtained as output.
[1879] Step 7:
[1880] The generated response or suggestion is then re-encrypted and sent to the device, where it is received and displayed on the application screen and optionally read aloud. The generated response text is given as input, and the output is the answer appropriate to the customer's request.
[1881] Step 8:
[1882] Based on a specific need, if a customer requests "I would like to reserve the next store," the server receives the request, searches for an appropriate store based on the location information and desired conditions, and performs the reservation procedure. The customer's request is given as input, and a reservation confirmation message is generated as output.
[1883] Step 9:
[1884] When a customer places an order, the terminal sends the details to the server, which automatically records who ordered what. After the meal is finished, the server displays the bill based on the recorded order data and allocates it to each customer. The order data is given as input, and the individual payment amounts are calculated as output.
[1885] Step 10:
[1886] Based on the customer's emotional state, the server adjusts the response content and quality of service to provide optimal service. The emotion analysis results are given as input, and the adjusted service content is provided as output.
[1887] In this way, the system can improve customer service in real time and also improve operational efficiency.
[1888] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1889] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1890] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1891] [Fourth embodiment]
[1892] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1893] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1894] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1895] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1896] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1897] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1898] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1899] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1900] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1901] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1902] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1903] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1904] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1905] System Overview
[1906] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI to assist the customer's conversation, answering any follow-up questions, recommending a second restaurant, and making reservations. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily comprised of a server, a terminal (tablet), and a user (customer).
[1907] System Details
[1908] 1. Setup
[1909] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[1910] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with mechanisms for voice recognition, image recognition, and generative AI models.
[1911] 2. Real-time analysis of customer behavior and eating habits
[1912] Device: The tablet's camera and microphone are used to collect audio and video of the customer in real time. The sensors are periodically checked to ensure they are working properly.
[1913] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[1914] 3. Data analysis and response generation
[1915] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the customer's conversation, and extracts the necessary information.
[1916] Server: Feeds the analysis results to the generative AI model to generate appropriate answers to customer questions and suggestions for additional orders.
[1917] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[1918] Specific processing examples
[1919] Conversation assistance
[1920] 1. User: A customer asks, "What's your recommendation today?"
[1921] 2. Terminal: Collects the question voice and sends it to the server.
[1922] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[1923] 4. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[1924] Reservation for the second restaurant
[1925] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[1926] 2. Terminal: Sends the request to the server.
[1927] 3. Server: Based on the customer's location and desired conditions, the server searches for information on nearby bars and lists available reservations.
[1928] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[1929] 5. User: The customer selects a store and confirms the reservation.
[1930] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[1931] 7. Terminal: Display a confirmation message to the customer.
[1932] Accounting responsibilities
[1933] 1. Terminal: During the meal, you enter additional orders into the terminal and send the details to the server.
[1934] 2. Server: Automatically records who ordered what using image recognition technology.
[1935] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1936] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[1937] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[1938] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[1939] The processing flow will be explained below.
[1940] Collection and analysis of customer audio and video data
[1941] Step 1:
[1942] Device: The tablet activates the camera and microphone to capture the customer's voice and video in real time, capturing details of the customer's movements and conversations.
[1943] Step 2:
[1944] Terminal: The collected audio and video data is compressed and securely sent to the server. The data transfer protocol is encrypted to protect customer privacy.
[1945] Step 3:
[1946] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[1947] Conversation assistance and question handling
[1948] Step 1:
[1949] User: A customer asks, "What's your recommendation today?"
[1950] Step 2:
[1951] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[1952] Step 3:
[1953] Server: Analyzes the received text data, understands the content of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[1954] Step 4:
[1955] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[1956] Step 5:
[1957] Server: Sends the generated answer to the device.
[1958] Step 6:
[1959] Terminal: Receives the answer from the server and displays it to the customer. At the same time, it reads the answer aloud.
[1960] Introduction and reservation of the second restaurant
[1961] Step 1:
[1962] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[1963] Step 2:
[1964] Terminal: Sends the request contents to the server.
[1965] Step 3:
[1966] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[1967] Step 4:
[1968] Server: Sends the listed store information to the terminal.
[1969] Step 5:
[1970] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[1971] Step 6:
[1972] User: The customer selects the desired store and confirms the reservation.
[1973] Step 7:
[1974] Terminal: Sends the selected information to the server.
[1975] Step 8:
[1976] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[1977] Step 9:
[1978] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[1979] Accounting responsibilities
[1980] Step 1:
[1981] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[1982] Step 2:
[1983] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[1984] Step 3:
[1985] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[1986] Step 4:
[1987] User: The customer checks the details of each order and enters them into the tablet device.
[1988] Step 5:
[1989] Terminal: Sends input to the server.
[1990] Step 6:
[1991] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[1992] Step 7:
[1993] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[1994] The above are the specific processing steps in the system of the present invention, which can reduce customer waiting times and significantly improve store operational efficiency.
[1995] Example 1
[1996] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1997] Conventional customer service systems in restaurants lack the technology to effectively utilize customer voice and video data to assist conversations, answer additional questions, recommend and reserve the next restaurant, record who ordered what, and allocate individual bills. Therefore, there is a need to improve the quality of customer service and solve the issues of employee shortages and work efficiency.
[1998] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1999] In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to a data processing device, means for the data processing device to analyze the audio and video data and supply the analysis results to a generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to introduce and make reservations for the next store to be visited, and means for recording who ordered what and dividing up the bill individually. This enables a system that improves the quality of customer service and solves issues such as employee shortages and work efficiency.
[2000] "Means for collecting customer audio and video data" refers to equipment installed to collect customer audio and video information in real time, such as devices including cameras and microphones.
[2001] "Means for transmitting collected audio and video data to a data processing device" refers to network communication functions and protocols for transmitting audio and video data collected from customers to a data processing device (server) in a secure manner.
[2002] "Means for analyzing audio and video data using a data processing device and providing the analysis results to a generative AI model" refers to a process and system that analyzes data using speech recognition and image recognition algorithms running on a server and provides the analysis results to a generative AI model.
[2003] "Means for the generative AI model to assist customer conversations based on the analysis results" refers to a function in which the generative AI model provides information to assist customer questions and conversations based on analyzed data.
[2004] "Means for the generative AI model to answer follow-up questions from customers" refers to a system in which the generative AI model generates and provides appropriate answers to further questions from customers.
[2005] "Means for the generative AI model to introduce and make reservations for the next store to visit" is a function in which the generative AI model suggests the next store to visit based on the customer's preferences and carries out the reservation process.
[2006] "A means of recording who ordered what and dividing the bill individually" is a system that records the menu items ordered by customers and calculates and divides the amount due for each customer at the time of payment.
[2007] System Overview
[2008] The system of the present invention uses terminals installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes a generative AI model to assist customer conversations, answer follow-up questions, and recommend and make reservations for the next restaurant to visit. It also has the ability to record who ordered what and apportion the bill individually. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[2009] Detailed configuration and operation
[2010] set up
[2011] Terminal: A tablet is installed at each customer's seat and equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, enabling communication with the server.
[2012] Server: The server stores customer information, menu information, and information about participating restaurants in a database. It also sets up voice recognition software (e.g., Google Cloud Speech-to-Text), image recognition software (e.g., AWS Rekognition), and generative AI models (e.g., OpenAI GPT-4), and sets the necessary API keys.
[2013] Collection of customer audio and video data
[2014] Device: The device uses the tablet's camera and microphone to collect real-time audio and video data from customers. It periodically checks the operation of sensors and issues alerts if there are any problems.
[2015] Terminal: The collected data is encrypted and sent to the data processing device (server). Security protocols (e.g. SSL / TLS) are used to ensure the safety of the transfer.
[2016] Data Analysis and Response Generation
[2017] Server: The server uses a speech recognition algorithm to convert the received audio data into text, and an image recognition algorithm to analyze the video data.
[2018] Server: The analyzed data is fed into a generative AI model, which generates appropriate answers and suggestions based on the customer's questions and conversations.
[2019] Terminal: Displays generated responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[2020] Specific processing examples
[2021] Conversation assistance
[2022] 1. User: A customer asks, "What's your recommendation today?"
[2023] 2. Device: The device collects the audio of this question and sends it to the server.
[2024] 3. Server: The server analyzes the voice data and generates menu recommendation information using a generative AI model.
[2025] 4. Terminal: The terminal displays the recommended menu information received from the server to the customer and reads it out loud.
[2026] Prompt Sentence Examples
[2027] "What's today's recommended menu?"
[2028] "What's the special tonight?"
[2029] Introduction and reservation of next store to visit
[2030] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[2031] 2. Terminal: The terminal sends the request to the server.
[2032] 3. Server: The server searches for nearby bars based on the customer's location and desired conditions, and lists available bars.
[2033] 4. Terminal: The terminal displays the list of suggested stores to the customer and confirms the reservation.
[2034] 5. User: The customer selects a store and confirms the reservation.
[2035] 6. Server: The server completes the reservation procedure with the selected store and sends a confirmation message to the terminal.
[2036] 7. Terminal: The terminal displays a confirmation message to the customer.
[2037] Prompt Sentence Examples
[2038] "Find recommended bars near you."
[2039] "Can you help me book a bar?"
[2040] Accounting responsibilities
[2041] 1. Terminal: When an additional order is entered during the meal, the terminal sends the details to the server.
[2042] 2. Server: The server uses image recognition technology to automatically record who ordered what.
[2043] 3. Terminal: Once the meal is finished, the terminal provides a screen where customers can enter who ordered what.
[2044] 4. Server: The server calculates the bill and calculates the individual amounts to be shared.
[2045] 5. Terminal: The terminal displays the payment amount for each customer and provides a confirmation screen.
[2046] Example prompts for implementing the invention
[2047] Conversation assistance prompts:
[2048] "What's today's recommended menu?"
[2049] "What's the special tonight?"
[2050] Prompt for next store introduction and reservation:
[2051] "Find recommended bars near you."
[2052] "Can you help me book a bar?"
[2053] Prompt for accounting allocation:
[2054] "Calculate the amount each customer pays."
[2055] "Please split the bill based on your order."
[2056] In this way, the system of the present invention improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. Each component works in cooperation with the others, significantly improving the quality of service in restaurants.
[2057] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2058] Step 1:
[2059] set up
[2060] Devices: Tablets are installed at each customer's seat, and the camera, microphone, and touch screen are set up. The dedicated application is installed and the network connection is checked.
[2061] Input: Initial tablet information
[2062] Output: Tablet configured and network connection confirmed
[2063] Step 2:
[2064] Collection of customer audio and video data
[2065] Device: The tablet's camera and microphone are used to collect audio and video from customers in real time. The collected data is encrypted and sent to a server.
[2066] Input: Customer audio and video
[2067] Output: Encrypted audio and video data
[2068] Step 3:
[2069] Sending data to the server
[2070] Terminal: Collected audio and video data is sent to the server using a security protocol (e.g., SSL / TLS).
[2071] Input: Encrypted audio and video data
[2072] Output: Audio and video data received by the server
[2073] Step 4:
[2074] Audio and video data analysis
[2075] Server: Uses voice recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data, and uses image recognition software (e.g., AWS Rekognition) to analyze the video data.
[2076] Input: Audio and video data received by the server
[2077] Output: Parsed character data and analysis results
[2078] Step 5:
[2079] Feed to generative AI models and generate responses
[2080] Server: Feeds the analysis results to a generative AI model (e.g., OpenAI GPT-4) to generate appropriate answers and suggestions for customer questions.
[2081] Input: Parsed text data and video information
[2082] Output: Generated answers and suggestions
[2083] Step 6:
[2084] Sending the response to the terminal
[2085] Server: Sends generated answers and suggestions to the device.
[2086] Input: Generated answers and suggestions
[2087] Output: Answers and suggestions sent to the terminal
[2088] Step 7:
[2089] Displayed to customers and read aloud
[2090] Terminal: Displays responses and suggestions received from the server to the customer and reads them aloud if necessary.
[2091] Input: Response and proposal received from the server
[2092] Output: Information displayed to the customer and spoken aloud
[2093] Step 8:
[2094] Introducing the next store to visit and making reservations
[2095] 1. User: A customer requests a reservation for their next store visit.
[2096] 2. Terminal: Sends the request to the server.
[2097] 3. Server: Searches for nearby store information based on the customer's location and desired conditions, and lists stores that can be reserved.
[2098] 4. Terminal: Shows the customer a list of suggested stores and confirms the reservation.
[2099] 5. User: The customer selects a store and confirms the reservation.
[2100] 6. Server: Makes a reservation at the selected store and sends a confirmation message to the terminal.
[2101] 7. Terminal: Display a confirmation message to the customer.
[2102] Input: Customer request, location, preferences
[2103] Output: Reservation confirmation message
[2104] Step 9:
[2105] Accounting responsibilities
[2106] 1. Terminal: Additional orders are entered during the meal and sent to the server.
[2107] 2. Server: Automatically records who ordered what using image recognition technology.
[2108] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[2109] 4. Server: Calculates the total amount and calculates the individual amounts to be shared.
[2110] 5. Terminal: Displays the payment amount for each customer and provides a confirmation screen.
[2111] Input: Customer additional order information, image data
[2112] Output: Recorded order information, calculated billing amount
[2113] (Application example 1)
[2114] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2115] There is a need to solve issues such as improving work efficiency at logistics centers, responding immediately to worker questions, proposing optimal work spots, and managing individual work progress.It is also important to achieve smooth and efficient work by performing these tasks in real time.
[2116] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2117] In this invention, the server includes a means for collecting voice and video data of workers, a means for transmitting the collected voice and video data to the server, a means for the generative AI model to propose optimal work procedures in response to questions from workers, a means for the generative AI model to propose and reserve work spots, and a means for recording who performed which work and reporting individual progress. This makes it possible to improve work efficiency, respond to questions, propose optimal spots, and manage individual progress in real time.
[2118] "Audio and video data of workers" refers to the audio made by workers in a logistics center, as well as video footage of their appearance and actions.
[2119] A "generative AI model" refers to an artificial intelligence system that uses a generative approach to generate diverse outputs based on input data.
[2120] "Distribution center" refers to a facility that receives, stores, ships, and distributes goods and materials.
[2121] "Work procedures" refer to specific instructions and methods for effectively and efficiently carrying out various tasks within a logistics center.
[2122] "Work Spot" refers to a location or area within a distribution center where specific work is performed.
[2123] "Progress report" refers to a report by a worker on the progress or achievement of the work he or she has done.
[2124] System Overview
[2125] The system of the present invention collects voice and video data of workers at a logistics center and transmits it to a server in real time, and then utilizes a generative AI model to assist workers with their questions, suggest work spots, and report on the progress of individual tasks. The system is primarily composed of a server, terminals, and users (workers).
[2126] System Details
[2127] 1. Setup
[2128] Terminal: A robot is installed in each work area within the logistics facility. The robot is equipped with a camera, microphone, and display, and has a dedicated application installed. The robot is capable of communicating with the server.
[2129] Server: The server registers data related to workers, work spots, and work content in a database. It is equipped with mechanisms for running voice recognition, image recognition, and generative AI models.
[2130] 2. Real-time analysis of work status
[2131] Terminal: The robot's camera and microphone collect the worker's voice and video in real time. The sensors periodically check whether the robot is working properly.
[2132] Device: The collected data is sent to the server. The data is transferred using a security protocol, so it remains safe.
[2133] 3. Data analysis and response generation
[2134] Server: Analyzes the received audio and video data using speech and image recognition algorithms, analyzes the worker's questions, and extracts the necessary information.
[2135] Server: Provides the analysis results to the generation AI model, generating appropriate answers to worker questions and proposing work procedures.
[2136] Terminal: Displays responses and suggestions sent from the server to the worker and reads them out loud if necessary.
[2137] Specific processing examples
[2138] Work procedure suggestions
[2139] 1. User: A worker asks, "Please tell me where to place the next package most efficiently."
[2140] 2. Terminal: Collects the question voice and sends it to the server.
[2141] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[2142] 4. Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[2143] Work Spot Reservation
[2144] 1. User: A worker requests that the next work spot be reserved.
[2145] 2. Terminal: Sends the request to the server.
[2146] 3. Server: Suggests the optimal location based on the work content and information on available work spots within the logistics center.
[2147] 4. Terminal: Shows the worker a list of suggested spots and confirms the reservation.
[2148] 5. User: The worker selects a spot and confirms the reservation.
[2149] 6. Server: Makes a reservation for the selected spot and sends a confirmation message to the terminal.
[2150] 7. Terminal: Display a confirmation message to the worker.
[2151] Progress reports for individual tasks
[2152] 1. Terminal: Enter additional work while working and send the content to the server.
[2153] 2. Server: Records who has done what work and manages progress.
[2154] 3. Terminal: Provides the worker with a progress report screen when the work is completed.
[2155] 4. Server: Organizes the progress of work and reports on individual tasks.
[2156] 5. Terminal: Displays the progress of each worker and provides a confirmation screen.
[2157] Hardware and software used
[2158] Camera and microphone: Installed on the robot to collect the worker's voice and video.
[2159] Robot: Moves around the distribution center and is positioned in each designated work area.
[2160] Server: Runs speech and image recognition and generative AI models (e.g., GPT-3.5-turbo).
[2161] Prompt Sentence Examples
[2162] "Please tell me where to place the next shipment for maximum efficiency."
[2163] "Please reserve the next work spot for me."
[2164] In this way, this system aims to improve work efficiency at logistics centers and provides real-time assistance to workers.
[2165] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2166] Step 1:
[2167] User: A worker asks, "Please tell me where to place the next package most efficiently."
[2168] Input: The worker's spoken question.
[2169] Output: Audio data is input to the robot's microphone.
[2170] Step 2:
[2171] Terminal: Collects voice questions and sends them to the server.
[2172] Input: Worker voice data.
[2173] Data processing: Audio data captured by the robot's microphone is converted into a digital signal and sent to a server via a secure communication protocol.
[2174] Output: Audio data transferred to the server.
[2175] Step 3:
[2176] Server: Analyzes the voice data and uses a generative AI model to search for and generate optimal work procedure information.
[2177] Input: The transmitted audio data.
[2178] Data computation: Speech data is converted into text using a speech recognition algorithm, and then fed into a generative AI model (e.g., GPT-3.5-turbo) for analysis.
[2179] Output: Text data containing optimal work procedure information.
[2180] Step 4:
[2181] Server: Based on the analysis results, the AI model generates work procedures and sends them to the device.
[2182] Input: Text data containing optimal work procedure information.
[2183] Data calculation: The generative AI model generates appropriate work procedures based on the analysis results.
[2184] Output: Work procedure information sent to the robot's terminal.
[2185] Step 5:
[2186] Terminal: The work procedure information received from the server is displayed to the worker and read aloud.
[2187] Input: Work procedure information sent from the server.
[2188] Data processing: Work procedure information is displayed on the robot's display and read aloud using a speech synthesis algorithm.
[2189] Output: Visual and audio feedback to the worker.
[2190] Step 6:
[2191] User: A worker requests to reserve the next work spot.
[2192] Input: A voice request to reserve a work spot.
[2193] Output: Audio data is input to the robot's microphone.
[2194] Step 7:
[2195] Terminal: Sends the request to the server.
[2196] Input: Worker voice data.
[2197] Data processing: Converts voice data into digital signals and sends them to a server via a secure communication protocol.
[2198] Output: Audio data transferred to the server.
[2199] Step 8:
[2200] Server: Suggests the optimal location based on the work content and information on available work spots within the distribution center.
[2201] Input: Voice request data from workers and work spot data within the distribution center.
[2202] Data Computation: Uses speech recognition algorithms to convert voice data into text, search for available work spots, and use generative AI models to suggest the best locations.
[2203] Output: Text data containing the suggested work spot information.
[2204] Step 9:
[2205] Terminal: Displays the proposed spot list received from the server to the worker and confirms the reservation.
[2206] Input: Suggested spot information sent from the server.
[2207] Data processing: Display a list of suggested spots on the robot's display and provide an interface for reservation confirmation.
[2208] Output: A confirmation screen that is displayed to the worker.
[2209] Step 10:
[2210] User: The worker selects a spot and confirms the reservation.
[2211] Input: Work spot information selected by the worker.
[2212] Output: The input selection information is obtained through the robot interface.
[2213] Step 11:
[2214] Terminal: Sends the selected information to the server.
[2215] Input: Worker selection information.
[2216] Data processing: Selected data is converted into a digital signal and sent to the server via a secure communication protocol.
[2217] Output: The selected data that is transferred to the server.
[2218] Step 12:
[2219] Server: Makes a reservation at the selected spot and sends a confirmation message to the device.
[2220] Input: Worker selection information.
[2221] Data calculation: Executes the reservation procedure and generates a confirmation message.
[2222] Output: Sends a confirmation message to the robot's terminal.
[2223] Step 13:
[2224] Terminal: Display a confirmation message to the worker.
[2225] Input: The confirmation message sent by the server.
[2226] Data processing: Display the message on the robot's display.
[2227] Output: A confirmation message that is displayed to the worker.
[2228] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2229] System Overview
[2230] The system of the present invention uses tablets installed at customer tables in restaurants to collect customer voice and video in real time, and utilizes generative AI and an emotion engine to assist customer conversations, answer follow-up questions, and recommend and make reservations for second restaurants. It also records who ordered what and has the ability to allocate the bill individually, recognizing user emotions in real time to improve the quality of service. The system is primarily composed of a server, a terminal (tablet), and a user (customer).
[2231] System Details
[2232] 1. Setup
[2233] Device: A tablet is installed at the customer's seat, and a camera, microphone, and touch screen are set up. A dedicated application is installed on the tablet, and it is possible to communicate with the server.
[2234] Server: The server registers information about customers, menus, and the partner restaurant (second restaurant) in a database. It is equipped with a system that runs voice recognition, image recognition, generative AI models, and an emotion engine.
[2235] 2. Real-time analysis of customer behavior and eating habits
[2236] Device: The tablet's camera and microphone are used to collect the customer's voice and video in real time.
[2237] Terminal: Collected data is sent to the server. The data transfer protocol is encrypted and secure.
[2238] 3. Data analysis and response generation
[2239] Server: Analyzes the received audio and video data using voice and image recognition algorithms, analyzes the customer's conversation and the progress of their meal, and extracts the necessary information.
[2240] Server: Feeds the analysis results to the generative AI model and emotion engine to generate appropriate answers to customer questions and suggestions for additional orders.
[2241] Device: Displays responses and suggestions sent from the server to the customer and reads them aloud if necessary.
[2242] 4. Use of Emotion Engine
[2243] Server: The emotion engine analyzes the customer's facial expressions and tone of voice based on collected audio and video data to grasp their emotional state (e.g., joy, displeasure, interest, anger, etc.) in real time.
[2244] Server: Based on the analysis results of the emotion engine, the generative AI model adjusts the responses and assistance provided accordingly. For example, if the customer shows signs of discomfort, the server softens the response and refrains from making meal suggestions.
[2245] Specific processing examples
[2246] Conversation assistance
[2247] 1. User: A customer asks, "What's your recommendation today?"
[2248] 2. Terminal: Collects voice questions, converts them into text, and sends them to the server.
[2249] 3. Server: Analyzes the voice data and uses a generative AI model to search for and generate recommended menu information.
[2250] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the generated response appropriately (e.g., adding more information if the customer is interested).
[2251] 5. Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[2252] Reservation for the second restaurant
[2253] 1. User: A customer requests, "Please make a reservation for the next bar I want to go to."
[2254] 2. Terminal: Sends the request to the server.
[2255] 3. Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. It lists available establishments and retrieves detailed information from the database.
[2256] 4. Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations (e.g., if the customer indicates they want to relax, prioritize quiet bars).
[2257] 5. Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[2258] 6. User: The customer selects the desired store and confirms the reservation.
[2259] 7. Terminal: Sends the selection information to the server.
[2260] 8. Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[2261] 9. Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[2262] Accounting responsibilities
[2263] 1. Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[2264] 2. Server: Automatically records who ordered what. Using image recognition technology, it automatically tracks who placed what order.
[2265] 3. Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[2266] 4. User: The customer checks the individual order details and enters them into the tablet device.
[2267] 5. Terminal: Sends the input to the server.
[2268] 6. Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user contributes based on the items ordered.
[2269] 7. Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each person's payment.
[2270] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[2271] The processing flow will be explained below.
[2272] Collection and analysis of customer audio and video data
[2273] Step 1:
[2274] Device: The tablet activates the camera and microphone to collect the customer's voice and video in real time, capturing the content of the conversation and the customer's behavior.
[2275] Step 2:
[2276] Terminal: Collected audio and video data is temporarily stored in internal memory and compressed, allowing for efficient data transfer.
[2277] Step 3:
[2278] Terminal: Compressed audio and video data is sent to the server. The data is encrypted to prevent information leakage during transmission.
[2279] Step 4:
[2280] Server: The received voice data is processed through a voice recognition algorithm to convert it into text data. At the same time, the video data is processed through an image recognition algorithm to analyze the customer's facial expressions and the progress of their meal.
[2281] Conversation assistance and question handling
[2282] Step 1:
[2283] User: A customer asks, "What's your recommendation today?"
[2284] Step 2:
[2285] Terminal: Collects voice questions, converts them into text, and sends them to the server.
[2286] Step 3:
[2287] Server: Analyzes the received text data, understands the intent of the question, and uses a generative AI model to search a database for the recommended menu for that day.
[2288] Step 4:
[2289] Server: Generates menu recommendations and adds specific details (e.g., ingredients, allergy information, etc.).
[2290] Step 5:
[2291] Server: Uses analyzed customer sentiment data to adjust the tone and content of the responses it generates, for example, providing more detailed information if the customer is interested.
[2292] Step 6:
[2293] Server: Sends the generated answer to the device.
[2294] Step 7:
[2295] Terminal: The recommended menu information received from the server is displayed to the customer and read aloud.
[2296] Introduction and reservation of the second restaurant
[2297] Step 1:
[2298] User: A customer requests, "Please make a reservation for the next bar we'll go to."
[2299] Step 2:
[2300] Terminal: Sends the request contents to the server.
[2301] Step 3:
[2302] Server: Searches for nearby bars and karaoke establishments based on the customer's location and preferences. Lists establishments that can be booked and retrieves detailed information from the database.
[2303] Step 4:
[2304] Server: The emotion engine analyzes the customer's emotional state and adjusts the store's recommendations accordingly. For example, if a customer indicates that they want to relax, it will prioritize quiet bars.
[2305] Step 5:
[2306] Server: Sends the listed store information to the terminal.
[2307] Step 6:
[2308] Terminal: A list of suggested stores is displayed to the customer. When the customer selects one, a reservation confirmation screen is displayed.
[2309] Step 7:
[2310] User: The customer selects the desired store and confirms the reservation.
[2311] Step 8:
[2312] Terminal: Sends the selected information to the server.
[2313] Step 9:
[2314] Server: Makes a reservation at the selected store and sends a reservation confirmation message to the terminal.
[2315] Step 10:
[2316] Terminal: Display a confirmation message to the customer to let them know that their booking is confirmed.
[2317] Accounting responsibilities
[2318] Step 1:
[2319] Terminal: When a customer orders additional food or drinks, the details are immediately sent to the server.
[2320] Step 2:
[2321] Server: Automatically records the received order details. Using image recognition technology, it automatically tracks who placed which order.
[2322] Step 3:
[2323] Terminal: Provides a screen where customers can enter who ordered what once they have finished their meal.
[2324] Step 4:
[2325] User: The customer checks the details of each order and enters them into the tablet device.
[2326] Step 5:
[2327] Terminal: Sends input to the server.
[2328] Step 6:
[2329] Server: Calculates the payment amount for each customer based on the total bill amount. Calculates the amount each user pays based on the items ordered.
[2330] Step 7:
[2331] Terminal: Displays the calculated individual payments to the customer and provides an interface for confirming each payment.
[2332] Step 8:
[2333] Server: The emotion engine analyzes the customer's emotional state at checkout and adjusts payment methods and responses as needed.
[2334] In this way, this system significantly improves customer convenience and solves the problems of staff shortages and operational efficiency in restaurants. In particular, by combining it with an emotion engine, it becomes possible to adjust services according to the customer's emotional state, providing a more personalized experience. Each component of the system works in conjunction with each other to significantly improve the quality of service in restaurants.
[2335] Example 2
[2336] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2337] Modern restaurants are required to improve customer convenience while simultaneously increasing operational efficiency. However, the number of employees is limited, and it is often difficult to provide adequate service, especially during busy periods. It is also difficult to grasp customers' emotional state in real time and provide service accordingly. Furthermore, it is time-consuming to accurately record who ordered what and assign individual tasks at the checkout. A system is needed to solve these issues and improve customer satisfaction.
[2338] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for encrypting the customer's voice and video data and transmitting it to the server, means for analyzing the data using voice recognition and image recognition algorithms and providing the analysis results to the generative AI model and emotion analysis engine, and means for the generative AI model to assist the customer in the conversation based on the analysis results. This enables secure transfer and analysis of the customer's voice and video data, immediate conversation support based on the analysis results, and adjustment of response content according to emotions. It also includes functions for introducing and making reservations for the next store and sharing individual billing responsibilities, greatly improving customer convenience.
[2339] "Audio and video data" refers to digital data that records the voice and video of a customer.
[2340] "Encryption" is a technology that converts data into a form that cannot be understood by third parties in order to transfer it securely.
[2341] A "server" is a computer system that handles back-end operations such as data management, analysis, and communication.
[2342] A "voice recognition algorithm" is a program for converting voice data into text data.
[2343] An "image recognition algorithm" is a program that analyzes video data to detect specific patterns or objects.
[2344] A "generative AI model" is an artificial intelligence model that generates natural language responses or suggestions based on input data.
[2345] An "emotion analysis engine" is a program for analyzing an individual's emotional state from audio and video data.
[2346] "Conversational assistance" means providing appropriate responses and suggestions to customer questions and requests.
[2347] The "means for answering additional questions" is a function for generating answers to additional questions from customers.
[2348] "Next store introduction and reservation" is a function that suggests the next store that the customer would like to visit and makes a reservation for it.
[2349] The "means for dividing the bill amount individually" is a function that records the order details for each customer and calculates the individual payment amount based on that.
[2350] The system of the present invention is designed to improve customer convenience in restaurants and is composed of multiple elements including terminals (tablets), servers, and users (customers). Each element will be described in detail below.
[2351] 1. Setup
[2352] Device:
[2353] The tablet is installed at the customer's seat and is equipped with a camera, microphone, and touch screen. A dedicated application is installed on the tablet, and it can communicate with the server. The tablet uses a Wi-Fi connection, and plugins and drivers are pre-configured. Normal operation is initiated by clicking the "Start" button on the main screen.
[2354] server:
[2355] The server registers customer information, menus, and data on affiliated restaurants in a database. The database uses MySQL, and an environment is set up in which voice recognition, image recognition, generative AI models, and an emotion analysis engine can operate. The necessary software and libraries are installed on the server in advance.
[2356] 2. Real-time analysis of customer behavior and eating habits
[2357] Device:
[2358] The tablet's built-in camera and microphone are used to collect customer audio and video data in real time. For example, when a customer speaks into the tablet, their audio and video are recorded.
[2359] Device:
[2360] The collected data is encrypted using the AES encryption algorithm and transmitted to the server via the HTTPS protocol, which ensures the security of the data.
[2361] 3. Data analysis and response generation
[2362] server:
[2363] The server uses a speech recognition algorithm (e.g., Reve) to convert the voice data into text. For example, the voice data "What's the recommendation today?" is converted into text data "What's the recommendation today?"
[2364] server:
[2365] Image recognition algorithms (e.g., OpenCV) are used to analyze video data and grasp customer facial expressions and movements in real time.
[2366] server:
[2367] A generative AI model (e.g., GPT-3) generates an appropriate response based on the analysis results. For example, in response to the text data "What's today's recommendation?", it generates a response such as "Today's recommendation is seafood pasta."
[2368] Device:
[2369] The response data from the server is sent to the tablet, displayed on the screen, and read aloud using a Text-to-Speech (TTS) engine.
[2370] 4. Use of sentiment analysis engines
[2371] server:
[2372] The emotion analysis engine analyzes the customer's emotional state from audio and video data. For example, if the customer is smiling, it is determined to be "happy."
[2373] server:
[2374] The generative AI model adjusts responses based on the results of sentiment analysis. For example, if a customer expresses displeasure, the next response will be generated in a softer tone. For example, a response such as, "May I suggest some other dishes so you can enjoy them?"
[2375] Examples of concrete examples and prompts
[2376] Conversation Assistance:
[2377] When a user asks, "What's the recommendation today?", the device collects the voice and sends it to the server, which analyzes the voice and generates a response such as, "Today's recommendation is seafood pasta."
[2378] 2nd restaurant reservation:
[2379] When a customer requests a reservation for the next bar they want to visit, the server will suggest a bar based on the customer's desired conditions and make the reservation.
[2380] Example prompt sentence:
[2381] "When a customer asks, 'What's the special today?' show them the current menu specials and read out the details."
[2382] With the above configuration, the present invention significantly improves customer convenience and improves the operational efficiency of restaurants. In particular, by combining it with an emotion analysis engine, it becomes possible to respond flexibly to customer emotions, thereby providing more personalized service.
[2383] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2384] Step 1:
[2385] Device:
[2386] A tablet is placed at the customer's seat, and the camera, microphone, and touchscreen are set up. A dedicated application is then launched. When the customer speaks into the tablet, the camera and microphone collect audio and video. The collected data is processed in real time, resulting in minimal latency.
[2387] Input: Customer audio and video data
[2388] Output: Encrypted audio and video data
[2389] For example, a customer asks, "What's the recommendation today?" Audio and video are captured and the application encrypts the data.
[2390] Step 2:
[2391] Device:
[2392] The collected data is encrypted using the AES encryption algorithm and sent to the server via the HTTPS protocol, ensuring data security and privacy.
[2393] Input: Unencrypted audio and video data
[2394] Output: Encrypted data sent to the server
[2395] As a specific example of operation, the encrypted data is sent to the server, and a message indicating that the transfer is complete is displayed.
[2396] Step 3:
[2397] server:
[2398] The voice data that arrives at the server is analyzed using a voice recognition algorithm and converted into text data. For example, the voice data "What's recommended today?" is converted into text.
[2399] Input: Encrypted audio data
[2400] Output: Text data
[2401] As a specific example of how it works, the voice recognition algorithm is activated and the result "What's recommended today?" is output.
[2402] Step 4:
[2403] server:
[2404] Image recognition algorithms are used to analyze the customer's facial expressions and movements from the incoming video data, thereby capturing the customer's real-time emotional state.
[2405] Input: Encrypted video data
[2406] Output: Emotional state data
[2407] As a specific example of operation, an image recognition algorithm is run to determine that "the customer is showing interest."
[2408] Step 5:
[2409] server:
[2410] A generative AI model is used to generate an appropriate response based on the analyzed text data and emotional state, for example, "Today's recommendation is seafood pasta."
[2411] Input: Text data and emotional state data
[2412] Output: Response message
[2413] As a specific example of how it works, the generative AI model generates the message "Today's recommendation is seafood pasta."
[2414] Step 6:
[2415] server:
[2416] The generated response message is sent to the terminal.
[2417] Input: Response message
[2418] Output: Message sent to the terminal
[2419] As a specific example of operation, a response message is sent from the server to the terminal.
[2420] Step 7:
[2421] Device:
[2422] The terminal receives the response message, displays it on the screen, and uses a Text-to-Speech (TTS) engine to read it aloud to the customer.
[2423] Input: Received response message
[2424] Output: Voice and text output to the customer
[2425] As a specific example of how it works, the message "Today's recommendation is seafood pasta" is displayed on the screen and simultaneously read aloud.
[2426] By having each step work in conjunction with each other, we aim to create a system that can respond to customers quickly and appropriately.
[2427] (Application example 2)
[2428] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2429] In recent years, restaurants have been required to improve the quality of customer service and streamline operations. However, staff shortages and inconsistent service quality have become problems. In particular, there are limitations to providing multilingual support and individualized service in real time, making it difficult to maintain customer satisfaction. It is also difficult to allocate individual amounts at the time of payment and provide service that responds to customer emotions.
[2430] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting customer audio and video data, means for transmitting the collected audio and video data to the server, means for the server to analyze the audio and video data and provide the analysis results to the generative AI model, means for the generative AI model to assist the customer in the conversation based on the analysis results, means for the generative AI model to answer any additional questions from the customer, means for the generative AI model to recommend and reserve the next restaurant, means for recording who ordered what and individually apportioning the bill, means for analyzing the customer's emotional state and adjusting the quality of service, means for displaying real-time video of the customer on a smartphone or other mobile device, and means for encrypting the collected data and securely transferring it to the server. This improves the quality of customer service and enables efficient business operations. It is also possible to provide multilingual support, real-time personalized services, and flexible service provision based on individual apportionment of the bill and emotions.
[2431] "Customer audio and video data" refers to audio produced by the customer and video information such as the customer's actions and facial expressions.
[2432] "Collecting means" refers to devices and methods for capturing audio and video data.
[2433] "Transmission means" refers to the communications technologies and protocols used to transmit collected data to the server.
[2434] "Server" refers to the computer system that analyzes collected data and runs the generative AI model and emotion engine.
[2435] "Means for analysis" refers to the algorithms and software used to analyze audio and video data and extract the necessary information.
[2436] "Generative AI model" refers to an artificial intelligence algorithm or program that generates responses or suggestions based on data it receives.
[2437] "Means for assisting conversation" refers to the method by which the generative AI model assists customer conversations based on the analysis results.
[2438] "Means for answering additional questions" refers to a method for generating and providing appropriate answers to new questions from customers.
[2439] "Means for introducing and making reservations at the next store" refers to a method for suggesting another store to the customer and making a reservation at that store if necessary.
[2440] "Means of recording who ordered what and allocating the bill individually" refers to a method of recording the order details by linking them to a specific customer and ultimately calculating the individual payment amount.
[2441] "Means for analyzing emotional state and adjusting quality of service" refers to a method for analyzing the emotions of a customer and providing service accordingly.
[2442] "Means for displaying real-time video on a smartphone or other mobile device" refers to methods and technologies for capturing video of a customer in real time and displaying it on a smartphone or other mobile device.
[2443] "Means for encrypting and securely transferring data to the server" refers to the technology and methods for encrypting collected data and transmitting it securely to the server while preventing unauthorized access from outside.
[2444] System Overview
[2445] The system of the present invention collects customer voice and video data and analyzes it on a server to improve the quality of service. The main components of the system are a server, a terminal, and a user, and is realized using various devices and software.
[2446] Hardware and Software Configuration
[2447] The main components running on the server and terminals are:
[2448] Hardware
[2449] Device: A mobile device such as a tablet or smartphone equipped with a camera and microphone to collect customer audio and video data.
[2450] Server: A powerful computer system responsible for analyzing data and running generative AI models.
[2451] Software and APIs
[2452] Speech Recognition API: Google Cloud Speech-to-Text API
[2453] Image Recognition API: Amazon Rekognition
[2454] Generative AI model: GPT-4 (OpenAI)
[2455] Emotion engine: IBM Watson Tone Analyzer
[2456] Database: Firebase Realtime Database
[2457] Communication protocol: HTTPS (SSL / TLS encryption)
[2458] Data processing and analysis flow
[2459] 1. Data Collection:
[2460] The device's camera and microphone are used to collect customer audio and video data.
[2461] The collected data is temporarily stored on the device.
[2462] 2. Data Transfer:
[2463] The collected data is sent to the server in real time using the HTTPS protocol.
[2464] Data is transferred securely via SSL / TLS encryption.
[2465] 3. Data Analysis:
[2466] The server converts the received voice data into text using the Google Cloud Speech-to-Text API.
[2467] The video data is analyzed using Amazon Rekognition to recognize customer facial expressions and behavior.
[2468] The analysis results are fed into a generative AI model (GPT-4) to generate responses and suggestions to customer questions and requests.
[2469] The emotion engine (IBM Watson Tone Analyzer) analyzes the customer's emotional state and adjusts the response content and service delivery method.
[2470] Provision of services
[2471] 1. Conversation assistance:
[2472] The server uses a generative AI model to generate appropriate answers to customer questions, such as suggesting menu recommendations in response to the question, "What's your recommendation today?"
[2473] The generated answers are displayed on the device screen and, if necessary, read aloud.
[2474] 2. Introduction and reservation of the following stores:
[2475] Based on the customer's request, the server searches for and suggests the next restaurant. The server processes the reservation and sends a confirmation message to the terminal.
[2476] 3. Accounting responsibilities:
[2477] The terminal records who has ordered what, and the server calculates the individual payment amount, which is displayed on the terminal for each customer to see.
[2478] Specific examples
[2479] For example, if a customer asks, "What dish do you recommend?" the system will process it as follows:
[2480] Prompt Sentence Examples
[2481] A customer asks, "What dish do you recommend?" Use the following information to generate a suitable answer:
[2482] Today's recommended dishes are "Omurice," "Steak," and "Salad."
[2483] The customer's emotional state is "interested."
[2484] answer:
[2485] Using this prompt, the generative AI model generates a response such as, "Today's recommended dishes are omelet rice, steak, and salad. They're all delicious. Let me know if you need more information," and displays it on the device. It also decides whether to provide more detailed information depending on the customer's emotional state.
[2486] The above is a specific embodiment for carrying out the invention. This system improves the quality of customer service and enables efficient business operations.
[2487] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2488] Step 1:
[2489] The user turns on the device (smartphone or tablet) and launches the dedicated application. The user then enables the camera and microphone. The camera and microphone collect the customer's audio and video data in real time. The customer's audio and video are captured as input and converted into digital data within the application.
[2490] Step 2:
[2491] The collected audio and video data is sent from the device to the server using the SSL / TLS encrypted HTTPS protocol. This data is securely received by the server. The collected audio and video data is sent as input, reaches the server, is safely received, and is output.
[2492] Step 3:
[2493] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. Voice data is given as input, and it is analyzed and converted into text data. The output is text data.
[2494] Step 4:
[2495] The server analyzes the received video data using the Amazon Rekognition API to recognize the customer's facial expressions and behavior. The video data is given as input, and it is analyzed to output information about facial expressions and behavior.
[2496] Step 5:
[2497] The analyzed audio and video data is integrated on the server and supplied to a generative AI model (GPT-4). The generative AI model creates prompts to generate appropriate answers to customer questions and requests. The textual customer utterances and emotiona...
Claims
1. a means for collecting customer audio and video data; means for transmitting the collected audio and video data to a server; A means for analyzing the audio and video data on a server and providing the analysis results to the generative AI model; A means for the generative AI model to assist customer conversations based on the analysis results, and A means for the generative AI model to answer follow-up customer questions; and A means for the generative AI model to recommend and book a second restaurant; and A way to record who ordered what and share the bill individually Including system.
2. A means to grasp the progress of customers' meals in real time using image recognition, including the means to propose additional orders based on that progress; The system of claim 1 .
3. Including means to recognize customer conversations in multiple languages and provide appropriate information and responses, The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A