system

The system addresses the complexity of current ordering systems by enabling voice-based ordering with AI-driven personalization, enhancing user experience and operational efficiency.

JP2026023371APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024125306
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Current ordering systems for restaurants and other establishments are cumbersome for elderly people and those with low digital literacy, leading to inefficient order processing and low customer satisfaction due to complex operations and lack of personalized recommendations.

Method used

A system that includes a microphone for voice orders, generative AI for data analysis, a display for confirmation, a database for order storage, and a POS system integration, allowing intuitive voice-based ordering with personalized menu recommendations based on past order history and customer attributes.

Benefits of technology

Enables efficient and intuitive ordering, improving customer satisfaction and operational efficiency by allowing voice-based ordering and personalized menu suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023371000001_ABST
    Figure 2026023371000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a microphone means for receiving an order by voice, a generation and AI means for analyzing the received voice, a display means for confirming and correcting the analyzed order contents, a means for storing the analyzed order contents in a database, and a means for linking the stored order contents to a POS system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In current ordering systems for restaurants and other establishments, ordering using a tablet device is complicated to operate, making it difficult to use, especially for elderly people and users with low digital literacy. In addition, it is difficult to process orders efficiently and browse menus, making it difficult to improve turnover and customer satisfaction. [Means for solving the problem]

[0005] The present invention is a system that includes a microphone means for accepting voice orders, a generation AI means for analyzing the accepted voice data, a display means for confirming and correcting the analyzed order details, a means for saving the analyzed order details in a database, and a means for linking the saved order data to a POS system.Furthermore, by including a means for analyzing customer voiceprint data and generating and displaying a recommended menu based on past order history and customer attribute information, a more intuitive and efficient ordering process is realized.

[0006] "Voice ordering" is the process by which a customer selects a product or service by speaking the words, and the system recognizes the selection.

[0007] The "microphone means" is a device for capturing the voice of the customer and transmitting the voice data to the inside of the system.

[0008] A "generative AI means" is a processing system that uses artificial intelligence technology to analyze received voice data and convert it into text data.

[0009] The "display means" is a device such as a display or screen for visually presenting the analyzed order details and confirmation messages to the customer.

[0010] "Means for saving in a database" refers to data storage for permanently saving the analyzed order details.

[0011] A "POS system" is a sales information management system used to manage order data and process accounting.

[0012] "Voiceprint data" refers to data used to analyze and identify the voice characteristics of individual customers.

[0013] "Order history" refers to data that records the contents of orders placed by customers in the past.

[0014] A "recommended menu" is a menu that recommends products and services suitable for a customer based on their past order history and customer attribute information. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention is a system for efficiently and intuitively placing orders using voice. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[0037] User operations

[0038] 1. Select the menu

[0039] The user selects items from a printed menu or a display screen.

[0040] For example, suppose a user wants to order "carbonara" and "cola."

[0041] 2. Voice ordering

[0042] The user orders the selected items by voice into the microphone, for example, saying, "One carbonara and one coke, please."

[0043] Terminal handling

[0044] 1. Audio capture

[0045] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[0046] 2. Display Recommendations

[0047] The terminal analyzes the user's voiceprint and displays a recommended menu on the screen based on past order history and attribute information.

[0048] For example, if a user has previously ordered "cheesecake," the system displays "Would you like dessert?"

[0049] Server Processing

[0050] 1. Audio analysis

[0051] The server analyzes the voice data and converts the content into text using generative AI.

[0052] For example, audio saying "One carbonara and one cola please" is converted into text data as "1 carbonara, 1 cola."

[0053] 2. Confirm your order details

[0054] The server may return a confirmation question to the user if necessary, generating a message such as "Would you like one carbonara and one coke?"

[0055] 3. Order Data

[0056] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[0057] 4. Order Linkage

[0058] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[0059] The kitchen displays an order for "one carbonara" and "one coke."

[0060] The cash register then connects the product information to prepare for payment.

[0061] Specific examples

[0062] 1. User Orders

[0063] A user says, "Two pizza margaritas and one Pepsi, please."

[0064] 2. Capturing and transmitting audio from your device

[0065] The microphone captures the user's speech and transmits the data to a server.

[0066] 3. Server audio analysis

[0067] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi."

[0068] 4. Server Order Confirmation

[0069] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two Margherita pizzas and one Pepsi?"

[0070] 5. User Verification

[0071] The user presses the confirmation button or answers "yes" aloud.

[0072] 6. Order data conversion and storage

[0073] The server stores the order data and links it to past order history and voiceprint data.

[0074] 7. Order Linkage

[0075] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[0076] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, it is possible to provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[0080] Step 2:

[0081] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0082] Step 3:

[0083] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[0084] Step 4:

[0085] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[0086] Step 5:

[0087] The server parses the text of the order and generates a confirmation message if necessary, such as "One carbonara and one coke, okay?"

[0088] Step 6:

[0089] The terminal displays a confirmation message on the display, allowing the user to visually confirm the order details.

[0090] Step 7:

[0091] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[0092] Step 8:

[0093] The server receives the user's response and saves the final order details in a database, including the customer ID, product name, and quantity.

[0094] Step 9:

[0095] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[0096] Step 10:

[0097] The device displays a recommended menu on the screen. For example, if a user has previously ordered dessert, the device might recommend, "Would you like a dessert?"

[0098] Step 11:

[0099] If the user wishes to place an additional order, the same process is repeated.

[0100] Step 12:

[0101] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[0102] These steps allow users to easily place orders by voice, and the store can efficiently manage order details and provide service quickly.

[0103] Example 1

[0104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0105] Conventional ordering systems require users to select products and manually input the exact order details, which is a cumbersome and time-consuming process. Stores also face issues with low order processing efficiency, with orders often being missed or delayed, especially during peak hours. Furthermore, there is a lack of personalized recommendations based on customers' past order history and attribute information, creating a need for improved customer satisfaction.

[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0107] In this invention, the server includes an input means for accepting voice orders, an AI generation means, and an output means, which makes it possible to process orders more efficiently and provide personalized services based on the customer's past order history and attribute information.

[0108] The "input means for accepting voice orders" refers to a data input device for capturing the user's voice. Specifically, it refers to a voice input device such as a microphone.

[0109] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting voice data into text data, including speech recognition models and natural language processing algorithms.

[0110] "Output means for confirming or correcting the analyzed order details" refers to a device that presents the analyzed order details to the user for confirmation or correction. Specifically, this refers to a display, voice response unit, etc.

[0111] The "means for saving the analyzed order details in a database" refers to a storage system for permanently saving the analyzed order details. Specifically, this refers to a relational database, cloud storage, etc.

[0112] "Means for linking stored order data to an information system" refers to communication means for transferring or sharing stored order data with other information systems, such as POS systems. Specifically, this refers to APIs and network interfaces.

[0113] "Means for analyzing customer voiceprint data and generating recommended menus based on past order history and customer attribute information" refers to an algorithm for generating recommended menus using customer voiceprint data, past order history, and attribute information. Specifically, this refers to a machine learning model or recommendation engine.

[0114] The present invention is a system for efficiently and intuitively placing orders using voice. To specifically implement this system, the program processing will be described in detail based on the roles of the server, terminal, and user.

[0115] User operations

[0116] The user selects items from a printed menu or a display screen. For example, if the user wants to order "carbonara" and "Coke," they would say into the microphone, "One carbonara and one Coke, please."

[0117] Terminal handling

[0118] A microphone is connected to the device, which captures the user's voice. This voice data is sent to the server in real time. The device then analyzes the user's voiceprint and displays recommended menu items on the display based on the user's past order history and attribute information. For example, if a user has previously ordered cheesecake, the device might ask, "Would you like dessert?"

[0119] Server Processing

[0120] The server analyzes the voice data and converts it into text using a generative AI model. For example, a user's speech, "One carbonara and one cola, please," is converted into text as "1 carbonara, 1 cola." The server then generates a confirmation message for the user, displaying "Is one carbonara and one cola okay?" Once the user confirms, the server saves the order data in a database. The server also connects the saved order data to the store's information system (POS system) and notifies the kitchen and cash register.

[0121] Specific examples

[0122] The user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures this voice and sends it to the server. The server uses a generative AI model to analyze the voice data and converts it into text: "Two Margherita Pizzas, one Pepsi." It then generates a confirmation message asking, "Are you sure your order is for two Margherita Pizzas and one Pepsi?" and displays it on the device's screen. Once the user confirms, the server saves this order data in a database and links the order data to the store's POS system. As a result, an order for "two Margherita Pizzas" and "one Pepsi" is displayed on the kitchen display, and the product information is sent to the cash register, completing the checkout process.

[0123] In this system, a generative AI model is used to analyze voice data. The model is input with the following prompt: "One carbonara and one coke, please." The model analyzes this prompt and generates the order details as text data.

[0124] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, stores can provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0126] Step 1:

[0127] The user selects a menu item and orders by voice.

[0128] Input: The user selects a menu item and decides what to order.

[0129] Action: The user speaks into the microphone, "One carbonara and one coke please."

[0130] Output: The microphone captures the user's voice.

[0131] Step 2:

[0132] The device captures the audio and sends it to the server

[0133] Input: User's spoken order.

[0134] How it works: The device's microphone captures audio data and converts it into a digital format.

[0135] Output: The device sends the converted audio data to the server in real time.

[0136] Step 3:

[0137] The server analyzes the voice data and converts it into text

[0138] Input: Audio data sent from the device.

[0139] How it works: The server inputs speech data into the generative AI model. For example, it analyzes speech data such as "One carbonara and one coke, please."

[0140] Data processing: The voice data is input as a prompt sentence into the generative AI model and converted into text data.

[0141] Output: Text data: "1 carbonara, 1 cola".

[0142] Step 4:

[0143] The server verifies the order and asks the user for confirmation.

[0144] Input: Text of the order.

[0145] What happens: The server generates a confirmation message saying "Would you like one carbonara and one coke?"

[0146] Output: A confirmation message will be printed to the terminal screen.

[0147] Step 5:

[0148] The user confirms the order details

[0149] Input: The confirmation message displayed on the device screen.

[0150] Action: The user presses the confirmation button or speaks "yes."

[0151] Output: The user's confirmation action is entered into the terminal.

[0152] Step 6:

[0153] The server saves the order data to a database

[0154] Input: The order details confirmed by the user.

[0155] What happens: The server records a data entry in the database, such as "Customer ID: 12345, Product: Carbonara, Quantity: 1."

[0156] Output: Order data is saved in the database.

[0157] Step 7:

[0158] The server connects the order data to the information system

[0159] Input: Order data stored in the database.

[0160] How it works: The server sends order data to the store's POS system.

[0161] Output: The kitchen display shows the order for "1 Carbonara" and "1 Coke", and the product information is sent to the cashier terminal, ready for payment.

[0162] (Application example 1)

[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0164] In recent years, the use of food delivery services has increased, but there is a demand for improved user convenience and efficient order management. Conventional systems require users to operate an application interface, which is cumbersome, especially for elderly people and users unfamiliar with technology. In addition, the cumbersome process of checking order details and managing order history affects the operational efficiency of food delivery services. To solve these issues, a system that allows users to order intuitively by voice is needed.

[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0166] In this invention, the server includes microphone means for accepting voice orders, AI generation means for analyzing the accepted voice data, display means for confirming and correcting the analyzed order details, means for saving the analyzed order details in a database, means for retrieving the saved order data from the database and linking it to the server, and generating and displaying a confirmation message, means for linking the order details to the food delivery service server and delivery management system, and means for linking the order details with past order history and voiceprint data to provide more appropriate service. This enables users to intuitively place food delivery orders by voice, streamlining the ordering process and improving the user experience.

[0167] "Microphone means for accepting voice orders" refers to devices or technology for detecting and capturing voice.

[0168] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting captured voice data into text and analyzing its meaning.

[0169] "Display means for confirming and correcting the analyzed order contents" refers to a device such as a display or monitor that visually presents the analysis results to the user and allows the user to confirm or correct the contents.

[0170] The "means for storing analyzed order details in a database" refers to a database system for electronically recording and managing analyzed order data.

[0171] "Means for retrieving saved order data from a database, linking to a server, and generating and displaying a confirmation message" refers to a system for retrieving saved order information from a database, sending it to a server, generating a confirmation message, and displaying it to the user.

[0172] "Means for linking order details to the food delivery service's server and delivery management system" refers to technology that transmits order information to the food delivery service's central server and delivery management system, enabling delivery operations to be carried out efficiently.

[0173] "Means for providing more appropriate services by linking with past order history and voiceprint data" refers to technology that utilizes a user's past order history and voiceprint data to provide personalized services such as recommendations and individual responses.

[0174] This invention is a system that uses voice to efficiently and intuitively order food delivery, and operates in cooperation with three parties: a server, a terminal, and a user.

[0175] User operations

[0176] 1. Select the menu

[0177] Users select products from a menu displayed on their smartphone screen or a printed menu.

[0178] For example, a user wants to order a "Pizza Margherita" and a "Pepsi."

[0179] 2. Voice ordering

[0180] The user orders the selected items by voice into the microphone on their smartphone, for example, saying, "Two Margherita pizzas and one Pepsi, please."

[0181] Terminal handling

[0182] 1. Audio capture

[0183] The microphone on the smartphone captures the user's voice, and this voice data is sent to the server in real time.

[0184] 2. Display Recommendations

[0185] The device analyzes the user's voiceprint and displays menu recommendations based on their past order history and attribute information. For example, if a user has previously ordered a salad, the device might suggest, "Would you like to try today's special salad?"

[0186] Server Processing

[0187] 1. Audio analysis

[0188] The server analyzes the voice data and converts it into text using a generative AI model. For example, a speech request such as "Two Margherita pizzas and one Pepsi, please" is converted into text data such as "Two Margherita pizzas, one Pepsi."

[0189] 2. Confirm your order details

[0190] The server generates a confirmation message for the user and displays it on the user's smartphone, such as "Are you sure you want to order two Margherita pizzas and one Pepsi?"

[0191] 3. Order Data

[0192] The server identifies the order details and saves information such as customer ID, product name, and quantity in a database. This constitutes the order data.

[0193] 4. Order Linkage

[0194] The server then connects this order data to the food delivery service's server and delivery management system, and prepares for delivery. An order for "two Margherita pizzas and one Pepsi" is then sent to restaurants within the delivery area.

[0195] Specific examples

[0196] For example, if a user orders by voice, such as "Two Margherita Pizzas and one Pepsi, please," the voice is captured by the smartphone and sent to the server. The server then converts the voice into text, such as "Two Margherita Pizzas, one Pepsi," and displays a confirmation message to the user. When the user presses the confirmation button, the order information is saved in a database, and this data is linked to the food delivery service's server and delivery management system.

[0197] Example prompts for generative AI models

[0198] A user says to their smartphone, "Two Margherita pizzas and one Pepsi, please." Describe the process for converting speech to text and saving the order to a database.

[0199] The above system will enable users to intuitively order food delivery using voice, which is expected to streamline the ordering process and improve the user experience.

[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0201] Step 1:

[0202] The user selects an item from the menu displayed on the smartphone display, specifically, "Pizza Margherita" and "Pepsi." The input is captured as the user's voice command.

[0203] Step 2:

[0204] The user orders the selected items by voice into the smartphone microphone. Specifically, the user orders "Two pizza margheritas and one Pepsi, please." The input is the user's voice data.

[0205] Step 3:

[0206] The device captures the user's voice using a microphone installed on the smartphone. The voice data is sent to the server in real time. The input is the captured voice data, and the output is the voice data sent to the server.

[0207] Step 4:

[0208] The server analyzes the received voice data and converts it into text using a generative AI model. The input is voice data, and the output is text data in the format "Pizza Margherita 2, Pepsi 1."

[0209] Step 5:

[0210] The server generates a message to reconfirm the order based on the analyzed text data and displays it on the user's smartphone. For example, the message might read, "Is your order two Margherita pizzas and one Pepsi?" The input is text data, and the output is a confirmation message.

[0211] Step 6:

[0212] The user presses the confirmation button or answers "yes" by voice. The device detects this confirmation action and connects to the server. The input is the user's confirmation action, and the output is the confirmation result data.

[0213] Step 7:

[0214] The server saves the order to a database, recording information such as customer ID, product name, and quantity. The input is the confirmed order data, and the output is the order record saved in the database.

[0215] Step 8:

[0216] The server retrieves the stored order data from the database, generates a confirmation message, and communicates with the food delivery service's server and delivery management system to prepare the delivery. The input is the order data retrieved from the database, and the output is the order data sent to the food delivery system.

[0217] Step 9:

[0218] The order details are sent to stores within the delivery area, and preparation begins. For example, an order of "two Margherita pizzas and one Pepsi" is displayed at the store. The input is the order data from the food delivery service's server, and the output is the start of cooking preparation at the store.

[0219] Step 10:

[0220] In order to provide more appropriate services by linking past order history and voiceprint data, the server generates a recommendation menu and displays it to the user the next time they use the service. The input is the past order history and voiceprint data, and the output is the recommendation menu.

[0221] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0222] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[0223] User operations

[0224] 1. Select the menu

[0225] The user selects items from a printed menu or a display screen.

[0226] For example, a user may want to order "carbonara" and "cola."

[0227] 2. Voice ordering

[0228] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0229] Terminal handling

[0230] 1. Audio capture

[0231] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[0232] 2. Sentiment Analysis

[0233] The device sends the captured voice data to an emotion engine, which analyzes the user's emotions, such as "joy," "anger," and "sadness."

[0234] 3. Display Recommendations

[0235] The terminal displays a recommended menu on the screen based on the user's emotions, voiceprint, past order history, and attribute information.

[0236] For example, if the user has the emotion "joy," recommendations such as "recommended desserts" will be displayed.

[0237] Server Processing

[0238] 1. Audio analysis

[0239] The server analyzes the voice data and converts the content into text using generative AI.

[0240] For example, audio saying "One carbonara and one cola please" can be converted into text "1 carbonara, 1 cola."

[0241] 2. Confirm your order details

[0242] The server generates a confirmation message if necessary based on the analysis results of the emotion engine.

[0243] For example, if the user has the emotion "anger," the system will display "Would you like to order one carbonara and one cola?"

[0244] 3. Order Data

[0245] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[0246] 4. Order Linkage

[0247] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[0248] The kitchen displays an order for "one carbonara" and "one coke."

[0249] The cash register then connects the product information to prepare for payment.

[0250] Specific examples

[0251] 1. User Orders

[0252] A user says, "Two pizza margaritas and one Pepsi, please."

[0253] 2. Capturing and transmitting audio from your device

[0254] The microphone captures the user's speech and transmits the data to a server.

[0255] 3. Server voice and emotion analysis

[0256] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi," and the emotion engine recognizes the user's emotion as "satisfied."

[0257] 4. Server Order Confirmation

[0258] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two pizza margaritas and one Pepsi?", increasing satisfaction.

[0259] 5. User Verification

[0260] The user presses the confirmation button or answers "yes" aloud.

[0261] 6. Order data conversion and storage

[0262] The server stores order data and links it to past order history and emotion data.

[0263] 7. Order Linkage

[0264] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[0265] This system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. By incorporating user sentiment analysis, a more personalized customer experience can be achieved.

[0266] The processing flow will be explained below.

[0267] Processing flow

[0268] Step 1:

[0269] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[0270] Step 2:

[0271] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0272] Step 3:

[0273] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[0274] Step 4:

[0275] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[0276] Step 5:

[0277] The server parses the text of the order and generates a message requiring confirmation, such as "Would you like one carbonara and one coke?"

[0278] Step 6:

[0279] The server uses an emotion engine to analyze emotions from the user's voice data and recognizes emotions such as "joy," "anger," and "sadness."

[0280] Step 7:

[0281] The device receives the confirmation message and the emotion analysis results sent from the server and displays them on the display. For example, if the user expresses anger, the device adds the phrase "We apologize for the inconvenience" to the confirmation message.

[0282] Step 8:

[0283] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[0284] Step 9:

[0285] The server receives the user's response and saves the final order details in a database. The customer ID, product name, quantity, and emotion data are recorded in the database.

[0286] Step 10:

[0287] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[0288] Step 11:

[0289] The device will display a recommended menu on the screen. For example, it will display a recommended dessert along with a message such as "Would you like some dessert?" based on the user's voiceprint and past order history.

[0290] Step 12:

[0291] If the user places an additional order, the same process is repeated. If the user places an additional order, the process starts again from step 2.

[0292] Step 13:

[0293] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[0294] These steps enable users to enjoy a more intuitive and personalized ordering experience through emotional confirmation messages and recommendations, while allowing merchants to efficiently manage orders and provide faster service.

[0295] Example 2

[0296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0297] Conventional ordering systems are often difficult to use for users who are unfamiliar with the operation or who have visual or hearing impairments. Furthermore, conventional systems do not take into account the user's emotions, making it difficult to provide services that meet individual needs. In particular, when ordering by voice, it is important to reliably understand the order and appropriately confirm it. For this reason, a system that can accurately recognize the user's voice and respond appropriately based on their emotions is required.

[0298] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice receiving means for accepting voice orders, a voice analysis means for analyzing the accepted voice data using a generative AI model, a display means for confirming and correcting the analyzed order details and the user's emotions, a data storage means for saving the analyzed order details and emotion data in a database, and an order linking means for linking the saved order data to a sales management system. This enables accurate recognition of voice orders and individual responses based on the user's emotions.

[0299] The "voice receiving means" is a device or function for capturing the voice uttered by the user and acquiring it as digital data.

[0300] "Voice analysis means" refers to a device or function that analyzes captured voice data using technologies such as generative AI models and converts the order details into text data.

[0301] The "display means" refers to a device or function that displays the analyzed order details and emotion analysis results to the user, allowing the user to confirm or correct them.

[0302] "Data storage means" refers to a device or function that stores the analyzed order details and emotion data in a database so that it can be used for later processing or analysis.

[0303] An "order linking means" is a device or function that links saved order data to other systems such as a sales management system and shares information with the kitchen, cash register, etc.

[0304] The "recommendation generating means" is a device or function for proposing the most suitable menu to a user based on the user's voiceprint data, past order history, and emotional data.

[0305] The "recommendation presentation means" is a device or function that displays a recommendation menu on a screen or the like based on the user's emotional data and attribute information.

[0306] A "generative AI model" is an artificial intelligence model that analyzes voice data and other input information and generates the required results in the form of text data or other information.

[0307] A "prompt sentence" is an input sentence used to prompt an AI model for a particular output.

[0308] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Below, we will explain in detail an embodiment of the system based on the roles of the server, terminal, and user.

[0309] User operations

[0310] The user selects an item from a printed menu or a display screen, then orders the item by voice into a microphone. For example, the user might say, "One carbonara and one coke, please." This voice command becomes the input for the system.

[0311] Terminal handling

[0312] A microphone built into the device captures the user's voice. This voice data is sent to a server in real time. The device is also equipped with an emotion engine that analyzes the captured voice data to recognize the user's emotions. The results of this emotion analysis are displayed as emotions such as "happiness," "anger," or "sadness." Based on this information, the device displays a recommended menu on the display. For example, if the user expresses the emotion of "happiness," it will display a recommended menu such as "recommended desserts."

[0313] Server Processing

[0314] The server analyzes the received voice data using a generative AI model and converts the order details into text data. For example, a voice saying "One carbonara and one cola, please" is converted into text as "One carbonara, one cola." The server then sends a confirmation message to the user if necessary based on the analysis results of the emotion engine. For example, if the user has the emotion "anger," it generates a message that displays "Is your order one carbonara and one cola?" The server saves the order details and emotion data in a database and links the order data to a sales management system. This link allows the order details to be notified to the kitchen and cash register, enabling prompt service.

[0315] Specific examples

[0316] For example, consider the case where a user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures the user's voice and sends that data to the server. The server uses a generative AI model to convert this voice data into text data: "Two Margherita Pizzas, one Pepsi." At the same time, the emotion engine recognizes the user's emotion as "Satisfied." The server then displays a confirmation message to the user, asking, "Are you sure you want two Margherita Pizzas and one Pepsi?" The order is confirmed when the user presses the confirmation button or answers "Yes" verbally. The server then saves the order data in a database and links it to past order history and emotion data. Finally, the server connects the order data to the store's sales management system, which notifies the kitchen and cashier of the order details.

[0317] Example prompt sentence:

[0318] "One carbonara and one coke please."

[0319] "Two pizza margaritas and one Pepsi, please."

[0320] As described above, this system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. Incorporating sentiment analysis can create a more personalized customer experience.

[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0322] Step 1: User speaks

[0323] Specific operation: The user speaks into the microphone the items they want to order. For example, "One carbonara and one coke, please."

[0324] Input: User's voice

[0325] Output: Audio data

[0326] Step 2: Your device captures audio

[0327] Specific operation: The device's built-in microphone captures the user's voice in real time.

[0328] Input: Audio data

[0329] Output: Digital audio data

[0330] Step 3: The device sends the audio data to the server

[0331] Specific operation: The device transmits the captured digital audio data to the server in real time.

[0332] Input: Digital audio data

[0333] Output: Audio data sent to the server

[0334] Step 4: The server analyzes the audio data

[0335] Specific operation: The server analyzes the voice data received using a generative AI model and converts the order details into text data.

[0336] Input: Audio data

[0337] Output: Text data (e.g. "1 Carbonara, 1 Cola")

[0338] Step 5: The device sends the voice data to the emotion engine

[0339] Specific operation: The device sends the captured voice data to the emotion engine, which analyzes the user's emotions.

[0340] Input: Audio data

[0341] Output: Emotion data (e.g., "joy," "anger," "sadness," etc.)

[0342] Step 6: The device sends the emotion analysis results to the server.

[0343] Specific operation: Emotion data analyzed by the emotion engine is sent to the server.

[0344] Input: Emotion data

[0345] Output: Emotion data sent to the server

[0346] Step 7: The server generates a message to confirm and modify the order details and emotion data.

[0347] Specific operation: The server generates a message based on the text data and emotion data, prompting the user to confirm or correct the message if necessary.

[0348] Input: Text data, emotion data

[0349] Output: Confirmation message (e.g. "Would you like to order one carbonara and one coke?")

[0350] Step 8: Your device will display a confirmation message

[0351] Specific operation: The terminal displays the confirmation message received from the server on the display, prompting the user to confirm.

[0352] Input:Confirmation message

[0353] Output: A confirmation message displayed on the display

[0354] Step 9: User confirms or corrects

[0355] Specific operation: The user checks the confirmation message displayed on the screen and confirms or modifies the order details by voice or button operation.

[0356] Input: User confirmation / correction (voice or button operation)

[0357] Output: Confirmed order details

[0358] Step 10: The server saves the confirmed order details and emotion data to the database.

[0359] Specific operation: The server stores the confirmed order details and emotion data in a database.

[0360] Input: Order details, emotion data

[0361] Output: Order data stored in a database

[0362] Step 11: The server connects the order data to the sales management system

[0363] Specific operation: The server connects the saved order data to the sales management system and notifies the kitchen and cash register.

[0364] Input: Order data stored in a database

[0365] Output: Order details linked to the sales management system and notified to the kitchen and cash register

[0366] Through these steps, the system accurately accepts user voice orders and provides optimal confirmation and recommendations based on emotion. This allows stores to efficiently manage order details and provide prompt service.

[0367] (Application example 2)

[0368] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0369] Conventional voice ordering systems simply convert voice to text data and accept orders. However, this method does not provide personalized recommendations based on user sentiment or past order history, making it difficult to improve the customer experience. Furthermore, order confirmation and correction are often insufficient, leading to ordering errors. This makes it difficult for stores to provide efficient service and can lead to lower customer satisfaction.

[0370] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0371] In this invention, the server includes a generating AI means for analyzing voice data, a display means for confirming and correcting the analyzed order details, a means for saving the analyzed order details in a database, a means for linking the saved order data to an information processing system, a sentiment analysis means for analyzing emotions from the voice data, and a means for displaying recommendation information based on the analyzed emotions. This makes it possible to provide personalized recommendations based on the user's emotions and past order history, streamline confirmation and correction of order details, and significantly improve user satisfaction.

[0372] "Voice ordering" is a method of accepting orders by inputting voice using a microphone and analyzing the voice data.

[0373] The "microphone means" is a device that captures sound as a digital or analog signal and sends it to the analysis means.

[0374] A "generative AI means" is an artificial intelligence system that converts voice data into text data and analyzes it.

[0375] The "display means" is a device that visually presents the analyzed order details and recommendation information to the user.

[0376] A "database" is an information system for storing and managing analyzed data and order data.

[0377] An "information processing system" is a computer system for linking order data with other systems.

[0378] "Emotion analysis means" is a system for analyzing and identifying user emotions from voice data.

[0379] "Recommendation information" is product and service information suggested based on a user's past order history and emotional data.

[0380] The present invention provides an ordering system that combines voice and emotion analysis, and its embodiments are specifically described below. In particular, this system efficiently accepts voice orders in brick-and-mortar stores and provides recommendations based on the user's emotions.

[0381] Server Processing

[0382] The server is responsible for the main data processing and collaboration. First, it analyzes the received voice data and converts it into text data using a generative AI model. This process uses the Google Speech API. Next, it performs sentiment analysis based on the converted text data. In this process, it calculates an emotion score using the TextBlob library. Specifically, it recognizes emotions such as "joy," "anger," and "sadness" from the voice. Once the analysis is complete, it stores the order data and emotion data in a database. Furthermore, the stored data is linked to the information processing system, and the order details are notified to the physical store's POS system, kitchen, and cash register.

[0383] Terminal handling

[0384] The device has the ability to capture voice input from the user and send it to a server in real time. Specifically, a microphone attached to the device captures voice and sends the voice data to the server. The device also receives the results of emotion analysis and displays menu recommendations based on the user's emotions. For example, if the user expresses the emotion "satisfied," the device will display "recommended desserts" on the display.

[0385] User operations

[0386] Users place orders simply and intuitively using voice. For example, they can complete an order by saying, "One pizza margherita and one coke, please." If the voice order is successful, the device displays a message allowing users to confirm or correct the order. If the user is satisfied, the system will suggest menu recommendations and encourage further orders.

[0387] Specific examples

[0388] As a specific example, suppose a user says, "One pizza margherita and one coke, please." This speech is captured by the device's microphone and sent to the server. The server uses a generative AI model to convert the speech into text and recognizes it as "One pizza margherita, one coke." Next, it uses sentiment analysis to identify the user's emotion as "satisfied." Based on this information, the server generates a recommendation for the user suggesting a "recommended dessert" and displays it on the device. The user confirms this and confirms the order by either saying "yes" again or pressing a button. This series of processes improves the user experience and also improves store operational efficiency.

[0389] Example prompt sentence:

[0390] The user orders by voice, "One pizza margherita and one coke please." If the user feels satisfied, the app will recommend a dessert.

[0391] As described above, this invention is a system that combines voice and emotion analysis to provide users with personalized recommendations and efficiently confirm and modify order details.

[0392] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0393] Step 1: Capture audio

[0394] The terminal uses a microphone to capture the user's voice order, such as "One pizza margherita and one coke please." This voice signal becomes the input to the terminal, and the captured voice data is output.

[0395] Step 2: Sending audio data

[0396] The captured audio data is sent by the device to the server in real time, where the input is the captured audio data and the output is the audio data arriving at the server.

[0397] Step 3: Audio analysis

[0398] The server converts the received voice data into text data using the Google Speech API. This generative AI model processes the voice data as input data and outputs the text data "Pizza Margherita 1, Coke 1."

[0399] Step 4: Sentiment analysis

[0400] The server calculates an emotion score using the TextBlob library based on the analyzed text data. The text data is input, and an emotion score is output, for example, as the emotion of "satisfied."

[0401] Step 5: Recommendation generation

[0402] The server generates recommendation information based on the results of the sentiment analysis. The input is the sentiment score, and recommendation information such as "recommended dessert" is generated based on the sentiment of "satisfaction."

[0403] Step 6: Send what you see

[0404] The server sends the generated recommendation information and a confirmation message of the order details to the terminal. The input is the recommendation information and the order confirmation message, and the output is a display instruction to the terminal.

[0405] Step 7: Display to the user

[0406] The device displays the recommendation information and confirmation messages received from the server. For example, the display might say, "Recommended dessert" or "Order details: One pizza margherita and one coke, is that OK?"

[0407] Step 8: User Verification

[0408] The user responds to the displayed confirmation message by saying "Yes" or by pressing a button. This input is sent to the terminal, and the terminal then sends the response data to the server.

[0409] Step 9: Save order data

[0410] The server receives an acknowledgement from the user and stores the order data in a database, with the input being the user response data and the output being the stored order data.

[0411] Step 10: Linking to information processing systems

[0412] The server connects the saved order data to an information processing system (such as a POS system or kitchen display). The input is the saved order data, and the output is a notification of the order details to the information processing system.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0416] [Second embodiment]

[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0429] The present invention is a system for efficiently and intuitively placing orders using voice. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[0430] User operations

[0431] 1. Select the menu

[0432] The user selects items from a printed menu or a display screen.

[0433] For example, suppose a user wants to order "carbonara" and "cola."

[0434] 2. Voice ordering

[0435] The user orders the selected items by voice into the microphone, for example, saying, "One carbonara and one coke, please."

[0436] Terminal handling

[0437] 1. Audio capture

[0438] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[0439] 2. Display Recommendations

[0440] The terminal analyzes the user's voiceprint and displays a recommended menu on the screen based on past order history and attribute information.

[0441] For example, if a user has previously ordered "cheesecake," the system displays "Would you like dessert?"

[0442] Server Processing

[0443] 1. Audio analysis

[0444] The server analyzes the voice data and converts the content into text using generative AI.

[0445] For example, audio saying "One carbonara and one cola please" is converted into text data as "1 carbonara, 1 cola."

[0446] 2. Confirm your order details

[0447] The server may return a confirmation question to the user if necessary, generating a message such as "Would you like one carbonara and one coke?"

[0448] 3. Order Data

[0449] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[0450] 4. Order Linkage

[0451] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[0452] The kitchen displays an order for "one carbonara" and "one coke."

[0453] The cash register then connects the product information to prepare for payment.

[0454] Specific examples

[0455] 1. User Orders

[0456] A user says, "Two pizza margaritas and one Pepsi, please."

[0457] 2. Capturing and transmitting audio from your device

[0458] The microphone captures the user's speech and transmits the data to a server.

[0459] 3. Server audio analysis

[0460] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi."

[0461] 4. Server Order Confirmation

[0462] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two Margherita pizzas and one Pepsi?"

[0463] 5. User Verification

[0464] The user presses the confirmation button or answers "yes" aloud.

[0465] 6. Order data conversion and storage

[0466] The server stores the order data and links it to past order history and voiceprint data.

[0467] 7. Order Linkage

[0468] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[0469] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, it is possible to provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0470] The processing flow will be explained below.

[0471] Step 1:

[0472] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[0473] Step 2:

[0474] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0475] Step 3:

[0476] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[0477] Step 4:

[0478] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[0479] Step 5:

[0480] The server parses the text of the order and generates a confirmation message if necessary, such as "One carbonara and one coke, okay?"

[0481] Step 6:

[0482] The terminal displays a confirmation message on the display, allowing the user to visually confirm the order details.

[0483] Step 7:

[0484] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[0485] Step 8:

[0486] The server receives the user's response and saves the final order details in a database, including the customer ID, product name, and quantity.

[0487] Step 9:

[0488] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[0489] Step 10:

[0490] The device displays a recommended menu on the screen. For example, if a user has previously ordered dessert, the device might recommend, "Would you like a dessert?"

[0491] Step 11:

[0492] If the user wishes to place an additional order, the same process is repeated.

[0493] Step 12:

[0494] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[0495] These steps allow users to easily place orders by voice, and the store can efficiently manage order details and provide service quickly.

[0496] Example 1

[0497] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0498] Conventional ordering systems require users to select products and manually input the exact order details, which is a cumbersome and time-consuming process. Stores also face issues with low order processing efficiency, with orders often being missed or delayed, especially during peak hours. Furthermore, there is a lack of personalized recommendations based on customers' past order history and attribute information, creating a need for improved customer satisfaction.

[0499] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0500] In this invention, the server includes an input means for accepting voice orders, an AI generation means, and an output means, which makes it possible to process orders more efficiently and provide personalized services based on the customer's past order history and attribute information.

[0501] The "input means for accepting voice orders" refers to a data input device for capturing the user's voice. Specifically, it refers to a voice input device such as a microphone.

[0502] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting voice data into text data, including speech recognition models and natural language processing algorithms.

[0503] "Output means for confirming or correcting the analyzed order details" refers to a device that presents the analyzed order details to the user for confirmation or correction. Specifically, this refers to a display, voice response unit, etc.

[0504] The "means for saving the analyzed order details in a database" refers to a storage system for permanently saving the analyzed order details. Specifically, this refers to a relational database, cloud storage, etc.

[0505] "Means for linking stored order data to an information system" refers to communication means for transferring or sharing stored order data with other information systems, such as POS systems. Specifically, this refers to APIs and network interfaces.

[0506] "Means for analyzing customer voiceprint data and generating recommended menus based on past order history and customer attribute information" refers to an algorithm for generating recommended menus using customer voiceprint data, past order history, and attribute information. Specifically, this refers to a machine learning model or recommendation engine.

[0507] The present invention is a system for efficiently and intuitively placing orders using voice. To specifically implement this system, the program processing will be described in detail based on the roles of the server, terminal, and user.

[0508] User operations

[0509] The user selects items from a printed menu or a display screen. For example, if the user wants to order "carbonara" and "Coke," they would say into the microphone, "One carbonara and one Coke, please."

[0510] Terminal handling

[0511] A microphone is connected to the device, which captures the user's voice. This voice data is sent to the server in real time. The device then analyzes the user's voiceprint and displays recommended menu items on the display based on the user's past order history and attribute information. For example, if a user has previously ordered cheesecake, the device might ask, "Would you like dessert?"

[0512] Server Processing

[0513] The server analyzes the voice data and converts it into text using a generative AI model. For example, a user's speech, "One carbonara and one cola, please," is converted into text as "1 carbonara, 1 cola." The server then generates a confirmation message for the user, displaying "Is one carbonara and one cola okay?" Once the user confirms, the server saves the order data in a database. The server also connects the saved order data to the store's information system (POS system) and notifies the kitchen and cash register.

[0514] Specific examples

[0515] The user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures this voice and sends it to the server. The server uses a generative AI model to analyze the voice data and converts it into text: "Two Margherita Pizzas, one Pepsi." It then generates a confirmation message asking, "Are you sure your order is for two Margherita Pizzas and one Pepsi?" and displays it on the device's screen. Once the user confirms, the server saves this order data in a database and links the order data to the store's POS system. As a result, an order for "two Margherita Pizzas" and "one Pepsi" is displayed on the kitchen display, and the product information is sent to the cash register, completing the checkout process.

[0516] In this system, a generative AI model is used to analyze voice data. The model is input with the following prompt: "One carbonara and one coke, please." The model analyzes this prompt and generates the order details as text data.

[0517] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, stores can provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0518] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0519] Step 1:

[0520] The user selects a menu item and orders by voice.

[0521] Input: The user selects a menu item and decides what to order.

[0522] Action: The user speaks into the microphone, "One carbonara and one coke please."

[0523] Output: The microphone captures the user's voice.

[0524] Step 2:

[0525] The device captures the audio and sends it to the server

[0526] Input: User's spoken order.

[0527] How it works: The device's microphone captures audio data and converts it into a digital format.

[0528] Output: The device sends the converted audio data to the server in real time.

[0529] Step 3:

[0530] The server analyzes the voice data and converts it into text

[0531] Input: Audio data sent from the device.

[0532] How it works: The server inputs speech data into the generative AI model. For example, it analyzes speech data such as "One carbonara and one coke, please."

[0533] Data processing: The voice data is input as a prompt sentence into the generative AI model and converted into text data.

[0534] Output: Text data: "1 carbonara, 1 cola".

[0535] Step 4:

[0536] The server verifies the order and asks the user for confirmation.

[0537] Input: Text of the order.

[0538] What happens: The server generates a confirmation message saying "Would you like one carbonara and one coke?"

[0539] Output: A confirmation message will be printed to the terminal screen.

[0540] Step 5:

[0541] The user confirms the order details

[0542] Input: The confirmation message displayed on the device screen.

[0543] Action: The user presses the confirmation button or speaks "yes."

[0544] Output: The user's confirmation action is entered into the terminal.

[0545] Step 6:

[0546] The server saves the order data to a database

[0547] Input: The order details confirmed by the user.

[0548] What happens: The server records a data entry in the database, such as "Customer ID: 12345, Product: Carbonara, Quantity: 1."

[0549] Output: Order data is saved in the database.

[0550] Step 7:

[0551] The server connects the order data to the information system

[0552] Input: Order data stored in the database.

[0553] How it works: The server sends order data to the store's POS system.

[0554] Output: The kitchen display shows the order for "1 Carbonara" and "1 Coke", and the product information is sent to the cashier terminal, ready for payment.

[0555] (Application example 1)

[0556] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0557] In recent years, the use of food delivery services has increased, but there is a demand for improved user convenience and efficient order management. Conventional systems require users to operate an application interface, which is cumbersome, especially for elderly people and users unfamiliar with technology. In addition, the cumbersome process of checking order details and managing order history affects the operational efficiency of food delivery services. To solve these issues, a system that allows users to order intuitively by voice is needed.

[0558] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0559] In this invention, the server includes microphone means for accepting voice orders, AI generation means for analyzing the accepted voice data, display means for confirming and correcting the analyzed order details, means for saving the analyzed order details in a database, means for retrieving the saved order data from the database and linking it to the server, and generating and displaying a confirmation message, means for linking the order details to the food delivery service server and delivery management system, and means for linking the order details with past order history and voiceprint data to provide more appropriate service. This enables users to intuitively place food delivery orders by voice, streamlining the ordering process and improving the user experience.

[0560] "Microphone means for accepting voice orders" refers to devices or technology for detecting and capturing voice.

[0561] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting captured voice data into text and analyzing its meaning.

[0562] "Display means for confirming and correcting the analyzed order contents" refers to a device such as a display or monitor that visually presents the analysis results to the user and allows the user to confirm or correct the contents.

[0563] The "means for storing analyzed order details in a database" refers to a database system for electronically recording and managing analyzed order data.

[0564] "Means for retrieving saved order data from a database, linking to a server, and generating and displaying a confirmation message" refers to a system for retrieving saved order information from a database, sending it to a server, generating a confirmation message, and displaying it to the user.

[0565] "Means for linking order details to the food delivery service's server and delivery management system" refers to technology that transmits order information to the food delivery service's central server and delivery management system, enabling delivery operations to be carried out efficiently.

[0566] "Means for providing more appropriate services by linking with past order history and voiceprint data" refers to technology that utilizes a user's past order history and voiceprint data to provide personalized services such as recommendations and individual responses.

[0567] This invention is a system that uses voice to efficiently and intuitively order food delivery, and operates in cooperation with three parties: a server, a terminal, and a user.

[0568] User operations

[0569] 1. Select the menu

[0570] Users select products from a menu displayed on their smartphone screen or a printed menu.

[0571] For example, a user wants to order a "Pizza Margherita" and a "Pepsi."

[0572] 2. Voice ordering

[0573] The user orders the selected items by voice into the microphone on their smartphone, for example, saying, "Two Margherita pizzas and one Pepsi, please."

[0574] Terminal handling

[0575] 1. Audio capture

[0576] The microphone on the smartphone captures the user's voice, and this voice data is sent to the server in real time.

[0577] 2. Display Recommendations

[0578] The device analyzes the user's voiceprint and displays menu recommendations based on their past order history and attribute information. For example, if a user has previously ordered a salad, the device might suggest, "Would you like to try today's special salad?"

[0579] Server Processing

[0580] 1. Audio analysis

[0581] The server analyzes the voice data and converts it into text using a generative AI model. For example, a speech request such as "Two Margherita pizzas and one Pepsi, please" is converted into text data such as "Two Margherita pizzas, one Pepsi."

[0582] 2. Confirm your order details

[0583] The server generates a confirmation message for the user and displays it on the user's smartphone, such as "Are you sure you want to order two Margherita pizzas and one Pepsi?"

[0584] 3. Order Data

[0585] The server identifies the order details and saves information such as customer ID, product name, and quantity in a database. This constitutes the order data.

[0586] 4. Order Linkage

[0587] The server then connects this order data to the food delivery service's server and delivery management system, and prepares for delivery. An order for "two Margherita pizzas and one Pepsi" is then sent to restaurants within the delivery area.

[0588] Specific examples

[0589] For example, if a user orders by voice, such as "Two Margherita Pizzas and one Pepsi, please," the voice is captured by the smartphone and sent to the server. The server then converts the voice into text, such as "Two Margherita Pizzas, one Pepsi," and displays a confirmation message to the user. When the user presses the confirmation button, the order information is saved in a database, and this data is linked to the food delivery service's server and delivery management system.

[0590] Example prompts for generative AI models

[0591] A user says to their smartphone, "Two Margherita pizzas and one Pepsi, please." Describe the process for converting speech to text and saving the order to a database.

[0592] The above system will enable users to intuitively order food delivery using voice, which is expected to streamline the ordering process and improve the user experience.

[0593] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0594] Step 1:

[0595] The user selects an item from the menu displayed on the smartphone display, specifically, "Pizza Margherita" and "Pepsi." The input is captured as the user's voice command.

[0596] Step 2:

[0597] The user orders the selected items by voice into the smartphone microphone. Specifically, the user orders "Two pizza margheritas and one Pepsi, please." The input is the user's voice data.

[0598] Step 3:

[0599] The device captures the user's voice using a microphone installed on the smartphone. The voice data is sent to the server in real time. The input is the captured voice data, and the output is the voice data sent to the server.

[0600] Step 4:

[0601] The server analyzes the received voice data and converts it into text using a generative AI model. The input is voice data, and the output is text data in the format "Pizza Margherita 2, Pepsi 1."

[0602] Step 5:

[0603] The server generates a message to reconfirm the order based on the analyzed text data and displays it on the user's smartphone. For example, the message might read, "Is your order two Margherita pizzas and one Pepsi?" The input is text data, and the output is a confirmation message.

[0604] Step 6:

[0605] The user presses the confirmation button or answers "yes" by voice. The device detects this confirmation action and connects to the server. The input is the user's confirmation action, and the output is the confirmation result data.

[0606] Step 7:

[0607] The server saves the order to a database, recording information such as customer ID, product name, and quantity. The input is the confirmed order data, and the output is the order record saved in the database.

[0608] Step 8:

[0609] The server retrieves the stored order data from the database, generates a confirmation message, and communicates with the food delivery service's server and delivery management system to prepare the delivery. The input is the order data retrieved from the database, and the output is the order data sent to the food delivery system.

[0610] Step 9:

[0611] The order details are sent to stores within the delivery area, and preparation begins. For example, an order of "two Margherita pizzas and one Pepsi" is displayed at the store. The input is the order data from the food delivery service's server, and the output is the start of cooking preparation at the store.

[0612] Step 10:

[0613] In order to provide more appropriate services by linking past order history and voiceprint data, the server generates a recommendation menu and displays it to the user the next time they use the service. The input is the past order history and voiceprint data, and the output is the recommendation menu.

[0614] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0615] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[0616] User operations

[0617] 1. Select the menu

[0618] The user selects items from a printed menu or a display screen.

[0619] For example, a user may want to order "carbonara" and "cola."

[0620] 2. Voice ordering

[0621] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0622] Terminal handling

[0623] 1. Audio capture

[0624] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[0625] 2. Sentiment Analysis

[0626] The device sends the captured voice data to an emotion engine, which analyzes the user's emotions, such as "joy," "anger," and "sadness."

[0627] 3. Display Recommendations

[0628] The terminal displays a recommended menu on the screen based on the user's emotions, voiceprint, past order history, and attribute information.

[0629] For example, if the user has the emotion "joy," recommendations such as "recommended desserts" will be displayed.

[0630] Server Processing

[0631] 1. Audio analysis

[0632] The server analyzes the voice data and converts the content into text using generative AI.

[0633] For example, audio saying "One carbonara and one cola please" can be converted into text "1 carbonara, 1 cola."

[0634] 2. Confirm your order details

[0635] The server generates a confirmation message if necessary based on the analysis results of the emotion engine.

[0636] For example, if the user has the emotion "anger," the system will display "Would you like to order one carbonara and one cola?"

[0637] 3. Order Data

[0638] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[0639] 4. Order Linkage

[0640] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[0641] The kitchen displays an order for "one carbonara" and "one coke."

[0642] The cash register then connects the product information to prepare for payment.

[0643] Specific examples

[0644] 1. User Orders

[0645] A user says, "Two pizza margaritas and one Pepsi, please."

[0646] 2. Capturing and transmitting audio from your device

[0647] The microphone captures the user's speech and transmits the data to a server.

[0648] 3. Server voice and emotion analysis

[0649] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi," and the emotion engine recognizes the user's emotion as "satisfied."

[0650] 4. Server Order Confirmation

[0651] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two pizza margaritas and one Pepsi?", increasing satisfaction.

[0652] 5. User Verification

[0653] The user presses the confirmation button or answers "yes" aloud.

[0654] 6. Order data conversion and storage

[0655] The server stores order data and links it to past order history and emotion data.

[0656] 7. Order Linkage

[0657] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[0658] This system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. By incorporating user sentiment analysis, a more personalized customer experience can be achieved.

[0659] The processing flow will be explained below.

[0660] Processing flow

[0661] Step 1:

[0662] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[0663] Step 2:

[0664] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0665] Step 3:

[0666] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[0667] Step 4:

[0668] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[0669] Step 5:

[0670] The server parses the text of the order and generates a message requiring confirmation, such as "Would you like one carbonara and one coke?"

[0671] Step 6:

[0672] The server uses an emotion engine to analyze emotions from the user's voice data and recognizes emotions such as "joy," "anger," and "sadness."

[0673] Step 7:

[0674] The device receives the confirmation message and the emotion analysis results sent from the server and displays them on the display. For example, if the user expresses anger, the device adds the phrase "We apologize for the inconvenience" to the confirmation message.

[0675] Step 8:

[0676] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[0677] Step 9:

[0678] The server receives the user's response and saves the final order details in a database. The customer ID, product name, quantity, and emotion data are recorded in the database.

[0679] Step 10:

[0680] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[0681] Step 11:

[0682] The device will display a recommended menu on the screen. For example, it will display a recommended dessert along with a message such as "Would you like some dessert?" based on the user's voiceprint and past order history.

[0683] Step 12:

[0684] If the user places an additional order, the same process is repeated. If the user places an additional order, the process starts again from step 2.

[0685] Step 13:

[0686] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[0687] These steps enable users to enjoy a more intuitive and personalized ordering experience through emotional confirmation messages and recommendations, while allowing merchants to efficiently manage orders and provide faster service.

[0688] Example 2

[0689] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0690] Conventional ordering systems are often difficult to use for users who are unfamiliar with the operation or who have visual or hearing impairments. Furthermore, conventional systems do not take into account the user's emotions, making it difficult to provide services that meet individual needs. In particular, when ordering by voice, it is important to reliably understand the order and appropriately confirm it. For this reason, a system that can accurately recognize the user's voice and respond appropriately based on their emotions is required.

[0691] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice receiving means for accepting voice orders, a voice analysis means for analyzing the accepted voice data using a generative AI model, a display means for confirming and correcting the analyzed order details and the user's emotions, a data storage means for saving the analyzed order details and emotion data in a database, and an order linking means for linking the saved order data to a sales management system. This enables accurate recognition of voice orders and individual responses based on the user's emotions.

[0692] The "voice receiving means" is a device or function for capturing the voice uttered by the user and acquiring it as digital data.

[0693] "Voice analysis means" refers to a device or function that analyzes captured voice data using technologies such as generative AI models and converts the order details into text data.

[0694] The "display means" refers to a device or function that displays the analyzed order details and emotion analysis results to the user, allowing the user to confirm or correct them.

[0695] "Data storage means" refers to a device or function that stores the analyzed order details and emotion data in a database so that it can be used for later processing or analysis.

[0696] An "order linking means" is a device or function that links saved order data to other systems such as a sales management system and shares information with the kitchen, cash register, etc.

[0697] The "recommendation generating means" is a device or function for proposing the most suitable menu to a user based on the user's voiceprint data, past order history, and emotional data.

[0698] The "recommendation presentation means" is a device or function that displays a recommendation menu on a screen or the like based on the user's emotional data and attribute information.

[0699] A "generative AI model" is an artificial intelligence model that analyzes voice data and other input information and generates the required results in the form of text data or other information.

[0700] A "prompt sentence" is an input sentence used to prompt an AI model for a particular output.

[0701] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Below, we will explain in detail an embodiment of the system based on the roles of the server, terminal, and user.

[0702] User operations

[0703] The user selects an item from a printed menu or a display screen, then orders the item by voice into a microphone. For example, the user might say, "One carbonara and one coke, please." This voice command becomes the input for the system.

[0704] Terminal handling

[0705] A microphone built into the device captures the user's voice. This voice data is sent to a server in real time. The device is also equipped with an emotion engine that analyzes the captured voice data to recognize the user's emotions. The results of this emotion analysis are displayed as emotions such as "happiness," "anger," or "sadness." Based on this information, the device displays a recommended menu on the display. For example, if the user expresses the emotion of "happiness," it will display a recommended menu such as "recommended desserts."

[0706] Server Processing

[0707] The server analyzes the received voice data using a generative AI model and converts the order details into text data. For example, a voice saying "One carbonara and one cola, please" is converted into text as "One carbonara, one cola." The server then sends a confirmation message to the user if necessary based on the analysis results of the emotion engine. For example, if the user has the emotion "anger," it generates a message that displays "Is your order one carbonara and one cola?" The server saves the order details and emotion data in a database and links the order data to a sales management system. This link allows the order details to be notified to the kitchen and cash register, enabling prompt service.

[0708] Specific examples

[0709] For example, consider the case where a user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures the user's voice and sends that data to the server. The server uses a generative AI model to convert this voice data into text data: "Two Margherita Pizzas, one Pepsi." At the same time, the emotion engine recognizes the user's emotion as "Satisfied." The server then displays a confirmation message to the user, asking, "Are you sure you want two Margherita Pizzas and one Pepsi?" The order is confirmed when the user presses the confirmation button or answers "Yes" verbally. The server then saves the order data in a database and links it to past order history and emotion data. Finally, the server connects the order data to the store's sales management system, which notifies the kitchen and cashier of the order details.

[0710] Example prompt sentence:

[0711] "One carbonara and one coke please."

[0712] "Two pizza margaritas and one Pepsi, please."

[0713] As described above, this system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. Incorporating sentiment analysis can create a more personalized customer experience.

[0714] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0715] Step 1: User speaks

[0716] Specific operation: The user speaks into the microphone the items they want to order. For example, "One carbonara and one coke, please."

[0717] Input: User's voice

[0718] Output: Audio data

[0719] Step 2: Your device captures audio

[0720] Specific operation: The device's built-in microphone captures the user's voice in real time.

[0721] Input: Audio data

[0722] Output: Digital audio data

[0723] Step 3: The device sends the audio data to the server

[0724] Specific operation: The device transmits the captured digital audio data to the server in real time.

[0725] Input: Digital audio data

[0726] Output: Audio data sent to the server

[0727] Step 4: The server analyzes the audio data

[0728] Specific operation: The server analyzes the voice data received using a generative AI model and converts the order details into text data.

[0729] Input: Audio data

[0730] Output: Text data (e.g. "1 Carbonara, 1 Cola")

[0731] Step 5: The device sends the voice data to the emotion engine

[0732] Specific operation: The device sends the captured voice data to the emotion engine, which analyzes the user's emotions.

[0733] Input: Audio data

[0734] Output: Emotion data (e.g., "joy," "anger," "sadness," etc.)

[0735] Step 6: The device sends the emotion analysis results to the server.

[0736] Specific operation: Emotion data analyzed by the emotion engine is sent to the server.

[0737] Input: Emotion data

[0738] Output: Emotion data sent to the server

[0739] Step 7: The server generates a message to confirm and modify the order details and emotion data.

[0740] Specific operation: The server generates a message based on the text data and emotion data, prompting the user to confirm or correct the message if necessary.

[0741] Input: Text data, emotion data

[0742] Output: Confirmation message (e.g. "Would you like to order one carbonara and one coke?")

[0743] Step 8: Your device will display a confirmation message

[0744] Specific operation: The terminal displays the confirmation message received from the server on the display, prompting the user to confirm.

[0745] Input:Confirmation message

[0746] Output: A confirmation message displayed on the display

[0747] Step 9: User confirms or corrects

[0748] Specific operation: The user checks the confirmation message displayed on the screen and confirms or modifies the order details by voice or button operation.

[0749] Input: User confirmation / correction (voice or button operation)

[0750] Output: Confirmed order details

[0751] Step 10: The server saves the confirmed order details and emotion data to the database.

[0752] Specific operation: The server stores the confirmed order details and emotion data in a database.

[0753] Input: Order details, emotion data

[0754] Output: Order data stored in a database

[0755] Step 11: The server connects the order data to the sales management system

[0756] Specific operation: The server connects the saved order data to the sales management system and notifies the kitchen and cash register.

[0757] Input: Order data stored in a database

[0758] Output: Order details linked to the sales management system and notified to the kitchen and cash register

[0759] Through these steps, the system accurately accepts user voice orders and provides optimal confirmation and recommendations based on emotion. This allows stores to efficiently manage order details and provide prompt service.

[0760] (Application example 2)

[0761] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0762] Conventional voice ordering systems simply convert voice to text data and accept orders. However, this method does not provide personalized recommendations based on user sentiment or past order history, making it difficult to improve the customer experience. Furthermore, order confirmation and correction are often insufficient, leading to ordering errors. This makes it difficult for stores to provide efficient service and can lead to lower customer satisfaction.

[0763] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0764] In this invention, the server includes a generating AI means for analyzing voice data, a display means for confirming and correcting the analyzed order details, a means for saving the analyzed order details in a database, a means for linking the saved order data to an information processing system, a sentiment analysis means for analyzing emotions from the voice data, and a means for displaying recommendation information based on the analyzed emotions. This makes it possible to provide personalized recommendations based on the user's emotions and past order history, streamline confirmation and correction of order details, and significantly improve user satisfaction.

[0765] "Voice ordering" is a method of accepting orders by inputting voice using a microphone and analyzing the voice data.

[0766] The "microphone means" is a device that captures sound as a digital or analog signal and sends it to the analysis means.

[0767] A "generative AI means" is an artificial intelligence system that converts voice data into text data and analyzes it.

[0768] The "display means" is a device that visually presents the analyzed order details and recommendation information to the user.

[0769] A "database" is an information system for storing and managing analyzed data and order data.

[0770] An "information processing system" is a computer system for linking order data with other systems.

[0771] "Emotion analysis means" is a system for analyzing and identifying user emotions from voice data.

[0772] "Recommendation information" is product and service information suggested based on a user's past order history and emotional data.

[0773] The present invention provides an ordering system that combines voice and emotion analysis, and its embodiments are specifically described below. In particular, this system efficiently accepts voice orders in brick-and-mortar stores and provides recommendations based on the user's emotions.

[0774] Server Processing

[0775] The server is responsible for the main data processing and collaboration. First, it analyzes the received voice data and converts it into text data using a generative AI model. This process uses the Google Speech API. Next, it performs sentiment analysis based on the converted text data. In this process, it calculates an emotion score using the TextBlob library. Specifically, it recognizes emotions such as "joy," "anger," and "sadness" from the voice. Once the analysis is complete, it stores the order data and emotion data in a database. Furthermore, the stored data is linked to the information processing system, and the order details are notified to the physical store's POS system, kitchen, and cash register.

[0776] Terminal handling

[0777] The device has the ability to capture voice input from the user and send it to a server in real time. Specifically, a microphone attached to the device captures voice and sends the voice data to the server. The device also receives the results of emotion analysis and displays menu recommendations based on the user's emotions. For example, if the user expresses the emotion "satisfied," the device will display "recommended desserts" on the display.

[0778] User operations

[0779] Users place orders simply and intuitively using voice. For example, they can complete an order by saying, "One pizza margherita and one coke, please." If the voice order is successful, the device displays a message allowing users to confirm or correct the order. If the user is satisfied, the system will suggest menu recommendations and encourage further orders.

[0780] Specific examples

[0781] As a specific example, suppose a user says, "One pizza margherita and one coke, please." This speech is captured by the device's microphone and sent to the server. The server uses a generative AI model to convert the speech into text and recognizes it as "One pizza margherita, one coke." Next, it uses sentiment analysis to identify the user's emotion as "satisfied." Based on this information, the server generates a recommendation for the user suggesting a "recommended dessert" and displays it on the device. The user confirms this and confirms the order by either saying "yes" again or pressing a button. This series of processes improves the user experience and also improves store operational efficiency.

[0782] Example prompt sentence:

[0783] The user orders by voice, "One pizza margherita and one coke please." If the user feels satisfied, the app will recommend a dessert.

[0784] As described above, this invention is a system that combines voice and emotion analysis to provide users with personalized recommendations and efficiently confirm and modify order details.

[0785] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0786] Step 1: Capture audio

[0787] The terminal uses a microphone to capture the user's voice order, such as "One pizza margherita and one coke please." This voice signal becomes the input to the terminal, and the captured voice data is output.

[0788] Step 2: Sending audio data

[0789] The captured audio data is sent by the device to the server in real time, where the input is the captured audio data and the output is the audio data arriving at the server.

[0790] Step 3: Audio analysis

[0791] The server converts the received voice data into text data using the Google Speech API. This generative AI model processes the voice data as input data and outputs the text data "Pizza Margherita 1, Coke 1."

[0792] Step 4: Sentiment analysis

[0793] The server calculates an emotion score using the TextBlob library based on the analyzed text data. The text data is input, and an emotion score is output, for example, as the emotion of "satisfied."

[0794] Step 5: Recommendation generation

[0795] The server generates recommendation information based on the results of the sentiment analysis. The input is the sentiment score, and recommendation information such as "recommended dessert" is generated based on the sentiment of "satisfaction."

[0796] Step 6: Send what you see

[0797] The server sends the generated recommendation information and a confirmation message of the order details to the terminal. The input is the recommendation information and the order confirmation message, and the output is a display instruction to the terminal.

[0798] Step 7: Display to the user

[0799] The device displays the recommendation information and confirmation messages received from the server. For example, the display might say, "Recommended dessert" or "Order details: One pizza margherita and one coke, is that OK?"

[0800] Step 8: User Verification

[0801] The user responds to the displayed confirmation message by saying "Yes" or by pressing a button. This input is sent to the terminal, and the terminal then sends the response data to the server.

[0802] Step 9: Save order data

[0803] The server receives an acknowledgement from the user and stores the order data in a database, with the input being the user response data and the output being the stored order data.

[0804] Step 10: Linking to information processing systems

[0805] The server connects the saved order data to an information processing system (such as a POS system or kitchen display). The input is the saved order data, and the output is a notification of the order details to the information processing system.

[0806] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0807] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0809] [Third embodiment]

[0810] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0811] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0812] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0813] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0814] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0815] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0816] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0817] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0818] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0819] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0820] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0821] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0822] The present invention is a system for efficiently and intuitively placing orders using voice. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[0823] User operations

[0824] 1. Select the menu

[0825] The user selects items from a printed menu or a display screen.

[0826] For example, suppose a user wants to order "carbonara" and "cola."

[0827] 2. Voice ordering

[0828] The user orders the selected items by voice into the microphone, for example, saying, "One carbonara and one coke, please."

[0829] Terminal handling

[0830] 1. Audio capture

[0831] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[0832] 2. Display Recommendations

[0833] The terminal analyzes the user's voiceprint and displays a recommended menu on the screen based on past order history and attribute information.

[0834] For example, if a user has previously ordered "cheesecake," the system displays "Would you like dessert?"

[0835] Server Processing

[0836] 1. Audio analysis

[0837] The server analyzes the voice data and converts the content into text using generative AI.

[0838] For example, audio saying "One carbonara and one cola please" is converted into text data as "1 carbonara, 1 cola."

[0839] 2. Confirm your order details

[0840] The server may return a confirmation question to the user if necessary, generating a message such as "Would you like one carbonara and one coke?"

[0841] 3. Order Data

[0842] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[0843] 4. Order Linkage

[0844] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[0845] The kitchen displays an order for "one carbonara" and "one coke."

[0846] The cash register then connects the product information to prepare for payment.

[0847] Specific examples

[0848] 1. User Orders

[0849] A user says, "Two pizza margaritas and one Pepsi, please."

[0850] 2. Capturing and transmitting audio from your device

[0851] The microphone captures the user's speech and transmits the data to a server.

[0852] 3. Server audio analysis

[0853] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi."

[0854] 4. Server Order Confirmation

[0855] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two Margherita pizzas and one Pepsi?"

[0856] 5. User Verification

[0857] The user presses the confirmation button or answers "yes" aloud.

[0858] 6. Order data conversion and storage

[0859] The server stores the order data and links it to past order history and voiceprint data.

[0860] 7. Order Linkage

[0861] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[0862] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, it is possible to provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0863] The processing flow will be explained below.

[0864] Step 1:

[0865] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[0866] Step 2:

[0867] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[0868] Step 3:

[0869] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[0870] Step 4:

[0871] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[0872] Step 5:

[0873] The server parses the text of the order and generates a confirmation message if necessary, such as "One carbonara and one coke, okay?"

[0874] Step 6:

[0875] The terminal displays a confirmation message on the display, allowing the user to visually confirm the order details.

[0876] Step 7:

[0877] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[0878] Step 8:

[0879] The server receives the user's response and saves the final order details in a database, including the customer ID, product name, and quantity.

[0880] Step 9:

[0881] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[0882] Step 10:

[0883] The device displays a recommended menu on the screen. For example, if a user has previously ordered dessert, the device might recommend, "Would you like a dessert?"

[0884] Step 11:

[0885] If the user wishes to place an additional order, the same process is repeated.

[0886] Step 12:

[0887] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[0888] These steps allow users to easily place orders by voice, and the store can efficiently manage order details and provide service quickly.

[0889] Example 1

[0890] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0891] Conventional ordering systems require users to select products and manually input the exact order details, which is a cumbersome and time-consuming process. Stores also face issues with low order processing efficiency, with orders often being missed or delayed, especially during peak hours. Furthermore, there is a lack of personalized recommendations based on customers' past order history and attribute information, creating a need for improved customer satisfaction.

[0892] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0893] In this invention, the server includes an input means for accepting voice orders, an AI generation means, and an output means, which makes it possible to process orders more efficiently and provide personalized services based on the customer's past order history and attribute information.

[0894] The "input means for accepting voice orders" refers to a data input device for capturing the user's voice. Specifically, it refers to a voice input device such as a microphone.

[0895] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting voice data into text data, including speech recognition models and natural language processing algorithms.

[0896] "Output means for confirming or correcting the analyzed order details" refers to a device that presents the analyzed order details to the user for confirmation or correction. Specifically, this refers to a display, voice response unit, etc.

[0897] The "means for saving the analyzed order details in a database" refers to a storage system for permanently saving the analyzed order details. Specifically, this refers to a relational database, cloud storage, etc.

[0898] "Means for linking stored order data to an information system" refers to communication means for transferring or sharing stored order data with other information systems, such as POS systems. Specifically, this refers to APIs and network interfaces.

[0899] "Means for analyzing customer voiceprint data and generating recommended menus based on past order history and customer attribute information" refers to an algorithm for generating recommended menus using customer voiceprint data, past order history, and attribute information. Specifically, this refers to a machine learning model or recommendation engine.

[0900] The present invention is a system for efficiently and intuitively placing orders using voice. To specifically implement this system, the program processing will be described in detail based on the roles of the server, terminal, and user.

[0901] User operations

[0902] The user selects items from a printed menu or a display screen. For example, if the user wants to order "carbonara" and "Coke," they would say into the microphone, "One carbonara and one Coke, please."

[0903] Terminal handling

[0904] A microphone is connected to the device, which captures the user's voice. This voice data is sent to the server in real time. The device then analyzes the user's voiceprint and displays recommended menu items on the display based on the user's past order history and attribute information. For example, if a user has previously ordered cheesecake, the device might ask, "Would you like dessert?"

[0905] Server Processing

[0906] The server analyzes the voice data and converts it into text using a generative AI model. For example, a user's speech, "One carbonara and one cola, please," is converted into text as "1 carbonara, 1 cola." The server then generates a confirmation message for the user, displaying "Is one carbonara and one cola okay?" Once the user confirms, the server saves the order data in a database. The server also connects the saved order data to the store's information system (POS system) and notifies the kitchen and cash register.

[0907] Specific examples

[0908] The user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures this voice and sends it to the server. The server uses a generative AI model to analyze the voice data and converts it into text: "Two Margherita Pizzas, one Pepsi." It then generates a confirmation message asking, "Are you sure your order is for two Margherita Pizzas and one Pepsi?" and displays it on the device's screen. Once the user confirms, the server saves this order data in a database and links the order data to the store's POS system. As a result, an order for "two Margherita Pizzas" and "one Pepsi" is displayed on the kitchen display, and the product information is sent to the cash register, completing the checkout process.

[0909] In this system, a generative AI model is used to analyze voice data. The model is input with the following prompt: "One carbonara and one coke, please." The model analyzes this prompt and generates the order details as text data.

[0910] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, stores can provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[0911] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0912] Step 1:

[0913] The user selects a menu item and orders by voice.

[0914] Input: The user selects a menu item and decides what to order.

[0915] Action: The user speaks into the microphone, "One carbonara and one coke please."

[0916] Output: The microphone captures the user's voice.

[0917] Step 2:

[0918] The device captures the audio and sends it to the server

[0919] Input: User's spoken order.

[0920] How it works: The device's microphone captures audio data and converts it into a digital format.

[0921] Output: The device sends the converted audio data to the server in real time.

[0922] Step 3:

[0923] The server analyzes the voice data and converts it into text

[0924] Input: Audio data sent from the device.

[0925] How it works: The server inputs speech data into the generative AI model. For example, it analyzes speech data such as "One carbonara and one coke, please."

[0926] Data processing: The voice data is input as a prompt sentence into the generative AI model and converted into text data.

[0927] Output: Text data: "1 carbonara, 1 cola".

[0928] Step 4:

[0929] The server verifies the order and asks the user for confirmation.

[0930] Input: Text of the order.

[0931] What happens: The server generates a confirmation message saying "Would you like one carbonara and one coke?"

[0932] Output: A confirmation message will be printed to the terminal screen.

[0933] Step 5:

[0934] The user confirms the order details

[0935] Input: The confirmation message displayed on the device screen.

[0936] Action: The user presses the confirmation button or speaks "yes."

[0937] Output: The user's confirmation action is entered into the terminal.

[0938] Step 6:

[0939] The server saves the order data to a database

[0940] Input: The order details confirmed by the user.

[0941] What happens: The server records a data entry in the database, such as "Customer ID: 12345, Product: Carbonara, Quantity: 1."

[0942] Output: Order data is saved in the database.

[0943] Step 7:

[0944] The server connects the order data to the information system

[0945] Input: Order data stored in the database.

[0946] How it works: The server sends order data to the store's POS system.

[0947] Output: The kitchen display shows the order for "1 Carbonara" and "1 Coke", and the product information is sent to the cashier terminal, ready for payment.

[0948] (Application example 1)

[0949] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0950] In recent years, the use of food delivery services has increased, but there is a demand for improved user convenience and efficient order management. Conventional systems require users to operate an application interface, which is cumbersome, especially for elderly people and users unfamiliar with technology. In addition, the cumbersome process of checking order details and managing order history affects the operational efficiency of food delivery services. To solve these issues, a system that allows users to order intuitively by voice is needed.

[0951] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0952] In this invention, the server includes microphone means for accepting voice orders, AI generation means for analyzing the accepted voice data, display means for confirming and correcting the analyzed order details, means for saving the analyzed order details in a database, means for retrieving the saved order data from the database and linking it to the server, and generating and displaying a confirmation message, means for linking the order details to the food delivery service server and delivery management system, and means for linking the order details with past order history and voiceprint data to provide more appropriate service. This enables users to intuitively place food delivery orders by voice, streamlining the ordering process and improving the user experience.

[0953] "Microphone means for accepting voice orders" refers to devices or technology for detecting and capturing voice.

[0954] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting captured voice data into text and analyzing its meaning.

[0955] "Display means for confirming and correcting the analyzed order contents" refers to a device such as a display or monitor that visually presents the analysis results to the user and allows the user to confirm or correct the contents.

[0956] The "means for storing analyzed order details in a database" refers to a database system for electronically recording and managing analyzed order data.

[0957] "Means for retrieving saved order data from a database, linking to a server, and generating and displaying a confirmation message" refers to a system for retrieving saved order information from a database, sending it to a server, generating a confirmation message, and displaying it to the user.

[0958] "Means for linking order details to the food delivery service's server and delivery management system" refers to technology that transmits order information to the food delivery service's central server and delivery management system, enabling delivery operations to be carried out efficiently.

[0959] "Means for providing more appropriate services by linking with past order history and voiceprint data" refers to technology that utilizes a user's past order history and voiceprint data to provide personalized services such as recommendations and individual responses.

[0960] This invention is a system that uses voice to efficiently and intuitively order food delivery, and operates in cooperation with three parties: a server, a terminal, and a user.

[0961] User operations

[0962] 1. Select the menu

[0963] Users select products from a menu displayed on their smartphone screen or a printed menu.

[0964] For example, a user wants to order a "Pizza Margherita" and a "Pepsi."

[0965] 2. Voice ordering

[0966] The user orders the selected items by voice into the microphone on their smartphone, for example, saying, "Two Margherita pizzas and one Pepsi, please."

[0967] Terminal handling

[0968] 1. Audio capture

[0969] The microphone on the smartphone captures the user's voice, and this voice data is sent to the server in real time.

[0970] 2. Display Recommendations

[0971] The device analyzes the user's voiceprint and displays menu recommendations based on their past order history and attribute information. For example, if a user has previously ordered a salad, the device might suggest, "Would you like to try today's special salad?"

[0972] Server Processing

[0973] 1. Audio analysis

[0974] The server analyzes the voice data and converts it into text using a generative AI model. For example, a speech request such as "Two Margherita pizzas and one Pepsi, please" is converted into text data such as "Two Margherita pizzas, one Pepsi."

[0975] 2. Confirm your order details

[0976] The server generates a confirmation message for the user and displays it on the user's smartphone, such as "Are you sure you want to order two Margherita pizzas and one Pepsi?"

[0977] 3. Order Data

[0978] The server identifies the order details and saves information such as customer ID, product name, and quantity in a database. This constitutes the order data.

[0979] 4. Order Linkage

[0980] The server then connects this order data to the food delivery service's server and delivery management system, and prepares for delivery. An order for "two Margherita pizzas and one Pepsi" is then sent to restaurants within the delivery area.

[0981] Specific examples

[0982] For example, if a user orders by voice, such as "Two Margherita Pizzas and one Pepsi, please," the voice is captured by the smartphone and sent to the server. The server then converts the voice into text, such as "Two Margherita Pizzas, one Pepsi," and displays a confirmation message to the user. When the user presses the confirmation button, the order information is saved in a database, and this data is linked to the food delivery service's server and delivery management system.

[0983] Example prompts for generative AI models

[0984] A user says to their smartphone, "Two Margherita pizzas and one Pepsi, please." Describe the process for converting speech to text and saving the order to a database.

[0985] The above system will enable users to intuitively order food delivery using voice, which is expected to streamline the ordering process and improve the user experience.

[0986] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0987] Step 1:

[0988] The user selects an item from the menu displayed on the smartphone display, specifically, "Pizza Margherita" and "Pepsi." The input is captured as the user's voice command.

[0989] Step 2:

[0990] The user orders the selected items by voice into the smartphone microphone. Specifically, the user orders "Two pizza margheritas and one Pepsi, please." The input is the user's voice data.

[0991] Step 3:

[0992] The device captures the user's voice using a microphone installed on the smartphone. The voice data is sent to the server in real time. The input is the captured voice data, and the output is the voice data sent to the server.

[0993] Step 4:

[0994] The server analyzes the received voice data and converts it into text using a generative AI model. The input is voice data, and the output is text data in the format "Pizza Margherita 2, Pepsi 1."

[0995] Step 5:

[0996] The server generates a message to reconfirm the order based on the analyzed text data and displays it on the user's smartphone. For example, the message might read, "Is your order two Margherita pizzas and one Pepsi?" The input is text data, and the output is a confirmation message.

[0997] Step 6:

[0998] The user presses the confirmation button or answers "yes" by voice. The device detects this confirmation action and connects to the server. The input is the user's confirmation action, and the output is the confirmation result data.

[0999] Step 7:

[1000] The server saves the order to a database, recording information such as customer ID, product name, and quantity. The input is the confirmed order data, and the output is the order record saved in the database.

[1001] Step 8:

[1002] The server retrieves the stored order data from the database, generates a confirmation message, and communicates with the food delivery service's server and delivery management system to prepare the delivery. The input is the order data retrieved from the database, and the output is the order data sent to the food delivery system.

[1003] Step 9:

[1004] The order details are sent to stores within the delivery area, and preparation begins. For example, an order of "two Margherita pizzas and one Pepsi" is displayed at the store. The input is the order data from the food delivery service's server, and the output is the start of cooking preparation at the store.

[1005] Step 10:

[1006] In order to provide more appropriate services by linking past order history and voiceprint data, the server generates a recommendation menu and displays it to the user the next time they use the service. The input is the past order history and voiceprint data, and the output is the recommendation menu.

[1007] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1008] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[1009] User operations

[1010] 1. Select the menu

[1011] The user selects items from a printed menu or a display screen.

[1012] For example, a user may want to order "carbonara" and "cola."

[1013] 2. Voice ordering

[1014] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[1015] Terminal handling

[1016] 1. Audio capture

[1017] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[1018] 2. Sentiment Analysis

[1019] The device sends the captured voice data to an emotion engine, which analyzes the user's emotions, such as "joy," "anger," and "sadness."

[1020] 3. Display Recommendations

[1021] The terminal displays a recommended menu on the screen based on the user's emotions, voiceprint, past order history, and attribute information.

[1022] For example, if the user has the emotion "joy," recommendations such as "recommended desserts" will be displayed.

[1023] Server Processing

[1024] 1. Audio analysis

[1025] The server analyzes the voice data and converts the content into text using generative AI.

[1026] For example, audio saying "One carbonara and one cola please" can be converted into text "1 carbonara, 1 cola."

[1027] 2. Confirm your order details

[1028] The server generates a confirmation message if necessary based on the analysis results of the emotion engine.

[1029] For example, if the user has the emotion "anger," the system will display "Would you like to order one carbonara and one cola?"

[1030] 3. Order Data

[1031] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[1032] 4. Order Linkage

[1033] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[1034] The kitchen displays an order for "one carbonara" and "one coke."

[1035] The cash register then connects the product information to prepare for payment.

[1036] Specific examples

[1037] 1. User Orders

[1038] A user says, "Two pizza margaritas and one Pepsi, please."

[1039] 2. Capturing and transmitting audio from your device

[1040] The microphone captures the user's speech and transmits the data to a server.

[1041] 3. Server voice and emotion analysis

[1042] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi," and the emotion engine recognizes the user's emotion as "satisfied."

[1043] 4. Server Order Confirmation

[1044] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two pizza margaritas and one Pepsi?", increasing satisfaction.

[1045] 5. User Verification

[1046] The user presses the confirmation button or answers "yes" aloud.

[1047] 6. Order data conversion and storage

[1048] The server stores order data and links it to past order history and emotion data.

[1049] 7. Order Linkage

[1050] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[1051] This system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. By incorporating user sentiment analysis, a more personalized customer experience can be achieved.

[1052] The processing flow will be explained below.

[1053] Processing flow

[1054] Step 1:

[1055] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[1056] Step 2:

[1057] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[1058] Step 3:

[1059] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[1060] Step 4:

[1061] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[1062] Step 5:

[1063] The server parses the text of the order and generates a message requiring confirmation, such as "Would you like one carbonara and one coke?"

[1064] Step 6:

[1065] The server uses an emotion engine to analyze emotions from the user's voice data and recognizes emotions such as "joy," "anger," and "sadness."

[1066] Step 7:

[1067] The device receives the confirmation message and the emotion analysis results sent from the server and displays them on the display. For example, if the user expresses anger, the device adds the phrase "We apologize for the inconvenience" to the confirmation message.

[1068] Step 8:

[1069] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[1070] Step 9:

[1071] The server receives the user's response and saves the final order details in a database. The customer ID, product name, quantity, and emotion data are recorded in the database.

[1072] Step 10:

[1073] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[1074] Step 11:

[1075] The device will display a recommended menu on the screen. For example, it will display a recommended dessert along with a message such as "Would you like some dessert?" based on the user's voiceprint and past order history.

[1076] Step 12:

[1077] If the user places an additional order, the same process is repeated. If the user places an additional order, the process starts again from step 2.

[1078] Step 13:

[1079] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[1080] These steps enable users to enjoy a more intuitive and personalized ordering experience through emotional confirmation messages and recommendations, while allowing merchants to efficiently manage orders and provide faster service.

[1081] Example 2

[1082] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1083] Conventional ordering systems are often difficult to use for users who are unfamiliar with the operation or who have visual or hearing impairments. Furthermore, conventional systems do not take into account the user's emotions, making it difficult to provide services that meet individual needs. In particular, when ordering by voice, it is important to reliably understand the order and appropriately confirm it. For this reason, a system that can accurately recognize the user's voice and respond appropriately based on their emotions is required.

[1084] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice receiving means for accepting voice orders, a voice analysis means for analyzing the accepted voice data using a generative AI model, a display means for confirming and correcting the analyzed order details and the user's emotions, a data storage means for saving the analyzed order details and emotion data in a database, and an order linking means for linking the saved order data to a sales management system. This enables accurate recognition of voice orders and individual responses based on the user's emotions.

[1085] The "voice receiving means" is a device or function for capturing the voice uttered by the user and acquiring it as digital data.

[1086] "Voice analysis means" refers to a device or function that analyzes captured voice data using technologies such as generative AI models and converts the order details into text data.

[1087] The "display means" refers to a device or function that displays the analyzed order details and emotion analysis results to the user, allowing the user to confirm or correct them.

[1088] "Data storage means" refers to a device or function that stores the analyzed order details and emotion data in a database so that it can be used for later processing or analysis.

[1089] An "order linking means" is a device or function that links saved order data to other systems such as a sales management system and shares information with the kitchen, cash register, etc.

[1090] The "recommendation generating means" is a device or function for proposing the most suitable menu to a user based on the user's voiceprint data, past order history, and emotional data.

[1091] The "recommendation presentation means" is a device or function that displays a recommendation menu on a screen or the like based on the user's emotional data and attribute information.

[1092] A "generative AI model" is an artificial intelligence model that analyzes voice data and other input information and generates the required results in the form of text data or other information.

[1093] A "prompt sentence" is an input sentence used to prompt an AI model for a particular output.

[1094] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Below, we will explain in detail an embodiment of the system based on the roles of the server, terminal, and user.

[1095] User operations

[1096] The user selects an item from a printed menu or a display screen, then orders the item by voice into a microphone. For example, the user might say, "One carbonara and one coke, please." This voice command becomes the input for the system.

[1097] Terminal handling

[1098] A microphone built into the device captures the user's voice. This voice data is sent to a server in real time. The device is also equipped with an emotion engine that analyzes the captured voice data to recognize the user's emotions. The results of this emotion analysis are displayed as emotions such as "happiness," "anger," or "sadness." Based on this information, the device displays a recommended menu on the display. For example, if the user expresses the emotion of "happiness," it will display a recommended menu such as "recommended desserts."

[1099] Server Processing

[1100] The server analyzes the received voice data using a generative AI model and converts the order details into text data. For example, a voice saying "One carbonara and one cola, please" is converted into text as "One carbonara, one cola." The server then sends a confirmation message to the user if necessary based on the analysis results of the emotion engine. For example, if the user has the emotion "anger," it generates a message that displays "Is your order one carbonara and one cola?" The server saves the order details and emotion data in a database and links the order data to a sales management system. This link allows the order details to be notified to the kitchen and cash register, enabling prompt service.

[1101] Specific examples

[1102] For example, consider the case where a user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures the user's voice and sends that data to the server. The server uses a generative AI model to convert this voice data into text data: "Two Margherita Pizzas, one Pepsi." At the same time, the emotion engine recognizes the user's emotion as "Satisfied." The server then displays a confirmation message to the user, asking, "Are you sure you want two Margherita Pizzas and one Pepsi?" The order is confirmed when the user presses the confirmation button or answers "Yes" verbally. The server then saves the order data in a database and links it to past order history and emotion data. Finally, the server connects the order data to the store's sales management system, which notifies the kitchen and cashier of the order details.

[1103] Example prompt sentence:

[1104] "One carbonara and one coke please."

[1105] "Two pizza margaritas and one Pepsi, please."

[1106] As described above, this system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. Incorporating sentiment analysis can create a more personalized customer experience.

[1107] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1108] Step 1: User speaks

[1109] Specific operation: The user speaks into the microphone the items they want to order. For example, "One carbonara and one coke, please."

[1110] Input: User's voice

[1111] Output: Audio data

[1112] Step 2: Your device captures audio

[1113] Specific operation: The device's built-in microphone captures the user's voice in real time.

[1114] Input: Audio data

[1115] Output: Digital audio data

[1116] Step 3: The device sends the audio data to the server

[1117] Specific operation: The device transmits the captured digital audio data to the server in real time.

[1118] Input: Digital audio data

[1119] Output: Audio data sent to the server

[1120] Step 4: The server analyzes the audio data

[1121] Specific operation: The server analyzes the voice data received using a generative AI model and converts the order details into text data.

[1122] Input: Audio data

[1123] Output: Text data (e.g. "1 Carbonara, 1 Cola")

[1124] Step 5: The device sends the voice data to the emotion engine

[1125] Specific operation: The device sends the captured voice data to the emotion engine, which analyzes the user's emotions.

[1126] Input: Audio data

[1127] Output: Emotion data (e.g., "joy," "anger," "sadness," etc.)

[1128] Step 6: The device sends the emotion analysis results to the server.

[1129] Specific operation: Emotion data analyzed by the emotion engine is sent to the server.

[1130] Input: Emotion data

[1131] Output: Emotion data sent to the server

[1132] Step 7: The server generates a message to confirm and modify the order details and emotion data.

[1133] Specific operation: The server generates a message based on the text data and emotion data, prompting the user to confirm or correct the message if necessary.

[1134] Input: Text data, emotion data

[1135] Output: Confirmation message (e.g. "Would you like to order one carbonara and one coke?")

[1136] Step 8: Your device will display a confirmation message

[1137] Specific operation: The terminal displays the confirmation message received from the server on the display, prompting the user to confirm.

[1138] Input:Confirmation message

[1139] Output: A confirmation message displayed on the display

[1140] Step 9: User confirms or corrects

[1141] Specific operation: The user checks the confirmation message displayed on the screen and confirms or modifies the order details by voice or button operation.

[1142] Input: User confirmation / correction (voice or button operation)

[1143] Output: Confirmed order details

[1144] Step 10: The server saves the confirmed order details and emotion data to the database.

[1145] Specific operation: The server stores the confirmed order details and emotion data in a database.

[1146] Input: Order details, emotion data

[1147] Output: Order data stored in a database

[1148] Step 11: The server connects the order data to the sales management system

[1149] Specific operation: The server connects the saved order data to the sales management system and notifies the kitchen and cash register.

[1150] Input: Order data stored in a database

[1151] Output: Order details linked to the sales management system and notified to the kitchen and cash register

[1152] Through these steps, the system accurately accepts user voice orders and provides optimal confirmation and recommendations based on emotion. This allows stores to efficiently manage order details and provide prompt service.

[1153] (Application example 2)

[1154] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1155] Conventional voice ordering systems simply convert voice to text data and accept orders. However, this method does not provide personalized recommendations based on user sentiment or past order history, making it difficult to improve the customer experience. Furthermore, order confirmation and correction are often insufficient, leading to ordering errors. This makes it difficult for stores to provide efficient service and can lead to lower customer satisfaction.

[1156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1157] In this invention, the server includes a generating AI means for analyzing voice data, a display means for confirming and correcting the analyzed order details, a means for saving the analyzed order details in a database, a means for linking the saved order data to an information processing system, a sentiment analysis means for analyzing emotions from the voice data, and a means for displaying recommendation information based on the analyzed emotions. This makes it possible to provide personalized recommendations based on the user's emotions and past order history, streamline confirmation and correction of order details, and significantly improve user satisfaction.

[1158] "Voice ordering" is a method of accepting orders by inputting voice using a microphone and analyzing the voice data.

[1159] The "microphone means" is a device that captures sound as a digital or analog signal and sends it to the analysis means.

[1160] A "generative AI means" is an artificial intelligence system that converts voice data into text data and analyzes it.

[1161] The "display means" is a device that visually presents the analyzed order details and recommendation information to the user.

[1162] A "database" is an information system for storing and managing analyzed data and order data.

[1163] An "information processing system" is a computer system for linking order data with other systems.

[1164] "Emotion analysis means" is a system for analyzing and identifying user emotions from voice data.

[1165] "Recommendation information" is product and service information suggested based on a user's past order history and emotional data.

[1166] The present invention provides an ordering system that combines voice and emotion analysis, and its embodiments are specifically described below. In particular, this system efficiently accepts voice orders in brick-and-mortar stores and provides recommendations based on the user's emotions.

[1167] Server Processing

[1168] The server is responsible for the main data processing and collaboration. First, it analyzes the received voice data and converts it into text data using a generative AI model. This process uses the Google Speech API. Next, it performs sentiment analysis based on the converted text data. In this process, it calculates an emotion score using the TextBlob library. Specifically, it recognizes emotions such as "joy," "anger," and "sadness" from the voice. Once the analysis is complete, it stores the order data and emotion data in a database. Furthermore, the stored data is linked to the information processing system, and the order details are notified to the physical store's POS system, kitchen, and cash register.

[1169] Terminal handling

[1170] The device has the ability to capture voice input from the user and send it to a server in real time. Specifically, a microphone attached to the device captures voice and sends the voice data to the server. The device also receives the results of emotion analysis and displays menu recommendations based on the user's emotions. For example, if the user expresses the emotion "satisfied," the device will display "recommended desserts" on the display.

[1171] User operations

[1172] Users place orders simply and intuitively using voice. For example, they can complete an order by saying, "One pizza margherita and one coke, please." If the voice order is successful, the device displays a message allowing users to confirm or correct the order. If the user is satisfied, the system will suggest menu recommendations and encourage further orders.

[1173] Specific examples

[1174] As a specific example, suppose a user says, "One pizza margherita and one coke, please." This speech is captured by the device's microphone and sent to the server. The server uses a generative AI model to convert the speech into text and recognizes it as "One pizza margherita, one coke." Next, it uses sentiment analysis to identify the user's emotion as "satisfied." Based on this information, the server generates a recommendation for the user suggesting a "recommended dessert" and displays it on the device. The user confirms this and confirms the order by either saying "yes" again or pressing a button. This series of processes improves the user experience and also improves store operational efficiency.

[1175] Example prompt sentence:

[1176] The user orders by voice, "One pizza margherita and one coke please." If the user feels satisfied, the app will recommend a dessert.

[1177] As described above, this invention is a system that combines voice and emotion analysis to provide users with personalized recommendations and efficiently confirm and modify order details.

[1178] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1179] Step 1: Capture audio

[1180] The terminal uses a microphone to capture the user's voice order, such as "One pizza margherita and one coke please." This voice signal becomes the input to the terminal, and the captured voice data is output.

[1181] Step 2: Sending audio data

[1182] The captured audio data is sent by the device to the server in real time, where the input is the captured audio data and the output is the audio data arriving at the server.

[1183] Step 3: Audio analysis

[1184] The server converts the received voice data into text data using the Google Speech API. This generative AI model processes the voice data as input data and outputs the text data "Pizza Margherita 1, Coke 1."

[1185] Step 4: Sentiment analysis

[1186] The server calculates an emotion score using the TextBlob library based on the analyzed text data. The text data is input, and an emotion score is output, for example, as the emotion of "satisfied."

[1187] Step 5: Recommendation generation

[1188] The server generates recommendation information based on the results of the sentiment analysis. The input is the sentiment score, and recommendation information such as "recommended dessert" is generated based on the sentiment of "satisfaction."

[1189] Step 6: Send what you see

[1190] The server sends the generated recommendation information and a confirmation message of the order details to the terminal. The input is the recommendation information and the order confirmation message, and the output is a display instruction to the terminal.

[1191] Step 7: Display to the user

[1192] The device displays the recommendation information and confirmation messages received from the server. For example, the display might say, "Recommended dessert" or "Order details: One pizza margherita and one coke, is that OK?"

[1193] Step 8: User Verification

[1194] The user responds to the displayed confirmation message by saying "Yes" or by pressing a button. This input is sent to the terminal, and the terminal then sends the response data to the server.

[1195] Step 9: Save order data

[1196] The server receives an acknowledgement from the user and stores the order data in a database, with the input being the user response data and the output being the stored order data.

[1197] Step 10: Linking to information processing systems

[1198] The server connects the saved order data to an information processing system (such as a POS system or kitchen display). The input is the saved order data, and the output is a notification of the order details to the information processing system.

[1199] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1200] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1201] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1202] [Fourth embodiment]

[1203] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1204] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1205] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1206] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1207] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1208] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1209] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1210] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1211] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1212] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1213] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1214] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1215] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1216] The present invention is a system for efficiently and intuitively placing orders using voice. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[1217] User operations

[1218] 1. Select the menu

[1219] The user selects items from a printed menu or a display screen.

[1220] For example, suppose a user wants to order "carbonara" and "cola."

[1221] 2. Voice ordering

[1222] The user orders the selected items by voice into the microphone, for example, saying, "One carbonara and one coke, please."

[1223] Terminal handling

[1224] 1. Audio capture

[1225] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[1226] 2. Display Recommendations

[1227] The terminal analyzes the user's voiceprint and displays a recommended menu on the screen based on past order history and attribute information.

[1228] For example, if a user has previously ordered "cheesecake," the system displays "Would you like dessert?"

[1229] Server Processing

[1230] 1. Audio analysis

[1231] The server analyzes the voice data and converts the content into text using generative AI.

[1232] For example, audio saying "One carbonara and one cola please" is converted into text data as "1 carbonara, 1 cola."

[1233] 2. Confirm your order details

[1234] The server may return a confirmation question to the user if necessary, generating a message such as "Would you like one carbonara and one coke?"

[1235] 3. Order Data

[1236] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[1237] 4. Order Linkage

[1238] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[1239] The kitchen displays an order for "one carbonara" and "one coke."

[1240] The cash register then connects the product information to prepare for payment.

[1241] Specific examples

[1242] 1. User Orders

[1243] A user says, "Two pizza margaritas and one Pepsi, please."

[1244] 2. Capturing and transmitting audio from your device

[1245] The microphone captures the user's speech and transmits the data to a server.

[1246] 3. Server audio analysis

[1247] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi."

[1248] 4. Server Order Confirmation

[1249] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two Margherita pizzas and one Pepsi?"

[1250] 5. User Verification

[1251] The user presses the confirmation button or answers "yes" aloud.

[1252] 6. Order data conversion and storage

[1253] The server stores the order data and links it to past order history and voiceprint data.

[1254] 7. Order Linkage

[1255] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[1256] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, it is possible to provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[1257] The processing flow will be explained below.

[1258] Step 1:

[1259] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[1260] Step 2:

[1261] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[1262] Step 3:

[1263] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[1264] Step 4:

[1265] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[1266] Step 5:

[1267] The server parses the text of the order and generates a confirmation message if necessary, such as "One carbonara and one coke, okay?"

[1268] Step 6:

[1269] The terminal displays a confirmation message on the display, allowing the user to visually confirm the order details.

[1270] Step 7:

[1271] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[1272] Step 8:

[1273] The server receives the user's response and saves the final order details in a database, including the customer ID, product name, and quantity.

[1274] Step 9:

[1275] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[1276] Step 10:

[1277] The device displays a recommended menu on the screen. For example, if a user has previously ordered dessert, the device might recommend, "Would you like a dessert?"

[1278] Step 11:

[1279] If the user wishes to place an additional order, the same process is repeated.

[1280] Step 12:

[1281] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[1282] These steps allow users to easily place orders by voice, and the store can efficiently manage order details and provide service quickly.

[1283] Example 1

[1284] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1285] Conventional ordering systems require users to select products and manually input the exact order details, which is a cumbersome and time-consuming process. Stores also face issues with low order processing efficiency, with orders often being missed or delayed, especially during peak hours. Furthermore, there is a lack of personalized recommendations based on customers' past order history and attribute information, creating a need for improved customer satisfaction.

[1286] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1287] In this invention, the server includes an input means for accepting voice orders, an AI generation means, and an output means, which makes it possible to process orders more efficiently and provide personalized services based on the customer's past order history and attribute information.

[1288] The "input means for accepting voice orders" refers to a data input device for capturing the user's voice. Specifically, it refers to a voice input device such as a microphone.

[1289] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting voice data into text data, including speech recognition models and natural language processing algorithms.

[1290] "Output means for confirming or correcting the analyzed order details" refers to a device that presents the analyzed order details to the user for confirmation or correction. Specifically, this refers to a display, voice response unit, etc.

[1291] The "means for saving the analyzed order details in a database" refers to a storage system for permanently saving the analyzed order details. Specifically, this refers to a relational database, cloud storage, etc.

[1292] "Means for linking stored order data to an information system" refers to communication means for transferring or sharing stored order data with other information systems, such as POS systems. Specifically, this refers to APIs and network interfaces.

[1293] "Means for analyzing customer voiceprint data and generating recommended menus based on past order history and customer attribute information" refers to an algorithm for generating recommended menus using customer voiceprint data, past order history, and attribute information. Specifically, this refers to a machine learning model or recommendation engine.

[1294] The present invention is a system for efficiently and intuitively placing orders using voice. To specifically implement this system, the program processing will be described in detail based on the roles of the server, terminal, and user.

[1295] User operations

[1296] The user selects items from a printed menu or a display screen. For example, if the user wants to order "carbonara" and "Coke," they would say into the microphone, "One carbonara and one Coke, please."

[1297] Terminal handling

[1298] A microphone is connected to the device, which captures the user's voice. This voice data is sent to the server in real time. The device then analyzes the user's voiceprint and displays recommended menu items on the display based on the user's past order history and attribute information. For example, if a user has previously ordered cheesecake, the device might ask, "Would you like dessert?"

[1299] Server Processing

[1300] The server analyzes the voice data and converts it into text using a generative AI model. For example, a user's speech, "One carbonara and one cola, please," is converted into text as "1 carbonara, 1 cola." The server then generates a confirmation message for the user, displaying "Is one carbonara and one cola okay?" Once the user confirms, the server saves the order data in a database. The server also connects the saved order data to the store's information system (POS system) and notifies the kitchen and cash register.

[1301] Specific examples

[1302] The user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures this voice and sends it to the server. The server uses a generative AI model to analyze the voice data and converts it into text: "Two Margherita Pizzas, one Pepsi." It then generates a confirmation message asking, "Are you sure your order is for two Margherita Pizzas and one Pepsi?" and displays it on the device's screen. Once the user confirms, the server saves this order data in a database and links the order data to the store's POS system. As a result, an order for "two Margherita Pizzas" and "one Pepsi" is displayed on the kitchen display, and the product information is sent to the cash register, completing the checkout process.

[1303] In this system, a generative AI model is used to analyze voice data. The model is input with the following prompt: "One carbonara and one coke, please." The model analyzes this prompt and generates the order details as text data.

[1304] This system allows users to intuitively place orders by voice, and stores can efficiently manage orders. By utilizing order history and voiceprints, stores can provide more appropriate service to returning customers, improving customer satisfaction and store sales.

[1305] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1306] Step 1:

[1307] The user selects a menu item and orders by voice.

[1308] Input: The user selects a menu item and decides what to order.

[1309] Action: The user speaks into the microphone, "One carbonara and one coke please."

[1310] Output: The microphone captures the user's voice.

[1311] Step 2:

[1312] The device captures the audio and sends it to the server

[1313] Input: User's spoken order.

[1314] How it works: The device's microphone captures audio data and converts it into a digital format.

[1315] Output: The device sends the converted audio data to the server in real time.

[1316] Step 3:

[1317] The server analyzes the voice data and converts it into text

[1318] Input: Audio data sent from the device.

[1319] How it works: The server inputs speech data into the generative AI model. For example, it analyzes speech data such as "One carbonara and one coke, please."

[1320] Data processing: The voice data is input as a prompt sentence into the generative AI model and converted into text data.

[1321] Output: Text data: "1 carbonara, 1 cola".

[1322] Step 4:

[1323] The server verifies the order and asks the user for confirmation.

[1324] Input: Text of the order.

[1325] What happens: The server generates a confirmation message saying "Would you like one carbonara and one coke?"

[1326] Output: A confirmation message will be printed to the terminal screen.

[1327] Step 5:

[1328] The user confirms the order details

[1329] Input: The confirmation message displayed on the device screen.

[1330] Action: The user presses the confirmation button or speaks "yes."

[1331] Output: The user's confirmation action is entered into the terminal.

[1332] Step 6:

[1333] The server saves the order data to a database

[1334] Input: The order details confirmed by the user.

[1335] What happens: The server records a data entry in the database, such as "Customer ID: 12345, Product: Carbonara, Quantity: 1."

[1336] Output: Order data is saved in the database.

[1337] Step 7:

[1338] The server connects the order data to the information system

[1339] Input: Order data stored in the database.

[1340] How it works: The server sends order data to the store's POS system.

[1341] Output: The kitchen display shows the order for "1 Carbonara" and "1 Coke", and the product information is sent to the cashier terminal, ready for payment.

[1342] (Application example 1)

[1343] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1344] In recent years, the use of food delivery services has increased, but there is a demand for improved user convenience and efficient order management. Conventional systems require users to operate an application interface, which is cumbersome, especially for elderly people and users unfamiliar with technology. In addition, the cumbersome process of checking order details and managing order history affects the operational efficiency of food delivery services. To solve these issues, a system that allows users to order intuitively by voice is needed.

[1345] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1346] In this invention, the server includes microphone means for accepting voice orders, AI generation means for analyzing the accepted voice data, display means for confirming and correcting the analyzed order details, means for saving the analyzed order details in a database, means for retrieving the saved order data from the database and linking it to the server, and generating and displaying a confirmation message, means for linking the order details to the food delivery service server and delivery management system, and means for linking the order details with past order history and voiceprint data to provide more appropriate service. This enables users to intuitively place food delivery orders by voice, streamlining the ordering process and improving the user experience.

[1347] "Microphone means for accepting voice orders" refers to devices or technology for detecting and capturing voice.

[1348] "Generative AI means for analyzing received voice data" refers to artificial intelligence technology for converting captured voice data into text and analyzing its meaning.

[1349] "Display means for confirming and correcting the analyzed order contents" refers to a device such as a display or monitor that visually presents the analysis results to the user and allows the user to confirm or correct the contents.

[1350] The "means for storing analyzed order details in a database" refers to a database system for electronically recording and managing analyzed order data.

[1351] "Means for retrieving saved order data from a database, linking to a server, and generating and displaying a confirmation message" refers to a system for retrieving saved order information from a database, sending it to a server, generating a confirmation message, and displaying it to the user.

[1352] "Means for linking order details to the food delivery service's server and delivery management system" refers to technology that transmits order information to the food delivery service's central server and delivery management system, enabling delivery operations to be carried out efficiently.

[1353] "Means for providing more appropriate services by linking with past order history and voiceprint data" refers to technology that utilizes a user's past order history and voiceprint data to provide personalized services such as recommendations and individual responses.

[1354] This invention is a system that uses voice to efficiently and intuitively order food delivery, and operates in cooperation with three parties: a server, a terminal, and a user.

[1355] User operations

[1356] 1. Select the menu

[1357] Users select products from a menu displayed on their smartphone screen or a printed menu.

[1358] For example, a user wants to order a "Pizza Margherita" and a "Pepsi."

[1359] 2. Voice ordering

[1360] The user orders the selected items by voice into the microphone on their smartphone, for example, saying, "Two Margherita pizzas and one Pepsi, please."

[1361] Terminal handling

[1362] 1. Audio capture

[1363] The microphone on the smartphone captures the user's voice, and this voice data is sent to the server in real time.

[1364] 2. Display Recommendations

[1365] The device analyzes the user's voiceprint and displays menu recommendations based on their past order history and attribute information. For example, if a user has previously ordered a salad, the device might suggest, "Would you like to try today's special salad?"

[1366] Server Processing

[1367] 1. Audio analysis

[1368] The server analyzes the voice data and converts it into text using a generative AI model. For example, a speech request such as "Two Margherita pizzas and one Pepsi, please" is converted into text data such as "Two Margherita pizzas, one Pepsi."

[1369] 2. Confirm your order details

[1370] The server generates a confirmation message for the user and displays it on the user's smartphone, such as "Are you sure you want to order two Margherita pizzas and one Pepsi?"

[1371] 3. Order Data

[1372] The server identifies the order details and saves information such as customer ID, product name, and quantity in a database. This constitutes the order data.

[1373] 4. Order Linkage

[1374] The server then connects this order data to the food delivery service's server and delivery management system, and prepares for delivery. An order for "two Margherita pizzas and one Pepsi" is then sent to restaurants within the delivery area.

[1375] Specific examples

[1376] For example, if a user orders by voice, such as "Two Margherita Pizzas and one Pepsi, please," the voice is captured by the smartphone and sent to the server. The server then converts the voice into text, such as "Two Margherita Pizzas, one Pepsi," and displays a confirmation message to the user. When the user presses the confirmation button, the order information is saved in a database, and this data is linked to the food delivery service's server and delivery management system.

[1377] Example prompts for generative AI models

[1378] A user says to their smartphone, "Two Margherita pizzas and one Pepsi, please." Describe the process for converting speech to text and saving the order to a database.

[1379] The above system will enable users to intuitively order food delivery using voice, which is expected to streamline the ordering process and improve the user experience.

[1380] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1381] Step 1:

[1382] The user selects an item from the menu displayed on the smartphone display, specifically, "Pizza Margherita" and "Pepsi." The input is captured as the user's voice command.

[1383] Step 2:

[1384] The user orders the selected items by voice into the smartphone microphone. Specifically, the user orders "Two pizza margheritas and one Pepsi, please." The input is the user's voice data.

[1385] Step 3:

[1386] The device captures the user's voice using a microphone installed on the smartphone. The voice data is sent to the server in real time. The input is the captured voice data, and the output is the voice data sent to the server.

[1387] Step 4:

[1388] The server analyzes the received voice data and converts it into text using a generative AI model. The input is voice data, and the output is text data in the format "Pizza Margherita 2, Pepsi 1."

[1389] Step 5:

[1390] The server generates a message to reconfirm the order based on the analyzed text data and displays it on the user's smartphone. For example, the message might read, "Is your order two Margherita pizzas and one Pepsi?" The input is text data, and the output is a confirmation message.

[1391] Step 6:

[1392] The user presses the confirmation button or answers "yes" by voice. The device detects this confirmation action and connects to the server. The input is the user's confirmation action, and the output is the confirmation result data.

[1393] Step 7:

[1394] The server saves the order to a database, recording information such as customer ID, product name, and quantity. The input is the confirmed order data, and the output is the order record saved in the database.

[1395] Step 8:

[1396] The server retrieves the stored order data from the database, generates a confirmation message, and communicates with the food delivery service's server and delivery management system to prepare the delivery. The input is the order data retrieved from the database, and the output is the order data sent to the food delivery system.

[1397] Step 9:

[1398] The order details are sent to stores within the delivery area, and preparation begins. For example, an order of "two Margherita pizzas and one Pepsi" is displayed at the store. The input is the order data from the food delivery service's server, and the output is the start of cooking preparation at the store.

[1399] Step 10:

[1400] In order to provide more appropriate services by linking past order history and voiceprint data, the server generates a recommendation menu and displays it to the user the next time they use the service. The input is the past order history and voiceprint data, and the output is the recommendation menu.

[1401] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1402] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system will be explained below based on the roles of the server, terminal, and user.

[1403] User operations

[1404] 1. Select the menu

[1405] The user selects items from a printed menu or a display screen.

[1406] For example, a user may want to order "carbonara" and "cola."

[1407] 2. Voice ordering

[1408] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[1409] Terminal handling

[1410] 1. Audio capture

[1411] The microphone on the device captures the user's voice, and this voice data is sent to the server in real time.

[1412] 2. Sentiment Analysis

[1413] The device sends the captured voice data to an emotion engine, which analyzes the user's emotions, such as "joy," "anger," and "sadness."

[1414] 3. Display Recommendations

[1415] The terminal displays a recommended menu on the screen based on the user's emotions, voiceprint, past order history, and attribute information.

[1416] For example, if the user has the emotion "joy," recommendations such as "recommended desserts" will be displayed.

[1417] Server Processing

[1418] 1. Audio analysis

[1419] The server analyzes the voice data and converts the content into text using generative AI.

[1420] For example, audio saying "One carbonara and one cola please" can be converted into text "1 carbonara, 1 cola."

[1421] 2. Confirm your order details

[1422] The server generates a confirmation message if necessary based on the analysis results of the emotion engine.

[1423] For example, if the user has the emotion "anger," the system will display "Would you like to order one carbonara and one cola?"

[1424] 3. Order Data

[1425] The server identifies the order details and saves information such as the customer ID, product name, and quantity in the database. This completes the order data.

[1426] 4. Order Linkage

[1427] The server connects this order data to the store's POS system and notifies the kitchen and cash register.

[1428] The kitchen displays an order for "one carbonara" and "one coke."

[1429] The cash register then connects the product information to prepare for payment.

[1430] Specific examples

[1431] 1. User Orders

[1432] A user says, "Two pizza margaritas and one Pepsi, please."

[1433] 2. Capturing and transmitting audio from your device

[1434] The microphone captures the user's speech and transmits the data to a server.

[1435] 3. Server voice and emotion analysis

[1436] The server converts this speech into text data: "2 Pizza Margheritas, 1 Pepsi," and the emotion engine recognizes the user's emotion as "satisfied."

[1437] 4. Server Order Confirmation

[1438] The server displays a confirmation message to the user on the screen asking, "Are you sure your order is two pizza margaritas and one Pepsi?", increasing satisfaction.

[1439] 5. User Verification

[1440] The user presses the confirmation button or answers "yes" aloud.

[1441] 6. Order data conversion and storage

[1442] The server stores order data and links it to past order history and emotion data.

[1443] 7. Order Linkage

[1444] The server connects the order data to the store's POS system, and the order details are notified to the kitchen and cash register.

[1445] This system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. By incorporating user sentiment analysis, a more personalized customer experience can be achieved.

[1446] The processing flow will be explained below.

[1447] Processing flow

[1448] Step 1:

[1449] The user selects an item from a menu (printed or displayed). For example, the user selects "carbonara" and "cola."

[1450] Step 2:

[1451] The user speaks their order into the microphone, saying, "One carbonara and one coke, please."

[1452] Step 3:

[1453] The microphone on the device captures the user's voice, and this captured voice data is sent to the server in real time.

[1454] Step 4:

[1455] The server analyzes the received voice data, and the generation AI converts the voice data into text data. For example, "One carbonara, one coke please" becomes "1 carbonara, 1 coke."

[1456] Step 5:

[1457] The server parses the text of the order and generates a message requiring confirmation, such as "Would you like one carbonara and one coke?"

[1458] Step 6:

[1459] The server uses an emotion engine to analyze emotions from the user's voice data and recognizes emotions such as "joy," "anger," and "sadness."

[1460] Step 7:

[1461] The device receives the confirmation message and the emotion analysis results sent from the server and displays them on the display. For example, if the user expresses anger, the device adds the phrase "We apologize for the inconvenience" to the confirmation message.

[1462] Step 8:

[1463] The user touches the confirmation button or answers "yes" aloud, and this answer is sent to the server.

[1464] Step 9:

[1465] The server receives the user's response and saves the final order details in a database. The customer ID, product name, quantity, and emotion data are recorded in the database.

[1466] Step 10:

[1467] The server then connects the saved order data to the POS system, which then displays the order for "one carbonara" and "one cola" on the kitchen display and sends the order information to the cash register.

[1468] Step 11:

[1469] The device will display a recommended menu on the screen. For example, it will display a recommended dessert along with a message such as "Would you like some dessert?" based on the user's voiceprint and past order history.

[1470] Step 12:

[1471] If the user places an additional order, the same process is repeated. If the user places an additional order, the process starts again from step 2.

[1472] Step 13:

[1473] After completing the entire order, the user pays at the cash register. Since the order information has already been sent to the cash register, the payment is made quickly.

[1474] These steps enable users to enjoy a more intuitive and personalized ordering experience through emotional confirmation messages and recommendations, while allowing merchants to efficiently manage orders and provide faster service.

[1475] Example 2

[1476] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1477] Conventional ordering systems are often difficult to use for users who are unfamiliar with the operation or who have visual or hearing impairments. Furthermore, conventional systems do not take into account the user's emotions, making it difficult to provide services that meet individual needs. In particular, when ordering by voice, it is important to reliably understand the order and appropriately confirm it. For this reason, a system that can accurately recognize the user's voice and respond appropriately based on their emotions is required.

[1478] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice receiving means for accepting voice orders, a voice analysis means for analyzing the accepted voice data using a generative AI model, a display means for confirming and correcting the analyzed order details and the user's emotions, a data storage means for saving the analyzed order details and emotion data in a database, and an order linking means for linking the saved order data to a sales management system. This enables accurate recognition of voice orders and individual responses based on the user's emotions.

[1479] The "voice receiving means" is a device or function for capturing the voice uttered by the user and acquiring it as digital data.

[1480] "Voice analysis means" refers to a device or function that analyzes captured voice data using technologies such as generative AI models and converts the order details into text data.

[1481] The "display means" refers to a device or function that displays the analyzed order details and emotion analysis results to the user, allowing the user to confirm or correct them.

[1482] "Data storage means" refers to a device or function that stores the analyzed order details and emotion data in a database so that it can be used for later processing or analysis.

[1483] An "order linking means" is a device or function that links saved order data to other systems such as a sales management system and shares information with the kitchen, cash register, etc.

[1484] The "recommendation generating means" is a device or function for proposing the most suitable menu to a user based on the user's voiceprint data, past order history, and emotional data.

[1485] The "recommendation presentation means" is a device or function that displays a recommendation menu on a screen or the like based on the user's emotional data and attribute information.

[1486] A "generative AI model" is an artificial intelligence model that analyzes voice data and other input information and generates the required results in the form of text data or other information.

[1487] A "prompt sentence" is an input sentence used to prompt an AI model for a particular output.

[1488] The present invention is a system that uses voice to allow efficient and intuitive ordering, and further improves the quality of service by combining it with an emotion engine that recognizes the user's emotions. Below, we will explain in detail an embodiment of the system based on the roles of the server, terminal, and user.

[1489] User operations

[1490] The user selects an item from a printed menu or a display screen, then orders the item by voice into a microphone. For example, the user might say, "One carbonara and one coke, please." This voice command becomes the input for the system.

[1491] Terminal handling

[1492] A microphone built into the device captures the user's voice. This voice data is sent to a server in real time. The device is also equipped with an emotion engine that analyzes the captured voice data to recognize the user's emotions. The results of this emotion analysis are displayed as emotions such as "happiness," "anger," or "sadness." Based on this information, the device displays a recommended menu on the display. For example, if the user expresses the emotion of "happiness," it will display a recommended menu such as "recommended desserts."

[1493] Server Processing

[1494] The server analyzes the received voice data using a generative AI model and converts the order details into text data. For example, a voice saying "One carbonara and one cola, please" is converted into text as "One carbonara, one cola." The server then sends a confirmation message to the user if necessary based on the analysis results of the emotion engine. For example, if the user has the emotion "anger," it generates a message that displays "Is your order one carbonara and one cola?" The server saves the order details and emotion data in a database and links the order data to a sales management system. This link allows the order details to be notified to the kitchen and cash register, enabling prompt service.

[1495] Specific examples

[1496] For example, consider the case where a user says, "Two Margherita Pizzas and one Pepsi, please." The device's microphone captures the user's voice and sends that data to the server. The server uses a generative AI model to convert this voice data into text data: "Two Margherita Pizzas, one Pepsi." At the same time, the emotion engine recognizes the user's emotion as "Satisfied." The server then displays a confirmation message to the user, asking, "Are you sure you want two Margherita Pizzas and one Pepsi?" The order is confirmed when the user presses the confirmation button or answers "Yes" verbally. The server then saves the order data in a database and links it to past order history and emotion data. Finally, the server connects the order data to the store's sales management system, which notifies the kitchen and cashier of the order details.

[1497] Example prompt sentence:

[1498] "One carbonara and one coke please."

[1499] "Two pizza margaritas and one Pepsi, please."

[1500] As described above, this system allows users to intuitively place orders by voice and provides optimal confirmation and recommendations based on emotions. It also allows stores to efficiently manage order details and provide service quickly. Incorporating sentiment analysis can create a more personalized customer experience.

[1501] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1502] Step 1: User speaks

[1503] Specific operation: The user speaks into the microphone the items they want to order. For example, "One carbonara and one coke, please."

[1504] Input: User's voice

[1505] Output: Audio data

[1506] Step 2: Your device captures audio

[1507] Specific operation: The device's built-in microphone captures the user's voice in real time.

[1508] Input: Audio data

[1509] Output: Digital audio data

[1510] Step 3: The device sends the audio data to the server

[1511] Specific operation: The device transmits the captured digital audio data to the server in real time.

[1512] Input: Digital audio data

[1513] Output: Audio data sent to the server

[1514] Step 4: The server analyzes the audio data

[1515] Specific operation: The server analyzes the voice data received using a generative AI model and converts the order details into text data.

[1516] Input: Audio data

[1517] Output: Text data (e.g. "1 Carbonara, 1 Cola")

[1518] Step 5: The device sends the voice data to the emotion engine

[1519] Specific operation: The device sends the captured voice data to the emotion engine, which analyzes the user's emotions.

[1520] Input: Audio data

[1521] Output: Emotion data (e.g., "joy," "anger," "sadness," etc.)

[1522] Step 6: The device sends the emotion analysis results to the server.

[1523] Specific operation: Emotion data analyzed by the emotion engine is sent to the server.

[1524] Input: Emotion data

[1525] Output: Emotion data sent to the server

[1526] Step 7: The server generates a message to confirm and modify the order details and emotion data.

[1527] Specific operation: The server generates a message based on the text data and emotion data, prompting the user to confirm or correct the message if necessary.

[1528] Input: Text data, emotion data

[1529] Output: Confirmation message (e.g. "Would you like to order one carbonara and one coke?")

[1530] Step 8: Your device will display a confirmation message

[1531] Specific operation: The terminal displays the confirmation message received from the server on the display, prompting the user to confirm.

[1532] Input:Confirmation message

[1533] Output: A confirmation message displayed on the display

[1534] Step 9: User confirms or corrects

[1535] Specific operation: The user checks the confirmation message displayed on the screen and confirms or modifies the order details by voice or button operation.

[1536] Input: User confirmation / correction (voice or button operation)

[1537] Output: Confirmed order details

[1538] Step 10: The server saves the confirmed order details and emotion data to the database.

[1539] Specific operation: The server stores the confirmed order details and emotion data in a database.

[1540] Input: Order details, emotion data

[1541] Output: Order data stored in a database

[1542] Step 11: The server connects the order data to the sales management system

[1543] Specific operation: The server connects the saved order data to the sales management system and notifies the kitchen and cash register.

[1544] Input: Order data stored in a database

[1545] Output: Order details linked to the sales management system and notified to the kitchen and cash register

[1546] Through these steps, the system accurately accepts user voice orders and provides optimal confirmation and recommendations based on emotion. This allows stores to efficiently manage order details and provide prompt service.

[1547] (Application example 2)

[1548] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1549] Conventional voice ordering systems simply convert voice to text data and accept orders. However, this method does not provide personalized recommendations based on user sentiment or past order history, making it difficult to improve the customer experience. Furthermore, order confirmation and correction are often insufficient, leading to ordering errors. This makes it difficult for stores to provide efficient service and can lead to lower customer satisfaction.

[1550] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1551] In this invention, the server includes a generating AI means for analyzing voice data, a display means for confirming and correcting the analyzed order details, a means for saving the analyzed order details in a database, a means for linking the saved order data to an information processing system, a sentiment analysis means for analyzing emotions from the voice data, and a means for displaying recommendation information based on the analyzed emotions. This makes it possible to provide personalized recommendations based on the user's emotions and past order history, streamline confirmation and correction of order details, and significantly improve user satisfaction.

[1552] "Voice ordering" is a method of accepting orders by inputting voice using a microphone and analyzing the voice data.

[1553] The "microphone means" is a device that captures sound as a digital or analog signal and sends it to the analysis means.

[1554] A "generative AI means" is an artificial intelligence system that converts voice data into text data and analyzes it.

[1555] The "display means" is a device that visually presents the analyzed order details and recommendation information to the user.

[1556] A "database" is an information system for storing and managing analyzed data and order data.

[1557] An "information processing system" is a computer system for linking order data with other systems.

[1558] "Emotion analysis means" is a system for analyzing and identifying user emotions from voice data.

[1559] "Recommendation information" is product and service information suggested based on a user's past order history and emotional data.

[1560] The present invention provides an ordering system that combines voice and emotion analysis, and its embodiments are specifically described below. In particular, this system efficiently accepts voice orders in brick-and-mortar stores and provides recommendations based on the user's emotions.

[1561] Server Processing

[1562] The server is responsible for the main data processing and collaboration. First, it analyzes the received voice data and converts it into text data using a generative AI model. This process uses the Google Speech API. Next, it performs sentiment analysis based on the converted text data. In this process, it calculates an emotion score using the TextBlob library. Specifically, it recognizes emotions such as "joy," "anger," and "sadness" from the voice. Once the analysis is complete, it stores the order data and emotion data in a database. Furthermore, the stored data is linked to the information processing system, and the order details are notified to the physical store's POS system, kitchen, and cash register.

[1563] Terminal handling

[1564] The device has the ability to capture voice input from the user and send it to a server in real time. Specifically, a microphone attached to the device captures voice and sends the voice data to the server. The device also receives the results of emotion analysis and displays menu recommendations based on the user's emotions. For example, if the user expresses the emotion "satisfied," the device will display "recommended desserts" on the display.

[1565] User operations

[1566] Users place orders simply and intuitively using voice. For example, they can complete an order by saying, "One pizza margherita and one coke, please." If the voice order is successful, the device displays a message allowing users to confirm or correct the order. If the user is satisfied, the system will suggest menu recommendations and encourage further orders.

[1567] Specific examples

[1568] As a specific example, suppose a user says, "One pizza margherita and one coke, please." This speech is captured by the device's microphone and sent to the server. The server uses a generative AI model to convert the speech into text and recognizes it as "One pizza margherita, one coke." Next, it uses sentiment analysis to identify the user's emotion as "satisfied." Based on this information, the server generates a recommendation for the user suggesting a "recommended dessert" and displays it on the device. The user confirms this and confirms the order by either saying "yes" again or pressing a button. This series of processes improves the user experience and also improves store operational efficiency.

[1569] Example prompt sentence:

[1570] The user orders by voice, "One pizza margherita and one coke please." If the user feels satisfied, the app will recommend a dessert.

[1571] As described above, this invention is a system that combines voice and emotion analysis to provide users with personalized recommendations and efficiently confirm and modify order details.

[1572] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1573] Step 1: Capture audio

[1574] The terminal uses a microphone to capture the user's voice order, such as "One pizza margherita and one coke please." This voice signal becomes the input to the terminal, and the captured voice data is output.

[1575] Step 2: Sending audio data

[1576] The captured audio data is sent by the device to the server in real time, where the input is the captured audio data and the output is the audio data arriving at the server.

[1577] Step 3: Audio analysis

[1578] The server converts the received voice data into text data using the Google Speech API. This generative AI model processes the voice data as input data and outputs the text data "Pizza Margherita 1, Coke 1."

[1579] Step 4: Sentiment analysis

[1580] The server calculates an emotion score using the TextBlob library based on the analyzed text data. The text data is input, and an emotion score is output, for example, as the emotion of "satisfied."

[1581] Step 5: Recommendation generation

[1582] The server generates recommendation information based on the results of the sentiment analysis. The input is the sentiment score, and recommendation information such as "recommended dessert" is generated based on the sentiment of "satisfaction."

[1583] Step 6: Send what you see

[1584] The server sends the generated recommendation information and a confirmation message of the order details to the terminal. The input is the recommendation information and the order confirmation message, and the output is a display instruction to the terminal.

[1585] Step 7: Display to the user

[1586] The device displays the recommendation information and confirmation messages received from the server. For example, the display might say, "Recommended dessert" or "Order details: One pizza margherita and one coke, is that OK?"

[1587] Step 8: User Verification

[1588] The user responds to the displayed confirmation message by saying "Yes" or by pressing a button. This input is sent to the terminal, and the terminal then sends the response data to the server.

[1589] Step 9: Save order data

[1590] The server receives an acknowledgement from the user and stores the order data in a database, with the input being the user response data and the output being the stored order data.

[1591] Step 10: Linking to information processing systems

[1592] The server connects the saved order data to an information processing system (such as a POS system or kitchen display). The input is the saved order data, and the output is a notification of the order details to the information processing system.

[1593] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1594] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1595] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1596] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1597] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1598] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1599] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1600] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1601] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1602] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1603] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1604] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1605] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1606] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1607] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1608] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1609] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1610] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1611] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1612] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1613] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1614] The following is further disclosed regarding the above embodiment.

[1615] (Claim 1)

[1616] a microphone means for accepting voice orders;

[1617] a generating AI means for analyzing the received voice data;

[1618] A display means for confirming and correcting the analyzed order details;

[1619] A means for storing the parsed order details in a database;

[1620] A system that includes a means for linking stored order data to a POS system.

[1621] (Claim 2)

[1622] 10. The system of claim 1, further comprising means for analyzing customer voiceprint data and generating menu recommendations based on past ordering history.

[1623] (Claim 3)

[1624] 2. The system according to claim 1, further comprising means for displaying a recommended menu based on customer attribute information.

[1625] "Example 1"

[1626] (Claim 1)

[1627] an input means for accepting voice orders;

[1628] a generating AI means for analyzing the received voice data;

[1629] An output means for confirming and correcting the analyzed order details;

[1630] A means for storing the parsed order details in a database;

[1631] A means for linking the stored order data to an information system;

[1632] A means for analyzing customer voiceprint data and generating recommended menus based on past order history and customer attribute information;

[1633] A system including:

[1634] (Claim 2)

[1635] 10. The system of claim 1, further comprising a user selecting a menu and placing an order by voice.

[1636] (Claim 3)

[1637] 10. The system of claim 1, further comprising means for generating an order confirmation message to reconfirm the order to the customer.

[1638] "Application Example 1"

[1639] (Claim 1)

[1640] a microphone means for accepting voice orders;

[1641] a generating AI means for analyzing the received voice data;

[1642] A display means for confirming and correcting the analyzed order details;

[1643] A means for storing the parsed order details in a database;

[1644] means for retrieving the stored order data from the database, linking it to the server, and generating and displaying a confirmation message;

[1645] A means for linking the order details to a food delivery service server and a delivery management system;

[1646] The system includes a means for linking past order history and voiceprint data to provide more appropriate service.

[1647] (Claim 2)

[1648] 10. The system of claim 1, further comprising means for analyzing customer voiceprint data and generating menu recommendations based on past ordering history.

[1649] (Claim 3)

[1650] 2. The system according to claim 1, further comprising means for displaying a recommended menu based on customer attribute information.

[1651] "Example 2: Combining Emotion Engines"

[1652] (Claim 1)

[1653] a voice receiving means for receiving voice orders;

[1654] A voice analysis means for analyzing the received voice data using a generative AI model;

[1655] A display means for confirming and correcting the analyzed order details and user's emotions;

[1656] a data storage means for storing the analyzed order details and emotion data in a database;

[1657] A system including an order linking means for linking stored order data to a sales management system.

[1658] (Claim 2)

[1659] 10. The system of claim 1, further comprising a recommendation generation means for analyzing customer voiceprint data and generating a recommended menu based on past order history and emotion data.

[1660] (Claim 3)

[1661] 2. The system according to claim 1, further comprising a recommendation presentation means for displaying a recommendation menu based on customer attribute information and emotion data.

[1662] "Application example 2 when combining emotion engines"

[1663] (Claim 1)

[1664] a microphone means for accepting voice orders;

[1665] a generating AI means for analyzing the received voice data;

[1666] A display means for confirming and correcting the analyzed order details;

[1667] A means for storing the parsed order details in a database;

[1668] A means for linking the stored order data to an information processing system;

[1669] emotion analysis means for analyzing emotions from voice data;

[1670] The system includes a means for displaying recommendation information based on the analyzed sentiment.

[1671] (Claim 2)

[1672] 10. The system of claim 1, further comprising means for analyzing customer voiceprint data and generating menu recommendations based on past order history and sentiment data.

[1673] (Claim 3)

[1674] 10. The system according to claim 1, further comprising means for displaying a recommendation menu based on customer attribute information and emotion data. [Explanation of symbols]

[1675] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a microphone means for accepting voice orders; a generating AI means for analyzing the received voice data; A display means for confirming and correcting the analyzed order details; A means for storing the parsed order details in a database; A system that includes a means for linking stored order data to a POS system.

2. The system according to claim 1 , further comprising means for analyzing customer voiceprint data and generating menu recommendations based on past ordering history.

3. The system according to claim 1 , further comprising means for displaying a recommended menu based on customer attribute information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A