Ordering method, ordering device

An AI-driven ordering method allows for automated food and beverage orders by matching user voice inputs with restaurant menus, addressing the lack of human-free ordering in existing systems and enhancing convenience and efficiency.

JP2026056495AActive Publication Date: 2026-04-01加藤 健資
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

There is no method for taking food and drink orders without human intervention in existing passenger seat management systems.

Method used

An ordering method that utilizes artificial intelligence to acquire user voice, match linguistic information with restaurant menu information, and identify similar menu items, transmitting this information to a cook terminal.

Benefits of technology

Enables food and beverage orders to be taken without human intervention, providing convenience and efficiency in ordering processes.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

We provide a method for taking food and drink orders without human intervention. [Solution] The ordering method includes the steps of: a computer device acquiring the user's voice; comparing the language information indicated by the voice with menu information which is the restaurant's menu, and requesting the AI ​​to identify one or more menu items similar to the language information; and transmitting the identified menu items to the chef's terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an ordering method.

Background Art

[0002] The statements in this section merely provide background information regarding the present disclosure and do not necessarily constitute prior art.

[0003] Patent Document 1 discloses a passenger seat management POS system characterized by including monitoring means such as a monitoring camera for monitoring passenger seats, displaying a video signal from the monitoring means for monitoring passenger seats on a display unit, and being able to confirm the situation of the passenger seats.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the inventor recognized that at least in the above embodiment, there is a disadvantage that there is no method for taking orders for food and drinks without a person.

Means for Solving the Problems

[0006] At least one aspect of the present disclosure is an ordering method, wherein the apparatus acquires the voice of a user, requests AI to collate the language information indicated by the voice with menu information that is the menu of a restaurant and identify one or more menu information similar to the language information, and transmits the identified menu information to a cook terminal. A method having is provided.

Effects of the Invention

[0007] This configuration has the advantage of at least being able to take food and beverage orders without human intervention.

[0008] These and other aspects, features, and advantages of the Disclosure will become apparent from the following detailed written description of preferred embodiments and aspects taken in conjunction with the following drawings, but variations and modifications thereof may be implemented without departing from the spirit and scope of the novel concepts of the Disclosure. An aspect of one embodiment in the Disclosure may be combined with or replaced by one or more aspects of another embodiment disclosed herein, insofar as they do not conflict. [Modes for carrying out the invention]

[0009] The following disclosure provides many different embodiments and examples for carrying out different features of the presented subject matter. For the sake of simplicity, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by or in contact with a second feature subsequently disclosed may include embodiments in which an additional feature is formed between the first and second features so that they do not come into direct contact, as well as embodiments in which the first and second features are formed so that they do not come into direct contact. Furthermore, the disclosure may repeat reference numbers and / or letters in various examples. This repetition is for the sake of brevity and clarity and does not require in itself to be related to the various embodiments and / or configurations described. Furthermore, when describing the first element as being "connected" or "joined" to the second element, such description includes embodiments in which the first and second elements are directly connected or joined to each other, as well as embodiments in which the first and second elements are indirectly connected or joined to each other by having one or more other elements interposed between them.

[0010] As used herein, the phrase "at least one of" encompasses all the exemplary variations. For example, the phrase "comprises at least one of A, B, or C" is synonymous with "consisting of A, B, C and combinations thereof," and encompasses all conceivable variations of A, B, C, A+B, A+C, B+C, and A+B+C.

[0011] In this disclosure, disclosures using an electronic operator or computer may include embodiments of methods, recording media, apparatus, or programs. Any statement used herein that “A is B” may be replaced with “A includes B” unless otherwise stated herein, to the extent that it does not conflict with or otherwise state otherwise herein.

[0012] The following embodiments can be used for the operating method in at least one embodiment. The following will be explained by reference to the description in JP6456303, which describes at least one embodiment in detail (beginning of reference).

[0013] As used herein, the term “computer” schematically includes, as known in the art, a processor, memory, at least one information storage / retrieval device such as a hard drive, disk drive or flash drive or memory stick, or other non-temporary computer-readable medium or non-temporary storage device, at least one input device such as a keyboard, mouse, point and touch device, touchscreen or microphone, and a display structure such as a well-known computer screen. In addition, a computer may include one or more network connections, such as wired or wireless connections. As known in the art, such a computer or computer system may more or less include, but is not limited to, for example, tablet computers or smart devices, other electronic media and electronic devices.

[0014] As used herein, the terms “cloud” or “cloud computing” refer to centralized and virtualized computing facilities where all computing resources are shared. For application systems and subsystems, since they all reside within the “cloud,” it is no longer possible to refer to specific machines.

[0015] As used herein, the term “Distributed Internet Services System” refers to a distributed internet services platform that translates internet applications for execution in various computing environments. A DIS system distributes internet applications, including content, data, and logic, to any number and type of device, to whatever extent appropriate, and along a network, via Component Distribution Servers / Asset Distribution Servers. Through a DIS, internet applications can be hosted and centrally managed as a service tailored to each user's needs, and can be locally cached and executed on the user's device or nearby locations while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-enabled, capable of enjoying and running distributed internet services. The Distributed Internet Service System is fully described in any one of the following patent families: U.S. Patents No. 7,136857, 7150015, 7181731, 7209921, 7430610, 7685183, 7685577, 7752214, 8326883, 8386525, 8443035, 8458142, 8458222, 8473468, 8527545, and 8650226, and U.S. Patent Publications 20120005205 ​​and 20130091252, all of which, like the present invention, are jointly owned by OP40 Holdings, Inc., and all of which are incorporated by reference. (End of quote)

[0016] Regarding the operating method used in at least one embodiment, the following embodiments can be taken for conventional internet systems that do not use a distributed internet. The following will be explained by reference to the description in JP7113047, which describes at least one embodiment in detail (beginning of reference).

[0017] Embodiments including those specifically disclosed herein can provide an automated response system based on artificial intelligence that is implemented in a manner that mimics actual human conversation, thereby enabling more natural communication with users while quickly and conveniently handling inquiries, reservations, delivery orders, and more.

[0018] The multiple electronic devices 110, 120, 130, and 140 may be fixed or mobile terminals implemented by a computer system. Examples of the multiple electronic devices 110, 120, 130, and 140 include AI speakers, smartphones, mobile phones, navigation systems, PCs (personal computers), notebook PCs, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablets, game consoles, wearable devices, IoT (Internet of Things) devices, VR (virtual reality) devices, and AR (augmented reality) devices. As an example, Figure 1 shows an AI speaker as electronic device 110, but in embodiments of the present invention, electronic device 110 may mean one of a variety of physical computer systems that can communicate with other electronic devices 120, 130, 140 and / or servers 150, 160 via a network 170 using substantially wireless or wired communication methods.

[0019] The communication method is not limited, and may include not only communication methods that utilize communication networks that can be included in network 170 (for example, mobile communication networks, wired internet, wireless internet, broadcasting networks, satellite networks, etc.), but also short-range wireless communication between devices. For example, network 170 may include one or more arbitrary networks such as PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Furthermore, network 170 may include, but is not limited to, one or more network topologies, including bus networks, star networks, ring networks, mesh networks, starbus networks, tree or hierarchical networks.

[0020] Servers 150 and 160 may each be implemented by one or more computer devices that communicate with a plurality of electronic devices 110, 120, 130, 140 via a network 170 to provide commands, codes, files, content, services, etc. For example, server 150 may be a system that provides a first service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170, and server 160 may also be a system that provides a second service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170. As a more specific example, server 150 may provide a service (such as an automatic response service, for example) targeted by the corresponding application as the first service to a plurality of electronic devices 110, 120, 130, 140 through an application that is a computer program installed and executed in a plurality of electronic devices 110, 120, 130, 140. As another example, server 160 may provide, as the second service, a service that distributes files for installation and execution of the above-described application to a plurality of electronic devices 110, 120, 130, 140.

[0021] FIG. 2 is a block diagram for explaining the internal configurations of an electronic device and a server in an embodiment of the present invention. In FIG. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are described as examples for an electronic device. Also, other electronic devices 120, 130, 140 and server 160 may have the same or similar internal configurations as the above-described electronic device 110 or server 150.

[0022] The electronic device 110 and the server 150 may include memory 211, 221, processors 212, 222, communication modules 213, 223, and input / output interfaces 214, 224. The memory 211, 221 may be a non-temporary computer-readable recording medium and may include non-temporary mass storage devices such as RAM (random access memory), ROM (read-only memory), disk drives, SSDs (solid-state drives), and flash memory. Here, non-temporary mass storage devices such as ROM, SSDs, flash memory, and disk drives may be included in the electronic device 110 and the server 150 as separate non-temporary storage devices distinct from the memory 211, 221. The memory 211, 221 may also store an operating system and at least one program code (for example, code for a browser installed and run on the electronic device 110, or code for an application installed on the electronic device 110 to provide a specific service). Such software components may be loaded from a computer-readable recording medium separate from the memory 211, 221. Such other computer-readable recording media may include computer-readable recording media such as floppy® drives, disks, tapes, DVD / CD-ROM drives, and memory cards. In other embodiments, software components may be loaded into memories 211, 221 through communication modules 213, 223 which are not computer-readable recording media. For example, at least one program may be loaded into memories 211, 221 based on a computer program (for example, the application described above) that is installed by a file provided over the network 170 by a file distribution system (for example, the server 160 described above) that distributes installation files for a developer or application.

[0023] The processors 212 and 222 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 212 and 222 by the memories 211 and 221 or the communication modules 213 and 223. For example, the processors 212 and 222 may be configured to execute instructions received according to program codes recorded in a recording device such as the memories 211 and 221.

[0024] The communication modules 213 and 223 may provide functions for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide functions for the electronic device 110 and / or the server 150 to communicate with other electronic devices (for example, the electronic device 120) or other servers (for example, the server 160). As an example, a request generated by the processor 212 of the electronic device 110 according to program codes recorded in a recording device such as the memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, control signals, instructions, contents, files, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 through the communication module 213 of the electronic device 110 via the communication module 223 and the network 170. For example, control signals, instructions, contents, files, etc. received through the communication module 213 from the server 150 may be transmitted to the processor 212 and the memory 211, and the contents and files, etc. may be recorded in a recording medium (the non-temporary recording device described above) that the electronic device 110 may further include.

[0025] The input / output interface 214 may be a means for interface with an input / output device 215. For example, an input device may include a keyboard, mouse, microphone, camera, etc., and an output device may include a display, speaker, haptic feedback device, etc. As another example, the input / output interface 214 may be a means for interface with a device that integrates input and output functions into one, such as a touchscreen. The input / output device 215 may consist of the electronic device 110 and one other device. Also, the input / output interface 224 of the server 150 may be a means for interface with an input or output device (not shown) that connects to or can be included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes instructions for a computer program loaded into memory 211, a service screen or content configured using data provided by the server 150 or electronic device 120 may be displayed on the display via the input / output interface 214.

[0026] Furthermore, in other embodiments, the electronic device 110 and the server 150 may include more components than those shown in Figure 2. However, it is not necessary to explicitly show most of the conventional components in the figure. For example, the electronic device 110 may be implemented to include at least some of the input / output devices 215 described above, and may further include other components such as transceivers, cameras, various sensors, and databases. As a more specific example, if the electronic device 110 is an AI speaker, the electronic device 110 may be implemented to further include a variety of components that are generally included in an AI speaker, such as various sensors, camera modules, various physical buttons, buttons using a touch panel, input / output ports, and vibrators for vibration. (End of quote)

[0027] The machine will now be described. According to at least one embodiment, the user terminal consists of a control unit, RAM, storage unit, graphics processing unit, communication interface, and interface unit, each connected by an internal bus.

[0028] According to at least one embodiment, the control unit consists of a CPU and ROM. The control unit executes programs stored in the storage unit and controls the user terminal. RAM is the work area of ​​the control unit. The storage unit is a memory area for saving programs and data. The control unit reads programs and data from RAM and processes them. By processing the programs and data loaded into RAM, the control unit outputs drawing commands to the graphics processing unit.

[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of this display unit functions as an input unit.

[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or via a wired connection, and can send and receive data with a server device via the communication network. The data received via the communication interface is loaded into RAM and processed by the control unit. External memory (e.g., an SD card) is connected to the interface unit.

[0031] According to at least one embodiment, the user terminal is not particularly limited as long as it is a computer device having a display screen and an input unit. Examples of user terminals include conventional mobile phones, tablet devices, smartphones, and desktop or notebook personal computers. A VR goggle, i.e., a screen (or two display panels, one for each eye) attached to a frame (or headset) that is fixed or attached to the head with a strap, may also be used. The user terminal has an audio output unit.

[0032] According to at least one embodiment, a user terminal can communicate with a server device via a communication network. It can transmit or receive information by establishing a communication connection via the communication network.

[0033] According to at least one embodiment, the server device comprises at least a control unit, RAM, a storage unit, and a communication interface, each connected by an internal bus.

[0034] According to at least one embodiment, the control unit consists of a CPU and ROM, executes programs stored in the storage unit, and controls the server device. The control unit also has an internal timer for timing. RAM is the work area of ​​the control unit. The storage unit is a memory area for saving programs and data. The control unit reads programs and data from RAM and performs program execution processing based on information received from the user terminal, etc.

[0035] This explains AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large-scale language models (LLMs), foundational models, and generative AI. Generative AI uses transformers and employs numerous attention mechanisms. It uses self-supervised learning and Extract Prediction. In this case, the AI ​​can predict the next word. Given a sentence, it predicts the next word from the sentence up to that point. It generates a large number of supervised learning problems. These allow for an AI that can predict the next word. Generative AI can predict grammatical structure, topic connections, and what kind of sentences people with a certain writing style are likely to write. Furthermore, by simply predicting the next sentence, generative AI can learn the underlying structure, causal relationships, and knowledge. Generative AI scales quickly, and its accuracy improves as the number of parameters increases. Ordinary statistics and machine learning overfit if the model parameters are too large compared to the sample size of the data. LLMs become more accurate as the number of parameters increases. One generative AI has 175 billion parameters. Generative AI is overlaid with supervised learning to facilitate smooth conversation. They are taught not to say anything strange. They write essays and act as call center operators.

[0036] According to at least one embodiment, Large Language Models (LLMs) are, in a non-inclusive sense, machine learning natural language processing models built using large datasets and deep learning techniques. Generally, they are adapted to various natural language processing (NLP) tasks such as text classification and generation, sentiment analysis, text summarization, and question answering using a technique called "fine-tuning," which involves training them on specific tasks. According to at least one embodiment, self-supervised learning is close to intrinsic human intelligence. When humans act, they are always predicting what will happen next and predicting the next input. In the process, they can learn the structure of the external world. Predicting the next word is considered intrinsic intelligence and is close to what is done in the cerebral cortex. According to at least one embodiment, Large Language Models memorize all the input information but generalize it to the extent necessary to predict the next word. They do not try to generalize all the information from the beginning. Large Language Models require capacity to remember information, and therefore require parameters. According to at least one embodiment, a large-scale language model incorporates eight models, each with 175 billion parameters or 220 billion parameters.

[0037] According to at least one embodiment, videos and images are represented as a collection of visual patches, which are small data units similar to text tokens in LLMs. Patches effectively represent models of visual data and are used as a highly scalable and effective representation for training generative models on various types of videos and images. Videos are transformed into patches by first compressing the video into a low-dimensional latent space and then decomposing the representation into spatiotemporal patches.

[0038] According to at least one embodiment, a video compression network is a network that reduces the dimensionality of visual data, taking raw video as input and outputting a temporally and spatially compressed latent representation. An AI is trained in this compressed latent space and then generates video within this compressed latent space.

[0039] According to at least one embodiment, Spacetime Latent Patches extract a series of spatiotemporal patches that function as transformer tokens when given a compressed input video. The patch-based representation allows Sora to be trained on videos and images of varying resolutions, lengths, and aspect ratios, and controls the size of the resulting video by arranging randomly initialized patches into a grid of appropriate size during inference.

[0040] According to at least one embodiment, the AI ​​is a diffusion model that, given a noisy patch (and conditioning information such as text prompts) as input, is trained to predict the original "clean" patch. The AI ​​is a diffusion transformer, which exhibits remarkable scaling properties in various domains, including language modeling, computer vision, and image generation. Diffusion transformers are also effective as video generation models. The AI's sample quality improves significantly as the amount of training computation increases.

[0041] According to at least one embodiment, the AI ​​applies caption regeneration technology to train a highly descriptive caption model, which is then used to generate text captions for all videos in the training set. Training highly descriptive captions improves not only the overall quality of the generated videos but also the fidelity of the text. GPT is used to convert short user prompts into long, detailed captions, which are then sent to the model. This allows the AI ​​to generate high-quality videos that precisely follow the user prompts.

[0042] According to at least one embodiment, the AI ​​can perform vectorization in natural language processing according to the following flow. First, as a preprocessing step, it performs a cleaning process on the given text. In the cleaning process, unnecessary words such as JavaScript code and HTML tags contained in the text are removed. These codes are used to display on the internet and are therefore not generally used in natural language processing. Next, the text is divided into words using morphological analysis. Morphological analysis is the classification of a natural language sentence written in characters into the smallest meaningful linguistic unit. Examples of morphological analysis tools include "MeCab", "JUMAN", and "JANOME". Normalization can be used. In normalization, words with the same meaning, such as variations in spelling, are unified into a single word. Stop words are words that are excluded from processing for reasons such as not being usable in natural language processing. Examples of stop words include particles and auxiliary verbs, which do not have meaning on their own. When calculating vectors, these may be removed, and only meaningful words may be targeted. Vectorization may also be performed without removing these stop words. Vectorization is the process of converting strings of words into vectors. Vectorization converts word data into numerical data. When converting words into vectors, methods called Bag of Words and distributed representations are used. Bag of Words is a method of vectorizing a sentence using the number of occurrences of each word in the given sentence. It focuses on how many times a word appears in the sentence, so the order of words and sentences is not considered. Distributed representation is a method of vectorizing by focusing on the meaning of words. By vectorizing the meaning of words, it is possible to give similar vectors to words with similar meanings and usages, and the relationships between words can also be represented by vectors. Vector representation allows for addition and subtraction of word meanings. In applied processing, natural language converted into numerical data can be used as input for machine learning. Specifically, vectorized natural language is fed into a classifier to perform text classification. Tools used in this process include TensorFlow, scikit-learn, and PyTorch.

[0043] The steps for acquiring user voice are disclosed. In at least one implementation, the device can be one or more machines or computers as described in this disclosure. A fixed terminal or a mobile terminal acquires the user's voice. As an example, a fixed terminal is a store terminal installed in a store and includes machines, computers, and tablet terminals as described in this disclosure. The store terminal is installed in the dining area of ​​a restaurant or in close proximity to the seats or tables where users (synonymous in this disclosure with customers) are eating and drinking. The fixed terminal prompts the user to voice their food and drink order. The method of the request may be to display it as text on the fixed terminal's display or to emit a voice message from the fixed terminal. The user voices their food and drink order to the fixed terminal. The fixed terminal acquires the user's voice information. In another embodiment, as an example, a mobile terminal includes a user terminal, a smartphone or tablet terminal owned by the user, and such terminals with the application installed. The user sits down at the seat where they will be eating and drinks and launches the application on the mobile terminal. The application is pre-installed on the mobile terminal. Alternatively, a display showing how to install the application to be installed is placed near the seat. The user launches the application. Based on the application's instructions, the mobile terminal prompts the user to voice their food and drink order. This request may be displayed as text on the mobile terminal's screen or as an audible message from the mobile terminal. The user voices their food and drink order to the mobile terminal. The mobile terminal acquires the user's voice information.

[0044] This document discloses a step in which the AI ​​is requested to match the linguistic information indicated by the voice with menu information, which is a restaurant menu, and to identify one or more menu items that are similar to the linguistic information. In at least one embodiment, the restaurant operator stores one or more menu items in the device. The menu information includes linguistic information, images, videos, etc., related to food and beverage menus. For example, linguistic information such as hamburger, cheeseburger, teriyaki burger, tomato, lettuce, etc., and images or videos corresponding to that linguistic information. Each piece of information can be stored associative information that evokes related information or concepts. For example, for the information "hamburger," information that evokes related or conceptual information such as beef, pork, and bread can be stored. This information may be used for AI analysis, as described later.

[0045] The device requests the AI ​​to identify one or more menu items similar to the linguistic information. In at least one embodiment, the operation of the AI ​​can take any of the embodiments described above. For example, the AI ​​resides in one or more of the following: a distributed internet service system, a fixed terminal, a mobile terminal, an application, or a server. The device provides the AI ​​with linguistic information acquired by the terminal. The AI ​​retrieves the stored menu information. The AI ​​performs vectorization on the linguistic information. Bags of Words can be used for the linguistic information. On the other hand, distributed representations may be more suitable for analyzing user utterances. The AI ​​also performs vectorization on the menu information. The AI ​​compares the vectors of the linguistic information with the vectors of the menu information, identifies menu items with vectors similar to the linguistic information vectors, and determines their order of similarity. In another embodiment, the AI ​​performs various natural language processing (NLP) tasks on both the linguistic information and the menu information, such as text classification, sentiment analysis, and text summarization. Through these tasks, the AI ​​compares the classified text of the linguistic information with the classified text of the menu information and identifies menu items with text that matches the classified text of the linguistic information. Alternatively, even if the text doesn't match, the AI ​​can identify menu information with a high degree of similarity and determine the order of those matches. If the AI ​​cannot identify menu information similar to the linguistic information, it can generate status information indicating that no information was found. For example, if a user voice-inputs "I want to eat beef and bread," the AI ​​will identify the menu information "hamburger," and will also identify menu information such as "cheeseburger" and "teriyaki burger."

[0046] The present invention discloses a step of transmitting identified menu information to a chef's terminal. In at least one embodiment, a computer or terminal serving as the chef's terminal is provided in the location where the chef of the restaurant is located, or in a location close to the kitchen. The terminal may be a fixed terminal or a mobile terminal. In the case of a mobile terminal, it can be a terminal that the chef can wear. The device transmits the menu information identified by the AI, or the top-ranking menu information in order of similarity, to the chef's terminal. The chef's terminal conveys the received menu information to the chef. This is done by displaying the information on the chef's terminal's display or by outputting the menu information by voice. After outputting the menu information, information regarding the process of making that menu or the recipe may be output along with it.

[0047] This document discloses the steps for obtaining confirmed menu information from the identified menu information to finalize an order. The device transmits the menu information identified by the AI, or multiple menu items in order of similarity, to the user terminal. The user terminal presents this information to the user and receives instructions to order these menus. Methods of presentation include displaying the menu information on the user terminal's display or outputting the menu information audibly. Methods of receiving instructions to order include displaying a button to finalize the order on the user terminal's display and determining that an order has been placed when this button is pressed. Alternatively, the user terminal obtains voice information from the user. The device requests the AI ​​to process the voice information using natural language processing. The AI ​​vectorizes the voice information and confirms its similarity to vectors indicating an order or vector information representing intent or emotion, thereby determining that it is an order instruction. The device obtains the menu information for which an order instruction has been received as confirmed menu information. Obtaining this information may involve generating new information as confirmed menu information or adding a confirmed status to the menu information. For example, the user terminal displays menu information such as "hamburger," "cheeseburger," and "teriyaki burger," and the device acquires the user's selected menu information as confirmed menu information.

[0048] The following describes a step of transmitting confirmed menu information to a chef's terminal. In at least one embodiment, the device transmits confirmed menu information to the chef's terminal. The chef's terminal then transmits the received confirmed menu information to the chef. This transmission method may involve displaying the information on the chef's terminal's display or outputting the confirmed menu information audibly. After outputting the confirmed menu information, information regarding the process and recipe for making that menu may be output along with it.

[0049] This document discloses a step in which the AI ​​is requested to match the linguistic information indicated by the voice with the arrangement information corresponding to the menu information and to identify one or more arrangement information similar to the linguistic information. In at least one embodiment, the restaurant operator stores one or more arrangement information in the device. The arrangement information includes linguistic information, images, videos, etc., that accompany the menu information. For example, the arrangement information for the menu information "hamburger" may include linguistic information such as the amount of lemon sauce, whether or not tomatoes are included, whether or not lettuce is included, and one selected cooking method from several cooking methods for the patty. The arrangement information also includes images, videos, etc., that correspond to that linguistic information. Each piece of information may be stored associative information that evokes related information or concepts. For example, for the information "lemon sauce," information that evokes the association or concept of "sour" may be stored, and for the information "tomato," information that evokes the association or concept of "healthy" or "vitamins" may be stored. As described later, this information may be used for AI analysis.

[0050] The device requests the AI ​​to identify one or more arrangement information items similar to the linguistic information. In at least one embodiment, the device provides the AI ​​with linguistic information acquired by the terminal. The AI ​​retrieves the stored arrangement information. The AI ​​vectorizes the linguistic information. Bags of Words can be used for the linguistic information. On the other hand, distributed representations may be more suitable for analyzing user utterances. The AI ​​also vectorizes the arrangement information. The AI ​​compares the vectors of the linguistic information with the vectors of the arrangement information, identifies arrangement information with vectors similar to the linguistic information vectors, and determines the order of their similarity. In another embodiment, the AI ​​performs various natural language processing (NLP) tasks on both the linguistic information and the arrangement information, such as text classification, sentiment analysis, and text summarization. Through these tasks, the AI ​​compares the classified text of the linguistic information with the classified text of the arrangement information and identifies menu information with text that matches the classified text of the linguistic information. Alternatively, even if the texts do not match, it identifies arrangement information with texts that have a high degree of similarity and determines the order of their similarity. If the AI ​​cannot identify any arrangement information similar to the linguistic information, it can generate status information indicating that no similar information was found. For example, if a user voice-inputs "I prefer sour food. I want a healthy meal," the AI ​​will identify arrangement information such as "extra lemon sauce" and "tomato."

[0051] The following describes a step of transmitting identified menu information and arrangement information to a chef's terminal. In at least one embodiment, the device transmits the identified menu information and arrangement information to the chef's terminal. The chef's terminal conveys the received identified menu information and arrangement information to the chef. This is done by displaying the information on the chef's terminal's display or by outputting the confirmed menu information by voice. After outputting the arrangement information, information regarding the process and recipe for making that menu can be output along with it.

[0052] This document discloses the steps for obtaining confirmed arrangement information from the identified arrangement information to finalize an order. The device transmits the arrangement information identified by the AI, or multiple arrangement information in order of similarity, to the user terminal. The user terminal presents this information to the user and receives instructions to order these arrangements. Methods of presentation include displaying the arrangement information on the user terminal's display or outputting the arrangement information by voice. Methods of receiving instructions to order include displaying a button to finalize the order on the user terminal's display and determining that an order has been placed when this button is pressed. Alternatively, the user terminal obtains voice information from the user. The device requests the AI ​​to process the voice information using natural language processing. The AI ​​vectorizes the voice information and confirms its similarity to vector information indicating an order or vector information representing intent or emotion, thereby determining that it is an order instruction. The device obtains the arrangement information for which an order instruction has been received as confirmed arrangement information. Obtaining this information may involve generating new information as confirmed arrangement information or adding the status information "confirmed" to the arrangement information. For example, the user terminal might present customization options such as "extra lemon sauce" and "tomato," and the device would then acquire the user's selected customization options as confirmed customization information.

[0053] The following describes the step of transmitting confirmed menu information and confirmed arrangement information to a chef's terminal. In at least one embodiment, the device transmits the confirmed menu information and confirmed arrangement information to the chef's terminal. The chef's terminal transmits the received confirmed menu information and confirmed arrangement information to the chef. The method of transmission is to display the information on the chef's terminal's display or to output the confirmed menu information and confirmed arrangement information by voice. After outputting the arrangement information, information regarding the process and recipe for making that menu can be output along with it.

[0054] The present invention discloses a step of instructing a self-propelled cart to move to a predetermined location by the approximate time elapsed for the required time corresponding to the menu information or arrangement information, or the confirmed menu information or confirmed arrangement information. In at least one embodiment, the restaurant operator causes the device to store the required time corresponding to the menu information or arrangement information. The required time can be arbitrarily determined by the operator. For example, for the menu information of hamburger, it is 20 minutes. Examples of required times for arrangement information include 5 minutes for well-done patty, 3 minutes for medium, 1 minute for adding tomatoes, etc. The required time corresponding to confirmed menu information or confirmed arrangement information is the required time corresponding to each piece of menu information or arrangement information.

[0055] The device generates order time information based on the time at which menu information or arrangement information, or confirmed menu information or confirmed arrangement information, was obtained. The device then adds the corresponding required time to the order time information to generate estimated completion time information.

[0056] This disclosure concerns a self-propelled cart. In at least one embodiment, the self-propelled cart is an unmanned, self-propelled vehicle or machine. The self-propelled cart has one or more spaces for carrying food. The self-propelled cart may have a computer or terminal as defined in this disclosure. That is, the self-propelled cart may communicate wirelessly with the device. In at least one embodiment, if an object is observed to be an unmanned, self-propelled vehicle or machine, it is considered a "self-propelled cart" as defined in this disclosure. The device instructs the self-propelled cart to move to a predetermined location at approximately the estimated completion time. As a method of instructing self-propulsion, the device wirelessly transmits a program or instruction to move the self-propelled cart to the predetermined location at approximately the estimated completion time. The self-propelled cart receives the program or instruction and moves to the predetermined location at approximately the estimated completion time. Examples of predetermined locations include the location of the cook, the kitchen, and the location where the finished food is handed over.

[0057] The present invention discloses a step of summing up numerical values ​​corresponding to menu information or arrangement information, or confirmed menu information or confirmed arrangement information. In at least one embodiment, the operator of a restaurant stores numerical values ​​corresponding to menu information or arrangement information in the device. The numerical values ​​represent at least the value of these menus or arrangements. For example, a hamburger might be 1000, and extra lemon sauce might be 50. That is, a hamburger costs 1000 yen and lemon sauce costs 50 yen. The device sums up the numerical values ​​corresponding to menu information or arrangement information, or confirmed menu information or confirmed arrangement information. For example, if a hamburger is ordered with extra lemon sauce, the sum would be 1050.

[0058] The following describes the step of presenting the aggregated figures to the user. In at least one embodiment, the device transmits the aggregated figures to the user terminal. In another embodiment, the device transmits the aggregated figures to an accounting device for the user to pay. Examples of accounting devices include accounting devices in restaurants, cash registers, user accounting applications, user electronic payment applications, and user terminals (including smartphones).

[0059] In at least one embodiment, any one of the embodiments described above has industrial applicability and advantages in the following cases: namely, restaurants, drive-throughs, etc. In the case of a drive-through, the user terminal in the step of acquiring the user's voice may correspond to a machine installed outdoors for the drive-through. In other embodiments, in the case of a drive-through, the user terminal may correspond to a device such as a smartphone or computer owned by the user, or these devices with a dedicated application installed.

[0060] The following is an overview of the embodiments described above.

[0061] The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is asked to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The steps include: sending the identified menu information to the chef's terminal, A method that has

[0062] According to this disclosure, at the very least, users can enjoy the convenience of ordering food and beverages by voice. Furthermore, restaurants can take orders based on user voice commands without human intervention, offering industrial applicability and advantages.

[0063] The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is asked to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. From the identified menu information, the step is to obtain the confirmed menu information to finalize the order, The steps include sending the confirmed menu information to the chef's terminal, A method that has

[0064] According to this disclosure, at the very least, users can enjoy the convenience of ordering food and beverages by voice. Furthermore, restaurants can take orders based on user voice commands without human intervention, offering industrial applicability and advantages.

[0065] The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is asked to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The process involves requesting the AI ​​to compare the linguistic information indicated by the voice with the arrangement information corresponding to the menu information, and to identify one or more arrangement information items that are similar to the linguistic information. The steps include: transmitting the identified menu information and arrangement information to the chef's terminal; A method that has

[0066] According to this disclosure, at the very least, users can enjoy the convenience of ordering food and beverages by voice. Furthermore, restaurants can take orders based on user voice commands without human intervention, offering industrial applicability and advantages.

[0067] The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is asked to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The process involves requesting the AI ​​to compare the linguistic information indicated by the voice with the arrangement information corresponding to the menu information, and to identify one or more arrangement information items that are similar to the linguistic information. From the identified menu information, the step is to obtain the confirmed menu information to finalize the order, From the identified arrangement information, the step is to obtain confirmed arrangement information to finalize the order, The steps include sending the confirmed menu information and confirmed arrangement information to the chef's terminal, A method that has

[0068] According to this disclosure, at the very least, users can enjoy the convenience of ordering food and beverages by voice. Furthermore, restaurants can take orders based on user voice commands without human intervention, offering industrial applicability and advantages.

[0069] Any of the above ordering methods, The steps include: instructing the self-propelled trolley to move to the designated location by approximately the time elapsed corresponding to the menu information or arrangement information; A further method

[0070] According to this disclosure, at a minimum, food and beverages ordered by the user can be quickly transported by a self-propelled cart. This has industrial applicability and advantages, as it enables the delivery of freshly prepared food to the user unmanned.

[0071] Any of the above ordering methods, A step of summing up numerical values ​​corresponding to menu information or arrangement information, The steps include presenting the combined figures to the user, A further method

[0072] According to this disclosure, there are industrial applications and advantages, at least in that it is possible to process payments for food and beverages ordered by users without human intervention.

[0073] In at least one embodiment, the step of "obtaining confirmed menu information for confirming an order from identified menu information" includes the following form of "a method by which the device receives an instruction to order." The restaurant operator defines action information relating to one or more physical actions in the device and stores it as specific action information. The action information includes information such as language information, images, and videos that accompany the physical action. For example, for the physical action of "tapping the table twice with fingers," it includes the language information and information such as images or videos showing examples of the action. Physical actions include any gestures such as raising a hand or nodding. Specific action information is information that the restaurant operator registers in the device from among the action information. The user terminal obtains the user's action information. One method of acquisition is to obtain a video or video of the user's actions using the camera on the user terminal.

[0074] The device requests the AI ​​to determine if the motion information is similar to specific motion information. The AI ​​represents the acquired motion information as a collection of visual patches, which are small data units similar to LLM text tokens, in the form of videos or images. The AI ​​takes raw video as input and outputs a temporally and spatially compressed latent representation in a network that reduces the dimensionality of the visual data about the motion information. The AI ​​also outputs a latent representation of specific motion information in the same manner as described above. The AI ​​compares the latent representation of the motion information with the latent representation of the specific motion information to determine if they are similar. For this determination, the restaurant operator can have the device store a similarity threshold for determining similarity. If the AI ​​determines that the motion information is similar to the specific motion information, it can create data or status information indicating that it is similar.

[0075] The device acquires the menu information to be ordered as confirmed menu information, provided that data or status information indicating similarity has been generated. Acquisition can be done by generating new information as confirmed menu information or by adding the status information of "confirmed" to the menu information. For example, the device presents "hamburger" as menu information to the user. The user taps the table twice with their finger. The device determines that the user's action information is specific action information and acquires "hamburger" as confirmed menu information.

[0076] In at least one embodiment, in the step of "acquiring confirmed arrangement information to confirm an order from among the identified arrangement information," the method by which the device "receives an instruction to place an order" can be the embodiment described above. That is, the operator of a restaurant defines action information relating to one or more physical actions in the device and stores it as specific action information. The user terminal acquires the user's action information. The AI ​​determines whether the action information is similar to confirmed action information. If data or status information indicating similarity is created, the device acquires the arrangement information for which an order instruction has been received as confirmed arrangement information. Acquisition can be done by generating new information as confirmed arrangement information or by adding status information indicating confirmation to the arrangement information. As an example, the device presents "extra lemon sauce" to the user as arrangement information. The user nods. The device determines that the user's action information is specific action information and acquires "extra lemon sauce" as confirmed arrangement information.

[0077] In at least one embodiment, a device or terminal for detecting user activity information may be provided separately from the user terminal, or in combination with the user terminal, as described in the above embodiment.

[0078] In at least one embodiment, instead of operation information and specific operation information, language information and specific language information may be used as the information to be determined in the above embodiment. That is, the restaurant operator defines language information equivalent to one or more passwords in the device and stores it as specific language information. The user terminal acquires the user's language information. The acquisition method may include using the microphone of the device or terminal. The AI ​​determines whether the language information is similar to the specific operation information using natural language processing. The device requests the AI ​​to make that determination. As already described, any one or more of the methods described herein can be used for natural language processing. If data or status information indicating that the language information is similar to the specific language information is created by the AI ​​or the device, the device acquires the menu information or arrangement information that has been instructed to be ordered as confirmed menu information or confirmed arrangement information. Acquisition may include generating new information as confirmed arrangement information or adding the status information of "confirmed" to the arrangement information. As an example, the device presents "extra lemon sauce" to the user as arrangement information. The user utters a password set by the restaurant operator (examples include the restaurant's trademark, emblem, brand name, nickname, or special catchphrase). The device determines that the user's language information is specific language information and acquires "extra lemon sauce" as confirmed customization information.

[0079] The following is an overview of the embodiments described above.

[0080] Any of the above ordering methods, The step to obtain confirmed menu information to finalize the order is: The device stores specific action information relating to one or more physical actions, A step in which the AI ​​is asked to determine whether the action information is similar to specific action information, It also includes.

[0081] Any of the above ordering methods, The step to obtain confirmed menu information to finalize the order is: When the AI ​​determines that the action information is similar to specific action information, it takes the step of making the menu information confirmed menu information. It also includes.

[0082] According to this disclosure, at the very least, users can confirm their orders by gesture, making it easier to confirm orders than by touching a touch panel as in conventional technologies. This has industrial applicability and advantages, as it improves the convenience of using the device or system.

[0083] The invention disclosed herein only needs to achieve at least one of the effects described above.

Claims

1. The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is requested to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The steps include: sending the identified menu information to the chef's terminal, A method that has

2. The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is requested to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. From the identified menu information, the step is to obtain the confirmed menu information to finalize the order, The steps include sending the confirmed menu information to the chef's terminal, A method that has

3. The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is requested to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The process involves requesting the AI ​​to compare the linguistic information indicated by the voice with the arrangement information corresponding to the menu information, and to identify one or more arrangement information items that are similar to the linguistic information. The steps include: transmitting the identified menu information and arrangement information to the chef's terminal; A method that has

4. The ordering method is, the device is, Steps to acquire the user's voice, The process involves a step in which the AI ​​is requested to match the linguistic information indicated by the voice with the menu information of a restaurant, and to identify one or more menu items that are similar to the linguistic information. The process involves requesting the AI ​​to compare the linguistic information indicated by the voice with the arrangement information corresponding to the menu information, and to identify one or more arrangement information items that are similar to the linguistic information. From the identified menu information, the step is to obtain the confirmed menu information to finalize the order, From the identified arrangement information, the step is to obtain confirmed arrangement information to finalize the order, The steps include sending the confirmed menu information and confirmed arrangement information to the chef's terminal, A method that has

5. An ordering method according to any one of claims 1 to 4, The steps include: instructing the self-propelled trolley to move to the designated location by approximately the time elapsed corresponding to the menu information or arrangement information; A further method

6. An ordering method according to any one of claims 1 to 4, A step of summing up numerical values ​​corresponding to menu information or arrangement information, The steps include presenting the combined figures to the user, A further method

7. An ordering method according to either claim 2 or 4, The step to obtain confirmed menu information to finalize the order is: The device stores specific action information relating to one or more physical actions, A step in which the AI ​​is asked to determine whether the action information is similar to specific action information, It also includes.

Citation Information

Patent Citations

  • Order receiving device and order receiving method for restaurant

    JP2005182140A

  • Data entry device, method and program

    JP2011103007A

  • Management apparatus, management method, and management program

    JP2019159378A

  • Information processing apparatus, information processing system and control program thereof

    JP2021149267A

  • On-demand coordinated food item delivery system

    JP2021500684A