Ordering method and ordering device

The method allows users to order food and drinks by voice, using AI to match audio inputs with menu items and send orders to a cook's terminal, addressing the lack of autonomous ordering in existing systems and enhancing user convenience and restaurant efficiency.

JP7678445B1Active Publication Date: 2025-05-16加藤 健資
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025022921
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-02-15
Publication Date
2025-05-16
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

Existing seat management POS systems lack the capability to order food and drinks without human intervention.

Method used

A method that utilizes voice recognition to match user audio inputs with restaurant menu information, using AI to identify similar menu items and send the selected items to a cook's terminal for preparation.

Benefits of technology

Enables users to order food and drinks independently without human assistance, improving convenience for users and reducing labor costs for restaurants.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide an ordering method capable of taking food and drink orders unmanned. [Solution] The ordering method includes the steps of: a device acquiring a user's voice; matching the language information indicated by the voice with menu information that is the restaurant's menu, and requesting an AI to identify one or more menu information items that are similar to the language information; and transmitting the identified menu information to a chef's terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an ordering method. [Background technology]

[0002] The statements in this section are merely intended to provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Patent document 1 discloses a seating management POS system that is characterized by having a monitoring means, such as a surveillance camera, for monitoring the seating area, and displaying a video signal from the monitoring means for monitoring the seating area on a display unit, allowing the status of the seating to be confirmed. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Unexamined Japanese Patent Publication No. 8-161403 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the inventors have recognized that at least the above embodiment has a drawback in that it does not provide a way to take food and drink orders in an unmanned manner. [Means for solving the problem]

[0006] At least one aspect of the present disclosure is a method for producing a method for manufacturing a semiconductor device comprising: 1. A method of ordering, comprising: acquiring a speech signal from a user; A step of requesting the AI ​​to match the linguistic information indicated by the voice with menu information that is a menu of a restaurant, and to identify one or more pieces of menu information that are similar to the linguistic information; transmitting the identified menu information to a chef terminal; Method for having to deliver. Effect of the Invention

[0007] This configuration has the advantage of at least allowing food and drink orders to be taken unattended.

[0008] These and other aspects, features, and advantages of the present disclosure will become apparent from the following detailed written description of the preferred embodiments and aspects taken in conjunction with the following drawings, variations and modifications of which may be made without departing from the spirit and scope of the novel concepts of the present disclosure. An aspect of one embodiment of the present disclosure may be combined with or substituted for one or more of the aspects of another embodiment of the present disclosure, unless inconsistent. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] In the following disclosure, many different embodiments and examples are provided for implementing different features of the presented subject matter. To simplify the disclosure, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by or in contact with a second feature subsequently disclosed may include embodiments in which the first feature and the second feature are formed in direct contact with each other, as well as embodiments in which an additional feature is formed between the first feature and the second feature such that the first feature and the second feature are not in direct contact with each other. Furthermore, the disclosure may repeat reference numbers and / or letters in various examples. Such repetition is for the sake of brevity and clarity, and does not, in itself, require a relationship between the various embodiments and / or configurations described. Furthermore, when a first element is described as "coupled" or "bonded" to a second element, such description includes embodiments in which the first and second elements are directly coupled or bonded to each other, as well as embodiments in which the first and second elements are indirectly coupled or bonded to each other with one or more other intervening elements therebetween.

[0010] As used herein, the phrase "at least one of" includes all exemplary variations. For example, the phrase "comprises at least one of A, B, or C" is equivalent to "consisting of A, B, C and combinations thereof," and includes all possible variations of A, B, C, A+B, A+C, B+C, and A+B+C.

[0011] In this disclosure, the disclosure of using an electronic operator or computer may include embodiments of a method, a recording medium, an apparatus, or a program. As used herein, the statement "A is B" may be replaced with "A includes B" unless there is a contradiction or otherwise stated in the specification.

[0012] The operating method used in at least one or more of the embodiments can take the following embodiments. The following description will be given with reference to JP6456303 (the following reference begins), which clearly explains at least one or more of the embodiments.

[0013] As used herein, the term "computer" generally includes a processor, memory, at least one information storage / retrieval device, such as a hard drive, disk drive or flash drive or memory stick, or other non-transitory computer-readable medium or non-transitory storage device, at least one input device, such as a keyboard, mouse, point and touch device, touch screen, or microphone, and a display structure, such as a well-known computer screen, as known in the art. In addition, a computer may include one or more network connections, such as wired or wireless connections. As known in the art, such computers or computer systems may include more or less of those listed above, including, for example, but not limited to, tablet computers and smart devices, as well as other electronic media and devices.

[0014] As used herein, the term "cloud" or "cloud computing" refers to a centralized and virtualized computing facility in which all computing resources are shared. Application systems and subsystems can no longer be pointed to specific machines, as they are all in the "cloud."

[0015] As used herein, the term "Distributed Internet Services System" refers to a distributed Internet services platform that transforms Internet applications to run in various computing environments. The DIS system distributes Internet applications, including content, data, and logic, to whatever extent appropriate and along the network to any number and type of devices via Component Distribution Servers / Asset Distribution Servers. Through the DIS, Internet applications can be hosted and centralized, with services based on each user's needs, cached and run locally on the user's device or nearby locations while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-enabled to enjoy and run distributed Internet services. The distributed internet services system is more fully described in any one of the following patent families: U.S. Patent Nos. 7,136,857, 7150015, 7181731, 7209921, 7430610, 7685183, 7685577, 7752214, 8326883, 8386525, 8443035, 8458142, 8458222, 8473468, 8527545, and 8650226, and U.S. Patent Publication Nos. 20120005205, and 20130091252, all of which, as well as the present invention, are commonly owned by OP40, Holdings, Inc., and all of which are incorporated by reference herein. (End of quote)

[0016] The operation method used in at least one or more embodiments can be implemented in the conventional Internet manner without using a distributed Internet. At least one or more embodiments are described with reference to JP7113047 (the following references are included), which clearly explains the invention.

[0017] Embodiments including those specifically disclosed in this specification can provide an automated response system that is based on artificial intelligence and is implemented in a form that resembles a real human conversation, thereby enabling more natural conversations with users while quickly and conveniently handling inquiries, reservations, delivery orders, etc.

[0018] The electronic devices 110, 120, 130, and 140 may be fixed terminals or mobile terminals realized by computer systems. Examples of the electronic devices 110, 120, 130, and 140 include an AI speaker, a smartphone, a mobile phone, a navigation system, a personal computer (PC), a notebook PC, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a tablet, a game console, a wearable device, an internet of things (IoT) device, a virtual reality (VR) device, and an augmented reality (AR) device. As an example, FIG. 1 shows an AI speaker as the electronic device 110, but in the embodiment of the present invention, the electronic device 110 may mean one of various physical computer systems that can communicate with other electronic devices 120, 130, and 140 and / or servers 150 and 160 via a network 170 using a substantially wireless or wired communication method.

[0019] The communication method is not limited, and may include not only a communication method using a communication network (such as a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.) that the network 170 can include, but also short-range wireless communication between devices. For example, the network 170 may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), the Internet, etc. Furthermore, the network 170 may include any one or more of a network topology including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or a hierarchical network, etc.

[0020] Each of the servers 150 and 160 may be realized by one or more computer devices that communicate with the multiple electronic devices 110, 120, 130, and 140 via the network 170 to provide instructions, codes, files, content, services, and the like. For example, the server 150 may be a system that provides a first service to the multiple electronic devices 110, 120, 130, and 140 connected via the network 170, and the server 160 may be a system that provides a second service to the multiple electronic devices 110, 120, 130, and 140 connected via the network 170. As a more specific example, the server 150 may provide a service (such as an auto-answer service, for example) targeted by an application, which is a computer program installed and executed in the multiple electronic devices 110, 120, 130, and 140, as a first service to the multiple electronic devices 110, 120, 130, and 140. As another example, the server 160 may provide, as a second service, a service of distributing files for installing and executing the above-mentioned application to the multiple electronic devices 110, 120, 130, and 140.

[0021] Fig. 2 is a block diagram for explaining the internal configuration of an electronic device and a server in an embodiment of the present invention. In Fig. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are explained as examples of electronic devices. In addition, other electronic devices 120, 130, 140 and server 160 may have the same or similar internal configuration as electronic device 110 or server 150 described above.

[0022] The electronic device 110 and the server 150 may include memories 211 and 221, processors 212 and 222, communication modules 213 and 223, and input / output interfaces 214 and 224. The memories 211 and 221 may be non-transitory computer-readable recording media and may include non-transitory mass storage devices such as random access memory (RAM), read only memory (ROM), disk drives, solid state drives (SSD), flash memories, and the like. Here, non-transitory mass storage devices such as ROM, SSD, flash memories, and disk drives may be included in the electronic device 110 and the server 150 as separate non-transitory storage devices separate from the memories 211 and 221. In addition, the memories 211 and 221 may store an operating system and at least one program code (for example, a code for a browser installed and executed in the electronic device 110, an application installed in the electronic device 110 to provide a specific service, and the like). Such software components may be loaded from a computer-readable recording medium separate from the memories 211 and 221. Such another computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. In another embodiment, the software components may be loaded into the memory 211, 221 through a communication module 213, 223 that is not a computer-readable recording medium. For example, at least one program may be loaded into the memory 211, 221 based on a computer program (for example, the above-mentioned application) that is installed by a file provided via the network 170 by a developer or a file distribution system that distributes an installation file of the application (for example, the above-mentioned server 160).

[0023] The processors 212, 222 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 212, 222 by the memory 211, 221 or the communication modules 213, 223. For example, the processors 212, 222 may be configured to execute instructions received according to program code stored in a storage device such as the memory 211, 221.

[0024] The communication modules 213 and 223 may provide a function for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide a function for the electronic device 110 and / or the server 150 to communicate with other electronic devices (for example, the electronic device 120) or other servers (for example, the server 160). For example, a request generated by the processor 212 of the electronic device 110 according to a program code recorded in a recording device such as the memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, a control signal, an instruction, a content, a file, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 through the communication module 213 of the electronic device 110 via the communication module 223 and the network 170. For example, control signals, instructions, content, files, etc. from the server 150 received through the communication module 213 may be transmitted to the processor 212 or memory 211, and the content, files, etc. may be recorded on a recording medium (the non-transitory recording device described above) that the electronic device 110 may further include.

[0025] The input / output interface 214 may be a means for interfacing with the input / output device 215. For example, the input device may include a device such as a keyboard, a mouse, a microphone, or a camera, and the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface 214 may be a means for interfacing with a device in which functions for input and output are integrated into one, such as a touch screen. The input / output device 215 may be configured as one device together with the electronic device 110. Also, the input / output interface 224 of the server 150 may be a means for interfacing with a device for input or output (not shown) that may be connected to the server 150 or included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes the instructions of the computer program loaded in the memory 211, a service screen or content configured using data provided by the server 150 or the electronic device 120 may be displayed on a display through the input / output interface 214.

[0026] Also, in other embodiments, the electronic device 110 and the server 150 may include more components than those in FIG. 2. However, it is not necessary to clearly show most of the conventional technical components in the figure. For example, the electronic device 110 may be realized to include at least some of the input / output devices 215 described above, and may further include other components such as a transceiver, a camera, various sensors, a database, etc. As a more specific example, if the electronic device 110 is an AI speaker, various components such as various sensors, a camera module, various physical buttons, a button using a touch panel, an input / output port, a vibrator for vibration, etc., which are generally included in an AI speaker, may be realized to be further included in the electronic device 110. (End of quote)

[0027] Machine Description: According to at least one embodiment, a user terminal includes a control unit, a RAM, a storage unit, a graphics processing unit, a communication interface, and an interface unit, each connected by an internal bus.

[0028] According to at least one embodiment, the control unit is composed of a CPU and a ROM. The control unit executes programs stored in the storage unit and controls the user terminal. The RAM is the work area of ​​the control unit. The storage unit is a memory area for saving programs and data. The control unit reads the programs and data from the RAM and processes them. The control unit processes the programs and data loaded into the RAM, thereby outputting drawing commands to the graphics processing unit.

[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of the display unit functions as the input unit.

[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or by wire, and can transmit and receive data to and from a server device via the communication network. Data received via the communication interface is loaded into RAM, and arithmetic processing is performed by the control unit. An external memory (e.g., an SD card, etc.) is connected to the interface unit.

[0031] According to at least one embodiment, the user terminal is a computing device having a display screen and an input section. Examples of user terminals include conventional mobile phones, tablet terminals, smartphones, and desktop and notebook personal computers. The user terminal may be comprised of a VR goggle, i.e., a screen (or two display panels, one for each eye) attached to a frame (or headset) that is strapped or attached to the head. The user terminal has an output section for audio.

[0032] According to at least one embodiment, the user terminal is communicatively connected to the server device via a communications network, and is capable of transmitting information or receiving information via the communications network.

[0033] According to at least one embodiment, the server device includes at least a control unit, a RAM, a storage unit, and a communication interface, each of which is connected by an internal bus.

[0034] According to at least one embodiment, the control unit is composed of a CPU and a ROM, executes a program stored in the storage unit, and controls the server device. The control unit also has an internal timer that measures time. The RAM is the work area of ​​the control unit. The storage unit is a memory area for saving programs and data. The control unit reads out the programs and data from the RAM, and performs program execution processing based on information received from the user terminal, etc.

[0035] We will explain AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large-scale language model, LLM, foundation model, and generative AI. Generative AI uses transformers and employs a number of mechanisms called attention. Self-supervised learning and Extract Prediction are used. In this case, the AI ​​can guess the next word. Given a sentence, it guesses the next word from the sentence up to the middle. A large number of supervised learning problems are created. These create an AI that can guess the next word. Generative AI can predict grammatical structure, topic connections, and that people with this style of writing are likely to write this kind of sentence. Furthermore, generative AI can learn the structure, causal relationships, and knowledge behind the next sentence just by guessing it. Generative AI has scalability, and the larger the number of parameters, the higher the accuracy. In ordinary statistics and machine learning, if the model parameters are made too large compared to the sample size of the data, it will overfit. In LLM, the accuracy increases the larger the number of parameters. One generative AI has 175 billion parameters, and the generative AI overlays supervised learning to make the dialogue smooth. They are instructed not to say anything strange. They write reviews and act as call center operators.

[0036] According to at least one embodiment, large language models (LLMs) are machine learning natural language processing models built non-comprehensively using large datasets and deep learning techniques. In general, a technique called "fine tuning" is used to train on a specific task, and they are adapted to various natural language processing (NLP) tasks such as text classification / generation, sentiment analysis, text summarization, and question answering. According to at least one embodiment, self-supervised learning is close to the intrinsic intelligence of humans. When humans behave, they always predict the next event and predict the next input. In the process, they can learn the structure of the outside world. Predicting the next word is considered to be an intrinsic intelligence and is close to what the cerebral cortex does. According to at least one embodiment, large language models memorize all the information that is input, but generalize it to the extent necessary to predict the next word. They do not generalize all the information from the beginning. Large language models require capacity to memorize information. In addition, parameters are required for this purpose. According to at least one embodiment, the large-scale language model includes eight models with 175 billion parameters and 220 billion parameters.

[0037] According to at least one embodiment, we represent videos and images as a collection of visual patches, which are small units of data similar to text tokens in LLMs. Patches effectively represent models of visual data and serve as a highly scalable and effective representation for training generative models on a wide variety of videos and images. We convert videos to patches by first compressing them into a low-dimensional latent space and then decomposing the representation into spatio-temporal patches.

[0038] According to at least one embodiment, a video compression network is a network that reduces the dimensionality of visual data, taking raw video as input and outputting a temporally and spatially compressed latent representation. An AI is trained on this compressed latent space and then generates video in this compressed latent space.

[0039] According to at least one embodiment, Spacetime Latent Patches extracts a set of spacetime patches that act as Transformer tokens given a compressed input video. The patch-based representation allows Sora to be trained on videos and images of different resolutions, lengths, and aspect ratios, and controls the size of the generated videos by placing randomly initialized patches on an appropriately sized grid at inference time.

[0040] According to at least one embodiment, the AI ​​is a diffusion model, trained to predict the original "clean" patch when fed a noisy patch (and conditioning information such as a text prompt). The AI ​​is a diffusion transformer, which shows remarkable scaling properties in a variety of domains, including language modeling, computer vision, and image generation. Diffusion transformers are also effective as video generation models. The AI ​​significantly improves the quality of the samples as the training computational effort increases.

[0041] According to at least one embodiment, the AI ​​applies caption regeneration techniques to train a highly descriptive caption model, which is then used to generate text captions for all videos in the training set. Training on highly descriptive captions improves the fidelity of the text as well as the overall quality of the generated videos. GPT is leveraged to convert short user prompts into long, detailed captions that are sent to the model. This allows the AI ​​to generate high-quality videos that accurately follow the user's prompts.

[0042] According to at least one embodiment, the AI ​​can perform vectorization in natural language processing along the following flow. First, a cleaning process is performed on the given text as preprocessing. In the cleaning process, unnecessary words such as JavaScript code and HTML tags contained in the text are deleted. These codes are used to display on the Internet, and are therefore not generally used in natural language processing. Next, the text is divided into words by morphological analysis. Morphological analysis is the classification of natural language sentences written in characters into the smallest meaningful linguistic units. As morphological analysis tools, "MeCab", "JUMAN", and "JANOME" can be used. In normalization, words with the same meaning, such as spelling variations, are unified into one word. Stop words are words that are not processed because they cannot be used in natural language processing. Examples of stop words include words that have no meaning by themselves, such as particles and auxiliary verbs. When calculating vectors, these may be removed and only meaningful words may be targeted. Vectorization may also be performed without removing these stop words. Vectorization is a process that converts words, which are strings of characters, into vectors. Vectorization converts word data into numerical data. When converting words into vectors, methods called bag of words or distributed representations are used. Bag of words is a method of vectorizing a sentence using the number of occurrences of words that appear in a given sentence. Since it focuses on how many words appear in a sentence, it does not take into account the order of words or sentences. Distributed representation is a method of vectorization that focuses on the meaning of words. By vectorizing the meaning of words, it is possible to give vectors that are close to words with similar meanings and usages, and the relationship between words can also be expressed as vectors. The expression in vectors makes it possible to add and subtract the meanings of words. In applied processing, natural language converted into numerical data can be used as input for machine learning. Specifically, vectorized natural language is input into a classifier to classify sentences.Tools used here include "TensorFlow," "scikit-learn," and "PyTorch."

[0043] A step of acquiring a user's voice is disclosed. In at least one embodiment, the device can be implemented by any one or more machines or computers in the present disclosure. A fixed terminal or a mobile terminal acquires voice from a user. As an example, the fixed terminal is a store terminal provided in a store, and includes a machine, a computer, and a tablet terminal in the present disclosure. The store terminal is installed in the eating and drinking space in a restaurant, or in the vicinity of a seat or table where a user (in the present disclosure, synonymous with a customer) eats and drinks. The fixed terminal requests the user to voice an order for food and drink. The request may be displayed in text on the display of the fixed terminal, or may be made by a voice to that effect from the fixed terminal. The user voices an order for food and drink to the fixed terminal. The fixed terminal acquires the user's voice information. As another embodiment, as an example, the mobile terminal includes a user terminal, a smartphone or tablet terminal owned by the user, or a terminal on which an application is installed. The user sits at a seat where he or she will eat and drink, and starts an application on the mobile terminal. The application is installed in the mobile terminal in advance. Alternatively, a display showing how to install the application to be installed is displayed near the seat. The user launches the application. Based on the instructions of the application, the mobile terminal requests the user to vocally order food and drink. The request may be made by displaying text on the display of the mobile terminal, or by vocalizing the request from the mobile terminal. The user vocally orders food and drink to the mobile terminal. The mobile terminal acquires the user's voice information.

[0044] The present invention discloses a step of requesting an AI to match the linguistic information indicated by the voice with menu information, which is a menu of a restaurant, and to identify one or more pieces of menu information similar to the linguistic information. In at least one embodiment, a restaurant operator causes a device to store one or more pieces of menu information. The menu information includes information such as linguistic information, images, and videos related to the food and drink menu. As an example, the linguistic information includes linguistic information such as hamburger, cheeseburger, teriyaki burger, tomato, and lettuce, and images and videos corresponding to the linguistic information. Each piece of information can be stored with associative information that evokes related related information or concepts. As an example, for the information "hamburger," information that evokes related or conceptual information such as beef, pork, and bread is stored. This information may be used for AI analysis, as described below.

[0045] The device requests the AI ​​to identify one or more pieces of menu information similar to the language information. In at least one embodiment, the operation of the AI ​​can take any of the above-mentioned embodiments. As an example, the AI ​​exists in one or more of a distributed Internet service system, a fixed terminal, a mobile terminal, an application, and a server. The device provides the language information acquired by the terminal to the AI. The AI ​​acquires the stored menu information. The AI ​​performs vectorization on the language information. Bag of Words can be used for the language information. On the other hand, distributed representation may be more suitable for analyzing user utterances. The AI ​​also performs vectorization on the menu information. The AI ​​compares the vector of the language information with the vector of the menu information, identifies menu information having a vector close to the vector of the language information, and identifies the order of high similarity. In another embodiment, the AI ​​performs various natural language processing (NLP) tasks such as text classification, sentiment analysis, and text summarization for each of the language information and the menu information. Through these tasks, the AI ​​compares the classified text of the language information with the classified text of the menu information, and identifies menu information having text that matches the classified text of the language information. Or, even if the text does not match, the AI ​​can identify menu information with highly similar text and determine the order of similarity. If the AI ​​cannot identify menu information similar to the language information, it can generate status information to the effect that no information has been identified. For example, if a user voice-inputs, "I want to eat beef and bread," the AI ​​will identify the menu information "hamburger," and also identify menu information such as "cheeseburger" and "teriyaki burger."

[0046] A step of transmitting the identified menu information to a chef terminal is disclosed. In at least one embodiment, a computer or terminal that is a chef terminal is provided in the location of the restaurant chef or in a location close to the kitchen. The terminal may be a fixed terminal or a mobile terminal. In the case of a mobile terminal, it may be a terminal that the chef can wear. The device transmits the menu information identified by the AI, or the menu information that is ranked highest in order of similarity, to the chef terminal. The chef terminal communicates the received menu information to the chef. The communication method is to display the information on the display of the chef terminal or to output the menu information by voice. After the menu information is output, information regarding the process and recipe for making the menu can be output together with the menu information.

[0047] The method discloses a step of acquiring confirmed menu information for confirming an order from among the identified menu information. The device transmits the menu information identified by the AI, or multiple menu information in order of high similarity, to a user terminal. The user terminal presents this information to the user and receives an instruction to order these menus. The method of presenting the information includes displaying the menu information on the display of the user terminal, or outputting the menu information by voice. The method of receiving the instruction to order includes displaying a button for confirming the order on the display of the user terminal, and determining that the instruction to order is an instruction to order when the button is pressed. Alternatively, the user terminal acquires voice information from the user. The device requests the AI ​​to perform natural language processing on the voice information. The AI ​​may vectorize the voice information, confirm that the vector indicating an order or the vector information having the intention or emotion is similar, and determine that the instruction is an instruction to order. The device acquires the menu information for which the instruction to order has been received as confirmed menu information. Acquiring may include a method of generating new information as confirmed menu information, or a method of adding status information of "confirmed" to the menu information. As an example, the user terminal presents menu information such as "hamburger," "cheeseburger," and "teriyaki burger," and the device obtains the menu information selected by the user as final menu information.

[0048] A step of transmitting the confirmed menu information to a cook terminal is disclosed. In at least one embodiment, the device transmits the confirmed menu information to a cook terminal. The cook terminal communicates the received confirmed menu information to the cook. The communication method may be to display information on a display of the cook terminal or to output the confirmed menu information by voice. After the confirmed menu information is output, information regarding the process or recipe for making the menu may be output.

[0049] The present invention discloses a step of requesting an AI to match the language information indicated by the voice with arrangement information corresponding to the menu information and identify one or more arrangement information similar to the language information. In at least one embodiment, an operator of a restaurant stores one or more arrangement information in the device. The arrangement information includes information such as language information, images, and videos accompanying the menu information. As an example, the arrangement information is language information such as the amount of lemon sauce, the presence or absence of tomatoes, the presence or absence of lettuce, and one selected from multiple ways of grilling the patty for the menu information of a hamburger. The arrangement information also includes images and videos corresponding to the language information. Each piece of information can be stored with associative information that evokes related information or concepts. As an example, information that evokes related or conceptual information such as sour is stored for information on lemon sauce, and information on tomatoes is stored with information on healthy and vitamins. These pieces of information may be used for analysis by the AI, as described later.

[0050] The device requests the AI ​​to identify one or more arrangement information similar to the language information. In at least one embodiment, the device provides the AI ​​with the language information acquired by the terminal. The AI ​​acquires the stored arrangement information. The AI ​​performs vectorization on the language information. Bag of Words can be used for the language information. On the other hand, distributed representation may be more suitable for analyzing user utterances. The AI ​​also performs vectorization on the arrangement information. The AI ​​compares the vector of the language information with the vector of the arrangement information, identifies arrangement information having a vector close to the vector of the language information, and identifies the order of high similarity. In another embodiment, the AI ​​performs various natural language processing (NLP) tasks such as text classification, sentiment analysis, and text summarization for each of the language information and the arrangement information. Through these tasks, the AI ​​compares the classified text of the language information with the classified text of the arrangement information, and identifies menu information having text that matches the classified text of the language information. Alternatively, even if the text does not match, it identifies arrangement information having text with a high degree of similarity, and identifies the order of high similarity. If the AI ​​cannot identify any composition information similar to the linguistic information, it can generate status information to that effect. For example, if a user voice-inputs, "Sour is good. Healthy food is good," the AI ​​will identify the composition information "more lemon sauce" and "tomato."

[0051] A step of transmitting the determined menu information and arrangement information to a cook terminal is disclosed. In at least one embodiment, the device transmits the determined menu information and arrangement information to a cook terminal. The cook terminal conveys the received determined menu information and arrangement information to the cook. The conveying method may be to display information on a display of the cook terminal or to output the confirmed menu information by voice. After outputting the arrangement information, information on the process or recipe for making the menu may be output together with the information.

[0052] The step of acquiring confirmed arrangement information for confirming an order from among the identified arrangement information is disclosed. The device transmits the arrangement information identified by the AI, or multiple arrangement information in order of high similarity, to a user terminal. The user terminal presents this information to the user and receives an instruction to order these arrangements. The method of presenting the information includes displaying the arrangement information on the display of the user terminal, or outputting the arrangement information by voice. The method of receiving the instruction to order includes displaying a button for confirming the order on the display of the user terminal, and determining that the instruction to order is an instruction to order when the button is pressed. Alternatively, the user terminal acquires voice information from the user. The device requests the AI ​​to perform natural language processing on the voice information. The AI ​​may vectorize the voice information, confirm that the vector indicating an order or the vector information having the intention or emotion is similar, and determine that the instruction is an instruction to order. The device acquires the arrangement information for which the instruction to order has been received as confirmed arrangement information. Acquiring may include a method of generating new information as confirmed arrangement information, or a method of adding status information of "confirmed" to the arrangement information. As an example, the user terminal presents arrangement information such as "more lemon sauce" and "tomato," and the device obtains the arrangement information selected by the user as final arrangement information.

[0053] A step of transmitting the confirmed menu information and the confirmed arrangement information to a cook terminal is disclosed. In at least one embodiment, the device transmits the confirmed menu information and the confirmed arrangement information to the cook terminal. The cook terminal conveys the received confirmed menu information and the confirmed arrangement information to the cook. The conveying method may be to display the information on the display of the cook terminal or to output the confirmed menu information and the confirmed arrangement information by voice. After the arrangement information is output, information on the process or recipe for making the menu may be output together with the information.

[0054] The present invention discloses a step of instructing a self-propelled cart to self-propel to a predetermined location by approximately the time corresponding to the menu information or arrangement information, or the confirmed menu information or confirmed arrangement information. In at least one embodiment, the operator of a restaurant stores the required time corresponding to the menu information or arrangement information in the device. The required time can be set arbitrarily by the operator. For example, it is 20 minutes for menu information such as a hamburger. Examples of the required time for arrangement information include 5 minutes if the patty is cooked well done, 3 minutes if it is medium, and 1 minute for adding tomatoes. The required time corresponding to the confirmed menu information or confirmed arrangement information is the required time corresponding to each menu information or arrangement information.

[0055] The device generates order time information based on the time when the menu information or arrangement information, or the confirmed menu information or the confirmed arrangement information is acquired. The device adds the corresponding required time to the order time information to generate estimated completion time information.

[0056] A self-propelled cart is disclosed. In at least one embodiment, the self-propelled cart is an unmanned vehicle or machine that moves by itself. The self-propelled cart has one or more spaces for placing food. The self-propelled cart can have a computer or terminal in the present disclosure. That is, the self-propelled cart can wirelessly communicate with the device. In at least one embodiment, when an object is observed, if the vehicle or machine is unmanned and moves by itself, it is considered to be a "self-propelled cart" in the present disclosure. The device instructs the self-propelled cart to move to a specified location at approximately the expected completion time. As a method of instructing the self-propelled cart, the device wirelessly transmits a program or instruction to self-propel the self-propelled cart to a specified location at approximately the expected completion time. The self-propelled cart receives the program or instruction and moves to the specified location at approximately the expected completion time. Examples of the specified location include a place where a chef is present, a kitchen, and a place where the completed food is handed over.

[0057] A step of adding up numerical values ​​corresponding to menu information or arrangement information, or confirmed menu information or confirmed arrangement information is disclosed. In at least one embodiment, a restaurant operator causes a device to store numerical values ​​corresponding to menu information or arrangement information. The numerical values ​​represent at least the value of these menus or arrangements. For example, 1000 for a hamburger and 50 for extra lemon sauce. That is, a hamburger is 1000 yen and lemon sauce is 50 yen. The device adds up the numerical values ​​corresponding to the menu information or arrangement information, or confirmed menu information or confirmed arrangement information. For example, if a hamburger is arranged with extra lemon sauce, the total value becomes 1050.

[0058] A step of presenting the total number to a user is disclosed. In at least one embodiment, the device transmits the total number to a user terminal. In another embodiment, the device transmits the total number to an accounting device through which the user makes a payment. Examples of accounting devices include accounting devices in restaurants, cash registers, applications for user payments, applications for user electronic payments, user terminals (including smartphones), etc.

[0059] In at least one embodiment, any one of the above-mentioned embodiments has industrial applicability and advantages in the following cases: restaurants, drive-throughs, etc. In the case of a drive-through, the user terminal in the step of acquiring the user's voice may correspond to a machine installed outdoors for the drive-through. In another embodiment, in the case of a drive-through, a device owned by a user, such as a smartphone or computer, or such a device with a dedicated application installed may correspond to the user terminal.

[0060] The following provides an overview of the above-described embodiment.

[0061] 1. A method of ordering, comprising: acquiring a speech signal from a user; A step of requesting the AI ​​to match the linguistic information indicated by the voice with menu information that is a menu of a restaurant, and to identify one or more pieces of menu information that are similar to the linguistic information; transmitting the identified menu information to a chef terminal; Method for having

[0062] According to the present disclosure, at least, there is the convenience of allowing a user to order food and drink by speaking. Furthermore, there is industrial applicability and advantage in that a restaurant can receive orders based on the user's speech without any attendant.

[0063] 1. A method of ordering, comprising: acquiring a speech signal from a user; A step of requesting the AI ​​to match the linguistic information indicated by the voice with menu information that is a menu of a restaurant, and to identify one or more pieces of menu information that are similar to the linguistic information; acquiring confirmed menu information for confirming the order from among the identified menu information; transmitting the confirmed menu information to a chef terminal; Method for having

[0064] According to the present disclosure, at least, there is the convenience of allowing a user to order food and drink by speaking. Furthermore, there is industrial applicability and advantage in that a restaurant can receive orders based on the user's speech without any attendant.

[0065] 1. A method of ordering, comprising: acquiring a speech signal from a user; A step of requesting the AI ​​to match the linguistic information indicated by the voice with menu information that is a menu of a restaurant, and to identify one or more pieces of menu information that are similar to the linguistic information; A step of requesting an AI to match the language information indicated by the voice with arrangement information corresponding to the menu information and identify one or more arrangement information similar to the language information; Transmitting the specified menu information and arrangement information to a chef terminal; Method for having

[0066] According to the present disclosure, at least, there is the convenience of allowing a user to order food and drink by speaking. Furthermore, there is industrial applicability and advantage in that a restaurant can receive orders based on the user's speech without any attendant.

[0067] 1. A method of ordering, comprising: acquiring a speech signal from a user; A step of requesting the AI ​​to match the linguistic information indicated by the voice with menu information that is a menu of a restaurant, and to identify one or more pieces of menu information that are similar to the linguistic information; A step of requesting an AI to match the language information indicated by the voice with arrangement information corresponding to the menu information and identify one or more arrangement information similar to the language information; acquiring confirmed menu information for confirming the order from among the identified menu information; obtaining final arrangement information for finalizing the order from among the identified arrangement information; A step of transmitting the confirmed menu information and the confirmed arrangement information to a chef terminal; Method for having

[0068] According to the present disclosure, at least, there is the convenience of allowing a user to order food and drink by speaking. Furthermore, there is industrial applicability and advantage in that a restaurant can receive orders based on the user's speech without any attendant.

[0069] Any of the above ordering methods, a step of instructing the self-propelled cart to move to a predetermined location by approximately the time when a required time corresponding to the menu information or the arrangement information has elapsed; Further comprising the method

[0070] According to the present disclosure, at least the food and drink ordered by the user can be quickly transported by a self-propelled cart, which has industrial applicability and advantages in that freshly cooked food can be delivered to the user unmanned.

[0071] Any of the above ordering methods, A step of adding up numerical values ​​corresponding to the menu information or arrangement information; presenting the combined number to a user; Further comprising the method

[0072] The present disclosure has industrial applicability and advantages, at least in that it allows unmanned accounting for food and beverage orders by users.

[0073] In at least one embodiment, in the "step of acquiring confirmed menu information for confirming an order from among the identified menu information", the "method of receiving an instruction to order" by the device includes the following forms: The restaurant operator defines motion information related to one or more physical motions in the device and stores it as specific motion information. The motion information includes information such as language information, images, and videos that accompany the physical motion. For example, for the physical motion of "tapping the table twice with a finger", the information includes the language information, and information such as an image or video showing an example of the motion. Physical motions include all kinds of gestures such as raising a hand and nodding. The specific motion information is information that the restaurant operator registers in the device from among the motion information. The user terminal acquires the user's motion information. One method of acquisition is to acquire video or video of the user's motion with a camera of the user terminal.

[0074] The device requests the AI ​​to determine whether the motion information is similar to the specific motion information. The AI ​​represents the acquired motion information as a collection of visual patch patches, which are small data units similar to the text tokens of LLM. The AI ​​is a network that reduces the dimension of visual data about the motion information, receiving raw video as input and outputting a temporally and spatially compressed latent representation. The AI ​​also outputs a latent representation for the specific motion information in a similar manner to the above. The AI ​​compares the latent representation of the motion information with the latent representation of the specific motion information and determines whether there is a similarity. Regarding this judgment, the restaurant operator can store in the device a similarity threshold value at which it is determined that there is a similarity. If the AI ​​determines that the motion information is similar to the specific motion information, it can create data or status information indicating the similarity.

[0075] The device acquires the menu information instructed to be ordered as finalized menu information, provided that data or status information indicating similarity has been generated. Acquiring may involve generating new information as finalized menu information, or adding status information indicating finalization to the menu information. As an example, the device presents "hamburger" to the user as menu information. The user taps his / her finger on the desk twice. The device determines that the user's motion information is specific motion information, and acquires "hamburger" as finalized menu information.

[0076] In at least one embodiment, in the "step of acquiring finalized arrangement information for confirming the order from among the specified arrangement information," the device can adopt the above embodiment as the "method of receiving an instruction to order." That is, the restaurant operator determines motion information related to one or more physical motions in the device and stores it as specific motion information. The user terminal acquires the user's motion information. The AI ​​judges whether the motion information is similar to the finalized motion information. If data or status information indicating similarity is created, the device acquires the arrangement information for which the order was instructed as the finalized arrangement information. The acquisition may be a method of generating new information as the finalized arrangement information, or a method of adding status information indicating that the arrangement information is finalized. As an example, the device presents "more lemon sauce" to the user as the arrangement information. The user nods. The device determines that the user's motion information is specific motion information, and acquires "more lemon sauce" as the finalized arrangement information.

[0077] In at least one embodiment, a separate device or terminal for detecting user motion information may be provided, either separately from the user terminal or in combination with the user terminal in the above embodiments.

[0078] In at least one embodiment, instead of the action information and the specific action information in the above embodiment, the language information and the specific language information may be used as the information to be determined. That is, the operator of the restaurant determines the language information corresponding to one or more passwords in the device and stores it as the specific language information. The user terminal acquires the user's language information. The acquisition method may be a method using a microphone of the device or the terminal. The AI ​​determines whether the language information is similar to the specific action information by natural language processing. The device requests the AI ​​to make the judgment. As already described, the natural language processing method may be any one or more methods in the present disclosure. When data or status information indicating that the language information is similar to the specific language information is created by the AI ​​or the device, the device acquires the menu information or arrangement information instructed to order as the confirmed menu information or the confirmed arrangement information. The acquisition may be a method of generating new information as the confirmed arrangement information or adding status information of confirmation to the arrangement information. As an example, the device presents "more lemon sauce" to the user as the arrangement information. The user speaks a password determined by the restaurant operator (examples include the restaurant's trademark, emblem, brand name, nickname, or special call). The device determines that the user's language information is specific language information, and acquires "extra lemon sauce" as the final arrangement information.

[0079] The following provides an overview of the above-described embodiment.

[0080] Any of the above ordering methods, The step of acquiring confirmed menu information for confirming the order includes: storing, by the device, specific motion information relating to one or more physical motions; Requesting the AI ​​to determine whether the motion information is similar to specific motion information; Further includes:

[0081] Any of the above ordering methods, The step of acquiring confirmed menu information for confirming the order includes: When the AI ​​determines that the motion information is similar to the specific motion information, the menu information is set as finalized menu information; Further includes:

[0082] According to the present disclosure, at least, a user can confirm an order by a gesture, which is easier than the conventional technique of touching a touch panel, and thus has industrial applicability and advantages in that it improves the convenience of using a device or system.

[0083] It is sufficient for the invention according to the present disclosure to achieve at least one of the effects described above.

Claims

1. 1. A method of ordering, comprising: storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify two or more pieces of menu information that are similar to the linguistic information; Present the two or more pieces of specified menu information to a user terminal; acquiring finalized menu information for finalizing the order from among the two or more pieces of identified menu information; Send the finalized menu information to the chef's terminal. method

2. 1. A method of ordering, comprising: Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify two or more pieces of menu information that are similar to the linguistic information; Present the two or more pieces of specified menu information to a user terminal; acquiring finalized menu information for finalizing the order from among the two or more pieces of identified menu information; Send the finalized menu information to the chef's terminal. method

3. 1. A method of ordering, comprising: storing or retrieving related information or associative information that is associated with the arrangement information corresponding to the menu of the restaurant; Get the user's language information, Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Among the identified arrangement information, obtain final arrangement information that finalizes the order; Send the final arrangement information to the chef's terminal. method.

4. 1. A method of ordering, comprising: storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtain finalized menu information for finalizing the order from among the identified menu information; Send the finalized menu information to the chef's terminal. method

5. 1. A method of ordering, comprising: Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtain finalized menu information for finalizing the order from among the identified menu information; Send the finalized menu information to the chef's terminal. method

6. An ordering device, comprising: storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify two or more pieces of menu information that are similar to the linguistic information; Present the two or more pieces of specified menu information to a user terminal; acquiring finalized menu information for finalizing the order from among the two or more pieces of identified menu information; Send the finalized menu information to the chef's terminal. Device

7. An ordering device, Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify two or more pieces of menu information that are similar to the linguistic information; Present the two or more pieces of specified menu information to a user terminal; acquiring finalized menu information for finalizing the order from among the two or more pieces of identified menu information; Send the finalized menu information to the chef's terminal. Device

8. An ordering device, storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtain finalized menu information for finalizing the order from among the identified menu information; Send the finalized menu information to the chef's terminal. Device

9. An ordering device, Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtain finalized menu information for finalizing the order from among the identified menu information; Send the finalized menu information to the chef's terminal. Device 10. A method of ordering, comprising: storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtaining finalized menu information for finalizing the order from among the identified menu information; method 11. A method of ordering, comprising: Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtaining finalized menu information for finalizing the order from among the identified menu information; method

12. An ordering device, comprising: storing associated information or concept-related information related to menu information of the restaurant; Get language information from the user, Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtaining finalized menu information for finalizing the order from among the identified menu information; Device

13. An ordering device, comprising: Get language information from the user, Searching for related information or associated information that is associated with the restaurant menu information; Requesting the AI ​​to compare the linguistic information with related or associative information and identify one or more menu information similar to the linguistic information; Present the identified menu information to the user terminal; Obtaining finalized menu information for finalizing the order from among the identified menu information; Device

14. An ordering device, comprising: storing or retrieving related information or associative information that is associated with the arrangement information corresponding to the menu of the restaurant; Get the user's language information, Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Among the identified arrangement information, obtain final arrangement information that finalizes the order; Send the final arrangement information to the chef's terminal. Device.

15. A method of ordering, comprising: storing or retrieving related information or associative information that is associated with the arrangement information corresponding to the menu of the restaurant; Get the user's language information, Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Obtaining final arrangement information for finalizing the order from among the identified arrangement information; method.

16. An ordering device, comprising: storing or retrieving related information or associative information that is associated with the arrangement information corresponding to the menu of the restaurant; Get the user's language information, Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Obtaining final arrangement information for finalizing the order from among the identified arrangement information; Device.

17. A method according to any one of claims 1, 2, 4, 5, 10 and 11, comprising: The device stores or retrieves the duration corresponding to the menu information. method 18. A method according to any one of claims 1, 2, 4, 5, 10 and 11, comprising: The apparatus further comprises: generating estimated completion time information by adding a required time corresponding to the confirmed menu information to the time when the confirmed menu information was acquired; method 19. An apparatus as claimed in any one of claims 6 to 9, 12 and 13, comprising: The device stores or retrieves the duration corresponding to the menu information. Device 20. An apparatus as claimed in any one of claims 6 to 9, 12 and 13, comprising: The apparatus further comprises: generating estimated completion time information by adding a required time corresponding to the confirmed menu information to the time when the confirmed menu information was acquired; Device 21. A method according to any one of claims 1 to 5, 10, 11 and 15, comprising: The linguistic information obtained from the user is text. method 22. A method according to any one of claims 1, 2, 4, 5, 10 and 11, comprising: The apparatus further comprises: storing associated information or concept-related information related to the arrangement information corresponding to the menu information; Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Among the identified arrangement information, obtain final arrangement information that finalizes the order; Send the finalized menu information to the chef's terminal. method 23. A method according to any one of claims 1, 2, 4, 5, 10 and 11, comprising: The apparatus further comprises: Searching for related information or associated information that is associated with a concept related to the arrangement information corresponding to the menu information; Requesting the AI ​​to compare the linguistic information with related or associated information of arrangement information and identify one or more arrangement information similar to the linguistic information; Present the identified arrangement information on the user terminal; Among the identified arrangement information, obtain final arrangement information that finalizes the order; Send the finalized menu information to the chef's terminal. method 24. An apparatus as claimed in any one of claims 6 to 9, 12, 13, 14 and 16, comprising: The linguistic information obtained from the user is text. Device 25. A method according to any one of claims 1, 2, 4, 5, 10 and 11, comprising: instructing the self-propelled cart to move to a predetermined location by the time when the required time corresponding to the menu information has elapsed; method

Citation Information

Patent Citations

  • Seat managing pos system

    JP1996161403A

  • Order receiving device and order receiving method for restaurant

    JP2005182140A

  • Management apparatus, management method, and management program

    JP2019159378A

  • On-demand coordinated food item delivery system

    JP2021500684A

  • Order reception apparatus, order reception system, and program

    JP2022190585A