How to order
The ordering method uses voice recognition and AI to automate food and drink orders, addressing the lack of unattended ordering in existing systems and enhancing efficiency and convenience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2026-04-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There is no method for taking food and drink orders without a person in existing systems.
An ordering method that utilizes voice recognition to acquire user input, compares it with menu information using AI, and sends the identified menu items to a chef's terminal for preparation.
Enables unattended food and drink ordering, improving efficiency and convenience by automating the ordering process.
Abstract
Description
Technical Field
[0001] The present invention relates to an ordering method.
Background Art
[0002] The statements in this section only provide background information related to the present disclosure and do not necessarily constitute prior art. not necessarily.
[0003] Patent Document 1 discloses a passenger seat management POS system characterized by including monitoring means such as a monitoring camera for monitoring passenger seats, displaying a video signal from the monitoring means for monitoring passenger seats on a display unit, and being able to confirm the situation of passenger seats. from the monitoring means for monitoring the passenger seats to the display unit, and being able to confirm the situation of the passenger seats. is disclosed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the inventor recognized that at least in the above-described embodiment, there is a drawback that there is no method for taking orders for food and drinks without a person. exists.
Means for Solving the Problems
[0006] At least one disclosure is an ordering method, wherein the apparatus acquires the voice of a user; requests an AI to collate the language information indicated by the voice with menu information that is the menu of a restaurant and identify one or more menu information that is similar to the language information; and The steps include: sending the identified menu information to the chef's terminal, A method that has To provide. [Effects of the Invention]
[0007] This configuration has the advantage of being able to take food and drink orders unattended. ru.
[0008] These and other aspects, features, and advantages of this disclosure are best taken in conjunction with the following drawings. As will become clear from the following detailed written description of the appropriate embodiments and aspects, their modifications and The modifications may be made without departing from the spirit and scope of the novel concepts of this disclosure. Aspects in one embodiment may, to the extent that they do not contradict, be other embodiments disclosed herein. It can be combined with or replaced with one or more of the embodiments in [the present invention]. [Modes for carrying out the invention]
[0009] In the following disclosures, there are many different implementations to carry out the different characteristics of the presented subject matter. This disclosure provides forms and examples. To simplify this disclosure, specific examples of components and their arrangement are given below. Disclosed. Of course, these are merely examples and are not intended to be limiting. For example, the first feature is covered by or adjacent to the second feature which is subsequently disclosed. The structure is formed such that the first feature and the second feature are in direct contact with each other. In addition, an additional feature is formed between the first feature and the second feature, and the first feature and the second feature Embodiments may include those in which the features do not come into direct contact. Furthermore, this disclosure includes, In various examples, reference numbers and / or characters may be repeated. The repetition is for the purpose of simplicity and clarity and is not required to have a relationship with various embodiments and / or the configurations described. Further, when the first element is described as "connected" or "coupled" to the second element , such description includes embodiments in which the first element and the second element are directly connected or coupled to each other and also includes embodiments in which the first element and the second element are indirectly connected or coupled to each other with one or more other elements intervening therebetween.
[0010] As used herein, the description "at least one of" encompasses all possible variations for illustration. For example, the description "at least one of A, B, and C (compris es at least one of A, B, or C)" is synonymous with "A, B, C, and combinations thereof (consisting of A, B, C and combinations thereof)". And it encompasses all possible variations of A, B, C, A + B, A + C, B + C, A + B + C.
[0011] In the present disclosure, the disclosure of using an electronic operator or a computer can include embodiments of a method, a recording medium , an apparatus, or a program. The description "A is B" used herein can be replaced with "A includes B" as long as there is no contradiction or unless otherwise stated in this specification .
[0012] Regarding the operation method used in at least one or more embodiments, the following embodiments can be taken . The description of JP6456303 that well explains at least one or more embodiments will be cited and explained (hereinafter, the citation begins).
[0013] As used herein, the term "computer" is known in the art to include, for example, a processor, a memory such as a hard drive, a disk drive or a flash drive or a memory stick, or other non - transient computer - readable media or non - transient storage devices, at least one information storage / search device such as a keyboard mouse, a pointing and touch device, a touch screen, or a microphone, at least one input device, and a display structure such as a well - known computer screen. Additionally, a computer may include one or more network connections such as a wired or wireless connection. As is known in the art, such a computer or computer system may include, more or less, the items listed above and is not limited, for example, to tablet computers or smart devices, but includes other electronic media and electronic devices.
[0014] As used herein, the term "cloud" or "cloud computing" refers to a centralized and virtualized computing facility where all computing resources are shared. For application systems and subsystems, since they are all within the "cloud", it is no longer possible to refer to a specific machine.
[0015] As used herein, the term "distributed Internet service system" To run internet applications in various computing environments, It refers to a decentralized internet service platform that replaces existing systems. The DIS system is a decentralized internet service platform. Component Distribution Server er) / Asset Distribution Server Through this, internet applications including content, data, and logic are to whatever extent appropriate It distributes to any number and any type of device along the network via DIS. And, internet applications are provided as services based on the needs of each user. Host and centrally manage it, maintaining its integrity while controlling the user's device or nearby It can be cached and executed locally in a specific location. Web-enabled computing data Vice can be upgraded with DIS software to provide decentralized internet services. It can become a DIS-compatible service that allows you to enjoy and execute tasks. The system is registered under US Patent Nos. 7136857, 7150015, and 7181731. , No. 7209921, No. 7430610, No. 7685183, No. 7685577 , No. 7752214, No. 8326883, No. 8386525, No. 8443035 , No. 8458142, No. 8458222, No. 8473468, No. 8527545 , and Patent No. 8650226, and U.S. Patent Publication No. 20120005205, This is fully described in any one of the patent family patents No. 20130091252, All of these, like the present invention, are shared by OP40 Holdings, Inc. It is owned by [company name], and all of these are incorporated by quotation. (End of quotation)
[0016] The operating method used in at least one embodiment utilizes a distributed internet. Regarding conventional internet methods that do not involve this, the following embodiments can be adopted. I will explain by referencing the description in JP7113047, which describes at least one embodiment in detail (hereinafter, start of use).
[0017] Embodiments including those specifically disclosed herein are based on artificial intelligence and actually involve humans We can provide an automated response system that is implemented in a way that resembles a conversation with the user, and by doing so, This enables more natural conversations with users while quickly handling inquiries, reservations, and delivery orders. Furthermore, it can be processed conveniently.
[0018] Multiple electronic devices 110, 120, 130, and 140 are implemented by a computer system. These may be fixed terminals or mobile terminals. Multiple electronic devices 110, 120, 130, 1 40 examples include AI speakers, smartphones, mobile phones, navigation systems, PCs ( personal computers, notebook PCs, digital broadcasting terminals, PDAs ( Personal Digital Assistant), PMP (Portable Multimedia Player, tablets, game consoles, wearables VR devices, IoT (Internet of Things) devices, VR (virtual reality) devices. AR (augmented reality) devices Examples include vices, etc. As an example, Figure 1 shows an AI speaker as electronic device 110. However, in embodiments of the present invention, the electronic device 110 is substantially wireless or wired. Using the formula, other electronic devices 120, 130, 140 and / Alternatively, various physical computer systems capable of communicating with servers 150 and 160 It can mean one of the "mu"s.
[0019] The communication method is not limited, and the network 170 can include any communication network (1 Examples include mobile communication networks, wired internet, wireless internet, broadcasting networks, and satellite networks. This may include not only communication methods that utilize ), but also short-range wireless communication between devices. Network 170 is a PAN (personal area network), L AN(local area network), CAN(campus area network) network), MAN(metropolitan area network), W AN(wide area network), BBN(broadband network) (ork), one or more arbitrary networks such as the Internet May include. Furthermore, network 170 includes bus networks, star networks, Ring network, mesh network, starbus network, tree or floor It may include any one or more network topologies, including layered networks. However, it is not limited to these.
[0020] Servers 150 and 160 are connected to multiple electronic devices 110, 120, 130, and 140, respectively. It communicates via network 170 to exchange commands, codes, files, content, and services. This may be implemented by one or more computer devices that provide services such as servers. Unit 150 is connected via network 170 to multiple electronic devices 110, 120, 13 0, 140 may be a system that provides the first service, and server 160 may also be a network Multiple electronic devices 110, 120, 130, and 140 connected via the 170 are second-party It may be a system that provides screws. As a more specific example, server 150 may be multiple Computers installed and run on electronic devices 110, 120, 130, and 140 Through the application, which is a computer program, the application aims to achieve its objectives. A service (for example, an automated response service) is designated as the first service by multiple electronic devices 1 It may be provided to 10, 120, 130, and 140. As another example, server 160 is as described above. The files for installing and running the application are on multiple electronic devices 110 The service distributed to 120, 130, and 140 may be offered as a second service.
[0021] Figure 2 illustrates the internal configuration of an electronic device and a server in one embodiment of the present invention. This is a block diagram. Figure 2 shows the internal configuration of electronic device 110 as an example of an electronic device. The internal configuration of server 150 will also be described. Furthermore, other electronic devices 120, 130, 140 and Server 160 are identical or similar to the aforementioned electronic equipment 110 or Server 150. It may have an internal structure.
[0022] Electronic equipment 110 and server 150 include memory 211, 221, processor 212, 2 22, communication modules 213, 223, and input / output interfaces 214, 224 May include: Memory 211, 221 is a non-temporary computer-readable recording medium. It can be a body, RAM (random access memory), ROM (re (add only memory), disk drive, SSD (solid state) Non-temporary drives, flash memory, etc. This may include high-capacity storage devices, such as ROM, SSD, flash memory, and disks. Non-temporary high-capacity recording devices such as drives are classified separately from memories 211 and 221. It may be included in electronic equipment 110 or server 150 as a non-temporary recording device. Mori 211 and 221 include an operating system and at least one program code. (For example, a browser installed and run on electronic device 110, An application installed on electronic device 110 for the purpose of providing a specific service. Code for which purpose may be recorded. Such software components may be stored in memory 21. 1,221 may be loaded from a computer-readable storage medium other than 221. Other computer-readable storage media include floppy disk drives, and Computer-readable discs, tapes, DVD / CD-ROM drives, memory cards, etc. It may include a removable recording medium. In other embodiments, the software components are Notes are sent via communication modules 213 and 223, which are not computer-readable recording media. It may be loaded into re211 and 221. For example, at least one program is developed A file distribution system that distributes installation files for a person or application ( For example, the files provided by the server 160 mentioned above via network 170 This is a computer program that is installed (for example, the application mentioned above) Based on (n), it may be loaded into memory 211 and 221.
[0023] Processors 212 and 222 perform basic arithmetic, logic, and input / output operations. This may be configured to process instructions for a computer program. The instructions are, The processor 212 is controlled by memory 211, 221 or communication modules 213, 223. , 222 may be provided. For example, processors 212, 222 may have memory 211, 22 The system executes the received instructions according to the program code recorded in a recording device like the one described in 1. It may be configured in such a way.
[0024] Communication modules 213 and 223 communicate with electronic equipment 110 via network 170. The B150 may provide a function for communication with each other, and the electronic equipment 110 and / or Or server 150 may be other electronic devices (for example, electronic device 120) or other servers ( For example, it may provide a function for communicating with server 160). The processor 212 of the device 110 receives the program code recorded in a recording device such as memory 211. The request generated according to the code is controlled by the communication module 213. It may be transmitted to server 150 via k170. Conversely, the process of server 150 Control signals, instructions, content, files, etc. provided in accordance with the control of sass222 via communication module 223 and network 170, the communication module 2 of electronic device 110 It may be received by the electronic device 110 through 13. For example, through the communication module 213 The received control signals, instructions, content, files, etc. from server 150 are processed by processor 2. The information is transmitted to 12 and memory 211, and the content and files are stored in the electronic device 110. Furthermore, it may be recorded on a recording medium that can contain (the non-temporary recording device described above).
[0025] The input / output interface 214 is for interfacing with the input / output device 215. The means may be any of the following. For example, the input device may be a keyboard, mouse, microphone, or camera. The output devices include displays, speakers, and haptic feedback devices. Any device may be included. Another example is the input / output interface 214. Interface with devices that integrate input and output functions into a single unit, such as a screen. It may also be a means for S. The input / output device 215 is an electronic device 110 and one device It may be configured as follows. Also, the input / output interface 224 of the server 150 is A device for input or output that can be connected to or included by server 150 (illustrated) It may be a means for interfacing with (without). As a more specific example, electronic devices Processor 212 receives the instructions of the computer program loaded into memory 211. In processing this, the system is configured using data provided by the server 150 and the electronic device 120. The service screen and content are displayed via the input / output interface 214. It may be displayed in section I.
[0026] In another embodiment, the electronic device 110 and the server 150 are components of Figure 2. It may include more components than that. However, most of the conventional components are not clearly illustrated. It is not necessary to show this. For example, the electronic device 110 is one of the input / output devices 215 described above. It may be implemented to include at least a part of, transceivers, cameras, various sensors, It may also include other components such as databases. As a more specific example, If electronic device 110 is an AI speaker, then the various sensors that an AI speaker typically includes... Camera module, various physical buttons, buttons using the touch panel, input / output A variety of components such as force ports and vibrators for vibration are further added to the electronic device 110. It may be implemented in such a way as to include it. (End of quote)
[0027] The machine will be described. According to at least one embodiment, the user terminal is a control unit. RAM, storage unit, graphics processing unit, communication interface, interface It consists of several parts, each connected by an internal bus.
[0028] According to at least one embodiment, the control unit consists of a CPU and ROM. The unit executes programs stored in the storage unit and controls user terminals. RA M is the work area of the control unit. The storage unit is used to store programs and data. This is the memory area. The control unit reads programs and data from RAM and performs processing. Now. The control unit processes the program and data loaded into RAM to control the drawing process. Output the command to the graphics processing unit.
[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. The control unit outputs drawing commands to the graphics processing unit. The graphics processing unit then outputs a video signal to display an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel functions as an input unit.
[0030] According to at least one embodiment, the communication interface is wireless or wired. It can connect to a network and send and receive data with server devices via the communication network. It is possible to communicate. Data received via the communication interface is loaded into RAM. The data is then processed by the control unit. The interface unit has external memory (for example, An SD card or similar is connected.
[0031] According to at least one embodiment, the user terminal has a display screen and an input unit. The computer device is not particularly limited. For example, a conventional mobile phone could be used as the user terminal. Telephones, tablet devices, smartphones, desktop and notebook personal computers Examples include computers. VR goggles, which are attached or secured to the head with a strap. A screen (or two displays) mounted on a frame (or headset) It may consist of an audio panel (one for each item). The user terminal has an audio output section. To possess.
[0032] According to at least one embodiment, a user terminal is connected to a server via a communication network. It is possible to establish a communication connection with the device. It establishes a communication connection via a communication network and transmits information. They can either do so or receive information.
[0033] According to at least one embodiment, the server device includes a control unit, RAM, storage unit and It is equipped with at least one communication interface, each connected by an internal bus.
[0034] According to at least one embodiment, the control unit consists of a CPU and ROM, and storage The control unit executes programs stored in the page unit and controls the server device. The control unit also operates on a time-based system. It has an internal timer for timing. RAM is the work area of the control unit. The memory section is a storage area for saving programs and data. The control unit is for saving programs and data. The program reads data from RAM and, based on information received from the user terminal, Perform the execution process.
[0035] This section explains AI. According to at least one embodiment, artificial intelligence is machine learning, deep learning This includes layered learning, generative AI, large-scale language models, LLMs, foundational models, and generative AI. It uses a transformer and employs numerous attention mechanisms. Self-supervised learning, Extrac Use prediction. In this case, the AI can predict the following word. Given a sentence... Then, you guess the next word from the sentence up to that point. This creates a large number of supervised learning problems. This allows for the creation of an AI that can predict the next word. The generative AI considers grammatical structure, topic connections, It is possible to predict what kind of sentences someone with this writing style is likely to write. Furthermore, generation AI can learn the underlying structure, cause and effect, and knowledge simply by guessing the next sentence. Yes. Generative AI scales quickly, and the accuracy increases with a larger number of parameters. Statistics and machine learning often involve setting model parameters large relative to the data sample size. Too many parameters will result in overfitting. With LLM, the accuracy increases as the number of parameters increases. It goes up. One generative AI has 175 billion parameters, and the generative AI can have smooth conversations. Teacher-led learning is superimposed in this way. Students are instructed not to say anything strange. (Referring to a reflection paper) Writing and acting as a call center operator.
[0036] According to at least one embodiment, Large Language Models (LLM) ) is non-comprehensively built using large datasets and deep learning techniques. This refers to a machine learning model for natural language processing. Generally, it is a model that is trained on a specific task. Using a technique called "fine-tuning," which involves fine-tuning, text classification and generation, and sentiment analysis are performed. It can be adapted to various natural language processing (NLP) tasks such as analysis, text summarization, and question answering. According to at least one embodiment, self-supervised learning is close to the intrinsic intelligence of humans. It is constantly predicting the next event that will occur when it takes action, and predicting the next input. In the process, we can learn about the structure of the external world. Predicting the next word is an essential form of intelligence, and the brain... I think it's similar to what happens in the cortex. According to at least one embodiment, large language models It memorizes the input information but generalizes it to the extent necessary to predict the next word. We don't try to generalize all the information from the start. Large-scale language models are designed to remember information. Capacity is required. Parameters are also needed for that. At least one implementation. According to the data, large-scale language models include models with 175 billion parameters and models with 220 billion parameters. It is equipped with eight of them.
[0037] According to at least one embodiment, videos and images are converted into small text tokens similar to those of LLM. It is represented as a collection of small data units called visual patches. A patch is a visual data Effectively represent the data model and train the generative model with various types of videos and images. It is used as a highly scalable and effective expression for rendering. First, video is used as a low-dimensional solution. The video is converted into patches by first compressing it into space, and then decomposing the representation into spatiotemporal patches.
[0038] According to at least one embodiment, a video compression network The WORK is a network that reduces the dimensionality of visual data, taking raw video as input. The AI then outputs a temporally and spatially compressed latent representation. It is trained in between, and then generates a video within this compressed latent space.
[0039] According to at least one embodiment, Spacetime Latent Patches are Given a compressed input video, a series of tokens that function as Transformer tokens Extract spatiotemporal patches. Using patch-based representations, Sora can be expressed at various resolutions and lengths. It can be trained with aspect ratio videos and images, and randomly initialized during inference. You can control the size of the generated video by arranging the switches in a grid of the appropriate size. .
[0040] According to at least one embodiment, the AI is a diffusion model, and the noise When many patches (and conditional information such as text prompts) are entered, the original " The AI is trained to predict "clean" patches. Formers are involved in a variety of fields, including language modeling, computer vision, and image generation. It exhibits remarkable scaling characteristics in certain regions. Diffusion transformers are used in video. It is also effective as a model. As the amount of computation required for training increases, the sample The quality of the product will improve significantly.
[0041] According to at least one embodiment, the AI applies caption regeneration technology and is very Train an explanatory caption model, then use it to train sets Generate text captions for all videos within [the specified location]. Highly descriptive captions are provided. The refinement improves not only the overall quality of the generated video but also the fidelity of the text. It leverages GPT. This converts short user prompts into long, detailed captions and sends them to the model. This allows the AI to generate high-quality videos that precisely follow the user's prompts. ru.
[0042] According to at least one embodiment, the AI processes natural language in the following manner Vectorization can be performed. First, as a preprocessing step, the given text is cleaned. The process will be carried out. During the cleaning process, JavaScript code and HTML contained within the text will be removed. Remove unnecessary words such as tags. These codes are not displayed on the internet. Because it is code used for that purpose, it is information that is not generally used in natural language processing. Next, the text is divided into words using morphological analysis. Morphological analysis is the process of dividing a sentence into words. This refers to classifying a written natural language sentence into the smallest meaningful linguistic units. For morphological analysis tools, you can use "MeCab," "JUMAN," and "JANOME." In morphology, words with the same meaning but varying spellings are unified into a single word. Stop words are: These are words that are excluded from processing for reasons such as not being usable in natural language processing. Examples of words include particles and auxiliary verbs, which do not have meaning on their own. This can be done. When calculating vectors, these can be removed, and only meaningful words will be included. It is also said that sometimes vectorization is performed without removing these stop words. Vectorization is the process of converting words, which are strings of characters, into vectors. Convert data into numerical data. When converting words into vectors, use Bag of Words or This will be done using a method called distributed representation. A Bag of Words is a set of words that appear in a given text. This method vectorizes a text using the frequency of each word. To focus on whether a single character appears, the order of words and sentences is not considered. Distributed representations are single characters. This method focuses on the meaning of words and vectorizes them. By vectorizing the meaning of words... Furthermore, it can give a similar vector to words that have similar meanings and usages, The relationships between words can also be represented by vectors. This vector representation allows for the representation of the meanings of words. Addition and subtraction are possible. The applied processing involves converting natural language into numerical data and using it as input for machine learning. It can be used to power. Specifically, vectorized natural language can be fed into a classifier to classify texts. We will implement this. The tools used here include "TensorFlow", "scikit-learn", and "PyTo Examples include "rch".
[0043] The steps for obtaining the user's voice are disclosed. The apparatus can be implemented using any one or more machines or computers as described in this disclosure. Fixed and mobile terminals acquire voice from the user. For example, a fixed terminal... Store terminals installed in stores, and machines, computers, and tablet terminals in this disclosure. Including the end. Store terminals are located in the dining area of restaurants and are used by users (in this disclosure, customers and The fixed terminal is installed close to the seats or tables where people eat and drink. We request that customers place their food orders by voice. The method of making the request is to display text on the fixed terminal's screen. The information may be displayed on the device, or an audio message to that effect may be emitted from the fixed terminal. Users will be able to eat and drink on the fixed terminal. The order is placed by voice. The fixed terminal acquires the user's voice information. Other embodiments include For example, mobile devices include user terminals, such as smartphones and tablets owned by users. This includes redline devices and those devices on which applications are installed. Users, I sit down at the table where I'm eating and drinking, and launch the application on my mobile device. It will be installed on the mobile terminal beforehand. Alternatively, it will be installed in a location close to the seat. A message will appear showing instructions on how to install the application. The user will then proceed to install the application. The application is launched. The mobile device, based on the application's instructions, responds to the user. Then, the user is asked to place their food and drink order by voice. The method of making the request is via the mobile terminal's display. The user may display the information in text (i) or emit a voice message to that effect from their mobile device. The user places food and drink orders by voice. The mobile terminal acquires the user's voice information.
[0044] The linguistic information from the audio is compared with the menu information from the restaurant's menu, and the linguistic information and We disclose the step of requesting the AI to identify one or more similar menu items. In at least one embodiment, the operator of a restaurant provides the device with one or more menu information It memorizes the information. Menu information includes language information, images, videos, etc., related to food and beverage menus. Includes information. For example, hamburger, cheeseburger, teriyaki burger, tomato, This includes linguistic information such as lettuce, and images and videos corresponding to that linguistic information. Information can store associative information that reminds us of related information or concepts. For example, regarding the information "hamburger," the associated or general information is beef, pork, and bread. Information that evokes thoughts or feelings is memorized. This information will be used for AI analysis, as will be explained later. There are cases where this occurs.
[0045] The device requests the AI to identify one or more menu items that are similar to the linguistic information. In at least one embodiment, the operation of the AI may take any of the embodiments described above. Yes, it is possible. For example, AI can be used in distributed internet service systems, fixed terminals, and mobile terminals. It exists in one or more of the following: terminal, application, or server. The device was acquired by the terminal. The language information is provided to the AI. The AI retrieves the memorized menu information. Vectorization will be performed for this. For linguistic information, Bag of Words can be used. On the other hand, distributed representations are sometimes more suitable for analyzing user statements. AI The menu information will also be vectorized. The AI will use vectors of language information and menu information. By comparing the vectors with those of the linguistic information, menu information with vectors similar to the linguistic information vectors is identified. Then, it identifies the order of similarity. In another embodiment, the AI uses language information and menus For each piece of information, various natural methods such as text classification, sentiment analysis, and text summarization are used. Perform language processing (NLP) tasks. Through these tasks, the AI classifies linguistic information. The text is compared with the categorized text of the menu information, and the categorized text of the language information is compared. Identify menu information that has text matching the list. Or, identify menu information that does not have text matching. Even if there are no matching items, the system identifies menu information with highly matching text and ranks them in descending order of matching degree. Determine. If the AI cannot identify menu information similar to the linguistic information, it will determine... It can generate status information indicating that it is not available. For example, if a user searches for "beef and bread"... If you say "I want to eat" using voice input, the AI will select "hamburger" as the menu item. It was decided that menu items such as "cheeseburger" and "teriyaki burger" would be specially listed. It will be determined.
[0046] The steps for transmitting identified menu information to the chef's terminal are disclosed. In the first embodiment, the location is where the chef of the restaurant is located, or in a location close to the kitchen. The kitchen is equipped with a computer or terminal that serves as the cook's terminal. The terminal can be a fixed terminal or a mobile terminal. That's also good. In the case of mobile devices, they can be worn by the person in charge of cooking. The device will be identified by AI. The selected menu information, or the top-ranked menu information in order of similarity, is sent to the chef's terminal. The chef's terminal transmits the received menu information to the chef. The method of transmission is as follows: The end display shows information, or outputs menu information via voice. After outputting the menu, information about the process and recipe for creating that menu will also be output. It is possible.
[0047] From the identified menu information, the step retrieves the confirmed menu information to finalize the order. We will disclose the following. The device uses menu information identified by AI, or multiple items in order of similarity. The menu information is sent to the user's terminal. The user's terminal receives this information from the user. Present the menu and receive instructions to order. The method of presentation is via the user terminal. The menu information is displayed on the screen or outputted via voice. Order The method for receiving instructions is to display a button to confirm the order on the user terminal's display. One way to interpret that being pressed is to determine that it is an order instruction. Alternatively, the user terminal is... The device obtains audio information from the source. The device then requests the AI to process the audio information using natural language processing. The AI vectorizes voice information, creating vectors that represent orders and vectors that convey intentions and emotions. One method is to confirm that it is similar to the report and then determine that it is an order instruction. The device is ordered The menu information that has been instructed to be obtained is acquired as confirmed menu information. Either generate new information as fixed menu information, or confirm the status as menu information. One method is to add information such as "hamburger" or "chi". For example, the user terminal may be labeled "hamburger" or "chi". The device displays menu information such as "Little Burger" and "Teriyaki Burger," and then acquires the user's selected menu information as confirmed menu information.
[0048] The following steps are disclosed regarding the transmission of confirmed menu information to the chef's terminal. In this embodiment, the device transmits confirmed menu information to the chef's terminal. The system transmits the received confirmed menu information to the chef. The method of transmission is via the chef's terminal display. Display information on the screen or output confirmed menu information via voice. Output confirmed menu information. After that, information about the process and recipe for creating that menu can be output along with it. can.
[0049] The language information indicated by the voice is compared with the arrangement information corresponding to the menu information, and the language information and This document discloses the step of requesting the AI to identify one or more similar arrangements. In at least one embodiment, the operator of a restaurant provides the device with one or more arrangement information It stores the information. The arrangement information includes language information, images, videos, and other information that accompanies the menu information. This includes information. For example, arrangement information is about the menu item "hamburger". The amount of sauce, whether or not tomato is included, whether or not tomato is included, and one of the several ways the pâté is cooked is selected. This includes linguistic information such as how to cook it. The arrangement information corresponds to that linguistic information. Includes images and videos. Each piece of information may evoke related information or concepts. It can store associative information. For example, for the information "lemon sauce," The information "tomato" evokes associations or concepts such as health and vitamins. The system memorizes the information it is trying to remember. This information may be used for AI analysis, as will be explained later. be.
[0050] The device requests the AI to identify one or more pieces of arrangement information that are similar to the linguistic information. In at least one embodiment, the device provides language information acquired by the terminal to the AI. The AI retrieves stored arrangement information. The AI performs vectorization on the linguistic information. For language information, Bag of Words can be used. On the other hand, user utterances Distributed representations are sometimes more suitable for analysis. AI also handles arrangement information. Vectorization is performed. The AI compares the vector of language information with the vector of arrangement information, and Identify arrangement information with vectors similar to the word information vectors, and rank them in descending order of similarity. Identify. In another embodiment, the AI, for each of the language information and the arrangement information, Performs various natural language processing (NLP) tasks such as text classification, sentiment analysis, and text summarization. These tasks allow the AI to classify linguistic information into text and arrange information. It compares the classified text with the text that matches the classified text in the linguistic information. Identify the menu information. Alternatively, even if the text does not match, select the text with the highest degree of match. The AI identifies arrangement information that has a certain degree of matching and determines the order in which it matches most closely. If no arrangement information similar to the report is identified, a status message indicating that no similar information was identified will be generated. It is possible. For example, a user might say, "I prefer sour things. I want healthy food." If you input this by voice, the AI will provide information about the customization, such as "extra lemon sauce" and "tomato." This will lead to identifying the cause.
[0051] Steps for transmitting identified menu information and arrangement information to the chef's terminal. Disclosed. In at least one embodiment, the device identifies menu information and The menu information is sent to the chef's terminal. The chef's terminal receives the identified menu information and The arrangement information is then conveyed to the chef. The method of conveying the information is by displaying it on the chef's terminal screen. The menu information is displayed or voice-over. After the arrangement information is displayed, the menu... It can output information about the process and recipe for making the product.
[0052] From the identified arrangement information, the step is to obtain the confirmed arrangement information to finalize the order. We will disclose the following. The device uses arrangement information identified by AI, or multiple arrangements in order of similarity. The arrangement information is sent to the user terminal. The user terminal receives this information from the user. Present these arrangements and receive instructions to order them. The method of presentation is via the user terminal. The display shows arrangement information, or outputs arrangement information via voice. Order The method for receiving instructions is to display a button to confirm the order on the user terminal's display. One way to interpret that being pressed is to determine that it is an order instruction. Alternatively, the user terminal is... The device obtains audio information from the source. The device then requests the AI to process the audio information using natural language processing. The AI vectorizes voice information, creating vectors that represent orders and vectors that convey intentions and emotions. One method is to confirm that it is similar to the report and then determine that it is an order instruction. The device is ordered The arrangement information that has been instructed to be obtained as confirmed arrangement information. Obtaining means confirming Either generate new information as fixed arrangement information, or confirm the status as arrangement information. One method is to add information such as "extra lemon sauce". For example, the user terminal might say "extra lemon sauce". The device presents the arrangement information "tomato" and confirms the arrangement information selected by the user. This information will be retrieved as arrangement data.
[0053] Steps for sending confirmed menu information and confirmed arrangement information to the chef's terminal. Disclosed. In at least one embodiment, the device provides confirmed menu information and confirmed allen The information is sent to the chef's terminal. The chef's terminal receives the confirmed menu information and confirmed order details. The microwave information is communicated to the cook. The method of communication is to display the information on the cook's terminal screen. Alternatively, it outputs confirmed menu information and confirmed arrangement information via voice. Afterward, it is possible to output information about the process and recipe for creating that menu item. Cut.
[0054] Menu information or arrangement information, or confirmed menu information or confirmed arrangement information The step involves instructing the self-propelled trolley to move to the designated location by approximately the time the corresponding required period has elapsed. The following is disclosed. In at least one embodiment, the operator of a restaurant provides the device with menu information. The system will store the time required for each report or arrangement. The time required will be determined at the operator's discretion. It is possible. For example, 20 minutes for menu information about a hamburger. As an example of the cooking time for the ingredients, if the patty is cooked well, it takes 5 minutes; if it's cooked medium, it takes 3 minutes. For example, adding tomatoes takes 1 minute. This corresponds to the confirmed menu information or confirmed arrangement information. The estimated time is the time required for each menu item or arrangement.
[0055] The device displays menu information or arrangement information, or confirmed menu information or confirmed arrangement. The time the information was acquired is generated as order time information. The device then uses the order time information to generate the order time information. Then, the corresponding required time is added to generate estimated completion time information.
[0056] A self-propelled trolley is disclosed. In at least one embodiment, the self-propelled trolley is unmanned and self It is a vehicle or machine that moves. A self-propelled cart has one or more spaces for carrying food. The trolley may have a computer or terminal as described in this disclosure. That is, a self-propelled trolley The vehicle can communicate with the device wirelessly. In at least one embodiment, the object is observed If, upon observation, it is an unmanned, self-propelled vehicle or machine, it will be considered a "self-propelled trolley" in this disclosure. The device will instruct the self-propelled trolley to move to the designated location at approximately the estimated completion time. As a legal requirement, the device must, at approximately the estimated completion time, implement a program or Instructions are transmitted wirelessly. The self-propelled trolley receives the program or instructions and, at approximately the estimated completion time, It moves under its own power to a designated location. Examples of designated locations include the cook's area, the kitchen, and the finished dish. There are designated pick-up and drop-off locations.
[0057] Menu information or arrangement information, or confirmed menu information or confirmed arrangement information The step of summing the corresponding numerical values is disclosed. In at least one embodiment, The restaurant operator stores numerical values corresponding to menu information or arrangement information in the device. The figures represent at least the price of these menu items or variations. For example, hamburger A hamburger costs 1000, and extra lemon sauce costs 50, etc. So a hamburger costs 1000. The price is 50 yen, and the lemon sauce is 50 yen. The device displays menu information, arrangement information, or confirmation. The numerical values corresponding to the menu information or confirmed arrangement information are summed up. For example, for a hamburger... If you add extra lemon sauce, the total value will be 1050.
[0058] The steps for presenting the aggregated figures to the user will be disclosed. At least one implementation will be disclosed. In this state, the device transmits the summed values to the user terminal. In another embodiment, The device then transmits the total amount to an accounting device for the user to perform the accounting. Examples include accounting equipment, cash registers, and user-facing accounting applications in food service establishments. , applications for users' electronic payments, user terminals (including smartphones) ) and so on.
[0059] In at least one embodiment, any one of the embodiments described above produces in the following case It has commercial applicability and advantages, namely, restaurants, drive-throughs, etc. In the case of EveThru, the user terminal in the step of acquiring the user's voice is dry In some embodiments, this may correspond to a machine installed outdoors for bus-through purposes. In the case of a drive-thru, the user's own device such as a smartphone or computer These devices, with a dedicated application installed, are equivalent to user terminals. There are cases where this occurs.
[0060] The following is an overview of the embodiments described above.
[0061] The ordering method is, the device is, Steps to acquire the user's voice, The linguistic information from the audio is compared with the menu information from the restaurant's menu, and the linguistic information is then compared with the menu information. The steps involve requesting the AI to identify one or more similar menu items, The steps include: sending the identified menu information to the chef's terminal, A method that has
[0062] According to this disclosure, at the very least, users can order food and drinks by voice, which offers convenience. Furthermore, restaurants can take orders unmanned based on the user's voice. It has industrial applicability and advantages.
[0063] The ordering method is, the device is, Steps to acquire the user's voice, The linguistic information from the audio is compared with the menu information from the restaurant's menu, and the linguistic information is then compared with the menu information. The steps involve requesting the AI to identify one or more similar menu items, Steps to obtain confirmed menu information from the identified menu information to finalize the order. and, The steps include sending the confirmed menu information to the chef's terminal, A method that has
[0064] According to this disclosure, at the very least, users can order food and drinks by voice, which offers convenience. Furthermore, restaurants can take orders unmanned based on the user's voice. It has industrial applicability and advantages.
[0065] The ordering method is, the device is, Steps to acquire the user's voice, The linguistic information from the audio is compared with the menu information from the restaurant's menu, and the linguistic information is then compared with the menu information. The steps involve requesting the AI to identify one or more similar menu items, The language information indicated by the audio is compared with the arrangement information corresponding to the menu information, and the language information is compared with the arrangement information. The steps involve requesting the AI to identify one or more similar arrangement pieces, The steps include: transmitting the identified menu information and arrangement information to the chef's terminal; A method that has
[0066] According to this disclosure, at the very least, users can order food and drinks by voice, which offers convenience. Furthermore, restaurants can take orders unmanned based on the user's voice. It has industrial applicability and advantages.
[0067] The ordering method is, the device is, Steps to acquire the user's voice, The linguistic information from the audio is compared with the menu information from the restaurant's menu, and the linguistic information is then compared with the menu information. The steps involve requesting the AI to identify one or more similar menu items, The language information indicated by the audio is compared with the arrangement information corresponding to the menu information, and the language information is compared with the arrangement information. The steps involve requesting the AI to identify one or more similar arrangement pieces, Steps to obtain confirmed menu information from the identified menu information to finalize the order. and, Steps to obtain confirmed arrangement information to finalize the order from the identified arrangement information. and, The steps include sending the confirmed menu information and confirmed arrangement information to the chef's terminal, A method that has
[0068] According to this disclosure, at the very least, users can order food and drinks by voice, which offers convenience. Furthermore, restaurants can take orders unmanned based on the user's voice. It has industrial applicability and advantages.
[0069] Any of the above ordering methods, By the approximate time elapsed corresponding to the menu information or arrangement information, the self-propelled trolley should be moved to the designated location. Steps to give instructions for self-propelled movement to a location, A further method
[0070] According to this disclosure, at the very least, the food and beverages ordered by the user are quickly transported by a self-propelled trolley. This makes it possible to deliver freshly prepared meals to users without human intervention. It has industrial applicability and advantages, meaning it can be done.
[0071] Any of the above ordering methods, A step of summing up numerical values corresponding to menu information or arrangement information, The steps include presenting the combined figures to the user, A further method
[0072] According to this disclosure, at least the food and beverages ordered by the user will be processed automatically. It has industrial applicability and advantages, as it allows for such applications.
[0073] In at least one embodiment, "from the identified menu information, confirm the order." In the step of obtaining confirmed menu information, the device receives the instruction to place an order. The "law" includes the following forms: The operator of a restaurant must install equipment that allows for one or more physical actions. Information is defined and stored as specific action information. Action information is linguistic information that accompanies physical movements. This includes information such as reports, images, and videos. For example, it includes information about a physical action such as "tapping a desk twice with your fingers." This includes information such as language information and images or videos showing examples of the actions. Physical actions are This includes all gestures such as raising your hand and nodding. Specific action information is included in the action information. Therefore, this is information that restaurant operators register with the device. The user terminal displays the user's activity information. Obtain the information. The method of acquisition is to capture video or video of the user's actions using the camera on the user's device. There is a way to obtain it.
[0074] The device requests the AI to determine if the operation information is similar to specific operation information. Regarding the acquired operational information, the video and images are converted into small data similar to LLM text tokens. It is represented as a collection of visual patches, which are units. The AI is about the operation information. A network that reduces the dimensionality of visual data, taking raw video as input and processing it temporally It outputs a spatially compressed latent representation. The AI also handles specific action information in the same way as above. The latent representation is output using the following method. The AI outputs the latent representation of the motion information and the latent representation of the specific motion information. They compare the two and determine if there is similarity. In this determination, the restaurant operator will consider the similarity to be the key factor. The device can be programmed to store a similarity threshold that should be used to determine if something is related. AI can store behavioral information. If it is determined that the information is similar to specific operational information, data or status information indicating that it is similar will be provided. It can be created.
[0075] The device will, on the condition that data or status information indicating similarity is generated, will fulfill the order. The menu information that has been instructed to be obtained is acquired as confirmed menu information. Either generate new information as fixed menu information, or confirm the status as menu information. One method is to add menu information. For example, the device will add "hamburger" to the menu information. The information is presented to the user. The user taps the desk twice with their fingers. The device responds to the user's actions. The system determines that the information is specific action information and retrieves "hamburger" as confirmed menu information. ru.
[0076] In at least one embodiment, "from among the identified arrangement information, confirm the order." In the step of obtaining confirmed arrangement information, the device receives instructions to place an order. The "Law" can adopt the above embodiment. That is, the operator of a restaurant may use the device, Define action information related to one or more physical movements and store it as specific action information. Finally, it acquires user action information. The AI determines whether the action information is similar to confirmed action information. The device will reject the order if similar data or status information is generated. The arrangement information received as instructed is acquired as confirmed arrangement information. Acquiring means confirming Either generate new information as arrangement information, or confirm the status as arrangement information. There are methods for adding information. For example, the device can be customized with "extra lemon sauce". The information is presented to the user. The user nods. The device identifies the user's actions. It is determined that this is operational information, and "extra lemon sauce" is obtained as confirmed customization information.
[0077] In at least one embodiment, in the above embodiment, separately from the user terminal, Alternatively, a separate device or terminal can be used in conjunction with the user terminal to detect the user's actions. It is permissible to set one up.
[0078] In at least one embodiment, in the above embodiment, operation information and specific operation information Alternatively, linguistic information and specific linguistic information may be used as the information to be judged. That is, food and drink. The store operator shall define one or more linguistic pieces of information equivalent to a password in the device, and as specific linguistic information It remembers. The user terminal obtains the user's language information. The method of obtaining this information is determined by the device or terminal. One possible method is to use a microphone. The AI will use natural language to determine if the linguistic information is similar to specific action information. The decision is made through word processing. The device requests the AI to make that decision. As already explained, you may use any one or more of the methods described in this disclosure. The device will use data or status information indicating that language information is similar to specific language information. Or, if created by a device, the menu information or arrangement information you received when ordering The information will be retrieved as confirmed menu information or confirmed arrangement information. Retrieving means confirming arrangement The status information either generates new information as range information or confirms it as arranged information. There are methods for providing information, such as the device providing information on the "extra lemon sauce" option. This information is presented to the user. The user utters a password set by the restaurant operator (1 Examples include trademarks and logos of restaurants, brand names and nicknames, and special catchphrases. The device determines that the user's language information is specific language information and confirms "extra lemon sauce". This information will be retrieved as fixed arrangement information.
[0079] The following is an overview of the embodiments described above.
[0080] Any of the above ordering methods, The step to obtain confirmed menu information to finalize the order is: The device stores specific action information relating to one or more physical actions, A step in which the AI is asked to determine whether the action information is similar to specific action information, It also includes.
[0081] Any of the above ordering methods, The step to obtain confirmed menu information to finalize the order is: When the AI determines that the action information is similar to specific action information, it confirms the menu information. The steps to report, It also includes.
[0082] According to this disclosure, at the very least, users can confirm their orders by gesture. This makes it simpler than the conventional method of touching a touch panel. The order can be confirmed. This improves the convenience of using the device or system. This has industrial applicability and advantages.
[0083] The invention relating to this disclosure is sufficient if it can achieve at least one of the effects described above. stomach.
Claims
1. A menu recommendation method in an ordering system, a) A step of obtaining a question specifying a category via voice input from the user, b) The steps of applying natural language processing to the above question, searching for categories from menu information, and categorizing them, c) A step of extracting menu items that belong to the categorized categories and recommending them to the user, A menu recommendation method characterized by including the following.
2. A menu recommendation method in an ordering system, a) A step of obtaining a question specifying a category via voice input from the user, b) A step of applying natural language processing to the above question and generating categories based on the metadata of the menu items, c) A step of extracting menu items corresponding to the generated categories and recommending them to the user, A menu recommendation method characterized by including the following.
3. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of presenting to the user an explanation of menu items corresponding to the category using speech synthesis.
4. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of interactively confirming the menu items recommended based on the user's voice response.
5. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of prioritizing menu items within the category based on the user's past order history.
6. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of prioritizing menu items within the category using trend analysis data based on the order history of other users.
7. A menu recommendation method according to claim 1 or claim 2, characterized in that the natural language processing includes the step of performing a user sentiment analysis and adjusting the recommendation of menu items within the category based on the sentiment.
8. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of translating menu items corresponding to the category into the user's preferred language and presenting them.
9. A menu recommendation method according to claim 1 or claim 2, characterized in that the recommendation step includes a step of adjusting menu items within the category based on environmental data including weather information or store congestion information.
10. A menu recommendation method according to claim 1 or claim 2, characterized by including the step of accessing an external coupon API and recommending menu items within the category, including applicable discount information.
Citation Information
Patent Citations
Voice-based dish ordering method, device and equipment
CN118155617A
Information recommendation system, information search device, information recommendation method, and program
WO2021250833A1
Seat managing pos system
JP1996161403A