Intelligent control method, device and equipment of unmanned beverage robot, and storage medium

CN122817409APending Publication Date: 2026-09-25BEIJING YINGZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611166254.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]本申请提供一种无人饮品机器人智能控制方法、装置、电子设备及存储介质,以至少解决相关技术中由于无人饮品机器人采用传统界面罗列数据,运维人员需要记忆多层操作路径,并在多个页面间切换查看不同类型的信息,识别故障时还要查阅操作手册,运维分析还需要查看报表,导致其操作效率低下的问题

Benefits of technology

本申请实施例中,获取运维用户输入的自然语言问题;确定与所述自然语言问题相关的无人饮品机器人的多源上下文信息,包括:设备实时运行状态数据、运营历史数据、知识片段和多轮历史对话上下文;基于所述自然语言问题的意图分类结果选择对应的提示词模板;按照选择的所述提示词模板对所述多源上下文信息和所述自然语言问题进行拼接,生成标准化输入提示词;基于所述标准化输入提示词调用大语言模型进行推理计算,生成结构化回答;在所述结构化回答中包括设备操控指令时,对所述设备操控指令进行解析与安全确认;在安全确认后通过远程控制接口下发所述设备操控指令至无人饮品机器人,控制所述无人饮品机器人执行对应运维操作。也就是说,本实施例中,运维用户无需记忆复杂的操作路径,通过自然语言对话即可完成设备状态查询、运营数据分析和设备操控,提高了操作效率;实时感知设备状态,提高了回答的准确性。本实施例通过多源上下文信息,大语言模型能够基于设备当前真实状态生成结构化回答,避免了大语言模型因缺乏实时数据而产生的幻觉问题。本实施例结合设备知识库和设备实时状态数据,大语言模型对常见故障提供准确的原因分析和处理建议,故障诊断辅助效果显著。本实施例通过用户二次确认和权限验证的双重安全机制,有效防止大语言模型误判导致的误操作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817409A_ABST
    Figure CN122817409A_ABST
Patent Text Reader

Abstract

The application provides an intelligent control method, device and equipment of an unmanned beverage robot and a storage medium. The method comprises: obtaining a natural language question input by an operation and maintenance user; determining multi-source context information of the unmanned beverage robot related to the natural language question; selecting a corresponding prompt word template based on an intention classification result of the natural language question; splicing the multi-source context information and the natural language question according to the selected prompt word template to generate a standardized input prompt word; calling a large language model based on the standardized input prompt word to perform inference calculation and generate a structured answer; when the structured answer includes a device control instruction, analyzing and safely confirming the device control instruction; after the safety confirmation, issuing the device control instruction to the unmanned beverage robot through a remote control interface to control the unmanned beverage robot to perform a corresponding operation and maintenance operation. The application improves operation efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and unmanned retail technology, and in particular to an intelligent control method, device, electronic device and storage medium for an unmanned beverage robot. Background Technology

[0002] With the rapid popularization of artificial intelligence and robotics technologies, unmanned beverage robots are being widely used in various public places. These robots are fully automated, unattended devices deployed in public places such as shopping malls, office buildings, and universities. Their operation and maintenance management involves multiple tasks, including status monitoring, material management, fault handling, and operational analysis. Currently, the industry's supporting operation and maintenance management software only displays data through traditional interfaces. Maintenance personnel need to frequently switch pages and memorize multiple operation paths to obtain complete equipment information; equipment anomalies only return fault codes, and troubleshooting requires manual consultation of manuals; operational data relies on manual verification and report analysis. The entire process suffers from cumbersome operation and maintenance, low fault handling efficiency, and delayed operational decision-making, resulting in low operational efficiency and low accuracy. Summary of the Invention

[0003] This application provides an intelligent control method, device, electronic device, and storage medium for an unmanned beverage robot, aiming to at least solve the problems in related technologies where unmanned beverage robots use traditional interfaces to display data, requiring maintenance personnel to memorize multiple operation paths, switch between multiple pages to view different types of information, consult operation manuals for fault identification, and view reports for maintenance analysis, resulting in low operational efficiency. The technical solution of this application is as follows: According to a first aspect of the embodiments of this application, an intelligent control method for an unmanned beverage robot is provided, the method being applied to an intelligent operation assistant system, the method comprising: Obtain natural language input from operations and maintenance users; Determine the multi-source contextual information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context; Select the corresponding prompt word template based on the intent classification result of the natural language question; The multi-source context information and the natural language question are concatenated according to the selected prompt word template to generate standardized input prompt words; Based on the standardized input prompts, a large language model is invoked to perform inference calculations and generate a structured answer; When the structured response includes device control commands, the device control commands are parsed and securely verified. After safety confirmation, the device control command is sent to the unmanned beverage robot through the remote control interface to control the unmanned beverage robot to perform the corresponding operation and maintenance.

[0004] Optionally, the determination of multi-source contextual information of the unmanned beverage robot related to the natural language problem includes: Based on the aforementioned natural language problem, real-time operational status data of the unmanned beverage robot is collected. Retrieve operational history data related to the natural language problem from the cloud database; Retrieve knowledge fragments related to the natural language problem from the knowledge base; Retrieve the context of multi-turn historical dialogues managed by the sliding window; The real-time operating status data, historical operating data, knowledge fragments, and multi-turn historical dialogue context of the device are fused to obtain multi-source context information.

[0005] Optionally, selecting the corresponding prompt word template based on the intent classification result of the natural language question includes: Semantic recognition is performed on the natural language problem to obtain the recognition result; The recognition results are classified according to the intent to obtain the corresponding intent classification results, wherein the intent classification results include at least one of the following: query, control, fault diagnosis and training. Select the corresponding prompt word template based on the intent classification results.

[0006] Optionally, the step of invoking a large language model to perform inference calculations based on the standardized input prompts to generate a structured answer includes: The standardized input prompts are input into a large language model via an application programming interface (API). The large language model loads multi-source contextual information from the standardized input prompts and performs semantic parsing in conjunction with the natural language question to determine the maintenance user's needs. Based on the maintenance user's needs, a structured response is generated, including at least one of the following: natural language response text, device control instructions, and associated shortcut operation matching. The structured response is then output.

[0007] Optionally, when the structured response includes device control commands, parsing and security verification of the device control commands includes: When the structured response includes device control instructions, the device control instructions are parsed to obtain the parsing result; Extract the operation type, operation target, and operation parameters from the parsing results; Initiate permission confirmation with the operation and maintenance user to execute the operation type, the operation target, and the operation parameters; Upon receiving confirmation from the operations and maintenance user, security is confirmed.

[0008] Optionally, the method further includes: When the structured response does not include device control instructions, the system displays formatted Q&A content to the maintenance user and pushes relevant operation suggestions that can be quickly executed.

[0009] Optionally, the method further includes: Obtain the execution result of the device control command, the execution result including success or failure and the reason for failure, and send the execution result to the operation and maintenance user.

[0010] Optionally, the effective question and answer content of the current round of dialogue is recorded and the historical dialogue context is updated; and the recorded effective question and answer content of the current round of dialogue is synchronized to the cloud so that the cloud updates the knowledge base based on the effective question and answer content of the current round of dialogue.

[0011] According to a second aspect of the embodiments of this application, an intelligent control device for an unmanned beverage robot is provided. The device is applied to an intelligent operation assistant system, and the device includes: The acquisition module is used to acquire natural language questions input by operation and maintenance users; The determination module is used to determine the multi-source context information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context; The selection module is used to select the corresponding prompt word template based on the intent classification result of the natural language question. The concatenation module is used to concatenate the multi-source context information and the natural language question according to the selected prompt word template to generate standardized input prompt words; The reasoning module is used to call a large language model to perform reasoning calculations based on the standardized input prompts and generate structured answers. The confirmation module is used to parse and securely confirm the device control instructions when the structured response includes device control instructions. The control module is used to send the device control commands to the unmanned beverage robot through a remote control interface after safety confirmation, so as to control the unmanned beverage robot to perform corresponding operation and maintenance operations.

[0012] Optionally, the determining module includes: The data acquisition module is used to collect real-time operating status data of the unmanned beverage robot based on the natural language question. The first retrieval module is used to retrieve operational history data related to the natural language problem from the cloud database; The second retrieval module is used to retrieve knowledge fragments related to the natural language problem from the knowledge base; The dialogue acquisition module is used to acquire the context of multi-turn historical dialogues managed by the sliding window. The fusion module is used to fuse the device's real-time operating status data, historical operating data, knowledge fragments, and multi-turn historical dialogue context to obtain multi-source context information.

[0013] Optionally, the selection module includes: The semantic recognition module is used to perform semantic recognition on the natural language question and obtain the recognition result; An intent classification module is used to classify the recognition results according to intent to obtain intent classification results, wherein the intent classification results include at least one of the following: query, manipulation, fault diagnosis and training. The template selection module is used to select the corresponding prompt word template based on the intent classification result.

[0014] Optionally, the inference module includes: The model input module is used to input the standardized input prompts into the large language model via an application programming interface (API). The model reasoning module is used to determine the maintenance user's needs by loading multi-source context information from the standardized input prompts into the large language model and performing semantic parsing in conjunction with the natural language question; based on the maintenance user's needs, it generates a structured response that includes at least one of the following: natural language response text, device control instructions, and related shortcut operation matching. The model output module is used to output the structured answer.

[0015] Optionally, the confirmation module includes: The parsing module is used to parse the device control instructions when the structured answer includes device control instructions, and obtain the parsing result; The extraction module is used to extract the operation type, operation target, and operation parameters from the parsing results; The permission confirmation module is used to initiate permission confirmation for the operation and maintenance user to execute the operation type, the operation target, and the operation parameters; The security determination module is used to confirm security when receiving confirmation from the operation and maintenance user.

[0016] Optionally, the confirmation module further includes: The archive synchronization module is used to archive the conversation record and synchronize it to the cloud when it receives the cancellation operation from the operation and maintenance user.

[0017] Optionally, the device further includes: The first sending module is used to display formatted Q&A content to the operation and maintenance user and push relevant operation suggestions that can be quickly executed when the structured answer does not include device control instructions.

[0018] Optionally, the device further includes: The result acquisition module is used to acquire the execution result of the device control command, and the execution result includes success or failure and the reason for failure; The second sending module is used to send the execution result to the operation and maintenance user.

[0019] Optionally, the device further includes: The record archiving module is used to record the valid questions and answers in the current round of dialogue and update the historical dialogue context; and The third sending module is used to synchronize the recorded valid question and answer content of the current round of dialogue to the cloud, so that the cloud can update the knowledge base based on the valid question and answer content of the current round of dialogue.

[0020] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising: It includes a processor, a memory; and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the intelligent control method for the unmanned beverage robot as described above.

[0021] According to a fourth aspect of the embodiments of this application, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor of an electronic device, implement the steps of the intelligent control method for an unmanned beverage robot as described above.

[0022] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program or instructions, which, when executed by a processor of an electronic device, implement the steps of the intelligent control method for an unmanned beverage robot as described above.

[0023] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects: In this embodiment, a natural language question input by the maintenance user is obtained; multi-source contextual information of the unmanned beverage robot related to the natural language question is determined, including: real-time operating status data of the device, historical operating data, knowledge fragments, and multi-turn historical dialogue context; a corresponding prompt word template is selected based on the intent classification result of the natural language question; the multi-source contextual information and the natural language question are concatenated according to the selected prompt word template to generate standardized input prompt words; a large language model is invoked based on the standardized input prompt words to perform inference calculations and generate a structured answer; when the structured answer includes device control instructions, the device control instructions are parsed and securely confirmed; after security confirmation, the device control instructions are sent to the unmanned beverage robot through a remote control interface to control the unmanned beverage robot to perform the corresponding maintenance operation. In other words, in this embodiment, the maintenance user does not need to remember complex operation paths and can complete device status queries, operational data analysis, and device control through natural language dialogue, improving operational efficiency; real-time perception of device status improves the accuracy of the answer. This embodiment, through multi-source contextual information, enables the large language model to generate structured answers based on the current real state of the device, avoiding the illusion problem caused by the lack of real-time data in the large language model. This embodiment combines a device knowledge base and real-time device status data. The large language model provides accurate cause analysis and handling suggestions for common faults, resulting in significant fault diagnosis assistance. This embodiment effectively prevents erroneous operations caused by misjudgments from the large language model through a dual security mechanism of secondary user confirmation and permission verification.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0025] The accompanying drawings, incorporated in and forming part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. They do not constitute an undue limitation of this application. To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0026] Figure 1 This is a flowchart of an intelligent control method for an unmanned beverage robot provided in an embodiment of this application.

[0027] Figure 2 This is a flowchart illustrating an application example of an intelligent control method for an unmanned beverage robot provided in this application embodiment.

[0028] Figure 3 This is a block diagram of an intelligent control device for an unmanned beverage robot provided in an embodiment of this application.

[0029] Figure 4 This is a block diagram of a determining module provided in an embodiment of this application.

[0030] Figure 5 This is a block diagram of a selection module provided in an embodiment of this application.

[0031] Figure 6 This is a block diagram of a confirmation module provided in an embodiment of this application.

[0032] Figure 7 This is a schematic diagram of the architecture of an intelligent operation assistant system provided in an embodiment of this application.

[0033] Figure 8 This is a block diagram of an electronic device provided in an embodiment of this application.

[0034] Figure 9 This is a block diagram of an intelligent control device for an unmanned beverage robot provided in an embodiment of this application. Detailed Implementation

[0035] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0036] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0037] Please see Figure 1 This is a flowchart of an intelligent control method for an unmanned beverage robot provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps: Step 101: Obtain the natural language questions input by the operations and maintenance user.

[0038] Step 102: Determine the multi-source contextual information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context.

[0039] Step 103: Select the corresponding prompt word template based on the intent classification result of the natural language question.

[0040] Step 104: Concatenate the multi-source context information and the natural language question according to the selected prompt word template to generate standardized input prompt words.

[0041] Step 105: Based on the standardized input prompts, call the large language model to perform inference calculations and generate a structured answer.

[0042] Step 106: When the structured response includes device control instructions, the device control instructions are parsed and securely verified.

[0043] Step 107: After safety confirmation, send the device control command to the unmanned beverage robot through the remote control interface to control the unmanned beverage robot to perform the corresponding operation and maintenance.

[0044] The intelligent control method for unmanned beverage robots described in this application can be applied to terminals, servers, or intelligent operation assistant systems for unmanned beverage robots, etc., without limitation. The terminal implementation device can be an electronic device such as a smartphone, laptop, tablet, desktop computer, personal digital assistant (PDA), and wearable device; the server can be an independent server, a server cluster, or a server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, or big data and artificial intelligence platforms, etc., without limitation.

[0045] The following is combined Figure 1 The specific implementation steps of the intelligent control method for an unmanned beverage robot provided in the embodiments of this application will be described in detail.

[0046] In step 101, the natural language question input by the operation and maintenance user is obtained.

[0047] In this step, the terminal (such as the front end of the intelligent operation assistant system) receives questions in natural language form from the operation and maintenance personnel. These questions serve as the initial input for this round of dialogue. The questions can include various operation and maintenance requests such as equipment information query, remote control request, fault diagnosis consultation, and operation and maintenance skills training.

[0048] In this application embodiment, "operation and maintenance user" is a general term. In specific applications, depending on the application scenario, it can be a user with different identities, such as equipment maintenance personnel, operation management personnel, operation training personnel, etc. This implementation does not impose any restrictions.

[0049] In step 102, multi-source context information of the unmanned beverage robot related to the natural language problem is determined. The multi-source context information includes: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context.

[0050] This step includes: 11) collecting real-time operating status data of the unmanned beverage robot based on the natural language question; wherein, the real-time operating status data may include: the remaining amount of each material (e.g., coffee beans, milk, purified water, paper cups, etc.), equipment operating parameters (e.g., temperature, pressure, output count, etc.) and equipment fault status (e.g., fault codes and descriptions, etc.).

[0051] 12) Retrieve historical operational data related to the natural language question from the cloud database, such as historical fault records and maintenance handling records. Specifically, Retrieval-Augmented Generation (RAG) can be used to retrieve recent operational data (sales, revenue, production records, material consumption trends, etc. for the current day / last 7 days / last 30 days) from the cloud database, and vector similarity matching can be used to filter the most relevant historical data segments to the current question (i.e., the natural language question). RAG is an enhancement technology framework for large language models. The process involves retrieval before generation. Before the large language model answers the user's natural language question, relevant reference information is retrieved from external data sources (such as cloud databases). This retrieved information is used as context and incorporated into prompts, which are then fed into the large language model for reasoning and answer generation. The purpose is to compensate for the shortcomings of the large language model's built-in knowledge lag and lack of private business knowledge, effectively reducing model illusion and ensuring that the model's output is based on real business data and documentation.

[0052] 13) Retrieve knowledge fragments related to the natural language problem from the knowledge base, such as maintenance manuals and operation and maintenance specifications.

[0053] Specifically, in this step, semantic retrieval (vector similarity matching) and a hybrid retrieval method based on BM25 sparse retrieval and vector dense retrieval can be used to retrieve the most relevant knowledge fragments to the current problem (i.e., the natural language problem) from knowledge bases such as equipment operation manuals, troubleshooting guides, and best practices. (For example, selecting the Top-K knowledge fragments, K=5, etc., but in practical applications, it is not limited to this, where K is a natural number not equal to zero). In other words, a hybrid retrieval method combining BM25 sparse retrieval and vector dense retrieval can be used to retrieve knowledge fragments from the knowledge base in parallel through keyword matching and semantic vector matching. The results from the two retrieval paths are weighted, fused, scored, and rearranged to select highly relevant knowledge fragments to construct a multi-source context.

[0054] Among them, BM25 sparse retrieval (keyword matching retrieval) is a traditional text retrieval algorithm that relies on keywords, technical terms, fault codes, and literal matching of equipment parameters to search the text library. It calculates text relevance scores based on word frequency and word importance, and can accurately match entity words such as operation and maintenance technical terms, fault codes, equipment models, and material names without losing literally relevant documents.

[0055] Dense vector retrieval (semantic vector retrieval) transforms user questions and knowledge base documents into high-dimensional semantic vectors through an embedding model, and calculates semantic similarity based on cosine similarity; it can understand deep semantics and match questions with different expressions but the same meaning (e.g., semantic association matching between "the machine does not produce beverages" and "the discharge pipe is blocked").

[0056] Hybrid retrieval based on BM25 sparse retrieval and vector dense retrieval, namely the RAG architecture, is implemented through the following process: First, parallel retrieval via two branches: User questions (natural language questions entered by operations and maintenance users) are simultaneously sent to two search branches, and the search is performed concurrently: Branch 1: Segment the question into words, extract keywords, and use the BM25 algorithm to perform sparse keyword retrieval in the operation and maintenance knowledge base, outputting document fragments with the highest keyword matching degree; Branch 2: Call the text embedding model to convert user questions and all knowledge fragments in the knowledge base into dense semantic vectors, perform similarity retrieval in the vector database, and output the document fragments with the highest semantic matching degree.

[0057] Then, the two search results are merged and scored: The BM25 score and vector similarity score were normalized separately, and weighted and fused with weight coefficients to calculate the comprehensive relevance score of each document. The weights can be adjusted according to business needs: increase the weight of BM25 for fault code and equipment parameter query scenarios; increase the weight of vector retrieval for fuzzy fault consultation and broad maintenance question scenarios.

[0058] This embodiment proposes a hybrid retrieval method that combines precise literal matching with deep semantic understanding, addressing the shortcomings of single retrieval methods. Specifically designed for the operation and maintenance of unmanned beverage robots, it can accurately locate proprietary information such as fault codes, materials, and equipment parameters, while also recognizing colloquial and vague maintenance questions. Furthermore, it improves retrieval recall accuracy, ensuring higher validity of reference knowledge fed into the large language model, reducing model illusions, and making fault diagnosis and equipment maintenance suggestions more aligned with actual equipment operating conditions. The process involves re-ranking and filtering to output Top-K knowledge fragments, where K is a non-zero natural number. All retrieved documents are sorted according to their combined scores, low-relevance redundant content is removed, and the highest-scoring knowledge fragments are selected as the retrieval results, incorporating multi-source contextual information.

[0059] 14) Obtain the context of multi-turn history dialogues managed by the sliding window.

[0060] In this step, the dialogue history of the most recent N rounds (e.g., N=10 by default) is retrieved, and the context length is managed through a sliding window mechanism to ensure dialogue continuity. In other words, a sliding window mechanism is used to manage the historical dialogue context, setting the maximum number of rounds N that the window can store (the default value is 10). The dialogue cache is updated after each round of interaction. When the cache is full, the oldest dialogue record is automatically removed, and only the most recent N rounds of question-and-answer content is retained for prompt word concatenation, balancing the semantic continuity of multi-round dialogues with input text length control.

[0061] In this embodiment, to maintain semantic coherence in multi-turn continuous question-and-answer sessions, a fixed-length sliding window mechanism is proposed to manage dialogue records, specifically including: 141) Setting dialog storage rules.

[0062] The system presets a maximum number of sliding window rounds N, with a default configuration value of 10 (this is not limited in actual applications and can be adjusted as needed). It only saves the question-and-answer records of recent interactions between the maintenance user and the intelligent assistant, discarding earlier dialogue content that exceeds the round limit, where N is a natural number that is not equal to zero.

[0063] 142) Sliding window update mechanism.

[0064] After each round of question-and-answer interaction is completed, the user's question and the structured response from the large language model are automatically stored in the dialogue cache. If the number of dialogue rounds in the cache reaches the set upper limit N, the earliest historical dialogue is automatically removed, and new dialogue records are added to the end of the cache. The window is then updated by sliding forward.

[0065] 143) Context reuse.

[0066] When constructing the Prompt prompt, the entire historical dialogue context stored in the window is simultaneously filled into the prompt template. The large language model can combine the previous question and answer content to understand the user's continuous requests such as coherent follow-up questions and progressive fault inquiries, avoiding irrelevant answers due to missing information in a single question, and ensuring the semantic coherence of multi-turn dialogues.

[0067] 15) The real-time operating status data, historical operating data, knowledge fragments and multi-round historical dialogue context of the device are fused to obtain multi-source context information.

[0068] In this step, the four types of information (i.e., real-time equipment operating status data, historical operating data, knowledge fragments, and multi-round historical dialogue context) are integrated and organized to obtain multi-source context information for the current round of dialogue. The method of integration and organization is well-known to those skilled in the art and will not be elaborated here.

[0069] One implementation method integrates the four types of information mentioned above into a multi-source contextual information architecture under a multi-source contextual RAG framework. The Retrieval Enhanced Generation (RAG) technology effectively addresses the issues of lagging knowledge updates and insufficient domain specialization in large language models by combining external knowledge bases with large language model (LLM) reasoning. Based on this, this embodiment proposes integrating RAG technology with real-time device status data into the operation and control scheme of unmanned beverage robots. This achieves, for the first time, the dynamic injection of real-time device status data of unmanned beverage robots into the LLM context. Combined with historical operation data retrieval and semantic retrieval from the device knowledge base, a multi-source contextual RAG architecture is constructed for unmanned beverage robot operation scenarios, enabling the LLM to generate accurate and real-time operational suggestions based on the current real-time status of the device.

[0070] In step 103, the corresponding prompt word template is selected based on the intent classification result of the natural language question.

[0071] The step includes: performing semantic recognition on the natural language question to obtain recognition results; classifying the recognition results according to intent to obtain corresponding intent classification results, wherein the intent classification results include at least one of query, manipulation, fault diagnosis and training; and selecting the corresponding prompt word template according to the intent classification results.

[0072] In this step, semantic recognition is performed on the natural language questions input by the operations and maintenance user to classify them into four categories: query, operation, fault diagnosis, and training. These classification results are then used as the basis for selecting the subsequent Prompt template. It should be noted that this embodiment uses four categories of intents as an example; in practical applications, it is not limited to this.

[0073] The following explanation uses four categories of intent as examples, but in practice, it is not limited to these: (1) Query type: equipment status query, operation data query, knowledge Q&A, etc., directly generating natural language answers, etc.

[0074] (2) Control type: Remote control commands for equipment (such as power on / off, parameter adjustment, cleaning start, etc.) require the extraction of structured commands (i.e. equipment control commands) and execution after user confirmation.

[0075] (3) Fault diagnosis: Based on the description of the fault phenomenon, combined with equipment status data and knowledge base, generate fault cause analysis and handling suggestions.

[0076] (4) Training: Operating procedure query, equipment knowledge learning, etc. Relevant content is retrieved from the knowledge base and presented in an easy-to-understand way.

[0077] In this embodiment, intent classification drives a differentiated prompt template engine. This embodiment pre-designs corresponding dedicated prompt templates for four different intents: query, control, fault diagnosis, and training. In application, the corresponding prompt template can be dynamically selected using an intent classifier or similar device, ensuring that large language models (such as LLM) generate consistent and accurate structured answers in different scenarios, significantly improving the quality and usability of the answers.

[0078] In step 104, the multi-source context information and the natural language question are concatenated according to the selected prompt word template to generate standardized input prompt words.

[0079] The Prompt template used in this step employs a structured template. In this step, according to the selected prompt template, real-time device operating status data, historical operational data, knowledge fragments, multi-turn historical dialogue context, and natural language questions are concatenated to generate standardized input prompts for the large language model. The Prompt templates are designed separately for different intent types (query / operation / fault diagnosis / training) to ensure that the LLM generates consistent and accurate responses across various scenarios.

[0080] In this embodiment, to constrain the structured response style and professional domain of the large language model's output, it can also be combined with preset roles (role definitions) to obtain standardized input prompts for the large language model. That is, role definitions (such as AI store manager identity, capability boundaries, etc.), real-time device operating status data (or corresponding device status summaries can be extracted), retrieved knowledge fragments, operational history data, multi-turn historical dialogue context, and the current user question are combined according to the fixed format of the selected prompt template to form standardized input prompts for generating the large language model.

[0081] Among them, role definition: specifies the role identity played by the large language model (such as LLM) (such as AI store manager, equipment maintenance expert, etc.) to constrain the structured answer style and professional field.

[0082] It should be noted that in this embodiment, the default role is defined as "AI Store Manager", but this role is not fixed and unique. The system supports configuring different role identities according to different usage scenarios. For example, it can be defined as "Equipment Maintenance Expert" when facing maintenance personnel, "Operation Data Analyst" when facing operation management personnel, and "Operation Training Instructor" when facing training scenarios.

[0083] In this embodiment, the role definition is an optional input field for the Prompt template, not a mandatory one. When no role is specified, the LLM will answer as a general assistant. However, in practice, a clear role definition can significantly improve the professionalism and consistency of the LLM's answers. Therefore, it is generally recommended to set a role when applying the template. In other words, the Prompt template supports dynamic role configuration to reflect the flexibility and scalability of this feature.

[0084] In step 105, a large language model is invoked based on the standardized input prompts to perform inference calculations and generate a structured answer.

[0085] In this step, the standardized input prompts are input into a large language model via an application programming interface (API) for inference operations. The large language model loads multi-source contextual information from the standardized input prompts and performs semantic parsing in conjunction with the natural language question to determine the maintenance user's needs. Based on the maintenance user's needs, a structured answer is generated, including at least one of the following: natural language response text, device control instructions, and associated shortcut operation matching. The structured answer is then output.

[0086] In this step, a Large Language Model (LLM) is invoked to infer the standardized input prompts and generate a structured answer that includes at least one of the following elements: (1) Natural language response text: Direct answers to user questions, with concise and clear language, suitable for maintenance personnel to read quickly; (2) Equipment control instructions: These are structured executable instructions (including operation type and parameters), which need to be confirmed by maintenance personnel before the system calls and executes the actual equipment actions through the remote control API. Among them, the equipment control instructions correspond to the "structured instruction parsing" and "equipment control instruction issuance" nodes.

[0087] (3) Operation suggestions: When the problem involves equipment operation, specific operation suggestions are generated. These are suggested text outputs, formatted as executable operation cards for display. The "Recommended related quick operations" nodes are displayed to the operation and maintenance personnel as quickly executable related operation suggestions pushed to them. The operation suggestions are presented in natural language, such as "It is recommended to check the remaining amount of bean hopper A" and "It is recommended to restart the bean grinding module". These are auxiliary suggestions generated by LLM based on context analysis. The operation and maintenance personnel can decide whether to execute them themselves. The system will not automatically trigger equipment actions.

[0088] (4) Recommended shortcut operations: Based on the current dialogue context, recommend 2 to 3 related shortcut operation buttons to reduce the user's operation path.

[0089] In other words, in this embodiment, when the structured response only includes a natural language response (i.e., includes textual descriptions of operational suggestions but no executable device control commands): the system directly displays the formatted response to the maintenance personnel, while recommending relevant quick operations for reference. The process ends here, and then the dialogue record is archived.

[0090] When the structured response includes equipment control instructions: the system enters the structured instruction parsing step, extracts the operation type and parameters, and waits for the maintenance personnel to confirm before issuing the equipment control instructions, i.e., steps 106 and 107.

[0091] Of course, a structured response can include all four elements mentioned above. For example, the structured response output by LLM includes the natural language response body, operation suggestions, relevant quick operation recommendations, and executable device control commands. In this case, the system prioritizes processing the device control command branch (execution step 106) while simultaneously displaying the natural language portion to the operations and maintenance personnel. Then, it processes the remaining natural language response body, operation suggestions, and relevant quick operation recommendations branches.

[0092] The standardized input prompts are input into the large language model for inference operations via the application programming interface (API). Combined with real-time device operating status data, historical operating data, knowledge fragments retrieved from the knowledge base, and historical dialogue context, a structured response text is generated, which includes at least one of the following: answer content (i.e., natural language response text), device control instructions, operation suggestions, and related quick operation recommendations.

[0093] In step 106, when the structured response includes device control instructions, the device control instructions are parsed and securely verified.

[0094] In this step, it is first determined whether the structured response includes device control instructions. If it does, the device control instructions are parsed to obtain the parsing result. The operation type, operation target, and operation parameters are extracted from the parsing result. Permission confirmation for executing the operation type, operation target, and operation parameters is initiated to the operation and maintenance user. When the operation and maintenance user confirms the operation, security is confirmed. Alternatively, when the operation and maintenance user cancels the operation, the dialogue record is archived and synchronized to the cloud.

[0095] When the structured response output by the LLM includes device control instructions, the system executes the following security procedures: 1) Structured Instruction Parsing: Extracting the operation type (e.g., "Start Cleaning", "Adjust Parameters"), operation target (e.g., "Coffee Brewing System"), and operation parameters (e.g., specific values) from the LLM response. 2) Permission Verification: Checking whether the current user account has the permission to execute the operation. 3) Secondary Confirmation by Maintenance User: Displaying an operation summary on the APP interface, requiring maintenance personnel to explicitly confirm before execution to prevent erroneous operations caused by LLM misjudgment.

[0096] It should be noted that in this embodiment, regardless of whether the operation and maintenance user confirms or cancels the operation, a dialogue record will be triggered after each scenario is completed and synchronized to the cloud for archiving.

[0097] In this step, parsing the device control instructions in the structured response involves splitting the text of the device control instructions into fields and extracting instruction information such as operation type, operation target, and operation parameters.

[0098] In this embodiment, the operation type refers to the category of actions that the unmanned beverage robot can perform, such as equipment start-up and shutdown, raw material replenishment, pipeline cleaning, refrigeration temperature adjustment, and modification of dispensing parameters. The operation target is to clearly define the equipment object to be controlled, which can be distinguished as a single robot, multiple devices in a cluster, or independent modules of the equipment (raw material silo, refrigeration module, cleaning components, etc.). The operation parameters are the quantitative execution conditions corresponding to the actions, including material filling amount, running time, temperature value, cleaning level, dispensing flow rate, and other values ​​and configuration information.

[0099] In other words, in this step, when determining if a device control command is involved, the command is parsed, and the operation type, operation target, and various operation parameters (i.e., natural language commands) are extracted from the parsing results. These extracted natural language commands are then converted into machine-readable structured fields. Afterward, a pop-up window is sent to the maintenance personnel to confirm the execution of the control command. If the maintenance personnel choose to cancel the operation, the conversation is recorded and archived to the cloud, and the process ends. If the maintenance personnel choose to confirm the operation, the system verifies the maintenance personnel's permissions, and after successful verification, calls the remote control API to issue the device control command to the unmanned beverage robot.

[0100] When it is determined that no device control commands are included, the formatted Q&A content will be directly displayed to the maintenance personnel; it can also push relevant operation suggestions that can be quickly executed to the maintenance personnel based on the current device status; and the complete dialogue record of this round will be archived to the cloud, and then jump to the end of this round of dialogue.

[0101] In step 107, after safety confirmation, the device control command is sent to the unmanned beverage robot through the remote control interface to control the unmanned beverage robot to perform the corresponding operation and maintenance.

[0102] In this step, after safety is confirmed, the device control command is issued. That is, the device control interface is called through the remote control API to send the device control command to the unmanned beverage robot, and control the unmanned beverage robot to perform the corresponding operation and maintenance operation.

[0103] This step implements an end-to-end linkage execution mechanism from natural language to device control. Specifically, it extracts executable operation instructions from LLM responses through structured instruction parsing, and combines these with security mechanisms such as secondary confirmation by maintenance users and permission verification to achieve end-to-end execution from natural language instructions to remote device control. This is an innovative application that extends LLM capabilities to the field of IoT device control.

[0104] In this embodiment, a natural language question input by the maintenance user is obtained; multi-source contextual information of the unmanned beverage robot related to the natural language question is determined, including: real-time operating status data of the device, historical operating data, knowledge fragments, and multi-turn historical dialogue context; a corresponding prompt word template is selected based on the intent classification result of the natural language question; the multi-source contextual information and the natural language question are concatenated according to the selected prompt word template to generate standardized input prompt words; a large language model is invoked based on the standardized input prompt words to perform inference calculations and generate a structured answer; when the structured answer includes device control instructions, the device control instructions are parsed and securely confirmed; after security confirmation, the device control instructions are sent to the unmanned beverage robot through a remote control interface to control the unmanned beverage robot to perform the corresponding maintenance operation. In other words, in this embodiment, the maintenance user does not need to remember complex operation paths and can complete device status queries, operational data analysis, and device control through natural language dialogue, improving operational efficiency; real-time perception of device status improves the accuracy of the answer. This embodiment, through multi-source contextual information, enables the large language model to generate structured answers based on the current real state of the device, avoiding the illusion problem caused by the lack of real-time data in the large language model. This embodiment combines a device knowledge base and real-time device status data. The large language model provides accurate cause analysis and handling suggestions for common faults, resulting in significant fault diagnosis assistance. This embodiment effectively prevents erroneous operations caused by misjudgments from the large language model through a dual security mechanism of secondary user confirmation and permission verification.

[0105] Optionally, in another embodiment, based on the above embodiments, the method may further include: obtaining the execution result of the device control command, the execution result including success or failure and the reason for failure, and sending the execution result to the operation and maintenance user.

[0106] In this step, after the unmanned beverage robot executes the equipment control command, it will send the execution result back to the system, clearly indicating whether the execution was successful, failed, and the corresponding reason for the failure, so that the operation and maintenance personnel can intuitively view the implementation status of the operation.

[0107] Optionally, in another embodiment, based on the above embodiments, the method may further include: recording the valid question and answer content of the current round of dialogue and updating the historical dialogue context; and synchronizing the recorded valid question and answer content of the current round of dialogue to the cloud, so that the cloud updates the knowledge base based on the valid question and answer content of the current round of dialogue.

[0108] In this embodiment, after each round of dialogue, the system uploads the complete dialogue record (user input, system response, and operation execution record) to a cloud database (such as a cloud log database) via a TLS 1.3 encrypted channel, supporting queries by account, time, device, intent type, and other dimensions. For valid responses from operations and maintenance personnel, the system automatically adds the corresponding question-and-answer pair to the knowledge base, enabling continuous optimization and updates to the knowledge base.

[0109] In other words, in this embodiment, all information such as user questions, LLM responses, command operations, and execution results in this round is summarized to generate archived data of valid question and answer content for this round of dialogue, and the global dialogue context history is updated synchronously to provide historical conversation basis for the prompt splicing of the next round of dialogue.

[0110] In this embodiment, high-quality questions and answers are automatically added to the knowledge base by the operation and maintenance personnel providing feedback on the effectiveness of the LLM's answers, thereby achieving continuous self-optimization of the knowledge base and enabling the intelligent operation assistant to continuously improve the accuracy of its answers as operational experience accumulates.

[0111] Please also see Figure 2 This is a schematic diagram illustrating the application of an intelligent control method for an unmanned beverage robot provided in an embodiment of this application, specifically including: Step 201: The operations and maintenance personnel input the natural language question.

[0112] In this step, maintenance personnel input their requests for the unmanned beverage robot in natural language form (i.e., natural language questions or user questions) at the front end of the intelligent operation assistant. The requests can cover three types of scenarios: equipment information query, remote control request, and equipment fault diagnosis, but are not limited to these. The system receives the text of the natural language question and performs basic text preprocessing, such as removing invalid spaces and special symbols, as the original input for this round of interaction.

[0113] Step 202: Intent classification (query, operation, fault diagnosis and training, etc.).

[0114] In this step, an intent classifier or text classification model can be used to perform semantic parsing on the input natural language question to complete intent segmentation. This embodiment takes four types of intent segmentation as an example: Query-based queries: device status query, operational data query, Q&A, etc., directly generating natural language answers; Control-related: Remote control commands for equipment (such as power on / off, parameter adjustment, cleaning start, etc.) need to be extracted into structured instructions and executed after user confirmation; Fault diagnosis: Based on the description of the fault phenomenon, combined with equipment status data and knowledge base, generate fault cause analysis and handling suggestions; Training: Operating procedure lookup, equipment knowledge learning, etc. Relevant content is retrieved from the knowledge base and presented in an easy-to-understand way.

[0115] Step 203: Multi-source context construction.

[0116] In this step, firstly, multiple types of information are collected in parallel using a hybrid retrieval architecture based on BM25 sparse retrieval and vector dense retrieval, specifically including: 1) Real-time acquisition of data on the remaining material quantity, operating parameters, current fault codes, and fault alarms of the unmanned beverage robot.

[0117] 2) Use RAG to retrieve relevant historical sales, revenue, production records, and other operational data related to the recall.

[0118] 3) Perform semantic retrieval of the knowledge base for the problem, and retrieve relevant fragments such as operation manuals and troubleshooting guides from the corresponding knowledge base through semantic retrieval (vector similarity matching) and BM25+ vector hybrid retrieval.

[0119] 4) Read the recent N rounds of historical dialogue context in the sliding window cache, that is, obtain the historical dialogue round records of the current session.

[0120] Afterwards, all the information collected above is formatted, redundant is filtered, and conflict is checked, and then integrated to form the complete multi-source context information used in this round of reasoning.

[0121] Step 204: Prompt assembly and LLM inference.

[0122] In this step, firstly, the corresponding Prompt template is retrieved based on the intent classification results. Multi-source contextual information and the user's question are then concatenated according to the preset format of the selected prompt word template to generate standardized prompt words for the large language model. The Prompt template is pre-designed for different intent types (such as query / operation / fault diagnosis / training) to ensure that the LLM generates consistent and accurate answers in different scenarios.

[0123] Then, the standardized prompt is transmitted to the server of the large language model (such as LLM) by calling the API through the large language model, and parameters such as inference temperature, maximum output length, and structured output constraints are configured to execute inference; The large language model combines multi-source contextual information from standardized prompts with user questions to generate structured answers. These structured answers can include: natural language responses, optional operation suggestions, related shortcut recommendations, and device control commands. See the above for details, which will not be repeated here.

[0124] Step 205: Determine whether the structured response includes device control instructions. If yes, proceed to step 206; otherwise, proceed to step 207.

[0125] The system parses the structured fields of the LLM output structured answer and identifies whether the answer contains device control instructions that can be sent to the unmanned beverage robot. If so, proceed to step 206; otherwise, proceed to step 207.

[0126] Step 206: Perform structured instruction parsing on the equipment control commands to obtain parsing results including: operation type, operation target, and operation parameters.

[0127] In this step, the device control command text is split into fields to extract the operation type (such as "start cleaning", "adjust parameters", etc.), operation target (such as "coffee brewing system") and operation parameters (such as specific values), and then step 208 is executed.

[0128] Step 207: Format and format the structured answers on the front end and display them to the operations and maintenance personnel. Then, directly jump to the dialogue record archiving stage, i.e., execute step 212.

[0129] Step 208: Determine whether the current user account of the operations and maintenance personnel has the permission to execute the corresponding operation in the parsing result. If yes, proceed to step 209; otherwise, proceed to step 210.

[0130] In this step, a pop-up window is displayed to the operations and maintenance personnel to show them the complete operation details, and the operations and maintenance personnel decide whether to execute the operation.

[0131] Its purpose is to verify permissions, checking whether the current user account has the authority to perform the operation. Secondary confirmation from operations and maintenance users involves displaying an operation summary on the app interface, requiring explicit confirmation from operations and maintenance personnel before execution, thus preventing erroneous operations caused by LLM misjudgments.

[0132] Step 209: Issue device control commands (remote control API call), then proceed to step 211.

[0133] In this step, the device control interface is called via remote control API to send device control commands to the unmanned beverage robot, thereby controlling the unmanned beverage robot to perform the corresponding operations.

[0134] In one embodiment, the device control command is encapsulated according to the device communication protocol, and the remote control API is called to send the encapsulated device control command to the device controller of the unmanned beverage robot so as to control the unmanned beverage robot to perform the corresponding operation.

[0135] Step 210: The operations and maintenance personnel select "Cancel" to directly archive the conversation record to the cloud and end this process; Step 211: Feedback on execution results (success, failure, and reason).

[0136] In this step, the unmanned beverage robot feeds back the operation results (including success or failure and the reason) to the LLM, and the LLM generates the final answer.

[0137] In other words, after the unmanned beverage robot receives the device control command and completes the action, it sends back the execution result through a remote API. If the execution is successful, it returns status update data; if the execution fails, it returns a fault code and the specific reason for the failure, so that the execution result can be pushed to the front end to be displayed to the operation and maintenance personnel.

[0138] It should be noted that if the operations and maintenance personnel choose to cancel the execution in this step, the control process will be terminated and the dialogue round record and archive will be executed directly. Step 212: Record the dialogue rounds and update the context history (i.e., the historical context).

[0139] In this step, after each round of dialogue, the system uploads the complete dialogue record (user input questions, system answers, operation execution records, etc.) to the cloud log database via a TLS 1.3 encrypted channel, supporting queries by account, time, device, intent type, and other dimensions. For valid answers provided by operations and maintenance personnel, the system automatically adds the corresponding questions and answers to the knowledge base, enabling continuous self-optimization of the knowledge base.

[0140] In other words, it summarizes complete information on user questions, LLM answers, command issuance, and execution results in the current round, and updates the historical dialogue cache in the sliding window, i.e., updates the historical context: if the number of dialogue rounds reaches the window limit N, the earliest round of dialogue is automatically removed to ensure semantic coherence in multi-round conversations, while controlling the context length.

[0141] The dialogue records are archived, and the full interaction log of this round is uploaded to the cloud for persistent storage; the Q&A content is deduplicated and its validity is filtered, and reusable operation and maintenance experience is automatically updated to the operation and maintenance knowledge base to achieve knowledge accumulation and iteration.

[0142] 213: Waiting for the next round of input.

[0143] This dialogue session has ended, and the system has returned to idle listening state, waiting for the operations and maintenance personnel to initiate the next natural language question.

[0144] In this embodiment, a multi-source context RAG architecture is constructed based on real-time device status data, operational history data, knowledge fragments from the knowledge base, and dialogue history context. This application is the first to dynamically inject real-time device status data of an unmanned beverage robot into the LLM context. Combined with operational history data retrieval and semantic retrieval from the device knowledge base, a multi-source context RAG architecture is built for unmanned beverage robot operation scenarios, enabling the LLM to generate accurate and real-time operational suggestions based on the current real-time status of the device. This embodiment also classifies user questions by intent, driving a differentiated Prompt template engine. For four different intents—query, operation, fault diagnosis, and training—corresponding dedicated Prompt templates are pre-designed. The intent classifier dynamically selects the corresponding dedicated Prompt template, ensuring that the LLM generates consistent and accurate structured answers in different scenarios, significantly improving answer quality and usability.

[0145] This application's embodiments extract executable operation instructions from LLM responses through structured instruction parsing. Combined with security mechanisms such as secondary confirmation by operation and maintenance users and permission verification, it realizes an end-to-end linkage execution mechanism from natural language instructions to remote device control. This is an innovative application that extends LLM capabilities to the field of IoT device control. This application automatically adds high-quality questions and answers to the knowledge base based on feedback from operation and maintenance personnel regarding the effectiveness of LLM responses, achieving continuous self-optimization of the knowledge base and enabling the intelligent operation assistant to continuously improve the accuracy of its answers as operational experience accumulates.

[0146] Optionally, in a specific application example, this embodiment of the application takes the use of an unmanned beverage robot operation and management APP (XBOT Claw) as an example for system verification. The numbers in this example are only illustrative and are not limited to this in actual applications. The process specifically includes: 1) Experimental setup: The intelligent operation assistant system of this application was integrated into the AI ​​store manager tab of the XBOT Claw APP to conduct actual use verification for 3 months (90 days) in the operation and management scenario of 30 test devices. A total of 47 operation and maintenance personnel participated in the test.

[0147] 2) Question and Answer Accuracy Verification: 1,000 dialogues were randomly selected from actual usage records, and professional maintenance personnel evaluated the accuracy of the answers. The accuracy rates were as follows: equipment status query: 97.3%; fault diagnosis: 91.2%; operational data analysis: 94.7%; operating procedure query: 96.1%; and overall accuracy: 94.8%.

[0148] 3) Verification of control command execution: During the test, a total of 1,247 device control commands were initiated (triggered through AI store manager dialogue). The success rate of structured command parsing was 99.1%, and the success rate of execution after user confirmation was 98.7%. No device misoperation events caused by LLM misjudgment occurred.

[0149] 4) Operation and maintenance efficiency verification: Compared with the control group that did not use the AI ​​store manager function, the number of daily APP operations by operation and maintenance personnel using the AI ​​store manager decreased by 37.2% (from an average of 143 times / day to 89.8 times / day), and the equipment anomaly response time was shortened from an average of 26 minutes to 11 minutes, a reduction of 57.7%.

[0150] 5) Knowledge base self-optimization verification: During the 3-month testing period, through the dialogue-driven knowledge accumulation mechanism, the number of knowledge base items increased from the initial 1,247 to 2,891, and the user satisfaction (out of 5) of LLM answers increased from the initial 3.91 to 4.68.

[0151] In this embodiment, maintenance personnel do not need to memorize complex operation paths; they can complete equipment status queries, operational data analysis, and equipment operation through natural language dialogue, significantly reducing the operational threshold for maintenance personnel and improving operational efficiency. In application testing, the onboarding time for newly hired maintenance personnel was shortened from an average of 3 days to 0.5 days, a reduction of approximately 83%.

[0152] In this embodiment, through a multi-source context RAG architecture, the LLM can generate answers based on the current real state of the device, achieving real-time device state awareness and high answer accuracy, thus avoiding the illusion problem caused by the lack of real-time data in general LLMs. In application testing, the accuracy rate of question-and-answer questions involving device state reached 97.3% (compared to 41.2% for the general LLM without injected device state data).

[0153] In this embodiment, by combining the device knowledge base and real-time status data, LLM can provide accurate cause analysis and handling suggestions for common faults, significantly improving the fault diagnosis assistance effect. In application testing, the first-time correct handling rate of Level 1 faults (which can be handled automatically) increased from 62.3% to 89.7%, and the average fault handling time was shortened from 38 minutes to 14 minutes.

[0154] In this embodiment, a dual security mechanism of secondary confirmation by the operation and maintenance user and permission verification effectively prevents erroneous operations caused by LLM misjudgment, thus ensuring operational security. During application testing, no device malfunctions caused by LLM misjudgment occurred.

[0155] In this embodiment of the application, through a dialogue-driven knowledge accumulation mechanism, the number of knowledge base items increased from the initial 1,247 to 2,891 during the 3-month testing period, an increase of 132%, and the user satisfaction rate of LLM answers increased from the initial 78.3% to 93.6%, thus achieving continuous optimization and updating of the knowledge base.

[0156] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to this application.

[0157] Please also see Figure 3 This is a block diagram of an intelligent control device for an unmanned beverage robot provided in an embodiment of this application. The device is applied to an intelligent operation assistant system and includes: an acquisition module 301, a determination module 302, a selection module 303, a splicing module 304, a reasoning module 305, a confirmation module 306, and a control module 307. The acquisition module 301 is used to acquire natural language questions input by the operation and maintenance user; The determination module 302 is used to determine the multi-source context information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments and multi-turn historical dialogue context; Selection module 303 is used to select the corresponding prompt word template based on the intent classification result of the natural language question; The splicing module 304 is used to splice the multi-source context information and the natural language question according to the selected prompt word template to generate standardized input prompt words; The reasoning module 305 is used to call a large language model to perform reasoning calculations based on the standardized input prompts and generate a structured answer; The confirmation module 306 is used to parse and securely confirm the device control instructions when the structured response includes device control instructions. The control module 307 is used to send the device control commands to the unmanned beverage robot through the remote control interface after safety confirmation, and control the unmanned beverage robot to perform corresponding operation and maintenance operations.

[0158] Optionally, the determining module 302 includes: a collection module 401, a first retrieval module 402, a second retrieval module 403, a dialogue acquisition module 404, and a fusion module 405, the structural block diagram of which is shown below. Figure 4 As shown, where, The data acquisition module 401 is used to acquire real-time operating status data of the unmanned beverage robot based on the natural language question. The first retrieval module 402 is used to retrieve operational history data related to the natural language problem from a cloud database; The second retrieval module 403 is used to retrieve knowledge fragments related to the natural language problem from the knowledge base; Dialogue acquisition module 404 is used to acquire the context of multi-turn historical dialogues managed by the sliding window; The fusion module 405 is used to fuse the real-time operating status data, historical operating data, knowledge fragments and multi-turn historical dialogue context of the device to obtain multi-source context information.

[0159] Optionally, the selection module 303 includes: a semantic recognition module 501, an intent classification module 502, and a template selection module 503, the structural block diagram of which is shown below. Figure 5 As shown, where, The semantic recognition module 501 is used to perform semantic recognition on the natural language question and obtain the recognition result; The intent classification module 502 is used to classify the recognition result according to the intent to obtain the intent classification result, wherein the intent classification result includes at least one of the following: query, control, fault diagnosis and training. The template selection module 503 is used to select the corresponding prompt word template based on the intent classification result.

[0160] Optionally, the inference module includes: a model input module, a model inference module, and a model output module, wherein, The model input module is used to input the standardized input prompts into the large language model via an application programming interface (API). The model reasoning module is used to determine the maintenance user's needs by loading multi-source context information from the standardized input prompts into the large language model and performing semantic parsing in conjunction with the natural language question; based on the maintenance user's needs, it generates a structured response that includes at least one of the following: natural language response text, device control instructions, and related shortcut operation matching. The model output module is used to output the structured answer.

[0161] Optionally, the confirmation module 306 includes: a parsing module 601, an extraction module 602, an authorization confirmation module 603, and a security determination module 604, the structure of which is shown in the figure below. Figure 6 As shown, where, The parsing module 601 is used to parse the device control instructions when the structured answer includes device control instructions, and obtain the parsing result; Extraction module 602 is used to extract the operation type, operation target and operation parameters from the parsing results; The permission confirmation module 603 is used to initiate permission confirmation for the operation and maintenance user to execute the operation type, the operation target and the operation parameters; The security determination module 604 is used to confirm security when it receives confirmation from the operation and maintenance user.

[0162] Optionally, the confirmation module further includes a record synchronization module, used to archive the dialogue record and synchronize it to the cloud when the cancellation operation of the operation and maintenance user is received.

[0163] Optionally, the device further includes: The first sending module is used to display formatted Q&A content to the operation and maintenance user and push relevant operation suggestions that can be quickly executed when the structured answer does not include device control instructions.

[0164] Optionally, the device further includes: The result acquisition module is used to acquire the execution result of the device control command, and the execution result includes success or failure and the reason for failure; The second sending module is used to send the execution result to the operation and maintenance user.

[0165] Optionally, the device further includes: The record update module is used to record the valid question and answer content of the current round of dialogue and update the historical dialogue context; and The third sending module is used to synchronize the recorded valid question and answer content of the current round of dialogue to the cloud, so that the cloud can update the knowledge base based on the valid question and answer content of the current round of dialogue.

[0166] Optionally, embodiments of this application also provide an electronic device, including: It includes a processor, a memory; and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the intelligent control method for the unmanned beverage robot as described above.

[0167] Optionally, embodiments of this application also provide a readable storage medium storing a program or instructions, which, when executed by a processor of an electronic device, implement the steps of the intelligent control method for the unmanned beverage robot described above.

[0168] Optionally, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor of an electronic device, implement the steps of the intelligent control method for the unmanned beverage robot described above.

[0169] Please also see Figure 7 This is a schematic diagram of the architecture of an intelligent operation assistant system provided in an embodiment of this application. In this embodiment, the intelligent control method of the unmanned beverage robot is applied to the intelligent operation assistant system. As shown in the figure, the system is divided into an input layer, a RAG construction layer, an LLM inference layer, and an output layer from top to bottom. The bottom layer is equipped with a unified APP interaction layer. The data interaction process of each layer includes: ① Input layer: This layer aggregates multiple types of basic input information: natural language questions raised by operations and maintenance personnel (i.e., the questions raised by operations and maintenance personnel in the diagram), real-time equipment status data of unmanned beverage robots (i.e., equipment status data), historical operation data (retrieving historical operation data related to natural language questions from this historical data knowledge base), knowledge fragments (retrieving relevant knowledge fragments from the operations and maintenance-specific knowledge base), and multi-turn historical dialogue contexts. These are distributed in parallel to the downstream RAG construction layer as the original data source for subsequent retrieval and prompt word assembly.

[0170] ② RAG construction layer: This layer, based on all the data from the input layer, performs retrieval enhancement and suggestion word construction in parallel through three processing branches: 21) Semantic retrieval + data aggregation: Perform semantic retrieval on real-time equipment operating conditions (real-time equipment operating status data) and historical operation data to complete the cleaning and integration of multi-source business data; 22) Knowledge base retrieval + dialogue history: A hybrid retrieval method is used to retrieve highly relevant knowledge fragments from the operation and maintenance knowledge base, while retrieving the historical dialogue context managed by the sliding window to supplement the dialogue continuity information; 23) Prompt Template Engine (role + context + instruction): Combines user question intent with exclusive prompt word templates, and concatenates system role settings, aggregated multi-source contexts, and user question requests and instructions according to prompt word template specifications to generate standardized input prompt words (i.e. input text) for a large language model. Finally, the input prompts for the standardized large language model are generated and sent to the inference layer of the large language model LLM.

[0171] ③ Large Language Model (LLM) Inference Layer: The standardized input prompts are input into a large language model via an application programming interface (API). The large language model loads multi-source contextual information from the standardized input prompts and performs semantic parsing in conjunction with the natural language question to determine the maintenance user's needs. Based on the maintenance user's needs, a structured response is generated, including at least one of the following: natural language response text, device control instructions, and associated shortcut operation matching. The structured response is then output.

[0172] In other words, the Large Language Model (LLM) performs three core operations in sequence: intent recognition, instruction reasoning, and answer generation. First, it performs intent recognition on the questions asked by the operation and maintenance users, then performs professional knowledge reasoning based on the external knowledge obtained from retrieval, and finally generates structured answer content, i.e., the reasoning result, and passes the reasoning result down to the output layer.

[0173] The process of calling the Large Language Model (LLM) to perform inference calculations on the constructed Prompt is detailed above and will not be repeated here.

[0174] ④ Output layer: The inference results output by the large language model are split into streams: 41) Convert the answer into a natural language response and complete the formatted visual display on the front end, i.e., natural language response and formatted display; 42) Extract operation and maintenance suggestions, and conduct structured analysis of instructions, i.e., operation suggestions, structured analysis; 43) Parse the obtained legitimate equipment control commands, issue and execute the equipment linkage control, that is, the equipment commands are executed in linkage; All processing results from the output layer are pushed to the bottom-level APP interaction layer.

[0175] APP interaction layer: Serving as a unified interactive portal for operations and maintenance personnel, the app integrates five major interactive modules: an AI store manager dialogue interface, quick problem recommendations, operation suggestion cards, remote control execution, and historical conversation viewing. This interactive layer centrally realizes full-process front-end interactive capabilities, including question-and-answer interaction, operation guidance, remote device control, and conversation record querying, thus achieving the interactive implementation of the entire intelligent operations assistant function.

[0176] In this embodiment, as shown in the architecture diagram of the intelligent operation assistant system for unmanned beverage robots based on a large language model, the system is divided into an input layer, a RAG construction layer, an LLM inference layer, an output layer, and an APP interaction layer. The input layer provides user questions, device operation data, and raw knowledge base data; the RAG construction layer retrieves multi-source information through hybrid retrieval and generates standardized prompts by matching historical dialogue templates; the LLM inference layer completes intent recognition and structured response generation; the output layer breaks down the response content, parses device control commands, and sends them to the device; finally, all interactive content is displayed in the APP interaction layer, and the interaction log synchronously updates the knowledge base for knowledge base updates and historical context maintenance, realizing intelligent question-and-answer and remote operation and maintenance control of the unmanned beverage robot.

[0177] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0179] Figure 8 This is a block diagram of an electronic device 800 provided in an embodiment of this application. For example, the electronic device 800 can be a mobile terminal or a server; in this embodiment, a mobile terminal is used as an example for explanation. For example, the electronic device 800 can be an unmanned beverage robot, etc.

[0180] Reference Figure 8 The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0181] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0182] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0183] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0184] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the maintenance user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the maintenance user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0185] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0186] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0187] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in the position of electronic device 800 or a component of electronic device 800, the presence or absence of contact between maintenance users and electronic device 800, the orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0188] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0189] In the embodiments, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the various processes of the intelligent control method embodiments of the unmanned beverage robot described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0190] In this embodiment, a readable storage medium is also provided, on which a program or instructions are stored. When executed by a processor of an electronic device, the program or instructions implement the steps of the intelligent control method for the unmanned beverage robot described above. The readable storage medium includes computer-readable storage media, such as ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices.

[0191] In this embodiment, a computer program product is also provided, including a computer program or instructions. When the computer program or instructions are executed by the processor 820 of the electronic device 800, the electronic device 800 executes the various processes of the intelligent control method embodiment of the unmanned beverage robot described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0192] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0193] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0194] Figure 9 This is a block diagram of a device 900 for intelligent control of an unmanned beverage robot, provided in an embodiment of this application. For example, device 900 can be provided as a server. (See also...) Figure 9 The apparatus 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by memory 932 for storing instructions, such as application programs, that can be executed by the processing component 922. The application programs stored in memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 922 is configured to execute instructions to perform the methods described above.

[0195] Device 900 may also include a power supply component 926 configured to perform power management of device 900, a wired or wireless network interface 950 configured to connect device 900 to a network, and an input / output (I / O) interface 958. Device 900 can operate on an operating system stored in memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0196] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0197] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An intelligent control method for an unmanned beverage robot, characterized in that, The method is applied to an intelligent operation assistant system, and the method includes: Obtain natural language input from operations and maintenance users; Determine the multi-source contextual information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context; Select the corresponding prompt word template based on the intent classification result of the natural language question; The multi-source context information and the natural language question are concatenated according to the selected prompt word template to generate standardized input prompt words; Based on the standardized input prompts, a large language model is invoked to perform inference calculations and generate a structured answer; When the structured response includes device control commands, the device control commands are parsed and securely verified. After safety confirmation, the device control command is sent to the unmanned beverage robot through the remote control interface to control the unmanned beverage robot to perform the corresponding operation and maintenance.

2. The intelligent control method for the unmanned beverage robot according to claim 1, characterized in that, The multi-source contextual information for determining the unmanned beverage robot related to the natural language problem includes: Based on the aforementioned natural language problem, real-time operational status data of the unmanned beverage robot is collected. Retrieve operational history data related to the natural language problem from the cloud database; Retrieve knowledge fragments related to the natural language problem from the knowledge base; Retrieve the context of multi-turn historical dialogues managed by the sliding window; The real-time operating status data, historical operating data, knowledge fragments, and multi-turn historical dialogue context of the device are fused to obtain multi-source context information.

3. The intelligent control method for the unmanned beverage robot according to claim 1, characterized in that, The selection of the corresponding prompt word template based on the intent classification result of the natural language problem includes: Semantic recognition is performed on the natural language problem to obtain the recognition result; The recognition results are classified according to the intent to obtain the corresponding intent classification results, wherein the intent classification results include at least one of the following: query, control, fault diagnosis and training. Select the corresponding prompt word template based on the intent classification results.

4. The intelligent control method for the unmanned beverage robot according to claim 1, characterized in that, The step of invoking a large language model based on the standardized input prompts to generate a structured answer includes: The standardized input prompts are input into a large language model via an application programming interface (API). The large language model loads multi-source contextual information from the standardized input prompts and performs semantic parsing in conjunction with the natural language question to determine the maintenance user's needs. Based on the maintenance user's needs, a structured answer is generated, including at least one of the following: natural language response text, device control instructions, and associated shortcut operation matching. The structured answer is then output.

5. The intelligent control method for the unmanned beverage robot according to claim 1, characterized in that, When the structured response includes device control commands, the device control commands are parsed and securely verified, including: When the structured response includes device control instructions, the device control instructions are parsed to obtain the parsing result; Extract the operation type, operation target, and operation parameters from the parsing results; Initiate permission confirmation with the operation and maintenance user to execute the operation type, the operation target, and the operation parameters; Upon receiving confirmation from the operations and maintenance user, security is confirmed.

6. The intelligent control method for the unmanned beverage robot according to claim 1, characterized in that, The method further includes: When the structured response does not include device control instructions, the system displays formatted Q&A content to the maintenance user and pushes relevant operation suggestions that can be quickly executed.

7. The intelligent control method for the unmanned beverage robot according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the execution result of the device control command, the execution result including success or failure and the reason for failure, and send the execution result to the maintenance user; and / or Record the valid questions and answers in this round of dialogue and update the historical dialogue context; and synchronize the recorded valid questions and answers in this round of dialogue to the cloud so that the cloud can update the knowledge base based on the valid questions and answers in this round of dialogue.

8. An intelligent control device for an unmanned beverage robot, characterized in that, The device is used in an intelligent operation assistant system, and the device includes: The acquisition module is used to acquire natural language questions input by operation and maintenance users; The determination module is used to determine the multi-source context information of the unmanned beverage robot related to the natural language problem, including: real-time operating status data of the device, historical operation data, knowledge fragments, and multi-turn historical dialogue context; The selection module is used to select the corresponding prompt word template based on the intent classification result of the natural language question. The concatenation module is used to concatenate the multi-source context information and the natural language question according to the selected prompt word template to generate standardized input prompt words; The reasoning module is used to call a large language model to perform reasoning calculations based on the standardized input prompts and generate structured answers. The confirmation module is used to parse and securely confirm the device control instructions when the structured response includes device control instructions. The control module is used to send the device control commands to the unmanned beverage robot through a remote control interface after safety confirmation, so as to control the unmanned beverage robot to perform corresponding operation and maintenance operations.

9. An electronic device, characterized in that, include: Including processor and memory; And a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the intelligent control method for the unmanned beverage robot as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The program or instructions stored on the readable storage medium implement the steps of the intelligent control method for the unmanned beverage robot as described in any one of claims 1 to 7 when executed by the processor of the electronic device.