Robot cluster operation and maintenance management method and system, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610710022.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]本申请实施例提供一种机器人集群的运维管理方法、系统、电子设备和存储介质,以解决现有技术中故障处理效率低的问题
[0006] This application provides a method, system, electronic device, and storage medium for the operation and maintenance management of a robot cluster, in order to solve the problem of low fault handling efficiency in the prior art.
Smart Images

Figure CN122601447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a method, system, electronic device, and storage medium for the operation and maintenance management of a robot swarm. Background Technology
[0002] With the development of industrial automation and smart logistics, automated warehousing systems based on robot swarms have become core infrastructure. In automated warehousing systems, the scale of robot swarms can typically reach hundreds or even thousands of units, and the stable operation of the robot swarm directly determines the overall efficiency of warehousing operations.
[0003] However, fault handling in existing robot swarms mainly relies on human intervention. For example, when a robot malfunctions, the abnormality can only be indicated by external audio-visual signals (such as flashing indicator lights and alarm sounds). For simple faults (such as low battery or lost QR codes), on-site personnel can make a preliminary judgment based on experience; but for complex faults (such as navigation deviation, abnormal communication in the robot scheduling system, or hidden hardware errors), external signals cannot directly pinpoint the cause.
[0004] Upon discovering a complex fault, an on-site engineer needs to connect to the robot, retrieve its log files using specialized tools, and then send the logs to the R&D engineer for analysis. The R&D engineer will analyze the logs and provide repair suggestions (such as restarting the robot, adjusting the path, or replacing hardware), which will then be executed by on-site personnel.
[0005] The above-mentioned fault handling process has a long time cycle. From the occurrence of the fault to the recovery, it needs to go through multiple stages, which takes a long time and seriously affects the continuity of warehousing operations. Moreover, it requires high technical skills from on-site personnel. Fault handling is passive and cannot predict potential risks in advance (such as insufficient power but no alarm is triggered, or network delay does not affect the task in the early stage). As a result, manual intervention is only carried out after the fault has expanded, which is not conducive to ensuring the stable operation and work efficiency of the robot cluster. Summary of the Invention
[0006] This application provides a method, system, electronic device, and storage medium for the operation and maintenance management of a robot cluster, in order to solve the problem of low fault handling efficiency in the prior art.
[0007] This application provides a method for the operation and maintenance management of a robot cluster, applied to an operation and maintenance intelligent agent, including: In response to a triggering event, a status query is initiated to the robot scheduling system. Triggering events include receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check. Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information; If the status information of the faulty robot indicates a network fault, a network fault notification message is sent to the terminal device. If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot. Based on the status information and log data of the faulty robot, generate and output the troubleshooting results.
[0008] In one implementation, based on the status information and log data of the faulty robot, fault diagnosis results are generated and output, including: The system matches the status information and log data of the faulty robot with pre-stored fault investigation rules, or calls a large model to analyze the log data of the faulty robot, and generates and outputs the fault investigation results.
[0009] In one implementation, based on the status information and log data of the faulty robot, fault diagnosis results are generated and output, including: Based on the status information of the faulty robot, the log data is filtered by keywords to generate valid fault logs related to the fault. Match valid fault logs with pre-stored fault investigation rules, or call a large model to analyze valid fault logs, and generate and output fault investigation results.
[0010] In one implementation, the steps of calling a large model to analyze the state information and valid log data of the faulty robot and generate fault diagnosis results include: Based on valid fault logs, pre-stored fault troubleshooting forms, and the status information of the faulty robot, generate input prompt words; Based on the input prompts, the system calls a large model, retrieves the fault causes and troubleshooting suggestions returned by the large model, and generates fault troubleshooting results.
[0011] In one implementation, the operations and maintenance agent runs in a virtualized container and interacts with robots, robot scheduling systems, and external communication platforms through preset interfaces.
[0012] In one implementation, the operation and maintenance agent is configured with long-term memory rules. These rules solidify read-only observation constraints, allowing status queries and log retrieval operations while prohibiting change operations on the robot or robot scheduling system.
[0013] In one implementation, after outputting the troubleshooting results, the following is also included: If the fault type identified by the fault investigation results is a Type I fault, a fault repair plan is generated based on the fault investigation results. Type I faults are those where the operation and maintenance intelligent agent has the necessary operation authorization and functional support. Execute the fault repair plan and output information about the executed fault repair plan.
[0014] In one implementation, after outputting the troubleshooting results, the following is also included: If the fault type identified by the fault investigation results is the second type of fault, a fault repair plan and a repair request are generated based on the fault investigation results. The second type of fault is a fault in which the operation and maintenance intelligent agent has functional support but no operation authorization. Output the fault repair plan and repair request; Receive confirmation information for the repair request and execute the fault repair plan.
[0015] In one implementation, after outputting the troubleshooting results, the following is also included: If the fault type identified by the fault investigation results is a third type of fault, a repair request will be generated and output. The third type of fault is a fault in which the operation and maintenance intelligent agent has no operation authorization and no functional support. Receive execution information and respond to the execution information to perform repair actions. The execution information includes repair actions generated based on the fault diagnosis results.
[0016] This application embodiment also provides a robot cluster operation and maintenance management system, including terminal equipment, robot scheduling system, robots and servers, with an operation and maintenance intelligent agent running in the server or terminal equipment; The operations and maintenance intelligent agent is configured as follows: In response to a triggering event, a status query is initiated to the robot scheduling system. Triggering events include receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check. Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information; If the status information of the faulty robot indicates a network fault, a network fault notification message is sent to the terminal device. If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot. Based on the status information and log data of the faulty robot, generate and output the troubleshooting results.
[0017] This application also provides an electronic device, including: Memory, used to store instructions; and The processor is used to execute the operation and maintenance management methods of the robot cluster by calling the instructions stored in the memory.
[0018] This application also provides a computer-readable storage medium storing computer-executable instructions. When the processor executes the computer-executable instructions, the above-mentioned operation and maintenance management method for the robot cluster is implemented.
[0019] This application also provides a computer program product, which, when read and executed by a computer, causes the above-mentioned operation and maintenance management method for the robot cluster to be executed. Attached Figure Description
[0020] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0021] Figure 1 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 1 .
[0022] Figure 2 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 2 .
[0023] Figure 3 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 3 .
[0024] Figure 4 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 4 .
[0025] Figure 5 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 5 .
[0026] Figure 6 A flowchart illustrating the operation and maintenance management method for the robot cluster provided in this application embodiment. Figure 6 .
[0027] Figure 7 This is a schematic diagram of the operation and maintenance management system for a robot cluster provided in an embodiment of this application. Detailed Implementation
[0028] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.
[0029] It should be noted that many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.
[0030] In this application, the use of terms such as "first," "second," etc., is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features.
[0031] The various embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0032] In large-scale automated warehousing scenarios based on flexible robots such as AGVs (Automated Guided Vehicles), AMRs (Autonomous Mobile Robots), CTUs (Carton Tote Unloaders), and four-way shuttles, the stable operation of the robot swarm determines the throughput efficiency of warehousing operations. In existing technologies, when robots malfunction, manual intervention is typically required. For simple faults, on-site personnel can make a preliminary judgment through external audio-visual signals; however, for complex faults, a lengthy process is required, involving on-site personnel accessing the robot to obtain logs, sending the logs to R&D engineers, R&D analysis and feedback of repair suggestions, and on-site personnel executing the repairs.
[0033] This traditional manual operation and maintenance model has significant drawbacks: First, it is time-consuming, involves multi-tasking collaboration, and has long fault recovery times, affecting operational continuity; second, it requires highly skilled on-site personnel who are familiar with robot log structures, interface calls, and fault troubleshooting processes, which is difficult to achieve in small and medium-sized warehouse scenarios; third, it lacks initiative, mostly responding passively after a fault occurs, and cannot anticipate potential risks in advance. To address these issues, this application provides a robot cluster operation and maintenance management method, system, electronic equipment, and storage medium to solve the problem of low fault handling efficiency in the prior art.
[0034] This application provides a method for the operation and maintenance management of a robot cluster, applied to an operation and maintenance intelligent agent, to realize the operation and maintenance management of the robot cluster. For example, the application scenario is an automated warehousing system, which can operate one or more different robots. The operating robots include, but are not limited to, handling robots that transport goods or pallets, bin robots that transport multiple bins simultaneously, or four-way shuttles in automated storage and retrieval systems.
[0035] Operations and maintenance agents can be installed on servers or terminal devices. For example, if the operations and maintenance agent is installed on a server, it can be added to instant messaging software such as WeChat, DingTalk, and Lark installed on terminal devices. Users can then send query requests to the operations and maintenance agent through the instant messaging software on their terminal devices.
[0036] If the operations and maintenance intelligent agent is installed on a terminal device, the agent window can be opened by creating an icon on the terminal device. Operators can then send query requests to the intelligent agent by inputting requests or commands through the touch components of the terminal device. It should be noted that other methods can also be used to send query requests to the operations and maintenance intelligent agent through the terminal device, and this application embodiment does not limit this method.
[0037] like Figure 1 As shown in the figure, the robot cluster operation and maintenance management method provided in this application includes steps S110 to S150.
[0038] Step S110: In response to the triggering event, initiate a status query to the robot scheduling system.
[0039] The triggering events include at least two types. One is receiving a query request from a terminal device. For example, an operator triggers a query request on a terminal device and sends it to the maintenance intelligence agent through the terminal device. This could be done via instant messaging software such as WeChat, DingTalk, or Lark installed on the terminal device. The other type is reaching a preset scheduled inspection cycle, where the maintenance intelligence agent triggers the event autonomously. For example, every five minutes, the maintenance intelligence agent initiates a status query to the robot scheduling system.
[0040] Status queries can include global robot status checks or target robot status checks, where there can be one or more target robots. Specifically, the content of the status query depends on the triggering scenario. For example, if the query request is for a specific device, it will be an individual status check of the target robot. Conversely, if the query request is for a global inspection, it will be a global robot status check, where "global robot" refers to all robots in the robot cluster.
[0041] Step S120: Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information.
[0042] The robot scheduling system maintains the real-time status of all robots, including their online status, task status, battery status, and location information. The operations and maintenance agent parses the returned robot status information; if it finds that the robot status information is marked as "faulty" or the parameters exceed the threshold (such as battery level below 10%), the robot is determined to be faulty.
[0043] In cases marked as faults, specific fault information is also displayed, such as network failure, position deviation, or task timeout. Task timeout may be due to various reasons, such as the robot not being in the specified task state, the specified task state not being started, or the specified task timeout.
[0044] Step S130: If the status information of the faulty robot indicates a network fault, then send a network fault prompt message to the terminal device.
[0045] If the status information of the faulty robot indicates a network failure, such as the robot scheduling system being unable to ping the robot's IP address (including network outage or abnormality of the upper-level scheduling system), the operation and maintenance intelligent agent will directly send a network failure prompt message to the terminal device, such as: "AGV-003's network is disconnected. Please check the network signal."
[0046] Step S140: If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot.
[0047] If the status information indicates a non-network fault, the operation and maintenance agent will initiate interaction with the faulty robot and retrieve the robot's hardware error logs and navigation logs through the SCP (Secure Copy Protocol) or HTTP (Hypertext Transfer Protocol) interface.
[0048] Non-network faults include, but are not limited to, robot task execution time exceeding the specified time, deviation from the target, encountering obstacles, QR code loss, and low battery.
[0049] For robot failures that time out during task execution, the operations and maintenance agent queries the scheduling system to check the status of other robots bound to the timed-out task. For example, based on task information, it obtains the mapping relationship between tasks and robots in the scheduling system to determine the other robots bound to execute the timed-out task. If the other robots bound to the timed-out task are not faulty, the status quo is maintained or the task priority is increased; if the other robots bound to the timed-out task are faulty, the specific fault handling process for that robot is initiated.
[0050] For malfunctions involving deviation, the maintenance intelligent agent queries the coordinate information reported by the fault robot to confirm the deviation and the current location.
[0051] When encountering obstacles, the operations and maintenance agent records the spatiotemporal information of the obstacle's appearance. This spatiotemporal information includes at least time information, spatial information, and associated context information. The time information is the timestamp when the obstacle was detected. The spatial information is the obstacle's position coordinates in the warehouse coordinate system and its corresponding area number (e.g., row 3 of area A). The associated context information is the relative distance between the obstacle and the robot, and the robot's current orientation.
[0052] In scenarios where robots use identification codes for navigation, if an identification code (such as a QR code) is lost, the maintenance agent records the error code and the current location information.
[0053] In the event of a low battery fault, the maintenance agent reads the robot's current battery percentage.
[0054] Step S150: Generate and output the troubleshooting results based on the status information and log data of the faulty robot.
[0055] The operations and maintenance intelligent agent generates and outputs troubleshooting results based on the acquired log data. For example, based on the status information of the faulty robot, it analyzes the log data of the target robot, generates a conclusion that the target robot's LiDAR occlusion caused the obstacle avoidance failure, and sends it to the operator's terminal device through the communication gateway.
[0056] This application embodiment proactively intervenes through an operation and maintenance intelligent agent, replacing the traditional manual polling and log acquisition process. This solves the technical problems of delayed fault detection and cumbersome manual log acquisition in the prior art, enabling rapid fault perception and preliminary location, and shortening the lead time for fault handling.
[0057] In some embodiments, fault diagnosis results are generated and output based on the log data of the faulty robot, including: matching the log data of the faulty robot with pre-stored fault diagnosis rules, or calling a large model to analyze the log data of the faulty robot and generate and output fault diagnosis results. For example, the large model call uses GPT (Generative Pre-trained Transformer), the input prompt word length does not exceed 2000 characters, and includes valid fault logs, fault diagnosis table index, and robot status information.
[0058] In some embodiments, after acquiring the log data of the faulty robot, the operations and maintenance agent performs a two-path analysis. The first path involves matching the log data against pre-stored fault diagnosis rules. These rules are stored in the agent's long-term memory as text files. For example, if the log contains an E-1024 error code (indicating motor overload), it is determined that the motor is overloaded. The second path involves calling a large-scale model to analyze the log data. When the pre-stored fault diagnosis rules cannot diagnose the fault, the operations and maintenance agent calls the API (Application Programming Interface) of the remote large-scale model. For example, for a deviation fault, the operations and maintenance agent inputs the coordinate offset data from the log into the large-scale model, which then analyzes the cause of the deviation.
[0059] In this way, deterministic matching based on fault diagnosis rules and generalized reasoning based on large models solve the technical problem that a single rule base cannot cover unknown faults, thereby improving the accuracy and coverage of fault diagnosis.
[0060] In some embodiments, such as Figure 2 As shown, based on the status information and log data of the faulty robot, the fault investigation results are generated and output, including steps S210 to S220.
[0061] Step S210: Based on the status information of the faulty robot, filter the log data by keywords to generate valid fault logs related to the fault.
[0062] After the operations and maintenance (O&M) agent retrieves log data from the robot itself, it filters the log data using keywords based on the robot's status information. Since the raw log data contains a large amount of redundant information (such as heartbeat packets and debugging information), the O&M agent utilizes its local file retrieval capabilities to extract valid fault logs related to the robot's status information. For example, in a fault scenario where a QR code is lost, the robot's status information is "QR code not found." The O&M agent filters out log segments containing "QR code not found" and its corresponding coordinates.
[0063] Step S220: Match the valid fault logs with the pre-stored fault investigation rules, or call the large model to analyze the valid fault logs, and generate and output the fault investigation results.
[0064] The operations and maintenance agent matches valid fault logs with pre-stored fault diagnosis tables, or calls a large model to analyze valid fault logs. For example, if a pre-stored fault diagnosis table records "error code 0x0A" (corresponding to a damaged QR code), the operations and maintenance agent will obtain the fault diagnosis result through matching. If no corresponding fault record is found, the agent combines valid log fragments, robot model, and pre-stored vehicle model fault diagnosis documents into input prompts, which are then input into the large model.
[0065] This application embodiment removes noise interference through log preprocessing, solving the technical problems of large original log data volume and difficulty in extracting key information, and improving the efficiency and accuracy of fault diagnosis results generation.
[0066] In some embodiments, such as Figure 3 As shown, the steps for calling a large model to analyze the status information and valid log data of the faulty robot and generate fault investigation results include steps S310 to S320.
[0067] Step S310: Generate input prompt words based on valid fault logs, pre-stored fault troubleshooting tables, and the status information of the faulty robot.
[0068] The operation and maintenance intelligent agent generates structured input prompts based on valid fault logs, pre-stored fault troubleshooting tables (i.e., local knowledge base), and the status information of the faulty robot (such as location and battery level). For example: Vehicle model CTU-200, currently located in row 3 of section A, the log shows "lift motor error", battery level 85%, please analyze the cause and provide repair suggestions.
[0069] Step S320: Based on the input prompt words, call the large model, obtain the fault cause and troubleshooting suggestions returned by the large model, and generate the fault troubleshooting results.
[0070] Based on the input prompts, the operations and maintenance (O&M) agent calls the interface of a large model deployed in the cloud. The large model, based on motor fault patterns from the training data, returns the cause of the fault (e.g., "Hall sensor failure in the lifting motor") and troubleshooting suggestions (e.g., "Check motor wiring or replace the sensor"). The O&M agent receives this result and generates the final fault diagnosis.
[0071] In this way, by leveraging the natural language understanding and reasoning capabilities of large models, we can solve the technical problem that complex fault causal relationships are difficult to deduce through hard-coded rules, and achieve expert-level fault diagnosis capabilities.
[0072] In some embodiments, the operation and maintenance intelligent agent runs in a virtualized container and interacts with robots, robot scheduling systems and external communication platforms through preset interfaces.
[0073] The operations and maintenance (O&M) agent runs in a virtualized container based on Docker technology, isolated from the host machine. The O&M agent interacts with robots, the robot scheduling system, and external communication platforms through pre-defined API interfaces. For example, the O&M agent can only query status through the interfaces provided by the robot scheduling system.
[0074] In this way, containerization and isolation can solve the technical problem that the operation and maintenance intelligent agent may be maliciously attacked or unauthorizedly operated due to code vulnerabilities, thereby reducing network security risks.
[0075] In some embodiments, the operation and maintenance agent is configured with long-term memory rules. The long-term memory rules solidify read-only observation constraints, allowing the execution of status query and log retrieval operations, and prohibiting the execution of change-type operations on the robot or robot scheduling system.
[0076] Specifically, the skill files for the operations and maintenance agent contain read-only interfaces, such as "get logs" and "query status," while strictly prohibiting interfaces that perform change operations such as "restart," "modify parameters," and "delete tasks." This way, even if the large model outputs a "suggest restarting the robot" command, the operations and maintenance agent will verify long-term memory rules before execution; if the command is a change operation, it will be automatically blocked.
[0077] This application's embodiments address the potential risks of misoperation in automated operations and maintenance by enforcing read-only rules, ensuring that the operations and maintenance process only involves observation and analysis without interfering with the normal operation of the robot, thus guaranteeing the stability and security of warehousing operations.
[0078] In some embodiments, such as Figure 4 As shown, after outputting the troubleshooting results, steps S410 to S420 are also included.
[0079] Step S410: If the fault type characterized by the fault investigation results is a Class I fault, generate a fault repair plan based on the fault investigation results.
[0080] The first type of fault refers to a fault for which the operation and maintenance intelligent agent has the necessary operational authorization and functional support. When the fault diagnosis result indicates a first-type fault, the operation and maintenance intelligent agent sends a fault repair plan to the robot scheduling system based on the pre-stored repair plan.
[0081] Step S420: Execute the fault repair plan and output the information that the fault repair plan has been executed.
[0082] For example, when troubleshooting indicates that the battery is too low and has not been scheduled for charging, the operations and maintenance agent, based on pre-stored repair plans, directly sends an emergency charging request to the scheduling system. If the scheduling system reports no available charging stations, the operations and maintenance agent further executes a resource allocation plan, that is, it sends a request to the scheduling system to suspend the charging tasks of robots with sufficient battery power, prioritizing the charging needs of robots with low battery power.
[0083] Among them, pausing the charging task of robots with sufficient power belongs to the internal resource allocation of the scheduling system. The operation and maintenance intelligent agent only sends a request and does not directly execute the change operation, which complies with the read-only rule.
[0084] For example, when a robot is detected to be stuck due to an obstacle, the operations and maintenance agent executes a pre-defined retry logic, setting it to wait 1-3 minutes and then check the obstacle status again, repeating this operation until a set number of times (e.g., 3 times) is reached. If the obstacle disappears, the task continues; if the obstacle still exists, the operations and maintenance agent notifies the scheduling system to attempt rerouting to avoid the obstacle.
[0085] After executing the above repair plan, the operation and maintenance intelligent agent sends a notification to the terminal device that charging scheduling has been performed or rerouting has been attempted. This achieves automated closed-loop processing of low-risk faults, resolving the inefficiency caused by the need for manual confirmation for simple faults and improving the working efficiency of the robot cluster.
[0086] In some embodiments, such as Figure 5 After outputting the troubleshooting results, steps S510 to S530 are also included.
[0087] Step S510: If the fault type characterized by the fault investigation results is a second type of fault, generate a fault repair plan and a repair request based on the fault investigation results.
[0088] The second type of fault refers to a fault where the operation and maintenance intelligent agent has functional support but lacks operation authorization.
[0089] Step S520: Output the fault repair plan and repair request; Step S530: Receive confirmation information for the repair request and execute the fault repair plan.
[0090] The message is output to the terminal device through the communication gateway. After the operator confirms it, the operation and maintenance intelligent agent receives the confirmation information and executes the rerouting instruction.
[0091] The second type of fault typically involves physical environmental intervention or high-risk operations, requiring manual confirmation. For example, when troubleshooting results indicate that the robot has deviated from its designated path, the operations and maintenance agent classifies this as a second-type fault. The agent first checks whether the robot model has an autonomous return-to-position function.
[0092] If the robot has an autonomous return function, the maintenance agent generates a repair plan (sends an autonomous return request to the robot) and a repair request, and outputs them to the terminal device. After receiving confirmation from the operator, it executes the return command and continuously checks the return status until successful; if the robot fails to return, it is classified as a third type of fault.
[0093] If the robot lacks autonomous repositioning capabilities, the maintenance agent generates a repair request, suggesting that on-site personnel straighten the stray robot. Simultaneously, the maintenance agent outputs auxiliary troubleshooting suggestions based on the navigation type: for QR code navigation, it suggests checking if the QR code's position or direction is off; for SLAM (Simultaneous Localization and Mapping) navigation, it suggests checking for environmental changes and determining if a rescan is necessary.
[0094] The re-scanning process specifically includes environmental data acquisition, map building, and map updating. Environmental data acquisition refers to controlling the robot to move along a preset or temporary path in the changed area and re-acquiring environmental data (such as point clouds and image frames) using LiDAR or vision sensors. Map building refers to constructing a local or global map (raster map or point cloud map) based on the acquired new environmental data using SLAM algorithms. Map updating refers to comparing the newly built map with the original map, replacing the map data in the changed areas, and updating the map library in the robot scheduling system.
[0095] The operations and maintenance intelligent agent outputs the above repair plan and request, and executes the corresponding action (if any) after receiving confirmation information or waits for human feedback. In this way, a balance can be achieved between automation and security control, solving the technical problem that medium-risk faults require manual review before execution, ensuring security while avoiding the inefficiency of full-process manual intervention.
[0096] In some embodiments, such as Figure 5 As shown, after outputting the troubleshooting results, steps S610 to S620 are also included.
[0097] In step S610, if the fault type represented by the fault investigation result is a third type of fault, a repair request is generated and output.
[0098] The third type of fault refers to faults where the operation and maintenance intelligent agent has no functional support and no operation authorization.
[0099] Step S620: Receive execution information and, in response to the execution information, perform repair actions. The execution information includes repair actions generated based on the fault diagnosis results.
[0100] The third type of fault requires on-site intervention by personnel. For example, when a robot using QR code navigation reports a lost QR code, the maintenance agent classifies it as a third-type fault. Based on the robot's current location and the reported error code, the maintenance agent generates a precise repair request, sends the specific QR code value to the on-site personnel for printing and pasting, and guides them to accurately apply the replacement.
[0101] For example, if the operations and maintenance intelligence detects that the scheduling system is unreachable, it will prioritize checking the network status. If the network is normal, it will wait 1-3 minutes and retry. If it is still unreachable after three attempts, it will generate a repair request and suggest that on-site personnel check whether the scheduling system is abnormal.
[0102] For example, if the operation and maintenance intelligent agent detects and confirms that the robot is disconnected from the network rather than that of the network infrastructure, and still cannot connect after waiting for a certain period of time, it generates a repair request and suggests that on-site personnel check for blind spots in the network signal.
[0103] The operation and maintenance intelligent agent generates and outputs the above repair request. After receiving the execution information sent by the terminal device (such as "QR code has been subsidized" or "scheduling system has been restarted"), the operation and maintenance intelligent agent responds to the execution information and performs verification actions, such as querying the robot status again to confirm that the fault has been resolved.
[0104] In this way, when the operation and maintenance intelligent agent is unable to complete the repair independently in special scenarios, it can send a repair request to the terminal device to ensure the integrity and reliability of the operation and maintenance process.
[0105] This application also provides an operation and maintenance management system for robot clusters, such as... Figure 7 As shown, the robot cluster's operation and maintenance management system includes terminal devices, a robot scheduling system, robots, and servers. An operation and maintenance intelligent agent runs within the server or terminal device. Terminal devices can be mobile phones, tablets, personal digital assistants (PDAs), desktop computers, etc., used to send query requests and receive alarms. The robot scheduling system is responsible for global task scheduling and provides API interfaces. The robot itself carries an embedded system that provides a log download interface.
[0106] When the operations and maintenance agent runs on a server, it operates as a Docker container and communicates with the terminal devices through an Nginx gateway. When a terminal device sends a command to trigger an event, the operations and maintenance agent calls the robot scheduling system's API to obtain the robot's status and processes it according to the aforementioned logic.
[0107] When the operations and maintenance (O&M) agent runs on a terminal device, it is installed directly on the terminal device held by the field engineer as a lightweight application or applet. The O&M agent on the terminal device side has the same logical functions as the O&M agent deployed on the server side, but its data interaction is more streamlined.
[0108] For example, the operation and maintenance intelligent agent can communicate directly with the robot scheduling system and the embedded system of the robot body through a local area network (LAN) or Wi-Fi, without having to go through the cloud Nginx gateway for forwarding.
[0109] For example, when a field engineer discovers a robot malfunction on-site, they can directly open the maintenance app on the terminal device. The maintenance agent can then directly pull the log data of the malfunctioning robot and perform rule matching or lightweight model inference on the local processor of the terminal device to generate troubleshooting results in real time.
[0110] This allows the computing power of the operation and maintenance intelligence agent to be pushed down to the edge, solving the problem of slow response caused by network latency in cloud deployments. It is suitable for single-point fault repair and industrial field environments with limited network coverage.
[0111] Specifically, the operations and maintenance intelligent agent is configured as follows: In response to a triggering event, a status query is initiated to the robot scheduling system. Triggering events include receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check.
[0112] Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information.
[0113] If the status information of the faulty robot indicates a network fault, a network fault notification message is sent to the terminal device.
[0114] If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot.
[0115] Based on the status information and log data of the faulty robot, generate and output the troubleshooting results.
[0116] In this way, by combining operation and maintenance management methods with hardware systems and proactively intervening through an intelligent operation and maintenance agent, the traditional manual polling and log acquisition process can be replaced. This solves the technical problems of delayed fault detection and cumbersome manual log acquisition in existing technologies, enabling rapid fault detection and preliminary location, and shortening the lead time for fault handling.
[0117] This application also provides an electronic device, which may be a physical server or a cloud server instance, including: a memory and a processor, wherein the memory is used to store instructions. The processor is used to invoke the instructions stored in the memory to execute the following methods: In response to a triggering event, a status query is initiated to the robot scheduling system. Triggering events include receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check.
[0118] Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information.
[0119] If the status information of the faulty robot indicates a network fault, a network fault notification message is sent to the terminal device.
[0120] If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot.
[0121] Based on the status information and log data of the faulty robot, generate and output the troubleshooting results.
[0122] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, the above-mentioned operation and maintenance management method for the robot cluster is executed.
[0123] This application also provides a computer program product, which, when read and executed by a computer, causes the above-mentioned operation and maintenance management method for the robot cluster to be executed.
[0124] The computer program product can be a software package or SDK (Software Development Kit), or it can be a computer program stored on a computer-readable storage medium (such as a USB flash drive, solid-state drive, optical disc, or cloud storage space).
Claims
1. A method for operation and maintenance management of a robot swarm, characterized in that, Applied to operational intelligence agents, including: In response to a triggering event, a status query is initiated to the robot scheduling system. The triggering event includes receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check. Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information; If the status information of the faulty robot indicates a network fault, then a network fault notification message is sent to the terminal device. If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot. Based on the status information and log data of the faulty robot, the fault investigation results are generated and output.
2. The method according to claim 1, characterized in that, The step of generating and outputting troubleshooting results based on the status information and log data of the faulty robot includes: The status information and log data of the faulty robot are matched with pre-stored fault investigation rules, or a large model is called to analyze the status information and log data of the faulty robot, and the fault investigation results are generated and output.
3. The method according to claim 2, characterized in that, The step of generating and outputting troubleshooting results based on the status information and log data of the faulty robot includes: Based on the status information of the faulty robot, the log data is filtered by keywords to generate valid fault logs related to the fault; The valid fault logs are matched with pre-stored fault investigation rules, or a large model is called to analyze the valid fault logs to generate and output the fault investigation results.
4. The method according to claim 3, characterized in that, The steps for analyzing the state information and valid log data of the faulty robot using a large model to generate fault diagnosis results include: Based on the valid fault logs, the pre-stored fault troubleshooting table, and the status information of the faulty robot, input prompt words are generated; Based on the input prompts, the large model is invoked to obtain the fault causes and troubleshooting suggestions returned by the large model, and the fault troubleshooting results are generated.
5. The method according to any one of claims 1-4, characterized in that, The operation and maintenance intelligent agent runs in a virtualized container and interacts with robots, robot scheduling systems and external communication platforms through preset interfaces.
6. The method according to any one of claims 1-4, characterized in that, The operation and maintenance intelligent agent is configured with long-term memory rules, which solidify read-only observation constraints, allowing status query and log retrieval operations, but prohibiting change operations on the robot or robot scheduling system.
7. The method according to any one of claims 1-4, characterized in that, After outputting the troubleshooting results, the following is also included: If the fault type represented by the fault investigation result is a first type of fault, a fault repair plan is generated based on the fault investigation result. The first type of fault is a fault in which the operation and maintenance intelligent agent has operation authorization and functional support. Execute the fault repair plan and output information indicating that the fault repair plan has been executed.
8. The method according to any one of claims 1-4, characterized in that, After outputting the troubleshooting results, the following is also included: If the fault type represented by the fault investigation result is a second type of fault, a fault repair plan and a repair request are generated based on the fault investigation result. The second type of fault is a fault in which the operation and maintenance intelligent agent has functional support but no operation authorization. Output the fault repair plan and the repair request; Upon receiving confirmation of the repair request, the fault repair plan is executed.
9. The method according to any one of claims 1-4, characterized in that, After outputting the troubleshooting results, the following is also included: If the fault type represented by the fault investigation result is a third type of fault, then a repair request is generated and output. The third type of fault is a fault in which the operation and maintenance intelligent agent has no operation authorization and no functional support. Receive execution information and, in response to the execution information, perform repair actions, wherein the execution information includes repair actions generated based on the fault diagnosis results.
10. A robot swarm operation and maintenance management system, characterized in that, It includes terminal equipment, robot scheduling system, robot and server, wherein the server or terminal equipment runs an operation and maintenance intelligent agent; The operation and maintenance intelligent agent is configured as follows: In response to a triggering event, a status query is initiated to the robot scheduling system. The triggering event includes receiving a query request from a terminal device or reaching a preset timed inspection cycle. The status query includes a global robot status check or a target robot status check. Obtain robot status information returned by the robot scheduling system, and identify faulty robots based on the status information; If the status information of the faulty robot indicates a network fault, then a network fault notification message is sent to the terminal device. If the status information of the faulty robot indicates a non-network fault, then an interaction is initiated with the body of the faulty robot to obtain the log data of the faulty robot. Based on the status information and log data of the faulty robot, the fault investigation results are generated and output.
11. An electronic device, characterized in that, include: Memory, used to store instructions; as well as A processor for executing the method of any one of claims 1 to 9 by invoking instructions stored in the memory.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-9.
13. A computer program product, characterized in that, When the computer reads and executes the computer program product, the method as described in any one of claims 1 to 9 is performed.