An operation and maintenance system, method, apparatus and electronic equipment
By using the monitoring terminal, network proxy terminal, and cloud platform of the operation and maintenance system, faults in the robot control terminal are handled automatically, solving the problem of low fault handling efficiency and achieving rapid fault handling and continuous monitoring.
Patent Information
- Application Number
- CN202411238731.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The troubleshooting process for robot control terminals is cumbersome, resulting in low efficiency and affecting the stable operation of industrial production.
An operation and maintenance system is provided, including a monitoring terminal, a network agent terminal, and a cloud platform. By acquiring operation and maintenance data from the robot control terminal, the system automatically identifies faults and generates fault handling instructions, enabling rapid fault handling and escalation.
Without affecting industrial production, a closed-loop system for rapid fault handling and detection is achieved, thus improving fault handling efficiency.
Smart Images

Figure CN119057780B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to an operation and maintenance system, method, apparatus and electronic device. Background Technology
[0002] Currently, in industrial production scenarios, robot systems can be used to replace manual labor in factories to perform tasks such as handling. A robot system can include multiple robots and a robot control terminal; the robot control terminal can be deployed in the factory's local computer room for scheduling the robots.
[0003] When a robot control terminal malfunctions, maintenance personnel need to rush to the site to troubleshoot the problem and collect information. This information is then fed back to the R&D personnel so they can determine the appropriate handling method. The R&D personnel then relay this method back to the maintenance personnel for on-site processing. As you can see, the information transmission and processing procedures for handling robot control terminal malfunctions are quite cumbersome. Often, the malfunction is only addressed after it has already caused serious negative impacts. Furthermore, the troubleshooting process can disrupt the continuous and stable operation of industrial production. Therefore, the efficiency of troubleshooting is relatively low. Summary of the Invention
[0004] The purpose of this application is to provide an operation and maintenance system, method, apparatus, and electronic device to solve the problem of low efficiency in fault handling of robot control terminals. The specific technical solution is as follows:
[0005] This application first provides an operation and maintenance system, which includes: a monitoring terminal, a network proxy terminal, and a cloud platform; wherein:
[0006] The monitoring terminal is used to acquire the operation and maintenance data of the robot control terminal to be monitored, and when it is determined that the robot control terminal is abnormal based on the operation and maintenance data, it sends the operation and maintenance data to the cloud platform through the network proxy terminal; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0007] The cloud platform is used to generate a fault handling instruction for the robot control terminal when a fault is determined based on the received operation and maintenance data, and to send the fault handling instruction to the robot control terminal through the network proxy terminal so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0008] Optionally, the cloud platform is specifically used to generate a fault handling instruction for handling the fault of the first specified type when it is determined that the fault of the robot control terminal belongs to a first specified type.
[0009] The cloud platform is also used to, when determining that the fault of the robot control terminal does not belong to the first specified type, obtain reference data associated with the operation and maintenance data from the log data and / or dump data of the robot control terminal, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data; and after receiving the upgrade file, send the upgrade file to the monitoring terminal through the network proxy terminal;
[0010] The monitoring terminal is also used to receive the upgrade file and send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file.
[0011] Optionally, the robot control terminal is upgraded in a rolling upgrade manner.
[0012] Optionally, the cloud platform is also used to receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal;
[0013] If a robot malfunction is determined based on the operational status data, a fault handling instruction is generated to handle the malfunction of the robot.
[0014] Optionally, the monitoring terminal is deployed on the server;
[0015] The monitoring terminal is also used to automatically shut down when it detects that the computing resources it occupies have reached a preset threshold.
[0016] Optionally, the system further includes: a monitoring and management terminal, used to perform at least one of the following steps:
[0017] When the monitoring terminal is detected to be closed, the monitoring terminal will be restarted.
[0018] The network proxy receives the monitoring terminal program package sent by the cloud platform to update the monitoring terminal deployed in the server.
[0019] The network proxy receives the upgrade file for the robot control terminal sent by the cloud platform and sends it to the monitoring terminal, so that the monitoring terminal upgrades the robot control terminal based on the upgrade file.
[0020] Optionally, the monitoring terminal is further configured to, before sending the operation and maintenance data to the cloud platform through the network proxy terminal when the robot control terminal is determined to be abnormal based on the operation and maintenance data, determine that the robot control terminal is abnormal for multiple different time periods corresponding to the same operation state in the operation and maintenance data if the multiple operation state data meet preset conditions.
[0021] The preset conditions include at least one of the following:
[0022] The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the multiple operational status data indicate that the number of consecutive anomalies is greater than a second threshold, and the multiple operational status data indicate that the duration of consecutive anomalies is greater than a first duration.
[0023] Optionally, the network agent is deployed in the robot's charging station.
[0024] This application embodiment also provides an operation and maintenance method applied to the monitoring terminal of an operation and maintenance system, wherein the operation and maintenance system further includes: a network proxy terminal and a cloud platform; the method includes:
[0025] Obtain the operation and maintenance data of the robot control terminal to be monitored; wherein the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0026] When the operation and maintenance data determines that the robot control terminal is abnormal, the operation and maintenance data is sent to the cloud platform through the network proxy terminal. This enables the cloud platform to generate a fault handling instruction for the robot control terminal when it determines that the robot control terminal is faulty based on the received operation and maintenance data. The cloud platform then sends the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0027] Optionally, the method further includes:
[0028] The system receives an upgrade file sent by the cloud platform for handling faults; wherein the upgrade file is generated by the maintenance personnel based on reference data; the reference data is data related to the maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type;
[0029] The upgrade file is sent to the robot control terminal so that the robot control terminal can perform an upgrade based on the upgrade file.
[0030] Optionally, the monitoring terminal is deployed on the server; the method further includes:
[0031] It automatically shuts down when it detects that the computing resources it is using have reached a preset threshold.
[0032] Optionally, before sending the maintenance data to the cloud platform through the network proxy when the maintenance data determines that the robot control terminal is abnormal, the method further includes:
[0033] For the same operating state in the operation and maintenance data, if the multiple operating state data meet the preset conditions, it is determined that the robot control terminal is abnormal.
[0034] The preset conditions include at least one of the following:
[0035] The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
[0036] This application also provides an operation and maintenance method applied to a cloud platform of an operation and maintenance system, wherein the operation and maintenance system further includes: a network proxy terminal and a monitoring terminal; the method includes:
[0037] The monitoring terminal receives maintenance data sent when it determines that the robot control terminal is abnormal; wherein the maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal.
[0038] If a fault is determined in the robot control terminal based on the maintenance data, a fault handling instruction is generated to handle the fault in the robot control terminal.
[0039] The fault handling instruction is sent from the network proxy to the robot control terminal, so that the robot control terminal can handle the fault according to the fault handling instruction.
[0040] Optionally, generating fault handling instructions for the robot control terminal includes:
[0041] Generate a fault handling instruction for handling the fault of the first specified type.
[0042] The method further includes:
[0043] If it is determined that the fault of the robot control terminal does not belong to the first specified type, obtain reference data related to the operation and maintenance data from the log data and / or dump data of the robot control terminal;
[0044] Send to the operations and maintenance client so that operations and maintenance personnel can generate an upgrade file for handling the fault based on the reference data;
[0045] Upon receiving the upgrade file, the upgrade file is sent to the monitoring terminal via the network proxy terminal; so that the monitoring terminal receives the upgrade file and sends the upgrade file to the robot control terminal, so that the robot control terminal can perform an upgrade based on the upgrade file.
[0046] Optionally, the method further includes:
[0047] Receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal;
[0048] If a robot malfunction is determined based on the operational status data, a fault handling instruction is generated to handle the malfunction of the robot.
[0049] This application embodiment also provides an operation and maintenance device, applied to the monitoring terminal of an operation and maintenance system, wherein the operation and maintenance system further includes: a network proxy terminal and a cloud platform; the device includes:
[0050] The operation and maintenance data acquisition module is used to acquire operation and maintenance data of the robot control terminal to be monitored; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0051] The operation and maintenance data sending module is used to send the operation and maintenance data to the cloud platform through the network proxy when it is determined that the robot control terminal is abnormal based on the operation and maintenance data. This enables the cloud platform to generate a fault handling instruction for the robot control terminal when it is determined that the robot control terminal is faulty based on the received operation and maintenance data. The cloud platform then sends the fault handling instruction to the robot control terminal through the network proxy, so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0052] Optionally, the device further includes:
[0053] The upgrade file receiving module is used to receive an upgrade file sent by the cloud platform for handling faults; wherein the upgrade file is generated by the operation and maintenance personnel based on reference data; the reference data is data related to the operation and maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type;
[0054] The upgrade file upgrade module is used to send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file.
[0055] Optionally, the monitoring terminal is deployed in the server; the device further includes:
[0056] The automatic shutdown module is used to automatically shut down when it detects that the computing resources it occupies have reached a preset threshold.
[0057] Optionally, the device further includes:
[0058] The anomaly determination module, configured to, when the operation and maintenance data determines that the robot control terminal is abnormal, before sending the operation and maintenance data to the cloud platform through the network proxy terminal, further includes:
[0059] For the same operating state in the operation and maintenance data, if the multiple operating state data meet the preset conditions, it is determined that the robot control terminal is abnormal.
[0060] The preset conditions include at least one of the following:
[0061] The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
[0062] This application embodiment also provides an operation and maintenance device applied to a cloud platform of an operation and maintenance system, wherein the operation and maintenance system further includes: a network proxy terminal and a monitoring terminal; the device includes:
[0063] The operation and maintenance data receiving module is used to receive operation and maintenance data sent by the monitoring terminal when it determines that the robot control terminal is abnormal; wherein, the operation and maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the operation and maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal.
[0064] The first instruction generation module is used to generate a fault handling instruction for handling the fault of the robot control terminal when the robot control terminal is determined to be faulty based on the operation and maintenance data.
[0065] The instruction sending module is used to send the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0066] Optionally, the first instruction generation module is specifically used to generate a fault handling instruction for handling the fault of the first specified type.
[0067] The device further includes:
[0068] The reference data acquisition module is used to acquire reference data associated with the operation and maintenance data from the log data and / or dump data of the robot control terminal when it is determined that the fault of the robot control terminal does not belong to the first specified type, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data.
[0069] The upgrade file sending module is used to send the upgrade file to the monitoring terminal through the network proxy terminal after receiving the upgrade file; so that the monitoring terminal receives the upgrade file and sends the upgrade file to the robot control terminal, so that the robot control terminal can perform an upgrade based on the upgrade file.
[0070] Optionally, the device further includes:
[0071] The operation status data receiving module is used to receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal;
[0072] The second instruction generation module is used to generate a fault handling instruction for the robot to handle the fault when the robot fault is determined based on the operating status data.
[0073] This application also provides an electronic device, including:
[0074] Memory, used to store computer programs;
[0075] When a processor executes a program stored in memory, it implements any of the above-described operation and maintenance methods.
[0076] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described operation and maintenance methods.
[0077] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the above-described operation and maintenance methods.
[0078] Beneficial effects of the embodiments in this application:
[0079] The robot operation and maintenance system provided in this application includes: a monitoring terminal, a network proxy terminal, and a cloud platform. The monitoring terminal can acquire operation and maintenance data of the robot control terminal to be monitored, and when it is determined that the robot control terminal is abnormal based on the operation and maintenance data, it sends the operation and maintenance data to the cloud platform through the network proxy terminal. The operation and maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server. The cloud platform can generate a fault handling instruction for the robot control terminal to handle the fault when it is determined that the robot control terminal is faulty based on the received operation and maintenance data, and send the fault handling instruction to the robot control terminal through the network proxy terminal so that the robot control terminal can handle the fault according to the fault handling instruction.
[0080] As can be seen, this solution enables automated and rapid information collection and transmission when a fault occurs in the robot control terminal, without affecting the operation of industrial production in the robot's environment. This allows for rapid fault handling, thus improving the efficiency of fault handling in the robot control terminal. Furthermore, the operation and maintenance system provided in this embodiment continuously monitors the robot control terminal, allowing for continuous detection of whether the resolved fault recurs after fault handling, achieving a closed-loop fault detection system.
[0081] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0082] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0083] Figure 1 This is a schematic diagram of the operation and maintenance system provided in the embodiments of this application;
[0084] Figure 2 This is a schematic diagram of the business architecture of the monitoring terminal in the operation and maintenance system provided in the embodiments of this application;
[0085] Figure 3 This is a schematic diagram of a process involved in the operation and maintenance system provided in the embodiments of this application;
[0086] Figure 4 This is a schematic diagram of rolling upgrades in an embodiment of this application;
[0087] Figure 5 A schematic diagram of the downlink network involved in the operation and maintenance system provided in the embodiments of this application;
[0088] Figure 6 This is a schematic diagram of the uplink network involved in the operation and maintenance system provided in the embodiments of this application;
[0089] Figure 7 This is a schematic diagram of the deployment architecture of the operation and maintenance system provided in the embodiments of this application;
[0090] Figure 8 This is a flowchart illustrating the operation and maintenance method provided in the embodiments of this application;
[0091] Figure 9 This is a flowchart illustrating another operation and maintenance method provided in an embodiment of this application;
[0092] Figure 10 This is a schematic diagram of the operation and maintenance device provided in the embodiments of this application;
[0093] Figure 11 This is a schematic diagram of another maintenance device provided in an embodiment of this application;
[0094] Figure 12 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0095] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0096] To improve the efficiency of fault handling in robot control systems, this application provides an operation and maintenance system, which includes: a monitoring terminal, a network proxy terminal, and a cloud platform; wherein:
[0097] The monitoring terminal is used to acquire the operation and maintenance data of the robot control terminal to be monitored, and when it is determined that the robot control terminal is abnormal based on the operation and maintenance data, it sends the operation and maintenance data to the cloud platform through the network proxy terminal; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0098] The cloud platform is used to generate fault handling instructions for the robot control terminal when a fault is determined based on the received operation and maintenance data. The fault handling instructions are then sent to the robot control terminal through a network proxy so that the robot control terminal can handle the fault according to the fault handling instructions.
[0099] This solution enables automated and rapid information collection and transmission when a robot control terminal malfunctions, without disrupting the industrial production process in the robot's environment. This facilitates quick fault handling and improves the efficiency of robot control terminal fault resolution. Furthermore, the maintenance system provided in this application continuously monitors the robot control terminal, allowing for ongoing detection of recurrence of resolved faults after they have been addressed, thus achieving a closed-loop fault detection system.
[0100] The operation and maintenance system provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0101] like Figure 1 As shown, the operation and maintenance system provided in this application embodiment may include:
[0102] Monitoring terminal 101, network proxy terminal 102, and cloud platform 103;
[0103] The operation and maintenance system proposed in this application embodiment can be used for the operation and maintenance of robot systems in scenarios such as industrial production.
[0104] The network proxy 102 can be a wired network card or a wireless network card. In one implementation, the network proxy 102 can be deployed in the robot's charging station to improve the level of integration. Multiple charging stations can exist in a robot system.
[0105] The monitoring terminal 101 is used to acquire the operation and maintenance data of the robot control terminal to be monitored, and when it is determined that the robot control terminal is abnormal based on the operation and maintenance data, it sends the operation and maintenance data to the cloud platform 103 through the network proxy terminal 102; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0106] The robot control terminal can be a server used to control the robot, or a control program running on that server. This server can install and execute various robot tasks and other programs. It can be deployed in a local data center. A robot system can contain multiple such servers, thereby achieving load balancing when receiving and processing robot control tasks.
[0107] The monitoring terminal 101 can be a program or device used to monitor the robot control terminal. In one implementation, the monitoring terminal 101 and the robot control terminal can be deployed on the same server. In the case of a robot system containing multiple servers, the monitoring terminal 101 can be deployed on each server.
[0108] like Figure 2 As shown, the monitoring terminal 101 can detect the system-level running status data of the server, which are parameters representing the overall running status of the server, such as CPU utilization, memory utilization, disk utilization, as well as thread information (e.g., the number of threads used, the running status of each process, etc.) and handle information (e.g., the number of handles used). Thread information can be obtained by detecting the PID (process ID).
[0109] For each robot business component running on the server, the monitoring terminal 101 can also detect the operating status data (i.e., component operating status data) of that robot business component. Each robot business component can be a component in the robot control terminal used to implement a business function. Robot business components in the robot control terminal can include: iWMS (Intelligence Warehouse Management System), RCMS (Robot Condition Monitoring System), RCS (Robot Control System), WCS (Warehouse Control System), AMS (Automated Manufacturing Systems), etc.
[0110] For each business component, component runtime status data for common basic detection items can be obtained. Common basic detection items include: memory usage, CPU usage, handle information (e.g., the number of handles used), database connection count, log backup information (e.g., whether log backup was successful), abnormal crash information (e.g., the number of crashes), and data migration task information (e.g., whether data migration was successful).
[0111] In addition, depending on the different business characteristics, it is also possible to detect the number of incomplete tasks of a specific business component, whether the message sending queue is backlogged, whether the message receiving queue is backlogged, interface exception information (e.g., the number of times the interface is exception), internal processing time, and other information for that specific business component.
[0112] The monitoring terminal 101 can also detect the middleware runtime status data running on the server. This middleware can be a general middleware that supports the operation of various robot business components on the server, such as databases and message middleware. For databases, it can detect configuration parameters, computing resources used by the database, database data volume, SQL (Structured Query Language), and the execution time of TOP SQL (the most time-consuming database query statement). For middleware, it can also detect information such as memory usage and CPU usage.
[0113] In addition, users can also manage the detection items to be detected by the monitoring terminal 101 according to their needs, that is, independently set the detection items to be detected, and perform routine configurations for the monitoring terminal 101, including configuring detection strategies, inference strategies, etc.
[0114] When the monitoring terminal 101 is deployed on a local server, it can use cgroups (control groups) to manage its own computing resources, such as CPU resources and memory. It can also automatically shut down when it detects that the computing resources it occupies have reached a preset threshold, so as to minimize the impact of the monitoring terminal 101 on the server performance and thus reduce the impact of the monitoring terminal 101 on the normal operation of the robot control terminal.
[0115] In one implementation, the monitoring terminal 101 can pre-set abnormal conditions for each parameter in the operation and maintenance data detected by the monitoring terminal 101. For example, the abnormal condition could be that the CPU utilization rate in the system-level operating status data exceeds a preset threshold, or that the memory occupied by a business component exceeds a preset threshold. Thus, the monitoring terminal 101 can confirm that the robot control terminal is abnormal when it detects that any parameter in the operation and maintenance data meets the abnormal condition.
[0116] After monitoring terminal 101 determines that the robot control terminal is abnormal, it can send operation and maintenance data to cloud platform 103 through network proxy terminal 102. The operation and maintenance data sent can be all the operation and maintenance data currently acquired by monitoring terminal 101, or operation and maintenance data related to the abnormality. For example, if a parameter in the component operation status data of a business component meets a preset condition, indicating that the business component is abnormal, then all the operation and maintenance data currently acquired for that business component can be sent to the cloud platform.
[0117] The operation and maintenance data acquired by the monitoring terminal 101 can be data generated within a certain time period. The monitoring terminal 101 can also periodically acquire the generated operation and maintenance data, and after each acquisition of operation and maintenance data, it determines whether the robot control terminal is abnormal based on the operation and maintenance data. When it is determined that the robot control terminal is abnormal, it can send the acquired operation and maintenance data, or the data in the acquired operation and maintenance data related to the abnormality, to the cloud platform 103.
[0118] In one implementation, the monitoring terminal 101 can first generate an event containing the operation and maintenance data to be sent, and then send the event to the cloud platform 103.
[0119] The cloud platform 103 is used to generate a fault handling instruction for the robot control terminal when a fault is determined based on the received operation and maintenance data, and to send the fault handling instruction to the robot control terminal through the network agent terminal 102 so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0120] In one implementation, the cloud platform 103 can determine the faults corresponding to the acquired maintenance data based on the pre-set correspondence between the maintenance dataset and the fault set. Furthermore, the fault set can include a "no fault" item, indicating that no fault has occurred at the robot control end. That is, when the "no fault" item is determined to correspond to the acquired maintenance data, it is determined that no fault has occurred at the robot control end.
[0121] In another implementation, maintenance data can be input into a pre-trained fault detection model to obtain the result representing the fault type output by the fault detection model. The fault detection model can be a neural network model trained using sample maintenance data and the corresponding labels representing the fault type of the sample maintenance data.
[0122] The cloud platform 103 can be pre-set with processing strategies for different fault types. After determining the fault type of the robot control terminal, the cloud platform 103 can generate fault processing instructions according to the corresponding processing strategy.
[0123] For example, when the fault type is determined to be: excessive memory leakage of a certain business component, excessive CPU utilization, or other resource out-of-bounds faults that affect the overall stability of the robot control terminal, the generated fault handling instruction can be used to instruct the shutdown or restart of the business component.
[0124] When the fault type is determined to be, such as excessive database data volume or excessive SQL latency, the generated fault handling instructions can be used to instruct the data to be migrated from the database. When the fault type is determined to be, such as data migration failure, the generated fault handling instructions can be used to instruct the data to be migrated again.
[0125] When the fault type is determined to be such as the system process of the robot control terminal being frozen or the system process of the robot control terminal being deadlocked, the generated fault handling instructions can be used to instruct the robot control terminal to retain a data snapshot and then restart, and to restore the data based on the data snapshot after restarting.
[0126] If the robot control terminal is in a virtual machine environment such as virtualization or containerization, the fault type may include insufficient system resources, that is, insufficient computing resources of the virtual machine used by the robot control terminal. In this case, the generated fault handling instructions can be used to instruct the robot control terminal to expand the computing resources it occupies.
[0127] If the fault type is determined to be: the robot control terminal takes too long to process and has insufficient processing capacity, the generated fault handling instruction can be used to instruct the robot control terminal to expand service nodes, that is, to enable idle servers for robot control.
[0128] After generating fault handling instructions, the cloud platform 103 can send the fault handling instructions to the robot control terminal through the network proxy terminal 102, so that the robot control terminal can handle the fault according to the fault handling instructions, thereby realizing the ability to automatically handle faults in the operation and maintenance system. The fault handling instructions can be sent by the cloud platform 103 to the monitoring terminal 101 through the network proxy terminal 102, and then the monitoring terminal 101 sends the fault handling instructions to the robot control terminal.
[0129] In one implementation, the cloud platform 103 can also be used to receive robot operating status data uploaded by the robot to be monitored through the network agent 102; and, if a robot fault is determined based on the operating status data, generate a fault handling instruction to handle the fault in the robot. The robot operating status data may include robot device-level information such as battery level, temperature, and pose, and may also include the robot's CPU utilization, memory utilization, and program errors.
[0130] The robot can monitor its operational status and upload operational status data based on its own alarm module. The robot can also generate events containing robot operational status data and then send the generated events to the cloud platform 103.
[0131] The robot can interact directly with the cloud platform 103 through the network proxy terminal 102. Alternatively, the monitoring terminal 101 can first obtain the robot's operating status data, and then send the operating status data to the cloud platform 103 through the network proxy terminal 102. The monitoring terminal 101 can then receive the fault handling instructions generated by the cloud platform 103 to handle the robot's faults and forward them to the robot.
[0132] The way cloud platform 103 generates fault handling instructions for robot fault handling is similar to the way fault handling instructions for robot control terminal are generated as described above, and will not be repeated here.
[0133] In this embodiment, when a fault occurs in the robot control terminal, information can be automatically and rapidly collected and transmitted without affecting the operation of industrial production in the robot's environment, thereby achieving rapid fault handling. Therefore, this solution improves the efficiency of fault handling in the robot control terminal. Furthermore, the operation and maintenance system provided in this embodiment can continuously monitor the robot control terminal, thus continuously detecting whether the handled fault recurs after fault handling, achieving a closed-loop fault detection system.
[0134] In another embodiment of this application, the cloud platform 103 is specifically used to generate a fault handling instruction for handling the fault of the first specified type when it is determined that the fault of the robot control terminal belongs to the first specified type.
[0135] The first specified type can be a fault type for which a corresponding fault handling strategy is pre-recorded in the cloud platform 103. In this case, the cloud platform 103 can directly generate a fault handling instruction for handling faults of the first specified type.
[0136] The cloud platform 103 is also used to obtain reference data related to operation and maintenance data from the log data and / or dump data of the robot control terminal when it is determined that the fault of the robot control terminal does not belong to the first specified type, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data; and after receiving the upgrade file, send the upgrade file to the monitoring terminal 101 through the network agent terminal.
[0137] If the fault in the robot control terminal is determined not to belong to the first specified type, it indicates that the preset processing strategy in cloud platform 103 cannot handle the fault. In this case, reference data related to operation and maintenance data can be obtained from the log data and / or dump data of the robot control terminal. Dump data refers to the data generated in memory when a program runs on a computer and then transferred to other storage media. For example, when a program encounters an error or crashes, the data in memory during the program's runtime can be dumped to a file for subsequent investigation of the cause of the problem. The data dumped in this process is called dump data.
[0138] Specifically, the reference data can be obtained by the monitoring terminal 101 from the storage device of the robot control terminal and then sent to the cloud platform 103, so that the cloud platform 103 can obtain the reference data. The reference data can be log data and / or dump data generated within a specified time period before the fault occurred.
[0139] After obtaining the reference data, cloud platform 103 can send the reference data to operations and maintenance personnel through the operations and maintenance client. This allows the personnel to perform fault analysis based on the reference data, determine the cause of the fault, and then generate upgrade files to address the fault. The operations and maintenance client can be a communication program, enterprise office software, enterprise office system, or other similar client.
[0140] The cloud platform 103 can handle robot faults in a similar way to the above. For example, if the determined robot fault does not belong to a specified type, it can obtain reference data related to operation and maintenance data from the robot's log data and / or dump data. This allows operation and maintenance personnel to generate robot upgrade files based on the reference data for the robot to handle the faults. The cloud platform 103 then receives the robot upgrade files and sends them to the robot through the network agent 102 to upgrade the robot.
[0141] For example, one processing flow of this application embodiment can be as follows: Figure 3As shown, firstly, the cloud platform 103 collects and aggregates events containing operation and maintenance data from the robot and its control terminal, persistently storing them on the cloud platform 103. Based on this data, the cloud platform 103 performs fault detection and diagnosis, automatically collecting log and dump data from the robot control terminal and generating fault handling instructions to send to the robot control terminal for pre-fault handling. Simultaneously, the cloud platform 103 can also notify operation and maintenance personnel of fault occurrences via telephone or by sending voice or text messages through enterprise office software. It can also push fault occurrence messages to the enterprise office system, enabling the system to notify the customer problem handling system to generate a customer problem ticket awaiting task execution. The cloud platform 103 allows for flexible configuration of different judgment conditions, alarm levels, alarm policies, and alarm push methods for different fault types. The cloud platform 103 can also generate fault diagnosis reports as needed, either periodically or at any time. These reports can include information categorized and statistically analyzed for faults occurring in the robot and its control terminal within a specific time period. Furthermore, the information in the fault diagnosis report can be used to further optimize and train the aforementioned fault detection model to improve the accuracy of the fault detection model.
[0142] In this embodiment, the cloud platform 103 used for the operation and maintenance of the robot and its control terminal can also be called an alarm monitoring center. One cloud platform 103 can be used to operate and maintain multiple robot systems.
[0143] Maintenance personnel develop upgrade files to handle faults, and finally upload the generated upgrade files as deliverables to the enterprise office system. The upgrade file can be a program installation package. The enterprise office system then pushes the upgrade file to the cloud platform 103, which in turn sends it to the monitoring terminal 101. The monitoring terminal 101 can install the upgrade file when the robot or robot control terminal meets the upgrade conditions, such as when the robot or robot control terminal is idle, to upgrade the program of the robot or robot control terminal. Afterwards, the monitoring terminal 101 and the cloud platform 103 can also detect whether the upgrade was successful and continuously monitor whether the same fault recurs.
[0144] The monitoring terminal 101 is also used to receive upgrade files and send upgrade files to the robot control terminal so that the robot control terminal can upgrade based on the upgrade files.
[0145] In one implementation method, the robot control terminal can be upgraded via either a canary release or a rolling upgrade. For a simple example, a rolling upgrade can be as follows: Figure 4As shown: Before the upgrade, 100% of the task traffic to be processed by the robot control terminal was handled by instances 1 and 2 in version V1. After installing the upgrade files to obtain the upgraded version V2, 5% of the task traffic can be initially allocated to instance 1 in version V2. After running for a preset time, the task traffic allocated to instance 1 in version V2 can be increased to 30%. After running for a period of time, if there are no abnormalities, the task traffic allocated to instance 1 in version V2 can be further increased to 50%, and no more task traffic can be allocated to instance 1 in version V1. After running for a period of time, 70% of the task traffic can be allocated to instances 1 and 2 in version V2. If the operation is stable, finally, no more task traffic can be allocated to version V1, and 100% of the task traffic can be allocated to instances 1 and 2 in version V2, completing the upgrade.
[0146] This embodiment improves the efficiency of fault handling in the robot control terminal. Furthermore, the cloud platform obtains reference data related to operation and maintenance data from the robot control terminal's log data and / or dump data, and sends it to the operation and maintenance client. This allows operation and maintenance personnel to generate upgrade files for fault handling based on the reference data. Upon receiving the upgrade files, the cloud platform sends them to the monitoring terminal via a network proxy. The monitoring terminal receives the upgrade files and sends them back to the robot control terminal, enabling the robot control terminal to upgrade based on the upgrade files. This allows operation and maintenance personnel to develop upgrade files for fault handling when automatic fault handling is not possible, and then automatically upgrade the robot control terminal to resolve the fault.
[0147] In another embodiment of this application, the monitoring terminal 101 is also used to determine that the robot control terminal is abnormal before sending the operation and maintenance data to the cloud platform 103 through the network proxy terminal 102 when the robot control terminal is determined to be abnormal based on the operation and maintenance data. If the multiple operation and maintenance data corresponding to the same operation and maintenance state at multiple different times meet the preset conditions, the robot control terminal is determined to be abnormal.
[0148] One operational status in the operation and maintenance data corresponds to each parameter that the monitoring terminal 101 needs to monitor.
[0149] The preset conditions include at least one of the following:
[0150] The first condition is that the frequency of anomalies in the multiple operational status data exceeds a first threshold; that is, within a specified time period, the number of anomalies in multiple operational status data for the same operational status exceeds the first threshold. The monitoring terminal 101 can periodically acquire operational status data corresponding to the same operational status in the operational data, and determine whether the operational status data indicates anomaly after each acquisition. For example, it can periodically acquire system-level memory usage; after each acquisition, if the memory usage rate is greater than a preset threshold, then the memory usage rate is determined to be anomaly.
[0151] The second condition is that the number of consecutive abnormalities among the multiple operational status data exceeds the second threshold; that is, for the same operational status, multiple consecutive operational status data all indicate anomalies.
[0152] Thirdly, among the multiple operational status data, the duration of consecutive abnormalities exceeds the first duration. When the operational status data acquired by the monitoring terminal 101 for a certain operational status first indicates an abnormality, the time of acquisition of the operational status data can be recorded as the time of abnormality occurrence. If every subsequent acquisition of operational status data corresponding to that operational status is abnormal, and the time interval between the most recent acquisition of the operational status and the time of abnormality occurrence reaches the first duration, then it can be determined that the robot control terminal is abnormal.
[0153] In one implementation, the operation and maintenance system may further include: a monitoring and management terminal, used to perform at least one of the following steps:
[0154] Restart monitoring terminal 101 when it is detected that monitoring terminal 101 is closed;
[0155] The network proxy 102 receives the program package sent by the cloud platform 103 to the monitoring terminal 101 to update the monitoring terminal 101 deployed on the server.
[0156] The network proxy terminal 102 receives the upgrade file for the robot control terminal sent by the cloud platform 103 and sends it to the monitoring terminal 101 so that the monitoring terminal 101 can upgrade the robot control terminal based on the upgrade file.
[0157] The monitoring and management terminal can be deployed on the server used to control the robot, or it can be a standalone device. The monitoring and management terminal can receive large files of data sent by the cloud platform 103, including the program package of the monitoring terminal 101 and the upgrade files of the robot control terminal. This allows for the initial installation and subsequent upgrades of the monitoring terminal 101, as well as the sending of the received upgrade files to the monitoring terminal 101 so that the monitoring terminal 101 can upgrade the robot control terminal based on the upgrade files.
[0158] like Figure 5As shown, the downlink data network involved in the operation and maintenance system provided in this embodiment may include: cloud platform 103, network agent terminal 102, robot, monitoring and management terminal and operation and maintenance application; the operation and maintenance application is an application used to perform operation and maintenance on the robot and robot control system, and may include the monitoring terminal 101 in this embodiment; the operation and maintenance application and the monitoring and management terminal can communicate through the local loopback address 127.0.0.1.
[0159] In the uplink data network involved in the operation and maintenance system provided in this embodiment, such as Figure 6 As shown, the robot and operation and maintenance application directly upload the operation status data to the cloud platform 103 through the network proxy. Since the data volume is small, the monitoring and management terminal does not need to be involved. Therefore, the monitoring and management terminal may not be included in the uplink data network.
[0160] The monitoring management terminal and monitoring terminal 101 can transmit different data through different links, which can avoid mutual interference and improve the operating efficiency of the operation and maintenance system.
[0161] In this embodiment, the efficiency of fault handling in the robot control terminal is improved. Furthermore, the monitoring terminal determines that the robot control terminal is abnormal only when multiple operating status data meet preset conditions, rather than considering the robot control terminal abnormal every time the acquired operating status data indicates an anomaly. This avoids frequently sending maintenance data to the cloud platform, and the cloud platform does not need to perform fault diagnosis frequently, thereby reducing the occupation of computing resources.
[0162] In one embodiment of this application, a deployment architecture diagram of the operation and maintenance system can be shown as follows: Figure 7 As shown: The monitoring management terminal can communicate with the monitoring terminal 101 via HTTP (Hypertext Transfer Protocol), and the monitoring terminal 101 communicates with the cloud platform 103 via HTTPS (Hypertext Transfer Protocol Secure) and the network proxy terminal 102. The monitoring terminal 101 can be deployed on the same server as the robot control terminal. The robot can also interact with the cloud platform 103 through the network proxy terminal 102. The cloud platform 103 can also send messages to the communication software to notify maintenance personnel of faults; the cloud platform 103 can also interact with the enterprise office and R&D management system, pushing fault occurrence messages to the enterprise office system so that the enterprise office system can notify the customer problem handling system to generate customer problem tickets pending task orders, and receive upgrade files uploaded by the enterprise office system.
[0163] In this embodiment, when a fault occurs at the robot control end, information can be automatically and rapidly collected and transmitted without affecting the operation of industrial production in the robot's environment, thereby achieving rapid fault handling. Therefore, this solution improves the efficiency of fault handling at the robot control end. Furthermore, the operation and maintenance system provided in this embodiment can continuously monitor the robot control end, thus continuously detecting whether the handled fault recurs after fault handling, achieving a closed-loop fault detection system. The entire process is highly automated and intelligent, greatly improving operation and maintenance, and enabling timely fault detection and handling.
[0164] This application also provides an operation and maintenance method applied to the monitoring end of an operation and maintenance system, which further includes a network proxy and a cloud platform; as shown in the embodiments of this application. Figure 8 As shown, the method includes:
[0165] S801, Obtain the operation and maintenance data of the robot control terminal to be monitored; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0166] S802, when an anomaly is determined to be in the robot control terminal based on the operation and maintenance data, the operation and maintenance data is sent to the cloud platform through the network proxy terminal, so that the cloud platform can generate a fault handling instruction for the robot control terminal when it is determined to be in fault based on the received operation and maintenance data, and send the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0167] Optionally, the method further includes:
[0168] Receive upgrade files sent by the cloud platform for handling faults; wherein the upgrade files are generated by the operation and maintenance personnel based on reference data; the reference data are the data related to the operation and maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type;
[0169] Send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file.
[0170] Optionally, the monitoring terminal is deployed on a server; the method also includes:
[0171] It automatically shuts down when it detects that the computing resources it is using have reached a preset threshold.
[0172] Optionally, before sending the maintenance data to the cloud platform via the network proxy when an anomaly is determined to be in the robot control terminal through maintenance data, the method further includes:
[0173] If the same operating state corresponds to multiple operating state data at different times in the operation and maintenance data, and the multiple operating state data meet the preset conditions, it is determined that the robot control terminal is abnormal.
[0174] The preset conditions include at least one of the following:
[0175] The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
[0176] This application also provides an operation and maintenance method applied to a cloud platform of an operation and maintenance system, which further includes: a network proxy terminal and a monitoring terminal; as shown in the embodiments of this application. Figure 9 As shown, the method includes:
[0177] S901, receive maintenance data sent by the monitoring terminal when it determines that the robot control terminal is abnormal; wherein, the maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal.
[0178] S902, when a fault is determined in the robot control terminal based on operation and maintenance data, a fault handling instruction is generated to handle the fault in the robot control terminal.
[0179] S903 sends fault handling instructions to the robot control terminal through the network agent, so that the robot control terminal can handle the fault according to the fault handling instructions.
[0180] Optionally, the above-mentioned fault handling instructions for generating fault handling at the robot control end include:
[0181] Generate fault handling instructions for handling faults of the first specified type;
[0182] The method also includes:
[0183] If it is determined that the fault in the robot control terminal does not belong to the first specified type, obtain reference data related to the operation and maintenance data from the log data and / or dump data of the robot control terminal;
[0184] Send to the operations and maintenance client so that operations and maintenance personnel can generate upgrade files based on reference data to handle the fault;
[0185] Upon receiving the upgrade file, the upgrade file is sent to the monitoring terminal via the network proxy terminal; this enables the monitoring terminal to receive the upgrade file and send it to the robot control terminal, allowing the robot control terminal to perform the upgrade based on the upgrade file.
[0186] Optionally, the method further includes:
[0187] Receive robot operation status data uploaded by the robot to be monitored through the network agent;
[0188] If a robot malfunction is determined based on operational status data, a fault handling instruction is generated to handle the malfunction.
[0189] This application embodiment also provides an operation and maintenance device applied to the monitoring terminal of an operation and maintenance system, which further includes: a network proxy terminal and a cloud platform; as shown in the example. Figure 10 As shown, the device includes:
[0190] The operation and maintenance data acquisition module 1001 is used to acquire operation and maintenance data of the robot control terminal to be monitored; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server.
[0191] The maintenance data sending module 1002 is used to send the maintenance data to the cloud platform through the network proxy when it is determined that the robot control terminal is abnormal based on the maintenance data, so that the cloud platform can generate a fault handling instruction for the robot control terminal to handle the fault when it is determined that the robot control terminal is faulty based on the received maintenance data, and send the fault handling instruction to the robot control terminal through the network proxy so that the robot control terminal can handle the fault according to the fault handling instruction.
[0192] Optionally, the device further includes:
[0193] The upgrade file receiving module is used to receive an upgrade file sent by the cloud platform for handling faults; wherein the upgrade file is generated by the operation and maintenance personnel based on reference data; the reference data is data related to the operation and maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type;
[0194] The upgrade file upgrade module is used to send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file.
[0195] Optionally, the monitoring terminal is deployed in the server; the device further includes:
[0196] The automatic shutdown module is used to automatically shut down when it detects that the computing resources it occupies have reached a preset threshold.
[0197] Optionally, the device further includes:
[0198] The anomaly determination module, configured to, when the operation and maintenance data determines that the robot control terminal is abnormal, before sending the operation and maintenance data to the cloud platform through the network proxy terminal, further includes:
[0199] For the same operating state in the operation and maintenance data, if the multiple operating state data meet the preset conditions, it is determined that the robot control terminal is abnormal.
[0200] The preset conditions include at least one of the following:
[0201] The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
[0202] This application also provides an operation and maintenance device applied to a cloud platform of an operation and maintenance system. The operation and maintenance system further includes a network proxy terminal and a monitoring terminal. Figure 11 As shown, the device includes:
[0203] The operation and maintenance data receiving module 1101 is used to receive operation and maintenance data sent by the monitoring terminal when it determines that the robot control terminal is abnormal; wherein, the operation and maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the operation and maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal.
[0204] The first instruction generation module 1102 is used to generate a fault handling instruction for the robot control terminal to handle the fault when the robot control terminal is determined to be faulty based on the operation and maintenance data.
[0205] The instruction sending module 1103 is used to send the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
[0206] Optionally, the first instruction generation module is specifically used to generate a fault handling instruction for handling the fault of the first specified type.
[0207] The device further includes:
[0208] The reference data acquisition module is used to acquire reference data associated with the operation and maintenance data from the log data and / or dump data of the robot control terminal when it is determined that the fault of the robot control terminal does not belong to the first specified type, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data.
[0209] The upgrade file sending module is used to send the upgrade file to the monitoring terminal through the network proxy terminal after receiving the upgrade file; so that the monitoring terminal receives the upgrade file and sends the upgrade file to the robot control terminal, so that the robot control terminal can perform an upgrade based on the upgrade file.
[0210] Optionally, the device further includes:
[0211] The operation status data receiving module is used to receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal;
[0212] The second instruction generation module is used to generate a fault handling instruction for the robot to handle the fault when the robot fault is determined based on the operating status data.
[0213] This application also provides an electronic device, such as... Figure 12 As shown, it includes:
[0214] Memory 1201 is used to store computer programs;
[0215] When the processor 1202 executes the program stored in the memory 1201, it implements the steps of any of the above-described operation and maintenance methods.
[0216] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 1202, the communication interface, and the memory 1201 communicating with each other via the communication bus.
[0217] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0218] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0219] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0220] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0221] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described operation and maintenance methods.
[0222] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the operation and maintenance methods described in the above embodiments.
[0223] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.
[0224] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0225] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the method embodiments are basically similar to the system embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the system embodiments.
[0226] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. An operation and maintenance system, characterized in that, The operation and maintenance system includes: a monitoring terminal, a network proxy terminal, and a cloud platform; wherein: The monitoring terminal is used to acquire the operation and maintenance data of the robot control terminal to be monitored, and when it is determined that the robot control terminal is abnormal based on the operation and maintenance data, it sends the operation and maintenance data to the cloud platform through the network proxy terminal; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server. The cloud platform is used to generate a fault handling instruction for the robot control terminal when a fault is determined based on the received operation and maintenance data, and to send the fault handling instruction to the robot control terminal through the network proxy terminal so that the robot control terminal can perform fault handling according to the fault handling instruction.
2. The system according to claim 1, wherein the cloud platform is specifically used to generate a fault handling instruction for handling the fault of the first specified type when it is determined that the fault of the robot control terminal belongs to a first specified type; The cloud platform is also used to, when determining that the fault of the robot control terminal does not belong to the first specified type, obtain reference data associated with the operation and maintenance data from the log data and / or dump data of the robot control terminal, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data; and after receiving the upgrade file, send the upgrade file to the monitoring terminal through the network proxy terminal; The monitoring terminal is also used to receive the upgrade file and send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file.
3. The system according to claim 2, characterized in that, The robot control terminal is upgraded via rolling upgrades.
4. The system according to claim 1, characterized in that, The cloud platform is also used to receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal; If a robot malfunction is determined based on the operational status data, a fault handling instruction is generated to handle the malfunction of the robot.
5. The system according to claim 1, characterized in that, The monitoring terminal is deployed on the server; The monitoring terminal is also used to automatically shut down when it detects that the computing resources it occupies have reached a preset threshold.
6. The system according to claim 5, characterized in that, The system further includes: a monitoring and management terminal, used to perform at least one of the following steps: When the monitoring terminal is detected to be closed, the monitoring terminal will be restarted. The network proxy receives the monitoring terminal program package sent by the cloud platform to update the monitoring terminal deployed in the server. The network proxy receives the upgrade file for the robot control terminal sent by the cloud platform and sends it to the monitoring terminal, so that the monitoring terminal upgrades the robot control terminal based on the upgrade file.
7. The system according to claim 1, characterized in that, The monitoring terminal is also used to determine that the robot control terminal is abnormal before sending the operation and maintenance data to the cloud platform through the network proxy terminal when the robot control terminal is determined to be abnormal based on the operation and maintenance data. If the multiple operation and maintenance data corresponding to the same operation and maintenance state at multiple different times in the operation and maintenance data meet the preset conditions, the monitoring terminal is also used to determine that the robot control terminal is abnormal. The preset conditions include at least one of the following: The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the multiple operational status data indicate that the number of consecutive anomalies is greater than a second threshold, and the multiple operational status data indicate that the duration of consecutive anomalies is greater than a first duration.
8. The system according to any one of claims 1-6, characterized in that, The network proxy is deployed in the robot's charging station.
9. An operation and maintenance method, characterized in that, The monitoring terminal is applied to an operations and maintenance system, which further includes a network proxy terminal and a cloud platform; the method includes: Obtain the operation and maintenance data of the robot control terminal to be monitored; wherein the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server. When the operation and maintenance data determines that the robot control terminal is abnormal, the operation and maintenance data is sent to the cloud platform through the network proxy terminal. This enables the cloud platform to generate a fault handling instruction for the robot control terminal when it determines that the robot control terminal is faulty based on the received operation and maintenance data. The cloud platform then sends the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
10. The method according to claim 9, characterized in that, The method further includes: The system receives an upgrade file sent by the cloud platform for handling faults; wherein the upgrade file is generated by the maintenance personnel based on reference data; the reference data is data related to the maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type; The upgrade file is sent to the robot control terminal so that the robot control terminal can perform an upgrade based on the upgrade file.
11. The method according to claim 9, characterized in that, The monitoring terminal is deployed on the server; the method further includes: It automatically shuts down when it detects that the computing resources it is using have reached a preset threshold.
12. The method according to claim 9, characterized in that, Before sending the maintenance data to the cloud platform via the network proxy when the maintenance data determines that the robot control terminal is abnormal, the method further includes: For the same operating state in the operation and maintenance data, if the multiple operating state data meet the preset conditions, it is determined that the robot control terminal is abnormal. The preset conditions include at least one of the following: The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
13. An operation and maintenance method, characterized in that, A cloud platform applied to an operations and maintenance system, wherein the operations and maintenance system further includes: a network proxy terminal and a monitoring terminal; the method includes: The monitoring terminal receives maintenance data sent when it determines that the robot control terminal is abnormal; wherein the maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal. If a fault is determined in the robot control terminal based on the maintenance data, a fault handling instruction is generated to handle the fault in the robot control terminal. The fault handling instruction is sent from the network proxy to the robot control terminal, so that the robot control terminal can handle the fault according to the fault handling instruction.
14. The method according to claim 13, characterized in that, The generation of fault handling instructions for the robot control terminal includes: Generate fault handling instructions for handling faults of the first specified type; The method further includes: If it is determined that the fault of the robot control terminal does not belong to the first specified type, reference data associated with the operation and maintenance data in the log data and / or dump data of the robot control terminal is obtained and sent to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data; Upon receiving the upgrade file, the upgrade file is sent to the monitoring terminal via the network proxy terminal; so that the monitoring terminal receives the upgrade file and sends the upgrade file to the robot control terminal, so that the robot control terminal can perform an upgrade based on the upgrade file.
15. The method according to claim 13, characterized in that, The method further includes: Receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal; If a robot malfunction is determined based on the operational status data, a fault handling instruction is generated to handle the malfunction of the robot.
16. An operation and maintenance device, characterized in that, A monitoring terminal for an operations and maintenance system, the operations and maintenance system further including: a network proxy terminal and a cloud platform; the device includes: The operation and maintenance data acquisition module is used to acquire operation and maintenance data of the robot control terminal to be monitored; wherein, the operation and maintenance data includes at least one of the following: system-level operation status data of the server used to control the robot, component operation status data of the robot business components running in the server, and middleware operation status data of the middleware running in the server. The operation and maintenance data sending module is used to send the operation and maintenance data to the cloud platform through the network proxy when it is determined that the robot control terminal is abnormal based on the operation and maintenance data. This enables the cloud platform to generate a fault handling instruction for the robot control terminal when it is determined that the robot control terminal is faulty based on the received operation and maintenance data. The cloud platform then sends the fault handling instruction to the robot control terminal through the network proxy, so that the robot control terminal can perform fault handling according to the fault handling instruction.
17. The apparatus according to claim 16, characterized in that, The device further includes: The upgrade file receiving module is used to receive an upgrade file sent by the cloud platform for handling faults; wherein the upgrade file is generated by the operation and maintenance personnel based on reference data; the reference data is data related to the operation and maintenance data in the log data and / or dump data of the robot control terminal obtained by the cloud platform when it determines that the fault of the robot control terminal does not belong to the first specified type; The upgrade file upgrade module is used to send the upgrade file to the robot control terminal so that the robot control terminal can upgrade based on the upgrade file; The monitoring terminal is deployed on the server; the device also includes: The automatic shutdown module is used to automatically shut down when it detects that the computing resources it is using have reached a preset threshold. The device further includes: An anomaly determination module is used to determine that the robot control terminal is abnormal before sending the operation and maintenance data to the cloud platform through the network proxy terminal when the operation and maintenance data determines that the robot control terminal is abnormal through the operation and maintenance data. If the multiple operation and maintenance data corresponding to the same operation state in the operation and maintenance data meet the preset conditions, the module determines that the robot control terminal is abnormal. The preset conditions include at least one of the following: The multiple operational status data indicate that the frequency of anomalies is greater than a first threshold, the number of consecutive anomalies in the multiple operational status data is greater than a second threshold, and the duration of consecutive anomalies in the multiple operational status data is greater than a first duration.
18. A maintenance device, characterized in that, A cloud platform applied to an operations and maintenance system, the operations and maintenance system further including: a network proxy terminal and a monitoring terminal; the device includes: The operation and maintenance data receiving module is used to receive operation and maintenance data sent by the monitoring terminal when it determines that the robot control terminal is abnormal; wherein, the operation and maintenance data includes at least one of the following: system-level operating status data of the server used to control the robot, component operating status data of the robot business components running in the server, and middleware operating status data of the middleware running in the server; and the operation and maintenance data is used by the monitoring terminal to determine whether the robot control terminal is abnormal. The first instruction generation module is used to generate a fault handling instruction for handling the fault of the robot control terminal when the robot control terminal is determined to be faulty based on the operation and maintenance data. The instruction sending module is used to send the fault handling instruction to the robot control terminal through the network proxy terminal, so that the robot control terminal can perform fault handling according to the fault handling instruction.
19. The apparatus according to claim 18, characterized in that, The first instruction generation module is specifically used to generate fault handling instructions for handling faults of a first specified type. The device further includes: The reference data acquisition module is used to acquire reference data associated with the operation and maintenance data from the log data and / or dump data of the robot control terminal when it is determined that the fault of the robot control terminal does not belong to the first specified type, and send it to the operation and maintenance client so that the operation and maintenance personnel can generate an upgrade file for handling the fault based on the reference data. The upgrade file sending module is used to send the upgrade file to the monitoring terminal through the network proxy terminal after receiving the upgrade file; so that the monitoring terminal receives the upgrade file and sends the upgrade file to the robot control terminal, so that the robot control terminal can perform an upgrade based on the upgrade file; The device further includes: The operation status data receiving module is used to receive robot operation status data uploaded by the robot to be monitored through the network proxy terminal; The second instruction generation module is used to generate a fault handling instruction for the robot to handle the fault when the robot fault is determined based on the operating status data.
20. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 9-12 or 13-15.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 9-12 or 13-15.
Citation Information
Patent Citations
Cloud power distribution room robot intelligent operation and maintenance duty management system
CN112055089A
Operation and maintenance information processing method and related device
CN116009506A